A method and system for automatically generating game advertisement videos based on artificial intelligence

By training agents and extracting features based on a hierarchical reward shaping mechanism, and combining this with a large language model to generate editing scripts, the automation and multimodal issues of game-based advertising video production are solved, achieving efficient and high-quality advertising video generation to meet the needs of rapid promotion.

CN122138017APending Publication Date: 2026-06-02安徽三七极光网络科技有限公司

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
安徽三七极光网络科技有限公司
Filing Date
2026-03-02
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Current game-style advertising video production relies on manual screen recording and editing, which is time-consuming and labor-intensive. It lacks intelligent recognition of exciting segments, and the editing lacks flexibility and innovation. It cannot make full use of multimodal information, resulting in low production efficiency, poor quality, and an inability to effectively attract players.

Method used

An agent based on a hierarchical reward shaping mechanism is trained to perform automated operation recording. Features are extracted using the CLIP multimodal model, and multimodal information is fused through a graph neural network to generate high-quality materials. An editing script is generated by combining a multimodal large language model, and the editing strategy is optimized through reinforcement learning to adaptively output multi-format advertising videos.

Benefits of technology

It achieves fully automated generation of game advertising videos, improving production efficiency and quality, generating more attractive and targeted advertising videos, meeting the needs of rapid promotion, reducing labor costs, accurately identifying and highlighting game highlights, and making full use of multimodal information to convey an immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122138017A_ABST
    Figure CN122138017A_ABST
Patent Text Reader

Abstract

This invention discloses an AI-based method and system for automatically generating game advertising videos. The method specifically includes: training an agent based on a hierarchical reward shaping mechanism; using the agent to perform automated operations and synchronous screen recording on a target game; extracting cross-modal features from the screen-recorded video and aligning and fusing them to obtain multimodal fusion features; identifying visual defects in the original material based on the multimodal fusion features, performing video rendering and special effects enhancement on the defective areas to form a high-quality material library; generating an editing script based on the high-quality material library and prior game information; and converting the editing script into video editing instructions and calling the editing engine to perform automated editing. This invention achieves fully automated generation of game advertising videos from automated operation recording to final adaptive output, improving production efficiency and quality, reducing labor costs, and generating more attractive and targeted advertising videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and system for automatically generating game advertising videos based on artificial intelligence. Background Technology

[0002] In today's booming gaming industry, gameplay-oriented advertising videos serve as a crucial means of attracting players and promoting games, and their production quality directly impacts the game's promotional effectiveness and market competitiveness. However, the current production of gameplay-oriented advertising videos faces numerous unresolved issues, severely hindering both production efficiency and quality.

[0003] Firstly, currently, the production of gameplay-based advertising videos primarily relies on manual screen recording and editing. Starting with recording gameplay, professionals are required to manually control the game and record the process. This process is not only time-consuming and labor-intensive but also demands a high level of proficiency from the operators. Subsequently, the editing stage also requires professional editors to meticulously edit the recorded footage, selecting suitable segments from numerous clips, adjusting the order of shots, adding special effects, etc. The entire process is tedious and complex, resulting in high production costs and long production cycles, making it difficult to meet the needs of rapid game promotion.

[0004] Secondly, games contain a wealth of information, with truly exciting and engaging gameplay moments often hidden within. Existing automated solutions have significant shortcomings in identifying and extracting these exciting moments. On one hand, they lack an effective intelligent recognition mechanism, making it difficult to automatically determine which moments are more attractive based on the game's characteristics and players' interests. On the other hand, existing solutions cannot accurately capture complex operations and hidden storylines within the game, resulting in generated advertising videos that fail to highlight the game's strengths and features, and thus fail to effectively attract the attention of potential players.

[0005] Furthermore, traditional automated editing methods are typically based on preset rules and templates. While they can achieve a certain degree of automation, they lack flexibility and innovation. These methods struggle to generate creative and rhythmic video effects tailored to the style of different games, promotional objectives, and the characteristics of the target audience. In conveying the game's selling points, traditional methods often fall short, failing to fully showcase the game's core advantages and unique charm to the audience through skillful editing, thus significantly diminishing the appeal and impact of the advertising video.

[0006] Finally, a game is a complex system containing multiple modalities of information, including images, sound, and text. These different modalities are interconnected and complementary, collectively creating the rich experiential content of a game. However, existing automated editing solutions have significant shortcomings in processing multimodal information, failing to fully utilize this information for intelligent editing. For example, during editing, they may focus only on the visual content, neglecting the emotions and semantics conveyed by audio and text information. This results in advertising videos that are not comprehensive or accurate in their information delivery, failing to provide viewers with an immersive experience. Summary of the Invention

[0007] The purpose of this invention is to provide an AI-based method and system for automatically generating game advertising videos, which realizes the fully automated generation of game advertising videos from automated operation recording to final adaptive output, improves production efficiency and quality, reduces labor costs, and generates more attractive and targeted advertising videos, thereby solving at least one of the aforementioned problems in the prior art.

[0008] In a first aspect, the present invention provides a method for automatically generating game advertising videos based on artificial intelligence, the method specifically comprising: The intelligent agent is trained based on a hierarchical reward shaping mechanism. The intelligent agent performs automated operations and synchronous screen recording on the target game. The hierarchical reward shaping mechanism includes basic game rewards, exploration rewards, skill rewards and story rewards. In response to the completion of recording, the CLIP multimodal model is used to extract cross-modal features from the screen recording video, and then the graph neural network is used for alignment and fusion to obtain multimodal fusion features; Visual defects in the original material are identified based on multimodal fusion features, and a pre-trained generative model is called to perform video rendering and special effects enhancement on the defective areas, forming a high-quality material library; Based on a high-quality resource library and prior game information, a multimodal large language model is used in conjunction with a dynamic Prompt project to generate editing scripts. The editing script is converted into video editing instructions and the editing engine is called to perform automated editing. At the same time, the editing strategy is fine-tuned through reinforcement learning based on historical user feedback. The edited video sequence is rendered and output, and various advertising video formats are adaptively generated according to the requirements of the target platform.

[0009] Secondly, this invention provides an artificial intelligence-based automatic generation system for game advertising videos, the system specifically comprising: The screen recording module is used to train an intelligent agent based on a hierarchical reward shaping mechanism. The intelligent agent performs automated operations and synchronous screen recording on the target game. The hierarchical reward shaping mechanism includes basic game rewards, exploration rewards, skill rewards, and story rewards. The feature fusion module is used to extract cross-modal features from the screen recording video in response to the completion of recording, and then perform alignment and fusion through a graph neural network to obtain multimodal fused features. The defect correction module is used to identify visual defects in the original material based on multimodal fusion features, and call a pre-trained generative model to perform video rendering and special effects enhancement on the defective areas to form a high-quality material library. The script generation module is used to generate editing scripts based on a high-quality material library and game prior information, by combining a multimodal large language model with a dynamic Prompt project. The video editing module is used to convert editing scripts into video editing instructions and call the editing engine to perform automated editing, while fine-tuning the editing strategy based on historical user feedback through reinforcement learning; The video output module is used to render and output the edited video sequence, and adaptively generate advertising videos in various formats according to the requirements of the target platform.

[0010] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, and a computer program stored in the memory, wherein when the computer program is executed on the processor, it implements the AI-based automatic generation method for game advertising videos as described in any of the above methods.

[0011] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the AI-based automatic generation method for game advertising videos as described in any of the above methods.

[0012] Compared with the prior art, the present invention has at least one of the following technical effects: 1. This invention realizes the fully automated generation of game advertising videos from automated operation recording to final adaptive output, which improves production efficiency and quality, reduces labor costs, and generates more attractive and targeted advertising videos.

[0013] 2. This invention automates game operation and recording by combining a general-purpose computer-controlled intelligent agent, eliminating the tedious process of manual screen recording and significantly saving time and labor costs. Simultaneously, it utilizes a multimodal large language model (MLLM) for intelligent editing, reducing the workload of manual editing, further shortening the production cycle, and improving production efficiency, thus meeting the needs of rapid game promotion.

[0014] 3. This invention trains an intelligent agent based on a hierarchical reward-based training mechanism, which includes various reward types such as basic game rewards, exploration rewards, skill rewards, and story rewards. Guided by these rewards, the intelligent agent can automatically explore various areas of the game, execute high-difficulty operation sequences, and activate hidden story events, thereby comprehensively capturing exciting gameplay segments. Utilizing a multimodal model for feature extraction and analysis of screen recording videos enables more accurate identification of attractive segments, ensuring that the generated advertising videos highlight the game's strengths and features.

[0015] 4. By analyzing information such as the core selling points, target audience, emotional tone, and key gameplay elements of a game, this invention uses a multimodal large language model to generate editing scripts that suit different game types and promotional objectives, making advertising videos more attractive and impactful, and effectively conveying the game's selling points.

[0016] 5. This invention utilizes the CLIP multimodal model to extract cross-modal features from screen-recorded videos and performs alignment and fusion through a graph neural network to obtain multimodal fusion features. In subsequent material processing and editing, the correlation and complementarity between various modal information such as images, sound, and text are fully considered, enabling a more comprehensive and accurate transmission of game information, providing viewers with an immersive experience, and improving the quality and effectiveness of advertising videos.

[0017] 6. This invention constructs an interaction graph and updates the graph convolutional network to deeply mine the semantic associations between entities of different modalities and generate more representative multimodal fusion features.

[0018] 7. This invention combines prior game information and dynamic Prompt engineering, and uses a multimodal large language model to generate creative editing scripts that conform to the game's characteristics and promotional goals.

[0019] 8. This invention utilizes reinforcement learning to optimize editing strategies based on user feedback data, making the generated advertising videos more aligned with user preferences and improving promotional effectiveness. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating an artificial intelligence-based method for automatically generating game advertisement videos according to an embodiment of the present invention. Figure 2This is a schematic diagram of the structure of an AI-based automatic game advertisement video generation system according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0022] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0023] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0024] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0025] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0026] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0027] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0028] In this application embodiment, the entity executing the process includes a terminal device. This terminal device includes, but is not limited to, devices capable of executing the methods disclosed in this application, such as servers, computers, smartphones, and tablets. Figure 1 A flowchart illustrating an embodiment of the artificial intelligence-based automatic generation method for game advertisement videos disclosed in this invention is shown below in detail: S101, Train the intelligent agent based on the hierarchical reward shaping mechanism, and perform automated operations and synchronous screen recording on the target game through the intelligent agent. The hierarchical reward shaping mechanism includes basic game rewards, exploration rewards, skill rewards and plot rewards. S102, in response to recording completion, uses the CLIP multimodal model to extract cross-modal features from the screen recording video, and performs alignment and fusion through a graph neural network to obtain multimodal fusion features; S103 identifies visual defects in the original material based on multimodal fusion features, and calls a pre-trained generative model to perform video rendering and special effects enhancement on the defective areas, forming a high-quality material library; S104, based on a high-quality material library and prior game information, generates editing scripts by combining a multimodal large language model with a dynamic Prompt project; S105 converts the editing script into video editing instructions and calls the editing engine to perform automated editing, while fine-tuning the editing strategy based on historical user feedback through reinforcement learning. S106 renders and outputs the edited video sequence, and adaptively generates advertising videos in various formats according to the requirements of the target platform.

[0029] In this embodiment, a tiered reward system is constructed, comprising basic game rewards, exploration rewards, skill rewards, and story rewards. Basic game rewards incentivize the agent to complete basic game tasks, such as clearing specific levels. Exploration rewards encourage the agent to explore unknown areas and hidden content in the game to discover more unique game scenes and elements. Skill rewards reward the agent for completing difficult maneuvers or complex combos, prompting the agent to improve its skill level. Story rewards reward the agent when it triggers key story nodes in the game, ensuring the complete presentation of the story. Based on this tiered reward system, the agent is trained. During training, the agent continuously tries and optimizes its operational strategies to maximize reward values. Once the agent is trained, it is applied to a target game. The agent can automatically control the game character to perform various operations while simultaneously recording the entire game process. This method avoids the tediousness of manually controlling and recording the game, and the trained agent can complete game operations in a relatively stable and efficient manner, providing a rich source of material for subsequent advertising video production.

[0030] After the screen recording video is completed, the CLIP multimodal model is used to extract cross-modal features from the video. The CLIP model has powerful multimodal understanding capabilities, able to process information from multiple modalities such as images and text simultaneously. In this step, the CLIP model extracts features from the screen recording video's image and audio (converted to text information through speech recognition), including visual information from the game screen and semantic information from the audio. Then, the extracted features from different modalities are input into a graph neural network. The graph neural network can model and analyze the relationships between features from different modalities, aligning and fusing features from different modalities through node and edge connections to obtain multimodal fused features. These fused features integrate information from various aspects of the game, more comprehensively reflecting the game's features and highlights.

[0031] Based on the obtained multimodal fusion features, visual defects are identified in the original screen recording footage. Analysis of the fusion features determines whether visual defects such as blurry images, color distortion, or incomplete scenes exist in the footage. Once the defective area is identified, a pre-trained generative model, such as a Generative Adversarial Network (GAN), is invoked. The generative model can generate content that matches the surrounding environment based on normal information around the defective area, performing video rendering and special effects enhancement on the defective area. For example, for blurry areas, the generative model can generate clear content for replacement; for areas with color distortion, it can adjust the colors to restore them to normal. After processing, a high-quality material library is formed, providing excellent material for subsequent advertising video editing.

[0032] Leveraging a high-quality resource library and prior game information, a multimodal large language model combined with a dynamic Prompt project generates editing scripts. Prior game information includes the game's genre, style, target audience, and key promotional points. The multimodal large language model understands the multimodal information in the game assets and the prior game information, generating a suitable editing script based on this information. The dynamic Prompt project dynamically adjusts the prompts input to the multimodal large language model according to different game characteristics and promotional needs, making the generated editing script more targeted and creative. For example, for an action game, the editing script might highlight exciting combat scenes and cool skill effects; for a puzzle game, it would emphasize the puzzle design and exploration process.

[0033] The generated editing scripts are converted into video editing instructions, which include information such as the editing order, duration, and effects to be added. Then, the video editing engine is invoked to automatically edit footage from a high-quality media library according to these instructions. During the editing process, the editing strategy is fine-tuned using reinforcement learning based on historical user feedback. Historical user feedback includes data such as user viewing time, likes, and comments on previous ad videos. By analyzing this data, we can understand users' preferences for different types of edited content. The reinforcement learning algorithm continuously adjusts the editing strategy based on historical user feedback, for example, adding editing techniques and content that users like and reducing parts that users are not interested in, thus making the generated ad video more in line with user tastes and increasing its attractiveness and appeal.

[0034] The edited video sequence is rendered and output. The rendering process further optimizes the video's image quality, color, and special effects, resulting in a better visual experience for the advertising video. Then, based on the requirements of the target platform, advertising videos in multiple formats are adaptively generated. Different game promotion platforms may have different requirements for video formats; for example, some platforms support MP4, while others may require AVI. By adaptively generating advertising videos in multiple formats, the playback needs of different platforms can be met, expanding the reach of the advertising video and improving the promotional effect of the game.

[0035] In some embodiments, step S101 above, which involves training the agent based on a hierarchical reward shaping mechanism, specifically includes: Construct the environment interaction interface of the target game and initialize the basic policy network, and establish a communication connection between the agent and the game environment through the basic policy network; Construct a multi-objective reward function that includes basic game rewards, exploration rewards, skill rewards, and story rewards; Based on the multi-objective reward function, the agent interacts with the game environment to generate experience data, and stores the experience data in the experience pool. A small batch of experience data is randomly sampled from the experience pool. The policy gradient is calculated using the sampled small batch of experience data, and the basic policy network parameters are updated using the backpropagation algorithm to guide the agent to optimize the policy in the direction of maximizing the cumulative reward.

[0036] In this embodiment, a dedicated environment interaction interface is constructed for the target game. This interface defines the set of operations the agent can perform and the types of information the game environment can provide back to the agent. For example, in action games, the set of operations might include actions such as moving, jumping, and attacking; the feedback information might cover the game character's state, current level information, and the surrounding environment. Through this interface, the agent can effectively communicate with the game environment. Simultaneously, a basic policy network is initialized. The basic policy network determines how the agent makes decisions when facing the game environment. The initialized basic policy network can be a simple neural network structure with a certain degree of randomness to allow for continuous learning and optimization during subsequent training. Through the basic policy network, the agent can select appropriate operations based on the information fed back from the game environment, thereby establishing an initial communication connection with the game environment.

[0037] To guide the agent in comprehensive and effective exploration and operation within the game, a multi-objective reward function is constructed, encompassing basic game rewards, exploration rewards, skill rewards, and story rewards. Basic game rewards are earned when the agent completes basic game tasks, such as clearing a level or defeating an enemy, incentivizing the agent to complete the game's basic flow. Exploration rewards encourage the agent to explore unknown areas and hidden content within the game; rewards are given when the agent discovers new map areas, unlocks hidden items, or triggers special events, helping it discover more exciting gameplay elements. Skill rewards reward the agent for performing high-difficulty maneuvers or complex combos, such as executing a flashy combo in a fighting game or achieving a high-speed drift through a corner in a racing game, prompting the agent to improve its operational skills and showcase the game's exciting aspects. Story rewards reward the agent when it triggers key plot points, ensuring the complete presentation of the story and allowing the agent to capture the game's narrative highlights. Through the organic combination of these four types of rewards, the multi-objective reward function can comprehensively evaluate the agent's performance in the game, providing clear guidance for the agent's training.

[0038] After constructing a multi-objective reward function, the agent begins interacting with the game environment. The agent makes operational decisions based on the basic policy network, and the game environment provides corresponding feedback based on the agent's actions, calculating the reward value obtained by the agent according to the multi-objective reward function. Each interaction generates a set of experience data, which includes information such as the agent's current state, the actions performed, the game environment's feedback, and the reward value obtained. For example, if the agent is in a certain position in the game and chooses to attack, the enemy's state changes after being attacked, and the agent receives a corresponding reward value based on the attack effect, this series of information constitutes a set of experience data. The agent continuously interacts with the game environment, generating a large amount of experience data, which is stored in an experience pool. The experience pool acts like a data warehouse, storing the agent's interaction experience in different states, providing rich data support for subsequent policy network training.

[0039] Once a certain amount of experience data has been stored in the experience pool, training of the policy network begins. Small batches of experience data are randomly sampled from the experience pool; these batches are representative and diverse, randomly selected from the entire pool. The policy gradient is calculated using these sampled small batches. The policy gradient reflects the direction of adjustment of the basic policy network parameters, guiding the agent on how to change its operational strategy to obtain greater cumulative rewards. For example, if experience data shows that the agent obtained a high reward by performing a certain action in a certain state, the policy gradient will tend to encourage the agent to perform that action more often in similar states. After calculating the policy gradient, the parameters of the basic policy network are updated using the backpropagation algorithm. The backpropagation algorithm adjusts the parameters of the neural network layer by layer from the output layer to the input layer based on the policy gradient, allowing the basic policy network to better adapt to the game environment and guiding the agent to optimize its strategy towards maximizing cumulative rewards. Through multiple such sampling and parameter update processes, the basic policy network is continuously optimized, and the agent's operational strategy becomes increasingly sophisticated, enabling it to complete various tasks more efficiently in the game, capture more exciting gameplay clips, and provide high-quality material for the automatic generation of subsequent game advertising videos.

[0040] Furthermore, the basic game reward is used to measure the progress of completing the basic game objectives, the exploration reward is calculated based on the state novelty assessment and used to incentivize the agent to explore unvisited areas, the skill reward is calculated based on the operation complexity assessment and used to incentivize the agent to execute high-difficulty operation sequences, and the plot reward is calculated based on the game plot trigger point detection and used to incentivize the agent to activate hidden plot events.

[0041] In this embodiment, the exploration reward is calculated based on the state novelty assessment, specifically including: constructing a target network with fixed random initialization and a trainable predictor network with the same structure, and using the mean square error between the output of the trainable predictor network to the current state and the output of the target network as the exploration reward value; when the agent accesses a novel state, if the prediction error is large, a high exploration reward is given, and when the state is fully explored, if the prediction error approaches zero, the exploration reward is decayed.

[0042] The skill reward is calculated based on the operational complexity assessment, specifically including: scoring the matching degree of the agent's continuous action sequence through a predefined high-difficulty operation template, or dynamically mining high-value operation patterns from historical high-reward trajectories and calculating similarity through a clustering learning algorithm.

[0043] The plot rewards are calculated based on the detection of game plot trigger points. Specifically, this includes monitoring multimodal information including in-game text prompts, character dialogue content, scene switching events, and background music changes. When a match is detected with preset plot keywords or event patterns, it is determined as a plot trigger and a corresponding reward value is assigned.

[0044] In this embodiment, basic game rewards are primarily used to measure progress towards basic game objectives. At the outset of game design, a series of basic objectives are defined, such as clearing specific levels, defeating a certain number of enemies, and collecting specific items. When the agent completes these basic objectives in the game, it receives corresponding basic game rewards. The reward value can be set according to the difficulty and importance of the objective; the higher the difficulty and the greater the importance of the objective, the greater the basic game reward upon completion. For example, in an adventure game, the basic game reward for defeating the final boss is far greater than the reward for defeating ordinary monsters. This method incentivizes the agent to operate towards completing the basic game objectives, ensuring the agent can smoothly advance the game and laying the foundation for subsequent capture of exciting gameplay moments and triggering plot events.

[0045] The exploration reward is calculated based on the novelty of the state, aiming to incentivize the agent to explore unvisited regions. In practice, a fixed, randomly initialized target network and a trainable predictor network with identical structure are first constructed. The target network provides a relatively stable reference standard, while the predictor network continuously learns and adjusts to adapt to changes in the game environment. During the agent's interaction with the game environment, the trainable predictor network predicts the output of the current state, then compares this prediction with the output of the target network, calculating the mean squared error between the two. This mean squared error is used as the exploration reward value.

[0046] When an agent encounters a novel state, its prediction differs significantly from the target network's output because the predictor network has never encountered a similar state before; this results in a large prediction error. In this case, the agent is given a high exploration reward. This high reward encourages the agent to continue exploring the new area to acquire more information. As the agent continues to explore and become familiar with the state, the predictor network gradually learns its characteristics, and the prediction error gradually decreases. When the state has been fully explored, the prediction error approaches zero, and the exploration reward diminishes accordingly. In this way, the agent is motivated to explore every corner of the game, discover hidden areas and gameplay elements, and support the generation of rich and diverse advertising video materials.

[0047] Skill rewards are calculated based on operational complexity assessment, aiming to incentivize agents to execute high-difficulty operation sequences. There are two implementation methods: One method involves scoring the agent's continuous action sequences based on the matching degree of predefined high-difficulty operation templates. During the game design phase, a series of high-difficulty operation templates are predefined according to the game's characteristics and gameplay. For example, in a fighting game, a series of flashy combos, or in a racing game, high-speed drifting through corners, can serve as high-difficulty operation templates. When the agent performs actions in the game, its continuous action sequences are matched against these predefined high-difficulty operation templates. The higher the matching degree, the closer the agent's actions are to high-difficulty operations, and the greater the skill reward. This method is simple and direct, clearly incentivizing the agent to complete specific high-difficulty operations. The other method uses clustering learning algorithms to dynamically mine high-value operation patterns from historical high-reward trajectories and calculate similarity. During the agent's interaction with the game environment, a large amount of historical trajectory data is recorded, including the agent's actions and corresponding rewards. Clustering learning algorithms are used to analyze and mine these historical high-reward trajectories to identify high-value operation patterns. When the agent performs actions in subsequent gameplay, its action sequences are compared with the discovered high-value action patterns for similarity calculation. The higher the similarity, the greater the skill reward. This method dynamically adjusts the reward mechanism based on the agent's actual performance in the game, more flexibly incentivizing the agent to explore and execute various high-difficulty actions, thus enhancing the game's excitement.

[0048] Story rewards are calculated based on game story trigger point detection and are used to incentivize agents to activate hidden story events. Game story is a crucial factor in attracting players, and hidden story events further enhance the game's fun and mystery. To detect game story trigger points, multimodal information including in-game text prompts, character dialogue, scene transitions, and background music changes needs to be monitored.

[0049] During gameplay, the agent continuously collects this multimodal information and matches it with preset plot keywords or event patterns. These preset plot keywords and event patterns are determined during the game design phase based on the plot content. For example, a character saying a specific line, entering a specific scene, or a change in background music might trigger a plot event. When a match with a preset plot keyword or event pattern is detected, it's considered a plot trigger, and the agent is awarded a corresponding reward. The reward value can be set based on the importance and excitement of the plot; important and exciting plot triggers receive higher reward values. This incentivizes the agent to actively seek out and activate hidden plot events, enriching the content of game promotional videos and attracting the attention of potential players.

[0050] In some embodiments, step S102 above, which involves extracting cross-modal features from the screen recording video using the CLIP multimodal model and aligning and fusing them through a graph neural network to obtain multimodal fused features, specifically includes: Multimodal data parsing and preprocessing are performed on screen recording videos to generate keyframe sequences, audio clips, operation event sequences, and game state variables; The keyframe sequence, audio clips, operation event sequence, and game state variables are input into the pre-trained ViFi-CLIP model for feature extraction, resulting in video frame feature vectors, audio feature vectors, and text semantic feature vectors. Semantic entity extraction is performed on video frame feature vectors, audio feature vectors, and text semantic feature vectors respectively to obtain fine-grained enhanced features containing visual entities, acoustic events, and semantic entities; Based on the fine-grained enhancement features, a cross-modal interaction graph with visual entities, acoustic events, and semantic entities as nodes is constructed, and the representation of each node is iteratively updated through a graph convolutional network to obtain multimodal fusion features.

[0051] In this embodiment, multimodal data parsing and preprocessing are performed on the screen recording video. For image information in the screen recording video, frames are extracted at certain time intervals or key action trigger points to generate a keyframe sequence. These keyframes represent important visual content in the video, providing a foundation for subsequent feature extraction. For audio information, the video is segmented according to its timeline to obtain audio segments. These audio segments can include various sound effects, background music, and character voices from the game, which play a crucial role in conveying the game's atmosphere and emotions. Simultaneously, the game operation process is recorded in detail to generate an operation event sequence, such as player key presses and mouse movements. These operation events reflect the player's behavioral patterns in the game. Furthermore, game state variables, such as the character's health, level, and current scene, need to be extracted. These variables reflect the current running state of the game. Through this multimodal data parsing and preprocessing of the screen recording video, various types of information in the video can be comprehensively obtained, preparing for subsequent feature extraction.

[0052] The pre-processed keyframe sequences, audio clips, action event sequences, and game state variables are input into a pre-trained ViFi-CLIP model for feature extraction. The ViFi-CLIP model is specifically designed for processing multimodal data and possesses powerful feature extraction capabilities. For keyframe sequences, the model analyzes various visual elements in the image, such as color, shape, and texture, and transforms them into video frame feature vectors. These feature vectors accurately describe the visual features of the image, providing visual information for subsequent cross-modal fusion. For audio clips, the model identifies acoustic features in the audio, such as frequency, pitch, and volume, and transforms them into audio feature vectors. Audio feature vectors reflect the acoustic characteristics of the audio, helping to understand the information conveyed by the audio. For action event sequences and game state variables, the model transforms them into textual semantic feature vectors. Action event sequences and game state variables can be transformed using natural language descriptions; the model can understand the semantics in this textual information and transform it into semantically representative feature vectors. Through feature extraction using the ViFi-CLIP model, we obtained video frame feature vectors, audio feature vectors, and text semantic feature vectors. These feature vectors represent the screen recording video from different modalities.

[0053] After obtaining video frame feature vectors, audio feature vectors, and text semantic feature vectors, semantic entity extraction is performed on each. For video frame feature vectors, image recognition techniques are used to analyze visual elements such as objects and scenes in the image, extracting fine-grained enhanced features containing visual entities. For example, in a role-playing game, visual entities such as characters, weapons, and monsters can be extracted. For audio feature vectors, audio analysis techniques are used to identify various sound events in the audio, such as explosions, footsteps, and dialogue, extracting fine-grained enhanced features containing acoustic events. These acoustic events can reflect specific scenes and actions in the game. For text semantic feature vectors, natural language processing techniques are used to analyze keywords and phrases in the text, extracting fine-grained enhanced features containing semantic entities. For example, semantic entities such as character level and health points can be extracted from the descriptive text of game state variables. Through semantic entity extraction, we obtain more fine-grained feature representations, which can more accurately describe the key content in different modalities.

[0054] Based on the extracted fine-grained enhancement features, a cross-modal interaction graph is constructed, with visual entities, acoustic events, and semantic entities as nodes. In this graph, entities from different modalities are connected through certain relationships. For example, a visual entity may be associated with an acoustic event because they appear simultaneously in the game; a semantic entity may have a semantic connection with either a visual entity or an acoustic event. Constructing such a cross-modal interaction graph visually demonstrates the relationships between different modalities. Then, a graph convolutional network iteratively updates the representations of each node in the cross-modal interaction graph. The graph convolutional network can utilize the information in the graph structure to propagate and aggregate node features. During the iterative update process, the features of each node are influenced by the features of its neighboring nodes, thus continuously fusing information from different modalities. After multiple iterations, the representations of each node can fully integrate multimodal feature information, ultimately obtaining multimodal fused features. These multimodal fused features can comprehensively and accurately describe various information in the screen recording video, providing strong support for the subsequent generation of game advertising videos.

[0055] In one possible implementation, the pre-training steps of the ViFi-CLIP model include: collecting publicly available game video libraries, official game promotional materials, and exciting game clips uploaded by players to form video data. For the collected video data, keyframes are extracted at regular time intervals to form a keyframe sequence. Simultaneously, audio segments are separated from the video, and preprocessing operations such as noise reduction are performed to improve audio quality. The text description information needs to describe the game scene, operation events, character states, etc., in detail and accurately, such as "the character releases a fire skill to attack monsters in a forest scene." To enhance the model's understanding and association of different modal information, the collected data is labeled. The labeled content includes visual entities in the video frames, such as characters, props, and scene elements; acoustic events in the audio, such as skill release sound effects and environmental sound effects; and semantic entities in the text, such as character names, skill names, and game state descriptions. Through labeling, the model can clearly understand the correspondence between different modal information, providing an accurate foundation for subsequent training.

[0056] The ViFi-CLIP model adopts a multimodal encoder architecture, which includes three main parts: a video encoder, an audio encoder, and a text encoder.

[0057] The video encoder employs a structure combining Convolutional Neural Networks (CNNs) and Transformers. The CNN part extracts local features from video frames by gradually reducing spatial resolution and increasing the number of channels through multiple convolutional and pooling operations, capturing visual information such as shape, color, and texture in the video frames. The Transformer part handles the temporal relationships between video frames, learning the temporal dependencies of video frames through a self-attention mechanism to obtain a global feature representation of the video.

[0058] The audio encoder uses a combination of a one-dimensional convolutional neural network (1D-CNN) and a long short-term memory network (LSTM). The 1D-CNN is used to extract local spectral features of audio segments, while the LSTM is used to process the temporal information of the audio, capturing features such as rhythm and pitch changes, and generating feature vectors for the audio.

[0059] The text encoder uses a pre-trained language model, such as BERT or GPT. These pre-trained models have been trained on a large amount of text data and can capture the semantic information of the text very well. The text description information is input into the pre-trained language model to obtain the semantic feature vector of the text.

[0060] The model architecture also incorporates a cross-modal interaction module to facilitate information exchange and fusion among features from different modalities. This module utilizes an attention mechanism to enable video, audio, and text features to mutually attend to and learn from each other, thereby enhancing the model's comprehensive understanding of multimodal information.

[0061] To enable the ViFi-CLIP model to learn the associations and semantic representations between different modal information, the following pre-training tasks are set: (1) Contrastive learning task: Pair up data from video, audio, and text modalities to form positive sample pairs, such as video-text positive sample pairs, audio-text positive sample pairs, and video-audio positive sample pairs. At the same time, randomly select data from different samples to form negative sample pairs. The goal of the model is to learn a feature space such that the distance between positive sample pairs in the feature space is as close as possible, and the distance between negative sample pairs in the feature space is as far as possible. Through contrastive learning, the model can learn the semantic consistency between different modal information and improve its ability to understand and represent multimodal data. (2) Masking language modeling task (for text modality): Randomly mask some words or phrases in the text description information and let the model predict the masked content based on the context information. This task can enhance the model's ability to understand and generate text semantics, and enable the model to better capture the semantic relationships and context information in the text. (3) Audio event prediction task (for audio modality): Randomly mask or add noise to audio segments, and let the model predict the masked or noise-affected audio events. Through this task, the model can learn key events and features in the audio, improving its ability to understand and process audio information. (4) Video frame prediction task (for video modality): Given the first few frames of a video, let the model predict the subsequent video frames. This task enables the model to learn the temporal dynamic information and motion patterns in the video, improving its ability to understand and generate video content.

[0062] For the CNN part of the video encoder, model parameters pre-trained on large-scale image datasets such as ImageNet are used for initialization. These parameters have learned rich image feature representations, which can accelerate the convergence of the model on video data. The parameters of the Transformer part are randomly initialized so that it can gradually learn the temporal relationships between video frames during training.

[0063] The 1D-CNN parameters in the audio encoder are also randomly initialized, and the parameters of the LSTM part can be initialized using parameters pre-trained on relevant audio datasets to improve the model's ability to process audio temporal information.

[0064] The text encoder directly uses the initial parameters of the pre-trained language model, which have been optimized on a large amount of text data and can well represent the semantic information of the text.

[0065] The parameters of the cross-modal interaction module are randomly initialized so that they can be dynamically adjusted during training based on the relationships between different modal features.

[0066] Simultaneously, setting appropriate hyperparameters such as learning rate, batch size, and number of training epochs is crucial. The learning rate determines the step size for updating model parameters; an excessively large learning rate may prevent the model from converging, while an excessively small learning rate will slow down the training process. The batch size affects the amount of data input to the model during each training iteration; an appropriate batch size can improve training efficiency and model stability. The number of training epochs determines the number of times the model learns from the entire training set and needs to be set reasonably based on the size of the dataset and the complexity of the model.

[0067] Preprocessed video, audio, and text data are grouped into training batches according to a set batch size and input into the ViFi-CLIP model for training. During the training of each batch, firstly, the video encoder processes the video frame sequence to generate video feature vectors; the audio encoder processes the audio segments to generate audio feature vectors; and the text encoder processes the text description information to generate text semantic feature vectors. Then, the cross-modal interaction module interacts and fuses the video, audio, and text feature vectors to obtain multimodal fused features. Based on the set pre-training tasks, the loss function values ​​of the model on different tasks are calculated. For example, for the contrastive learning task, the distance between positive and negative sample pairs in the feature space is calculated, and the contrastive loss is calculated based on the distance; for the masked language modeling task, the cross-entropy loss of the model predicting the masked words is calculated, etc. The loss function values ​​of each pre-training task are weighted and summed to obtain the total loss function value. The gradient of the model parameters is calculated using the backpropagation algorithm, and the model parameters are updated using an optimizer (such as the Adam optimizer) based on the learning rate and gradient. By repeatedly performing this process and gradually adjusting the model's parameters, the loss function value of the model on the training set is continuously reduced, thereby improving the model's ability to understand and represent multimodal information.

[0068] Furthermore, the step of inputting keyframe sequences, audio segments, operation event sequences, and game state variables into a pre-trained ViFi-CLIP model for feature extraction to obtain video frame feature vectors, audio feature vectors, and text semantic feature vectors specifically includes: The keyframe sequence is input into the image encoder branch of the pre-trained ViFi-CLIP model. Each frame image is independently encoded by the image encoder branch to generate a frame-level visual embedding vector. The temporal differential attention module performs feature pooling operation on the frame-level visual embedding vector in the temporal dimension to obtain the video frame feature vector. The audio segment is input into the audio encoder branch of the pre-trained ViFi-CLIP model. The audio encoder branch converts the original audio waveform into a Mel spectrogram representation and extracts acoustic features to obtain an audio feature vector. Each operation event in the operation event sequence is converted into a structured text description, and each set of state variables in the game state variable sequence is converted into a numerical text representation. The structured text description and numerical text representation are input into the text encoder branch of the pre-trained ViFi-CLIP model to extract semantic features and obtain text semantic feature vectors.

[0069] In this embodiment, the keyframe sequence generated through multimodal data parsing and preprocessing is input into the image encoder branch of the pre-trained ViFi-CLIP model. The image encoder branch possesses powerful image feature extraction capabilities, independently encoding each frame. During processing, it analyzes various visual elements in the image, such as color distribution, shape features, and texture details, and transforms this visual information into frame-level visual embedding vectors. These frame-level visual embedding vectors can initially represent the visual features of each frame, but they exist independently and lack temporal correlation information.

[0070] To enable video frame features to reflect the dynamic changes of the video over time, a temporal differential attention module is introduced. This module performs temporal feature pooling on the frame-level visual embedding vectors. Specifically, it analyzes the visual differences between adjacent frames, assigns different weights to each frame based on these differences, and then aggregates the frame-level visual embedding vectors of all frames according to their weights. In this way, frames with significant changes in the video can be highlighted while preserving the overall visual information, ultimately obtaining video frame feature vectors that accurately reflect the temporal and spatial characteristics of the video frames.

[0071] For each audio segment, it is fed into the audio encoder branch of the pre-trained ViFi-CLIP model. The audio encoder branch first converts the raw audio waveform into a Mel spectrogram representation. A Mel spectrogram is a graphical representation that reflects the changes in audio frequency over time. It simulates the human ear's perception of sound frequencies and can more effectively extract important information from the audio.

[0072] After obtaining the Mel spectrogram, the audio encoder branch further extracts acoustic features. It analyzes the frequency distribution, energy variations, and other characteristics in the Mel spectrogram, transforming this acoustic information into representative audio feature vectors. These audio feature vectors accurately describe the acoustic characteristics of audio segments, such as pitch, volume, and timbre, providing crucial audio information for subsequent multimodal fusion.

[0073] The operation event sequence and game state variable sequence are non-textual data. In order to extract semantic features using the text encoder branch of the ViFi-CLIP model, they need to be converted into text form first.

[0074] For a sequence of action events, each action event is converted into a structured text description. For example, in a role-playing game, a player's action event might be "the character uses skill A to attack monster B". Converting it into a structured text description clearly expresses information such as the subject, action, and object of the action.

[0075] For a sequence of game state variables, each set of state variables is converted into a numerical text representation. For example, if the character's health is 100, level is 5, and the scene is "forest", this information can be converted into a numerical text representation such as "Health: 100, Level: 5, Scene: Forest".

[0076] After converting the sequence of operation events and the sequence of game state variables into text, the structured text descriptions and numerical text representations are input into the text encoder branch of the pre-trained ViFi-CLIP model. The text encoder branch performs semantic analysis on these texts, understanding keywords, phrases, and sentence structures, extracting semantic information, and converting it into text semantic feature vectors. These text semantic feature vectors accurately reflect the semantic content inherent in the operation events and game states, providing semantic support for subsequent multimodal fusion.

[0077] Furthermore, the construction of a cross-modal interaction graph with visual entities, acoustic events, and semantic entities as nodes based on fine-grained enhancement features, and the iterative updating of the node representations through a graph convolutional network to obtain multimodal fusion features, specifically includes: Each visual entity, each acoustic event, and each semantic entity in the fine-grained enhancement features is treated as an independent node in the cross-modal interaction graph, and a corresponding feature vector is initialized for each node. Calculate the semantic similarity between any two nodes, whereby the semantic similarity is used to measure the degree of semantic association between entities of different modalities; A cross-modal interaction graph is constructed based on semantic similarity, with visual entity nodes, acoustic event nodes, and semantic entity nodes as basic units. Edges are established for node pairs with semantic similarity exceeding a preset similarity threshold, and semantic similarity is used as the initial weight of the edges. The constructed cross-modal interaction graph is input into a multi-layer graph convolutional network. In each layer of graph convolution, each node aggregates the features of all its neighboring nodes and fuses the aggregated neighbor information with the node's own features to obtain the updated feature vectors of each node. Perform global pooling on all updated node feature vectors to generate multimodal fusion features.

[0078] In this embodiment, the fine-grained enhancement features are processed, and each visual entity, acoustic event, and semantic entity contained therein is treated as an independent node in the cross-modal interaction graph. For example, in a game screen recording, visual entities might be game characters, items, scenes, etc.; acoustic events might be game sound effects, character dialogues, etc.; and semantic entities might be textual descriptions related to game operations and plot. A corresponding feature vector is initialized for each node. These feature vectors are preliminary feature representations of each entity, containing basic feature information of the entity in its respective modality, laying the foundation for subsequent interaction and fusion.

[0079] Calculate the semantic similarity between any two nodes. Semantic similarity measures the degree of semantic association between entities of different modalities. For example, a visual entity might be the protagonist in a game, and a semantic entity might be the text describing the protagonist's skills. By calculating their semantic similarity, we can determine the degree of semantic association between the text description and the visual entity of the protagonist. Calculating semantic similarity comprehensively considers multiple dimensions of the node feature vectors and uses specific semantic analysis methods to analyze their semantic similarity, thereby obtaining a specific similarity value.

[0080] Based on the calculated semantic similarity, a cross-modal interaction graph is constructed, using visual entity nodes, acoustic event nodes, and semantic entity nodes as basic units. A preset similarity threshold is set. For node pairs whose semantic similarity exceeds this threshold, an edge is established between the two nodes, with their semantic similarity serving as the initial weight of the edge. For example, if the semantic similarity between the visual entity "magic attack" and the acoustic event "magic release sound effect" exceeds the preset threshold, then an edge is established between these two nodes, with the weight being their calculated semantic similarity. In this way, entities from different modalities are connected, forming a cross-modal interaction graph that reflects the semantic relationships between multimodal information.

[0081] The constructed cross-modal interaction graph is input into a multi-layer graph convolutional network. In each graph convolutional operation, for each node in the graph, it aggregates the features of all its neighboring nodes. For example, a visual entity node has multiple neighboring nodes, including acoustic event nodes and semantic entity nodes; the features of these neighboring nodes are aggregated in a specific way. Then, the aggregated neighbor information is fused with the node's own features to obtain updated feature vectors for each node. Through iterative processing by the multi-layer graph convolutional network, each node can continuously absorb information from its surrounding neighbors, thereby gradually enriching its own feature representation and better reflecting the interaction and fusion between multimodal information.

[0082] After iteratively updating the multi-layer graph convolutional network, global pooling is performed on all updated node feature vectors. Global pooling comprehensively processes the feature vectors of all nodes, extracting the most representative feature information to generate a multimodal fusion feature. This multimodal fusion feature integrates information from multiple modalities, including visual, acoustic, and semantics, comprehensively and accurately reflecting various information in the game recording video, providing strong support for subsequent game advertising video generation.

[0083] This embodiment constructs a cross-modal interaction graph based on fine-grained enhancement features, and iteratively updates the representation of each node through a graph convolutional network, ultimately obtaining multimodal fusion features. This effectively integrates information from different modalities and enhances the ability to utilize multimodal information in the production of game advertising videos.

[0084] In some embodiments, step S103 above, which involves identifying visual defects in the original material based on multimodal fusion features and calling a pre-trained generative model to perform video outlining and special effects enhancement on the defective areas, specifically includes: The multimodal fusion features are input into the pre-trained multi-visual artifact detection model. The artifact-aware dynamic feature extractor in the multi-visual artifact detection model extracts the spatial features related to defects in each video frame, and identifies the various types of visual defects and their corresponding defect areas in the original material. For each defective region, spatiotemporal range analysis, contextual content analysis, and semantic importance assessment are performed. Combined with semantic entity information in multimodal fusion features and game context, targeted video completion strategies and special effects enhancement strategies are generated. For defect areas that need to be expanded or have their backgrounds supplemented, a pre-trained video background drawing generator based on a diffusion model is invoked to generate background extension content that is consistent with the style of the original video and spatiotemporally coherent, based on the contextual content of the defect area and the scene semantics in the multimodal fusion features. For defective areas that are determined to require enhanced visual appeal or to highlight game selling points, a pre-trained text-driven video effects generation model is invoked. Based on the semantic entity information in the multimodal fusion features and the preset selling point prompts, visual effects that semantically match the game content are generated.

[0085] In this embodiment, multimodal fusion features are input into a pre-trained multi-visual artifact detection model. This multi-visual artifact detection model possesses powerful feature analysis capabilities. Its artifact-aware dynamic feature extractor performs detailed defect-related spatial feature extraction for each video frame of the screen recording. It analyzes potential anomalies from various dimensions of the video frame, accurately identifying various types of visual defects in the original material through comparison and judgment with normal visual features. These defects include blurring, color distortion, and ghosting caused by image jitter, while simultaneously determining the specific location of these defects within the video frame. This process fully utilizes the rich information contained in the multimodal fusion features, enabling a more comprehensive and accurate discovery of visual problems in the original material.

[0086] After identifying visual defects and their regions, each defect region undergoes multi-dimensional analysis. Spatiotemporal range analysis aims to clarify the specific manifestation of the defect in time and space, such as whether the defect persists within a specific time period or appears in a specific location on the screen. Contextual content analysis studies the characteristics of the normal video content surrounding the defect region, understanding its scene, objects, and other information. Semantic importance assessment determines the importance of the defect region to the overall video content. Combining semantic entity information from multimodal fusion features—which includes semantic descriptions of various objects and actions in the video, as well as game context information such as the game's plot background and operation scenes—targeted video completion and special effects enhancement strategies are generated. For example, if the defect region is part of an important game scene that is crucial for showcasing the game's features, the completion strategy might focus on restoring the scene as realistically as possible; if the defect region is used to highlight a specific exciting operation in the game, the special effects enhancement strategy might consider adding effects that accentuate the operation.

[0087] For defect areas that require expanded field of view or background augmentation, a pre-trained video background rendering generator based on a diffusion model is invoked. The diffusion model possesses powerful generation capabilities, able to generate high-quality content based on given conditions. The video background rendering generator relies on the contextual content of the defect area and the scene semantics from multimodal fusion features. The contextual content provides information about the surrounding normal video, while the scene semantics clarifies the game scene type and characteristics to which the area belongs. Using this information, the video background rendering generator can generate background extension content that is consistent in style and spatiotemporally coherent with the original video. For example, if the original video is a scene of a character fighting in a forest, and the defect area is a missing part of the forest background, the video background rendering generator will generate appropriate background extension content based on information such as the surrounding trees, lighting, and the overall style of the forest in the game scene, making the entire image look complete and natural.

[0088] For areas identified as needing enhanced visual appeal or to highlight game selling points, a pre-trained text-driven video effects generation model is invoked. This model generates corresponding visual effects based on textual information. The semantic entity information in the multimodal fusion features contains semantic descriptions of various elements in the video, while pre-defined selling point keywords clearly define the game's key features and advantages. The text-driven video effects generation model combines this semantic entity information and selling point keywords to generate visual effects that semantically match the game content. For example, if a game's selling point is a character's powerful magical attack, when a defect related to this magical attack is identified, the model will generate visual effects such as dazzling magical light or stunning explosion effects based on the selling point keyword "powerful magical attack" and the semantic representation of this magic in the video, thereby highlighting the game's selling point and enhancing the video's appeal.

[0089] This embodiment can accurately identify visual defects in the original material based on multimodal fusion features, and enhance the defective areas with video rendering and special effects by calling the corresponding pre-trained generative model, effectively improving the quality of the original material and laying the foundation for generating high-quality game advertising videos in the future.

[0090] In one possible implementation, the multi-visual artifact detection model employs an architecture combining multiple advanced technologies. Its core component is an artifact-aware dynamic feature extractor, which consists of multiple convolutional neural network (CNN) layers and an attention mechanism module. The CNN layers automatically learn local features in video frames, progressively extracting feature information from low to high levels through different levels of convolutional operations. The attention mechanism module enhances the model's focus on important features by assigning different weights based on the importance of features in different regions, allowing the model to focus more intently on features related to visual defects. Following the artifact-aware dynamic feature extractor, a classification layer and a regression layer are connected. The classification layer determines the presence and type of visual defects in the video frame, while the regression layer accurately predicts the location and extent of the defect region. Through this architectural design, the model can comprehensively and accurately detect various visual defects in videos.

[0091] The pre-training steps of the multi-visual artifact detection model include: extensively collecting materials from various game gameplay videos, covering different game types (such as role-playing, strategy, and competitive games), different visual styles (such as realistic, cartoon, and pixel art), and videos from different recording devices and environments; for each video frame, labeling the types of visual defects and their corresponding defect regions to form a visual defect dataset. This visual defect dataset is then input into the multi-visual artifact detection model for training. During the training of each batch of data, the model first extracts features from the video frames using an artifact-aware dynamic feature extractor, obtaining the feature representation of each frame. Then, the classification layer determines whether visual defects exist in the video frames and their types based on these features, while the regression layer predicts the location and extent of the defect regions. The loss function value between the model's prediction results and the labeled data is calculated. For classification tasks, the cross-entropy loss function is used to measure the difference between the predicted defect type and the true type; for regression tasks, the mean squared error loss function is used to measure the difference between the predicted defect region location and extent and the true value. The classification loss and regression loss are weighted and summed to obtain the total loss function value. Based on the total loss function value, the gradient of the model parameters is calculated using the backpropagation algorithm, and the model parameters are updated using an optimizer (such as a stochastic gradient descent optimizer or its variant) according to the learning rate and gradient. By continuously repeating this process, the model parameters are gradually adjusted, causing the loss function value of the model on the training set to continuously decrease, thereby improving the model's ability to detect visual defects.

[0092] In one possible implementation, the pre-training steps of the video outlining generator based on the diffusion model include: constructing the core architecture of the video outlining generator based on the diffusion model, which mainly consists of a forward diffusion process and a reverse denoising process. In the forward diffusion process, Gaussian noise is gradually added to the original video frame, and after a certain number of steps, the original frame is transformed into a pure noise image. The reverse denoising process is the target of model learning, that is, gradually removing noise from the pure noise image to recover an image similar to the original video frame or meeting specific requirements. In the model architecture, an encoder-decoder structure is adopted. The encoder is used to extract features from the input video frame, downsampling the video frame through a multi-layer convolutional neural network (CNN), gradually reducing the spatial resolution and increasing the number of channels, thereby capturing feature information at different levels. The decoder is responsible for reconstructing the image from the features extracted by the encoder, gradually restoring the spatial resolution of the image through upsampling operations, and simultaneously combining the noise information from the diffusion process to generate the outlining video frame. Furthermore, an attention mechanism is introduced into the model, enabling the model to focus on important features in the video frame related to the outlining region, improving the quality and coherence of the generated content.

[0093] To train a model that accurately generates spatiotemporally coherent background extension content consistent with the style of the original video, a suitable loss function needs to be determined. The loss function mainly includes reconstruction loss, perceptual loss, and diffusion loss. Reconstruction loss measures the pixel-level difference between the generated out-of-frame video and the real target frame, typically calculated using mean squared error (MSE) or mean absolute error (MAE). By minimizing the reconstruction loss, the model can restore the pixel information of the target frame as accurately as possible. Perceptual loss, based on a pre-trained deep neural network (such as the VGG network), calculates the difference between the generated frame and the real frame in the high-level feature space. Perceptual loss captures the semantic information of the image, making the generated out-of-frame content more visually consistent with human perception habits, improving the quality and realism of the generated image. Diffusion loss guides the model to gradually remove noise during the reverse denoising process, ensuring the generation process conforms to the principles of a diffusion model. By calculating the difference between the noise estimate and the real noise of the generated frame at different diffusion steps, the model is guided to learn the correct denoising strategy. The three loss functions are weighted and summed to obtain the total loss function, which is then optimized to train the model.

[0094] To train a video rendering generator based on a diffusion model, a large amount of game video data with rich scenes and diverse content needs to be collected, covering different game types such as role-playing, strategy, and action games, to ensure that the model can learn the features of various game scenes. The collected video data is preprocessed, and the preprocessed video frame data is input into the diffusion model-based video rendering generator in batches for training. During the training of each batch of data, a forward diffusion process is first performed, gradually adding Gaussian noise to the input video frames according to the set noise scheduling parameters to obtain noisy images at different diffusion steps. Then, the noisy images are input into the model for a reverse denoising process. Based on the features extracted by the encoder and the noise information at the current diffusion step, the model gradually generates rendered video frames through the decoder. The total loss function value between the generated frames and the real target frames is calculated. Based on the total loss function value, the gradient of the model parameters is calculated using the backpropagation algorithm, and the model parameters are updated using an optimizer (such as the Adam optimizer) based on the learning rate and gradient. By repeatedly performing this process and gradually adjusting the model's parameters, the loss function value of the model on the training set is continuously reduced, thereby improving the model's ability to generate video background content.

[0095] In one possible implementation, the text-driven video effects generation model employs an encoder-decoder structure. The encoder consists of a text encoder and a video encoder. The text encoder uses a pre-trained language model, such as BERT, to encode the input text description information into a high-dimensional semantic feature vector, capturing semantic information and key instructions within the text. The video encoder uses a convolutional neural network (CNN) to extract features from the input video frames. Through multiple convolutional and pooling operations, it gradually reduces the spatial resolution and increases the number of channels to extract visual features from the video frames, including the shape, color, and position of game elements. The decoder is responsible for generating video effects based on the features extracted by the text and video encoders. A generative adversarial network (GAN) architecture is used. The generator generates effect images based on text and video features, while the discriminator determines whether the generated effect images are realistic and match the text description and video content. An attention mechanism is introduced into the generator, enabling the model to focus on key regions in the video frames related to the text description, improving the accuracy and relevance of effect generation. Meanwhile, a temporal modeling module, such as a Long Short-Term Memory (LSTM) network or a Temporal Convolutional Network (TCN), is added to the model to process the temporal information of the video, ensuring that the generated special effects are coherent in the time dimension and conform to the development of the game scene.

[0096] To train a model to accurately generate visual effects that match video content based on text descriptions, a suitable loss function needs to be determined. The main loss functions include adversarial loss, reconstruction loss, and semantic matching loss. Adversarial loss guides the generator and discriminator in adversarial training. The generator aims to generate realistic effect images that the discriminator cannot distinguish between generated and real images; the discriminator aims to determine as accurately as possible whether the input image is generated or real. Optimizing the adversarial loss makes the generated effect images more realistic and natural. Reconstruction loss measures the pixel-level difference between the generated effect image and the real effect image (if labeled real effect data is available), typically calculated using mean squared error (MSE) or mean absolute error (MAE). By minimizing the reconstruction loss, the model can restore the pixel information of the effect image as accurately as possible. Semantic matching loss ensures that the generated effect is semantically consistent with the text description. By calculating the similarity between the feature vector of the generated effect and the feature vector of the text description, such as cosine similarity, the model is guided to generate effects that semantically match the text. The three loss functions are weighted and summed to obtain the total loss function, which is then used to train the model.

[0097] To pre-train a text-driven video effects generation model, it is necessary to collect a wide range of video data covering various game types, styles, and scenes, along with matching text descriptions. Both the video data and text descriptions are then preprocessed. The preprocessed video frame data and corresponding text descriptions are input into the text-driven video effects generation model in batches for training. During training with each batch of data, firstly, the text encoder encodes the text descriptions into semantic feature vectors, and the video encoder extracts features from the video frames to obtain visual feature vectors. Then, the generator generates effects images based on the text and video feature vectors. The generated effects images are then input into a discriminator along with real video frames (if labeled with real effects data is available). The discriminator determines whether the input image is generated or real and calculates the adversarial loss. Simultaneously, the reconstruction loss between the generated and real effects images, as well as the semantic matching loss between the generated effects feature vectors and the text description feature vectors, are calculated. Based on the total loss function value, the gradient of the model parameters is calculated using the backpropagation algorithm, and the model parameters are updated using an optimizer (such as the Adam optimizer) based on the learning rate and gradient. By repeatedly performing this process and gradually adjusting the model's parameters, the loss function value of the model on the training set is continuously reduced, thereby improving the model's ability to generate video effects.

[0098] In some embodiments, step S104 above, which involves generating an editing script based on a high-quality material library and prior game information, using a multimodal large language model combined with a dynamic Prompt project, specifically includes: Obtain prior information about the target game, including game introduction, gameplay instructions, official promotional images, trailer videos, and user review summaries; By analyzing prior information using a multimodal large language model, prior game features containing core selling points, target audience, emotional tone, and key gameplay elements are extracted. Build a dynamic Prompt template library for different game types and promotional goals, and match the most suitable Prompt template based on the game's prior characteristics; Candidate material segments are retrieved from a high-quality material library according to segment type weight and content tags, and a multimodal input representation containing keyframe images, audio waveforms, text descriptions, sentiment tags, duration information, and performance ratings is generated for each candidate material segment. Multimodal input representations, Prompt templates, and game prior features are input into a multimodal large language model for fusion processing to form an editing script.

[0099] In this embodiment, the prior information of the target game includes: a game introduction, which allows the model to quickly understand the game's general background and theme; gameplay instructions, which detail the game's operation methods and core gameplay mechanics; official promotional images, which visually present the game's style and visual elements; trailer videos, which showcase exciting clips and unique features of the game; and user review summaries, reflecting players' genuine feelings and concerns about the game. By collecting this multifaceted prior information, a sufficient data foundation is provided for subsequent model analysis.

[0100] A multimodal large language model is used to perform deep analysis of the acquired prior information. This model possesses powerful natural language processing and multimodal information understanding capabilities, enabling it to accurately extract key game prior features from the prior information. These features include the game's core selling points, such as unique gameplay, innovative systems, or exquisite graphics; the target audience, clearly defining the main types of players the game is aimed at, such as age, gender, and gaming preferences; the emotional tone, determining the emotional atmosphere the game wants to convey to players—whether it's tense and exciting, lighthearted and enjoyable, or deeply moving; and key gameplay elements, such as unique in-game items and important levels. By extracting these features, the model can comprehensively grasp the game's characteristics and promotional focus.

[0101] To better guide the multimodal large language model in generating editing scripts that meet game requirements and promotional goals, a dynamic Prompt template library was constructed, tailored to different game genres and promotional objectives. This library is meticulously categorized and designed according to various game genres and promotional scenarios, containing multiple Prompt templates with different styles and focuses. After extracting prior game features, the most suitable Prompt template is matched from the library based on these features. For example, if the game is an action-adventure game targeting young male players, and the promotional focus is on the game's exciting combat and flashy skills, the model will match a Prompt template from the library that emphasizes action elements and combat scenes, providing clear guidance for subsequent script generation.

[0102] Candidate clips are retrieved from a high-quality material library based on clip type weights and content tags. Clip type weights are set according to the importance of different types of materials in the ad video; for example, clips featuring exciting gameplay or unique scenes may have higher weights. Content tags provide detailed descriptions of the clips, such as scene type, gameplay type, and emotional atmosphere. This method accurately selects representative candidate clips that are relevant to the game's prior features. Then, a multimodal input representation is generated for each candidate clip, including keyframe images, audio waveforms, text descriptions, sentiment tags, duration information, and an appeal score. Keyframe images visually display the core visuals of the clip; audio waveforms reflect the clip's audio characteristics; text descriptions briefly explain the clip's content; sentiment tags reflect the emotional atmosphere conveyed by the clip; duration information clarifies the clip's duration; and the appeal score evaluates the quality and attractiveness of the material. These multimodal input representations provide the model with comprehensive and detailed material information.

[0103] Finally, the generated multimodal input representation, the matched Prompt template, and the extracted game prior features are input into a multimodal large language model for fusion processing. The multimodal large language model comprehensively considers this information, and based on the guidance of the Prompt template and the requirements of the game prior features, it rationally sorts and combines candidate clips, while adding necessary transition effects and text descriptions to form a complete editing script. This script not only conforms to the game's promotional goals and stylistic characteristics but also makes full use of high-quality materials from the resource library, highlighting the game's core selling points and exciting content, providing clear and accurate guidance for subsequent automated editing.

[0104] In some embodiments, step S105 above, converting the editing script into video editing instructions and calling the editing engine to perform automated editing, specifically includes: The script parser performs syntax checks and semantic verification on the editing script, generating a sequence of editing instructions. The editing command sequence is input into the video editing engine for editing operations to obtain the initial video sequence; The pre-trained video quality assessment model is invoked to perform a quality self-check on the initial video sequence. If the quality score is lower than the preset score threshold, a re-editing process is triggered.

[0105] In this embodiment, the script parser possesses powerful text processing capabilities, performing a comprehensive and meticulous grammar check on the generated editing script. During the grammar check, it analyzes the statement structure, symbol usage, and other aspects of the script line by line according to preset script grammar rules to ensure compliance with standards. For example, it checks for spelling errors, incomplete or redundant sentence components, and other grammatical issues. After completing the grammar check, the script parser further performs semantic verification. It considers the business logic and common scenarios of game advertising video production to determine whether the editing intent expressed by the script is reasonable and clear. For example, it verifies whether the descriptions of operations such as material switching and adding effects in the script meet actual editing needs and whether there are logical contradictions or semantic content that is difficult to achieve. After grammar and semantic verification, the script parser converts the editing script into a series of accurate and standardized editing instruction sequences. These editing instruction sequences are arranged in a specific format and order, clearly indicating the specific tasks that the video editing engine needs to perform in subsequent operations, such as from which time point to insert which material, what effects to add, and the parameter settings for the effects.

[0106] After generating the editing instruction sequence, it is input into the video editing engine. The video editing engine is a powerful and professional software tool capable of accurately recognizing and parsing the input editing instruction sequence. Based on the instructions in the sequence, the video editing engine begins executing the corresponding editing operations. It accurately locates the required footage segments from a high-quality media library and performs operations such as splicing, cropping, and sorting the footage according to the instructions. Simultaneously, for any special effects requested in the instructions, the video editing engine accurately adds and adjusts them according to preset effect parameters. After completing all editing operations, the video editing engine outputs an initial video sequence. This initial video sequence is generated according to the preliminary intent of the editing script, but it may still contain some potential quality issues, requiring further quality evaluation and optimization.

[0107] To ensure high-quality generated advertising videos, a pre-trained video quality assessment model is invoked to perform a quality self-check on the initial video sequence. This model, trained on a large number of high-quality advertising video samples, possesses the ability to comprehensively evaluate video quality. During the evaluation process, the model analyzes the initial video sequence from multiple dimensions, including image sharpness, color saturation, audio quality, editing rhythm, and content coherence. Based on the evaluation results of these dimensions, the model provides a comprehensive quality score for the initial video sequence. This score is then compared to a preset score threshold. If the quality score is lower than the preset threshold, it indicates that the initial video sequence has quality issues in some aspects and cannot achieve the expected promotional effect. In this case, the system automatically triggers a re-editing process. The re-editing process re-examines the editing script and editing instruction sequence, analyzes potential quality issues, and makes corresponding adjustments and optimizations to the editing script. Then, the adjusted editing script is converted into an editing instruction sequence again, input into the video editing engine for editing operations, generates a new initial video sequence, and performs a quality self-check until the quality score reaches or exceeds the preset score threshold.

[0108] This embodiment can accurately convert editing scripts into video editing instructions and call the video editing engine to perform automated editing operations. At the same time, through quality self-checking and re-editing processes, it ensures that the generated advertising videos have high quality and meet the needs of game promotion.

[0109] In some embodiments, step S105 above, which involves fine-tuning the editing strategy based on historical user feedback using reinforcement learning, specifically includes: Collect user interaction data during the viewing process, clean and normalize it, and generate reward signals. The interaction data includes video click-through rate, completion rate, number of likes, number of comments, conversion rate, and user retention time. Based on reward signals, a state space is set that includes game type, promotional goals, metadata features of the current edited video, and historical user feedback statistics, and an action space that includes segment type selection weights, special effects type preferences, transition style preferences, and duration allocation ratios. A strategy network and a value network are then constructed. Extract historical generated videos and their corresponding user feedback data from the advertising video library to construct a training sample set; Based on the state space and action space, the training sample set is used as input, and the proximal policy optimization algorithm is used to iteratively train the policy network and value network. The policy network parameters are updated by maximizing the expected cumulative reward. The generation strategy for subsequent editing scripts is adjusted based on the trained policy network and value network.

[0110] In this embodiment, a comprehensive data collection system is constructed to collect various interaction data from users watching game advertising videos. This interaction data covers several important metrics, including video click-through rate (CTR), reflecting the user's initial interest in the advertising video; completion rate, reflecting the user's willingness to watch the entire video; number of likes, directly indicating the user's approval of the video content; number of comments, reflecting user participation and discussion; conversion rate, the proportion of users who go from watching the video to actually performing game-related actions (such as downloading or registering), a key indicator for measuring advertising effectiveness; and user retention time, reflecting the video's ability to attract and retain user attention. After collecting this raw data, it is cleaned and normalized. The cleaning process mainly removes abnormal and erroneous data to ensure the accuracy and reliability of the data; normalization unifies data with different dimensions to the same range, facilitating subsequent analysis and calculation. After processing, a reward signal is generated based on this interaction data. This reward signal will serve as important feedback information in reinforcement learning, guiding the adjustment of editing strategies.

[0111] Based on the generated reward signal, the state space and action space are further defined. The state space contains information across multiple dimensions to comprehensively describe the current ad video. These include: game type (different game types have different characteristics and target audiences, requiring different editing strategies); promotional objectives (e.g., promoting a new game, increasing user activity, or encouraging in-game spending; different objectives require different editing focuses and techniques); metadata characteristics of the current edited video, such as video length, visual style, and music rhythm; and historical user feedback statistics (analyzing past user feedback data to summarize user preferences and needs for different video types). The action space defines the various operations that can be performed during editing, including segment type selection weights (the proportion of different game segments (e.g., combat, story, social segments) in the edit); special effects type preferences (e.g., the frequency of use of lighting effects, particle effects, blur effects); transition style preferences (e.g., the choice of fade-in / fade-out, flash cut, rotation, etc.); and duration allocation ratios (the reasonable allocation of duration across different parts of the video). After setting up the state space and action space, a policy network and a value network are constructed. The policy network is responsible for selecting the optimal action based on the current state, i.e., determining the best editing strategy; the value network is used to evaluate the value of the current state, helping the policy network make more reasonable decisions.

[0112] To effectively train the policy network and value network, a suitable training sample set needs to be constructed. A large number of historical generated videos and their corresponding user feedback data are extracted from an advertising video library as the sample source. For each historical generated video, its relevant editing parameters, video features, and user feedback data are recorded. This data is then organized and labeled according to a specific format to form training samples. The training samples should contain rich information so that the network can learn the impact of various actions on the reward signal under different states, thereby better adjusting the editing strategy.

[0113] Using a pre-constructed training sample set as input, the policy network and value network are iteratively trained using a proximal policy optimization algorithm. In each iteration, based on the current state space and action space, the policy network generates an action, thus determining an editing policy. Then, an edited video sample is generated based on this action, and the corresponding reward signal is obtained. The value network evaluates the value of the current state based on the reward signal. By continuously repeating this process, the expected cumulative reward is calculated, and the parameters of the policy network and value network are updated using the proximal policy optimization algorithm. This allows the policy network to gradually learn the policy of selecting the optimal action in different states to maximize the expected cumulative reward. After multiple iterations of training, the performance of the policy network and value network will be significantly improved, enabling more accurate adjustments to the editing policy based on user feedback.

[0114] Based on the trained policy network and value network, the generation strategy for subsequent editing scripts is adjusted. The policy network can select optimal actions such as segment type selection weights, effect type preferences, transition style preferences, and duration allocation ratios based on state information such as the current game type, promotional goals, video metadata characteristics, and historical user feedback statistics. Based on these actions, editing scripts that better meet user needs and preferences are generated, thereby improving the quality and attractiveness of game advertising videos and better achieving the game's promotional goals.

[0115] This embodiment can effectively fine-tune the editing strategy based on historical user feedback using reinforcement learning, making the generation of game advertising videos more intelligent and personalized, and meeting the needs and expectations of different users.

[0116] Reference Figure 2 An embodiment of the present invention provides an artificial intelligence-based automatic generation system 2 for game advertising videos, the system 2 specifically comprising: The screen recording module 201 is used to train an intelligent agent based on a hierarchical reward shaping mechanism. The intelligent agent performs automated operations and synchronous screen recording on the target game. The hierarchical reward shaping mechanism includes basic game rewards, exploration rewards, skill rewards and story rewards. The feature fusion module 202 is used to extract cross-modal features from the screen recording video using the CLIP multimodal model in response to the completion of recording, and to perform alignment and fusion through a graph neural network to obtain multimodal fused features; The defect correction module 203 is used to identify visual defects in the original material based on multimodal fusion features, and call a pre-trained generative model to perform video rendering and special effects enhancement on the defect area to form a high-quality material library. The script generation module 204 is used to generate editing scripts based on a high-quality material library and game prior information, by combining a multimodal large language model with a dynamic Prompt project. The video editing module 205 is used to convert editing scripts into video editing instructions and call the editing engine to perform automated editing, while fine-tuning the editing strategy based on historical user feedback through reinforcement learning; The video output module 206 is used to render and output the edited video sequence, and adaptively generate advertising videos in various formats according to the requirements of the target platform.

[0117] It is understandable that, such as Figure 1 The content shown in the embodiments of the AI-based game ad video automatic generation method is applicable to the embodiments of this AI-based game ad video automatic generation system. The specific functions implemented in the embodiments of this AI-based game ad video automatic generation system are the same as those shown in the figure. Figure 1 The embodiment of the AI-based automatic generation method for game advertisement videos shown is the same, and the beneficial effects achieved are the same as those described above. Figure 1 The beneficial effects achieved by the illustrated embodiment of the AI-based automatic generation method for game advertising videos are the same.

[0118] It should be noted that the information interaction and execution process between the above systems are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.

[0119] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0120] Reference Figure 3 The present invention also provides a computer device 3, including: a memory 302 and a processor 301, and a computer program 303 stored on the memory 302. When the computer program 303 is executed on the processor 301, it implements the method for automatically generating game advertising videos based on artificial intelligence as described in any of the above methods.

[0121] The computer device 3 may be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device 3 may include, but is not limited to, a processor 301 and a memory 302. Those skilled in the art will understand that... Figure 3 The computer device 3 is merely an example and does not constitute a limitation on the computer device 3. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.

[0122] The processor 301 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0123] In some embodiments, the memory 302 may be an internal storage unit of the computer device 3, such as a hard disk or memory of the computer device 3. In other embodiments, the memory 302 may be an external storage device of the computer device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 3. Furthermore, the memory 302 may include both internal and external storage units of the computer device 3. The memory 302 is used to store the operating system, applications, boot loader, data, and other programs, such as the program code of the computer program. The memory 302 can also be used to temporarily store data that has been output or will be output.

[0124] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the AI-based automatic generation method for game advertising videos as described in any of the above methods.

[0125] In this embodiment, if the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0126] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0127] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0128] In the embodiments disclosed in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0129] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

Claims

1. A method for automatically generating game advertisement videos based on artificial intelligence, characterized in that, The method specifically includes: The intelligent agent is trained based on a hierarchical reward shaping mechanism. The intelligent agent performs automated operations and synchronous screen recording on the target game. The hierarchical reward shaping mechanism includes basic game rewards, exploration rewards, skill rewards and story rewards. In response to the completion of recording, the CLIP multimodal model is used to extract cross-modal features from the screen recording video, and then the graph neural network is used for alignment and fusion to obtain multimodal fusion features; Visual defects in the original material are identified based on multimodal fusion features, and a pre-trained generative model is called to perform video rendering and special effects enhancement on the defective areas, forming a high-quality material library; Based on a high-quality resource library and prior game information, a multimodal large language model is used in conjunction with a dynamic Prompt project to generate editing scripts. The editing script is converted into video editing instructions and the editing engine is called to perform automated editing. At the same time, the editing strategy is fine-tuned through reinforcement learning based on historical user feedback. The edited video sequence is rendered and output, and various advertising video formats are adaptively generated according to the requirements of the target platform.

2. The method according to claim 1, characterized in that, The training of the agent based on the hierarchical reward shaping mechanism specifically includes: Construct the environment interaction interface of the target game and initialize the basic policy network, and establish a communication connection between the agent and the game environment through the basic policy network; Construct a multi-objective reward function that includes basic game rewards, exploration rewards, skill rewards, and story rewards; Based on the multi-objective reward function, the agent interacts with the game environment to generate experience data, and stores the experience data in the experience pool. A small batch of experience data is randomly sampled from the experience pool. The policy gradient is calculated using the sampled small batch of experience data, and the basic policy network parameters are updated using the backpropagation algorithm to guide the agent to optimize the policy in the direction of maximizing the cumulative reward.

3. The method according to claim 2, characterized in that, The basic game reward is used to measure the progress of completing the basic game objectives. The exploration reward is calculated based on the state novelty assessment and is used to incentivize the agent to explore unvisited areas. The skill reward is calculated based on the operation complexity assessment and is used to incentivize the agent to execute high-difficulty operation sequences. The plot reward is calculated based on the game plot trigger point detection and is used to incentivize the agent to activate hidden plot events.

4. The method according to claim 1, characterized in that, The method involves using the CLIP multimodal model to extract cross-modal features from screen-recorded videos and then aligning and fusing them using a graph neural network to obtain multimodal fused features. Specifically, this includes: Multimodal data parsing and preprocessing are performed on screen recording videos to generate keyframe sequences, audio clips, operation event sequences, and game state variables; The keyframe sequence, audio clips, operation event sequence, and game state variables are input into the pre-trained ViFi-CLIP model for feature extraction, resulting in video frame feature vectors, audio feature vectors, and text semantic feature vectors. Semantic entity extraction is performed on video frame feature vectors, audio feature vectors, and text semantic feature vectors respectively to obtain fine-grained enhanced features containing visual entities, acoustic events, and semantic entities; Based on the fine-grained enhancement features, a cross-modal interaction graph with visual entities, acoustic events, and semantic entities as nodes is constructed, and the representation of each node is iteratively updated through a graph convolutional network to obtain multimodal fusion features.

5. The method according to claim 4, characterized in that, The process of inputting keyframe sequences, audio segments, operation event sequences, and game state variables into a pre-trained ViFi-CLIP model for feature extraction to obtain video frame feature vectors, audio feature vectors, and text semantic feature vectors specifically includes: The keyframe sequence is input into the image encoder branch of the pre-trained ViFi-CLIP model. Each frame image is independently encoded by the image encoder branch to generate a frame-level visual embedding vector. The temporal differential attention module performs feature pooling operation on the frame-level visual embedding vector in the temporal dimension to obtain the video frame feature vector. The audio segment is input into the audio encoder branch of the pre-trained ViFi-CLIP model. The audio encoder branch converts the original audio waveform into a Mel spectrogram representation and extracts acoustic features to obtain an audio feature vector. Each operation event in the operation event sequence is converted into a structured text description, and each set of state variables in the game state variable sequence is converted into a numerical text representation. The structured text description and numerical text representation are input into the text encoder branch of the pre-trained ViFi-CLIP model to extract semantic features and obtain text semantic feature vectors.

6. The method according to claim 4, characterized in that, The process involves constructing a cross-modal interaction graph with visual entities, acoustic events, and semantic entities as nodes based on fine-grained enhancement features, and iteratively updating the representation of each node through a graph convolutional network to obtain multimodal fusion features. Specifically, this includes: Each visual entity, each acoustic event, and each semantic entity in the fine-grained enhancement features is treated as an independent node in the cross-modal interaction graph, and a corresponding feature vector is initialized for each node. Calculate the semantic similarity between any two nodes, whereby the semantic similarity is used to measure the degree of semantic association between entities of different modalities; A cross-modal interaction graph is constructed based on semantic similarity, with visual entity nodes, acoustic event nodes, and semantic entity nodes as basic units. Edges are established for node pairs with semantic similarity exceeding a preset similarity threshold, and semantic similarity is used as the initial weight of the edges. The constructed cross-modal interaction graph is input into a multi-layer graph convolutional network. In each layer of graph convolution, each node aggregates the features of all its neighboring nodes and fuses the aggregated neighbor information with the node's own features to obtain the updated feature vectors of each node. Perform global pooling on all updated node feature vectors to generate multimodal fusion features.

7. The method according to claim 1, characterized in that, The process of identifying visual defects in the original material based on multimodal fusion features and then using a pre-trained generative model to perform video rendering and special effects enhancement on the defective areas specifically includes: The multimodal fusion features are input into the pre-trained multi-visual artifact detection model. The artifact-aware dynamic feature extractor in the multi-visual artifact detection model extracts the spatial features related to defects in each video frame, and identifies the various types of visual defects and their corresponding defect areas in the original material. For each defective region, spatiotemporal range analysis, contextual content analysis, and semantic importance assessment are performed. Combined with semantic entity information in multimodal fusion features and game context, targeted video completion strategies and special effects enhancement strategies are generated. For defect areas that need to be expanded or have their backgrounds supplemented, a pre-trained video background drawing generator based on a diffusion model is invoked to generate background extension content that is consistent with the style of the original video and spatiotemporally coherent, based on the contextual content of the defect area and the scene semantics in the multimodal fusion features. For defective areas that are determined to require enhanced visual appeal or to highlight game selling points, a pre-trained text-driven video effects generation model is invoked. Based on the semantic entity information in the multimodal fusion features and the preset selling point prompts, visual effects that semantically match the game content are generated.

8. The method according to claim 1, characterized in that, The process of generating editing scripts based on a high-quality resource library and prior game information, using a multimodal large language model combined with a dynamic Prompt project, specifically includes: Obtain prior information about the target game, including game introduction, gameplay instructions, official promotional images, trailer videos, and user review summaries; By analyzing prior information using a multimodal large language model, prior game features containing core selling points, target audience, emotional tone, and key gameplay elements are extracted. Build a dynamic Prompt template library for different game types and promotional goals, and match the most suitable Prompt template based on the game's prior characteristics; Candidate material segments are retrieved from a high-quality material library according to segment type weight and content tags, and a multimodal input representation containing keyframe images, audio waveforms, text descriptions, sentiment tags, duration information, and performance ratings is generated for each candidate material segment. Multimodal input representations, Prompt templates, and game prior features are input into a multimodal large language model for fusion processing to form an editing script.

9. The method according to any one of claims 1 to 8, characterized in that, The aforementioned fine-tuning of the editing strategy based on historical user feedback through reinforcement learning specifically includes: Collect user interaction data during the viewing process, clean and normalize it, and generate reward signals. The interaction data includes video click-through rate, completion rate, number of likes, number of comments, conversion rate, and user retention time. Based on reward signals, a state space is set that includes game type, promotional goals, metadata features of the current edited video, and historical user feedback statistics, and an action space that includes segment type selection weights, special effects type preferences, transition style preferences, and duration allocation ratios. A strategy network and a value network are then constructed. Extract historical generated videos and their corresponding user feedback data from the advertising video library to construct a training sample set; Based on the state space and action space, the training sample set is used as input, and the proximal policy optimization algorithm is used to iteratively train the policy network and value network. The policy network parameters are updated by maximizing the expected cumulative reward. The generation strategy for subsequent editing scripts is adjusted based on the trained policy network and value network.

10. An AI-based automatic game advertising video generation system, characterized in that, The system specifically includes: The screen recording module is used to train an intelligent agent based on a hierarchical reward shaping mechanism. The intelligent agent performs automated operations and synchronous screen recording on the target game. The hierarchical reward shaping mechanism includes basic game rewards, exploration rewards, skill rewards, and story rewards. The feature fusion module is used to extract cross-modal features from the screen recording video in response to the completion of recording, and then perform alignment and fusion through a graph neural network to obtain multimodal fused features. The defect correction module is used to identify visual defects in the original material based on multimodal fusion features, and call a pre-trained generative model to perform video rendering and special effects enhancement on the defective areas to form a high-quality material library. The script generation module is used to generate editing scripts based on a high-quality material library and game prior information, by combining a multimodal large language model with a dynamic Prompt project. The video editing module is used to convert editing scripts into video editing instructions and call the editing engine to perform automated editing, while fine-tuning the editing strategy based on historical user feedback through reinforcement learning; The video output module is used to render and output the edited video sequence, and adaptively generate advertising videos in various formats according to the requirements of the target platform.