Method, device, equipment, medium and product for extracting game material from a movie or television

By performing multimodal analysis and completion of film and television videos, game elements are extracted and completed, solving the problems of high cost and long cycle when adapting film and television works into game works, and realizing a high-efficiency and low-cost adaptation process.

CN122493360APending Publication Date: 2026-07-31BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING INST OF TECH
Filing Date
2026-04-29
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing technologies, adapting film and television works into game works presents problems such as high economic costs and long design cycles.

Method used

By performing multimodal analysis on film and television videos, game elements, including story elements, character elements, and scene elements, are extracted and completed. The necessary completion is then performed using a multimodal model, thereby improving the efficiency of game design.

Benefits of technology

It has achieved efficient and low-cost conversion of film and television works into game works, and shortened the production cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493360A_ABST
    Figure CN122493360A_ABST
Patent Text Reader

Abstract

This invention provides a method, apparatus, device, medium, and product for extracting game materials from film and television works, relating to the field of data processing technology. It addresses the drawbacks of high economic costs and long design cycles in adapting film and television works into games, thereby improving the efficiency of this adaptation process. The method includes: performing multimodal analysis on the film and television video to obtain structured multimodal data, which includes data information from multiple storyboard videos; dividing the multimodal data based on its contextual relevance to obtain multiple game plot units, each corresponding to at least one storyboard video; extracting game elements for each game plot unit, including at least one of the following: story elements, character elements, and scene elements; performing multimodal completion processing on the extracted game elements to obtain supplemented game elements, which are then used as game materials to generate a game corresponding to the film and television video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus, device, medium, and product for extracting game materials from film and television. Background Technology

[0002] With the rapid development of the gaming industry, there is an increasing number of games adapted from classic film and television works (such as movies, TV series, animations, web series, documentaries, and stage play recordings), attracting a large number of film and television fans. The core of these adapted games lies in transforming the storylines, characters, and scenes from the films and television shows into interactive narrative elements, character dialogues, and background images.

[0003] However, the design of game elements in existing adapted games mainly relies on manual operation, involving content analysis, character extraction, and scene extraction from film and television works to obtain the corresponding game. This results in high economic costs and long design cycles. Summary of the Invention

[0004] This invention provides a method, apparatus, device, medium, and product for extracting game materials from films and television programs, in order to solve the problems of high economic costs and long design cycles in the existing technology of adapting films and television programs into games, thereby improving the efficiency of adapting films and television programs into games and shortening the cycle.

[0005] This invention provides a method for extracting game materials from films and television programs, comprising the following steps.

[0006] Multimodal analysis is performed on film and television videos to obtain structured multimodal data. The multimodal data includes data information from multiple storyboard videos. The data information of each storyboard video includes at least one of the following: storyboard video sequence, storyboard keyframe sequence, storyboard audio clip, and storyboard text information. Based on the contextual relevance of multimodal data, the multimodal data is divided into multiple game plot units, each of which corresponds to at least one storyboard video. For each game plot unit, extract game elements, which include at least one of the following: story elements, character elements, and scene elements; The extracted game elements are processed using multimodal completion to obtain the supplemented game elements. These supplemented game elements are then used as game assets to generate game works that correspond to the film and television videos.

[0007] According to the present invention, a method for extracting game materials from film and television content involves performing multimodal analysis on film and television videos to obtain structured multimodal data, including: Perform shot boundary detection on film and television videos, and divide the film and television videos into multiple shot videos, with each shot video corresponding to a shot video sequence; Keyframes are extracted from each storyboard video to obtain the storyboard keyframe sequence corresponding to each storyboard video; The audio trajectory corresponding to each storyboard video is segmented to obtain the storyboard audio segment corresponding to each storyboard video. The audio segments corresponding to each storyboard video are converted into audio text, and the subtitles in each storyboard video are recognized to obtain subtitle text, thus obtaining the storyboard text information corresponding to each storyboard video. The storyboard text information includes audio text and subtitle text. Multimodal data is obtained based on the storyboard video sequence, storyboard keyframe sequence, storyboard audio clip, and storyboard text information corresponding to each storyboard video.

[0008] According to a method for extracting game materials from film and television provided by the present invention, the multimodal data is divided based on the contextual relevance of the multimodal data to obtain multiple game plot units, including: The data from multiple storyboard videos are analyzed, and the multiple storyboard videos are divided at target points to obtain multiple game plot units. The target points include at least one of the following: video scene changes, topic changes, character appearance changes, plot conflict points, and emotional curve turning points.

[0009] According to the present invention, a method for extracting game materials from film and television is provided. The game elements include story elements, which include at least one of the following: world view elements, script elements, dialogue elements, character relationship elements, and emotional elements. For each game plot unit, extract game elements, including: Based on the data information of at least one storyboard video corresponding to each game plot unit, determine the plot information of each game plot unit. The plot information includes at least one of the following: plot name, story summary, plot preconditions, plot post-plot results, plot dependencies, and world view elements. Based on the storyboard text information of at least one storyboard video corresponding to each game plot unit, determine script elements, dialogue elements, and character relationship elements; Sentiment analysis was performed on the storyboard text information of at least one storyboard video corresponding to each game plot unit to identify emotional elements.

[0010] According to the present invention, a method for extracting game materials from film and television is provided, wherein the game elements include character elements, and the character elements include at least one of the following: basic information and character feature information; For each game plot unit, extract game elements, including: Based on the data information of at least one storyboard video corresponding to each game plot unit, determine the basic information of each character. The basic information includes at least one of the following: character name, age, gender, personality, and backstory. Based on the keyframe sequence of at least one storyboard video corresponding to each game plot unit, the character feature information of each character is determined. The character feature information is used to ensure that the same character remains consistent in different storyboard videos. The character feature information includes at least one of the following: language features, action features, facial expression features, and appearance features.

[0011] According to a method for extracting game materials from film and television according to the present invention, the game elements include scene elements, and the scene elements include at least one of the following: scene identifier, background image, and scene attribute label; For each game plot unit, extract game elements, including: Based on the keyframe sequence of at least one storyboard video corresponding to each game plot unit, the background of the keyframes is analyzed to determine scene elements.

[0012] According to a method for extracting game elements from film and television provided by the present invention, the extracted game elements are subjected to multimodal completion processing to obtain the completed game elements, including: The story elements are analyzed to determine the completeness parameters of the story elements, and supplementary story elements that satisfy the worldview elements are generated based on the completeness parameters. Analyze the character elements and generate appearance variant images for each character. The appearance variant images include at least one of the following: expression variant image, action variant image, pose variant image, age variant image, clothing variant image, perspective variant image, lighting variant image, and style variant image. Analyze scene elements and generate background variant images for each scene. The background variant images include at least one of the following: weather variant image, season variant image, time period variant image, state variant image, and view distance variant image.

[0013] The present invention also provides an apparatus for extracting game materials from film and television, comprising the following modules: a parsing module, a segmentation module, an extraction module, and a supplementation module; The parsing module is used to perform multimodal parsing on film and television videos to obtain structured multimodal data. The multimodal data includes data information from multiple storyboard videos. The data information of each storyboard video includes at least one of the following: storyboard video sequence, storyboard keyframe sequence, storyboard audio clip, and storyboard text information. The segmentation module is used to segment multimodal data based on the contextual relevance of the multimodal data to obtain multiple game plot units, each of which corresponds to at least one storyboard video. The extraction module is used to extract game elements for each game plot unit. Game elements include at least one of the following: story elements, character elements, and scene elements. The supplementary module is used to perform multimodal completion processing on the extracted game elements to obtain supplemented game elements. These supplemented game elements are then used as game assets to generate game works that correspond to film and video content.

[0014] According to the present invention, an apparatus for extracting game materials from film and television is provided, and the parsing module is specifically used for: Perform shot boundary detection on film and television videos, and divide the film and television videos into multiple shot videos, with each shot video corresponding to a shot video sequence; Keyframes are extracted from each storyboard video to obtain the storyboard keyframe sequence corresponding to each storyboard video; The audio trajectory corresponding to each storyboard video is segmented to obtain the storyboard audio segment corresponding to each storyboard video. The audio segments corresponding to each storyboard video are converted into audio text, and the subtitles in each storyboard video are recognized to obtain subtitle text, thus obtaining the storyboard text information corresponding to each storyboard video. The storyboard text information includes audio text and subtitle text. Multimodal data is obtained based on the storyboard video sequence, storyboard keyframe sequence, storyboard audio clip, and storyboard text information corresponding to each storyboard video.

[0015] According to the present invention, an apparatus for extracting game materials from film and television is provided, comprising modules specifically used for: The data from multiple storyboard videos are analyzed, and the multiple storyboard videos are divided at target points to obtain multiple game plot units. The target points include at least one of the following: video scene changes, topic changes, character appearance changes, plot conflict points, and emotional curve turning points.

[0016] According to the present invention, an apparatus for extracting game materials from film and television is provided. The game elements include story elements, which in turn include at least one of the following: world-building elements, script elements, dialogue elements, character relationship elements, and emotional elements. The extraction module is specifically used for: Based on the data information of at least one storyboard video corresponding to each game plot unit, determine the plot information of each game plot unit. The plot information includes at least one of the following: plot name, story summary, plot preconditions, plot post-plot results, plot dependencies, and world view elements. Based on the storyboard text information of at least one storyboard video corresponding to each game plot unit, determine script elements, dialogue elements, and character relationship elements; Sentiment analysis was performed on the storyboard text information of at least one storyboard video corresponding to each game plot unit to identify emotional elements.

[0017] According to the present invention, an apparatus for extracting game materials from film and television is provided. The game elements include character elements, and the character elements include at least one of the following: basic information and character feature information; the extraction module is specifically used for: Based on the data information of at least one storyboard video corresponding to each game plot unit, determine the basic information of each character. The basic information includes at least one of the following: character name, age, gender, personality, and backstory. Based on the keyframe sequence of at least one storyboard video corresponding to each game plot unit, the character feature information of each character is determined. The character feature information is used to ensure that the same character remains consistent in different storyboard videos. The character feature information includes at least one of the following: language features, action features, facial expression features, and appearance features.

[0018] According to the present invention, an apparatus for extracting game materials from film and television is provided. The game elements include scene elements, and the scene elements include at least one of the following: scene identifier, background image, and scene attribute label; the extraction module is specifically used for: Based on the keyframe sequence of at least one storyboard video corresponding to each game plot unit, the background of the keyframes is analyzed to determine scene elements.

[0019] According to the present invention, an apparatus for extracting game materials from film and television is provided, and a supplementary module is specifically used for: The story elements are analyzed to determine the completeness parameters of the story elements, and supplementary story elements that satisfy the worldview elements are generated based on the completeness parameters. Analyze the character elements and generate appearance variant images for each character. The appearance variant images include at least one of the following: expression variant image, action variant image, pose variant image, age variant image, clothing variant image, perspective variant image, lighting variant image, and style variant image. Analyze scene elements and generate background variant images for each scene. The background variant images include at least one of the following: weather variant image, season variant image, time period variant image, state variant image, and view distance variant image.

[0020] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement any of the methods described above for extracting game materials from movies and television.

[0021] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the methods described above for extracting game materials from film and television.

[0022] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the methods described above for extracting game materials from film and television.

[0023] This invention provides a method, apparatus, device, medium, and product for extracting game materials from film and television. By performing multimodal analysis on film and television videos, structured multimodal data is obtained, which yields data information for multiple storyboard videos. Then, based on the contextual relevance of the multimodal data, the data is divided into multiple game plot units. For each game plot unit, game elements such as story elements, character elements, and scene elements are extracted. The extracted game elements are then subjected to multimodal completion processing to obtain supplemented game elements. These supplemented game elements can then be used as game materials to generate a game corresponding to the film and television video. Based on this, when adapting film and television works into game works, manual operations such as content analysis, character extraction, and scene extraction are unnecessary. This application efficiently extracts and supplements game elements such as story elements, character elements, and scene elements from film and television videos, thereby improving the efficiency of adapting film and television works into game works, shortening the game production cycle, and reducing production costs. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0025] Figure 1 This is one of the flowcharts illustrating the method for extracting game materials from movies and television provided by the present invention.

[0026] Figure 2 This is the second flowchart illustrating the method for extracting game materials from film and television provided by the present invention.

[0027] Figure 3 This is the third flowchart illustrating the method for extracting game materials from film and television provided by the present invention.

[0028] Figure 4 This is the fourth flowchart illustrating the method for extracting game materials from film and television provided by the present invention.

[0029] Figure 5 This is the fifth flowchart illustrating the method for extracting game materials from film and television provided by the present invention.

[0030] Figure 6 This is the sixth flowchart illustrating the method for extracting game materials from film and television provided by the present invention.

[0031] Figure 7 This is a schematic diagram of the device for extracting game materials from movies and TV shows provided by the present invention.

[0032] Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0034] Analysis of film and television works can be conducted on multiple levels, primarily including three aspects: content analysis, character extraction, and scene extraction. Content analysis can focus on video frame segmentation, shot detection, and subtitle / audio text extraction. Character extraction can focus on detecting and identifying main characters and their physical features. Scene extraction can utilize virtual space construction technology to construct 3D virtual scenes from 2D film and television scenes. Furthermore, in terms of game element extraction technology, interactive content generation technology can be used to extract multimodal elements such as scenes and characters from long videos and generate interactive content.

[0035] However, these methods are not suitable for comprehensively extracting and completing game elements from film and television works. Film and television content analysis technology excels at processing the temporal structure and basic extraction of videos, but its processing depth is limited to shot segmentation and simple text extraction, and does not involve the high-level semantic extraction (such as interactive script logic, character relationship evolution) and content completion (such as fictional side plots) required for game development. Film and television character extraction technology mainly relies on face detection and feature extraction, focusing on static appearance or identity recognition, and does not involve the extraction and completion of attributes such as character background stories, voice patterns, action sequences, facial expressions, and the generation of varied art materials (such as different ages and clothing). Film and television scene extraction technology focuses on generating interactive scenes or adjusting displays from video frames, and does not involve dynamic background image completion or audio-driven narrative interaction information extraction. These technologies can only solve specific single problems, stopping at forming single elements such as character appearance and scene images, and cannot form the complete game elements that cover story elements, character elements, and scene elements required for game development. Furthermore, the technology for extracting interactive content from videos only supports basic modal input or specific encapsulation, and is not suitable for film and television works that need to process video frame sequences, audio tracks, subtitle text and dynamic timing information simultaneously, and perform multimodal completion.

[0036] With the development of the digital entertainment industry, efficiently transforming existing films and television series into games is of great significance. For intellectual property (IP), adapting film and television works into games allows for in-depth exploration of IP value, expansion of interactive narrative forms, and promotion of film and television dissemination. For users and creators, it can reduce the cost and technical barriers of game creation, enabling non-professionals to create and express themselves based on their favorite films and television series, producing experiential interactive visual novels and promoting the rapid digitization and dissemination of creative content. However, current technology cannot automatically extract and complete structured game elements from films and television series; the process of adapting film and television works into games still heavily relies on manual labor, resulting in long production cycles and high costs.

[0037] To address the aforementioned issues, this application aims to achieve multimodal analysis of film and television works, and automatically extract game elements such as story, characters, and scenes from video frame sequences, audio, and subtitle information of the films and television works. It utilizes multimodal models to perform necessary completion, thereby improving the efficiency of game design and reducing the production cost and creative threshold of games.

[0038] The following is combined Figures 1 to 8 This invention describes a method, apparatus, device, medium, and product for extracting game assets from film and television. This application transforms video and audio-mixed film and television works into structured video sequence-multimodal data pairs, extracts game story, character, and scene elements from them, and uses a multimodal model for completion, thereby improving the efficiency of game element design.

[0039] Figure 1This is one of the flowcharts illustrating the method for extracting game materials from film and television provided by the present invention, such as... Figure 1 As shown, the method includes the following: Step 101: Perform multimodal analysis on the film and television videos to obtain structured multimodal data.

[0040] The multimodal data includes data information from multiple storyboard videos. The data information for each storyboard video includes at least one of the following: storyboard video sequence, storyboard keyframe sequence, storyboard audio clip, and storyboard text information.

[0041] Step 102: Based on the contextual relevance of the multimodal data, divide the multimodal data to obtain multiple game plot units.

[0042] Each game plot unit corresponds to at least one storyboard video.

[0043] Step 103: Extract game elements for each game plot unit.

[0044] The game elements include at least one of the following: story elements, character elements, and scene elements.

[0045] Step 104: Perform multimodal completion processing on the extracted game elements to obtain the completed game elements.

[0046] The supplemented game elements are used as game assets to generate game works that correspond to film and television videos.

[0047] In this application embodiment, video content includes, but is not limited to, movies, TV series, animations, web series, documentaries, stage play recordings, and other video content with a complete narrative structure.

[0048] In this application embodiment, the game elements that can be extracted and completed include, but are not limited to: story elements, character elements, scene elements, game rules, branching dialogue systems, etc. This application will use the extraction and completion of story elements, character elements, and scene elements as examples for illustrative description.

[0049] This invention provides a method, apparatus, device, medium, and product for extracting game materials from film and television. By performing multimodal analysis on film and television videos, structured multimodal data is obtained, which yields data information for multiple storyboard videos. Then, based on the contextual relevance of the multimodal data, the data is divided into multiple game plot units. For each game plot unit, game elements such as story elements, character elements, and scene elements are extracted. The extracted game elements are then subjected to multimodal completion processing to obtain supplemented game elements. These supplemented game elements can then be used as game materials to generate a game corresponding to the film and television video. Based on this, when adapting film and television works into game works, manual operations such as content analysis, character extraction, and scene extraction are unnecessary. This application efficiently extracts and supplements game elements such as story elements, character elements, and scene elements from film and television videos, thereby improving the efficiency of adapting film and television works into game works, shortening the game production cycle, and reducing production costs.

[0050] Figure 2 This is a second flowchart illustrating the method for extracting game materials from film and television provided by the present invention, as shown below. Figure 2 As shown, "performing multimodal analysis on film and television videos to obtain structured multimodal data" includes the following: Step 201: Perform shot boundary detection on the video and divide the video into multiple shot videos.

[0051] One storyboard video corresponds to one storyboard video sequence.

[0052] Step 202: Extract keyframes from each storyboard video to obtain the storyboard keyframe sequence corresponding to each storyboard video.

[0053] Step 203: Segment the audio trajectory corresponding to each storyboard video to obtain the storyboard audio segment corresponding to each storyboard video.

[0054] Step 204: Convert the audio segments corresponding to each storyboard video into audio text, and recognize the subtitles in each storyboard video to obtain subtitle text, thus obtaining the storyboard text information corresponding to each storyboard video.

[0055] The storyboard text information includes audio text and subtitle text.

[0056] Step 205: Based on the storyboard video sequence, storyboard keyframe sequence, storyboard audio clip, and storyboard text information corresponding to each storyboard video, obtain multimodal data.

[0057] It should be noted that film and television videos are usually presented in a mixed video and audio format, with text information typically appearing in subtitles, dialogue, or narration. To extract visual novel game elements from film and television videos, it is first necessary to separate the video frame sequence from the audio / text information, transforming the video into a structured video sequence-multimodal data pair sequence.

[0058] First, shot boundary detection is performed on the acquired video files, dividing the video into independent shot sequences (multiple shot videos). Shot boundary detection can employ traditional methods based on color histogram difference, brightness variance, edge feature differences, etc., or it can use deep learning-based shot boundary detection models (including but not limited to TransNet, TransNetV2, and models based on 3D convolutional neural networks). Next, the audio trajectory corresponding to each shot video is segmented to obtain the corresponding shot audio segment; a speech recognition model is used to convert the shot audio segment into audio text. Simultaneously, an OCR model is used to extract subtitle text from the shot videos, obtaining the shot text information corresponding to each shot video, and synchronizing the audio and video frames.

[0059] In one possible implementation, a multimodal understanding model can be used to further analyze the storyboard video sequence, storyboard keyframe sequence, storyboard audio clips, and storyboard text information. The input consists of the storyboard video sequence, storyboard keyframe sequence, storyboard audio clips, and storyboard text information corresponding to all storyboard videos. The multimodal understanding model first classifies the storyboard audio clips and storyboard text information into dialogue, narration, and sound effects.

[0060] Subsequently, the multimodal understanding model identifies the corresponding film and television characters for the dialogue text. Based on the character's position in the frame, lip-sync, audio waveform, and dialogue content, it establishes a correspondence between the dialogue and the film and television character. The output dialogue text format is "Character: Dialogue Content," the narration text format is "Narration: Narration Content," and the sound effect format is "Sound Effect: Description." Simultaneously, the multimodal understanding model generates descriptive text for the storyboard keyframe sequence. Finally, each storyboard video generates a corresponding data unit (i.e., the data information for each storyboard video), including the storyboard keyframe sequence, audio clips, and text. The storyboard videos are then sorted chronologically to form a video sequence-multimodal data pair sequence (i.e., multimodal data).

[0061] In one possible implementation, multimodal data is divided based on the contextual relevance of the multimodal data to obtain multiple game plot units. This includes: analyzing data from multiple storyboard videos, dividing the multiple storyboard videos at target points to obtain multiple game plot units. The target points include at least one of the following: video scene changes, topic changes, character appearance changes, plot conflict points, and emotional curve turning points.

[0062] In this embodiment of the application, the video sequence-multimodal data pair sequence can be segmented into logically independent game plot units, and each game plot unit is represented in the form of a video sequence-multimodal data pair subsequence.

[0063] The descriptive text, dialogue, narration text, and audio of the storyboard keyframe sequence are analyzed using a large language model. The sequence is segmented at points where the scene and topic change, forming a series of game plot units, which can be numbered sequentially.

[0064] In one possible implementation, game elements include story elements, which include at least one of the following: world-building elements, script elements, dialogue elements, character relationship elements, and emotional elements.

[0065] Figure 3 This is the third flowchart illustrating the method for extracting game materials from film and television provided by the present invention, as shown below. Figure 3 As shown, "Extracting game elements for each game plot unit" includes the following: Step 301: Determine the plot information of each game plot unit based on the data information of at least one storyboard video corresponding to each game plot unit.

[0066] The plot information includes at least one of the following: plot name, story summary, plot preconditions, plot post-plot results, plot dependencies, and world-building elements.

[0067] Step 302: Based on the storyboard text information of at least one storyboard video corresponding to each game plot unit, determine the script elements, dialogue elements, and character relationship elements.

[0068] Step 303: Perform sentiment analysis on the storyboard text information of at least one storyboard video corresponding to each game plot unit to determine the sentiment elements.

[0069] In one possible implementation, for each game plot unit, the main storyline of the game needs to be defined, and a plot name and story summary are generated for each game plot unit using a large language model.

[0070] For example, if game plot unit A is obtaining clues (pre-plot condition) and game plot unit B is solving a puzzle (post-plot result), then the relationship between game plot unit A and game plot unit B is that game plot unit B depends on game plot unit A.

[0071] Furthermore, based on the plot information of the game's plot units, the longest dependency path is identified as the main storyline. Plots on the main storyline that contain the core plot are considered main storylines, while plots on branching paths that do not involve the core plot are considered side stories. Simultaneously, a large language model is used to perform semantic analysis on the entire text to extract elements of the film's worldview.

[0072] In one possible implementation, for each game plot unit, it is also necessary to filter out all dialogue texts in the corresponding video sequence-multimodal data pair (storyboard video). Each dialogue text contains fields for the speaking character and the content of the speech, which can be used as script elements and dialogue elements.

[0073] In one possible implementation, for each game plot unit, a large language model can be used to extract the characters and their relationships from the text. Characters can be represented as a list, including character name, frequency of occurrence, and nicknames. Relationships can be represented as "Zhang San - Li Si: Allies". First, task prompts for extracting character and relationship entities are defined, providing multiple input text fragments and extracted character and relationship examples. Then, the large language model is called to perform information extraction.

[0074] In one possible implementation, for each game plot unit, a large language model can be used to perform sentiment analysis on all the text of that plot. By comprehensively understanding the vocabulary, tone of the dialogue, and context, the emotional tags expressed in the dialogue (such as happiness, sadness, anger, tension, calmness, etc.) can be determined, and the extracted sentiment tags can be added to the dialogue as sentiment elements.

[0075] In one possible implementation, game elements include character elements, which include at least one of the following: basic information and character characteristic information.

[0076] Figure 4 This is the fourth flowchart illustrating the method for extracting game materials from film and television provided by the present invention, as shown below. Figure 4 As shown, "Extracting game elements for each game plot unit" includes the following: Step 401: Based on the data information of at least one storyboard video corresponding to each game plot unit, determine the basic information of each character.

[0077] The basic information includes at least one of the following: character name, age, gender, personality, and backstory.

[0078] Step 402: Determine the character feature information of each character based on the keyframe sequence of at least one storyboard video corresponding to each game plot unit.

[0079] Among them, character feature information is used to ensure that the same character remains consistent in different storyboard videos. Character feature information includes at least one of the following: language features, action features, facial expression features, and appearance features.

[0080] In one possible implementation, game characters include player characters and non-player characters. All characters need to have character background information (basic information) and character dynamic design. Player characters also need to have character characteristic information, language characteristics, action characteristics, facial expression characteristics, appearance characteristics and other attributes. Non-player characters need to have their own unique role in the game narrative.

[0081] The extraction of character background information can be achieved using a named entity recognition model to identify character entities and resolve coreferences in the full text, identifying and associating mentions of the same character across different plot points. Semantic analysis is then performed on the text containing character descriptions to obtain the character background information.

[0082] The extraction of character dynamic design can input the keyframe sequence of the storyboard into the multimodal understanding model to directly analyze and extract the dynamic appearance features of each character (such as facial feature changes, clothing, hairstyle, color scheme, etc.) and generate appearance description text.

[0083] Furthermore, cross-shot character alignment is required. Before alignment, characters appear independently in different shots, leading to inconsistencies. Cross-shot character alignment aims to solve this problem by identifying the same character across different shots. A multimodal understanding model is used to extract local appearance features (such as facial features, clothing, and hairstyle) from character instance image sequences, converting them into appearance feature vectors. A vector database of character appearances is established. For each newly appearing character, the similarity between its feature vector and existing character feature vectors in the database is calculated. If the similarity is greater than a threshold, it is determined to be the same character; otherwise, a new character entry is created. After each successful match, the new features are fused to update the character's average feature vector.

[0084] In one possible implementation, for element extraction of the player character, a text style analysis model can be used to analyze all dialogue texts of the player character to extract its word usage habits (e.g., colloquialism, formality), sentence structure characteristics, high-frequency words, or catchphrases. For action features, a pose estimation model can be used to analyze all character instance image sequences of the player character to summarize the character's signature actions. For facial expression features, a facial expression recognition model can be used to summarize the character's typical expressions.

[0085] In one possible implementation, for the extraction of elements from non-player characters, a large language model can be used to analyze the dialogue text between non-player characters and player characters, and to classify the roles of non-player characters (including but not limited to quest givers, allies, guards, merchants, enemies, etc.).

[0086] In one possible implementation, game elements include scene elements, which include at least one of the following: scene identifier, background image, and scene attribute label. "Extracting game elements for each game plot unit" includes: analyzing the background of the keyframes based on a keyframe sequence from at least one storyboard video corresponding to each game plot unit to determine the scene elements.

[0087] Specifically, the keyframe sequence corresponding to each storyboard video can be input into the multimodal understanding model to directly analyze the background, identify visual elements of the scene (such as city streets, indoor environments, forests, etc.), and further annotate scene attribute labels, including day / night time and weather conditions. Finally, the main scenes are represented in the form of a scene list, including scene identifiers, feature vectors of the background image sequence representing the scene, and a set of labels (such as indoor environment, night scene, rainy day).

[0088] Figure 5 This is the fifth flowchart illustrating the method for extracting game materials from film and television provided by the present invention, as shown below. Figure 5 As shown, "performing multimodal completion on the extracted game elements to obtain the completed game elements" includes the following: Step 501: Analyze the story elements, determine the completeness parameters of the story elements, and generate supplementary story elements that satisfy the worldview elements based on the completeness parameters.

[0089] Step 502: Analyze the character elements and generate a variant image of each character's appearance.

[0090] Among them, appearance variant images include at least one of the following: expression variant images, action variant images, posture variant images, age variant images, clothing variant images, perspective variant images, lighting variant images, and style variant images.

[0091] Step 503: Analyze the scene elements and generate a background variant image for each scene.

[0092] The background variant image includes at least one of the following: weather variant image, seasonal variant image, time period variant image, state variant image, and line-of-sight variant image.

[0093] In one possible implementation, the extracted game elements can be supplemented to fill in missing parts of the original film or television work, ensuring element integrity, based on the characteristics of different game types. A multimodal generation model (such as a text-to-image or text-to-audio generation model) is used as input to generate the content.

[0094] Specifically, for story elements, the extracted story elements can be analyzed using a large language model to determine the completeness of the main and side plots. If there are logical gaps or undeveloped branches (such as background events not mentioned by the characters), then necessary fictional but world-building plots can be generated. For example, inputting "Based on the existing world-building, complete the side story of character A" will generate an interactive dialogue script.

[0095] For character elements, a multimodal generative model can be used to complete the missing appearance variations of the extracted character elements. For example, inputting the character description text "Generate different facial expression images of character B (happy, sad)", or "Generate different age stage images of character C (youth, middle age)", or "Generate different clothing images of character D (formal, casual)", the output will be the corresponding art asset image sequence.

[0096] For scene elements, a multimodal generative model can be used to complete the background variants from the extracted scene elements. For example, given the input scene description "Generate a rainy day variant image based on the existing nighttime indoor environment", the output is an enhanced background image that supports dynamic switching within the game.

[0097] In a complete example Figure 6 This is the sixth flowchart illustrating the method for extracting game materials from film and television provided by the present invention, as shown below. Figure 6 As shown, the entire process consists of four main steps, starting with the input of film and television works (including video, audio, and subtitles), going through multimodal parsing, plot segmentation, game element extraction, and multimodal completion, and finally outputting a structured element library or asset package.

[0098] Step 1: The multimodal parsing module first performs shot detection and segmentation. A complete film / video work (including video, audio, and subtitles) is input, and a shot boundary detection algorithm is used to divide the video into independent shot sequences. Keyframe extraction is then performed, extracting keyframe sequences from each shot sequence. These keyframes represent important visual information within the shot. An Automatic Speech Recognition (ASR) model is used to convert the audio into text, and Optical Character Recognition (OCR) technology is used to extract subtitle text. Finally, the keyframes, audio clips, and text information are synchronized in time to ensure they are aligned on the timeline. A multimodal understanding model is then used to analyze the synchronized data units, generating structured data containing keyframe sequences, audio clips, and text, outputting a structured semantic description of the shot.

[0099] Step 2: Using the plot segmentation module, the descriptive text, dialogue text, and audio information in the structured data are analyzed using a large language model. The data is then segmented at the target points to form multiple logically independent game plot units.

[0100] Step 3: Using the game element extraction module, a large language model is used to generate a plot name and story summary for each game plot unit. Based on the plot information of the game plot unit, the longest dependency path is identified as the main storyline. Dialogue text is then filtered from the main text to identify character entities and their relationships, and sentiment analysis is performed. The output includes story elements, character elements, and scene elements.

[0101] Step 4: Using the multimodal completion module, analyze the completeness of the main and side plots using a large language model to generate fictional plots or interactive dialogue scripts that conform to the world setting (story completion and scripting). Using the multimodal generation model, generate missing appearance variant images of the character based on the character description text, including different expressions, different age groups, or different clothing (character portrait variant completion). Using the multimodal generation model, generate background variant images of the scene based on the scene description (scene background variant completion).

[0102] Ultimately, a structured output (element library / asset package) is obtained, which includes the completed story elements, character elements, and scene elements, forming a structured element library or asset package that can be used by game developers.

[0103] The apparatus for extracting game materials from movies and TV shows provided by the present invention is described below. The apparatus for extracting game materials from movies and TV shows described below can be referred to in correspondence with the method for extracting game materials from movies and TV shows described above.

[0104] Figure 7 This is a schematic diagram of the device for extracting game materials from movies and television provided by the present invention, as shown below. Figure 7 As shown, the device for extracting game materials from film and television includes the following modules: parsing module 701, segmentation module 702, extraction module 703, and supplementation module 704; The parsing module is used to perform multimodal parsing on film and television videos to obtain structured multimodal data. The multimodal data includes data information from multiple storyboard videos. The data information of each storyboard video includes at least one of the following: storyboard video sequence, storyboard keyframe sequence, storyboard audio clip, and storyboard text information. The segmentation module is used to segment multimodal data based on the contextual relevance of the multimodal data to obtain multiple game plot units, each of which corresponds to at least one storyboard video. The extraction module is used to extract game elements for each game plot unit. Game elements include at least one of the following: story elements, character elements, and scene elements. The supplementary module is used to perform multimodal completion processing on the extracted game elements to obtain supplemented game elements. These supplemented game elements are then used as game assets to generate game works that correspond to film and video content.

[0105] According to the present invention, an apparatus for extracting game materials from film and television is provided, and the parsing module is specifically used for: Perform shot boundary detection on film and television videos, and divide the film and television videos into multiple shot videos, with each shot video corresponding to a shot video sequence; Keyframes are extracted from each storyboard video to obtain the storyboard keyframe sequence corresponding to each storyboard video; The audio trajectory corresponding to each storyboard video is segmented to obtain the storyboard audio segment corresponding to each storyboard video. The audio segments corresponding to each storyboard video are converted into audio text, and the subtitles in each storyboard video are recognized to obtain subtitle text, thus obtaining the storyboard text information corresponding to each storyboard video. The storyboard text information includes audio text and subtitle text. Multimodal data is obtained based on the storyboard video sequence, storyboard keyframe sequence, storyboard audio clip, and storyboard text information corresponding to each storyboard video.

[0106] According to the present invention, an apparatus for extracting game materials from film and television is provided, comprising modules specifically used for: The data from multiple storyboard videos are analyzed, and the multiple storyboard videos are divided at target points to obtain multiple game plot units. The target points include at least one of the following: video scene changes, topic changes, character appearance changes, plot conflict points, and emotional curve turning points.

[0107] According to the present invention, an apparatus for extracting game materials from film and television is provided. The game elements include story elements, which in turn include at least one of the following: world-building elements, script elements, dialogue elements, character relationship elements, and emotional elements. The extraction module is specifically used for: Based on the data information of at least one storyboard video corresponding to each game plot unit, determine the plot information of each game plot unit. The plot information includes at least one of the following: plot name, story summary, plot preconditions, plot post-plot results, plot dependencies, and world view elements. Based on the storyboard text information of at least one storyboard video corresponding to each game plot unit, determine script elements, dialogue elements, and character relationship elements; Sentiment analysis was performed on the storyboard text information of at least one storyboard video corresponding to each game plot unit to identify emotional elements.

[0108] According to the present invention, an apparatus for extracting game materials from film and television is provided. The game elements include character elements, and the character elements include at least one of the following: basic information and character feature information; the extraction module is specifically used for: Based on the data information of at least one storyboard video corresponding to each game plot unit, determine the basic information of each character. The basic information includes at least one of the following: character name, age, gender, personality, and backstory. Based on the keyframe sequence of at least one storyboard video corresponding to each game plot unit, the character feature information of each character is determined. The character feature information is used to ensure that the same character remains consistent in different storyboard videos. The character feature information includes at least one of the following: language features, action features, facial expression features, and appearance features.

[0109] According to the present invention, an apparatus for extracting game materials from film and television is provided. The game elements include scene elements, and the scene elements include at least one of the following: scene identifier, background image, and scene attribute label; the extraction module is specifically used for: Based on the keyframe sequence of at least one storyboard video corresponding to each game plot unit, the background of the keyframes is analyzed to determine scene elements.

[0110] According to the present invention, an apparatus for extracting game materials from film and television is provided, and a supplementary module is specifically used for: The story elements are analyzed to determine the completeness parameters of the story elements, and supplementary story elements that satisfy the worldview elements are generated based on the completeness parameters. Analyze the character elements and generate appearance variant images for each character. The appearance variant images include at least one of the following: expression variant image, action variant image, pose variant image, age variant image, clothing variant image, perspective variant image, lighting variant image, and style variant image. Analyze scene elements and generate background variant images for each scene. The background variant images include at least one of the following: weather variant image, season variant image, time period variant image, state variant image, and view distance variant image.

[0111] Figure 8 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 8As shown, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communications bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other through the communications bus 840. The processor 810 can call logic instructions in the memory 830 to execute a method for extracting game assets from film and television. This method includes: performing multimodal analysis on the film and television video to obtain structured multimodal data, the multimodal data including data information from multiple storyboard videos, each storyboard video including at least one of the following: storyboard video sequence, storyboard keyframe sequence, storyboard audio clip, and storyboard text information; dividing the multimodal data based on the contextual relevance of the multimodal data to obtain multiple game plot units, each game plot unit corresponding to at least one storyboard video; extracting game elements for each game plot unit, game elements including at least one of the following: story elements, character elements, and scene elements; performing multimodal completion processing on the extracted game elements to obtain supplemented game elements, the supplemented game elements being used as game assets to generate a game corresponding to the film and television video.

[0112] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0113] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the methods provided above for extracting game materials from film and television. The method includes: performing multimodal analysis on the film and television video to obtain structured multimodal data, the multimodal data including data information of multiple storyboard videos, each storyboard video including at least one of the following: storyboard video sequence, storyboard keyframe sequence, storyboard audio clip, and storyboard text information; dividing the multimodal data based on the contextual relevance of the multimodal data to obtain multiple game plot units, each game plot unit corresponding to at least one storyboard video; extracting game elements for each game plot unit, the game elements including at least one of the following: story elements, character elements, and scene elements; performing multimodal completion processing on the extracted game elements to obtain supplemented game elements, the supplemented game elements being used as game materials to generate a game corresponding to the film and television video.

[0114] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a method for extracting game materials from film and television provided by the methods described above. This method includes: performing multimodal analysis on the film and television video to obtain structured multimodal data, the multimodal data including data information of multiple storyboard videos, each storyboard video including at least one of the following: storyboard video sequence, storyboard keyframe sequence, storyboard audio clip, and storyboard text information; dividing the multimodal data based on the contextual relevance of the multimodal data to obtain multiple game plot units, each game plot unit corresponding to at least one storyboard video; extracting game elements for each game plot unit, the game elements including at least one of the following: story elements, character elements, and scene elements; performing multimodal completion processing on the extracted game elements to obtain supplemented game elements, the supplemented game elements being used as game materials to generate a game corresponding to the film and television video.

[0115] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0116] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for extracting game materials from film and television, characterized in that, include: Multimodal analysis is performed on film and television videos to obtain structured multimodal data. The multimodal data includes data information of multiple storyboard videos. The data information of each storyboard video includes at least one of the following: storyboard video sequence, storyboard keyframe sequence, storyboard audio clip, and storyboard text information. Based on the contextual relevance of the multimodal data, the multimodal data is divided into multiple game plot units, and each game plot unit corresponds to at least one storyboard video. For each game plot unit, game elements are extracted, and the game elements include at least one of the following: story elements, character elements, and scene elements; The extracted game elements are subjected to multimodal completion processing to obtain supplemented game elements, which are used as game materials to generate game works corresponding to the film and television videos.

2. The method of claim 1, wherein, The process of performing multimodal analysis on film and television videos to obtain structured multimodal data includes: Shot boundary detection is performed on the video footage to divide it into multiple split-scene videos, with each split-scene video corresponding to a split-scene video sequence. Keyframes are extracted from each of the storyboard videos to obtain the storyboard keyframe sequence corresponding to each storyboard video; The audio trajectory corresponding to each of the storyboard videos is segmented to obtain the storyboard audio segment corresponding to each of the storyboard videos; The audio segment corresponding to each storyboard video is converted into audio text, and the subtitles in each storyboard video are identified to obtain subtitle text, thereby obtaining the storyboard text information corresponding to each storyboard video. The storyboard text information includes the audio text and the subtitle text. The multimodal data is obtained based on the storyboard video sequence, the storyboard keyframe sequence, the storyboard audio clip, and the storyboard text information corresponding to each storyboard video.

3. The method of claim 1, wherein, Based on the contextual relevance of the multimodal data, the multimodal data is divided into multiple game plot units, including: The data of the multiple storyboard videos are analyzed, and the multiple storyboard videos are divided at target points to obtain the multiple game plot units. The target points include at least one of the following: video scene changes, topic changes, character appearance changes, plot conflict points, and emotional curve turning points.

4. The method of claim 1-3, wherein, The game elements include story elements, which include at least one of the following: world-building elements, script elements, dialogue elements, character relationship elements, and emotional elements. For each game plot unit, the extraction of game elements includes: Based on the data information of at least one storyboard video corresponding to each game plot unit, the plot information of each game plot unit is determined, and the plot information includes at least one of the following: plot name, story summary, plot preconditions, plot post-plot results, plot dependencies, and the world view elements; Based on the storyboard text information of at least one storyboard video corresponding to each game plot unit, the script elements, the dialogue elements, and the character relationship elements are determined; Sentiment analysis is performed on the storyboard text information of at least one storyboard video corresponding to each game plot unit to determine the emotional elements.

5. The method of claim 1-3, wherein, The game elements include character elements, and the character elements include at least one of the following: basic information and character characteristic information; For each game plot unit, the extraction of game elements includes: Based on the data information of at least one storyboard video corresponding to each game plot unit, the basic information of each character is determined, and the basic information includes at least one of the following: character name, age, gender, personality, and background story; Based on the keyframe sequence of at least one storyboard video corresponding to each game plot unit, character feature information for each character is determined. The character feature information is used to ensure that the same character remains consistent in different storyboard videos. The character feature information includes at least one of the following: language features, action features, facial expression features, and appearance features.

6. The method of claim 1-3, wherein, The game elements include scene elements, which include at least one of the following: scene identifier, background image, and scene attribute label; For each game plot unit, the extraction of game elements includes: Based on the keyframe sequence of at least one storyboard video corresponding to each game plot unit, the background of the keyframes is analyzed to determine the scene elements.

7. The method of claim 1-3, wherein, The process of performing multimodal completion on the extracted game elements to obtain the completed game elements includes: The story elements are analyzed to determine the completeness parameters of the story elements, and supplementary story elements that satisfy the worldview elements are generated based on the completeness parameters. The character elements are analyzed to generate appearance variant images for each character. The appearance variant images include at least one of the following: expression variant image, action variant image, posture variant image, age variant image, clothing variant image, perspective variant image, lighting variant image, and style variant image. The scene elements are analyzed to generate a background variant image for each scene. The background variant image includes at least one of the following: weather variant image, season variant image, time period variant image, state variant image, and view distance variant image.

8. An apparatus for extracting game material from a movie, characterized by comprising: include: The module includes a parsing module, a partitioning module, an extraction module, and a supplementary module. The parsing module is used to perform multimodal parsing on film and television videos to obtain structured multimodal data. The multimodal data includes data information of multiple storyboard videos. The data information of each storyboard video includes at least one of the following: storyboard video sequence, storyboard keyframe sequence, storyboard audio clip, and storyboard text information. The segmentation module is used to segment the multimodal data based on the contextual relevance of the multimodal data to obtain multiple game plot units, and each game plot unit corresponds to at least one storyboard video. The extraction module is used to extract game elements for each game plot unit, and the game elements include at least one of the following: story elements, character elements, and scene elements; The supplementary module is used to perform multimodal completion processing on the extracted game elements to obtain supplemented game elements, which are used as game materials to generate game works corresponding to the film and television videos.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the method for extracting game materials from movies and television programs as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for extracting game materials from movies and television as described in any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for extracting game materials from movies and television as described in any one of claims 1 to 7.