A method for generating a graphic work, related devices, equipment and storage medium
By acquiring and retrieving original text information, and utilizing large language models and image generation models, graphic and text works are automatically generated. This solves the problems of long creation cycles and low efficiency in traditional graphic and text creation, and achieves efficient and rich graphic and text work generation, thereby improving creation efficiency and content quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2025-01-10
- Publication Date
- 2026-07-14
AI Technical Summary
Current technologies rely on traditional methods for creating graphic works, resulting in long creation cycles, low efficiency, and a lack of systematic methods for integrating multiple elements.
By acquiring and retrieving original text information, and utilizing large language models and image generation models, we can automatically generate image-text sequences, integrate multiple elements, and produce image-text works that are both logical and literary.
It significantly reduces the time cost of manual creation, improves creation efficiency, and generates richer and higher-quality content. It enhances narrative ability and artistic appeal, thereby improving creation efficiency, artistic appeal, and content quality.
Smart Images

Figure CN122391396A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, related apparatus, device, and storage medium for generating graphic works. Background Technology
[0002] As a creative form that integrates visual elements (such as paintings and images) with text to convey rich literary meaning and artistic appeal, illustrated works, especially picture books, have received much attention in the field of reading.
[0003] Currently, the content generation of different types of graphic works often relies on their respective traditional creative methods, lacking a systematic approach that integrates multiple elements and technologies. For example, picture book creation often focuses on combining hand-drawn illustrations with simple text creation; the creative process is relatively independent, the creation cycle is long, and the creation efficiency is low. Therefore, an effective method is urgently needed to address these issues. Summary of the Invention
[0004] This application provides a method, related apparatus, equipment, and storage medium for generating graphic works, which solves the problems of long creation cycles and low creation efficiency caused by manual creation of graphic works in the prior art.
[0005] This application provides a method for generating graphic works, including:
[0006] The process involves acquiring raw text information and retrieving text information. The raw text information is used to represent content elements, while the retrieved text information is information related to the content elements obtained based on the raw text information.
[0007] Based on the original text information and the retrieved text information, generate the corresponding target text material;
[0008] Generate target text information corresponding to the target text material based on the large language model;
[0009] An image-text sequence corresponding to target text information is generated according to an image generation model. The image-text sequence includes M image-text information items, where the representation information of M image content items corresponds to the content of the target text information. The M image-text information items include text content, which is a portion of the target text information. M is an integer greater than or equal to 1. Another aspect of this application provides an apparatus for generating image-text works, comprising:
[0010] The acquisition module is used to acquire raw text information and retrieve text information. The raw text information is used to represent content elements, and the retrieved text information is information related to the content elements obtained based on the raw text information.
[0011] The generation module is used to generate corresponding target text materials based on the original text information and the retrieved text information;
[0012] The generation module is also used to generate target text information corresponding to the target text material based on the large language model;
[0013] The generation module is also used to generate a graphic-text sequence corresponding to the target text information based on the image generation model. The graphic-text sequence includes M graphic-text information, the representation information of the M screen contents in the M graphic-text information corresponds to the content of the target text information, and the M graphic-text information includes text content, which is a part of the target text information, and M is an integer greater than or equal to 1.
[0014] In another aspect, this application provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the methods described above.
[0015] Another aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the methods described above.
[0016] Another aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the methods described above.
[0017] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:
[0018] This application provides a method for generating graphic and textual works. By acquiring original text information and retrieving text information, it can integrate information related to the content elements of graphic and textual works from different channels and types, avoiding the limitations of a single information source and making the generated content richer and more comprehensive. Taking picture book creation as an example, the traditional method of combining hand-drawn illustrations with simple text creation is time-consuming. However, this application utilizes artificial intelligence technology to automatically perform information retrieval, text generation, and image generation, greatly reducing the time cost of manual creation. Especially in projects involving a large number of graphic and textual works or those with timeliness requirements, it can quickly respond to needs and produce works in a timely manner. The large language model, trained on a large amount of text data, can generate target text information that is logical, coherent, and literary. Compared with simple hand-drawn text, its language expression is more accurate and vivid, better able to portray characters, tell plots, and depict scenes, enhancing the narrative ability and artistic appeal of graphic and textual works. In summary, the graphic and textual work generation method of this application, by integrating multiple elements and technologies, demonstrates significant advantages in terms of creation efficiency and content quality. Attached Figure Description
[0019] Figure 1A schematic diagram illustrating the application of the graphic and textual work generation method provided in this application embodiment to a game picture book scene;
[0020] Figure 2 A schematic diagram of the graphic works provided in the embodiments of this application;
[0021] Figure 3 The method for generating graphic works provided in this application is applied to a game picture book generation scene architecture diagram;
[0022] Figure 4 A schematic diagram of client-side reporting provided in an embodiment of this application;
[0023] Figure 5 A schematic diagram illustrating the interaction between the client and server provided in an embodiment of this application;
[0024] Figure 6 A schematic diagram of the battlefield terrain story of game A provided in this application embodiment;
[0025] Figure 7 A schematic diagram of the lobby page of a certain version of Game A provided in this application embodiment;
[0026] Figure 8 A flowchart illustrating the method for generating graphic works provided in this application embodiment;
[0027] Figure 9 A schematic diagram illustrating a method for generating graphic works provided in this application embodiment;
[0028] Figure 10 A schematic diagram illustrating the model training process provided in the embodiments of this application;
[0029] Figure 11 A schematic diagram illustrating the processing procedure of the large language model provided in the embodiments of this application;
[0030] Figure 12 A schematic diagram illustrating the processing procedure of the large language model provided in the embodiments of this application;
[0031] Figure 13 A schematic diagram illustrating the generation of image and text sequences provided in an embodiment of this application;
[0032] Figure 14 A structural diagram of the graphic and textual work generation device provided in the embodiments of this application. Detailed Implementation
[0033] This application provides a method for generating graphic and textual works. By using a large language model and an image generation model, graphic and textual works can be generated, effectively improving the efficiency of graphic and textual work creation and enhancing content quality.
[0034] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0035] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0036] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0037] Before introducing the specific methods of this application, we will first provide an illustrative example of the application scenarios for generating graphic works in this application. It should be understood that the following application scenarios are merely examples and are not limited to these.
[0038] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating the application of the graphic and textual work generation method provided in this application to a game picture book scene. A user wants to generate a picture book for hero character 'a' in game A. The user... Figure 1 The interface diagram of the large language model shown in (A) allows users to input "the story of hero character a in game A" through the information input window 111 and click the send control 112 to send the input information to the server.
[0039] The server extracts the original text information "Hero character a in game A" from the user's input. Based on this original text information, the server performs a search to obtain related search text information: "Detailed introduction of hero character a's skills in game A and their connection to the background story. Details of hero character a's growth process in game A, such as specific training scenarios during their growth in the Mercenary Guild, and other significant events they experienced. Information about the Rose Earl family, including their status, sphere of influence, and relationships with other forces in game A, to enrich the story background. Detailed information about the Fallen Legion, such as the legion's composition, leader information, and combat style." The specific details of defeating Adorn and the Knights Templar enhance the story's plausibility and tension. Game A's settings and rules related to mercenary missions, such as mission issuance mechanisms, contract terms, and reward distribution methods, make the story more consistent with the game's worldview. Information about the Rose Knights, including their composition, combat abilities, and influence in Game A, allows for better depiction of battle scenes. The backstories and skill characteristics of other similar hero characters in Game A are used to contrast and highlight the uniqueness of hero character A. Game A's settings regarding the impact of battlefield environment and terrain on combat make the battle descriptions more realistic.
[0040] Based on the original text information and the retrieved text information, the server generates the corresponding target text material: "Hero character A in Game A is a codenamed mercenary trained by the Mercenary Guild. An orphan, she learned combat skills and served the guild to pay off debts, performing various missions, such as disguising herself as a maid to protect others. While protecting Count Rose's youngest daughter, Adriss, they encountered a powerful Fallen Legion. Adriss led the Rose Knights into battle, and Hero character A upheld her duty, experiencing fierce battles alongside her. Her skills are consistent with her background, allowing her to quickly enter the battlefield and assassinate key targets. Detailed skill characteristics, such as rapid movement and high burst damage, how to utilize these skills in combat, and the connection between these skills and her mercenary identity and experiences. Background information related to the Count Rose family, the Rose Knights, and the Fallen Legion, and the roles and relationships of these elements in the story. Mercenary mission-related settings in Game A, such as mission issuance, contracts, and rewards, and how they affect Hero character A's behavior and decisions. Comparative information with other similar heroes, highlighting Hero character A's characteristics, and the impact of battlefield environment settings on combat, making the entire story more complete and richer."
[0041] like Figure 1As shown in (B), the server generates target text information corresponding to the target text material based on the large language model. 121 "Hero character A is a hero character in game A. Her backstory depicts the image of a mercenary. According to the description in xx encyclopedia, Hero character A is not her real name, but a code name given to her by the Mercenary Guild. She is a special mercenary trained by the Mercenary Guild, a training method adopted by the Mercenary Guild to meet the various specific requirements of employers for mercenaries. Hero character A was raised by the guild as an orphan and was taught combat skills. After adulthood, she served the guild as a contract mercenary and took a portion of the mission rewards to repay the guild's investment and interest until the debt was paid off. She has performed various missions, including impersonating a noble maid to provide protection. In one mission..." In the mission, Hero Character A is tasked with protecting Count Rose's youngest daughter, Adrienne. What initially seemed like a simple task becomes exceptionally difficult as they face the Fallen Legion, who have defeated Adorn and the Knights Templar. Adrienne bravely leads the Knights Rose into battle, with Hero Character A standing by her side, experiencing the fierce fighting alongside her. Hero Character A's principles are unwavering; the word "escape" doesn't exist in her dictionary. Even with the escalating difficulty of the mission, she doesn't choose to cancel, but instead upholds her duties and principles, bravely facing all challenges. This story showcases Hero Character A's ruthlessness and efficiency as a mercenary. In the game, Hero Character A's skill design aligns with her backstory; she possesses rapid movement and high-damage output, enabling her to quickly infiltrate enemy lines and precisely assassinate key targets.
[0042] like Figure 1 As shown in (C), the user can click the picture book generation control 131, and the server will generate the corresponding image and text sequence based on the target text information generated by the large language model. For example... Figure 1 As shown in (D), the M graphic and textual information pieces constitute the game picture book.
[0043] Preferred, such as Figure 2 As shown, the M pieces of graphic and textual information form a game picture book that can be published as a book.
[0044] Please see Figure 3 , Figure 3 The diagram illustrates the architecture of a game picture book generation scenario where the graphic and textual artwork generation method provided in this application is applied. This scenario includes a terminal 310, a network 320, a server 330, and a database 340.
[0045] Terminal 310 includes a human-computer interaction screen, a processor, and a memory. The human-computer interaction screen is used to display a chat window of a large language model; it also provides a human-computer interaction interface to receive information entered by the user in the information input window. The processor is used to generate interaction instructions in response to the above human-computer interaction operations and send the interaction instructions to the server. The memory is used to store relevant attribute data.
[0046] The terminals 310 involved in this application include, but are not limited to, mobile phones, tablets, laptops, desktop computers, smart voice interaction devices, virtual reality devices, smart home appliances, vehicle terminals, aircraft, etc.
[0047] Run client 3101 in terminal 310. Taking the large language model client as an example, client 3101 is deployed on terminal 310. Client 3101 can run on terminal 310 through a browser, or through a standalone application (APP) or a mini-program, etc.
[0048] Network 320 uses standard communication technologies and / or protocols, typically the Internet, but can also be any network, including but not limited to Bluetooth, local area network (LAN), metropolitan area network (MAN), wide area network (WAN), mobile, private network, or any combination of virtual private network. In some embodiments, custom or dedicated data communication technologies may be used to replace or supplement the aforementioned data communication technologies.
[0049] Server 330 includes a processor. The server 330 involved in this application can be a standalone physical server, a server cluster or distributed system consisting of at least one physical server, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence (AI) platforms.
[0050] Database 340 is used to store information related to game A, including attribute information of all characters in game A, storyline, scene map, etc.
[0051] The user enters "the story of hero character a in game A" in the information input window of the terminal and clicks the send control.
[0052] As in step S301, terminal 310 receives user input information and generates an information retrieval instruction based on the input information.
[0053] As in step S302, terminal 310 sends the information acquisition instruction carrying the input information to server 330 via network 320.
[0054] As in step S303, server 330 receives information acquisition instruction, parses the information acquisition instruction, and obtains the original text information "hero character a in game A" from the information acquisition instruction.
[0055] As in step S304, server 330 retrieves the searched text information from database 340 based on the original text information.
[0056] As in step S305, server 330 generates corresponding target text material based on the original text information and the retrieved text information.
[0057] As in step S306, server 330 generates target text information corresponding to the target text material based on the large language model.
[0058] As in step S307, the server 330 sends the generated target text information to the terminal 310 via the network 320.
[0059] As in step S308, terminal 310 displays the target text information.
[0060] When a user clicks the picture book generation control in the terminal, as in step S309, the terminal 310 receives the user's click action and generates a picture book generation instruction.
[0061] As in step S310, terminal 310 sends the picture book generation instruction to server 330 via network 320.
[0062] As in step S311, server 330 receives the picture book generation instruction and generates a picture-text sequence corresponding to the target text information according to the image generation model. The picture-text sequence includes M picture-text information, the representation information of the M picture content in the M picture-text information corresponds to the content of the target text information, and the M picture-text information includes text content, the text content is part of the content in the target text information, and M is an integer greater than or equal to 1.
[0063] As in step S312, server 330 sends the image and text sequence to terminal 310 via network 320.
[0064] As in step S313, terminal 310 receives the image and text sequence and displays it.
[0065] In the context of game picture book creation, the method provided in this application brings comprehensive improvement to the generation of game picture books, showing significant advantages in terms of content quality, creation efficiency, artistic appeal and creative inspiration. It helps to promote the development of the field of game picture book creation and provides a better content carrier for the dissemination of game culture and player experience.
[0066] Please see Figure 4 , Figure 4 This diagram illustrates a client-side reporting method provided in an embodiment of this application. Game A is a mobile game; from a hardware perspective, Game A is an app running on a mobile phone, occupying approximately 1GB of storage and requiring a network speed of 2GHz or higher to function. The method for generating graphic works described herein is deployed on a server-side platform. After the mobile game app clicks "match" on the client side, it requests a response from the cloud server.
[0067] Please see Figure 5 , Figure 5 This diagram illustrates the client-server interaction provided in an embodiment of this application. The server is responsible for forwarding data from the gateway using protocols. The lobby server (gamesvr) is the game login page and store, etc. If players from different regions (or even different countries) are matching together, the server can ensure the lobby server is shared. Load management (loadsvr) minimizes the load on different servers. If there are additional requirements (such as various Asian Games versions, requiring the smallest isolated servers), there is a relay server, which is strongly related to the game; each region will have its own relay server. Finally, data is transmitted via protocols.
[0068] The combat suit refers to in-game data transmitted via a protocol. From the map and heroes to minions, monsters, and even the dragon, the combat suit describes elements of the game. These can be explained to a child from a world-building perspective. Please refer to [link / reference]. Figure 6 , Figure 6 A schematic diagram of the battlefield terrain story of game A provided in this application embodiment is shown.
[0069] After logging in, players will see a startup animation and a video button in the upper right corner of the lobby. Please refer to [link / reference]. Figure 7 , Figure 7 The image shows the lobby page of a certain version of game A. Each startup animation tells a story, involving the background, origin, and quests of each hero. These can all be expressed through picture books and illustrations, especially some of the cartoon characters in the game, such as Alice, who have won the hearts of gamers with their cute appearance and voice.
[0070] Please see Figure 8 , Figure 8A flowchart illustrating a method for generating graphic and textual works is provided. It should be noted that the method for generating graphic and textual works provided in this application embodiment can be applied to a server, and this application embodiment does not impose any limitations. The method includes:
[0071] S810, Obtaining raw text information and retrieving text information.
[0072] The original text information is used to represent content elements, while the retrieved text information is information related to the content elements obtained based on the original text information.
[0073] Understandably, raw text information is primarily used to represent content elements, and its sources are quite diverse. In the example of game-based storybook generation, user input is a common source. For instance, a user might enter "the story of hero character a in game A" into the information input window of a terminal device (such as a large language model client interface running on a mobile phone or tablet), which becomes the starting point for raw text information. Besides this, it may also come from system recommendations. For example, the game system might push key content (such as popular game characters, classic game scenes, etc.) to the user based on popular elements, past browsing or usage records, etc. Additionally, it can be extracted from external text materials, such as summaries of game-related content on game strategy websites, or interesting topics shared by players on game forums. Key expressions related to game elements can be filtered from these text data and used as raw text information.
[0074] Retrieving text information involves further mining content elements from the original text, making its relevance crucial. After receiving the original text information on the server side, search engine technology is used to perform the retrieval. If retrieving from a local database (such as one storing extensive information about game A, including character attributes, storylines, and map details), appropriate database queries are written to filter relevant information from numerous documents and records. For example, it might search for detailed skill descriptions for hero character A, details of their in-game progression, and other related game elements. Online searches might involve interacting with the game's official API to obtain authoritative and up-to-date information, or scraping high-quality user-shared content from authoritative game communities and forums as retrieval text information. The goal is to comprehensively collect all kinds of detailed information closely related to the content elements pointed to by the original text information. Retrieving text information greatly enriches the material base of text and image works. It allows the subsequently generated target text materials to go beyond the simple content covered by the original text information, enabling the expansion and deepening of core content elements from multiple dimensions.
[0075] Content elements can encompass many aspects, not just characters, but also unique scenes, special items, or specific plot clues within a game. The original text information is used to extract these elements through precise vocabulary or phrases, becoming the foundation for subsequent creation. For example, when using a scene as a content element, the original text information "the xxxx scene in game A" will guide subsequent exploration of information about the terrain, creatures, and stories associated with that xxxx scene to enrich the content of the graphic and textual work.
[0076] This step, through reasonable information collection and expansion, provides a rich and valuable information source for the entire graphic and text creation process, ensuring that subsequent stages can be carried out in an orderly and high-quality manner, thereby helping to generate graphic and text works that are rich in content, logically rigorous, and attractive.
[0077] S820. Based on the original text information and the retrieved text information, generate the corresponding target text material.
[0078] Understandably, the acquired raw text information and retrieved text information are integrated. For example, the raw text information "hero character a in game A" and the retrieved text information about the character's skills, background story, etc., are combined according to a preset text structure template (such as a prompt template). The preset template could be "[hero character name], which has [skill characteristics], and the background story is [detailed background]".
[0079] Optimize the generated target text materials. Check the logical coherence of the text materials to ensure that information in different parts flows naturally. For example, when describing character skills and backstory, make the transitions between them smooth and natural, avoiding abrupt shifts. Polish the language in the text materials to make them more accurate and vivid. For example, transform some technical terms into more colloquial expressions (if applicable), or add descriptive vocabulary to enrich character portrayal and plot content.
[0080] S830. Generate target text information corresponding to the target text material based on the large language model.
[0081] Understandably, it's important to choose a suitable large language model, such as GPT-4. The model parameters should be set according to actual needs, such as setting the maximum length of the generated text and a temperature parameter (to control the randomness of the generated text; higher temperatures result in greater randomness). If more deterministic and accurate text is desired, the temperature parameter can be set lower, such as around 0.5; if more creative variations are needed, the temperature parameter can be increased appropriately, such as to 0.8. The target text material should then be input into the large language model.
[0082] The model understands and analyzes the input material based on its pre-trained knowledge and algorithms. Taking the target text material "Hero character a in game A, whose skills include..." as an example, the large language model will generate more detailed, vivid and logical target text information based on its understanding of concepts such as game characters and skills, as well as its grasp of language structure and logic.
[0083] S840. Generate the image-text sequence corresponding to the target text information based on the image generation model.
[0084] The image-text sequence includes M image-text information items. The representation information of the M image-text items corresponds to the content of the target text information. The M image-text items include text content, which is a part of the target text information. M is an integer greater than or equal to 1.
[0085] Understandably, the target text information is first preprocessed, such as through word segmentation to identify keywords like nouns, verbs, and adjectives, in order to better understand the text content. Then, based on semantics and logic, the target text information is divided into M target text paragraphs. For example, taking the generated target text information about hero character 'a' as an example, it can be divided into paragraphs such as "Background introduction of hero character 'a'", "Skill characteristics of hero character 'a'", and "An important mission experience of hero character 'a'".
[0086] Each target text paragraph is input into an image generation model, such as DALL-E3. The image generation model generates corresponding image content based on the content of the text paragraph. For the paragraph "Background introduction of hero character A", an image might be generated showing hero character A training hard in the mercenary guild, showing her resolute expression and the harsh training environment around her; for the paragraph "Skill characteristics of hero character A", the generated image might highlight the cool special effects and powerful force when she releases her skills.
[0087] While generating the image, some key text content from the target text paragraph is extracted as text content in the image-text sequence and placed in appropriate positions on the image. For example, a short text description is added below or next to the image, such as "Hero character A grows up in the mercenary guild and learns various combat skills" or "Hero character A releases a powerful first skill, causing damage to the enemy," to ensure close matching and mutual complementarity between the image and text.
[0088] For example, suppose a user wants to generate a picture book about a scene in game A (such as the Mysterious Forest scene). First, the original text information "The Mysterious Forest scene in game A" is obtained. Then, a search is performed to obtain searched text information, such as the special terrain of the Mysterious Forest, possible monsters, and game plot related to the scene.
[0089] Based on this information, target text materials are generated, such as "The mysterious forest scene in game A, where there are hidden caves and huge ancient trees, where you may encounter forest spirits and ferocious beasts, and players have the task of finding mysterious treasures."
[0090] Input the target text material into a large language model to generate target text information, such as: "In the mysterious forest of Game A, huge and ancient trees block out the sun, with sunlight only filtering through layers of leaves to cast dappled shadows. Caves hidden deep in the forest exude a mysterious aura, as if concealing countless unknown secrets. Spirited forest spirits often roam the forest, possessing the ability to manipulate the forces of nature and using magic to protect it; while ferocious beasts lurk in the shadows, ready to attack at any moment. Players enter this forest, shouldering the heavy responsibility of finding mysterious treasures. They must explore carefully, solve the forest's puzzles, and avoid dangers in order to find the legendary treasure and begin a new chapter of adventure."
[0091] Next, the target text information is preprocessed and divided into paragraphs, such as "environmental description of the mysterious forest", "elves and beasts in the forest", and "player's treasure hunt quest".
[0092] Finally, each paragraph is input into the image generation model to generate a text-image sequence. For the "Environmental Description of the Mysterious Forest" paragraph, an image is generated showing the lush scenery of the mysterious forest, ancient trees, and dappled sunlight, with the text "In the mysterious forest, ancient trees block out the sun, and sunlight is dappled" added below the image; for the "Forest Spirits and Beasts" paragraph, an image is generated showing spirits casting spells in the forest and beasts lurking in the grass, with the text "Forest spirits cast spells, and ferocious beasts lurk in the shadows"; for the "Player's Treasure Hunt" paragraph, an image is generated showing the player carefully exploring the forest, searching for clues to treasure, with the text "The player shoulders a heavy responsibility, exploring the forest to find treasure." Through this process, a complete text-image sequence of a game book about the mysterious forest scene in Game A is finally generated.
[0093] For easier understanding, please refer to Figure 9 , Figure 9The diagram illustrates the method for generating a text-image work. First, original text information 910 is acquired. Second, based on the original text information 910, retrieved text information 920 related to the content elements represented by the original text information. Next, the original text information 910 and the retrieved text information 920 are input into a specific prompt template to obtain target text material 930. Then, the target text material 940 is input into a large language model to obtain target text information 950. Finally, the target text information 950 is input into an image generation model to obtain a text-image sequence 960.
[0094] The method provided in this application fully utilizes the advantages of original text information, retrieved text information, large language models, and image generation models to achieve efficient generation of high-quality graphic works. Furthermore, it can be flexibly adjusted and expanded according to different game elements and user needs to meet diverse game picture book creation requirements.
[0095] In an optional embodiment of the graphic and textual work generation method provided in the above embodiments of this application, the method further includes:
[0096] Based on the retrieved text information and target text material, the large language model is trained through context learning to adjust its parameters.
[0097] Understandably, after obtaining the retrieved text information and generating the target text material according to the above method, the work of contextual learning and training of the large language model based on them begins.
[0098] First, ensure that the selected large language model is in a trainable and adjustable state, and has the necessary functional configurations, such as interfaces for receiving external input data for training. Simultaneously, clearly define the basic architecture, parameter structure, and existing pre-training status of the current large language model. This will help in more targeted training and parameter adjustments based on new data (i.e., retrieved text information and target text materials).
[0099] The retrieved text information and target text materials are integrated and processed to conform to the format requirements of the large language model's input data. For example, if the large language model requires the input text data to be presented in a specific encoding method (such as UTF-8 encoding) and has specified text length limits, text delimiters, and other format requirements, then the data needs to be converted and organized accordingly. If the integrated text data volume is large, in order to improve training efficiency and ensure the stability of the training process, this data can be divided into multiple training batches. For example, based on criteria such as the number of text paragraphs or the number of characters, the data can be evenly divided into several small batches, with each batch having a moderate data size, ensuring that the model can effectively process the data in a single training iteration without causing memory overflow problems due to excessive data volume.
[0100] The integrated retrieved text information and target text material are input into the large language model in a dialogue format. This dialogue format can simulate real human-computer dialogue scenarios. For example, the retrieved text information is used as background knowledge, and the user asks the large language model questions related to the target text material. For example, "Given [related content of the retrieved text information], what do you think about [the situation described in the target text material, such as the rationality of hero character A's behavior in a certain task]?" This allows the large language model to learn the context based on these inputs and generate candidate text materials as responses.
[0101] Upon receiving such input, the large language model, based on its pre-trained knowledge and understanding of the retrieved text information and target text material, performs feature extraction and semantic analysis through internal components such as multi-head attention mechanisms and feedforward neural networks. This generates candidate text material that is logically consistent and relevant to the input context. For example, in response to the question about hero character A, the large language model might generate candidate text material that further explains why hero character A's behavior in this task is reasonable based on their skill characteristics, personality traits, and the environmental circumstances, providing detailed reasons and relevant analysis.
[0102] Next, the previously generated candidate text materials are input into the large language model again in a dialogue format. At this time, the large language model will perform comparative learning and further contextual learning based on the candidate text materials, the previously input search text information, and the target text materials.
[0103] In this process, the large language model compares the candidate text materials with the original target text materials, analyzing their differences in language expression, logical reasoning, and content completeness. Simultaneously, it incorporates the rich background knowledge provided by the retrieved text information to adjust and optimize its parameters, ensuring that the subsequently generated content better meets expectations. For example, if there is a discrepancy between the description of hero character A's skill application in the candidate text materials and the actual effect of the skill in the retrieved text information, the large language model will correct this discrepancy by adjusting its internal parameters (such as weight parameters related to semantic understanding and logical connections), so that the next generated content more accurately reflects the application of hero character A's skill in a specific context and its fit with the overall story background.
[0104] Based on the comparative and contextual learning observed above, specific strategies for adjusting the parameters of the large language model are formulated. Generally, the goal of the adjustment is to enable the large language model to more accurately, reasonably, and richly output text content related to the game picture book when generating target text information based on similar target text materials, and to better integrate various relevant elements contained in the retrieved text information.
[0105] For example, if it is found that when a large language model generates story content about a hero character 'a', the emotional description of 'a' in a specific battle scene is not subtle enough and does not match the character's personality traits in the retrieved text information, then it is necessary to adjust the parameters related to emotional expression and text detail generation to enhance the model's performance in this regard.
[0106] For different types of parameters in a large language model, appropriate adjustment methods are adopted. Common parameters, such as weight matrix parameters (found in fully connected layers, attention mechanisms, etc.), can be updated and adjusted using gradient descent-based optimization algorithms (such as stochastic gradient descent, Adam optimization algorithm, etc.). Based on the loss function calculated during the learning process (which can be constructed by comparing the difference between the generated candidate text materials and the expected reasonable text), the gradients corresponding to each parameter are calculated. Then, the weight parameters are updated according to the set learning rate (the learning rate can be adjusted according to the actual training situation; an initial small value, such as 0.001, can be set, and then appropriately changed according to the training convergence). This optimizes the model in the direction of reducing the loss function.
[0107] The same principle applies to adjusting some bias parameters (such as the bias terms of each neuron) to change the activation threshold and other characteristics of the neuron, thereby optimizing the output performance of the model.
[0108] During parameter tuning, continuously monitor the model's performance on a validation dataset (a subset of existing game-related text data can be extracted as a validation set; the data distribution of the validation set should be as similar as possible to the training data, and it should not be used in the training process, but only for evaluating model performance). For example, determine if model performance has improved by calculating similarity metrics between the generated text and the actual expected text (such as BLEU scores, ROUGE scores, etc.) and evaluating the logical rationality of the text (through manual sampling or using natural language processing tools for grammatical and semantic logic detection). If it is found that after a certain number of rounds of parameter tuning, the model performance no longer improves or even declines (i.e., overfitting occurs), stop the parameter tuning process and save the model parameters corresponding to the current optimal performance.
[0109] It should be noted that in actual operation, the entire training and parameter adjustment process may require multiple trials and optimizations. The specific implementation details should be flexibly adjusted according to different game content, text characteristics and the model's own situation in order to achieve the best training effect and the quality of the generated graphic works.
[0110] For easier understanding, please refer to Figure 10 , Figure 10 A schematic diagram illustrating the model training process is shown. (For example...) Figure 10 As shown in (A), first, the model parameters are set. For example, the selected model is chatGpt-3.5-turbo, the randomness is set to 1.0, and the single response limit is set to 2000. Figure 10 As shown in (B), the retrieved text information and target text material are then pre-loaded by calling the API, and the large language model is trained through contextual learning to obtain prompt words full of prior knowledge.
[0111] The method provided in this application embodiment trains a large language model through contextual learning based on retrieved text information and target text materials, and adjusts its parameters so that the large language model can better adapt to specific graphic and text creation tasks (such as game picture book creation), improve the quality of subsequent generated target text information, and thus lay a good text foundation for the entire graphic and text creation process.
[0112] In an optional embodiment of the graphic and textual work generation method provided in the above embodiments of this application, context learning training is performed on a large language model based on retrieved text information and target text materials, including:
[0113] Based on the dialogue format, the retrieved text information and target text material are input into the large language model to obtain candidate text material. The large language model is used to learn the context based on the retrieved text information and target text material.
[0114] Based on the dialogue format, candidate text materials are input into the large language model, enabling the large language model to perform comparative learning and contextual learning based on the candidate text materials, retrieved text information, and target text materials.
[0115] Understandably, the process begins with organizing the retrieved text information and target text material into a conversational input format. Next, this constructed conversational input data is sent to a large language model. Upon receiving this information, the large language model processes it using its internal neural network structure. Based on its pre-trained knowledge and understanding of the current input retrieved text information and target text material, it performs correlation analysis on different parts of the text using a multi-head attention mechanism, extracting key semantic features. For example, it focuses on the skill description of hero character A and its plot-related aspects within the story framework, analyzing the potential impact and role of skills within that plot. Then, after further processing by a feedforward neural network, candidate text material is generated as a response to the user's question. For instance, the large language model might generate detailed descriptions about how hero character A can better utilize skills in the task, the challenges they might encounter, and coping strategies. These descriptions, based on skill details in the retrieved text information and the task framework in the target text material, supplement and expand upon the original story framework, forming candidate text material.
[0116] The candidate text materials generated in the previous step are reconstructed into conversational input and fed into the large language model. Upon receiving the candidate text material input, the large language model compares it with the previously retrieved text information and the target text material. During the comparison, analysis is performed from multiple dimensions, including the accuracy of language expression, logical coherence, and content completeness. For example, it checks whether the description of hero character 'a' in the candidate text material is consistent with the character setting in the retrieved text information, and whether the story development logic conforms to the overall framework set by the target text material. Simultaneously, the large language model utilizes context learning mechanisms to re-examine various detailed elements in the retrieved text information (such as character traits, relevant rules in the game world, etc.) and the story direction in the target text material, adjusting its parameters based on the new content in the candidate text material.
[0117] In this process, the neural network within the large language model calculates parameter updates using the backpropagation algorithm. For example, if a candidate text describes the emotional changes of hero character A in a way that doesn't match the character's personality in the retrieved text, the model calculates the gradients of parameters related to emotional expression and then updates these parameters according to a set learning rate to optimize the model's accuracy in handling similar emotional descriptions. Simultaneously, parameters related to logical reasoning, such as those used to judge the plausibility of plot development, are also adjusted based on comparison results, enabling the model to better construct reasonable and engaging story content based on the retrieved text and the target text.
[0118] For easier understanding, please refer to Figure 11 , Figure 11 The processing flow of the large language model is illustrated. First, the original text information 1110 is acquired. Second, based on the original text information 1110, retrieval text information 1120 related to the content elements represented by the original text information is obtained. Next, the original text information 1110 and the retrieval text information 1120 are input into a specific prompt template to obtain the target text material 1130. Then, the target text material 1130 and the retrieval text information 1120 are repeatedly input into the large language model to obtain the target text information 1150. Finally, the target text information 1150 is input into the image generation model to obtain the image-text sequence 1160.
[0119] The method provided in this application embodiment, through a two-stage input and learning process based on dialogue, enables the large language model to continuously optimize its understanding and processing capabilities for information related to the creation of specific graphic works, improve the quality of the generated text, make it more in line with the needs of graphic works (such as game picture books), lay a solid foundation for the subsequent generation of high-quality target text information, and gradually adapt to the creative style and logical requirements of specific game content during continuous training.
[0120] In an optional embodiment of the graphic and textual work generation method provided in the above embodiments of this application, adjusting the parameters of the large language model includes:
[0121] The parameters of the prediction layer of the large language model are adjusted.
[0122] Understandably, training data is input into a large language model, processed by the preceding layers, and then passed to the prediction layer. Based on the current parameter settings of the prediction layer, a corresponding predicted output (such as the probability distribution of the text) is generated. Then, by comparing the predicted output with the target output (i.e., the expected correct text description) in the training data, the loss function value is calculated. Common loss functions, such as cross-entropy loss, quantify the degree of difference between the predicted result and the true result.
[0123] Using the backpropagation algorithm, starting from the prediction layer, the gradients for each parameter (weight matrix parameters and bias parameters) are calculated along the backpropagation path of the model. The gradient indicates in which direction the parameters should be adjusted in the current state to reduce the loss function value, that is, to make the model's predictions closer to the true results. For example, if the text generated by the model describing a hero character's skill 'a' deviates significantly from the accurate description in the training data, the gradient calculated through backpropagation will instruct how the weight parameters corresponding to skill-related words in the prediction layer should be adjusted—whether to increase or decrease their weight values—to improve the probability of selecting the correct skill description words.
[0124] The parameters of the prediction layer are updated based on the calculated gradient and the set learning rate. The learning rate is a pre-defined hyperparameter that determines the step size of each parameter update. If the learning rate is set too large, the parameters may skip the optimal value during the adjustment process, making it difficult for the model to converge or even degrading its performance; if the learning rate is set too small, the parameter updates will be very slow, and the training time will increase significantly. Generally, a moderate learning rate can be chosen initially, such as around 0.001, and then adjusted appropriately according to the model's convergence during training.
[0125] The weight matrix parameters are updated using the formula "new weight = old weight - learning rate × gradient"; a similar update method is used for the bias parameters. After each parameter update, the training data is input into the model again for the next round of training. This process is repeated continuously to gradually optimize the parameters of the prediction layer, making the text generated by the model increasingly meet the requirements of game picture book creation, that is, accurately describing game characters, plots, and other content, and having good logic, coherence, and consistency with the game's world setting.
[0126] After each round of parameter update training (or at regular intervals), the performance of the large language model is evaluated using validation data. The validation data is input into the model with updated parameters, and the previously defined evaluation metrics (such as perplexity, similarity, logicality, and coherence) are calculated. The trends of these metrics as the number of training rounds increases are observed.
[0127] If, with increasing training epochs, the model's performance metrics on validation data (such as gradually decreasing perplexity, gradually increasing similarity, and improving logical coherence) continuously improve, it indicates that the parameter adjustments to the prediction layer are effective, and the model is converging in a better direction. Parameter adjustments and training can continue. However, if performance metrics stop improving or even begin to decline, this may indicate overfitting—the model has become too adapted to the training data and performs poorly on new data (validation data). In this case, the parameter adjustments should be stopped, and the optimal prediction layer parameter settings should be saved as the final adjusted parameters for the large language model's prediction layer. This will allow for subsequent generation of high-quality target text information, better serving the creation of graphic works such as games and picture books.
[0128] For easier understanding, please refer to Figure 12 , Figure 12 The processing flow of the large language model is illustrated. First, the original text information 1210 is acquired. Second, based on the original text information 1210, retrieval text information 1220 related to the content elements represented by the original text information is obtained. Next, the original text information 1210 and the retrieval text information 1220 are input into a specific prompt template to obtain the target text material 1230. Then, the target text material 1230 and the retrieval text information 1220 are repeatedly input into the large language model for fine-tuning 1240.
[0129] In fine-tuning a large language model, the existing ChatGPT model is first loaded. At the start of training, most of the model's parameters are frozen, with parameters only released in the last prediction layer. This is a common training strategy aimed at making targeted adjustments to the model based on the existing model. For example, a pre-trained ChatGPT model has already learned a large amount of language knowledge and semantic understanding, but for a specific game picture book generation task, it needs to be fine-tuned to adapt to the new task without destroying the original knowledge structure. Here, due to the relatively small amount of data prepared (compared to large-scale pre-training data), the epoch is set to 1 (meaning the entire training dataset is passed through the model once), the step is 2000 (meaning 2000 parameter update steps are performed during training), and the learning rate is set to 0.0001 (the learning rate determines the step size of each parameter update; a smaller learning rate helps the model converge more stably, but may require more training steps).
[0130] Gradually loosening parameters and performance evaluation: After completing one training iteration and saving the model, the parameters of each layer are gradually loosened, allowing the model to be adjusted across more layers. This is because as training progresses, the model may need to optimize more of its internal structure to improve performance. During this process, an early stop mechanism is employed. This involves continuously evaluating the model's performance on the validation set (e.g., the matching degree between generated text and expected game picture book content, text quality, etc.) during training. When the model's performance stops improving or even begins to decline, training is stopped, and the model parameters corresponding to the epoch with the best performance are retained. Simultaneously, during fine-tuning, a grid search is used to find the optimal parameter combination, including the learning rate (tried different values such as 0.001, 0.0005, and 0.0001), optimizer (different optimization algorithms such as 'adam' and 'sgd'), and number of epochs (1, 5, 10, etc.). By comprehensively searching and evaluating these parameter combinations, the most suitable parameter settings for the game picture book generation task are determined, thereby improving the model's generation effect and performance. For example, during fine-tuning, a grid search method is used to find the optimal parameters for learning rate, epochs, etc.
[0131] Learning rate: lr = [0.001, 0.0005, 0.0001];
[0132] Optimizer: optimizer = ['adam', 'sgd'];
[0133] Epoch count: [1,5,10] #The process will stop early depending on the results.
[0134] The method provided in this application makes targeted adjustments to the prediction layer parameters of a large language model, which can optimize the performance of the large language model in specific graphic and text creation tasks, making its output text more in line with creative needs and improving the quality and effect of the entire graphic and text creation process. In actual operation, it may be necessary to flexibly adjust and optimize the details such as parameter settings, training rounds, and evaluation frequency in the above steps according to the specific characteristics of the large language model, the actual requirements of game picture book creation, and data conditions.
[0135] In an optional embodiment of the graphic and textual work generation method provided in the above embodiments of this application, generating a graphic and textual sequence corresponding to the target text information according to an image generation model includes:
[0136] The target text information is segmented into words to obtain M target text paragraphs;
[0137] M image-text information items are generated based on M target text paragraphs, wherein the representation information of the M screen contents in the M image-text information items corresponds to the content of the M target text paragraphs.
[0138] Understandably, the first step is to obtain the target text information generated by a large language model. This text information details the content related to the graphic work, such as the story of game characters and scene characteristics in a game picture book. This target text information is then processed using word segmentation tools or techniques from natural language processing. Common word segmentation methods include rule-based segmentation, statistical segmentation, and deep learning-based segmentation methods.
[0139] For example, consider the following target text information about hero character a in game A: "Hero character a is a hero character in game A. Her backstory depicts her as a mercenary. She grew up under the tutelage of the Mercenary Guild, learning various combat skills. After reaching adulthood, she began to perform various missions. In a mission to protect the youngest daughter of Count Rose, Adriss, she encountered a powerful Fallen Legion. With her superb combat abilities and unwavering faith, she fought alongside Adriss and the Rose Knights, successfully completing the mission." This text can be segmented into words such as "hero character a," "is," "game A," "in," "one," "hero character," "her," "backstory," "depict," "one," "mercenary," "of," and "image," facilitating further analysis of the text's semantic structure.
[0140] After word segmentation, the text is divided into paragraphs based on factors such as semantic logic, plot development, and content relevance, resulting in M target text paragraphs. For example, the text information of the hero character 'a' can be divided into the following paragraphs according to different content sections:
[0141] The first paragraph states: "Hero character A is a hero character in game A. Her backstory depicts her as a mercenary. She grew up under the tutelage of the Mercenary Guild, learning various combat skills." This paragraph mainly focuses on the character's basic background and growth experience, introducing her origins and early ability development.
[0142] The second paragraph reads: "After reaching adulthood, he began to undertake various missions. During a mission to protect Count Rose's youngest daughter, Etricia, he encountered a powerful legion of Fallen Ones." This paragraph focuses on describing the character's entry into the mission execution phase and introduces a key mission scenario.
[0143] The third paragraph states: "With her superb combat skills and unwavering faith, she fought alongside the Rose Knights led by Adris and successfully completed the mission." This paragraph focuses on the character's specific performance in the key mission and the final result.
[0144] The M target text paragraphs (M being an example of 3 paragraphs) are divided here. Each paragraph has a relatively independent and complete theme, providing a clear text foundation for generating the corresponding image and text sequence later.
[0145] The M predefined target text segments are sequentially input into the image generation model. Currently used image generation models (such as DALL-3, DALL-E, and StableDiffusion) all have the ability to generate corresponding images based on text descriptions. However, it is necessary to ensure that the text conforms to the model's formatting requirements during input. For example, some models may require certain encoding conversions of the text content (such as UTF-8 encoding), and there may be restrictions on text length and special characters. Therefore, the format of each target text segment needs to be adjusted accordingly to ensure that it can be accurately received by the image generation model.
[0146] After receiving input for each target text segment, the image generation model performs semantic understanding and feature extraction on the text content based on its internal neural network architecture and pre-trained knowledge. Taking the second text segment describing the key mission scene of hero character A as an example, the image generation model will identify the characters involved (hero character A, Count Rose's youngest daughter Adriana, members of the Rose Knights, etc.), events (performing a protection mission, encountering the Fallen Legion, etc.), and key contextual elements (mission scene, powerful enemies, etc.). Then, through generative adversarial networks (GANs), variational autoencoders (VAEs), and other related technologies (different image generation models rely on different core technologies) to generate the corresponding visual content.
[0147] The generated image might depict hero character A and members of the Rose Knights standing ready against the approaching Fallen Legion. The image would showcase a tense atmosphere, the expressions of the characters, and a scene setting consistent with the game's world view. Furthermore, the information represented by this image would closely correspond to the content described in the second target text paragraph, intuitively presenting the plot scene described in the text.
[0148] While generating the image, key text content is extracted from the corresponding target text paragraph as the text portion of the image-text sequence. For example, for the generated task scene image above, a concise text statement that summarizes the core content of the scene, such as "Hero character A and the Rose Knights prepare to deal with the attack of the Fallen Legion," can be extracted and placed in an appropriate position on the image, such as below or beside the image, forming a well-matched image-text unit.
[0149] Following the same method, a corresponding image and matching text content are generated for each target text paragraph, ultimately forming M image-text information pieces. The representational information of the image content in each image-text information piece corresponds to the content of the corresponding target text paragraph. Together, these image-text information pieces constitute a complete image-text work, which can vividly and figuratively display relevant content such as the story of hero character 'a' in a game picture book, allowing readers to better understand and experience the entire story scene and plot development through a combination of images and text.
[0150] Please see Figure 13 , Figure 13 The diagram illustrates the generation of the image-text sequence. The target text input to the image generation model is "These are low-resolution pixel art game items: a bronze shield, a quiver, an iron helmet, and a pair of leather boots. These designs are suitable for retro-style RPG games." The image generation model generates four images of the game items, along with the text information.
[0151] The method provided in this application embodiment can generate corresponding graphic and text sequences in an orderly manner based on target text information, realize the organic combination of text and images, provide strong support for generating high-quality graphic and text works (such as game picture books), and in actual operation, the specific details of word segmentation, paragraph division, image generation and other steps can be flexibly adjusted and optimized according to different graphic and text works themes, style requirements and the characteristics of the selected image generation model.
[0152] In an optional embodiment of the graphic and textual work generation method provided in the above embodiments of this application, the method further includes:
[0153] Determine the visual style based on the target text information.
[0154] Generate a text-image sequence based on M target text paragraphs, including:
[0155] Based on the visual style, generate M image and text information from M target text paragraphs.
[0156] Understandably, the first step is to extract key elements from the target text information generated by the large language model that can suggest the style of the visuals. These elements can cover multiple aspects, such as:
[0157] Story background related: If the target text describes a game world with a mysterious and fantasy theme, such as magic, mysterious creatures, ancient ruins, etc. (e.g., "In the mysterious forest of game A, there is an ancient castle that emits a strange light, and elves who can cast magic often appear around it"), then it may tend to have a fantasy style of visuals, and the visuals will show gorgeous magic effects, fantastic creature designs, and mysterious scene settings.
[0158] In terms of character traits: If the text emphasizes the character's bravery and heroic spirit, and is set in a plot full of intense battles (such as "Hero character A, wielding a greatsword and wearing heavy armor, charges into battle without fear of the enemy's fierce attacks, each swing of the sword bringing a splash of blood"), then a rugged, realistic, and dynamic visual style is more suitable. The color tone of the visuals may be heavy, and the lines may be hard, in order to highlight the intensity of the battle and the character's power.
[0159] Emotional Atmosphere: When the text creates a warm and romantic emotional atmosphere (such as "On a quiet night in a small town, hero character A and his beloved stroll along a moonlit street, with the surrounding flowers emitting a faint fragrance"), the corresponding visual style may be soft and beautiful, with warm and gentle colors and delicate brushstrokes to convey this romantic and beautiful feeling.
[0160] A style library containing various common visual styles can be pre-built, with corresponding feature descriptions and example images for each style (e.g., fantasy style example images feature dazzling magical light and unique creature designs; realistic style example images emphasize detail and closely resemble real-world scenes or characters). Key elements extracted from the target text are compared and matched with the style features in the library to find the most suitable visual style. For example, if the text contains numerous magical elements and the story background is full of fantasy settings, comparing it with the style library determines its visual style to be fantasy.
[0161] In addition to style libraries, rules for style setting can be established. For example, if the game scene mentioned in the text is set in ancient history and involves court elements, but lacks obvious fantasy or magic elements, it can be classified as a classical realistic style according to the rules. The visuals will emphasize details such as ancient architecture and clothing, with a rustic and elegant color palette. Based on these rules and the specific content of the target text, the visual style can be accurately determined.
[0162] When inputting M target text paragraphs into the image generation model, specific cue words indicating the visual style are added simultaneously. For example, after determining the visual style to be "Japanese anime style," cue words such as "presented in Japanese anime style" are added before or after each target text paragraph. After receiving the text input with style cue words, the image generation model adjusts the style of the generated image based on the image features of Japanese anime style learned during its pre-training (such as characters having large eyes, bright and contrasting colors, and simple and smooth lines).
[0163] Depending on the style of the image, the corresponding model parameters are adjusted. For example, if the style is determined to be oil painting, parameters related to brushstroke texture, color saturation, and lighting effects may need to be adjusted to make the generated image more closely resemble the characteristics of oil painting: thick paint, rich colors, and textured brushstrokes. By consulting model documentation or conducting style parameter debugging and verification in the early stages, the specific parameter adjustment values corresponding to each style are determined. Then, before generating the image corresponding to each target text paragraph, the model parameters are set accordingly to ensure that the generated image conforms to the established image style.
[0164] After passing the visual style requirements to the image generation model in the above manner, image generation is performed for each of the M target text paragraphs. For example, for a target text paragraph describing a hero character A exploring a mysterious forest in a game, and which has been determined to be in a fantasy style, the image generation model might generate a colorful forest scene with trees of peculiar shapes, magical runes shimmering with mysterious light floating in the air, and hero character A wearing clothing with mysterious patterns. The overall scene is full of fantasy elements, and the content of this scene closely corresponds to the exploration scene described in the target text paragraph.
[0165] After generating the images, key text content is extracted from the corresponding target text paragraphs and used as the text portion of the image-text sequence. This text is then placed in appropriate positions within the images to form well-matched image-text units. Following this process, images that meet the stylistic requirements of the visuals and matching text content are generated sequentially for M target text paragraphs, ultimately forming an image-text sequence. These sequences not only correspond to the target text information in terms of content but also maintain a unified and harmonious visual style, creating a unique and text-content-appropriate visual effect for the final image-text work (such as a game picture book), thereby enhancing the work's artistic appeal and overall quality.
[0166] The method provided in this application determines the visual style based on the target text information and generates a text-image sequence based on this style, ensuring that the entire text-image work maintains a consistent visual style. A consistent visual style helps to strengthen the thematic expression of the text-image work.
[0167] In an optional embodiment of the graphic and textual work generation method provided in the above embodiments of this application, a corresponding target text material is generated based on the original text information and the retrieved text information, including:
[0168] The original text information and the retrieved text information are combined according to the preset prompt template to obtain the target text material.
[0169] Understandably, the first step is to design a pre-defined prompt template. This template should have a clear structure, be able to accommodate key content elements from both the original and retrieved text information, and be arranged in a logical order. For example, for a game-based picture book creation scenario, the template could be designed to include different sections such as "Character Information," "Background Story," "Skills and Abilities," and "Related Plot Clues." Each section should have reserved space for subsequent filling in of relevant information.
[0170] Taking "Hero character a in game A" as an example, the "Character Information" section in the template can be filled in with basic information such as character name and identity; the "Background Story" section is used to place search text information related to the character's growth experience and background; the "Skills and Abilities" section is used to describe the character's skill details obtained from the search text information; and the "Related Plot Clues" section can record information such as important events or mission plots related to the character.
[0171] When designing templates, it's essential to consider both versatility and flexibility. Versatility means the template can be applied to different types of original and retrieved text information. Whether it's information about game characters, game scenes, or game items, it can find a suitable place within the template for integration. For example, for the original text information of a game scene, the "Background Story" section of the template can be used to describe the scene's historical origins, geographical features, and other background information, while the "Related Plot Clues" section can record game events or tasks that occurred within that scene.
[0172] Flexibility is reflected in the ability to adjust the template appropriately according to the specific needs of creating graphic works. For example, if a particular game element is found to require a more detailed description during the creation process, a corresponding sub-section can be added to the template or the content scope of a section can be expanded.
[0173] Key content is extracted from the acquired raw and retrieved text information and organized according to the requirements of a preset template. Taking "hero character a in game A" as an example, the raw text information identifies the core character as "hero character a". Detailed background stories about this character are extracted from the retrieved text information, such as "Hero character a was raised by a guild as an orphan and taught combat skills. As an adult, she serves the guild as a contracted mercenary, taking a portion of her mission rewards to repay the guild's investment and interest until the debt is fully repaid." This information is then organized and placed in the "Background Story" section of the template.
[0174] Next, extract the character's skill information, such as "Her passive skill increases attack speed and fires multiple arrows; her first skill enhances attack and deals damage to enemies in front of her; her second skill deals magic damage and slows enemies; her third skill is a global control skill that can stun and damage enemy heroes." After organizing this skill information, fill it into the "Skills and Abilities" section.
[0175] The prepared original text information and searched text information are combined according to a preset template to form the target text material. For example, after filling in the information according to the above template, the target text material may appear in the following form: "Character: Hero Character A. Background Story: Hero Character A was raised by the guild as an orphan and taught combat skills. After adulthood, she served the guild as a contract mercenary and took a portion of the mission rewards to repay the guild's investment and interest on her until the debt was paid off. Skills and Abilities: Passive skill increases attack speed and fires multiple arrows; Skill 1 enhances attack and deals damage to enemies in front; Skill 2 deals magic damage and slows enemies; Skill 3 is a map-wide control skill that can stun and damage enemy heroes. Related Plot Clues: She has performed missions such as disguising herself as a noble maid to provide protection. In one mission, she was responsible for protecting Count Rose's youngest daughter, Adriss. Facing the powerful Fallen Legion, she fought alongside Adriss and the Rose Knights."
[0176] The method provided in this application combines the original text information and the retrieved text information according to a preset prompt template, which can generate target text materials efficiently and orderly, laying a good foundation for the subsequent generation of target text information based on a large language model. In actual operation, the preset prompt template can be optimized and adjusted according to different game content, creative style and graphic work type, etc., to better adapt to various creative needs.
[0177] In an optional embodiment of the graphic and textual work generation method provided in the above embodiments of this application, obtaining original text information and retrieving text information includes:
[0178] Obtain the original text information;
[0179] The retrieved text information is obtained from the original text information, where the original text information and the retrieved text information correspond to the same content elements.
[0180] Understandably, in scenarios involving the creation of graphic works such as game-based picture books, users interact with the system through a corresponding client (such as a large language model client, which may be presented as a browser, a standalone app, or a mini-program) running on a terminal device (which could be a mobile phone, tablet, or computer). Users input relevant content into the information input window provided on the client interface, which serves as the source of the original text information. For example, if a user wants to generate a picture book about a specific scene in game A, they would enter "the mysterious forest scene in game A" in the input window. This input becomes the original text information, clearly defining the core content element around which the subsequent graphic work is generated: the mysterious forest scene.
[0181] In addition to user-inputted content, the system can also recommend or provide preset content as raw text information based on certain rules. For example, based on a user's browsing history and usage preferences (e.g., if the user frequently views character-related content in game A or frequently uses the character-related image and text creation function), the system can recommend popular game A character names (such as "hero character b in game A") as raw text information. Alternatively, in certain functional modules, the system can preset some representative game elements (such as classic main storyline scene names in game A) and display them to the user, who can then choose one as the raw text information to initiate the image and text creation process.
[0182] After receiving the original text information on the server side (assuming the server is responsible for handling the core logic of generating text and image works and data retrieval operations), it will retrieve the retrieved text information from the local database based on the original text information. The local database stores a massive amount of game-related data, covering detailed information such as all character attribute information, storyline, scene maps, and item functions for game A.
[0183] For example, when the original text information is "Hero character a in game A", the server uses a database query to search for the corresponding text information in the database. The retrieved information includes detailed skill descriptions for hero character a (including skill effects, cooldown times, and other specific parameters), their complete growth path in the game (such as which faction they belonged to, what important training stages they underwent, etc.), their relationships with other characters (allies or rivals, what collaborations or conflicts they had), and the main game story missions they participated in (the objectives, process, and results of each mission), among many other aspects. All this retrieved information is closely related to the original text information's reference to "hero character a," providing a comprehensive supplement and expanded description.
[0184] To ensure the retrieved text information is more comprehensive and up-to-date, the server can also integrate online resources. On one hand, it can interact with the game's official API to obtain the latest information on game elements, such as adjustments to character skills after updates or new character backstories, and incorporate this information into the retrieved text. On the other hand, it can also scrape high-quality user-shared content from authoritative game communities, forums, or professional game strategy websites. For example, in-depth analyses of hero character A's strengths and weaknesses in specific battle scenarios, or information shared by players on forums about hidden character stories, can be filtered and organized and included as part of the retrieved text, further enriching the information reserves related to the content elements corresponding to the original text information.
[0185] After obtaining a large amount of potential search text information through various channels, it is necessary to filter and integrate it to ensure that the search text information highly corresponds to the original text information and has practical value. For information obtained from local databases, online resources, etc., duplicate, redundant, or content elements that are not strongly related to the content elements pointed to by the original text information should be removed. For example, if some scattered information about other unrelated characters in game A is found, and this information does not help describe "hero character a", it should be discarded.
[0186] Then, the filtered searched text information is integrated according to a certain logical order and category, which facilitates the subsequent generation of target text materials based on the original text information and these searched text information. For example, the searched text information for hero character 'a' can be organized into categories such as skills, background story, and plot participation, making it clear and easy for subsequent text processing and creation processes to utilize this information. This ensures that the entire graphic and text creation process can revolve around the same content element (such as "hero character 'a' in game A"), obtaining sufficient and effective information support from different angles, thereby generating high-quality graphic and text works.
[0187] The method provided in this application embodiment can systematically complete the operation of obtaining original text information and corresponding retrieved text information, providing a solid information foundation for the subsequent generation of graphic works. In practical applications, the specific operation details of obtaining information can be flexibly adjusted and optimized according to different types of graphic works, application scenarios, and data sources involved.
[0188] The following is a detailed description of the device for generating graphic works in this application. For example... Figure 14 , Figure 14 Figure (A) shows a structural diagram of a graphic work generation device 1400, which includes:
[0189] The acquisition module 1410 is used to acquire original text information and retrieve text information. The original text information is used to represent content elements, and the retrieved text information is information related to the content elements obtained based on the original text information.
[0190] The generation module 1420 is used to generate corresponding target text materials based on the original text information and the retrieved text information;
[0191] The generation module 1420 is also used to generate target text information corresponding to the target text material based on the large language model;
[0192] The generation module 1420 is also used to generate a graphic-text sequence corresponding to the target text information according to the image generation model. The graphic-text sequence includes M graphic-text information, the representation information of the M screen contents in the M graphic-text information corresponds to the content of the target text information, and the M graphic-text information includes text content, the text content is a part of the target text information, and M is an integer greater than or equal to 1.
[0193] In another embodiment of this application, please refer to Figure 14 In section (B), the graphic and textual artwork generation device 1400 also includes:
[0194] Training module 1430 is used to train the large language model through contextual learning based on the retrieved text information and target text material, so as to adjust the parameters of the large language model.
[0195] In another embodiment of this application, the training module 1430 is further configured to:
[0196] Based on the dialogue format, the retrieved text information and target text material are input into the large language model to obtain candidate text material. The large language model is used to learn the context based on the retrieved text information and target text material.
[0197] Based on the dialogue format, candidate text materials are input into the large language model, enabling the large language model to perform comparative learning and contextual learning based on the candidate text materials, retrieved text information, and target text materials.
[0198] In another embodiment of this application, the training module 1430 is further configured to:
[0199] The parameters of the prediction layer of the large language model are adjusted.
[0200] In another embodiment of this application, the generation module 1420 is further configured to:
[0201] The target text information is segmented into words to obtain M target text paragraphs;
[0202] M image-text information items are generated based on M target text paragraphs, wherein the representation information of the M screen contents in the M image-text information items corresponds to the content of the M target text paragraphs.
[0203] In another embodiment of this application, the generation module 1420 is further configured to:
[0204] Determine the visual style based on the target text information.
[0205] Generate M image-text information items based on M target text paragraphs, including:
[0206] Based on the visual style, generate M image and text information from M target text paragraphs.
[0207] In another embodiment of this application, the generation module 1420 is further configured to:
[0208] The original text information and the retrieved text information are combined according to the preset prompt template to obtain the target text material.
[0209] In another embodiment of this application, the acquisition module 1410 is further configured to:
[0210] Obtain the original text information;
[0211] The retrieved text information is obtained from the original text information, where the original text information and the retrieved text information correspond to the same content elements.
[0212] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0213] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0214] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0215] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0216] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0217] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a server or terminal device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing computer programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0218] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for generating graphic works, characterized in that, include: The method involves acquiring original text information and retrieving text information, wherein the original text information is used to characterize content elements, and the retrieved text information is information related to the content elements obtained based on the original text information. Based on the original text information and the retrieved text information, the corresponding target text material is generated; Generate target text information corresponding to the target text material based on the large language model; The target text information is generated according to the image generation model, and the image-text sequence includes M image-text information. The representation information of the M screen contents in the M image-text information corresponds to the content of the target text information. The M image-text information includes text content, which is a part of the target text information. M is an integer greater than or equal to 1.
2. The method as described in claim 1, characterized in that, The method further includes: Based on the retrieved text information and the target text material, the large language model is trained using context learning to adjust its parameters.
3. The method as described in claim 2, characterized in that, The step of training the large language model based on the retrieved text information and the target text material through context learning includes: Based on a dialogue format, the retrieved text information and the target text material are input into the large language model to obtain candidate text materials. The large language model is used to perform context learning based on the retrieved text information and the target text material. Based on the dialogue format, the candidate text materials are input into the large language model, so that the large language model can perform comparative learning and contextual learning based on the candidate text materials, the retrieved text information and the target text materials.
4. The method as described in claim 2 or 3, characterized in that, The adjustment of the parameters of the large language model includes: The parameters of the prediction layer of the large language model are adjusted.
5. The method according to any one of claims 1-4, characterized in that, The step of generating the image-text sequence corresponding to the target text information according to the image generation model includes: The target text information is segmented into words to obtain M target text paragraphs; M image-text information items are generated based on the M target text paragraphs, wherein the representation information of the M screen contents in the M image-text information items corresponds to the content of the M target text paragraphs.
6. The method as described in claim 5, characterized in that, The method further includes: Based on the target text information, determine the visual style. The step of generating M image-text information items based on the M target text paragraphs includes: Based on the aforementioned visual style, M image and text information items are generated from the M target text paragraphs.
7. The method according to any one of claims 1-6, characterized in that, The step of generating corresponding target text material based on the original text information and the retrieved text information includes: The original text information and the retrieved text information are combined according to a preset prompt template to obtain the target text material.
8. The method according to any one of claims 1-6, characterized in that, The acquisition of original text information and the retrieval of text information include: Obtain the original text information; The retrieved text information is obtained based on the original text information, wherein the original text information and the retrieved text information correspond to the same content element.
9. A device for generating graphic works, characterized in that, include: The acquisition module is used to acquire original text information and retrieve text information, wherein the original text information is used to represent content elements, and the retrieved text information is information related to the content elements obtained based on the original text information; The generation module is used to generate corresponding target text materials based on the original text information and the retrieved text information; The generation module is also used to generate target text information corresponding to the target text material based on the large language model; The generation module is further configured to generate a graphic-text sequence corresponding to the target text information according to the image generation model, wherein the graphic-text sequence includes M graphic-text information, the representation information of the M screen contents in the M graphic-text information corresponds to the content of the target text information, and the M graphic-text information includes text content, the text content being a part of the target text information, and M being an integer greater than or equal to 1.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it performs the step of generating the graphic work as described in any one of claims 1 to 8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it performs the steps of generating the graphic work as described in any one of claims 1 to 8.
12. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program performs the steps of generating the graphic work as described in any one of claims 1 to 8.