Model training method and device, scene image generation method and device, equipment and medium
By training the model using a character table and descriptive text of existing story scene images, the problems of inconsistent generation and character inconsistency were solved, achieving more accurate and consistent story scene image generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NETEASE (HANGZHOU) NETWORK CO LTD
- Filing Date
- 2024-10-31
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies have issues such as mismatch between generated story scene images and descriptions during model training, and inconsistencies in the same character across multiple scene images.
By acquiring a character list of a pre-defined story world, including character images and appearance attributes, and combining it with existing story scene images and descriptive text, an initial model is trained to generate a target story scene diffusion model, which is used to generate consistent story scene images based on new descriptive text.
The generated story scene images are more consistent with the descriptions, and the same character remains consistent across different images, improving the accuracy and consistency of the model's generation.
Smart Images

Figure CN121982140A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and more specifically, to a model training method, a scene image generation method, an apparatus, a device, and a medium. Background Technology
[0002] The story visualization task refers to a task in which, given a story text containing multiple sentences and images of characters appearing in the story, the model generates images frame by frame based on each sentence describing a story scene to depict the entire story unfolding.
[0003] In related technologies, models are trained based on multiple character images and general descriptions of the characters. However, in the application stage, models trained in this way are prone to problems such as the generated story scene images not matching the sentences describing the story scene, and inconsistencies in the same character across multiple story scene images. Summary of the Invention
[0004] The purpose of this application is to address the shortcomings of the prior art by providing a model training method, a scene image generation method, an apparatus, a device, and a medium to solve the aforementioned technical problems in the related technologies.
[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows:
[0006] In a first aspect, embodiments of this application provide a model training method, including:
[0007] Obtain a character list for a preset story world, the character list including multiple character images in the preset story world, and the appearance attributes of a first character in each character image;
[0008] Based on multiple existing story scene images of the preset story world and the character list, determine the appearance description text of the second character and the existing event description text associated with the second character in each of the existing story scene images;
[0009] Based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, and multiple existing story scene images, an initial story scene diffusion model is trained to obtain a target story scene diffusion model; wherein, the target story scene diffusion model is used to generate a new story scene image corresponding to the new story scene description text based on the new story scene description text.
[0010] Secondly, embodiments of this application also provide a scene image generation method, including:
[0011] Obtain new story scene description text;
[0012] A target scene image generation model is used to generate a new story scene image corresponding to the new story scene description text based on the new story scene description text, wherein the target story scene diffusion model is a model trained using the method described in any of the first aspects above.
[0013] Thirdly, embodiments of this application also provide a model training apparatus, including:
[0014] The acquisition module is used to acquire a character table of a preset story world. The character table includes multiple character images in the preset story world, as well as the appearance attributes of the first character in each character image.
[0015] The determination module is used to determine the appearance description text of the second character and the existing event description text associated with the second character in each of the existing story scene images, based on multiple existing story scene images of the preset story world and the character table.
[0016] The training module is used to train an initial story scene diffusion model based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, and multiple existing story scene images, to obtain a target story scene diffusion model; wherein, the target story scene diffusion model is used to generate a new story scene image corresponding to the new story scene description text based on the new story scene description text.
[0017] Fourthly, embodiments of this application also provide a scene image generation apparatus, including:
[0018] The acquisition module is used to acquire new story scene description text;
[0019] The generation module is used to generate a new story scene image corresponding to the new story scene description text based on the new story scene description text using a target scene image generation model, wherein the target story scene diffusion model is a model trained using the method described in any of the first aspects above.
[0020] Fifthly, embodiments of this application also provide an electronic device, including: a memory and a processor, wherein the memory stores a computer program executable by the processor, and the processor executes the computer program to implement the model training method described in any of the first aspects above, or the scene image generation method described in any of the second aspects above.
[0021] Sixthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when read and executed, implements the model training method described in any of the first aspects above, or the scene image generation method described in any of the second aspects above.
[0022] The beneficial effects of this application are as follows: This application provides a model training method, comprising: obtaining a character table of a preset story world, the character table including multiple character images in the preset story world, and the appearance attributes of a first character in each character image; determining, based on multiple existing story scene images of the preset story world and the character table, the appearance description text of a second character in each existing story scene image and the existing event description text associated with the second character; training an initial story scene diffusion model based on the appearance description text of the second character in each existing story scene image, the existing event description text associated with the second character, and multiple existing story scene images to obtain a target story scene diffusion model; wherein, the target story scene diffusion model is used to generate a new story scene image corresponding to a new story scene description text based on the new story scene description text. The rich and detailed information, such as the appearance description text of the second character in each existing story scene image and the existing event description text associated with the second character, participates in model training. When the target story scene diffusion model is applied, the new story scene image generated is more consistent with the new story scene description text, and the same character in the new story scene image can remain consistent. Attached Figure Description
[0023] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 A flowchart illustrating a model training method provided in this application embodiment. Figure 1 ;
[0025] Figure 2 A flowchart illustrating a model training method provided in this application embodiment. Figure 2 ;
[0026] Figure 3 A flowchart illustrating a model training method provided in this application embodiment. Figure 3 ;
[0027] Figure 4 A flowchart illustrating a model training method provided in this application embodiment. Figure 4 ;
[0028] Figure 5 A flowchart illustrating a model training method provided in this application embodiment. Figure 5 ;
[0029] Figure 6 A flowchart illustrating a model training method provided in this application embodiment. Figure 6 ;
[0030] Figure 7 A flowchart illustrating a model training method provided in this application embodiment. Figure 7 ;
[0031] Figure 8 A flowchart illustrating a model training method provided in this application embodiment. Figure 8 ;
[0032] Figure 9 A flowchart illustrating a model training method provided in this application embodiment. Figure 9 ;
[0033] Figure 10 A flowchart illustrating a scene image generation method provided in an embodiment of this application;
[0034] Figure 11 This is a schematic diagram of the structure of a model training device provided in an embodiment of this application;
[0035] Figure 12 This is a schematic diagram of the structure of a scene image generation device provided in an embodiment of this application;
[0036] Figure 13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of this application, but not all embodiments.
[0038] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0039] In the description of this application, it should be noted that if the terms "upper", "lower", etc. appear to indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship that the product of this application is usually placed in, it is only for the convenience of describing this application and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.
[0040] Furthermore, the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Additionally, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0041] It should be noted that, where there is no conflict, the features in the embodiments of this application can be combined with each other.
[0042] The present application provides a model training method and a scene image generation method, which are applied to an electronic device. The electronic device can be a server or a terminal device. The terminal device can be any of the following: computer equipment, laptop computer, desktop computer, smartphone, etc.
[0043] The following is an explanation of a model training method provided in the embodiments of this application.
[0044] Figure 1 A flowchart illustrating a model training method provided in this application embodiment. Figure 1 ,like Figure 1 As shown, the method may include:
[0045] S101. Obtain the character table of the preset story world. The character table includes multiple character images in the preset story world, as well as the appearance attributes of the first character in each character image.
[0046] The pre-set story world can be understood as the world in which the story takes place; for example, an animated film can correspond to a story world.
[0047] In some implementations, multiple character images for multiple first characters in a given preset story world can be collected from a publicly available network, based on a first set of characters in that preset story world.
[0048] It should be noted that the multiple character images are frontal images of multiple primary characters in the preset story world. The appearance attributes of the primary character in each character image are used to characterize the detailed appearance description information of the primary character, such as: the character's outline, clothing style and color, expression, skin color, hair color, facial features, etc.
[0049] Optionally, the character table may also include: the name of the first character in each character image; and a correspondence between multiple character images, the appearance attributes of the first character in each character image, and the name of the first character in each character image.
[0050] S102. Based on multiple existing story scene images and character lists in the preset story world, determine the appearance description text of the second character and the existing event description text associated with the second character in each existing story scene image.
[0051] In this context, the second character refers to a subset of the multiple first characters within a pre-defined story world; that is, the first characters in the character list include the second character. Existing story scene images refer to existing scene images with a storyline in which the second character participates. For example, if the multiple first characters include: Character A, Character B, Character C, and Character D, the second character could include: Character D; that is, Character D in the existing story scene images participates in the storyline. Since the second character in the existing story scene images is a subset of the characters in the character list, the description of the appearance of the second character in the existing story scene images can be determined from the character list.
[0052] In some implementations, matching is performed on multiple existing story scene images and character tables in a preset story world to obtain matching results; based on the matching results, the appearance description text of the second character in the existing story scene images and the existing event description text associated with the second character can be further determined.
[0053] It is worth noting that the appearance description text of the second character is used to represent detailed appearance information of the second character. The existing event description text associated with the second character is used to represent the specific events performed by the second character and the relationships between the second characters.
[0054] S103. Based on the appearance description text of the second character in each existing story scene image, the existing event description text associated with the second character, and multiple existing story scene images, train the initial story scene diffusion model to obtain the target story scene diffusion model.
[0055] Among them, the target story scene diffusion model is used to generate a new story scene image corresponding to the new story scene description text based on the new story scene description text.
[0056] In some implementations, an initial story scene diffusion model is used to process the appearance description text of the second character and the existing event description text associated with the second character in each existing story scene image to generate multiple result story scene images. Based on the multiple result story scene images and the multiple existing story scene images, the model parameters of the initial story scene diffusion model are updated until the training termination condition is met to obtain the target story scene diffusion model.
[0057] It should be noted that the new story scene description text may include appearance description text for a third character and event description text associated with that third character. The third character may also be one of several first characters. The final generated new story scene images are specific to the preset story world. There can be multiple new story scene images, and the appearance of the third character remains consistent across all images. The story depicted in the new story scene images is consistent with the story represented by the event description text associated with the third character.
[0058] In summary, this application provides a model training method, comprising: obtaining a character table of a preset story world, the character table including multiple character images in the preset story world, and the appearance attributes of a first character in each character image; determining, based on multiple existing story scene images of the preset story world and the character table, the appearance description text of a second character in each existing story scene image and the existing event description text associated with the second character; training an initial story scene diffusion model based on the appearance description text of the second character in each existing story scene image, the existing event description text associated with the second character, and multiple existing story scene images to obtain a target story scene diffusion model; wherein, the target story scene diffusion model is used to generate a new story scene image corresponding to a new story scene description text based on the new story scene description text. The rich and detailed information, such as the appearance description text of the second character in each existing story scene image and the existing event description text associated with the second character, participates in model training. When the target story scene diffusion model is applied, the new story scene image generated is more consistent with the new story scene description text, and the same character in the new story scene image can remain consistent.
[0059] It should be noted that the appearance description text of the second character is determined based on a character table containing multiple character images and the appearance attributes of the first character in each image. This makes the appearance description text of the second character rich in appearance attributes, more detailed and comprehensive than the general character description text in related technologies. For example, in related technologies, the general character description text can be "a cat"; while in this embodiment, the appearance description text of the second character can be "an orange fat cat with a round face, short limbs, and a long tail," and more specifically, it can include size. Moreover, during training, in addition to the appearance description text of the second character, it is also based on the existing event description text associated with the second character. This allows the trained target story scene diffusion model to solve the technical problems existing in related technologies when applied.
[0060] Figure 2 A flowchart illustrating a model training method provided in this application embodiment. Figure 2 ,like Figure 2 As shown, the process of obtaining the character list of the preset story world in S101 can include:
[0061] S201. Based on multiple character images, determine the character description text corresponding to each character image.
[0062] In some implementations, a VLM (Vision-Language Model) is used, given a character appearance description instruction. c This process describes multiple character images separately, resulting in detailed character description text for each image. The character description text for each image can be represented as C. * .
[0063] It is worth noting that the character description text corresponding to each character image can be determined sequentially, or the character description text corresponding to each character image can be determined simultaneously. This application embodiment does not impose specific limitations on this.
[0064] S202. Parse the character description text corresponding to each character image to obtain multiple appearance attributes of the first character in each character image.
[0065] In some implementations, a scene graph parser is used to parse the character description text corresponding to each character image to obtain multiple appearance attributes of the first character in each character image.
[0066] Similarly, the character description text corresponding to each character image can be parsed sequentially, or the character description text corresponding to each character image can be parsed simultaneously. This application embodiment does not impose specific limitations on this.
[0067] S203. Construct a character table based on multiple character images and multiple appearance attributes of the first character in each character image.
[0068] One character image can correspond to multiple appearance attributes of the first character in that character image.
[0069] In this embodiment of the application, multiple character images are used as indexes for the character table, and multiple appearance attributes of the first character in each character image are the values corresponding to the indexes to construct the character table.
[0070] Optionally, Figure 3 A flowchart illustrating a model training method provided in this application embodiment. Figure 3 ,like Figure 3 As shown, the process of constructing a character table based on multiple character images and multiple appearance attributes of the first character in each character image in S203 above may include:
[0071] S301. Using the first character in each character image as the character node and the appearance attributes of the first character in each character image as attribute nodes, construct a network of multiple character attribute networks corresponding to the first character in each character image.
[0072] The first role can be represented as O. * The multiple appearance attributes corresponding to the first character can be represented as: The first character can be considered a character node, and the multiple appearance attributes of the first character can be considered multiple character attribute nodes. Among them, the character node can also be called the main node.
[0073] It should be noted that, taking the construction of multiple character attribute networks corresponding to a first character in a character image as an example, for a first character in a character image, there is one character node. This character node has multiple appearance attributes. Based on this character node and one of its corresponding appearance attributes, a character attribute network can be constructed for this character node. Then, multiple character attribute networks for this character node can be constructed, that is, multiple character attribute networks corresponding to a first character in a character image. Using the same method, multiple character attribute networks corresponding to the first character in each character image can be constructed.
[0074] Optionally, a character attribute network constructed from a character node and its corresponding appearance attribute can be represented as:
[0075] S302. Based on the multiple character attribute networks corresponding to the first character in each character image, construct the character network of the first character in each character image.
[0076] Specifically, based on the attribute networks of multiple characters corresponding to a first character, a character network for that first character can be constructed. This character network can be represented as follows: Similarly, the character network of the first character in each character image can be constructed, thus obtaining the character network of multiple first characters.
[0077] S303. Using multiple character images as indexes, construct a character table based on the character network of the first character in each character image.
[0078] In this embodiment, a character table is constructed using multiple character images as indexes and the character net of the first character in each image as the corresponding index value. A character image can be represented as: I * Of course, the value corresponding to the index in the character table can also include the name of the first character in each character image.
[0079] In practical applications, the role attributes of the first role can be defined as attribute nodes connected to the main node.
[0080] Optionally, Figure 4 A flowchart illustrating a model training method provided in this application embodiment. Figure 4 ,like Figure 4 As shown, the process in S102 above, which determines the appearance description text of the second character and the existing event description text associated with the second character in each existing story scene image based on multiple existing story scene images and a character list in a preset story world, may include:
[0081] S401. Based on multiple existing story scene images, determine the description information of each existing story scene image.
[0082] In some implementations, a VLM model is used for multiple existing story scene images, given an instruction to describe the image. e This involves determining the descriptive information for each existing story scene image. The descriptive information for each existing story scene image can be represented as F. c Each existing story scene image can be represented as F.
[0083] S402. Parse the description information of each existing story scene image to obtain the description information of the second character in each existing story scene image, as well as the events related to the second character.
[0084] Among them, a scene graph parser is used to analyze the descriptive information F of each existing story scene image. c The process involves parsing to obtain the descriptive information of the second character in each existing story scene image, as well as the events related to the second character. The descriptive information of the second character can be represented as char. * Events related to the second role can be represented as R. *,*.
[0085] The above process can be represented as:
[0086]
[0087] Among them, char j This can be represented as the description information of the j-th second character, where Parser represents the scene graph parser, and R... j,* F represents the event related to the j-th second role. c Instruct represents descriptive information for each existing story scene image. e The instruction represents a description of a given image, where F indicates an existing story scene image.
[0088] S403. Based on the description information of the second character in each existing story scene image, the events related to the second character, and the character table, determine the appearance description text of the second character in each existing story scene image and the description text of the existing events associated with the second character.
[0089] In some implementations, a matching process is performed based on the description information of the second character in each existing story scene image and a character table to obtain a matching result; based on the matching result and the character table, the appearance description text of the second character in each existing story scene image is determined; based on the matching result and events related to the second character, the existing event description text associated with the second character is determined.
[0090] Optionally, Figure 5 A flowchart illustrating a model training method provided in this application embodiment. Figure 5 ,like Figure 5 As shown, the process in S403 above, which determines the appearance description text of the second character and the description text of existing events associated with the second character in each existing story scene image based on the description information of the second character in each existing story scene image, the events related to the second character, and the character table, may include:
[0091] S501. From multiple character images in the character list, find the name of the first character in the target character image that matches the description information of the second character in the existing story scene image.
[0092] Among them, the description information of the second character in the existing story scene images is textual coarse-grained description information, that is, general description information.
[0093] In this embodiment, a cross-modal model CLIP (Contrastive Language-Image Pre-training) is used to search for the target character image with the highest similarity to the descriptive information of the second character in an existing story scene image from the index of the character table, i.e., from multiple character images in the character table. The character table includes: multiple character images, the name of the first character in each character image, and appearance attributes. By querying the character table based on the target character image, the name of the first character in the target character image can be determined.
[0094] The above process can be represented as:
[0095]
[0096] Where Sim represents the CLIP similarity matching operation, char j It is F c The coarse-grained text category of the j-th second character, i.e., the descriptive information of the j-th second character; O i `argmax` represents the index of the character table, i.e., multiple character images in the character table; `argmax` indicates the operation of taking the maximum value from the index. The name of the first character in the matched target character image.
[0097] S502. Based on the name of the first character in the target character image and the character table, determine the appearance description text of the second character in each existing story scene image.
[0098] The character table includes: multiple character images, the name of the first character in each character image, and a character web, where the character web is used to represent the appearance attributes of the first character in each character image.
[0099] In this embodiment of the application, by searching the character table based on the description information of the second character in the existing story scene images, that is, the general description information of the second character, the appearance description text of the second character in each existing story scene image can be obtained, that is, the detailed appearance description information of the second character.
[0100] S503. Based on the name of the first character in the target character image, the events related to the second character, and the relationship between the first character in the target character image, determine the existing event description text associated with the second character.
[0101] In some implementations, a connection operation is performed on the name of the first character in the target character image, the events related to the second character, and the relationship between the first characters in the target character image to obtain the existing event description text associated with the second character.
[0102] The above process can be represented as:
[0103]
[0104] in, R represents the name of the first character in the target character image. j,j′ This refers to events related to the second character, i.e. and The events that occur together represent the relationship between the first characters in the target character image.
[0105] Optionally, Figure 6 A flowchart illustrating a model training method provided in this application embodiment. Figure 6 ,like Figure 6 As shown, the process in S502 above, which determines the appearance description text of the second character in each existing story scene image based on the name of the first character in the target character image and the character table, may include:
[0106] S601. Determine the appearance attributes of the first character in the target character image based on the name of the first character in the target character image and the character table.
[0107] S602. Based on the name of the target character image and the appearance attributes of the first character in the target character image, determine the appearance description text of the second character in each existing story scene image.
[0108] In some implementations, based on the name of the first character in the target character image, the corresponding character network is searched from the character table. The searched character network can represent the appearance attributes of the first character in the target character image. A connection operation is performed based on the name of the first character in the target character image and the appearance attributes of the first character in the target character image to determine the appearance description text of the second character in each existing story scene image.
[0109] The process of finding the corresponding character network described above can be represented as:
[0110]
[0111] in, The name of the first character in the target character image. For all characters in the character list, The target character network is the appearance attributes of the first character in the target character image. This is an operation for looking up the character table.
[0112] The process of determining the appearance description text of the second character in each existing story scene image can be represented as:
[0113]
[0114] in, The name of the first character in the target character image. The appearance attributes represented by the found character network. Let be the appearance description text for the j-th second character, and ⊕ represents the connection operation.
[0115] Optionally, the process in S103 above, which trains the initial story scene diffusion model based on the appearance description text of the second character in each existing story scene image, the existing event description text associated with the second character, and multiple existing story scene images, to obtain the target story scene diffusion model, may include:
[0116] Based on the appearance description text of the second character in each existing story scene image, the existing event description text associated with the second character, the preset style description text of each existing story scene image, and multiple existing story scene images, the initial story scene diffusion model is trained to obtain the target story scene diffusion model.
[0117] It should be noted that the preset style description text for each existing story scene image can be represented as W. s W s It can be a common attribute of all secondary roles and time.
[0118] Among them, the appearance description text of the second character in each existing story scene image, the existing event description text associated with the second character, and the preset style description text of each existing story scene image can be referred to as the existing story scene description text of each existing story scene image. The existing story scene description text of each existing story scene image can be represented as: n represents the total number of second characters appearing in all existing story scene images. W is the appearance description text for the j-th second character. e W is the existing event description text associated with the second role. s Preset style description text for each existing story scene image.
[0119] In this embodiment of the application, the concept of a role atlas G is defined as follows:<O,E,A> The components of G include character nodes O, events E, and attributes A. The story containing a given set of characters is treated as a whole world, and a character graph G is used as a structured knowledge graph to represent this world. Characters, events, relationships, and attributes of characters and events are selected as components of the character graph, defining the character graph as: G =<O,E,A> In this graph, O represents the set of roles, which is defined as the main node in the graph, i.e., the role node. E represents the event, which is defined as the edge connecting the main node in the graph. A represents the attribute related to the role or the event. The role attribute is defined as the attribute node connected to the main node, and the event attribute is defined as the attribute node connected to the edge.
[0120] Optionally, Figure 7 A flowchart illustrating a model training method provided in this application embodiment. Figure 7 ,like Figure 7 As shown, the process in S103 above, which trains the initial story scene diffusion model to obtain the target story scene diffusion model based on the appearance description text of the second character in each existing story scene image, the existing event description text associated with the second character, and multiple existing story scene images, may include:
[0121] S701. Calculate the probability of the second character appearing in any pixel in each existing story scene image.
[0122] Specifically, a Gaussian distribution is used to initialize the prior probability density function of the j-th second character's position in any pixel of an existing story scene image, thus obtaining the probability of the second character appearing in any pixel of each existing story scene image. Furthermore, the prior probability density function can be expressed as p j (x,y).
[0123] S702. Based on the appearance description text of the second character, the probability is corrected to obtain the corrected probability.
[0124] In some implementations, the appearance description text of the second character is input into a character knowledge encoder, and the probability is corrected based on the output of the character knowledge encoder to obtain a corrected probability. The modified probability can be expressed as: p′ j (x,y).
[0125] S703. Based on the corrected probabilities, modify the cross-attention weights in the initial story scene diffusion model to obtain the modified cross-attention weights.
[0126] In this embodiment, the spatial guidance corresponding to the second role is obtained based on the modified probability sampling, and the cross-attention weights in the initial story scene diffusion model are modified using the spatial guidance to obtain the modified cross-attention weights.
[0127] Specifically, the spatial guidance can be a matrix. Based on the set threshold of the role region, the first target part in the spatial guidance that is greater than or equal to the threshold is determined. The cross-attention weight is also a matrix, where the spatial guidance and the cross-attention weight have a corresponding relationship. Based on the first target part in the spatial guidance, the second target part in the cross-attention weight is determined. The weight of the region corresponding to the second target part is enhanced according to the preset adjustment value. The remaining part of the cross-attention weight that is not the second target part is suppressed according to the preset adjustment value, so as to obtain the modified cross-attention weight.
[0128] The above process can be represented as:
[0129]
[0130] in, For cross-attention weights, For spatial guidance, β is the threshold, and s is the preset adjustment value. Specifically, s = s(t) = α·(ln(t+1)+1) is a correction magnitude that is positively correlated with the time step t.
[0131] S704. Based on the appearance description text of the second character in each existing story scene image, the existing event description text associated with the second character, the modified cross-attention weights, and multiple existing story scene images, train the initial story scene diffusion model to obtain the target story scene diffusion model.
[0132] Optionally, Figure 8 A flowchart illustrating a model training method provided in this application embodiment. Figure 8 ,like Figure 8 As shown, the process in S704 above, which trains the initial story scene diffusion model based on the appearance description text of the second character in each existing story scene image, the existing event description text associated with the second character, the modified cross-attention weights, and multiple existing story scene images, to obtain the target story scene diffusion model, may include:
[0133] S801. Generate the resulting story scene image based on the appearance description text of the second character in each existing story scene image, the existing event description text associated with the second character, and the modified cross-attention weight.
[0134] In particular, after the modified cross-attention weights are changed, the resulting story scene images determined based on the appearance description text of the second character in each existing story scene image, the existing event description text associated with the second character, and the preset style description text of each existing story scene image will also change.
[0135] Among them, the appearance description text of the second character in each existing story scene image, the existing event description text associated with the second character, and the preset style description text of each existing story scene image are the existing story scene description text of the existing story scene images, denoted as T. g .
[0136] In some implementations, a variational encoder is used to encode existing story scene images to obtain existing story image features f; the existing story image features f are sampled according to the time step t used to obtain noisy existing story image features f. t The existing story scene description text T is generated using a text encoder based on the existing story scene images. g Encode to obtain the text features τ(T) of the existing story scene. g The initial story scene diffusion model is adopted, based on the modified cross-attention weights according to f. t , τ(T g ), t, output predicted noise, and use the predicted noise to denoise the existing story scene image to obtain the result story scene image.
[0137] S802. Using a loss function, update the model parameters of the initial story scene diffusion model based on the resulting story scene image and multiple existing story scene images until the training termination condition is met, thus obtaining the target story scene diffusion model.
[0138] In this embodiment of the application, the noise sampled from the standard normal Gaussian distribution is used as the standard noise ∈; the loss function between the standard noise and the predicted noise is calculated, and the model parameters θ of the initial story scene diffusion model are updated according to the calculation result of the loss function until the training termination condition is met, so as to obtain the target story scene diffusion model.
[0139] The above process can be represented as:
[0140]
[0141] It should be noted that E is a variational encoder, responsible for encoding the existing story scene image F to obtain the existing story image features f, T g It consists of existing story scene description text for each existing story scene image, ∈ is derived from a standard normal Gaussian distribution. The sampled standard noise, t is the sampling time step, ft It is based on the features of an existing story image with added noise after sampling at time step t, where τ is the text encoder, and for T... g After encoding, the existing story scene text features are obtained, ∈ θ This is the initial story scene diffusion model, and θ represents the model parameters that need to be updated. θ With noisy image features f t Text features τ(T) g The noise to be removed at time t is taken as input and the sampling time step t as input. θ (f t ,τ(T g ),t), and calculate the squared error with the previously sampled standard noise ∈, that is, calculate the loss function.
[0142] In practical applications, the training termination condition is met when the loss function converges, or when the number of iterations of the model parameters is greater than or equal to a preset threshold.
[0143] In this embodiment of the application, the cross-attention weights are modified during the model training process, so that when the target story scene diffusion model is applied, the same characters in the new story scene images output by the trained model have visual identity consistency, and multiple characters in a story also have consistency in different new story scene images.
[0144] Optionally, Figure 9 A flowchart illustrating a model training method provided in this application embodiment. Figure 9 ,like Figure 9 As shown, the process in S702 above, which corrects the probability based on the appearance description text of the second character to obtain the corrected probability, may include:
[0145] S901. Using a role knowledge encoder, the appearance description text of the second role is processed to obtain the mean offset and variance correction.
[0146] The appearance description text for the second character can be represented as follows:
[0147] In some implementations, the appearance description text of the second character is input into the character knowledge encoder, which outputs a mean offset and a variance correction. The mean offset can be represented as Δμ, and the variance correction can be represented as γ.
[0148] The above process can be represented as: Φ→{(Δμ)} x ,Δμ y The role knowledge encoder can be represented as Φ, where γ} is the denoted character.
[0149] In this embodiment, the character knowledge encoder includes a text encoder, a first neural network model, and a second neural network model. The text encoder is connected to both the first and second neural network models. The appearance description text of the second character is input into the text encoder, which outputs the appearance text features of the second character. The appearance text features of the second character are input into the first neural network model, which outputs a mean offset. The appearance text features of the second character are input into the second neural network model, which outputs a variance correction.
[0150] Optionally, the first neural network model and the second neural network model can be MLP (Multilayer Perceptron), which is a feedforward artificial neural network model that maps multiple input datasets to a single output dataset.
[0151] S902. Based on the mean offset and variance correction, the probability is corrected to obtain the corrected probability.
[0152] In this embodiment of the application, the above-mentioned correction process can be represented as:
[0153]
[0154] Where, μ x and μ y Let σ represent the mean along the x-axis and y-axis, respectively. 2 It's the variance. Assuming that the characters in each story scene are horizontally and evenly arranged on the image, let μ be... y =0, σ 2 =I. j represents the index of the second character, n is the total number of second characters in the existing story scene images, I represents the identity matrix, and Δμ x and Δμ y γ represents the mean offset along the x-axis and y-axis, and γ is the variance correction.
[0155] After model training, this application was tested on the newly constructed benchmark TBC-Bench (a new story visualization benchmark). The test dataset was constructed as follows: For story visualization of a single character, 5 stories were generated for each character, each story contained 9 story description sentences, and each sentence generated 5 different story scene images. For story generation of multiple characters, 10 stories containing multi-character interactions were generated for each animation set, each story contained 9 story description sentences, and each sentence generated 5 different scenes.
[0156] For a single role, CLIP-I (a metric that evaluates the consistency and quality between the generated image and the description by matching the image with the text description) and DINO-I (an evaluation metric) are used as metrics for role consistency, while CLIP-T (primarily used to evaluate the quality of the image generation model) is used as a metric for text-image semantic consistency. For multiple roles, Frame Accuracy (used to evaluate the accuracy of role generation) and Character F1 (used to evaluate the recognition rate of role generation) are used as metrics for correct role generation, while CLIP-T is used as a metric for text-image semantic consistency. We also include the total number of parameters required for one forward computation of the model.
[0157] It should be noted that the test results for story scene images generated for a single character are shown in Table 1:
[0158] Table 1
[0159]
[0160]
[0161] In addition, the test results for story scene images generated by multiple characters are shown in Table 2:
[0162] Table 2
[0163]
[0164]
[0165] The unit for the number of model parameters is M.
[0166] As shown in Table 1, among single-character story visualization methods, the adapter-based method performs better in terms of textual semantics, but it cannot accurately reproduce the character's ID (IDentity) features. The other two object-customization-based methods can effectively customize the appearance features of a given character, but they suffer from poor textual semantic consistency and require a large number of parameters. The model proposed in this application can achieve uniform high-fidelity and high semantic consistency in story visualization within a single model. In multi-character story visualization, existing methods generally exhibit poor textual semantic consistency and are prone to character confusion. However, the model proposed in this application can correctly generate multiple characters while maintaining high textual semantic information.
[0167] Optionally, Figure 10 This is a flowchart illustrating a scene image generation method provided in an embodiment of this application, as shown below. Figure 10 As shown, the method may include:
[0168] S1001. Obtain new story scene description text.
[0169] It should be noted that the new story scene description text includes: the appearance description text of the third character and the event description text associated with the third character. Alternatively, the new story scene description text includes: the appearance description text of the third character, the event description text associated with the third character, and the style description text.
[0170] S1002. Using the target scene image generation model, generate a new story scene image corresponding to the new story scene description text based on the new story scene description text.
[0171] The target story scene diffusion model is a model trained using the above-mentioned model training method.
[0172] In some implementations, multiple new story scene description texts can be obtained, and these sequentially ordered texts are used to describe a complete storyline. A target scene image generation model can generate a new story scene image based on each new story scene description text. Similarly, multiple sequentially ordered story scene images can be generated based on the sequentially ordered texts. These sequentially ordered story scene images, played frame-by-frame, form a video with a complete storyline. The complete storyline represented by the concatenation of the multiple new story scene images is consistent with the complete storyline described by the multiple new story scene description texts.
[0173] In this embodiment, a new story scene description text is input into a target scene image generation model, which then outputs a new story scene image. This new story scene image better matches the new story scene description text, and the same character can remain consistent within the new story scene image, making the generated new story scene image more accurate and reliable.
[0174] In summary, the embodiments of this application provide a model training method and a scene image generation method. The rich and detailed information such as the appearance description text of the second character and the existing event description text associated with the second character in each existing story scene image participates in the model training. When the target story scene diffusion model is applied, the new story scene image generated is more consistent with the new story scene description text. Moreover, the same character in the new story scene image can remain consistent.
[0175] Furthermore, this application proposes a novel knowledge graph called a Character-Graph (CG), which can structurally represent all knowledge in the story world. By representing each scene in the story world using a Character-Graph, fine-grained visual scene semantics are transformed into detailed text descriptions. Knowledge enhancement using the Character-Graph allows the diffusion model to better model the generation of scene images in the story, achieving consistent generation of text semantics and character IDs. When generating story scenes with multiple characters, knowledge-enhanced spatial guidance is introduced to adjust the incorrectly allocated cross-attention weights within the diffusion model, achieving correct generation of story scenes with multi-character interactions.
[0176] This application also proposes an image generator (StoryWeaver) that uses a character graph for knowledge enhancement, namely a target story scene diffusion model, to achieve high-quality story visualization. Through knowledge enhancement space guidance, it corrects the multi-character identity mixing problem caused by incorrect cross-attention allocation in the initial story scene diffusion model, thereby improving the performance of multi-character story visualization.
[0177] This application also proposes a new story visualization benchmark called TBC-Bench. Experimental comparisons were conducted on this benchmark, and the Story Weaver proposed in this application achieved the best performance in terms of character identity customization, complex scene generation, and text semantic alignment.
[0178] The following describes the model training apparatus, scene image generation apparatus, electronic device and storage medium used to execute the model training apparatus, scene image generation apparatus and other related equipment provided in this application. For the specific implementation process and technical effects, please refer to the relevant content of the above-mentioned model training method and scene image generation apparatus, which will not be repeated below.
[0179] Figure 11 This is a schematic diagram of the structure of a model training device provided in an embodiment of this application, as shown below. Figure 11 As shown, the device includes:
[0180] The acquisition module 101 is used to acquire a character table of a preset story world. The character table includes multiple character images in the preset story world, as well as the appearance attributes of the first character in each character image.
[0181] The determining module 102 is used to determine the appearance description text of the second character and the existing event description text associated with the second character in each of the existing story scene images, based on the multiple existing story scene images of the preset story world and the character table.
[0182] Training module 103 is used to train an initial story scene diffusion model based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, and multiple existing story scene images to obtain a target story scene diffusion model; wherein, the target story scene diffusion model is used to generate a new story scene image corresponding to the new story scene description text based on the new story scene description text.
[0183] Optionally, the acquisition module is specifically used to determine the character description text corresponding to each of the multiple character images; parse the character description text corresponding to each of the character images to obtain multiple appearance attributes of the first character in each of the character images; and construct the character table based on the multiple character images and the multiple appearance attributes of the first character in each of the character images.
[0184] Optionally, the acquisition module 101 is specifically used to construct multiple character attribute networks corresponding to the first character in each of the character images, using the first character in each of the character images as character nodes and each appearance attribute of the first character in each of the character images as attribute nodes; construct a character network of the first character in each of the character images based on the multiple character attribute networks corresponding to the first character in each of the character images; and construct the character table based on the character network of the first character in each of the character images, using multiple character images as indexes.
[0185] Optionally, the determining module 102 is specifically configured to: determine the description information of each of the existing story scene images based on the plurality of existing story scene images; parse the description information of each of the existing story scene images to obtain the description information of the second character in each of the existing story scene images, and the events related to the second character; and determine the appearance description text of the second character in each of the existing story scene images and the description text of the existing events associated with the second character based on the description information of the second character in each of the existing story scene images, the events related to the second character, and the character table.
[0186] Optionally, the determining module 102 is specifically configured to: search for the name of the first character in a target character image that matches the description information of the second character in the existing story scene image from multiple character images in the character table; determine the appearance description text of the second character in each of the existing story scene images based on the name of the first character in the target character image and the character table; and determine the existing event description text associated with the second character based on the name of the first character in the target character image, the events related to the second character, and the relationship between the first characters in the target character image.
[0187] Optionally, the determining module is specifically configured to determine the appearance attributes of the first character in the target character image based on the name of the first character in the target character image and the character table; and to determine the appearance description text of the second character in each of the existing story scene images based on the name of the target character image and the appearance attributes of the first character in the target character image.
[0188] Optionally, the training module 103 is used to train the initial story scene diffusion model based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, the preset style description text of each of the existing story scene images, and multiple existing story scene images, to obtain the target story scene diffusion model.
[0189] Optionally, the training module 103 is used to calculate the probability of the second character appearing in any pixel in each of the existing story scene images; to correct the probability based on the appearance description text of the second character, to obtain a corrected probability; to modify the cross-attention weights in the initial story scene diffusion model based on the corrected probability, to obtain modified cross-attention weights; and to train the initial story scene diffusion model based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, the modified cross-attention weights, and multiple existing story scene images, to obtain the target story scene diffusion model.
[0190] Optionally, the training module 103 is used to generate a result story scene image based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, and the modified cross-attention weights; and to update the model parameters of the initial story scene diffusion model using a loss function based on the result story scene image and multiple existing story scene images until the training termination condition is met, thereby obtaining the target story scene diffusion model.
[0191] Optionally, the training module 103 is used to process the appearance description text of the second character using a character knowledge encoder to obtain a mean offset and a variance correction; and to correct the probability based on the mean offset and the variance correction to obtain the corrected probability.
[0192] Figure 12 This is a schematic diagram of the structure of a scene image generation device provided in an embodiment of this application, as shown below. Figure 12 As shown, the device includes:
[0193] Module 201 is used to acquire new story scene description text;
[0194] The generation module 202 is used to generate a new story scene image corresponding to the new story scene description text based on the new story scene description text using a target scene image generation model, wherein the target story scene diffusion model is a model trained using the above-mentioned model training method.
[0195] Optionally, the new story scene description text includes: appearance description text of the third character, event description text associated with the third character, and style description text.
[0196] The above-described device is used to execute the method provided in the foregoing embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.
[0197] These modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more digital signal processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). Alternatively, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a system-on-a-chip (SOC).
[0198] Figure 13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, such as... Figure 13 As shown, the electronic device includes: a processor 301 and a memory 302.
[0199] The memory 302 is used to store programs, and the processor 301 calls the programs stored in the memory 302 to execute the above-described model training method embodiment and scene image generation method embodiment. The specific implementation methods and technical effects are similar, and will not be described in detail here.
[0200] For example, the model training method may include:
[0201] Obtain a character list for a preset story world, the character list including multiple character images in the preset story world, and the appearance attributes of a first character in each character image;
[0202] Based on multiple existing story scene images of the preset story world and the character list, determine the appearance description text of the second character and the existing event description text associated with the second character in each of the existing story scene images;
[0203] Based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, and multiple existing story scene images, an initial story scene diffusion model is trained to obtain a target story scene diffusion model; wherein, the target story scene diffusion model is used to generate a new story scene image corresponding to the new story scene description text based on the new story scene description text.
[0204] Optionally, obtaining the character list of the preset story world includes:
[0205] Based on the multiple character images, determine the character description text corresponding to each character image;
[0206] The character description text corresponding to each of the character images is parsed to obtain multiple appearance attributes of the first character in each of the character images;
[0207] The character table is constructed based on multiple character images and multiple appearance attributes of the first character in each character image.
[0208] Optionally, constructing the character table based on multiple character images and multiple appearance attributes of the first character in each character image includes:
[0209] Using the first character in each of the character images as the character node and each appearance attribute of the first character in each of the character images as the attribute node, construct multiple character attribute networks corresponding to the first character in each of the character images;
[0210] Based on the multiple character attribute networks corresponding to the first character in each of the character images, construct the character network of the first character in each of the character images;
[0211] Using multiple character images as indexes, and based on the character network of the first character in each character image, the character table is constructed.
[0212] Optionally, determining the appearance description text of the second character and the existing event description text associated with the second character in each of the existing story scene images based on the multiple existing story scene images of the preset story world and the character table includes:
[0213] Based on the multiple existing story scene images, determine the descriptive information of each existing story scene image;
[0214] The description information of each of the existing story scene images is parsed to obtain the description information of the second character in each of the existing story scene images, as well as the events related to the second character;
[0215] Based on the description information of the second character in each of the existing story scene images, the events related to the second character, and the character table, determine the appearance description text of the second character in each of the existing story scene images and the description text of the existing events associated with the second character.
[0216] Optionally, determining the appearance description text of the second character and the existing event description text associated with the second character in each of the existing story scene images, based on the description information of the second character in each of the existing story scene images, the events related to the second character, and the character table, includes:
[0217] From the multiple character images in the character list, find the name of the first character in the target character image that matches the description information of the second character in the existing story scene image;
[0218] Based on the name of the first character in the target character image and the character table, determine the appearance description text of the second character in each of the existing story scene images;
[0219] Based on the name of the first character in the target character image, the events related to the second character, and the relationship between the first character in the target character image, determine the existing event description text associated with the second character.
[0220] Optionally, determining the appearance description text of the second character in each of the existing story scene images based on the name of the first character in the target character image and the character table includes:
[0221] Based on the name of the first character in the target character image and the character table, determine the appearance attributes of the first character in the target character image;
[0222] Based on the name of the target character image and the appearance attributes of the first character in the target character image, determine the appearance description text of the second character in each of the existing story scene images.
[0223] Optionally, the step of training an initial story scene diffusion model to obtain a target story scene diffusion model based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, and multiple existing story scene images includes:
[0224] The initial story scene diffusion model is trained based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, the preset style description text of each of the existing story scene images, and multiple existing story scene images to obtain the target story scene diffusion model.
[0225] Optionally, the step of training an initial story scene diffusion model to obtain a target story scene diffusion model based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, and multiple existing story scene images includes:
[0226] Calculate the probability of the second character appearing in any pixel in each of the existing story scene images;
[0227] Based on the appearance description text of the second character, the probability is corrected to obtain the corrected probability;
[0228] Based on the corrected probabilities, the cross-attention weights in the initial story scene diffusion model are modified to obtain the modified cross-attention weights;
[0229] Based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, the modified cross-attention weights, and multiple existing story scene images, the initial story scene diffusion model is trained to obtain the target story scene diffusion model.
[0230] Optionally, the step of training the initial story scene diffusion model based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, the modified cross-attention weights, and multiple existing story scene images to obtain the target story scene diffusion model includes:
[0231] Based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, and the modified cross-attention weight, a result story scene image is generated;
[0232] Using a loss function, the model parameters of the initial story scene diffusion model are updated based on the resulting story scene image and multiple existing story scene images until the training termination condition is met, thus obtaining the target story scene diffusion model.
[0233] Optionally, the step of correcting the probability based on the appearance description text of the second character to obtain a corrected probability includes:
[0234] A character knowledge encoder is used to process the appearance description text of the second character to obtain the mean offset and variance correction.
[0235] The probability is corrected based on the mean offset and the variance correction to obtain the corrected probability.
[0236] For example, the scene image generation method may include:
[0237] Obtain new story scene description text;
[0238] A target scene image generation model is used to generate a new story scene image corresponding to the new story scene description text based on the new story scene description text, wherein the target story scene diffusion model is a model trained using the above-mentioned model training method.
[0239] Optionally, the new story scene description text includes: appearance description text of the third character, event description text associated with the third character, and style description text.
[0240] In summary, the embodiments of this application provide a model training method and a scene image generation method. The rich and detailed information such as the appearance description text of the second character and the existing event description text associated with the second character in each existing story scene image participates in the model training. When the target story scene diffusion model is applied, the new story scene image generated is more consistent with the new story scene description text. Moreover, the same character in the new story scene image can remain consistent.
[0241] Optionally, this application also provides a program product, such as a computer-readable storage medium, including a program that, when executed by a processor, performs the above-described method embodiments.
[0242] For example, the model training method may include:
[0243] Obtain a character list for a preset story world, the character list including multiple character images in the preset story world, and the appearance attributes of a first character in each character image;
[0244] Based on multiple existing story scene images of the preset story world and the character list, determine the appearance description text of the second character and the existing event description text associated with the second character in each of the existing story scene images;
[0245] Based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, and multiple existing story scene images, an initial story scene diffusion model is trained to obtain a target story scene diffusion model; wherein, the target story scene diffusion model is used to generate a new story scene image corresponding to the new story scene description text based on the new story scene description text.
[0246] Optionally, obtaining the character list of the preset story world includes:
[0247] Based on the multiple character images, determine the character description text corresponding to each character image;
[0248] The character description text corresponding to each of the character images is parsed to obtain multiple appearance attributes of the first character in each of the character images;
[0249] The character table is constructed based on multiple character images and multiple appearance attributes of the first character in each character image.
[0250] Optionally, constructing the character table based on multiple character images and multiple appearance attributes of the first character in each character image includes:
[0251] Using the first character in each of the character images as the character node and each appearance attribute of the first character in each of the character images as the attribute node, construct multiple character attribute networks corresponding to the first character in each of the character images;
[0252] Based on the multiple character attribute networks corresponding to the first character in each of the character images, construct the character network of the first character in each of the character images;
[0253] Using multiple character images as indexes, and based on the character network of the first character in each character image, the character table is constructed.
[0254] Optionally, determining the appearance description text of the second character and the existing event description text associated with the second character in each of the existing story scene images based on the multiple existing story scene images of the preset story world and the character table includes:
[0255] Based on the multiple existing story scene images, determine the descriptive information of each existing story scene image;
[0256] The description information of each of the existing story scene images is parsed to obtain the description information of the second character in each of the existing story scene images, as well as the events related to the second character;
[0257] Based on the description information of the second character in each of the existing story scene images, the events related to the second character, and the character table, determine the appearance description text of the second character in each of the existing story scene images and the description text of the existing events associated with the second character.
[0258] Optionally, determining the appearance description text of the second character and the existing event description text associated with the second character in each of the existing story scene images, based on the description information of the second character in each of the existing story scene images, the events related to the second character, and the character table, includes:
[0259] From the multiple character images in the character list, find the name of the first character in the target character image that matches the description information of the second character in the existing story scene image;
[0260] Based on the name of the first character in the target character image and the character table, determine the appearance description text of the second character in each of the existing story scene images;
[0261] Based on the name of the first character in the target character image, the events related to the second character, and the relationship between the first character in the target character image, determine the existing event description text associated with the second character.
[0262] Optionally, determining the appearance description text of the second character in each of the existing story scene images based on the name of the first character in the target character image and the character table includes:
[0263] Based on the name of the first character in the target character image and the character table, determine the appearance attributes of the first character in the target character image;
[0264] Based on the name of the target character image and the appearance attributes of the first character in the target character image, determine the appearance description text of the second character in each of the existing story scene images.
[0265] Optionally, the step of training an initial story scene diffusion model to obtain a target story scene diffusion model based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, and multiple existing story scene images includes:
[0266] The initial story scene diffusion model is trained based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, the preset style description text of each of the existing story scene images, and multiple existing story scene images to obtain the target story scene diffusion model.
[0267] Optionally, the step of training an initial story scene diffusion model to obtain a target story scene diffusion model based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, and multiple existing story scene images includes:
[0268] Calculate the probability of the second character appearing in any pixel in each of the existing story scene images;
[0269] Based on the appearance description text of the second character, the probability is corrected to obtain the corrected probability;
[0270] Based on the corrected probabilities, the cross-attention weights in the initial story scene diffusion model are modified to obtain the modified cross-attention weights;
[0271] Based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, the modified cross-attention weights, and multiple existing story scene images, the initial story scene diffusion model is trained to obtain the target story scene diffusion model.
[0272] Optionally, the step of training the initial story scene diffusion model based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, the modified cross-attention weights, and multiple existing story scene images to obtain the target story scene diffusion model includes:
[0273] Based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, and the modified cross-attention weight, a result story scene image is generated;
[0274] Using a loss function, the model parameters of the initial story scene diffusion model are updated based on the resulting story scene image and multiple existing story scene images until the training termination condition is met, thus obtaining the target story scene diffusion model.
[0275] Optionally, the step of correcting the probability based on the appearance description text of the second character to obtain a corrected probability includes:
[0276] A character knowledge encoder is used to process the appearance description text of the second character to obtain the mean offset and variance correction.
[0277] The probability is corrected based on the mean offset and the variance correction to obtain the corrected probability.
[0278] For example, the scene image generation method may include:
[0279] Obtain new story scene description text;
[0280] A target scene image generation model is used to generate a new story scene image corresponding to the new story scene description text based on the new story scene description text, wherein the target story scene diffusion model is a model trained using the above-mentioned model training method.
[0281] Optionally, the new story scene description text includes: appearance description text of the third character, event description text associated with the third character, and style description text.
[0282] In summary, the embodiments of this application provide a model training method and a scene image generation method. The rich and detailed information such as the appearance description text of the second character and the existing event description text associated with the second character in each existing story scene image participates in the model training. When the target story scene diffusion model is applied, the new story scene image generated is more consistent with the new story scene description text. Moreover, the same character in the new story scene image can remain consistent.
[0283] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0284] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0285] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in a combination of hardware and software functional units.
[0286] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0287] The above are merely preferred embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A model training method, characterized in that, include: Obtain a character list for a preset story world, the character list including multiple character images in the preset story world, and the appearance attributes of a first character in each character image; Based on multiple existing story scene images of the preset story world and the character list, determine the appearance description text of the second character and the existing event description text associated with the second character in each of the existing story scene images; Based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, and multiple existing story scene images, an initial story scene diffusion model is trained to obtain a target story scene diffusion model; wherein, the target story scene diffusion model is used to generate a new story scene image corresponding to the new story scene description text based on the new story scene description text.
2. The method according to claim 1, characterized in that, The character list for obtaining the preset story world includes: Based on the multiple character images, determine the character description text corresponding to each character image; The character description text corresponding to each of the character images is parsed to obtain multiple appearance attributes of the first character in each of the character images; The character table is constructed based on multiple character images and multiple appearance attributes of the first character in each character image.
3. The method according to claim 2, characterized in that, The step of constructing the character table based on multiple character images and multiple appearance attributes of the first character in each character image includes: Using the first character in each of the character images as the character node and each appearance attribute of the first character in each of the character images as the attribute node, construct multiple character attribute networks corresponding to the first character in each of the character images; Based on the multiple character attribute networks corresponding to the first character in each of the character images, construct the character network of the first character in each of the character images; Using multiple character images as indexes, and based on the character network of the first character in each character image, the character table is constructed.
4. The method according to claim 1, characterized in that, The step of determining the appearance description text of the second character and the existing event description text associated with the second character in each of the existing story scene images, based on multiple existing story scene images of the preset story world and the character table, includes: Based on the multiple existing story scene images, determine the descriptive information of each existing story scene image; The description information of each of the existing story scene images is parsed to obtain the description information of the second character in each of the existing story scene images, as well as the events related to the second character; Based on the description information of the second character in each of the existing story scene images, the events related to the second character, and the character table, determine the appearance description text of the second character in each of the existing story scene images and the description text of the existing events associated with the second character.
5. The method according to claim 4, characterized in that, The step of determining the appearance description text of the second character and the existing event description text associated with the second character in each of the existing story scene images, based on the description information of the second character in each of the existing story scene images, the events related to the second character, and the character table, includes: From the multiple character images in the character list, find the name of the first character in the target character image that matches the description information of the second character in the existing story scene image; Based on the name of the first character in the target character image and the character table, determine the appearance description text of the second character in each of the existing story scene images; Based on the name of the first character in the target character image, the events related to the second character, and the relationship between the first character in the target character image, determine the existing event description text associated with the second character.
6. The method according to claim 5, characterized in that, The step of determining the appearance description text of the second character in each of the existing story scene images based on the name of the first character in the target character image and the character table includes: Based on the name of the first character in the target character image and the character table, determine the appearance attributes of the first character in the target character image; Based on the name of the target character image and the appearance attributes of the first character in the target character image, determine the appearance description text of the second character in each of the existing story scene images.
7. The method according to claim 1, characterized in that, The step of training an initial story scene diffusion model based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, and multiple existing story scene images to obtain a target story scene diffusion model includes: The initial story scene diffusion model is trained based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, the preset style description text of each of the existing story scene images, and multiple existing story scene images to obtain the target story scene diffusion model.
8. The method according to claim 1, characterized in that, The step of training an initial story scene diffusion model based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, and multiple existing story scene images to obtain a target story scene diffusion model includes: Calculate the probability of the second character appearing in any pixel in each of the existing story scene images; Based on the appearance description text of the second character, the probability is corrected to obtain the corrected probability; Based on the corrected probabilities, the cross-attention weights in the initial story scene diffusion model are modified to obtain the modified cross-attention weights; Based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, the modified cross-attention weights, and multiple existing story scene images, the initial story scene diffusion model is trained to obtain the target story scene diffusion model.
9. The method according to claim 8, characterized in that, The initial story scene diffusion model is trained based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, the modified cross-attention weights, and multiple existing story scene images to obtain the target story scene diffusion model, including: Based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, and the modified cross-attention weight, a result story scene image is generated; Using a loss function, the model parameters of the initial story scene diffusion model are updated based on the resulting story scene image and multiple existing story scene images until the training termination condition is met, thus obtaining the target story scene diffusion model.
10. The method according to claim 8, characterized in that, The step of correcting the probability based on the appearance description text of the second character to obtain the corrected probability includes: A character knowledge encoder is used to process the appearance description text of the second character to obtain the mean offset and variance correction. The probability is corrected based on the mean offset and the variance correction to obtain the corrected probability.
11. A method for generating scene images, characterized in that, include: Obtain new story scene description text; A target scene image generation model is used to generate a new story scene image corresponding to the new story scene description text based on the new story scene description text, wherein the target story scene diffusion model is a model trained using the method described in any one of claims 1-10.
12. The method according to claim 11, characterized in that, The new story scene description text includes: appearance description text of the third character, event description text associated with the third character, and style description text.
13. A model training device, characterized in that, include: The acquisition module is used to acquire a character table of a preset story world. The character table includes multiple character images in the preset story world, as well as the appearance attributes of the first character in each character image. The determination module is used to determine the appearance description text of the second character and the existing event description text associated with the second character in each of the existing story scene images, based on multiple existing story scene images of the preset story world and the character table. The training module is used to train an initial story scene diffusion model based on the appearance description text of the second character in each of the existing story scene images, the existing event description text associated with the second character, and multiple existing story scene images, to obtain a target story scene diffusion model; wherein, the target story scene diffusion model is used to generate a new story scene image corresponding to the new story scene description text based on the new story scene description text.
14. A scene image generation apparatus, characterized in that, include: The acquisition module is used to acquire new story scene description text; A generation module is used to generate a new story scene image corresponding to the new story scene description text based on the new story scene description text using a target scene image generation model, wherein the target story scene diffusion model is a model trained using the method described in any one of claims 1-10.
15. An electronic device, characterized in that, include: The memory and the processor, wherein the memory stores a computer program executable by the processor, and the processor executes the computer program to implement the model training method according to any one of claims 1-10, or the scene image generation method according to any one of claims 11-12.
16. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when read and executed, implements the model training method according to any one of claims 1-10, or the scene image generation method according to any one of claims 11-12.