Image generation method, device and non-volatile computer readable storage medium
By receiving user-input attribute information and preset object templates, and combining them with a machine learning model, images of virtual objects are automatically generated, solving the problem of low image generation efficiency in existing technologies and achieving efficient and high-quality image generation.
Patent Information
- Application Number
- CN202311597187.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2043-11-27
AI Technical Summary
Current technologies have low image generation efficiency and poor image quality, while incurring high time and labor costs.
By receiving virtual object attribute information input by the user, and combining it with preset object templates and machine learning models, the system automatically generates relevant images of the virtual objects. It also uses object template priority and masked attribute information to handle conflicts and generate high-quality images.
It improves the efficiency and quality of image generation, reduces time and labor costs, and generates high-quality images that meet user needs.
Smart Images

Figure CN117437320B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to an image generation method, an image generation apparatus, and a non-volatile computer-readable storage medium. Background Technology
[0002] In computer applications such as games and interactive stories, it is often necessary to use art resources such as images to visually represent virtual objects such as characters and scenes. Examples include character portraits and background images.
[0003] In related technologies, producers of interactive multimedia content need to provide professional illustrators with specific image generation requirements, relying on manual drawing of corresponding art resources. Summary of the Invention
[0004] The inventors of this disclosure have discovered the following problems in the aforementioned related technologies: high time and labor costs lead to a decrease in image generation efficiency and generation effect.
[0005] In view of this, this disclosure proposes an image generation technology solution that can improve image generation efficiency and generation effect.
[0006] According to some embodiments of this disclosure, an image generation method is provided, including: receiving a first set of attribute information of a virtual object input by a user; generating a related image of the virtual object according to the first set of attribute information and an object template, wherein the object template is used to determine at least one of the pose information and composition information of the virtual object in the related image, and the object template is matched with the first set of attribute information.
[0007] In some embodiments, if there is a conflict between the first attribute information set and the object template, the object template has a higher priority than the first attribute information set.
[0008] In some embodiments, generating a related image of a virtual object based on a first set of attribute information and an object template includes: masking attribute information that conflicts with the object template in the first set of attribute information; and generating a related image based on the masking result and the object template.
[0009] In some embodiments, generating a related image of a virtual object based on a first set of attribute information and an object template includes: generating a related image based on the first attribute information, the object template, and a stored second set of attribute information, wherein the second set of attribute information is used to indicate the specified requirements of the related image, and in the event of a conflict between the first set of attribute information and the second set of attribute information, the second set of attribute information has a higher priority than the first set of attribute information.
[0010] In some embodiments, the second attribute information set includes negative attribute information, and generating a related image based on the stored second attribute information set includes: if there is attribute information in the first attribute information set that conflicts with the negative attribute information, masking the attribute information; and generating a related image based on the masking result and the second attribute information set.
[0011] In some embodiments, the second set of attribute information includes positive attribute information used to improve the image quality of the relevant images.
[0012] In some embodiments, the second set of attribute information is not visible to the user.
[0013] In some embodiments, the second set of attribute information is used to indicate the image style type of the relevant image.
[0014] In some embodiments, generating a related image of a virtual object based on a first set of attribute information and an object template includes generating a related image based on an image style type selected by the user.
[0015] In some embodiments, generating a related image of a virtual object based on a first set of attribute information and an object template includes: generating multiple candidate images based on the first set of attribute information, wherein the resolution of the multiple candidate images is lower than a threshold; and, in response to a user's selection operation among the multiple candidate images, performing a high-resolution operation on the selected candidate image to generate a related image, wherein the resolution of the related image is higher than that of the multiple candidate images.
[0016] In some embodiments, the image generation method further includes: determining the dynamic information of a virtual object in a related image based on content information corresponding to the related image, wherein the virtual object includes a human object; and generating a dynamic image of the human object based on the related image using a digital human model based on the dynamic information.
[0017] In some embodiments, generating a dynamic image of a person object using a digital human model based on relevant images according to dynamic information includes: adjusting the facial angle of the person object in the relevant images according to dynamic information; and generating the dynamic image based on the adjusted relevant images.
[0018] In some embodiments, the virtual object includes a person object, and the pose information is used to indicate that the face of the person object in the relevant image is facing the user and the eyes of the person object are facing the user.
[0019] In some embodiments, generating a related image of a virtual object based on a first set of attribute information and an object template includes: in response to a user's selection operation among multiple candidate object templates, determining an object template from the multiple candidate object templates, wherein the multiple candidate object templates match the first set of attribute information.
[0020] In some embodiments, the relevant image corresponds to the current frame image of the interactive multimedia content, and the object template is selected based on the content to be represented by the current frame image.
[0021] In some embodiments, the image generation method further includes: when the virtual object includes a person object, identifying the facial region of the person object in the relevant image; and segmenting the relevant image based on the facial region to generate a portrait of the person object.
[0022] In some embodiments, the portion of the related image other than the virtual object is transparent, and the image generation method further includes: fusing the related image with the background image to generate a frame image in interactive multimedia content.
[0023] In some embodiments, the image generation method further includes: segmenting the image region where the virtual object is located from the relevant image; and fusing the image region with the background image to generate a frame image in interactive multimedia content.
[0024] In some embodiments, the virtual object includes a character object and a background object, and the composition information includes at least one of the virtual object's position information and size information in the relevant image.
[0025] According to some other embodiments of this disclosure, an image generation apparatus is provided, comprising: a receiving unit for receiving a first set of attribute information of a virtual object input by a user; and a generation unit for generating a related image of the virtual object based on the first set of attribute information and an object template, wherein the object template is used to determine at least one of the pose information and composition information of the virtual object in the related image, and the object template is matched with the first set of attribute information.
[0026] In some embodiments, if there is a conflict between the first attribute information set and the object template, the object template has a higher priority than the first attribute information set.
[0027] In some embodiments, if the generation unit has attribute information in the first attribute information set that conflicts with the object template, it masks the attribute information; and generates a relevant image based on the masking result and the object template.
[0028] In some embodiments, the generation unit generates a related image based on the first attribute information, the object template, and the stored second attribute information set. The second attribute information set is used to indicate the specified requirements of the related image. In the event of a conflict between the first attribute information set and the second attribute information set, the second attribute information set has a higher priority than the first attribute information set.
[0029] In some embodiments, the second attribute information set includes negative attribute information. If the generation unit has attribute information in the first attribute information set that conflicts with the negative attribute information, it masks the attribute information and generates a relevant image based on the masking result and the second attribute information set.
[0030] In some embodiments, the second set of attribute information includes positive attribute information used to improve the image quality of the relevant images.
[0031] In some embodiments, the second set of attribute information is not visible to the user.
[0032] In some embodiments, the second set of attribute information is used to indicate the image style type of the relevant image.
[0033] In some embodiments, the generation unit generates a relevant image based on the image style type selected by the user.
[0034] In some embodiments, the generating unit generates multiple candidate images based on a first set of attribute information. The resolution of the multiple candidate images is lower than a threshold. In response to a user's selection operation among the multiple candidate images, a high-resolution operation is performed on the selected candidate image to generate a related image. The resolution of the related image is higher than that of the multiple candidate images.
[0035] In some embodiments, the generation unit determines the dynamic information of the virtual object in the relevant image based on the content information corresponding to the relevant image. The virtual object includes a human figure. Based on the dynamic information, a dynamic image of the human figure is generated using a digital human model on the basis of the relevant image.
[0036] In some embodiments, the generation unit adjusts the facial angle of a person in a related image based on dynamic information; and generates a dynamic image based on the adjusted related image.
[0037] In some embodiments, the virtual object includes a person object, and the pose information is used to indicate that the face of the person object in the relevant image is facing the user and the eyes of the person object are facing the user.
[0038] In some embodiments, the generation unit responds to a user's selection operation among multiple candidate object templates, determines an object template from the multiple candidate object templates, and the multiple candidate object templates match a first set of attribute information.
[0039] In some embodiments, the relevant image corresponds to the current frame image of the interactive multimedia content, and the object template is selected based on the content to be represented by the current frame image.
[0040] In some embodiments, when the virtual object includes a person object, the generation unit identifies the facial region of the person object in the relevant image, and segments the relevant image based on the facial region to generate the person object's avatar.
[0041] In some embodiments, the portion of the related image other than the virtual object is transparent, and the generation unit merges the related image with the background image to generate a frame image in the interactive multimedia content.
[0042] In some embodiments, the generation unit segments the image region where the virtual object is located from the relevant image, and merges the image region with the background image to generate a frame image in the interactive multimedia content.
[0043] In some embodiments, the virtual object includes a character object and a background object, and the composition information includes at least one of the virtual object's position information and size information in the relevant image.
[0044] According to further embodiments of the present disclosure, an image generation apparatus is provided, including: a memory; and a processor coupled to the memory, the processor being configured to execute the image generation method of any of the above embodiments based on instructions stored in the memory device.
[0045] According to further embodiments of the present disclosure, a non-volatile computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the image generation method of any of the above embodiments.
[0046] In the above embodiments, by combining user-input attribute information and preset object templates, relevant images of virtual objects are automatically generated. This reduces time and labor costs, thereby improving image generation efficiency and quality. Attached Figure Description
[0047] The accompanying drawings, which form part of this specification, illustrate embodiments of this disclosure and, together with the specification, serve to explain the principles of this disclosure.
[0048] This disclosure will become clearer with reference to the accompanying drawings and the following detailed description:
[0049] Figure 1 Flowcharts illustrating some embodiments of the image generation method of this disclosure;
[0050] Figure 2 Flowcharts illustrating other embodiments of the image generation method of this disclosure;
[0051] Figure 3a Schematic diagrams illustrating some embodiments of the image generation method for human objects disclosed herein;
[0052] Figure 3b Schematic diagrams illustrating some embodiments of the image generation method for background objects of this disclosure;
[0053] Figure 3c Schematic diagrams illustrating some embodiments of the candidate image generation method of this disclosure;
[0054] Figure 4 Block diagrams showing some embodiments of the image generation apparatus of this disclosure;
[0055] Figure 5 Block diagrams showing other embodiments of the image generation apparatus of this disclosure;
[0056] Figure 6 Block diagrams illustrating further embodiments of the image generation apparatus of this disclosure are shown. Detailed Implementation
[0057] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.
[0058] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.
[0059] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.
[0060] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, they should be considered part of the specification.
[0061] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.
[0062] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0063] As mentioned earlier, manually drawing art resources for interactive multimedia content is time-consuming and labor-intensive, resulting in high time and labor costs, which reduces image generation efficiency and quality.
[0064] To address the aforementioned technical issues, this disclosure utilizes preset object templates to define preset character poses and image compositions. Combined with simple descriptive information provided by the user in the input box and preset prompts, an image is automatically generated through a machine learning model. Users can also select a desired image from multiple generated images for high-resolution processing to generate character or background images suitable for interactive multimedia content.
[0065] In this way, by combining user-input attribute information and preset object templates, high-quality images of virtual objects can be generated quickly, thereby improving image generation efficiency and quality. For example, the technical solution of this disclosure can be implemented through the following embodiments.
[0066] For example, the technical solution of this disclosure can be implemented through the following embodiments.
[0067] Figure 1 Flowcharts illustrating some embodiments of the image generation method of this disclosure are shown.
[0068] like Figure 1 As shown, in step 110, the first set of attribute information of the virtual object input by the user is received.
[0069] For example, virtual objects may include character objects, background objects, etc. in multimedia interactive content; multimedia interactive content may include games, interactive stories, interactive comics, and interactive dramas, etc.; the first attribute information set may include multiple first attribute information used to describe the characteristics of virtual objects.
[0070] In step 120, a relevant image of the virtual object is generated based on the first attribute information set and the object template. The object template is used to determine at least one of the pose information and composition information of the virtual object in the relevant image, and the object template is matched with the first attribute information set.
[0071] In the above embodiments, by combining user-input attribute information and preset object templates, relevant images of virtual objects are automatically generated. This reduces time and labor costs, thereby improving image generation efficiency and quality.
[0072] The following examples illustrate how to automatically generate images related to virtual objects.
[0073] Figure 2 Flowcharts illustrating some other embodiments of the image generation method of this disclosure are shown.
[0074] like Figure 2As shown, the first set of attribute information input by the user can be used as a Prompt message and input into a machine learning model to automatically generate relevant images of virtual objects. For example, the first set of attribute information can be masked in steps 1210-1230, and then the relevant images of virtual objects can be automatically generated using a machine learning model.
[0075] For example, machine learning models can be used for models like LLM (Large Language Model) that can generate image content based on user-provided prompt information. This allows for improvements in image generation efficiency and effectiveness.
[0076] In some embodiments, if there is a conflict between the first attribute information set and the object template, the object template has a higher priority than the first attribute information set. For example, in step 1210, if there is first attribute information in the first attribute information set that conflicts with the object template, this first attribute information is masked.
[0077] For example, when the virtual object is a character object, the object template can be used to determine the character object's posture information, etc. For instance, the first attribute information set includes the first attribute information "character jumps", but the posture information defined by the object template includes "character stands". In this case, "character jumps" is hidden in the Prompt information, and "character stands" is added after the Prompt information and then input into the machine learning model.
[0078] For example, when the virtual object is a background object, the object template can be used to determine the composition information of the background object. The composition information may include at least one of the virtual object's position and size information in the relevant image. For example, if the first attribute information set includes the first attribute information "the mountain peak is located in the center of the image", but the composition information defined by the object template includes "the mountain peak is located in the upper left of the image", then "the mountain peak is located in the center of the image" is masked in the Prompt information, and "the mountain peak is located in the upper left of the image" is added after the Prompt information before being input into the machine learning model.
[0079] In some embodiments, the object template is matched with the first set of attribute information. For example, the type and art style of the virtual object corresponding to the object template can be matched with the type and art style of the virtual object corresponding to the first set of attribute information.
[0080] For example, if the first set of attribute information describes a character object, then the object template is a template for the character object; if the first set of attribute information describes a background object, then the object template is a template for the background object. For example, if the character object described by the first set of attribute information is an ancient character, then the style of the object template is ancient style; if the character object described by the first set of attribute information is a science fiction character, then the style of the object template is science fiction style.
[0081] In some embodiments, a related image is generated based on first attribute information, an object template, and a stored set of second attribute information. For example, the second attribute information set is used to indicate specific requirements for the related image. In the event of a conflict between the first and second attribute information sets, the second attribute information set takes precedence over the first attribute information set.
[0082] For example, in step 1220, if there is first attribute information in the first attribute information set that conflicts with negative attribute information, the first attribute information is masked.
[0083] In some embodiments, the second set of attribute information includes negative attribute information. For example, negative attribute information is prompt information that needs to be masked, which may include prompt information such as "bad hands" or "exposed body" that does not meet the specified requirements for image generation.
[0084] For example, the first attribute information set includes the first attribute information "unusual hand shape" which conflicts with the negative prompt word "bad hand shape" that needs to be masked; in this case, "unusual hand shape" in the Prompt information is masked before being input into the machine learning model.
[0085] In some embodiments, the second set of attribute information includes positive attribute information used to improve the image quality of the relevant images. For example, positive attribute information may include terms such as "high quality" or "excellent," and this positive attribute information, along with the user-provided first attribute information, can be used as prompt information input into a machine learning model.
[0086] In some embodiments, the second attribute information set is invisible to the user. For example, the second attribute information set includes negative attribute information and Prompt information indicating the image style type of the relevant image (such as cyberpunk style, Chinese martial arts style, modern urban style, etc.). This information is hidden in the background, invisible to the user and requires no user intervention, thereby improving the efficiency and quality of image generation.
[0087] In some embodiments, relevant images can also be generated based on the image style type selected by the user. For example, if the user selects a modern urban style, a relevant image can be generated based on the object template corresponding to the modern urban style. In this way, matching images can be generated according to the user's actual needs, thereby improving the efficiency and quality of image generation.
[0088] In some embodiments, steps 1210 and 1220 described above may not be executed in any particular order, or they may be executed in parallel. For example, after the masking process of steps 1210 and 1220, the prompt information can be determined as "a smiling man wearing glasses and modern clothing"; this prompt information can be input into a machine learning model to generate a relevant image.
[0089] For example, in step 1230, a relevant image is generated based on the masking result and the object template. The masked Prompt information and its matching object template can be input into the machine learning model; the machine learning model generates the relevant image based on the Prompt information and its matching object template.
[0090] The following is through Figures 3a-3c Some embodiments of the present invention provide exemplary illustrations of how to automatically generate images related to virtual objects.
[0091] Figure 3a Schematic diagrams illustrating some embodiments of the image generation method for human objects disclosed herein.
[0092] like Figure 3a As shown, for example, if a user wants to generate an image of a virtual character, they can input a set of first attribute information 31a containing multiple first attribute information through an input box on the page. For example, the first attribute information set 31a input by the user includes "a smiling man wearing glasses, dressed in modern clothing, with four hands, leaping."
[0093] For example, if the type of the virtual object is determined to be a character object based on the first attribute information set 31a input by the user, then the corresponding character object template 32a can be automatically matched as the current object template; if the image style type is determined to be modern urban style based on the first attribute information set 31a input by the user, then the corresponding modern urban style object template 32a can be automatically matched as the current object template.
[0094] In some embodiments, in response to a user's selection operation among multiple candidate object templates, an object template is determined from the multiple candidate object templates, and the multiple candidate object templates are matched with a first set of attribute information. For example, multiple image style types of templates can be provided to the user for selection; alternatively, the image style type can be automatically determined using a machine learning model based on the historical context and related plot of the interactive multimedia content in which the virtual character object is located, thereby improving image generation efficiency.
[0095] For example, object template 32a limits the posture information of the character object to "character standing", which conflicts with "jumping" in the first attribute information set 31a; in this case, "jumping" is hidden in the Prompt information and "character standing" is added to the Prompt information.
[0096] In this way, the prompts for user input can be automatically and effectively controlled through pre-made object templates, thereby improving the quality of image generation.
[0097] For example, the first attribute information "four hands" in the first attribute information set 31a conflicts with the negative prompt word "bad hands" in the second attribute information set 34a that needs to be blocked; in this case, "four hands" in the Prompt information is blocked.
[0098] For example, the second attribute information set 34a includes positive attribute information used to improve the image quality of related images. For example, the positive attribute information may include terms such as "high quality" or "excellent," and this positive attribute information, along with the unmasked first attribute information, can be added to the Prompt information.
[0099] In this way, the prompts for user input can be automatically and effectively controlled through the hidden set of second attribute information in the background, thereby improving the image generation effect.
[0100] For example, after the aforementioned masking and addition processing is performed on the first attribute information set 31a, the prompt information is determined to be "a high-quality image of a standing, smiling man wearing glasses and modern clothing"; this prompt information can be input into the machine learning model 33a. In this way, the machine learning model can generate a related image 35a of the person object based on the user-input first attribute information set 31a, under the control of the pre-made object template 32a and the second attribute information set 34a.
[0101] Figure 3b Schematic diagrams illustrating some embodiments of the image generation method for background objects of this disclosure.
[0102] like Figure 3bAs shown, for example, if a user wants to generate an image related to a virtual background object, they can input a set of first attribute information 31b containing multiple first attribute information through an input box on the page. For example, the first attribute information set 31b input by the user includes "a forest in front of a mountain peak located in the center of the image under a night sky".
[0103] For example, if the type of the virtual object is determined to be a background object based on the first attribute information set 31b input by the user, the corresponding background object template 32b can be automatically matched as the current object template; if the image style type is determined to be modern style based on the first attribute information set 31b input by the user, the corresponding modern style object template 32b can be automatically matched as the current object template.
[0104] For example, users can be provided with templates of various image styles to choose from; or the image style type can be automatically determined using machine learning models based on the historical context and related plot of the interactive multimedia content in which the virtual background object is located, thereby improving image generation efficiency.
[0105] For example, object template 32b can specify the size information of the generated image, such as the width and height, and can also specify the position of the main target. If object template 32b specifies the position of the main target in the image as "located in the upper left of the image", which conflicts with "located in the center of the image" in the first attribute information set 31b, then "located in the center of the image" is hidden in the Prompt information, and "located in the upper left of the image" is added to the Prompt information.
[0106] In this way, the prompts for user input can be automatically and effectively controlled through pre-made object templates, thereby improving the quality of image generation.
[0107] For example, the second attribute information set 34b includes positive attribute information used to improve the image quality of related images. For example, the positive attribute information may include "high quality," "excellent," etc., and this positive attribute information, together with the unmasked first attribute information, can be added to the Prompt information.
[0108] In this way, the prompts for user input can be automatically and effectively controlled through the hidden set of second attribute information in the background, thereby improving the image generation effect.
[0109] For example, after performing the aforementioned masking and addition processing on the first attribute information set 31b, the prompt information is determined to be "a high-quality image of a forest in front of a mountain peak in the upper left corner of the image under a night sky"; this prompt information can be input into the machine learning model 33b. In this way, the machine learning model can generate a relevant image 35b of the background object based on the user-input first attribute information set 31b, under the control of the pre-made object template 32b and the second attribute information set 34b.
[0110] In some embodiments, multiple candidate images are generated based on a first set of attribute information, and the resolution of the multiple candidate images is lower than a threshold. In response to a user's selection operation among the multiple candidate images, a high-resolution operation is performed on the selected candidate image to generate a related image, the resolution of which is higher than that of the multiple candidate images. For example, this can be achieved through... Figure 3c The embodiments described above implement the technical solution for generating candidate images.
[0111] Figure 3c Schematic diagrams illustrating some embodiments of the candidate image generation method of this disclosure.
[0112] like Figure 3c As shown, if a user wants to generate images related to a virtual character, they can input a set of first attribute information, such as "a smiling man wearing glasses, modern clothing, and standing," through an input box on the page. The Prompt information corresponding to this first attribute information set is then input into machine learning model 33a to generate low-resolution related images 351, 352, and 353. In this way, the resolution of related images 351, 352, and 353 is all below a threshold, thereby improving the efficiency of image generation.
[0113] like Figure 3b As shown, images 351, 352, and 353 all conform to the user's definition of "a smiling man wearing glasses and modern clothing, standing" in the first set of attribute information, but there are differences between them. For example, the man in image 351 is wearing a short-sleeved T-shirt, the man in image 352 is wearing a suit, and the man in image 353 is wearing a long-sleeved T-shirt.
[0114] For example, a user can select the desired image from relevant images 351, 352, and 353; in response to the user selecting relevant image 352, high-resolution operations can be performed on relevant image 352 (such as through artificial intelligence algorithms) to generate relevant image 35a with a resolution higher than a threshold.
[0115] This approach allows for the rapid generation of multiple low-resolution candidate images for the user to choose from. The selected image is then processed at high resolution to obtain the desired high-quality image. This improves both the efficiency of image generation and the quality of the generated images, ensuring they better meet practical needs.
[0116] In some embodiments, the dynamic information of a virtual object in a related image is determined based on content information corresponding to the related image, the virtual object including a human figure; based on the dynamic information, a dynamic image of the human figure is generated using a digital human model on the basis of the related image.
[0117] For example, when the virtual object is a character in a game, there may be scenarios where the character performs a skill. In this case, corresponding action information can be generated as dynamic information based on the character's skill; based on this action information, and combined with a digital human model on the basis of relevant images, dynamic images of the character performing the skill can be generated.
[0118] For example, when the virtual object is a human character, there might be a scene where the human character is speaking. In this case, lip movements can be generated based on the human character's lines as dynamic information, and a dynamic image of the human character speaking can be generated based on the lip movements in the relevant image.
[0119] In some embodiments, the facial angle of a person in a relevant image is adjusted based on dynamic information; a dynamic image is then generated based on the adjusted relevant image. For example, the virtual object includes a person, and pose information is used to indicate that the face of the person in the relevant image is facing the user, and that the eyes of the person are facing the user.
[0120] For example, in scenarios depicting a person speaking, the person's face often needs to be turned towards the user off-screen to clearly convey their intention. In such cases, the angle of the person's face in the relevant image can be adjusted so that the adjusted face appears... Figure 3a The related image 35a is facing the user directly.
[0121] This ensures that the generated images can be combined with digital human technology to create dynamic images of people, thereby improving the quality of image generation.
[0122] In some embodiments, the relevant image corresponds to the current frame image of the interactive multimedia content, and the object template is selected based on the content to be represented by the current frame image.
[0123] For example, a virtual object might be a character in a game that consists of multiple frames, each used to display different plot content. For instance, if the current frame of the game is intended to depict a character standing and speaking, an object template can be automatically matched, specifying the character's posture as standing and their face facing the user.
[0124] This allows for the automatic matching of object templates that match the current content, thereby improving the efficiency and quality of image generation.
[0125] In some embodiments, where the virtual object includes a person object, the facial region of the person object is identified in the relevant image; based on the facial region, the relevant image is segmented to generate the avatar of the person object.
[0126] For example, if the virtual object is a character in a game, and we need to generate the character's portrait as an art asset in the game, we can improve the efficiency and quality of art asset generation by segmenting the portrait from the existing images of the character.
[0127] In some embodiments, portions of the related image other than the virtual object are transparent. The related image is then blended with a background image to generate frame images for interactive multimedia content. For example, Figure 3a The background portion of the related image 35a of the generated character object is transparent, and the related image 35a can be directly overlaid on the background image to generate frame images that conform to the plot of interactive multimedia content, thereby improving the efficiency and quality of image generation.
[0128] In some embodiments, the image region containing the virtual object is segmented from the relevant image; the image region is then fused with the background image to generate a frame image in the interactive multimedia content. For example, the human object in the relevant image 35a can be segmented to obtain a human image containing only the human object; this human image is then directly overlaid on the background image to generate a frame image that conforms to the plot of the interactive multimedia content, thereby improving the efficiency and quality of image generation.
[0129] In some embodiments, a calculator learning model can be used to automatically generate a set of attribute information about a character based on prior knowledge such as the plot content and character settings of the interactive multimedia content; and then generate a related image of the character based on the set of attribute information.
[0130] For example, if the user is a game creator, after completing the setting of game chapters and character outlines, LLM can be used to understand the game's plot, character settings, and background settings, automatically generating a set of first attribute information to describe the character's characteristics. Based on this first attribute information set, combined with preset character poses and scene composition information, the composition and prompt information of the image can be controlled to generate the relevant image. In this way, the automated deployment of the text-to-image process can be achieved, thereby improving the efficiency and quality of image generation.
[0131] In some embodiments, if the generated Prompt information is not in a language the user is familiar with, the user can directly modify the Prompt information in their own language. Artificial intelligence technology is then used to automatically translate the user-modified Prompt information into its original language. This allows users to easily adjust the generated image using their familiar language, thereby improving image generation efficiency.
[0132] In the above embodiments, a machine learning model is used to automatically generate images of virtual objects by combining user-input attribute information and preset object templates. This allows users to generate high-quality images that meet their needs without requiring artistic skills, thereby reducing the time and labor costs of image generation and improving image generation efficiency and quality.
[0133] Figure 4 Block diagrams illustrating some embodiments of the image generation apparatus of this disclosure are shown.
[0134] like Figure 4 As shown, the image generation device 4 includes: a receiving unit 41, for receiving a first set of attribute information of a virtual object input by a user; and a generation unit 42, for generating a related image of the virtual object based on the first set of attribute information and an object template, wherein the object template is used to determine at least one of the pose information and composition information of the virtual object in the related image, and the object template matches the first set of attribute information.
[0135] In some embodiments, if there is a conflict between the first attribute information set and the object template, the object template has a higher priority than the first attribute information set.
[0136] In some embodiments, if the generation unit 42 has attribute information in the first attribute information set that conflicts with the object template, it masks the attribute information; and generates a related image based on the masking result and the object template.
[0137] In some embodiments, the generation unit 42 generates a related image based on the first attribute information, the object template, and the stored second attribute information set. The second attribute information set is used to indicate the specified requirements of the related image. In the event of a conflict between the first attribute information set and the second attribute information set, the second attribute information set has a higher priority than the first attribute information set.
[0138] In some embodiments, the second attribute information set includes negative attribute information. If there is attribute information in the first attribute information set that conflicts with the negative attribute information, the generation unit 42 masks the attribute information and generates a related image based on the masking result and the second attribute information set.
[0139] In some embodiments, the second set of attribute information includes positive attribute information used to improve the image quality of the relevant images.
[0140] In some embodiments, the second set of attribute information is not visible to the user.
[0141] In some embodiments, the second set of attribute information is used to indicate the image style type of the relevant image.
[0142] In some embodiments, the generation unit 42 generates a relevant image based on the image style type selected by the user.
[0143] In some embodiments, the generation unit 42 generates multiple candidate images based on a first set of attribute information. The resolution of the multiple candidate images is lower than a threshold. In response to a user's selection operation among the multiple candidate images, a high-resolution operation is performed on the selected candidate image to generate a related image. The resolution of the related image is higher than that of the multiple candidate images.
[0144] In some embodiments, the generation unit 42 determines the dynamic information of the virtual object in the relevant image based on the content information corresponding to the relevant image. The virtual object includes a human figure. Based on the dynamic information, the generation unit 42 generates a dynamic image of the human figure using a digital human model on the basis of the relevant image.
[0145] In some embodiments, the generation unit 42 adjusts the facial angle of the person in the relevant image according to the dynamic information; and generates a dynamic image based on the adjusted relevant image.
[0146] In some embodiments, the virtual object includes a person object, and the pose information is used to indicate that the face of the person object in the relevant image is facing the user and the eyes of the person object are facing the user.
[0147] In some embodiments, the generation unit 42, in response to a user's selection operation among multiple candidate object templates, determines an object template from the multiple candidate object templates, and the multiple candidate object templates match a first set of attribute information.
[0148] In some embodiments, the relevant image corresponds to the current frame image of the interactive multimedia content, and the object template is selected based on the content to be represented by the current frame image.
[0149] In some embodiments, when the virtual object includes a person object, the generation unit 42 identifies the facial region of the person object in the relevant image, and segments the relevant image based on the facial region to generate the avatar of the person object.
[0150] In some embodiments, the portion of the related image other than the virtual object is transparent, and the generation unit 42 merges the related image with the background image to generate a frame image in the interactive multimedia content.
[0151] In some embodiments, the generation unit 42 segments the image region where the virtual object is located from the relevant image and merges the image region with the background image to generate a frame image in the interactive multimedia content.
[0152] In some embodiments, the virtual object includes a character object and a background object, and the composition information includes at least one of the virtual object's position information and size information in the relevant image.
[0153] In the above embodiments, by combining user-input attribute information and preset object templates, relevant images of virtual objects are automatically generated. This reduces time and labor costs, thereby improving image generation efficiency and quality.
[0154] Figure 5 Block diagrams illustrating other embodiments of the image generation apparatus of this disclosure are shown.
[0155] like Figure 5 As shown, the image generation apparatus 5 of this embodiment includes a memory 51 and a processor 52 coupled to the memory 51. The processor 52 is configured to execute the image generation method of any embodiment of this disclosure based on instructions stored in the memory 51.
[0156] The memory 51 may include, for example, system memory, fixed non-volatile storage media, etc. The system memory stores, for example, the operating system, application programs, a boot loader, a database, and other programs.
[0157] Figure 6 Block diagrams illustrating further embodiments of the image generation apparatus of this disclosure are shown.
[0158] like Figure 6As shown, the image generation apparatus 6 of this embodiment includes a memory 610 and a processor 620 coupled to the memory 610. The processor 620 is configured to execute the image generation method of any of the foregoing embodiments based on instructions stored in the memory 610.
[0159] The memory 610 may include, for example, system memory, fixed non-volatile storage media, etc. The system memory may store, for example, the operating system, application programs, a boot loader, and other programs.
[0160] The image generation device 6 may also include an input / output interface 630, a network interface 640, and a storage interface 650. These interfaces 630, 640, and 650, as well as the memory 610 and processor 620, can be connected, for example, via a bus 660. The input / output interface 630 provides a connection interface for input / output devices such as a monitor, mouse, keyboard, touchscreen, microphone, and speakers. The network interface 640 provides a connection interface for various networked devices. The storage interface 650 provides a connection interface for external storage devices such as SD cards and USB flash drives.
[0161] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable non-transitory storage media containing computer-usable program code, including but not limited to disk storage, CD-ROM, optical storage, etc.
[0162] The image generation method, image generation apparatus, and non-volatile computer-readable storage medium according to this disclosure have been described in detail. To avoid obscuring the concept of this disclosure, some details known in the art have not been described. Those skilled in the art will fully understand how to implement the technical solutions disclosed herein based on the above description.
[0163] The methods and systems of this disclosure may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the methods is for illustrative purposes only, and the steps of the methods of this disclosure are not limited to the order specifically described above unless otherwise specifically stated. Furthermore, in some embodiments, this disclosure may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the methods according to this disclosure. Thus, this disclosure also covers recording media storing programs for performing the methods according to this disclosure.
[0164] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.
Claims
1. An image generation method, comprising: Receives the first set of attribute information of the virtual object input by the user; Based on the first set of attribute information and the object template, a relevant image of the virtual object is generated. The object template is used to determine at least one of the pose information and composition information of the virtual object in the relevant image. The object template matches the first set of attribute information. The step of generating the relevant image of the virtual object based on the first attribute information set and the object template includes: The relevant image is generated based on the first attribute information, the object template, and the stored second attribute information set. The second attribute information set is used to indicate the specified requirements of the relevant image. In the event of a conflict between the first attribute information set and the second attribute information set, the second attribute information set has a higher priority than the first attribute information set.
2. The image generation method according to claim 1, wherein, In the event of a conflict between the first set of attribute information and the object template, the priority of the object template is higher than the priority of the first set of attribute information.
3. The image generation method according to claim 2, wherein, The step of generating the relevant image of the virtual object based on the first attribute information set and the object template includes: If there are attribute information in the first attribute information set that conflict with the object template, then the attribute information is masked. Based on the masking results and the object template, the relevant image is generated.
4. The image generation method according to claim 1, wherein, The second set of attribute information includes negative attribute information. The step of generating the relevant image based on the stored set of second attribute information includes: If there is attribute information in the first attribute information set that conflicts with the negative attribute information, then the attribute information is masked. The relevant image is generated based on the masking result and the second set of attribute information.
5. The image generation method according to claim 1, wherein, The second set of attribute information includes positive attribute information used to improve the image quality of the relevant images.
6. The image generation method according to claim 1, wherein, The second set of attribute information is not visible to the user.
7. The image generation method according to claim 1, wherein, The second set of attribute information is used to indicate the image style type of the related image.
8. The image generation method according to any one of claims 1-3, wherein, The step of generating the relevant image of the virtual object based on the first attribute information set and the object template includes: The relevant image is generated based on the image style type selected by the user.
9. The image generation method according to any one of claims 1-3, wherein, The step of generating the relevant image of the virtual object based on the first attribute information set and the object template includes: Based on the first set of attribute information, multiple candidate images are generated, wherein the resolution of the multiple candidate images is lower than a threshold. In response to the user's selection operation among the multiple candidate images, a high-resolution operation is performed on the selected candidate image to generate the related image, the related image having a higher resolution than the multiple candidate images.
10. The image generation method according to any one of claims 1-3, further comprising: Based on the content information corresponding to the relevant image, the dynamic information of the virtual object in the relevant image is determined, and the virtual object includes a human figure object; Based on the dynamic information, a dynamic image of the person is generated using a digital human model on the basis of the relevant image.
11. The image generation method according to claim 10, wherein, The step of generating a dynamic image of the human object based on the relevant image using a digital human model according to the dynamic information includes: Based on the dynamic information, adjust the facial angle of the person in the relevant image; The dynamic image is generated based on the adjusted relevant images.
12. The image generation method according to any one of claims 1-3, wherein, The virtual object includes a human figure, and the pose information is used to indicate that the face of the human figure in the relevant image is facing the user, and the eyes of the human figure are facing the user.
13. The image generation method according to any one of claims 1-3, wherein, The step of generating the relevant image of the virtual object based on the first attribute information set and the object template includes: In response to the user's selection operation among multiple candidate object templates, the object template is determined from the multiple candidate object templates, and the multiple candidate object templates match the first attribute information set.
14. The image generation method according to claim 13, wherein, The relevant image corresponds to the current frame image of the interactive multimedia content, and the object template is selected according to the content to be represented by the current frame image.
15. The image generation method according to any one of claims 1-3, further comprising: When the virtual object includes a human figure, the facial region of the human figure is identified in the relevant image; The relevant image is segmented based on the facial region to generate the avatar of the person.
16. The image generation method according to any one of claims 1-3, wherein, The portion of the image other than the virtual object is transparent. The image generation method further includes: The relevant images are fused with the background image to generate frame images in interactive multimedia content.
17. The image generation method according to any one of claims 1-3, further comprising: Segment the image region containing the virtual object from the relevant images; The image region is fused with the background image to generate a frame image in interactive multimedia content.
18. The image generation method according to any one of claims 1-3, wherein, The virtual object includes a character object and a background object, and the composition information includes at least one of the position information and size information of the virtual object in the relevant image.
19. An image generation apparatus, comprising: The receiving unit is used to receive the first set of attribute information of the virtual object input by the user; The generation unit is configured to generate a relevant image of the virtual object based on the first attribute information set and the object template, wherein the object template is used to determine at least one of the pose information and composition information of the virtual object in the relevant image, and the object template matches the first attribute information set. The generation unit generates the relevant image based on the first attribute information, the object template, and the stored second attribute information set. The second attribute information set is used to indicate the specified requirements of the relevant image. In the event of a conflict between the first attribute information set and the second attribute information set, the second attribute information set has a higher priority than the first attribute information set.
20. An image generation apparatus, comprising: Memory; and A processor coupled to the memory, the processor being configured to execute the image generation method of any one of claims 1-18 based on instructions stored in the memory.
21. A non-volatile computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image generation method according to any one of claims 1-18.
Citation Information
Patent Citations
Image optimization method and system based on artificial intelligence
CN109767397A
Image fusion transformation
CN109993716A
Data processing method and device, electronic equipment and storage medium
CN112734883A
Image special effect adding method and device, electronic equipment and storage medium
CN113643411A
NLP tool to dynamically create movies / animated scenes
US20060217979A1