Image generation method and apparatus, and nonvolatile computer-readable storage medium
By combining preset object templates and user-input attribute information, the machine learning model is used to automatically generate images, which solves the problem of low image generation efficiency and effect in the prior art, and achieves efficient and low-cost image generation.
Patent Information
- Application Number
- PCT/CN2024/108339
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-27
- Filing Date
- 2024-07-30
- Publication Date
- 2025-06-05
AI Technical Summary
In the prior art, image generation efficiency and generation effect are limited by time and labor costs, resulting in a decrease in image generation efficiency and effect.
By utilizing preset object templates and user-input attribute information, images are automatically generated in combination with machine learning models. The user can select the desired image from the multiple generated images for high resolution processing to generate high-quality virtual object images.
It improves image generation efficiency and generation effect, reduces time and labor costs, and allows users to generate high-quality images without having an art foundation.
Smart Images

Figure CN2024108339_05062025_PF_FP_ABST
Abstract
Description
Image generation method, device, and non-volatile computer-readable storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is based on the application with CN application number 202311597187.3 and application date November 27, 2023, and claims its priority. The disclosed content of the CN application is hereby introduced as a whole into this application. Technical Field
[0003] The present disclosure relates to the field of computer technology, and in particular to an image generating method, an image generating device, and a non-volatile computer-readable storage medium. Background Art
[0004] In computer applications such as games and interactive stories, it is often necessary to visually display virtual objects such as characters and scenes through images and other art resources. For example, character portraits and background images are needed.
[0005] In related technologies, producers of interactive multimedia content need to provide specific image generation requirements to professional painters, and rely on manual work to draw the corresponding art resources.
[0006] Summary of the Invention
[0007] According to some embodiments of the present disclosure, an image generation method is provided, including: receiving a first attribute information set of a virtual object input by a user; generating a related image of the virtual object based on the first attribute information set and an object template, wherein the object template is used to determine at least one of the posture information and composition information of the virtual object in the related image, and the object template matches the first attribute information set.
[0008] In some embodiments, when there is a conflict between the first attribute information set and the object template, the priority of the object template is higher than the priority of the first attribute information set.
[0009] In some embodiments, generating a related image of a virtual object based on a first attribute information set and an object template includes: when there is attribute information in the first attribute information set that conflicts with the object template, masking the attribute information; and generating a related image based on the masking result and the object template.
[0010] In some embodiments, generating a related image of a virtual object based on a first attribute information set and an object template includes: generating a related image based on the first attribute information, the object template and a stored second attribute information set, the second attribute information set being used to indicate specified requirements for the related image, and in the event of a conflict between the first attribute information set and the second attribute information set, the second attribute information set has a higher priority than the first attribute information set.
[0011] In some embodiments, the second attribute information set includes negative attribute information, and generating a related image based on the stored second attribute information set includes: when there is attribute information in the first attribute information set that conflicts with the negative attribute information, masking the attribute information; and generating a related image based on the masking result and the second attribute information set.
[0012] In some embodiments, the second attribute information set includes positive attribute information for improving image quality of the associated image.
[0013] In some embodiments, the second set of attribute information is not visible to the user.
[0014] In some embodiments, the second attribute information set is used to indicate the image style type of the related image.
[0015] In some embodiments, generating a related image of the virtual object based on the first attribute information set and the object template includes: generating a related image based on an image style type selected by a user.
[0016] In some embodiments, generating a related image of a virtual object based on a first attribute information set and an object template includes: generating multiple candidate images based on the first attribute information set, and the resolution of the multiple candidate images is lower than a threshold; in response to a user's selection operation among the multiple candidate images, performing a high-resolution operation on the selected candidate image to generate a related image, and the resolution of the related image is higher than that of the multiple candidate images.
[0017] In some embodiments, the image generation method further includes: determining the dynamic information of the virtual object in the relevant image based on the content information corresponding to the relevant image, the virtual object including the human object; and generating a dynamic image of the human object based on the relevant image using a digital human model according to the dynamic information.
[0018] In some embodiments, generating a dynamic image of a human object using a digital human model based on relevant images according to dynamic information includes: adjusting the facial angle of the human object in the relevant images according to the dynamic information; and generating a dynamic image based on the adjusted relevant images.
[0019] In some embodiments, the virtual object includes a character object, and the posture information is used to indicate that the face of the character object in the relevant image is facing the user, and the eyes of the character object are facing the user.
[0020] In some embodiments, generating a related image of a virtual object based on a first attribute information set and an object template includes: determining an object template from multiple candidate object templates in response to a user's selection operation among multiple candidate object templates, and multiple candidate object templates match the first attribute information set.
[0021] In some embodiments, the relevant image corresponds to a current frame image of the interactive multimedia content, and the object template is selected according to the content that the current frame image is intended to represent.
[0022] In some embodiments, the image generation method further includes: when the virtual object includes a human object, identifying a facial area of the human object in the relevant image; and segmenting the relevant image according to the facial area to generate an avatar of the human object.
[0023] In some embodiments, the portion of the relevant image other than the virtual object is transparent, and the image generation method further includes: fusing the relevant image with the background image to generate a frame image in the interactive multimedia content.
[0024] In some embodiments, the image generation method further includes: segmenting an image region where the virtual object is located from the relevant image; and fusing the image region with the background image to generate a frame image in the interactive multimedia content.
[0025] In some embodiments, the virtual object includes a character object and a background object, and the composition information includes at least one of position information and size information of the virtual object in the relevant image.
[0026] According to some other embodiments of the present disclosure, an image generating device is provided, including: a receiving unit for receiving a first attribute information set of a virtual object input by a user; a generating unit for generating a related image of the virtual object based on the first attribute information set and an object template, the object template being used to determine at least one of posture information and composition information of the virtual object in the related image, and the object template matching the first attribute information set.
[0027] In some embodiments, when there is a conflict between the first attribute information set and the object template, the priority of the object template is higher than the priority of the first attribute information set.
[0028] In some embodiments, when there is attribute information in the first attribute information set that conflicts with the object template, the generating unit masks the attribute information; and generates a related image based on the masking result and the object template.
[0029] In some embodiments, the generation unit generates a related image based on the first attribute information, the object template and the stored second attribute information set, where the second attribute information set is used to indicate the specified requirements of the related image. In the event of a conflict between the first attribute information set and the second attribute information set, the second attribute information set has a higher priority than the first attribute information set.
[0030] In some embodiments, the second attribute information set includes negative attribute information. When attribute information that conflicts with the negative attribute information exists in the first attribute information set, the generation unit masks the attribute information and generates a related image based on the masking result and the second attribute information set.
[0031] In some embodiments, the second attribute information set includes positive attribute information for improving image quality of the associated image.
[0032] In some embodiments, the second set of attribute information is not visible to the user.
[0033] In some embodiments, the second attribute information set is used to indicate the image style type of the related image.
[0034] In some embodiments, the generating unit generates relevant images according to the image style type selected by the user.
[0035] In some embodiments, the generation unit generates multiple candidate images based on the first attribute information set, and the resolution of the multiple candidate images is lower than a threshold. In response to the user's selection operation among the multiple candidate images, the generation unit performs high-resolution operation on the selected candidate image to generate a related image, and the resolution of the related image is higher than the multiple candidate images.
[0036] In some embodiments, the generation unit determines the dynamic information of the virtual object in the relevant image based on the content information corresponding to the relevant image. The virtual object includes a human object. According to the dynamic information, a dynamic image of the human object is generated based on the relevant image using a digital human model.
[0037] In some embodiments, the generating unit adjusts the facial angle of the human object in the relevant image according to the dynamic information; and generates a dynamic image according to the adjusted relevant image.
[0038] In some embodiments, the virtual object includes a character object, and the posture information is used to indicate that the face of the character object in the relevant image is facing the user, and the eyes of the character object are facing the user.
[0039] In some embodiments, the generating unit determines the object template from the plurality of candidate object templates in response to a user's selection operation among the plurality of candidate object templates, and the plurality of candidate object templates match the first attribute information set.
[0040] In some embodiments, the relevant image corresponds to a current frame image of the interactive multimedia content, and the object template is selected according to the content that the current frame image is intended to represent.
[0041] In some embodiments, when the virtual object includes a human object, the generation unit identifies a facial region of the human object in the relevant image, and segments the relevant image according to the facial region to generate an avatar of the human object.
[0042] In some embodiments, the portion of the relevant image other than the virtual object is transparent, and the generation unit merges the relevant image with the background image to generate a frame image in the interactive multimedia content.
[0043] In some embodiments, the generating unit segments an image region where the virtual object is located from the relevant image, and merges the image region with the background image to generate a frame image in the interactive multimedia content.
[0044] In some embodiments, the virtual object includes a character object and a background object, and the composition information includes at least one of position information and size information of the virtual object in the relevant image.
[0045] According to some further embodiments of the present disclosure, an image generating device is provided, comprising: a memory; and a processor coupled to the memory, wherein the processor is configured to execute the image generating method in any of the above embodiments based on instructions stored in the memory device.
[0046] According to some further embodiments of the present disclosure, a non-volatile computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the image generating method in any of the above embodiments is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0048] The present disclosure can be more clearly understood from the following detailed description with reference to the accompanying drawings:
[0049] FIG1 shows a flowchart of some embodiments of the image generation method of the present disclosure;
[0050] FIG2 shows a flowchart of another embodiment of the image generation method of the present disclosure;
[0051] FIG3 a shows a schematic diagram of some embodiments of the method for generating an image of a human object disclosed herein;
[0052] FIG3 b is a schematic diagram showing some embodiments of the method for generating an image of a background object disclosed herein;
[0053] FIG3 c shows a schematic diagram of some embodiments of the candidate image generation method disclosed herein;
[0054] FIG4 shows a block diagram of some embodiments of the image generation apparatus of the present disclosure;
[0055] FIG5 is a block diagram showing some other embodiments of the image generating apparatus disclosed herein;
[0056] FIG6 is a block diagram showing some further embodiments of the image generating apparatus disclosed herein. DETAILED DESCRIPTION
[0057] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangement of components and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure.
[0058] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0059] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.
[0060] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.
[0061] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.
[0062] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0063] The inventors of the present disclosure have discovered that the above-mentioned related technologies have the following problems: high time and labor costs lead to decreased image generation efficiency and generation effects.
[0064] In view of this, the present disclosure proposes an image generation technology solution that can improve image generation efficiency and generation effect.
[0065] As mentioned above, manually drawing art resources in interactive multimedia content is time-consuming and labor-intensive, resulting in high time and labor costs, and reducing image generation efficiency and effects.
[0066] To address the aforementioned technical issues, the present invention utilizes pre-defined object templates to create preset character poses and image compositions. This, combined with a simple user-provided description in an input box and pre-defined prompts, automatically generates images using a machine learning model. Users can also select a desired image from the generated images for high-resolution processing, generating character or background images suitable for interactive multimedia content.
[0067] In this way, by combining the attribute information input by the user and the preset object template, high-quality images of the virtual object can be quickly generated, thereby improving the efficiency and effect of image generation. For example, the technical solution of the present disclosure can be implemented through the following embodiments.
[0068] For example, the technical solutions of the present disclosure can be implemented through the following embodiments.
[0069] FIG1 shows a flowchart of some embodiments of the image generation method of the present disclosure.
[0070] As shown in FIG1 , in step 110 , a first attribute information set of a virtual object input by a user is received.
[0071] For example, virtual objects may include character objects, background objects, etc. in multimedia interactive content; multimedia interactive content may include games, interactive stories, interactive comics, and interactive dramas, etc.; the first attribute information set may include multiple first attribute information used to describe the characteristics of the virtual object.
[0072] In step 120, a related image of the virtual object is generated based on the first attribute information set and the object template. The object template is used to determine at least one of posture information and composition information of the virtual object in the related image, and the object template matches the first attribute information set.
[0073] In the above embodiment, the attribute information input by the user and the preset object template are combined to automatically generate the relevant image of the virtual object. This can reduce time and labor costs, thereby improving the efficiency and effect of image generation.
[0074] The following uses some embodiments to illustrate how to automatically generate images related to virtual objects.
[0075] FIG2 shows a flowchart of other embodiments of the image generating method disclosed herein.
[0076] As shown in Figure 2, the first attribute information set input by the user can be used as prompt information and input into the machine learning model to automatically generate the relevant image of the virtual object. For example, the first attribute information set can be masked in steps 1210 to 1230, and then the relevant image of the virtual object can be automatically generated using the machine learning model.
[0077] For example, the machine learning model can be a large language model (LLM) that can generate image content based on prompt information provided by the user. In this way, the generation efficiency and efficiency of the image can be improved through the machine learning model.
[0078] In some embodiments, if the first attribute information set conflicts with the object template, the object template has a higher priority than the first attribute information set. For example, in step 1210, if the first attribute information set contains first attribute information that conflicts with the object template, the first attribute information is masked.
[0079] For example, if the virtual object is a human object, the object template can be used to determine the posture information of the human object. For example, the first attribute information set includes the first attribute information "human leaping", but the posture information defined by the object template includes "human standing". In this case, "human leaping" is masked from the Prompt information, and "human standing" is added to the Prompt information before being input into the machine learning model.
[0080] For example, if the virtual object is a background object, the object template can be used to determine the background object's composition information. Composition information can include at least one of the virtual object's position and size information within the relevant image. For example, if the first attribute information set includes the first attribute information "mountain peak located in the center of the image," but the object template defines the composition information as "mountain peak located in the upper left of the image," in this case, the prompt message "mountain peak located in the center of the image" is masked and "mountain peak located in the upper left of the image" is added to the prompt message before being input into the machine learning model.
[0081] In some embodiments, the object template matches the first attribute information set. For example, the type and art style of the virtual object corresponding to the object template may match the type and art style of the virtual object corresponding to the first attribute information set.
[0082] For example, if the first attribute information set describes a person object, the object template is a template for the person object; if the first attribute information set describes a background object, the object template is a template for the background object. For example, if the person object described by the first attribute information set is an ancient figure, the style of the object template is an ancient style; if the person object described by the first attribute information set is a science fiction character, the style of the object template is a science fiction style.
[0083] In some embodiments, a related image is generated based on the first attribute information, the object template, and a stored second attribute information set. For example, the second attribute information set is used to indicate specific requirements for the related image. If there is a conflict between the first attribute information set and the second attribute information set, the second attribute information set takes precedence over the first attribute information set.
[0084] For example, in step 1220, when there is first attribute information in the first attribute information set that conflicts with negative attribute information, the first attribute information is masked.
[0085] In some embodiments, the second attribute information set includes negative attribute information. For example, negative attribute information is prompt information that needs to be blocked, and may include prompt information such as "bad hands" and "exposed body" that does not meet the specified requirements for image generation.
[0086] For example, the first attribute information set includes the first attribute information "specially shaped hands" which conflicts with the negative prompt word "bad hands" that needs to be blocked; in this case, the "specially shaped hands" in the Prompt information is blocked and then input into the machine learning model.
[0087] In some embodiments, the second attribute information set includes positive attribute information used to improve the image quality of the associated image. For example, the positive attribute information may include "high quality" or "excellent work." This positive attribute information and the first attribute information provided by the user may be input into the machine learning model as prompt information.
[0088] In some embodiments, the second attribute information set is invisible to the user. For example, the second attribute information set includes negative attribute information and prompt information indicating the image style type of the relevant image (such as cyberpunk style, Chinese martial arts style, modern urban style, etc.). This information is hidden in the background and is invisible to the user and does not require any operation, thereby improving the efficiency and quality of image generation.
[0089] In some embodiments, related images can also be generated based on the image style type selected by the user. For example, if the user selects a modern urban style, related images can be generated based on the object template corresponding to the modern urban style. In this way, matching images can be generated based on the user's actual needs, thereby improving the efficiency and quality of image generation.
[0090] In some embodiments, steps 1210 and 1220 may be executed in any order and may be executed in parallel. For example, after the masking process of steps 1210 and 1220, the prompt information may be determined to be "a smiling man wearing glasses and modern clothing"; this prompt information may be input into a machine learning model to generate a relevant image.
[0091] For example, in step 1230, a related image is generated based on the masking result and the object template. The masked Prompt information and the Prompt information can be input into a machine learning model; the machine learning model generates a related image based on the Prompt information and the object template that matches it.
[0092] The following uses some embodiments in FIG. 3 a to FIG. 3 c to exemplify how to automatically generate images related to virtual objects.
[0093] FIG3 a shows a schematic diagram of some embodiments of the method for generating an image of a human object according to the present disclosure.
[0094] As shown in Figure 3a, for example, if a user wants to generate an image related to a virtual character object, they can enter a first attribute information set 31a containing multiple pieces of first attribute information through the input box on the page. For example, the first attribute information set 31a entered by the user includes "a smiling man wearing glasses, modern clothing, four arms, jumping, and a smile."
[0095] For example, based on the first attribute information set 31a input by the user, if the type of the virtual object is determined to be a character object, the object template 32a corresponding to the character object can be automatically matched as the current object template; and based on the first attribute information set 31a input by the user, if the image style type is determined to be a modern urban style, the object template 32a corresponding to the modern urban style can be automatically matched as the current object template.
[0096] In some embodiments, in response to a user selecting from multiple candidate object templates, an object template is determined from the multiple candidate object templates, and the multiple candidate object templates match the first attribute information set. For example, templates of multiple image styles may be provided to the user for selection. Alternatively, a machine learning model may be used to automatically determine the image style based on the historical context and relevant plot of the interactive multimedia content containing the avatar object, thereby improving image generation efficiency.
[0097] For example, the object template 32a limits the posture information of the character object to "character standing", which conflicts with "jumping" in the first attribute information set 31a; in this case, "jumping" is masked in the Prompt information, and "character standing" is added to the Prompt information.
[0098] In this way, the prompt information input by the user can be automatically and effectively controlled through the prefabricated object template, thereby improving the effect of image generation.
[0099] For example, the first attribute information "four hands" in the first attribute information set 31a conflicts with the negative prompt word "bad hands" that needs to be blocked in the second attribute information set 34a; in this case, "four hands" in the Prompt information is blocked.
[0100] For example, the second attribute information set 34a includes positive attribute information for improving the image quality of the related image. For example, the positive attribute information may include "high quality" or "excellent work", and the positive attribute information and the unmasked first attribute information may be added to the Prompt information.
[0101] In this way, the prompt information input by the user can be automatically and effectively controlled through the second attribute information set hidden in the background, thereby improving the effect of image generation.
[0102] For example, after performing the aforementioned masking and addition processing on first attribute information set 31a, the prompt information is determined to be "a high-quality image of a man wearing glasses, modern clothing, standing, and smiling." This prompt information can be input into machine learning model 33a. In this way, the machine learning model can generate a related image 35a of a person object based on the first attribute information set 31a input by the user, under the control of a pre-made object template 32a and a second attribute information set 34a.
[0103] FIG3 b is a schematic diagram showing some embodiments of a method for generating an image of a background object according to the present disclosure.
[0104] As shown in Figure 3b, for example, if a user wants to generate an image related to a virtual background object, they can enter a first attribute information set 31b containing multiple pieces of first attribute information through the input box on the page. For example, the first attribute information set 31b entered by the user includes "under the night sky, there is a forest in front of the mountain in the center of the picture."
[0105] For example, based on the first attribute information set 31b input by the user, if the type of the virtual object is determined to be a background object, the object template 32b corresponding to the background object can be automatically matched as the current object template; and based on the first attribute information set 31b input by the user, if the image style type is determined to be a modern style, the object template 32b corresponding to the modern style can be automatically matched as the current object template.
[0106] For example, templates of various image style types may be provided to users for selection. The image style type may also be automatically determined using a machine learning model based on the historical context and relevant plot of the interactive multimedia content in which the virtual background object is located, thereby improving image generation efficiency.
[0107] For example, object template 32b can define the size information of the generated image, such as the width and height, and can also define the position of the primary target. For example, if object template 32b defines the position of the primary target in the image as "located in the upper left of the screen," this conflicts with the "mountain located in the center of the screen" in first attribute information set 31b. In this case, "mountain located in the center of the screen" is masked from the Prompt information, and "mountain located in the upper left of the screen" is added to the Prompt information.
[0108] In this way, the prompt information input by the user can be automatically and effectively controlled through the prefabricated object template, thereby improving the effect of image generation.
[0109] For example, the second attribute information set 34b includes positive attribute information for improving the image quality of the related image. For example, the positive attribute information may include "high quality" or "excellent work", and the positive attribute information and the unshielded first attribute information may be added to the Prompt information.
[0110] In this way, the prompt information input by the user can be automatically and effectively controlled through the second attribute information set hidden in the background, thereby improving the effect of image generation.
[0111] For example, after performing the aforementioned masking and addition processing on first attribute information set 31b, the prompt information is determined to be "a high-quality image of a forest in front of a mountain in the upper left corner of the image under a night sky." This prompt information can be input into machine learning model 33b. In this way, the machine learning model can generate an image 35b related to the background object based on the user-input first attribute information set 31b and under the control of the pre-fabricated object template 32b and the second attribute information set 34b.
[0112] In some embodiments, multiple candidate images are generated based on the first attribute information set, where the resolution of the multiple candidate images is lower than a threshold. In response to a user selecting one of the multiple candidate images, a high-resolution operation is performed on the selected candidate image to generate a related image, where the resolution of the related image is higher than that of the multiple candidate images. For example, the technical solution for generating candidate images can be implemented through the embodiment shown in FIG3c.
[0113] FIG3 c shows a schematic diagram of some embodiments of the candidate image generation method of the present disclosure.
[0114] As shown in Figure 3c, if a user wants to generate images related to a virtual character object, they can enter a first attribute information set containing multiple first attribute information items, such as "a standing, smiling man wearing glasses and modern clothing" through the input box on the page. The prompt information corresponding to this first attribute information set is then input into the machine learning model 33a, generating low-resolution related images 351, 352, and 353. This ensures that the resolution of related images 351, 352, and 353 is below the threshold, thereby improving image generation efficiency.
[0115] As shown in Figure 3b, related images 351, 352, and 353 all meet the user's definition of "a standing, smiling man wearing glasses and modern clothing" in the first attribute information set, but they differ from each other. For example, the man in related image 351 is wearing a short-sleeved T-shirt, the man in image 352 is wearing a suit, and the man in image 353 is wearing a long-sleeved T-shirt.
[0116] For example, the user can select the image he wants from the related images 351, 352 and 353; in response to the user selecting the related image 352, high-resolution operation can be performed on the related image 352 (such as can be achieved through artificial intelligence algorithms, etc.) to generate a related image 35a with a resolution higher than a threshold.
[0117] In this way, multiple candidate low-resolution images can be quickly generated for the user to select, and then the selected images can be processed to obtain the desired high-quality images. This improves the efficiency of image generation while making the generated images more in line with actual needs, thereby improving the quality of the generated images.
[0118] In some embodiments, dynamic information of a virtual object in a relevant image is determined based on content information corresponding to the relevant image, and the virtual object includes a human object; based on the dynamic information, a dynamic image of the human object is generated using a digital human model based on the relevant image.
[0119] For example, if the virtual object is a character in a game, there may be a scene where the character performs a skill. In this case, corresponding action information can be generated based on the character's skill as dynamic information. Based on this action information, the digital human model can be combined with the relevant images to generate a dynamic image of the character performing the skill.
[0120] For example, when the virtual object is a person, there may be a scene where the person speaks. In this case, the corresponding lip shape can be generated as dynamic information based on the person's lines, and a dynamic image of the person speaking can be generated based on the lip shape based on the relevant images.
[0121] In some embodiments, the facial angle of a person in a related image is adjusted based on the dynamic information, and a dynamic image is generated based on the adjusted related image. For example, the virtual object includes a person, and the posture information indicates that the face of the person in the related image is facing the user, and the eyes of the person are facing the user.
[0122] For example, in a scene showing a person speaking, it is often necessary to have the person's face facing the user outside the screen to clearly convey the person's speaking intention. In this case, the angle of the person's face in the relevant image can be adjusted so that the adjusted face is facing the user, as in the relevant image 35a in Figure 3a.
[0123] In this way, it is ensured that the generated relevant images can be combined with digital human technology to generate dynamic images of the characters, thereby improving the quality of image generation.
[0124] In some embodiments, the relevant image corresponds to a current frame image of the interactive multimedia content, and the object template is selected according to the content that the current frame image is intended to represent.
[0125] For example, the virtual object is a character in a game that includes multiple frames, each depicting a different plot. For example, if the current frame depicts the character standing and speaking, an object template can be automatically matched to the character's pose, with the character's face facing the user.
[0126] In this way, object templates that match the current content can be automatically matched, thereby improving the efficiency and quality of image generation.
[0127] In some embodiments, when the virtual object includes a human object, a facial region of the human object is identified in a related image; and the related image is segmented based on the facial region to generate an avatar of the human object.
[0128] For example, if the virtual object is a character in a game, and you need to generate an avatar of the character as an art resource in the game, you can improve the efficiency and quality of art resource generation by segmenting the avatar from the already generated image of the character.
[0129] In some embodiments, the portion of the associated image other than the virtual object is transparent. The associated image is fused with the background image to generate a frame image for the interactive multimedia content. For example, the background portion of the associated image 35a of the human object generated in FIG3a is transparent. The associated image 35a can be directly overlaid on the background image to generate a frame image that matches the plot of the interactive multimedia content, thereby improving the efficiency and quality of image generation.
[0130] In some embodiments, the image region containing the virtual object is segmented from the relevant image; the image region is then fused with the background image to generate a frame image for the interactive multimedia content. For example, the human object in the relevant image 35a can be segmented to obtain a human image containing only the human object; this human image is then directly overlaid on the background image to generate a frame image that matches the plot of the interactive multimedia content, thereby improving the efficiency and quality of image generation.
[0131] In some embodiments, a computer learning model can be used to automatically generate a set of attribute information about the character based on prior knowledge such as the plot content and character settings of the interactive multimedia content; and then generate a related image of the character based on the attribute information set.
[0132] For example, if a user is a game creator and has completed the game chapter and character profile settings, they can use LLM to understand the game's plot content, character settings, background settings, and other prior knowledge, and automatically generate a first attribute information set to describe the character's characteristics. Based on this first attribute information set, combined with preset information such as the character's posture and screen composition, the image composition and prompt information are controlled to generate relevant images. This can achieve automated deployment of the text-to-image process, thereby improving the efficiency and quality of image generation.
[0133] In some embodiments, if the generated prompt message is not in a language familiar to the user, the user can directly modify the prompt message in a language familiar to the user. Using artificial intelligence technology, the modified prompt message is automatically translated into the original language of the prompt message. This allows the user to conveniently adjust image generation in a language familiar to the user, thereby improving image generation efficiency.
[0134] In the above embodiment, a machine learning model is used to automatically generate images related to virtual objects, combining user-entered attribute information with pre-set object templates. This allows users to generate high-quality images that meet their needs without requiring any prior art knowledge, thereby reducing the time and labor costs of image generation and improving image generation efficiency and effectiveness.
[0135] FIG4 shows a block diagram of some embodiments of the image generating apparatus of the present disclosure.
[0136] As shown in Figure 4, the image generating device 4 includes: a receiving unit 41, used to receive a first attribute information set of a virtual object input by a user; a generating unit 42, used to generate a related image of the virtual object based on the first attribute information set and an object template, the object template is used to determine at least one of the posture information and composition information of the virtual object in the related image, and the object template matches the first attribute information set.
[0137] In some embodiments, when there is a conflict between the first attribute information set and the object template, the priority of the object template is higher than the priority of the first attribute information set.
[0138] In some embodiments, when there is attribute information in the first attribute information set that conflicts with the object template, the generating unit 42 masks the attribute information; and generates a related image based on the masking result and the object template.
[0139] In some embodiments, the generation unit 42 generates a related image based on the first attribute information, the object template and the stored second attribute information set, where the second attribute information set is used to indicate the specified requirements of the related image. In the event of a conflict between the first attribute information set and the second attribute information set, the second attribute information set has a higher priority than the first attribute information set.
[0140] In some embodiments, the second attribute information set includes negative attribute information. When there is attribute information in the first attribute information set that conflicts with the negative attribute information, the generation unit 42 masks the attribute information and generates a related image based on the masking result and the second attribute information set.
[0141] In some embodiments, the second attribute information set includes positive attribute information for improving image quality of the associated image.
[0142] In some embodiments, the second set of attribute information is not visible to the user.
[0143] In some embodiments, the second attribute information set is used to indicate the image style type of the related image.
[0144] In some embodiments, the generating unit 42 generates relevant images according to the image style type selected by the user.
[0145] In some embodiments, the generation unit 42 generates multiple candidate images based on the first attribute information set, and the resolution of the multiple candidate images is lower than a threshold. In response to the user's selection operation among the multiple candidate images, the generation unit 42 performs high-resolution operation on the selected candidate image to generate a related image, and the resolution of the related image is higher than the multiple candidate images.
[0146] In some embodiments, the generation unit 42 determines the dynamic information of the virtual object in the relevant image based on the content information corresponding to the relevant image. The virtual object includes a human object. According to the dynamic information, a dynamic image of the human object is generated based on the relevant image using a digital human model.
[0147] In some embodiments, the generating unit 42 adjusts the facial angle of the human object in the relevant image according to the dynamic information; and generates a dynamic image according to the adjusted relevant image.
[0148] In some embodiments, the virtual object includes a character object, and the posture information is used to indicate that the face of the character object in the relevant image is facing the user, and the eyes of the character object are facing the user.
[0149] In some embodiments, the generating unit 42 determines an object template from a plurality of candidate object templates in response to a user's selection operation among a plurality of candidate object templates, and the plurality of candidate object templates match the first attribute information set.
[0150] In some embodiments, the relevant image corresponds to a current frame image of the interactive multimedia content, and the object template is selected according to the content that the current frame image is intended to represent.
[0151] In some embodiments, when the virtual object includes a human object, the generation unit 42 identifies a facial region of the human object in the relevant image, and segments the relevant image according to the facial region to generate an avatar of the human object.
[0152] In some embodiments, the portion of the relevant image other than the virtual object is transparent, and the generating unit 42 merges the relevant image with the background image to generate a frame image in the interactive multimedia content.
[0153] In some embodiments, the generating unit 42 segments the image region where the virtual object is located from the relevant image, and merges the image region with the background image to generate a frame image in the interactive multimedia content.
[0154] In some embodiments, the virtual object includes a character object and a background object, and the composition information includes at least one of position information and size information of the virtual object in the relevant image.
[0155] In the above embodiment, the attribute information input by the user and the preset object template are combined to automatically generate the relevant image of the virtual object. This can reduce time and labor costs, thereby improving the efficiency and effect of image generation.
[0156] FIG5 is a block diagram showing some other embodiments of the image generating apparatus disclosed herein.
[0157] As shown in FIG5 , the image generating device 5 of this embodiment includes: a memory 51 and a processor 52 coupled to the memory 51 . The processor 52 is configured to execute the image generating method in any one embodiment of the present disclosure based on instructions stored in the memory 51 .
[0158] The memory 51 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory may store, for example, an operating system, application programs, a boot loader, a database, and other programs.
[0159] FIG6 is a block diagram showing some further embodiments of the image generating apparatus disclosed herein.
[0160] As shown in FIG6 , the image generating device 6 of this embodiment includes: a memory 610 and a processor 620 coupled to the memory 610 . The processor 620 is configured to execute the image generating method of any of the aforementioned embodiments based on instructions stored in the memory 610 .
[0161] The memory 610 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory may store, for example, an operating system, application programs, a boot loader, and other programs.
[0162] The image generation device 6 may further include an input / output interface 630, a network interface 640, a storage interface 650, and the like. These interfaces 630, 640, 650, as well as the memory 610 and the processor 620, may be connected, for example, via a bus 660. The input / output interface 630 provides a connection interface for input / output devices such as a display, mouse, keyboard, touch screen, microphone, and speakers. The network interface 640 provides a connection interface for various networked devices. The storage interface 650 provides a connection interface for external storage devices such as SD cards and USB flash drives.
[0163] Those skilled in the art will appreciate that embodiments of the present disclosure may be provided as methods, systems, or computer program products. Thus, the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable non-transitory storage media, including but not limited to magnetic disk storage, CD-ROMs, optical storage, and the like, containing computer-usable program code.
[0164] The image generation method, image generation device, and non-volatile computer-readable storage medium according to the present disclosure have been described in detail. To avoid obscuring the concepts of the present disclosure, some details known in the art have been omitted. Based on the above description, those skilled in the art will fully understand how to implement the technical solutions disclosed herein.
[0165] The methods and systems of the present disclosure may be implemented in many ways. For example, the methods and systems of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above unless otherwise specified. In addition, in some embodiments, the present disclosure may also be implemented as programs recorded in a recording medium, which include machine-readable instructions for implementing the methods according to the present disclosure. Therefore, the present disclosure also covers recording media that store programs for executing the methods according to the present disclosure.
[0166] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art will appreciate that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Those skilled in the art will appreciate that modifications may be made to the above embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. A method for generating an image, comprising: Receive a first attribute information set of a virtual object input by a user; A related image of the virtual object is generated according to the first attribute information set and an object template, wherein the object template is used to determine at least one of posture information and composition information of the virtual object in the related image, and the object template matches the first attribute information set.
2. The image generation method according to claim 1, wherein: In the case that there is a conflict between the first attribute information set and the object template, the priority of the object template is higher than the priority of the first attribute information set.
3. The image generation method according to claim 2, wherein: The generating the related image of the virtual object according to the first attribute information set and the object template comprises: In the case where there is attribute information in the first attribute information set that conflicts with the object template, shielding the attribute information; The related image is generated according to the masking result and the object template.
4. The image generation method according to any one of claims 1 to 3, wherein: The generating the related image of the virtual object according to the first attribute information set and the object template comprises: The related image is generated according to the first attribute information, the object template and the stored second attribute information set, wherein the second attribute information set is used to indicate the specified requirements of the related image, and in case of a conflict between the first attribute information set and the second attribute information set, the second attribute information set has a higher priority than the first attribute information set.
5. The image generation method according to claim 4, wherein: The second attribute information set includes negative attribute information, Generating the related image according to the stored second attribute information set includes: When there is attribute information in the first attribute information set that conflicts with the negative attribute information, shielding the attribute information; The related image is generated according to the masking result and the second attribute information set.
6. The image generation method according to claim 4 or 5, wherein: The second attribute information set includes positive attribute information for improving the image quality of the related image.
7. The image generation method according to any one of claims 4 to 6, wherein: The second attribute information set is invisible to the user.
8. The image generation method according to any one of claims 4 to 7, wherein: The second attribute information set is used to indicate the image style type of the related image.
9. The image generation method according to any one of claims 1 to 8, wherein: The generating the related image of the virtual object according to the first attribute information set and the object template comprises: The related image is generated according to the image style type selected by the user.
10. The image generation method according to any one of claims 1 to 9, wherein: The generating the related image of the virtual object according to the first attribute information set and the object template comprises: generating a plurality of candidate images according to the first attribute information set, wherein the resolution of the plurality of candidate images is lower than a threshold; In response to the user's selection operation among the multiple candidate images, a high-resolution operation is performed on the selected candidate image to generate the related image, where the resolution of the related image is higher than that of the multiple candidate images.
11. The image generation method according to any one of claims 1 to 10, further comprising: Determining dynamic information of the virtual object in the relevant image according to content information corresponding to the relevant image, the virtual object including a character object; According to the dynamic information, a dynamic image of the human object is generated by using a digital human model based on the relevant image.
12. The image generation method according to claim 11, wherein: The step of generating the dynamic image of the person object by using the digital human model based on the relevant image according to the dynamic information comprises: According to the dynamic information, adjusting the facial angle of the person object in the relevant image; The dynamic image is generated according to the adjusted related image.
13. The image generation method according to any one of claims 1 to 12, wherein: The virtual object includes a character object, and the posture information is used to indicate that the face of the character object in the relevant image is facing the user, and the eyes of the character object are facing the user.
14. The image generation method according to any one of claims 1 to 13, wherein: The generating the related image of the virtual object according to the first attribute information set and the object template comprises: In response to a selection operation by the user among a plurality of candidate object templates, the object template is determined from the plurality of candidate object templates, and the plurality of candidate object templates match the first attribute information set.
15. The image generation method according to claim 14, wherein: The related image corresponds to a current frame image of the interactive multimedia content, and the object template is selected according to the content that the current frame image is intended to represent.
16. The image generation method according to any one of claims 1 to 15, further comprising: In a case where the virtual object includes a human object, identifying a facial region of the human object in the relevant image; The relevant image is segmented according to the facial region to generate an image of the person object.
17. The image generation method according to any one of claims 1 to 16, wherein: The portion other than the virtual object in the relevant image is transparent, The image generation method further includes: The related image is merged with the background image to generate a frame image in the interactive multimedia content.
18. The image generation method according to any one of claims 1 to 17, further comprising: Segmenting the image region where the virtual object is located from the relevant image; The image region is merged with a background image to generate a frame image in interactive multimedia content.
19. The image generation method according to any one of claims 1 to 18, wherein: The virtual object includes a person object and a background object, and the composition information includes at least one of position information and size information of the virtual object in the relevant image.
20. An image generating device, comprising: A receiving unit, configured to receive a first attribute information set of a virtual object input by a user; A generating unit is used to generate a related image of the virtual object based on the first attribute information set and an object template, wherein the object template is used to determine at least one of the posture information and composition information of the virtual object in the related image, and the object template matches the first attribute information set.
21. An image generating device, comprising: Memory; and A processor coupled to the memory, wherein the processor is configured to execute the image generation method according to any one of claims 1 to 19 based on instructions stored in the memory.
22. A non-volatile computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the image generating method according to any one of claims 1 to 19 is implemented.
23. A computer program product comprising instructions which, when executed by a processor, cause the processor to perform the image generation method according to any one of claims 1 to 19.
Citation Information
Patent Citations
Multimedia content generation method and apparatus, and device / terminal / server
CN109496295A
Image optimization method and system based on artificial intelligence
CN109767397A
Image fusion transformation
CN109993716A
Data processing method and device, electronic equipment and storage medium
CN112734883A
Image special effect adding method and device, electronic equipment and storage medium
CN113643411A