Image generation method, apparatus and electronic device

CN116645452BActive Publication Date: 2026-09-29NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310450910.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-20
Publication Date
2026-09-29
Estimated Expiration
2043-04-20

AI Technical Summary

Technical Problem

[0002]相关技术中,多人编辑同一个图像,一般是可以对图像进行划分,每人负责编辑其中的一个部分,最后将划分的各个部分进行拼接;但是拼接后的图像可能存在风格不一致、衔接不自然等问题,难以实现完美的融合效果

Benefits of technology

[0011]本发明提供了一种图像生成方法、装置和电子设备,生成虚拟画布区域;接收第一终端设备发送的第一绘图信息;第一绘图信息包括:在虚拟画布区域的第一子区域中绘制的图像内容;接收第二终端设备发送的第二绘图信息;第二绘图信息包括:在虚拟画布区域的第二子区域中绘制的图像内容;对第一绘图信息和第二绘图信息进行融合处理,并生成虚拟画布区域中目标区域显示的目标图像。在多人协作绘画时,自动将多人绘制的多个绘图信息进行融合,生成目标区域显示的目标图像,提高了协同绘画的效率;目标区域至少包括第一子区域和第二子区域之间的空白区域,进而丰富了图像创意效果,满足了多人协作绘图的需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116645452B_ABST
    Figure CN116645452B_ABST
Patent Text Reader

Abstract

The application provides an image generation method, device and electronic equipment, generates a virtual canvas region, receives first drawing information sent by a first terminal device, the first drawing information comprising image content drawn in a first sub-region of the virtual canvas region, receives second drawing information sent by a second terminal device, the second drawing information comprising image content drawn in a second sub-region of the virtual canvas region, performs fusion processing on the first drawing information and the second drawing information, and generates a target image displayed in a target region of the virtual canvas region. When multiple people collaborate to draw, multiple drawing information drawn by the multiple people is automatically fused to generate a target image displayed in a target region, improving the efficiency of collaborative drawing. The target region at least comprises a blank region between the first sub-region and the second sub-region, thereby enriching the image creative effect and meeting the demand of multiple people collaborating to draw.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an image generation method, apparatus, and electronic device. Background Technology

[0002] In related technologies, when multiple people edit the same image, the image can generally be divided, with each person responsible for editing a portion, and then the divided portions are stitched together. However, the stitched image may have inconsistent styles and unnatural transitions, making it difficult to achieve a perfect fusion effect. Alternatively, a version control system can be used, allowing multiple people to edit copies of the same image separately, and then merge the different versions to form the final image. However, this method requires manually merging different versions, which is prone to conflicts, and the merging process can be tedious and time-consuming. Real-time collaborative image editing software can also be used, allowing multiple people to edit the same image online simultaneously, achieving real-time synchronization. However, this method only enables real-time editing, and the creative effects are not ideal. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide an image generation method, apparatus and electronic device, in which the server can automatically merge multiple drawing information drawn by multiple people to generate a target image displayed in the target area when multiple people are collaboratively drawing, so as to improve the efficiency of collaborative creation.

[0004] In a first aspect, embodiments of the present invention provide an image generation method. This method is applied to a server, which is communicatively connected to multiple terminal devices, including at least a first terminal device and a second terminal device. The method includes: generating a virtual canvas region; wherein the virtual canvas region is used to draw a target image, and initially, the virtual canvas region is a blank area; receiving first drawing information sent by the first terminal device; wherein the first drawing information includes image content drawn in a first sub-region of the virtual canvas region; receiving second drawing information sent by the second terminal device; wherein the second drawing information includes image content drawn in a second sub-region of the virtual canvas region; the second sub-region and the first sub-region are different sub-regions within the virtual canvas region; inputting the first drawing information and the second drawing information into a pre-trained image generation model, so as to fuse the first drawing information and the second drawing information through the image generation model, and generate a target image displayed in a target area within the virtual canvas region; wherein the target area includes at least a blank area between the first sub-region and the second sub-region.

[0005] Secondly, embodiments of the present invention provide an image generation method, which provides a drawing interface through a first terminal device, the first terminal device being communicatively connected to a server. The method includes: responding to a first drawing operation on a first drawing area in the drawing interface, determining first drawing information corresponding to the first drawing operation, displaying image content corresponding to the first drawing information in the first drawing area of ​​the drawing interface, and sending the first drawing information to the server; receiving image content corresponding to second drawing information sent by the server, and displaying image content corresponding to the second drawing information in the second drawing area of ​​the drawing interface, wherein the image content corresponding to the second drawing information is generated by the server based on the second drawing information sent by the second terminal device; receiving a target image sent by the server, and displaying the target image in the target drawing area of ​​the drawing interface, wherein the target image is image content generated by the server using the first drawing information, the second drawing information, and a pre-trained image generation model, wherein the target drawing area includes at least a blank area between the first drawing area and the second drawing area.

[0006] Thirdly, embodiments of the present invention provide an image generation apparatus. The apparatus is disposed on a server, and the server is communicatively connected to multiple terminal devices, including at least a first terminal device and a second terminal device. The apparatus includes: a canvas generation module for generating a virtual canvas area; wherein the virtual canvas area is used to draw a target image, and in an initial state, the virtual canvas area is a blank area; a first receiving module for receiving first drawing information sent by the first terminal device; wherein the first drawing information includes image content drawn in a first sub-region of the virtual canvas area; a second receiving module for receiving second drawing information sent by the second terminal device; wherein the second drawing information includes image content drawn in a second sub-region of the virtual canvas area; the second sub-region and the first sub-region are different sub-regions in the virtual canvas area; and an image generation module for inputting the first drawing information and the second drawing information into a pre-trained image generation model, so as to fuse the first drawing information and the second drawing information through the image generation model and generate a target image displayed in a target area of ​​the virtual canvas area; wherein the target area includes at least a blank area between the first sub-region and the second sub-region.

[0007] Fourthly, embodiments of the present invention provide an image generation apparatus that provides a drawing interface via a first terminal device. The first terminal device is communicatively connected to a server. The apparatus includes: a first drawing information sending module, configured to respond to a first drawing operation on a first drawing area in the drawing interface, determine first drawing information corresponding to the first drawing operation, display image content corresponding to the first drawing information in the first drawing area of ​​the drawing interface, and send the first drawing information to the server; a second drawing information receiving module, configured to receive image content corresponding to the second drawing information sent by the server, and display the image content corresponding to the second drawing information in the second drawing area of ​​the drawing interface, wherein the image content corresponding to the second drawing information is generated by the server based on the second drawing information sent by the second terminal device; and a target image receiving module, configured to receive a target image sent by the server, and display the target image in a target drawing area of ​​the drawing interface, wherein the target image is image content generated by the server using the first drawing information, the second drawing information, and a pre-trained image generation model, wherein the target drawing area includes at least a blank area between the first drawing area and the second drawing area.

[0008] Fifthly, embodiments of the present invention provide an electronic device, including a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the image generation method of either the first aspect or the second aspect.

[0009] In a sixth aspect, embodiments of the present invention provide a computer-readable storage medium storing computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the image generation method of either the first aspect or the second aspect.

[0010] The embodiments of the present invention bring the following beneficial effects:

[0011] This invention provides an image generation method, apparatus, and electronic device, which generate a virtual canvas area; receive first drawing information sent by a first terminal device; the first drawing information includes image content drawn in a first sub-region of the virtual canvas area; receive second drawing information sent by a second terminal device; the second drawing information includes image content drawn in a second sub-region of the virtual canvas area; fuse the first and second drawing information to generate a target image displayed in the target area of ​​the virtual canvas area. In collaborative drawing, this method automatically fuses multiple drawing information entries from multiple users to generate a target image displayed in the target area, improving the efficiency of collaborative drawing; the target area includes at least the blank area between the first and second sub-regions, thereby enriching the creative effect of the image and meeting the needs of collaborative drawing.

[0012] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.

[0013] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0014] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0015] Figure 1 A flowchart of an image generation method provided in an embodiment of the present invention;

[0016] Figure 2 This is a schematic diagram of a virtual canvas area provided in an embodiment of the present invention;

[0017] Figure 3 This is a schematic diagram of another virtual canvas area provided in an embodiment of the present invention;

[0018] Figure 4 A schematic diagram of model training provided in an embodiment of the present invention;

[0019] Figure 5 This is a schematic flowchart of an image generation method provided in an embodiment of the present invention;

[0020] Figure 6 A flowchart illustrating another image generation method provided in an embodiment of the present invention;

[0021] Figure 7 A flowchart of another image generation method provided in an embodiment of the present invention;

[0022] Figure 8 A schematic diagram of a drawing interface provided in an embodiment of the present invention;

[0023] Figure 9 This is a schematic diagram of the structure of an image generation device provided in an embodiment of the present invention;

[0024] Figure 10 This is a schematic diagram of another image generation device provided in an embodiment of the present invention;

[0025] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] With the development of technology and the booming creative industries, collaborative creation by multiple people has become an increasingly common need. Traditional collaborative methods have certain limitations in terms of real-time performance, collaborative effects, and creative integration. Regarding multiple people editing the same image, related technologies generally involve dividing the image, with each person responsible for editing a portion, and then stitching the divided portions together. However, the stitched image may have issues such as inconsistent styles and unnatural transitions, making it difficult to achieve a good image effect. Alternatively, a version control system can be used, allowing multiple people to edit copies of the same image separately, and then merge the different versions to form the final image. However, this method requires manually merging different versions, which is prone to conflicts, and the merging process can be tedious and time-consuming. Real-time collaborative image editing software, such as Adobe Creative Cloud and Figma, can also be used, allowing multiple people to edit the same image online simultaneously, achieving real-time synchronization. However, this method only enables real-time editing, and the creative effect is not ideal. Based on this, the embodiments of the present invention provide an image generation method, apparatus, and electronic device, which can be applied to devices such as mobile phones, computers, laptops, computers, and servers.

[0028] To facilitate understanding of this embodiment, a detailed description of an image generation method disclosed in this invention will be provided first. This method is applied to a server, which is communicatively connected to multiple terminal devices, including at least a first terminal device and a second terminal device. Figure 1 As shown, the method includes the following steps:

[0029] Step S102: Generate a virtual canvas area; wherein, the virtual canvas area is used to draw the target image, and in the initial state, the virtual canvas area is a blank area;

[0030] The aforementioned virtual canvas area corresponds to the drawing interfaces provided by multiple terminal devices. In other words, the virtual canvas provides a drawing interface for terminal devices, allowing different users to perform drawing operations on various terminal devices. Specifically, when a user opens the drawing software on a terminal device, the server generates a virtual canvas area, and the terminal device displays the corresponding drawing interface. When the user is not performing any drawing operations, the virtual canvas area is blank, and the drawing interface is also blank.

[0031] The target image mentioned above is typically a fused image generated from multiple drawings by multiple users after the drawings are completed. Alternatively, it can be a fused image generated from multiple drawings by multiple users after the drawings are completed, based on multiple drawings and user-indicated fusion commands. Furthermore, the target image may include one or more.

[0032] Step S104: Receive first drawing information sent by the first terminal device; wherein, the first drawing information includes: image content drawn in a first sub-region of the virtual canvas area;

[0033] Step S106: Receive second drawing information sent by the second terminal device; wherein, the second drawing information includes: image content drawn in a second sub-region of the virtual canvas area; the second sub-region and the first sub-region are different sub-regions in the virtual canvas area;

[0034] The aforementioned first drawing information provides a drawing interface through a first terminal device, and is determined in response to a first drawing operation on a first drawing area of ​​that drawing interface, which corresponds to the aforementioned first sub-area. Similarly, the aforementioned second drawing information provides a drawing interface through a second terminal device, and is determined in response to a second drawing operation on a second drawing area of ​​that drawing interface. The image content drawn in the first and second sub-areas can be user-drawn images, image materials from a resource library, or intelligently generated images, etc.

[0035] Specifically, in collaborative creation, different users will draw drawing information on different terminal devices, and different users will draw drawing information in different areas of the same drawing interface. The terminal devices will send the drawing information drawn by the users to the server.

[0036] Step S108: Input the first drawing information and the second drawing information into the pre-trained image generation model, so as to fuse the first drawing information and the second drawing information through the image generation model and generate a target image displayed in the target area of ​​the virtual canvas area; wherein, the target area includes at least the blank area between the first sub-region and the second sub-region.

[0037] The target area mentioned above may include only the blank area between the first sub-region and the second sub-region; it may also include a portion of the first sub-region, a portion of the second sub-region, and the blank area between the first and second sub-regions; or it may be a user-specified area, such as the area surrounding the first sub-region, or the area surrounding the second sub-region, or the area surrounding the first and second sub-regions.

[0038] If the target area can consist only of the blank area between the first and second sub-regions, the image content of the target area in the aforementioned virtual canvas area includes at least a portion of the image content from the first and second drawing information. Specifically, the image content in the target area near the first sub-region includes a significant amount of the first drawing information; similarly, the image content in the target area near the second sub-region includes a significant amount of the second drawing information. In other words, during the fusion process and generation of the target image for display, image content from the first drawing information is collected in the target area near the first sub-region, and image content from the second drawing information is collected in the target area near the second sub-region. Fusion is then performed based on the collected image content.

[0039] If the target area includes a portion of a first sub-region, a portion of a second sub-region, and a blank area between the first and second sub-regions, the target image displayed on the portion of the first sub-region typically has a certain degree of transparency (allowing the image content of the first drawing information in the portion of the first sub-region to be seen through the target image), or the image content of the target image displayed on the portion of the first sub-region is typically blended with the image content of the first drawing information in the portion of the first sub-region. Similarly, the target image displayed on a portion of the second sub-region also typically has a certain degree of transparency (allowing the image content of the second drawing information in the portion of the second sub-region to be seen through the target image), or the image content of the target image displayed on a portion of the second sub-region is typically blended with the image content of the second drawing information in the portion of the second sub-region.

[0040] Specifically, the first drawing information and the second drawing information are input into a pre-trained image generation model, and the first drawing information and the second drawing information are fused to generate a fusion feature of the first drawing information and the second drawing information; then, a target image displayed in the target area of ​​the virtual canvas region is generated based on the fusion feature; wherein, the target image includes image features of at least part of the image content from multiple drawing information.

[0041] For example, if the first drawing information includes the ocean and the second drawing information includes a beach, a transition image between the ocean and the beach can be automatically generated in the middle area between the first and second drawing information. Or, if the first drawing information includes grassland and the second drawing information includes trees, blue sky and white clouds or the moon and stars at night can be automatically generated in the area above the first and second drawing information.

[0042] This invention provides an image generation method that generates a virtual canvas area. The virtual canvas area is used to draw a target image, and initially, it is a blank area. The method involves receiving first drawing information from a first terminal device, where the first drawing information includes image content drawn in a first sub-region of the virtual canvas area. It also involves receiving second drawing information from a second terminal device, where the second drawing information includes image content drawn in a second sub-region of the virtual canvas area. The second sub-region and the first sub-region are different sub-regions within the virtual canvas area. The first and second drawing information are input into a pre-trained image generation model to fuse the first and second drawing information, generating a target image displayed in the target area of ​​the virtual canvas. In collaborative drawing, this method automatically fuses multiple drawing information entries from multiple users to generate a target image displayed in the target area, improving the efficiency of collaborative drawing. The target area includes at least the blank area between the first and second sub-regions, thus enriching the creative effect of the image and meeting the needs of collaborative drawing.

[0043] The above method also includes displaying the target image in the blank area between the first sub-region and the second sub-region of the target region.

[0044] For example, such as Figure 2 As shown, the target area is the blank area between the first sub-region and the second sub-region, and the target image is displayed in this blank area.

[0045] The target area mentioned above includes the blank area between the first sub-region and the second sub-region. This target area can be automatically determined by the server or indicated by the user.

[0046] In another embodiment, the method further includes: determining a target region in the virtual canvas region, wherein the target region includes a portion of the first sub-region, a portion of the second sub-region, and a blank region between the first sub-region and the second sub-region.

[0047] The target image is displayed in a portion of the first sub-region, a portion of the second sub-region, and the blank area between the first and second sub-regions within the target area. The image content in the target area is continuous with the image content in the first and second sub-regions, respectively.

[0048] For example, such as Figure 3 As shown, the target area includes a portion of the first sub-region, a portion of the second sub-region, and a blank area between the first and second sub-regions, and the target image is displayed in this target area.

[0049] The continuity of image content in the target area with image content in the first sub-region and the second sub-region typically means that the image content in the target area is connected to, merged with, or overlaps with the image content in the first and second sub-regions. This method enables a transition effect between the newly generated target image and the image content created collaboratively by multiple people, further improving the image quality of collaborative drawing.

[0050] Specifically, the target area mentioned above can be an area automatically determined by the server. For example, the target area can be determined in the virtual canvas area, including a portion of the first sub-region, a portion of the second sub-region, and the blank area between the first and second sub-regions.

[0051] The target area described above can also be a user-specified area. For example, if a user draws an area in the drawing interface of a terminal device, that drawn area can be designated as the target area. This target area can be the blank area between the first and second sub-regions; it can also include a portion of the first sub-region, a portion of the second sub-region, and the blank area between the first and second sub-regions; furthermore, it can include blank areas at the edges of the first sub-region, or blank areas at the edges of the second sub-region, or a portion of the first sub-region and the blank area at the edges of the first sub-region (the portion and the blank area at the edges are connected), or a portion of the second sub-region and the blank area at the edges of the second sub-region (the portion and the blank area at the edges are connected).

[0052] In the image content of the first image region in the target image, the image content in the first drawing information is greater than the image content in the second drawing information; in the image content of the second image region in the target image, the image content of the second drawing information is greater than the image content of the first drawing information; the distance between the first image region and the first sub-region is less than the distance between the second image region and the first sub-region.

[0053] The aforementioned first image region typically refers to the region near the first sub-region, and the aforementioned second image region typically refers to the region near the second sub-region. In other words, in the final generated target image, the image content near the first sub-region contains more image content from the first drawing information than from the second drawing information. Or, the closer the image content is to the first sub-region in the target image, the more image content it contains from the first drawing information. Similarly, in the final generated target image, the image content near the second sub-region contains more image content from the second drawing information than from the first drawing information. Or, the closer the image content is to the second sub-region in the target image, the more image content it contains from the second drawing information. This target image can produce a transition effect between the first and second drawing information, or an edge transition effect between the first and second drawing information, further improving the natural transition and creative effect of collaborative creation.

[0054] In one possible implementation, the target image includes a first target image. The step of inputting the first drawing information and the second drawing information into a pre-trained image generation model to fuse the first drawing information and the second drawing information through the image generation model, and generating the image content of the target area in the virtual canvas region to obtain the target image, can be implemented in one possible way:

[0055] If the user does not specify the display location of the target image, the target area will be determined as the first target area. In other words, the server will automatically determine the first target area and generate the first target image to be displayed in the first target area.

[0056] The first drawing information and the second drawing information are respectively input into the image encoder of the pre-trained image generation model to obtain the first image feature of the first drawing information and the second image feature of the second drawing information. The first image feature and the second image feature are input into the fusion layer of the image generation model. In the multimodal feature space of the image generation model, the first image feature and the second image feature are fused to generate the first fused feature of the image content of the first target area in the virtual canvas area. The multimodal feature space stores the correspondence between different types of features and is used to fuse different types of features. The first fused feature is input into the image decoder of the image generation model to generate the first target image displayed in the first target area of ​​the virtual canvas area. The first target image includes at least a portion of the image content of the first drawing information and at least a portion of the image content of the second drawing information.

[0057] The image generation model described above includes a multimodal feature space, which stores the correspondence between text features and various image features, and is used to fuse different types of features. The image generation model can fuse different images, and of course, it can also fuse images and text.

[0058] Specifically, the image generation model mentioned above includes modules such as an encoder, a fusion module, and a decoder. Specifically, the encoder in the image generation model obtains image features of the drawing information. Then, feature fusion is performed on the image features in a multimodal feature space (Latent Space) to obtain fused features of multiple drawing information. Since the multimodal feature space stores the correspondence between various image features, multiple image features can be mixed in the fusion module. For example, multiple image features can be directly concatenated and added to obtain the fused feature, or weighted fusion can be performed on the image features according to preset weight values ​​to obtain the aforementioned fused feature, or image features can be fused through fully connected layers, attention mechanisms, etc. Specifically, multiple image features can be directly added, or the weights of each feature can be determined based on the distance between each feature, and finally, a weighted average of the features is performed based on the weights to obtain the fused feature.

[0059] The image generation model described above can be built based on machine learning networks. Machine learning networks can include Convolutional Neural Networks (CNNs), Deconvolutional Neural Networks (DNs), Deep Neural Networks (DNNs), Deep Convolutional Inverse Graphics Networks (DCIGNs), Generative Adversarial Networks (GANs), and Variational Autoencoder (VAE) models, among others. A Variational Autoencoder is a generative model that learns latent representations of data and can generate new data from these representations. VAEs introduce stochasticity between the encoder and decoder, enabling them to learn a continuous latent space.

[0060] The image decoder mentioned above can be a generator in a GAN, U-Net, or a decoder in a Variational Autoencoder (VAE), among others. U-Net is an improved FCN (Fully Convolutional Network) structure. Specifically, U-Net is a fully convolutional network structure with skip connections, which makes the network more efficient at merging features from different levels. The VAE decoder is responsible for mapping the latent vectors back to the original data space. For image data, the VAE decoder is typically a neural network using transposed convolutional layers or upsampling layers, which decodes the latent vectors into an output image with the same size as the input image.

[0061] The target image typically includes features of the image content involved in the fusion. For example, if multiple plotting information includes an image of a tree and an image of a mountain, the target image obtained after fusing these images will generally include both trees and mountains.

[0062] Additionally, it should be noted that the images generated by the aforementioned image generation models typically possess a certain image style. The specific style type of the image is usually determined by the training samples. If the image style of the training samples is the original art style when training the model parameters, the image generated by the aforementioned image generation model will be closer to the original art style; if the image style of the training samples is the realistic style when training the model parameters, the image generated by the aforementioned image generation model will be closer to the realistic style, and so on.

[0063] In this process, image features from different sources are fused in a unified manner within the multimodal feature space to obtain entirely new fused features. Since the multimodal feature space is trained from a large number of existing image features, the fused features represent the desired image features within this space. Subsequently, a pre-trained image decoder is used to reconstruct these features, yielding the final image that incorporates the required feature information—the target image described above.

[0064] The above embodiments describe the steps for generating the first target image when the user does not specify image fusion information. In this embodiment, before inputting the first drawing information and the second drawing information into the pre-trained image generation model, the method further includes:

[0065] Receive image fusion information sent by any one of multiple terminal devices; wherein the image fusion information is used to indicate a second target area and the specified image content included in the second target area.

[0066] If the user specifies the display location of the target image, the target area will be designated as the second target area. In other words, the server will determine the second target area based on the user's instructions and generate a second target image to be displayed in the second target area.

[0067] The aforementioned image fusion information includes text or image content drawn in a designated sub-region of the virtual canvas area. This designated sub-region can be an editing area or a drawing area. Specifically, the second target area can be indicated by the text content in the image fusion information. For example, the text content may include location information (such as blank areas at the edges of the first sub-region), and the second target area can be determined by this location information. Alternatively, the second target area can be indicated by the designated sub-region where the image fusion information is located. For example, the designated sub-region where the image fusion information is located may be the second target area. Furthermore, the designated image content included in the second target area can also be indicated by the text content in the image fusion information. For example, the text content may include image information (such as a stream), and the designated image content in the second target area (which includes an image of a stream) can be determined by this image information.

[0068] In practice, any user can draw image blending information within the drawing interface. Possible methods include inputting image blending information in any editing area of ​​the drawing interface, drawing image blending information in any drawing area of ​​the drawing interface, or inputting or drawing image blending information in a preset area of ​​the drawing interface (such as a user-defined area). In this method, after completing the drawing information, the user can also draw image blending information to indicate the display area and image content of the target image. This allows users to independently control the generated target image without having to draw the image content themselves, further improving the efficiency of collaborative creation and meeting the personalized needs of the target image.

[0069] Furthermore, one possible approach to the steps described above, which involve inputting the first and second drawing information into a pre-trained image generation model to fuse the first and second drawing information and generate a target image displayed in the target area of ​​the virtual canvas, is as follows:

[0070] The first drawing information, the second drawing information, and the image fusion information are input into a pre-trained image generation model to fuse the first drawing information, the second drawing information, and the image fusion information through the image generation model, and generate a second target image displayed in the second target area of ​​the virtual canvas area; wherein, the second target image includes the specified image content indicated by the image fusion information, as well as at least a portion of the image content of the first drawing information and at least a portion of the image content of the second drawing information.

[0071] For example, if the first drawing information includes grass and people, the second drawing information includes trees, and the image fusion information includes text content such as "night stars" or "night stars" located around the first and second sub-regions, then night stars can be automatically generated in the area above or around the first and second drawing information. As another example, if the first drawing information includes buildings, the second drawing information includes roads, and the image fusion information includes text content such as "greenery" in the middle, or the image fusion information includes text content such as "greenery" located between the first and second sub-regions, then greenery can be generated between the first and second drawing information.

[0072] Specifically, the first and second drawing information are input into the image encoder of the pre-trained image generation model to obtain the first image features of the first drawing information and the second image features of the second drawing information; the image fusion information is input into the text encoder of the image generation model to obtain the text features of the image fusion information; the first image features, the second image features, and the text features are input into the fusion layer of the image generation model, and feature fusion is performed on the first image features, the second image features, and the text features in the multimodal feature space of the image generation model to generate the second fused features of the image content of the second target area in the virtual canvas area; the second fused features are input into the image decoder of the image generation model to generate the second target image displayed in the second target area of ​​the virtual canvas area.

[0073] The image encoders described above can use convolutional neural networks (CNNs) such as EfficientNet or variational autoencoders (VAEs). The text encoders described above can use Transformer-based language models such as Bidirectional Encoder Representations from Transformers (BERT), GPT-3 (Generative Pre-training Transformer), or T5 (text-to-text transfer transformer). The aforementioned prior sub-models are used to map (transform) text features into image features.

[0074] Specifically, separate encoders are used to extract text and image features. For the text encoder, language models such as GPT (Generative Pre-training Transformer) can be used; for the image encoder, convolutional neural networks (CNNs) such as ResNet or VGG can be used, but are not limited to.

[0075] In practical implementation, before inputting the first and second drawing information into the image encoder of the image generation model, multiple images are typically preprocessed. Preprocessing multiple images ensures input consistency and reduces computational burden. Preprocessing operations include image scaling and normalization to ensure all images have the same format. After processing, the multiple images are input into the image encoder separately, and image feature extraction is performed to obtain the image features of each image. For text preprocessing, the goal is to convert the input natural language text into a form that the model can understand, including text segmentation and word quantization. The preprocessed text is then input into the text encoder, where text feature extraction is performed to obtain the text features.

[0076] Specifically, after inputting the image fusion information into the text encoder of the image generation model to obtain the text features of the image fusion information, the above method further includes: inputting the text features into the prior sub-model of the image generation model to obtain the mapped text features. That is, mapping the text features to the corresponding position points of the image features in the multimodal feature space.

[0077] The text and image features in the multimodal feature space are learned by training the image generation model using pre-defined sample data, which includes multiple matching image and text samples. This sample data typically comes from multimodal datasets with extensive annotations, such as the COCO (Common Objects in Context) dataset and the Visual Genome (VG) dataset, where each image sample is matched with a corresponding text description.

[0078] When fusing features, various methods can be employed, such as weighted summation, concatenation, or attention mechanisms. For example, image features and text features (mapped text features) can be directly added together, or feature weights can be determined based on the distance between features, and then the image and text features can be fused using these weights to obtain the fused features. Furthermore, attention mechanisms can be used to weight and fuse features, allowing the model to focus on more important feature components. Typically, the second target image is based on the image corresponding to the user-input text, and then image features from multiple user-uploaded images are fused into this image.

[0079] Specifically, in one possible approach, in the multimodal feature space of the image generation model, the text feature vector (corresponding to the text features mentioned above) and the image feature vector (corresponding to the image features mentioned above) are concatenated along the same dimension. Assuming the dimension of the text feature vector is d1 and the dimension of the image feature vector is d2, then after the concatenation operation, the dimension of the fused feature vector will become (d1+d2).

[0080] In another approach, different weights are assigned to text features and image features, and then the weighted features are summed to obtain the fused features. This is to highlight the more important feature components during the weighted summation process, thus obtaining the fused features. For example, if α and β are the weights of text features and image features, respectively, then the fused feature would be α text features + β image features. Typically, the weights are learned through a training process. They can also be manually defined by the user at input time, such as favoring text features or image features more in feature fusion.

[0081] Another approach involves mapping textual and image features to a shared latent space (i.e., a multimodal feature space), allowing these features to be compared and fused within the same space. This can be achieved by learning two independent mapping functions (such as linear transformations, neural networks, etc.) to map textual and image features to a latent space of the same dimension. Once mapped to the same latent space, any fusion method (such as weighted summation, concatenation, etc.) can be used to integrate the two features.

[0082] In another approach, attention mechanisms can be used to determine the most relevant target features among text and image features within a specific context. By applying attention weights to each target feature, the relative contributions of text and image features in a given context can be dynamically adjusted. This method is often combined with other fusion methods to achieve more complex feature fusion.

[0083] In another approach, bidirectional mapping is a method for calculating the interaction between two features, which can be used for fusing text and image features. A new fused feature can be calculated by multiplying the text features by a matrix M and then multiplying the result by the image features. The matrix M can be learned during training.

[0084] The image samples mentioned above can be images of the same style or different styles. When the samples are images of the same style, the image generation model can generate images of a fixed style. However, when the samples contain images of different styles, the image generation model has wider applicability and can generate images of multiple styles.

[0085] One possible implementation:

[0086] In the multimodal feature space of the image generation model, the weight values ​​of the first image feature, the second image feature, and the text feature are determined. Based on the weight values, the first image feature, the second image feature, and the text feature are weighted and averaged to obtain the second fused feature.

[0087] Specifically, in the multimodal feature space of the image generation model, a target region is determined, which includes multiple locations between the first image feature and the second image feature; for each location, a first distance between the location and the first image feature, a second distance between the location and the second image feature, and a third distance between the location and the text feature are determined; based on the first distance, the second distance, and the third distance, the weight values ​​of the first image feature at the location, the weight values ​​of the second image feature at the location, and the weight values ​​of the text feature at the location are calculated.

[0088] The features stored in the aforementioned multimodal feature space are multidimensional, with both the first and second image features including at least one feature vector. Specifically, multiple positions can be determined on the line connecting the first and second image features. These positions are used to store the fused features of the first image, the second image, and the text. For each position, a first distance from the position to the first image feature, a second distance from the position to the second image feature, and a third distance from the position to the text feature are determined. The sum of the first, second, and third distances is calculated. The ratio of the second distance to the sum of distances is determined as the weight value of the first image feature at that position, the ratio of the first distance to the sum of distances is determined as the weight value of the second image feature at that position, and the ratio of the third distance to the sum of distances is determined as the weight value of the text feature at that position.

[0089] Specifically, for each location, the weight value of the first image feature at that location is calculated as a first product of the first image feature and the first image feature itself; the weight value of the second image feature at that location is calculated as a second product of the second image feature and the second image feature itself; and the weight value of the text feature at that location is calculated as a third product of the text feature and the text feature itself. The sum of the first, second, and third product values ​​is then calculated to obtain the fused feature at that location.

[0090] For example, in the final target image, the portion closer to the first image uses feature representations that are closer to the features of the first image, and the portion closer to the second image uses feature representations that are closer to the features of the second image. This allows the generated target image to have a more natural transition between the first and second images.

[0091] Furthermore, the above image generation model is trained in the following manner: multiple matching image samples and text samples are acquired, wherein the image samples and text samples have matching identifiers; the image samples are input into the image encoder of the image generation model to be trained to obtain the image sample features of the image samples; the image sample features of the image samples and the text sample features of the text samples are input into the fusion layer of the image generation model to be trained to output the pairing results of the image samples and text samples; wherein, the fusion layer is used to learn the correspondence between the text sample features and the image sample features; based on the loss function and the identifiers carried by the image samples and text samples, the loss value of the pairing results is calculated; based on the loss value, the model parameters of the image generation model are updated until the loss value meets the preset loss threshold.

[0092] To train the text encoder and image encoder, as well as the multimodal feature space, in the image generation model, a multi-task learning approach can be employed. First, for the text encoder, a Masked Language Model (MLM) task is used to learn the text representation. For the image encoder, an image classification or object detection task is used to learn the image representation. Simultaneously, a contrastive learning task is introduced, which maximizes the similarity between positive sample pairs while minimizing the similarity between negative sample pairs. Specifically, the fusion layer outputs the pairing results of image and text samples. If the image and text in the pairing results do not match, the model parameters need to be updated and training continues; if the image and text in the pairing results match, training stops, resulting in the trained text encoder, image encoder, and fusion layer.

[0093] In addition, the purpose of the above training is to enable the model to learn the correspondence between text features and image features. After training is completed, a model that can map text features and image features to a unified feature space (Latent Space) can be obtained, namely the multimodal feature space mentioned above.

[0094] In practice, to enable image generation models to understand the correspondence between text features and image features, multimodal learning methods can be used. Multimodal learning aims to extract information from multiple types of data sources (such as images and text). A multimodal feature space refers to a representation that fuses features from different modalities (such as text and images) into a single latent space. To obtain a model that can map text features and image features to a unified feature space, the following steps can be taken:

[0095] Extracting features from text and images separately: First, separate models are used to extract text and image features. For text, a pre-trained language model such as GPT can be used; for images, a pre-trained convolutional neural network (CNN) such as ResNet or VGG can be used, but is not limited to. Then, a fusion layer (such as a fully connected layer or an attention mechanism) is used to combine the text and image features. The goal of the fusion layer is to learn how to map both types of features to a shared latent space. The representation in this shared space can capture the associations and interdependencies between text and image features.

[0096] Finally, to train this image generation model, a loss function needs to be designed so that the model can learn to understand the correspondence between text features and image features during training. Common loss functions include contrastive loss and triplet loss. These loss functions encourage the model to bring related text features and image features closer together in the latent space, while keeping unrelated features further apart. By minimizing the loss function, the model can learn how to map text and image features to a shared latent space.

[0097] Optional, such as Figure 4 In the exemplary implementation shown, the image encoder and text encoder can be trained first through CLIP (Contrastive Language-Image Pre-Training, which trains a visual pre-trained model with strong transferability using the supervision signal of the text), and then a prior sub-model can be trained based on the paired text features and image features.

[0098] The above method further includes: inputting the image sample features of the target image sample into the image decoder of the image generation model to be trained to obtain the training image corresponding to the target image sample; calculating the loss value between the training image and the target image sample; and updating the parameters of the image decoder based on the loss value until the loss value meets the preset loss threshold.

[0099] To train the decoder module (i.e., the image decoder), it is necessary to reconstruct the real image from image features, such as... Figure 4 As shown. This process is similar to that of an autoencoder, reconstructing the input image from an intermediate feature layer, but not exactly the same. The image encoder trained based on the above method can obtain image features zi from the CLIP image x, and input the image features zi into the image decoder to reconstruct an image with the same semantics as the CLIP image x, but not completely identical to x.

[0100] Furthermore, after the step of inputting the text sample into the text encoder of the image generation model to be trained to obtain the text sample features of the text sample, the above method also includes: inputting the text sample features of the target text sample into the prior sub-model of the image generation model to be trained to obtain the mapped text sample features; calculating the loss value between the image sample features of the matching image sample that matches the target text sample and the mapped text sample features; updating the model parameters in the prior sub-model based on the loss value until the loss value meets the preset loss threshold.

[0101] Take the trained text encoder and input text y to obtain the text encoding zt. Similarly, take the trained image encoder and input image x to obtain the image encoding zi. The goal is for the prior module to obtain the corresponding zi from zt. Assuming the feature output of zt after the prior sub-model is zi′, we want zi′ to be as close to zi as possible. Specifically, we update the prior module (i.e., the prior sub-model) by calculating the loss value using a loss function.

[0102] Finally, the trained prior sub-model will be concatenated with the text encoder, which will generate the corresponding image-coded features zi (i.e., the mapped text features) based on the input text y.

[0103] Furthermore, the above method also includes: inputting the mapped text sample features into the image decoder of the image generation model to be trained to obtain the training image corresponding to the text sample; calculating the loss value between the training image and the image sample matched with the text sample; and updating the parameters of the image decoder based on the loss value until the loss value meets the preset loss threshold.

[0104] In addition to generating images from images, images can also be generated from text. To train a decoder module (i.e., an image decoder), the corresponding image can be restored from text features. The text encoder trained based on the above method can obtain mapped text features from the text, and input the mapped text features into the image decoder to restore an image with the same semantics as the text.

[0105] See Figure 5The flowchart of the image generation method shown is as follows: In one specific implementation, the text description is processed by a text encoder to obtain text features, the user-uploaded image is processed by an image encoder to obtain image features, the text features and image features are fused to obtain fused features, and finally processed by an image decoder to obtain the final image (i.e., the target image mentioned above).

[0106] The method further includes: if the first drawing information includes text content for drawing an image in a first sub-region of the virtual canvas area, inputting the first drawing information into a pre-trained image generation model to generate an image corresponding to the text content displayed in the first sub-region of the virtual canvas area; displaying the image corresponding to the text content in the first sub-region; then sending the image corresponding to the text content to each terminal device, and displaying the image corresponding to the text content in the first drawing area of ​​the drawing interface provided by the terminal device; wherein the first drawing area corresponds to the first sub-region.

[0107] The aforementioned first drawing information may be text content entered by the user in the editing area of ​​the drawing interface. This text content is used to instruct the display of the image corresponding to the text content in the first sub-area of ​​the virtual canvas area.

[0108] Specifically, the first drawing information is input into the text encoder of the image generation model to obtain the text features; the text features are input into the prior sub-model of the image generation model to obtain the text features; and the text features are input into the image decoder of the image generation model to generate the image corresponding to the text.

[0109] See Figure 6 As shown, if a user only uploads text, the text features can be directly obtained through a text editor. These text features are then used to obtain the image features of the text through a prior sub-model. Finally, an image decoder is used to obtain the image corresponding to the text. For example, if the text above is "Generate a white cat with a red background in the first sub-region," then the final generated image will include a white cat with a red background, displayed in the first sub-region.

[0110] In addition, after the steps of inputting the first drawing information and the second drawing information into the pre-trained image generation model to fuse the first drawing information and the second drawing information through the image generation model and generate the target image displayed in the target area of ​​the virtual canvas area, the above method further includes: sending the target image to each terminal device and displaying the target image in a designated drawing area of ​​the drawing interface provided by the terminal device; wherein the designated drawing area corresponds to the target area.

[0111] Specifically, after the server generates the target image, it will display the target image in the drawing interface provided by each terminal device, so that each user can see the target image automatically generated based on the content of the user's drawing and the fusion information indicated.

[0112] In the above approach, an image generation model is used to map text features and image features to a unified feature space, and then fuse them within that space to generate new image content. Through multi-task learning and contrastive learning, the model can effectively capture the correlation between text descriptions and image content. Simultaneously, through server-side deployment and a real-time update mechanism, the function of multiple users collaboratively editing the same image online is realized. Compared to traditional image editing methods, this technology has significant advantages in automation, collaboration, and creative expression.

[0113] To map and fuse text and image features into a unified feature space, a fusion layer needs to be designed to combine these two types of features. Two common feature fusion methods are: using fully connected layers and attention mechanisms.

[0114] 1. Fully Connected Layer: A fully connected layer is a simple and direct feature fusion method. Assume that text features and image features have been extracted. The specific fusion steps are as follows:

[0115] a. First, concatenate the text feature vector (dimension d_t) and the image feature vector (dimension d_i). For example, if the text feature vector has dimension d_t and the image feature vector has dimension d_i, then the concatenated feature vector will have dimension d_t + d_i and will be concatenated with the image feature vector.

[0116] b. The concatenated feature vectors are mapped to the shared latent space through a fully connected layer. The weight matrix of the fully connected layer has a size of (d_t+d_i)*d_l, where d_l is the dimension of the shared latent space.

[0117] c. An activation function (such as ReLU or tanh) can be added after the fully connected layer to introduce nonlinearity.

[0118] 2. Attention Mechanism: The attention mechanism is a more complex and flexible feature fusion method. It allows the model to automatically learn how to assign different weights based on the input text and image features. The specific fusion steps are as follows:

[0119] a. First, calculate the similarity matrix between the text feature vectors and the image feature vectors. Dot product or cosine similarity can typically be used as the similarity measure.

[0120] b. Then, normalize the similarity matrix (e.g., using the softmax function) to obtain the attention weight matrix.

[0121] c. Next, the text feature vector and image feature vector are weighted and summed using an attention weight matrix to obtain a fused feature vector. This fused feature vector contains both textual and image information.

[0122] d. The fused feature vectors can be mapped to a shared latent space through a fully connected layer, and an activation function can be added after the fully connected layer to introduce non-linearity.

[0123] In summary, both fully connected layers and attention mechanisms can achieve the fusion of textual and image features. Fully connected layers are relatively simple, while attention mechanisms offer greater flexibility.

[0124] Furthermore, in the process of enabling collaborative drawing among multiple users, different types of neural networks can be used, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), or variational autoencoders (VAEs). These networks can achieve image feature extraction, fusion, and generation to varying degrees, or can be combined to achieve more efficient results. Each of these networks has its own advantages in achieving image feature extraction, fusion, and generation, and combining them can lead to more efficient results.

[0125] To improve the real-time nature and convenience of multi-user collaboration, real-time collaboration functionality can be incorporated into the technical implementation, allowing multiple users to edit simultaneously on the same canvas. Simultaneously, a version control system should be introduced to facilitate users in revisiting historical operations and viewing and restoring previous editing states.

[0126] When training image generation models, transfer learning techniques can be considered. This involves applying a pre-trained neural network model to a specific task, thereby reducing training time and computational resources. Through transfer learning, high-performance models can be obtained in a shorter time, thus improving the effectiveness of collaborative drawing processes.

[0127] To enhance the personalized experience of multi-user collaboration, collaborative filtering technology and recommendation systems can be employed. Based on a user's drawing history and preferences, these systems can recommend materials, colors, and drawing techniques that are similar to the user's style. This will help users better unleash their creativity and improve the overall collaborative effect.

[0128] The technology in this patent can be integrated with existing image editing software (such as Adobe Photoshop, GIMP, etc.) as a plugin or extension module, thereby facilitating multi-user collaboration and drawing in a familiar environment. In practical applications, the appropriate solution can be selected based on actual needs, scenarios, and resource conditions.

[0129] This invention provides an image generation method, which provides a drawing interface through a first terminal device. The first terminal device is communicatively connected to a server, such as... Figure 7 As shown, the method includes the following steps:

[0130] Step S702: Respond to the first drawing operation for the first drawing area in the drawing interface, determine the first drawing information corresponding to the first drawing operation, and display the first drawing information in the first drawing area of ​​the drawing interface;

[0131] Step S704: Receive the image content corresponding to the second drawing information sent by the server, and display the image content corresponding to the second drawing information in the second drawing area of ​​the drawing interface. The image content corresponding to the second drawing information is generated by the server based on the second drawing information sent by the second terminal device.

[0132] Step S706: Receive the target image sent by the server and display the target image in the target drawing area of ​​the drawing interface. The target image is the image content generated by the server using the first drawing information, the second drawing information, and the pre-trained image generation model. The target drawing area includes at least the blank area between the first drawing area and the second drawing area.

[0133] The image content in the target drawing area is continuous with the image content in the first drawing area and the image content in the second drawing area, respectively.

[0134] In one possible approach, before performing the first and second drawing operations, the user can determine the user drawing area (i.e., the first drawing area and the second drawing area mentioned above, corresponding to the first sub-area and the second sub-area mentioned above) and the smart drawing area (i.e., the designated drawing area, corresponding to the designated sub-area mentioned above).

[0135] The user-selected and intelligent drawing area can be determined by the user, or the server can automatically determine the user's drawing area based on the drawing display area after the drawing is completed in the drawing interface. The intelligent drawing area may or may not overlap with the user-selected drawing area.

[0136] In practice, users can input text and images in any area of ​​the drawing interface, and the area where the text or image is displayed can be defined as the user's drawing area. Alternatively, the user can pre-define the drawing area and then input text or images within that area.

[0137] The first and second drawing information mentioned above can be text, text from existing materials, or images from existing materials.

[0138] For example, such as Figure 8 As shown, colleague A's drawing area is the user's corresponding user drawing area, and colleague B's drawing area is the user's corresponding user drawing area. Colleague A can draw in their respective drawing area, and colleague B can draw in their respective drawing area. Finally, the merged image (target image) will be displayed in the intelligent drawing area.

[0139] In the above method, when multiple people are drawing collaboratively, the multiple drawing information drawn by multiple people is automatically merged to generate a target image displayed in the target area, which improves the efficiency of collaborative drawing; the target area includes at least the blank area between the first sub-region and the second sub-region, thereby enriching the creative effect of the image and meeting the needs of multiple people drawing collaboratively.

[0140] The above method has the following beneficial effects: Improved creative efficiency: Through artificial intelligence algorithms, target images are automatically filled in during multi-person collaboration, reducing the time and effort required for manual image editing and adjustment, thereby greatly improving creative efficiency. For example, in fields such as game design, film production, and advertising design, artists and designers can use this technology to complete complex scene designs and character creation more quickly.

[0141] Enhanced collaboration: Multiple people can collaboratively edit the same image online in real time, without switching between different platforms or tools or transferring files. This online collaboration method greatly improves the ease of cross-regional and cross-departmental cooperation. For example, members of a design team located around the world can easily collaborate on the same canvas using this technology, thereby reducing communication costs.

[0142] Enriching Creative Expression: Intelligent painting algorithms can automatically generate images with a certain artistic style and creativity based on user input, thus enriching creative possibilities. For example, in the process of illustration creation, artists can use AI-generated images as creative references to inspire them to create more attractive and innovative works.

[0143] Personalized customization: Based on the user's text description and input, image content that meets the user's personalized needs can be generated. For example, the user can enter a specific text description, such as "a forest at night, a fiery sunset," and the system will generate a corresponding scene based on the description, achieving a personalized visual effect.

[0144] Lowering the learning barrier: For ordinary users without professional drawing skills, this technology lowers the learning threshold for image creation. Users only need to input text descriptions to generate images with a certain artistic value. This allows more people to participate in the image creation process, broadening the audience for image creation.

[0145] Scalability: The model and algorithm of this technical solution have excellent scalability and can be applied to various scenarios, such as animation production, architectural design, and fashion design. By training the model in a targeted manner, it can adapt to the image creation needs of different fields.

[0146] Corresponding to the above method embodiments, this invention provides an image generation apparatus, which is disposed on a server. The server is communicatively connected to multiple terminal devices, including at least a first terminal device and a second terminal device; for example... Figure 9 As shown, the device includes:

[0147] Canvas generation module 91 is used to generate a virtual canvas area; wherein, the virtual canvas area is used to draw the target image, and in the initial state, the virtual canvas area is a blank area;

[0148] The first receiving module 92 is used to receive first drawing information sent by the first terminal device; wherein, the first drawing information includes: image content drawn in a first sub-region of the virtual canvas area;

[0149] The second receiving module 93 is used to receive second drawing information sent by the second terminal device; wherein, the second drawing information includes: image content drawn in a second sub-region of the virtual canvas area; the second sub-region and the first sub-region are different sub-regions in the virtual canvas area;

[0150] The image generation module 94 is used to input the first drawing information and the second drawing information into the pre-trained image generation model, so as to fuse the first drawing information and the second drawing information through the image generation model and generate a target image displayed in the target area of ​​the virtual canvas area; wherein, the target area includes at least the blank area between the first sub-region and the second sub-region.

[0151] In the above method, when multiple people are drawing collaboratively, the multiple drawing information drawn by multiple people is automatically merged to generate a target image displayed in the target area, which improves the efficiency of collaborative drawing; the target area includes at least the blank area between the first sub-region and the second sub-region, thereby enriching the creative effect of the image and meeting the needs of multiple people drawing collaboratively.

[0152] The aforementioned device further includes a first target image display module, used to display the target image in the blank area between the first sub-region and the second sub-region of the target region.

[0153] The aforementioned device further includes a target area determination module, used to: determine a target area in a virtual canvas area, wherein the target area includes a portion of a first sub-area, a portion of a second sub-area, and a blank area between the first and second sub-areas.

[0154] The aforementioned device further includes a second target image display module, used to display the target image in a portion of a first sub-region, a portion of a second sub-region, and a blank area between the first and second sub-regions within the target region.

[0155] The image content in the target region is continuous with the image content in the first sub-region and the image content in the second sub-region, respectively.

[0156] In the image content of the first image region in the target image, the image content in the first drawing information is greater than the image content in the second drawing information; in the image content of the second image region in the target image, the image content in the second drawing information is greater than the image content in the first drawing information; the distance between the first image region and the first sub-region is less than the distance between the second image region and the first sub-region.

[0157] The image generation module described above is further configured to: input the first drawing information and the second drawing information into the image encoder of the pre-trained image generation model to obtain the first image feature of the first drawing information and the second image feature of the second drawing information; input the first image feature and the second image feature into the fusion layer of the image generation model, and perform feature fusion on the first image feature and the second image feature in the multimodal feature space of the image generation model to generate the first fused feature of the image content of the first target region in the virtual canvas area; wherein, the multimodal feature space stores the correspondence between different types of features for fusing different types of features; input the first fused feature into the image decoder of the image generation model to generate the first target image displayed in the first target region in the virtual canvas area; wherein, the first target image includes at least a portion of the image content of the first drawing information and at least a portion of the image content of the second drawing information.

[0158] The first target region mentioned above is the region between the first sub-region and the second sub-region.

[0159] The aforementioned device further includes a third receiving module, configured to: receive image fusion information sent by any one of the multiple terminal devices; wherein the image fusion information is used to indicate a second target area and specified image content included in the second target area.

[0160] The image generation module described above is further configured to: input the first drawing information, the second drawing information, and the image fusion information into a pre-trained image generation model, so as to perform fusion processing on the first drawing information, the second drawing information, and the image fusion information through the image generation model, and generate a second target image displayed in the second target area of ​​the virtual canvas area; wherein the second target image includes the specified image content indicated by the image fusion information, as well as at least a portion of the image content of the first drawing information and at least a portion of the image content of the second drawing information.

[0161] The aforementioned image fusion information includes: text content drawn in a designated sub-region of the virtual canvas area; the aforementioned image generation module is further configured to: input the first drawing information and the second drawing information into the image encoder of the pre-trained image generation model respectively to obtain the first image feature of the first drawing information and the second image feature of the second drawing information; input the image fusion information into the text encoder of the image generation model to obtain the text feature of the image fusion information; input the first image feature, the second image feature and the text feature into the fusion layer of the image generation model, and perform feature fusion on the first image feature, the second image feature and the text feature in the multimodal feature space of the image generation model to generate the second fusion feature of the image content of the second target area in the virtual canvas area; input the second fusion feature into the image decoder of the image generation model to generate the second target image displayed in the second target area of ​​the virtual canvas area.

[0162] The aforementioned device also includes a mapping module, used to: input text features into a prior sub-model of an image generation model to obtain mapped text features.

[0163] The aforementioned device further includes a text-to-image generation module, configured to: if the first drawing information includes text content for drawing an image in a first sub-region of the virtual canvas area, input the first drawing information into a pre-trained image generation model to generate an image corresponding to the text content displayed in the first sub-region of the virtual canvas area through the image generation model; and display the image corresponding to the text content in the first sub-region.

[0164] The aforementioned device further includes a first sending module, configured to: send an image corresponding to the text content to each terminal device, and display the image corresponding to the text content in a first drawing area of ​​the drawing interface provided by the terminal device; wherein the first drawing area corresponds to a first sub-area.

[0165] The aforementioned device further includes a second sending module, used for: sending the target image to each terminal device, and displaying the target image in a designated drawing area of ​​the drawing interface provided by the terminal device; wherein the designated drawing area corresponds to the target area.

[0166] The image generation apparatus provided in this embodiment of the invention has the same technical features as the image generation method provided in the above embodiments, so it can also solve the same technical problems and achieve the same technical effects.

[0167] Corresponding to the above method embodiments, this embodiment of the invention provides an image generation device that provides a drawing interface through a first terminal device. The first terminal device is communicatively connected to a server, such as... Figure 10 As shown, the device includes:

[0168] The first drawing information sending module 1001 is used to respond to a first drawing operation for a first drawing area in the drawing interface, determine the first drawing information corresponding to the first drawing operation, display the image content corresponding to the first drawing information in the first drawing area of ​​the drawing interface, and send the first drawing information to the server.

[0169] The second drawing information receiving module 1002 is used to receive the image content corresponding to the second drawing information sent by the server, and display the image content corresponding to the second drawing information in the second drawing area of ​​the drawing interface. The image content corresponding to the second drawing information is generated by the server according to the second drawing information sent by the second terminal device.

[0170] The target image receiving module 1003 is used to receive the target image sent by the server and display the target image in the target drawing area of ​​the drawing interface. The target image is the image content generated by the server through the first drawing information, the second drawing information and the pre-trained image generation model. The target drawing area includes at least the blank area between the first drawing area and the second drawing area.

[0171] The image content in the target drawing area is continuous with the image content in the first drawing area and the image content in the second drawing area, respectively.

[0172] In the above method, when multiple people are drawing collaboratively, the multiple drawing information drawn by multiple people is automatically merged to generate a target image displayed in the target area, which improves the efficiency of collaborative drawing; the target area includes at least the blank area between the first sub-region and the second sub-region, thereby enriching the creative effect of the image and meeting the needs of multiple people drawing collaboratively.

[0173] The image generation apparatus provided in this embodiment of the invention has the same technical features as the image generation method provided in the above embodiments, so it can also solve the same technical problems and achieve the same technical effects.

[0174] This embodiment also provides an electronic device, including a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above-described image generation method. This electronic device can be a server or a terminal device.

[0175] See Figure 11 As shown, the electronic device includes a processor 100 and a memory 101. The memory 101 stores machine-executable instructions that can be executed by the processor 100. The processor 100 executes the machine-executable instructions to implement the above-described image generation method.

[0176] Furthermore, Figure 11The electronic device shown also includes a bus 102 and a communication interface 103, with the processor 100, the communication interface 103 and the memory 101 connected via the bus 102.

[0177] The memory 101 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 103 (which can be wired or wireless), such as the Internet, wide area network, local area network, or metropolitan area network. The bus 102 may be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 11 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0178] Processor 100 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 100 or by instructions in software form. Processor 100 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a readily available storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 101, and the processor 100 reads the information from memory 101 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments.

[0179] The processor in the aforementioned electronic device, by executing machine-executable instructions, with the server as the execution entity, can implement the following operations in the aforementioned image generation method:

[0180] A virtual canvas region is generated; wherein the virtual canvas region is used to draw the target image, and in the initial state, the virtual canvas region is a blank area; first drawing information is received from a first terminal device; wherein the first drawing information includes: image content drawn in a first sub-region of the virtual canvas region; second drawing information is received from a second terminal device; wherein the second drawing information includes: image content drawn in a second sub-region of the virtual canvas region; the second sub-region and the first sub-region are different sub-regions in the virtual canvas region; the first drawing information and the second drawing information are input into a pre-trained image generation model, so as to fuse the first drawing information and the second drawing information through the image generation model, and generate a target image displayed in the target area of ​​the virtual canvas region; wherein the target area includes at least the blank area between the first sub-region and the second sub-region.

[0181] The above method also includes displaying the target image in the blank area between the first sub-region and the second sub-region of the target region.

[0182] The above method also includes: determining a target area in the virtual canvas area, wherein the target area includes a portion of the first sub-area, a portion of the second sub-area, and a blank area between the first sub-area and the second sub-area.

[0183] The target image is displayed in a portion of the first sub-region, a portion of the second sub-region, and the blank area between the first and second sub-regions within the target area.

[0184] The image content in the target region is continuous with the image content in the first sub-region and the image content in the second sub-region, respectively.

[0185] In the image content of the first image region in the target image, the image content in the first drawing information is greater than the image content in the second drawing information; in the image content of the second image region in the target image, the image content in the second drawing information is greater than the image content in the first drawing information; the distance between the first image region and the first sub-region is less than the distance between the second image region and the first sub-region.

[0186] The steps described above, which involve inputting the first drawing information and the second drawing information into a pre-trained image generation model to fuse the first drawing information and the second drawing information and generate a target image displayed in the target area of ​​the virtual canvas region, include: inputting the first drawing information and the second drawing information into the image encoder of the pre-trained image generation model to obtain a first image feature of the first drawing information and a second image feature of the second drawing information; inputting the first image feature and the second image feature into the fusion layer of the image generation model, and performing feature fusion on the first image feature and the second image feature in the multimodal feature space of the image generation model to generate a first fused feature of the image content of the first target area in the virtual canvas region; wherein the multimodal feature space stores the correspondence between different types of features for fusing different types of features; and inputting the first fused feature into the image decoder of the image generation model to generate a first target image displayed in the first target area of ​​the virtual canvas region; wherein the first target image includes at least a portion of the image content of the first drawing information and at least a portion of the image content of the second drawing information.

[0187] Before the steps of inputting the first drawing information and the second drawing information into the pre-trained image generation model, the method further includes: receiving image fusion information sent by any of the multiple terminal devices; wherein the image fusion information is used to indicate a second target region and the specified image content included in the second target region.

[0188] The above-described step of inputting the first drawing information and the second drawing information into a pre-trained image generation model to fuse the first drawing information and the second drawing information through the image generation model and generate a target image displayed in the target area of ​​the virtual canvas region includes: inputting the first drawing information, the second drawing information and the image fusion information into a pre-trained image generation model to fuse the first drawing information, the second drawing information and the image fusion information through the image generation model and generate a second target image displayed in the second target area of ​​the virtual canvas region; wherein the second target image includes the specified image content indicated by the image fusion information, as well as at least a portion of the image content of the first drawing information and at least a portion of the image content of the second drawing information.

[0189] The aforementioned image fusion information includes: text content drawn in a designated sub-region of the virtual canvas area; the step of inputting the first drawing information, the second drawing information, and the image fusion information into a pre-trained image generation model, so as to perform fusion processing on the first drawing information, the second drawing information, and the image fusion information through the image generation model, and generating a second target image displayed in the second target region of the virtual canvas area, includes: inputting the first drawing information and the second drawing information into the image encoder of the pre-trained image generation model respectively to obtain the first image feature of the first drawing information and the second image feature of the second drawing information; inputting the image fusion information into the text encoder of the image generation model to obtain the text feature of the image fusion information; inputting the first image feature, the second image feature, and the text feature into the fusion layer of the image generation model, and performing feature fusion on the first image feature, the second image feature, and the text feature in the multimodal feature space of the image generation model to generate the second fusion feature of the image content of the second target region in the virtual canvas area; inputting the second fusion feature into the image decoder of the image generation model to generate the second target image displayed in the second target region of the virtual canvas area.

[0190] After the steps described above, which involve inputting image fusion information into the text encoder of the image generation model to obtain text features of the image fusion information, the method further includes: inputting the text features into the prior sub-model of the image generation model to obtain mapped text features.

[0191] The method further includes: if the first drawing information includes text content for drawing an image in a first sub-region of the virtual canvas area, inputting the first drawing information into a pre-trained image generation model to generate an image corresponding to the text content displayed in the first sub-region of the virtual canvas area through the image generation model; and displaying the image corresponding to the text content in the first sub-region.

[0192] After the step of displaying the image corresponding to the text content in the first sub-region, the method further includes: sending the image corresponding to the text content to each terminal device, and displaying the image corresponding to the text content in the first drawing area of ​​the drawing interface provided by the terminal device; wherein the first drawing area corresponds to the first sub-region.

[0193] After the steps of inputting the first drawing information and the second drawing information into a pre-trained image generation model to fuse the first drawing information and the second drawing information through the image generation model and generate the image content of the target area in the virtual canvas area to obtain the target image, the method further includes: sending the target image to each terminal device and displaying the target image in a designated drawing area of ​​the drawing interface provided by the terminal device; wherein the designated drawing area corresponds to the target area.

[0194] The terminal device is the main execution entity:

[0195] In response to a first drawing operation on a first drawing area in the drawing interface, the system determines the first drawing information corresponding to the first drawing operation, displays the image content corresponding to the first drawing information in the first drawing area of ​​the drawing interface, and sends the first drawing information to the server; it receives the image content corresponding to the second drawing information sent by the server, and displays the image content corresponding to the second drawing information in the second drawing area of ​​the drawing interface, wherein the image content corresponding to the second drawing information is generated by the server based on the second drawing information sent by the second terminal device; it receives a target image sent by the server, and displays the target image in the target drawing area of ​​the drawing interface, wherein the target image is the image content generated by the server using the first drawing information, the second drawing information, and a pre-trained image generation model, wherein the target drawing area includes at least the blank area between the first drawing area and the second drawing area.

[0196] The image content in the target drawing area is continuous with the image content in the first drawing area and the image content in the second drawing area, respectively.

[0197] In the above method, when multiple people are drawing collaboratively, the multiple drawing information drawn by multiple people is automatically merged to generate a target image displayed in the target area, which improves the efficiency of collaborative drawing; the target area includes at least the blank area between the first sub-region and the second sub-region, thereby enriching the creative effect of the image and meeting the needs of multiple people drawing collaboratively.

[0198] In addition, the image generation model can automatically fuse image features and text features with different image content by storing a large number of correspondences between text and image features in the multimodal feature space. This generates fused images that meet the user's personalized needs, improves the fusion effect and efficiency of fused images, enriches the creative effects of fused images, and meets the needs of multi-person collaborative drawing.

[0199] This embodiment also provides a machine-readable storage medium storing machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions cause the processor to implement the above-described image generation method.

[0200] The machine-executable instructions stored in the aforementioned machine-readable storage medium, by executing these machine-executable instructions with the server as the execution entity, can realize the following operations in the aforementioned image generation method:

[0201] A virtual canvas region is generated; wherein the virtual canvas region is used to draw the target image, and in the initial state, the virtual canvas region is a blank area; first drawing information is received from a first terminal device; wherein the first drawing information includes: image content drawn in a first sub-region of the virtual canvas region; second drawing information is received from a second terminal device; wherein the second drawing information includes: image content drawn in a second sub-region of the virtual canvas region; the second sub-region and the first sub-region are different sub-regions in the virtual canvas region; the first drawing information and the second drawing information are input into a pre-trained image generation model, so as to fuse the first drawing information and the second drawing information through the image generation model, and generate a target image displayed in the target area of ​​the virtual canvas region; wherein the target area includes at least the blank area between the first sub-region and the second sub-region.

[0202] The above method also includes displaying the target image in the blank area between the first sub-region and the second sub-region of the target region.

[0203] The above method also includes: determining a target area in the virtual canvas area, wherein the target area includes a portion of the first sub-area, a portion of the second sub-area, and a blank area between the first sub-area and the second sub-area.

[0204] The target image is displayed in a portion of the first sub-region, a portion of the second sub-region, and the blank area between the first and second sub-regions within the target area.

[0205] The image content in the target region is continuous with the image content in the first sub-region and the image content in the second sub-region, respectively.

[0206] In the image content of the first image region in the target image, the image content in the first drawing information is greater than the image content in the second drawing information; in the image content of the second image region in the target image, the image content in the second drawing information is greater than the image content in the first drawing information; the distance between the first image region and the first sub-region is less than the distance between the second image region and the first sub-region.

[0207] The steps described above, which involve inputting the first drawing information and the second drawing information into a pre-trained image generation model to fuse the first drawing information and the second drawing information and generate a target image displayed in the target area of ​​the virtual canvas region, include: inputting the first drawing information and the second drawing information into the image encoder of the pre-trained image generation model to obtain a first image feature of the first drawing information and a second image feature of the second drawing information; inputting the first image feature and the second image feature into the fusion layer of the image generation model, and performing feature fusion on the first image feature and the second image feature in the multimodal feature space of the image generation model to generate a first fused feature of the image content of the first target area in the virtual canvas region; wherein the multimodal feature space stores the correspondence between different types of features for fusing different types of features; and inputting the first fused feature into the image decoder of the image generation model to generate a first target image displayed in the first target area of ​​the virtual canvas region; wherein the first target image includes at least a portion of the image content of the first drawing information and at least a portion of the image content of the second drawing information.

[0208] Before the steps of inputting the first drawing information and the second drawing information into the pre-trained image generation model, the method further includes: receiving image fusion information sent by any of the multiple terminal devices; wherein the image fusion information is used to indicate a second target region and the specified image content included in the second target region.

[0209] The above-described step of inputting the first drawing information and the second drawing information into a pre-trained image generation model to fuse the first drawing information and the second drawing information through the image generation model and generate a target image displayed in the target area of ​​the virtual canvas region includes: inputting the first drawing information, the second drawing information and the image fusion information into a pre-trained image generation model to fuse the first drawing information, the second drawing information and the image fusion information through the image generation model and generate a second target image displayed in the second target area of ​​the virtual canvas region; wherein the second target image includes the specified image content indicated by the image fusion information, as well as at least a portion of the image content of the first drawing information and at least a portion of the image content of the second drawing information.

[0210] The aforementioned image fusion information includes: text content drawn in a designated sub-region of the virtual canvas area; the step of inputting the first drawing information, the second drawing information, and the image fusion information into a pre-trained image generation model, so as to perform fusion processing on the first drawing information, the second drawing information, and the image fusion information through the image generation model, and generating a second target image displayed in the second target region of the virtual canvas area, includes: inputting the first drawing information and the second drawing information into the image encoder of the pre-trained image generation model respectively to obtain the first image feature of the first drawing information and the second image feature of the second drawing information; inputting the image fusion information into the text encoder of the image generation model to obtain the text feature of the image fusion information; inputting the first image feature, the second image feature, and the text feature into the fusion layer of the image generation model, and performing feature fusion on the first image feature, the second image feature, and the text feature in the multimodal feature space of the image generation model to generate the second fusion feature of the image content of the second target region in the virtual canvas area; inputting the second fusion feature into the image decoder of the image generation model to generate the second target image displayed in the second target region of the virtual canvas area.

[0211] After the steps described above, which involve inputting image fusion information into the text encoder of the image generation model to obtain text features of the image fusion information, the method further includes: inputting the text features into the prior sub-model of the image generation model to obtain mapped text features.

[0212] The method further includes: if the first drawing information includes text content for drawing an image in a first sub-region of the virtual canvas area, inputting the first drawing information into a pre-trained image generation model to generate an image corresponding to the text content displayed in the first sub-region of the virtual canvas area through the image generation model; and displaying the image corresponding to the text content in the first sub-region.

[0213] After the step of displaying the image corresponding to the text content in the first sub-region, the method further includes: sending the image corresponding to the text content to each terminal device, and displaying the image corresponding to the text content in the first drawing area of ​​the drawing interface provided by the terminal device; wherein the first drawing area corresponds to the first sub-region.

[0214] After the steps of inputting the first drawing information and the second drawing information into a pre-trained image generation model to fuse the first drawing information and the second drawing information through the image generation model and generate the image content of the target area in the virtual canvas area to obtain the target image, the method further includes: sending the target image to each terminal device and displaying the target image in a designated drawing area of ​​the drawing interface provided by the terminal device; wherein the designated drawing area corresponds to the target area.

[0215] The terminal device is the main execution entity:

[0216] In response to a first drawing operation on a first drawing area in the drawing interface, the system determines the first drawing information corresponding to the first drawing operation, displays the image content corresponding to the first drawing information in the first drawing area of ​​the drawing interface, and sends the first drawing information to the server; it receives the image content corresponding to the second drawing information sent by the server, and displays the image content corresponding to the second drawing information in the second drawing area of ​​the drawing interface, wherein the image content corresponding to the second drawing information is generated by the server based on the second drawing information sent by the second terminal device; it receives a target image sent by the server, and displays the target image in the target drawing area of ​​the drawing interface, wherein the target image is the image content generated by the server using the first drawing information, the second drawing information, and a pre-trained image generation model, wherein the target drawing area includes at least the blank area between the first drawing area and the second drawing area.

[0217] The image content in the target drawing area is continuous with the image content in the first drawing area and the image content in the second drawing area, respectively.

[0218] In the above method, when multiple people are drawing collaboratively, the multiple drawing information drawn by multiple people is automatically merged to generate a target image displayed in the target area, which improves the efficiency of collaborative drawing; the target area includes at least the blank area between the first sub-region and the second sub-region, thereby enriching the creative effect of the image and meeting the needs of multiple people drawing collaboratively.

[0219] In addition, the image generation model can automatically fuse image features and text features with different image content by storing a large number of correspondences between text and image features in the multimodal feature space. This generates fused images that meet the user's personalized needs, improves the fusion effect and efficiency of fused images, enriches the creative effects of fused images, and meets the needs of multi-person collaborative drawing.

[0220] The computer program product of the image generation method, apparatus and system provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the preceding method embodiments. For specific implementation, please refer to the method embodiments, which will not be repeated here.

[0221] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and apparatus described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0222] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.

[0223] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0224] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0225] Finally, it should be noted that the above embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An image generation method characterized by, The method is applied to a server in communication connection with a plurality of terminal devices, the plurality of terminal devices at least comprising a first terminal device and a second terminal device; the method comprises: generating a virtual canvas region; wherein the virtual canvas region is used for drawing a target image, and in an initial state, the virtual canvas region is a blank region; receiving first drawing information sent by the first terminal device; wherein the first drawing information comprises image content drawn in a first sub-region of the virtual canvas region; receiving second drawing information sent by the second terminal device; wherein the second drawing information comprises image content drawn in a second sub-region of the virtual canvas region; the second sub-region and the first sub-region are different sub-regions in the virtual canvas region; inputting the first drawing information and the second drawing information into a pre-trained image generation model to perform fusion processing on the first drawing information and the second drawing information by the image generation model, and generating a target image displayed in a target region in the virtual canvas region; wherein the target region at least comprises a blank region between the first sub-region and the second sub-region. The step of inputting the first drawing information and the second drawing information into a pre-trained image generation model to perform fusion processing on the first drawing information and the second drawing information by the image generation model, and generating a target image displayed in a target region in the virtual canvas region, comprises: inputting the first drawing information and the second drawing information into an image encoder of the pre-trained image generation model respectively to obtain first image features of the first drawing information and second image features of the second drawing information; inputting the first image features and the second image features into a fusion layer of the image generation model, performing feature fusion on the first image features and the second image features in a multi-modal feature space of the image generation model to generate first fusion features of image content of a first target region in the virtual canvas region; wherein the multi-modal feature space stores a corresponding relationship between different types of features for fusing different types of features; inputting the first fusion features into an image decoder of the image generation model to generate a first target image displayed in the first target region in the virtual canvas region; wherein the first target image comprises at least part of image content of the first drawing information and at least part of image content of the second drawing information.

2. The method of claim 1, wherein, The method further comprises: displaying the target image in the blank region between the first sub-region and the second sub-region in the target region.

3. The method of claim 1, wherein, The method further comprises: determining a target region in the virtual canvas region, wherein the target region comprises a part of the first sub-region, a part of the second sub-region, and a blank region between the first sub-region and the second sub-region.

4. The method of claim 3, wherein, The target image is displayed in a partial region of the first sub-region, a partial region of the second sub-region, and a blank region between the first sub-region and the second sub-region in the target region.

5. The method of claim 1, wherein, The image content in the target region is continuous with the image content in the first sub-region and the image content in the second sub-region, respectively.

6. The method of claim 1, wherein, In the image content of the first image region in the target image, the image content of the first drawing information is greater than the image content of the second drawing information; in the image content of the second image region in the target image, the image content of the second drawing information is greater than the image content of the first drawing information. The distance between the first image region and the first sub-region is less than the distance between the second image region and the first sub-region.

7. The method of claim 1, wherein, Before the step of inputting the first drawing information and the second drawing information into the image generation model trained in advance, the method further comprises: Receiving image fusion information sent by any terminal device in the plurality of terminal devices; wherein the image fusion information is used to indicate a second target region, and the second target region includes specified image content.

8. The method of claim 7, wherein, The step of inputting the first drawing information and the second drawing information into the image generation model trained in advance to perform fusion processing on the first drawing information and the second drawing information by the image generation model, and generating a target image displayed in a target region in the virtual canvas region, comprises: The step of inputting the first drawing information, the second drawing information and the image fusion information into the image generation model trained in advance to perform fusion processing on the first drawing information, the second drawing information and the image fusion information by the image generation model, and generating a second target image displayed in a second target region in the virtual canvas region, comprises:

9. The method of claim 8, wherein, The image fusion information includes text content drawn in a specified sub-region of the virtual canvas region. The step of inputting the first drawing information, the second drawing information and the image fusion information into the image generation model trained in advance to perform fusion processing on the first drawing information, the second drawing information and the image fusion information by the image generation model, and generating a second target image displayed in a second target region in the virtual canvas region, comprises: The first drawing information and the second drawing information are respectively input into an image encoder of the image generation model trained in advance to obtain first image features of the first drawing information and second image features of the second drawing information. The image fusion information is input into a text encoder of the image generation model to obtain text features of the image fusion information. The image fusion information is input into a text encoder of the image generation model to obtain the text features of the image fusion information. inputting the first image feature, the second image feature and the text feature into a fusion layer of the image generation model, performing feature fusion on the first image feature, the second image feature and the text feature in a multi-modal feature space of the image generation model, and generating a second fusion feature of image content of a second target region in the virtual canvas region; inputting the second fusion feature into an image decoder of the image generation model to generate the second target image displayed in the second target region in the virtual canvas region.

10. The method of claim 9, wherein, After the step of inputting the image fusion information into a text encoder of the image generation model to obtain a text feature of the image fusion information, the method further comprises: inputting the text feature into a prior sub-model of the image generation model to obtain the mapped text feature.

11. The method of claim 1, wherein, The method further comprises: if the first drawing information includes text content for drawing an image in a first sub-region of the virtual canvas region, inputting the first drawing information into the pre-trained image generation model to generate an image corresponding to the text content displayed in the first sub-region of the virtual canvas region through the image generation model; displaying the image corresponding to the text content in the first sub-region.

12. The method of claim 11, wherein, After the step of displaying the image corresponding to the text content in the first sub-region, the method further comprises: sending the image corresponding to the text content to each of the terminal devices, and displaying the image corresponding to the text content in a first drawing area of a drawing interface provided by the terminal device; wherein the first drawing area corresponds to the first sub-region.

13. The method of claim 1, wherein, After the step of inputting the first drawing information and the second drawing information into the pre-trained image generation model to perform fusion processing on the first drawing information and the second drawing information through the image generation model, and generating a target image displayed in a target region in the virtual canvas region, the method further comprises: sending the target image to each of the terminal devices, and displaying the target image in a specified drawing area of a drawing interface provided by the terminal device; wherein the specified drawing area corresponds to the target region.

14. An image generation method characterized by, providing a drawing interface through a first terminal device, the first terminal device being in communication connection with a server, and the method comprising: in response to a first drawing operation on a first drawing area in the drawing interface, determining first drawing information corresponding to the first drawing operation, displaying image content corresponding to the first drawing information in the first drawing area of the drawing interface, and sending the first drawing information to the server; receiving image content corresponding to second drawing information sent by the server, and displaying the image content corresponding to the second drawing information in a second drawing area of the drawing interface, wherein the image content corresponding to the second drawing information is generated by the server according to second drawing information sent by a second terminal device; receive the target image sent by the server, and display the target image in a target drawing area of the drawing interface, wherein the target image is an image content generated by the server through the first drawing information, the second drawing information, and a pre-trained image generation model, and the target drawing area at least includes a blank area between the first drawing area and the second drawing area; The step of generating the image content by the server through the first drawing information, the second drawing information, and a pre-trained image generation model comprises: The server inputs the first drawing information and the second drawing information into an image encoder of the pre-trained image generation model respectively to obtain first image features of the first drawing information and second image features of the second drawing information; The first image features and the second image features are input into a fusion layer of the image generation model, and the first image features and the second image features are fused in a multi-modal feature space of the image generation model to generate first fusion features of image content of a first target area in a virtual canvas area; wherein the multi-modal feature space stores a corresponding relationship between different types of features for fusing different types of features; and the virtual canvas area is used for drawing a target image, and in an initial state, the virtual canvas area is a blank area. The first fusion features are input into an image decoder of the image generation model to generate a first target image displayed in the first target area in the virtual canvas area; wherein the first target image includes at least part of image content of the first drawing information and at least part of image content of the second drawing information.

15. The method of claim 14, wherein, The image content in the target drawing area is continuous with the image content in the first drawing area and the image content in the second drawing area respectively.

16. An image generation apparatus characterized by comprising: The device is arranged in a server, the server is in communication connection with a plurality of terminal devices, and the plurality of terminal devices at least include a first terminal device and a second terminal device; the device comprises: a canvas generation module configured to generate a virtual canvas area; wherein the virtual canvas area is used for drawing a target image, and in an initial state, the canvas area is a blank area; a first receiving module configured to receive first drawing information sent by the first terminal device; wherein the first drawing information includes image content drawn in a first sub-area of the virtual canvas area; a second receiving module configured to receive second drawing information sent by the second terminal device; wherein the second drawing information includes image content drawn in a second sub-area of the virtual canvas area; and the second sub-area and the first sub-area are different sub-areas in the virtual canvas area. The image generation module is configured to input the first drawing information and the second drawing information into a pre-trained image generation model, perform fusion processing on the first drawing information and the second drawing information by the image generation model, and generate a target image displayed in a target region in the virtual canvas region; wherein the target region at least includes a blank region between the first sub-region and the second sub-region. The image generation module is further configured to input the first drawing information and the second drawing information into an image encoder of a pre-trained image generation model respectively to obtain first image features of the first drawing information and second image features of the second drawing information; input the first image features and the second image features into a fusion layer of the image generation model, perform feature fusion on the first image features and the second image features in a multi-modal feature space of the image generation model to generate first fusion features of image content of a first target region in the virtual canvas region; wherein the multi-modal feature space stores a corresponding relationship between different types of features for fusing different types of features; and input the first fusion features into an image decoder of the image generation model to generate a first target image displayed in the first target region in the virtual canvas region; wherein the first target image includes at least part of image content of the first drawing information and at least part of image content of the second drawing information.

17. An image generation apparatus characterized by comprising: A drawing interface is provided by a first terminal device, the first terminal device is in communication connection with a server, and the device includes: A first drawing information sending module is configured to, in response to a first drawing operation on a first drawing region in the drawing interface, determine first drawing information corresponding to the first drawing operation, display image content corresponding to the first drawing information in the first drawing region of the drawing interface, and send the first drawing information to the server; A second drawing information receiving module is configured to receive image content corresponding to second drawing information sent by the server and display the image content corresponding to the second drawing information in a second drawing region of the drawing interface, wherein the image content corresponding to the second drawing information is generated by the server according to second drawing information sent by a second terminal device; A target image receiving module is configured to receive a target image sent by the server and display the target image in a target drawing region of the drawing interface, wherein the target image is image content generated by the server through the first drawing information, the second drawing information, and a pre-trained image generation model, and the target drawing region at least includes a blank region between the first drawing region and the second drawing region. The server is further configured to: input the first drawing information and the second drawing information into an image encoder of a pre-trained image generation model respectively to obtain first image features of the first drawing information and second image features of the second drawing information; input the first image features and the second image features into a fusion layer of the image generation model, perform feature fusion on the first image features and the second image features in a multi-modal feature space of the image generation model, and generate first fusion features of image content of a first target region in a virtual canvas region; the multi-modal feature space stores a corresponding relationship between different types of features and is used for fusing different types of features; the virtual canvas region is used for drawing a target image, and in an initial state, the virtual canvas region is a blank region; input the first fusion features into an image decoder of the image generation model to generate a first target image displayed in the first target region in the virtual canvas region; the first target image includes at least part of image content of the first drawing information and at least part of image content of the second drawing information.

18. An electronic device, comprising: A device includes a processor and a memory, the memory storing computer executable instructions that are executable by the processor, the processor executing the computer executable instructions to implement the image generation method of any one of claims 1-13, or to implement the image generation method of any one of claims 14-15.

19. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer executable instructions, when the computer executable instructions are invoked and executed by a processor, the computer executable instructions cause the processor to implement the image generation method of any one of claims 1-13, or to implement the image generation method of any one of claims 14-15.

Citation Information

Patent Citations

  • Method and equipment for carrying out multi-person drawing on web browser, and computer program product

    CN108335342A

  • Image generation method and device, computer equipment and storage medium

    CN112116681A