Image generation method and device, program product, storage medium and electronic equipment

By acquiring and encoding a variety of information to generate scene maps, and updating features with the multimodal attention module, the problem of low accuracy in scene map generation in the prior art is solved, and higher generation accuracy and detail restoration is achieved.

CN120279138APending Publication Date: 2025-07-08ZHEJIANG TMALL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510333585.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, when generating scene maps for products, the generation accuracy of scene maps is low.

Method used

By obtaining the image information, text information, graphic information and spatial structure information of the target product, encode these information using multiple encoders, generate a scene map containing the target product, and use the multimodal attention module in the image generation model to update and decode the feature.

Benefits of technology

It improves the accuracy of the generation of scene graphs, can better understand and restore the details and semantics of scene graphs, and realizes accurate scene graph generation based on multiple control conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279138A_ABST
    Figure CN120279138A_ABST
Patent Text Reader

Abstract

The invention discloses an image generation method and device, a program product, a storage medium and electronic equipment. Relates to the field of artificial intelligence, and the method comprises the steps: obtaining target information which comprises image information of a target commodity and at least one of the following information: text information used for describing a to-be-generated scene, graphic and text information of an article matched with the target commodity, and space structure information of the to-be-generated scene; and encoding the target information through a plurality of encoders in the image generation model to obtain a target feature, and generating a scene graph including the target commodity based on the target feature, the plurality of encoders being used for processing different information in the target information. The problem of low generation accuracy of the scene graph when the scene graph containing the commodity image is generated for the commodity in the related technology is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular, to a method, apparatus, program product, storage medium, and electronic device for generating images. Background Art

[0002] In the fields of e-commerce and online retail, generating high-quality scene graphs for products has become one of the key technologies to enhance the user experience and increase the attractiveness of products. Scene graphs can be applied to different channels such as main images, product details, and content. They can place the product in an actual usage environment to help consumers better understand the size, function, and aesthetic value of the product. Currently, in related technologies, scene graphs corresponding to products are usually generated in a text-to-image manner based on single information (such as product information), resulting in the problem of low accuracy in generating scene graphs.

[0003] For the above problems, no effective solutions have been proposed yet. Summary of the Invention

[0004] Embodiments of the present application provide a method, apparatus, program product, storage medium, and electronic device for generating images, so as to at least solve the technical problem of low accuracy in generating scene graphs containing product images for products in related technologies.

[0005] According to one aspect of the embodiments of the present application, a method for generating an image is provided, including: obtaining target information, where the target information includes image information of a target product, and at least one of the following information: text information for describing a scene to be generated, graphic and text information of an item paired with the target product, and spatial structure information of the scene to be generated; encoding the target information through a plurality of encoders in an image generation model to obtain target features, and generating a scene graph containing the target product based on the target features, where the plurality of encoders are used to process different information in the target information.

[0006] Further, when the target information includes image information, text information, graphic and text information, and spatial structure information, encoding the target information through a plurality of encoders in an image generation model to obtain target features includes: encoding the text information and spatial structure information through a first encoder to obtain first features; encoding the image information through a second encoder to obtain second features; encoding the graphic and text information through a third encoder to obtain third features; and determining the target features according to the first features, second features, and third features.

[0007] Further, the image information includes a first image and a second image. The first image refers to a scene canvas image containing the target commodity, and the second image refers to a mask image of the target commodity. Encoding the image information through a second encoder to obtain a second feature includes: encoding the first image through the second encoder to obtain a first sub-feature; performing block encoding on the second image to obtain a second sub-feature; splicing the first sub-feature and the second sub-feature to obtain the second feature.

[0008] Further, determining a target feature according to the first feature, the second feature, and the third feature includes: determining a first position encoding corresponding to the first feature, a second position encoding corresponding to the second feature, and a third position encoding corresponding to the third feature through a position encoding module in the image generation model; determining a target first feature according to the first feature and the first position encoding, a target second feature according to the second feature and the second position encoding, and a target third feature according to the third feature and the third position encoding through the position encoding module; determining the target feature according to the target first feature, the target second feature, and the target third feature.

[0009] Further, the image information includes a first image, and the first image refers to a scene canvas image containing the target commodity. The graphic information includes the position information of the items paired with the target commodity in the scene canvas image. Determining a first position encoding corresponding to the first feature, a second position encoding corresponding to the second feature, and a third position encoding corresponding to the third feature through a position encoding module in the image generation model includes: encoding the preset position information to obtain the first position encoding; determining the target position of the target commodity in the scene canvas image according to the first image, and encoding the target position to obtain the second position encoding; determining the central position of the item in the scene canvas image according to the position information of the item in the scene canvas image, and encoding the central position to obtain the third position encoding.

[0010] Further, generating a scene graph containing the target commodity based on the target feature includes: updating the initial image feature according to the target feature and the initial image feature through a multi-modal attention module in the image generation model to obtain an updated image feature; decoding the updated image feature to obtain the scene graph.

[0011] Further, the multi-modal attention module in the image generation model updates the initial image features based on the target features and the initial image features. The updated image features obtained include: normalizing the input features of the multi-modal attention module through a normalization layer to obtain a fourth feature, where the input features are determined based on the target features and the initial image features; processing the fourth feature through a linear layer to obtain a fifth feature, and processing the fourth feature through a multi-modal attention layer to obtain a sixth feature; processing the fifth feature and the sixth feature through a projection layer to obtain a seventh feature; performing a residual connection on the seventh feature and the input features to obtain the output features of the multi-modal attention module, and determining the updated image features based on the output features.

[0012] Further, the image generation model is obtained in the following manner: obtaining a training sample set, where the training samples in the training sample set include sample information, and the sample information includes image information of the sample commodity and at least one of the following information: text information for describing the sample scene to be generated, graphic and text information of the items paired with the sample commodity, and spatial structure information of the sample scene to be generated. The true label of the training sample is a sample scene graph including the sample commodity; training an initial image generation model based on the training sample set to obtain the image generation model.

[0013] According to another aspect of the embodiments of the present application, there is also provided a method for generating an image, including: obtaining target information uploaded by a client, where the target information includes image information of a target commodity and at least one of the following information: text information for describing the scene to be generated, graphic and text information of the items paired with the target commodity, and spatial structure information of the scene to be generated; encoding the target information through multiple encoders in the image generation model in a cloud server to obtain target features, and generating a scene graph including the target commodity based on the target features, where the multiple encoders are used to process different information in the target information; and feeding back the scene graph to the client.

[0014] According to another aspect of the embodiments of the present application, there is also provided an image generation device, including: a first acquisition unit for acquiring target information, where the target information includes image information of a target commodity and at least one of the following information: text information for describing the scene to be generated, graphic and text information of the items paired with the target commodity, and spatial structure information of the scene to be generated; and a generation unit for encoding the target information through multiple encoders in the image generation model to obtain target features, and generating a scene graph including the target commodity based on the target features, where the multiple encoders are used to process different information in the target information.

[0015] Further, the generation unit includes: a first processing subunit, configured to encode and process text information and spatial structure information through a first encoder to obtain a first feature; a second processing subunit, configured to encode and process image information through a second encoder to obtain a second feature; a third processing subunit, configured to encode and process text-image information through a third encoder to obtain a third feature; and a determination subunit, configured to determine a target feature according to the first feature, the second feature, and the third feature.

[0016] Further, the image information includes a first image and a second image. The first image refers to a scene canvas image including a target commodity, and the second image refers to a mask image of the target commodity. The second processing subunit includes: a first processing module, configured to encode and process the first image through a second encoder to obtain a first sub-feature; a second processing module, configured to perform block encoding processing on the second image to obtain a second sub-feature; and a third processing module, configured to splice the first sub-feature and the second sub-feature to obtain a second feature.

[0017] Further, the determination subunit includes: a first determination module, configured to determine a first position encoding corresponding to the first feature, a second position encoding corresponding to the second feature, and a third position encoding corresponding to the third feature through a position encoding module in the image generation model; a second determination module, configured to determine a target first feature according to the first feature and the first position encoding, determine a target second feature according to the second feature and the second position encoding, and determine a target third feature according to the third feature and the third position encoding through the position encoding module; and a third determination module, configured to determine a target feature according to the target first feature, the target second feature, and the target third feature.

[0018] Further, the image information includes a first image, which refers to a scene canvas image including a target commodity, and the text-image information includes the position information of an item paired with the target commodity in the scene canvas image. The first determination module includes: a first processing sub-module, configured to encode and process preset position information to obtain a first position encoding; a second processing sub-module, configured to determine a target position of the target commodity in the scene canvas image according to the first image, and encode and process the target position to obtain a second position encoding; and a third processing sub-module, configured to determine a central position of the item in the scene canvas image according to the position information of the item in the scene canvas image, and encode and process the central position to obtain a third position encoding.

[0019] Further, the generation unit includes: a fourth processing subunit, configured to update and process an initial image feature according to the target feature and the initial image feature through a multi-modal attention module in the image generation model to obtain an updated image feature; and a fifth processing subunit, configured to perform decoding processing on the updated image feature to obtain a scene graph.

[0020] Further, the fourth processing sub-unit includes: a fourth processing module, configured to normalize the input features of the multi-modal attention module through a normalization layer to obtain fourth features, where the input features are determined based on the target features and the initial image features; a fifth processing module, configured to process the fourth features through a linear layer to obtain fifth features, and process the fourth features through a multi-modal attention layer to obtain sixth features; a sixth processing module, configured to process the fifth features and the sixth features through a projection layer to obtain seventh features; a fourth determination module, configured to perform a residual connection between the seventh features and the input features to obtain the output features of the multi-modal attention module, and determine the updated image features according to the output features.

[0021] Further, the image generation device includes: a second acquisition unit, configured to acquire a training sample set, where the training samples in the training sample set include sample information, and the sample information includes image information of the sample commodity, and at least one of the following information: text information for describing the sample scene to be generated, graphic and text information of the items paired with the sample commodity, and spatial structure information of the sample scene to be generated, and the true label of the training sample is a sample scene graph including the sample commodity; a training unit, configured to train an initial image generation model based on the training sample set to obtain an image generation model.

[0022] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including: a memory storing an executable program; a processor configured to run the program, where when the program runs, it executes the image generation method of any one of the above.

[0023] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium storing a program, where when the program runs, it controls the device where the storage medium is located to execute the image generation method of any one of the above.

[0024] According to another aspect of the embodiments of the present invention, there is also provided a computer program product including a computer program, where when the computer program is executed by a processor, it implements the image generation method of any one of the above.

[0025] In the embodiments of the present application, by obtaining target information, where the target information includes image information of a target commodity and at least one of the following information: text information for describing a to-be-generated scene, graphic and text information of an item paired with the target commodity, and spatial structure information of the to-be-generated scene; encoding the target information through multiple encoders in an image generation model to obtain target features, and generating a scene graph including the target commodity based on the target features. The multiple encoders are used to process different information in the target information. By obtaining the target information and processing the target information through the image generation model, key information related to the to-be-generated scene graph is injected into the image generation model as control conditions, so that the generation process of the scene graph can be accurately controlled. By designing corresponding encoders for different information in the target information, the image generation model can extract more accurate features specifically, so as to better understand and restore the details and semantics of the scene graph, and improve the generation accuracy of the scene graph. The purpose of generating a scene graph based on multiple control conditions is achieved, thereby achieving the technical effect of improving the generation accuracy of the scene graph, and further solving the technical problem of low generation accuracy of the scene graph when generating a scene graph including a commodity image for a commodity in the related art. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0027] Figure 1 is a schematic diagram of a computer terminal provided in Embodiment 1 of the present application;

[0028] Figure 2 is a flowchart of a method for generating an image provided in Embodiment 1 of the present application;

[0029] Figure 3 is a schematic diagram of target information and a scene graph provided in Embodiment 1 of the present application;

[0030] Figure 4 is a schematic diagram of the working process of an image generation model provided in Embodiment 1 of the present application;

[0031] Figure 5 is a schematic diagram of a multi-modal attention module provided in Embodiment 1 of the present application;

[0032] Figure 6 is a schematic diagram of a scene graph provided according to the related art Figure 1 ;

[0033] Figure 7 is a schematic diagram of a scene graph provided in Embodiment 1 of the present applicationFigure 1 ;

[0034] Figure 8 is a schematic diagram of a scene graph provided according to the related art Figure 2 ;

[0035] Figure 9 is a schematic diagram of a scene graph provided according to Embodiment 1 of the present application Figure 2 ;

[0036] Figure 10 is a schematic diagram of a scene graph provided according to the related art Figure 3 ;

[0037] Figure 11 is a schematic diagram of a scene graph provided according to Embodiment 1 of the present application Figure 3 ;

[0038] Figure 12 is a flowchart of a method for generating an image provided according to Embodiment 2 of the present application;

[0039] Figure 13 is a schematic diagram of an image generation device provided according to Embodiment 3 of the present application;

[0040] Figure 14 is a block diagram of the structure of an electronic device provided according to Embodiment 4 of the present application. Detailed implementation manners

[0041] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0042] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0043] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in the relevant regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0044] Embodiment 1

[0045] According to an embodiment of the present application, there is also provided a method for generating an image. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0046] The method embodiment provided by the first embodiment of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal (or mobile device) for implementing the method for generating an image is shown. As Figure 1 shown, the computer terminal (or mobile device) 10 may include a set of processors 102 (the set of processors 102 may include, but is not limited to, processing devices such as a microcontroller unit (MCU) or a field programmable gate array (FPGA), and the set of processors 102 may include a set of processors, Figure 1 which are shown as 102a, 102b,..., 102n in the figure), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.

[0047] It should be noted that one or more of the above-mentioned processors 102 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computer terminal 10 (or mobile device).

[0048] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the image generation method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned image generation method. The memory 104 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 can further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.

[0049] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network can include the wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0050] The display can be a touch-screen liquid crystal display, which enables the user to interact with the user interface of the computer terminal 10 (or mobile device).

[0051] In the fields of e-commerce and online retail, generating high-quality scene graphs for products has become one of the key technologies to enhance the user experience and increase the attractiveness of products. Scene graphs can be applied to different channels such as main pictures, product details, and content. They can place products in the actual usage environment to help consumers better understand the size, function, and aesthetic value of products. Currently, in related technologies, scene graphs corresponding to products are usually generated in a text-to-image manner based on a single piece of information (such as product information), resulting in the problem of low accuracy in generating scene graphs.

[0052] Under the above technical background, the present application provides a method for generating an image as shown in Figure 2 . Figure 2 It is a flowchart of a method for generating an image according to Embodiment 1 of the present application. The method includes:

[0053] Step S201, obtain target information, where the target information includes image information of a target commodity, and at least one of the following information: text information for describing a scene to be generated, graphic and text information of an item to be paired with the target commodity, and spatial structure information of the scene to be generated.

[0054] Optionally, devices such as an electronic device, an application system, and a server can be used as the execution subject of the present application. In this embodiment, the target processing system is used as the execution subject to execute the above method for generating an image.

[0055] Optionally, when a user expects to generate a scene graph corresponding to a certain commodity (i.e., the target commodity), the user can input the target information corresponding to the commodity through the front-end human-computer interaction interface for the target processing system to obtain. Generally speaking, a scene graph corresponding to a certain commodity can be understood as a scene graph with the commodity as the main commodity (i.e., the core element), and this scene graph serves the commodity and is used for scene-based display of the commodity. Specifically, the scene graph corresponding to a certain commodity includes the commodity and the relevant background for displaying the commodity. The aforementioned relevant background for displaying the commodity is the scene to be generated, and this scene can also be understood as the application scene of the target commodity. For example, if the target commodity is a sofa, the scene to be generated can be a room, which may also include common furnishings in the room (such as a carpet, a coffee table, a bookshelf, and a vase, etc.).

[0056] The target commodity can be a commodity supported for sale on an e-commerce platform. For example, sofas, TVs, desks, beds, etc., which vary according to different actual application requirements. Therefore, in this embodiment, the commodity type of the target commodity is not specifically limited.

[0057] Optionally, the target information includes image information of the target commodity, and at least one of the following information: text information for describing the scene to be generated, graphic and text information of the item to be paired with the target commodity, and spatial structure information of the scene to be generated.

[0058] In an alternative embodiment, the target information includes at least one of the image information of the target product, the text information (or graphic and text information, or spatial structure information) for describing the scene to be generated, and the remaining information. For example, when the target information includes the aforementioned image information and text information, the remaining information is the graphic and text information and the spatial structure information. Another example is that when the target information includes the aforementioned image information and graphic and text information, the remaining information is the text information and the spatial structure information. Still another example is that when the target information includes the aforementioned image information and spatial structure information, the remaining information is the text information and the graphic and text information.

[0059] In an alternative embodiment, the target information includes the image information of the target product, the text information for describing the scene to be generated, the graphic and text information of the items paired with the target product, and the spatial structure information of the scene to be generated.

[0060] Optionally, the image information of the target product includes the image content related to the target product. For example, the image information of the target product includes the product image of the target product. Another example is that the image information includes a first image and a second image. The first image refers to the scene canvas image containing the target product, and the second image refers to the mask image of the target product. For example, Figure 3 is a schematic diagram of the target information and the scene graph provided in the first embodiment of the present application. Assuming that the target product is a black sofa, the first image corresponding to the target product is as Figure 3 shown in the upper image in the image information. In this image, the target product is placed in the center of the scene canvas. The second image corresponding to the target product is as Figure 3 shown in the lower image in the image information.

[0061] Optionally, the text information for describing the scene to be generated is an important reference information for generating the scene graph. The text information can provide high-level semantic content such as the overall style, color preference, and functional area of the scene. For example, an alternative text information is as follows:

[0062] "This picture shows a warm indoor scene, in which there is a unique, round black leather chair. It looks soft and attractive. The chair is placed in front of a large window, and the natural light coming in through the window creates a warm and inviting atmosphere.

[0063] There is an open book lying flat on the chair, and the pages are slightly unfolded, suggesting that it has just been read recently. There is a modern glass cabinet in the background, with a golden trim on the edge, displaying various ornaments such as small sculptures and possible collections, adding a touch of elegance to the room.

[0064] A fashionable area rug is laid on the floor, with a striking black-and-white pattern that forms a sharp contrast with the chair. There is also a small black-and-white cushion with an abstract design on the rug, further enhancing the visual interest of the space.

[0065] Through the window, some traces of the urban environment can be seen, showing brick buildings and reflective glass, emphasizing the modernity of the interior space. Overall, this scene conveys a feeling of combining comfort and modern style, which is very suitable for relaxation or reading.

[0066] Optionally, the items paired with the target product can be understood as the items used to assist in pairing the target product in the scene. For example, such items can be hanging paintings, ornaments, throw pillows, etc. They are equivalent to the secondary elements in the scene. Although these items are not the core of the scene, they play an important role in creating the overall atmosphere. The graphic information of the items can include the item image and the position information of the item in the scene canvas image. For example, an optional piece of graphic information is as follows:

[0067] "pillow":

[0068] {

[0069] "box":[63.982547760009766,

[0070] 687.0017700195312,

[0071] 441.14111328125, 798.326904296875

[0073] ,

[0074] "type":"full_preserve",

[0075] "object_url":

Image content

[0076] }

[0077] Among them, the content after "box" represents the position information of the corner points of the outer rectangular frame of the item image in the scene canvas image. For example, the first and second numbers in "box" represent the X-axis and Y-axis coordinates of the upper left corner of the rectangular frame, and the third and fourth numbers in "box" represent the X-axis and Y-axis coordinates of the lower right corner of the rectangular frame. This position information is equivalent to the expected position information. The

Image content

[0078] Optionally, the spatial structure information of the scene includes the relative position information of structural elements such as the ground and walls in the scene to help the model generate a reasonable spatial layout. Generally speaking, the spatial structure information can also be understood as the housing type structure information. The spatial structure information can be used as the input data of the model in a parametric form. For example, an optional spatial structure information is shown as follows:

[0079] floor_wall_pts:[(0, 0.75), (0.4, 0.55), (1, 0.675)];

[0080] Among them, the three coordinates after "floor_wall_pts" represent the intersection coordinates of the floor and the wall in the scene graph. Through the parametric description of the spatial structure, the spatial layout information of the scene can be provided for the image generation model. This information is input into the model in a parametric form, making the generated scene graph more reasonable and accurate in spatial structure and achieving accurate control of the wall and floor areas in the scene.

[0081] Step S202, encode the target information through multiple encoders in the image generation model to obtain target features, and generate a scene graph containing the target commodity based on the target features, where the multiple encoders are used to process different information in the target information.

[0082] Optionally, the target processing system can input the target information into the image generation model. The image generation model includes multiple encoders, and the multiple encoders are used to process different information in the target information. For example, the text encoder is used to process the text information describing the scene to be generated and the spatial structure information of the scene to be generated, one image encoder is used to process the image information of the target commodity, and another image encoder is used to process the image content in the graphic and text information of the item matched with the target commodity.

[0083] After splicing the results output by the multiple encoders, target features are obtained. The target generation model can generate a scene graph containing the target commodity according to the target features. For example, the target commodity scene graph is obtained by processing the target features through the multi-modal attention module.

[0084] Optionally, the scene graph includes the target product and the relevant background for displaying the target product. The image-based generation method exemplifies the image content of the scene graph according to the application scenario it aims at. For example, if it is for indoor products, the generated scene graph can be an indoor space, and the target products in the scene graph can be indoor furniture, lighting fixtures, and decorations, etc., to display the indoor decoration effects of different styles and design concepts. If it is for automotive products, the generated scene graph can be a road or a parking lot, and the target products in the scene graph can be cars of different brands and models, to display the appearance, interior, and performance of the cars. If it is for clothing matching, the generated scene graph can be a street or a shopping mall, and the target products in the scene graph can be clothes of different styles and colors, to display the matching effects and styles of the clothes. If it is for catering, the generated scene graph can be a restaurant or a coffee shop, and the target products in the scene graph can be different kinds of food and beverages, to display the appearance and taste of different foods. The specific content of the scene graph can be generated personalized according to actual needs, so no specific limitation is made here.

[0085] For example, assume that the target information input into the image generation model includes the image information, text information, graphic and text information, and spatial structure information as shown in Figure 3 , then the scene graph output by the image generation model is as shown in the rightmost image in Figure 3 .

[0086] In this solution, by obtaining the target information and processing the target information through the image generation model, a variety of key information related to the to-be-generated scene graph is injected into the image generation model as control conditions, so as to accurately control the generation process of the scene graph. By designing corresponding encoders for different information in the target information, the image generation model can extract more accurate features specifically, so as to better understand and restore the details and semantics of the scene graph, and improve the generation accuracy of the scene graph. The purpose of generating the scene graph based on multiple control conditions is achieved, thus achieving the technical effect of improving the generation accuracy of the scene graph, and further solving the technical problem of low generation accuracy of the scene graph when generating a scene graph containing a product image for a product in the related art.

[0087] How to obtain the target feature is crucial. Therefore, in the image generation method provided in the first embodiment of this application, when the target information includes image information, text information, graphic-text information, and spatial structure information, the target information is encoded by multiple encoders in the image generation model to obtain the target feature, including: encoding the text information and spatial structure information through the first encoder to obtain the first feature; encoding the image information through the second encoder to obtain the second feature; encoding the graphic-text information through the third encoder to obtain the third feature; and determining the target feature according to the first feature, the second feature, and the third feature.

[0088] Optionally, the first encoder is a text encoder for processing text-type information. In some embodiments, the first encoder is a single text encoder, and in other embodiments, the first encoder is composed of multiple text encoders.

[0089] For example, in order to make full use of the text information, the first text encoder and the second text encoder are combined as the first encoder, and the first text encoder is different from the second text encoder. For example, the first text encoder can adopt the encoder structure in a pre-trained language model. This pre-trained language model adopts an encoder-decoder structure and is mainly used to process various natural language processing tasks. The encoder in this model contains multiple encoder layers, and the encoder layer contains a self-attention mechanism and a feed-forward network, which are used to convert the input text into a fixed-size context vector. That is, the first text encoder focuses on extracting fine-grained information in the text, such as descriptive words and phrases. The second text encoder can adopt the encoder structure in a multi-modal pre-trained model. This multi-modal pre-trained model is used to encode text and images into the same vector space to achieve alignment and retrieval between images and texts. The encoder in this model processes the text sequence through a causal attention mechanism and uses the feature vector of the last token as the representation of the text. That is, the second text encoder focuses on extracting high-level semantic information, such as the overall style and theme of the scene. By combining the two, the text information can be extracted both finely and at a high level, so as to obtain a more comprehensive text feature representation, provide richer semantic guidance for the image generation model, and enable the generated scene graph to respond more accurately to the text description.

[0090] In the process of encoding text information and spatial structure information through the first encoder to obtain the first feature, the text information can be input into the first text encoder and the second text encoder respectively to obtain the first text feature and the second text feature corresponding to the text information, and then the first text feature and the second text feature are concatenated to obtain the text feature corresponding to the text information. The spatial structure information is input into the first text encoder and the second text encoder respectively to obtain the first spatial structure feature and the second spatial structure feature corresponding to the spatial structure information, and then the first spatial structure feature and the second spatial structure feature are concatenated to obtain the spatial structure feature corresponding to the spatial structure information. Then, the aforementioned text feature and spatial structure feature are concatenated to obtain the first feature.

[0091] Optionally, the second encoder is an image encoder for processing image-type information. For example, the second encoder can adopt the encoder structure in a generative model that combines an autoencoder and probability modeling. This generative model maps the input data to a probability distribution (usually a Gaussian distribution) in the latent space through the encoder and samples from it to generate new data. The loss function of this generative model consists of the reconstruction error and the KL (Kullback-Leibler Divergence) divergence, which is used to constrain the distribution of the latent space. This design enables this generative model not only to compress and reconstruct data but also to generate new samples similar to the training data. In scene graph generation, the image information of the target commodity can be encoded through the encoder in this generative model to retain its key features. The advantage of this encoder is that it can compress the image into a low-dimensional latent space while retaining the image details.

[0092] Optionally, the third encoder is an image encoder for processing image-type information. For example, in order to implement the perspective adaptation function of the accessory in the generation process, the encoder structure in a multi-modal pre-training model is used as the encoder structure of the third encoder to encode the image content in the text-image information. This model is used for the alignment task of images and texts, aiming to improve the alignment ability between images and texts through improved training methods and architectures. This model includes a text encoder and a visual encoder, that is, the third encoder adopts the encoder structure of the visual encoder in this model. Compared with the second encoder, the difference of the third encoder is that the features it extracts are more abstract high-level features, which are convenient for perspective adaptation generation.

[0093] The graphic information of an item may include the item image of the item and the position information of the item in the scene canvas image. In the process of encoding the graphic information through a third encoder to obtain a third feature, the image generation model may process the image content (i.e., the item image) in the graphic information of the item through the third encoder to obtain the item feature corresponding to the item, and thus use the item feature as the third feature. The aforementioned third encoder can extract the detailed semantic information of the accessory items (i.e., paired items), and at the same time allows the similarity requirements for these items to be appropriately reduced during the generation process. This enables the generated accessory items to be flexibly adjusted according to the perspective and layout of the scene while maintaining semantic consistency.

[0094] Optionally, there may be multiple items paired with the target product. In the case where there are multiple items paired with the target product, the item features of the multiple items are obtained separately through the third encoder, and the item features of the multiple items are concatenated to obtain the third feature.

[0095] In an alternative embodiment, after obtaining the first feature, the second feature, and the third feature, the first feature, the second feature, and the third feature may be concatenated to obtain the target feature.

[0096] In an alternative embodiment, after obtaining the first feature, the second feature, and the third feature, corresponding position encodings may be added to each feature, and then the first feature, the second feature, and the third feature after adding the position encodings are concatenated to obtain the target feature.

[0097] It should be noted that through the above method, the injection of pixel alignment information, text information, and semantic information in the image generation model is realized. By combining the main product information, background accessory information, scene house type structure information, and scene description information as conditional data, the accurate control of the scene graph in all aspects is realized. The model can more accurately restore the details of the scene graph, including the shape of the main product, the similarity of background objects, and the rationality of the overall environment, so that the generated scene graph achieves better effects in terms of details, semantic consistency, and overall layout. In addition, the present application can realize the implantation of multiple objects in the same model. The model can flexibly achieve the complete conformal shape of the main object and the adaptive perspective effect of the accessory items, and the present application also has strong scalability and can expand more control conditions according to actual application requirements.

[0098] To generate a more accurate second feature, in the image generation method provided in the first embodiment of the present application, the image information includes a first image and a second image. The first image refers to a scene canvas image containing the target commodity, and the second image refers to a mask image of the target commodity. Encoding the image information through a second encoder to obtain the second feature includes: encoding the first image through the second encoder to obtain a first sub-feature; performing block encoding on the second image to obtain a second sub-feature; and splicing the first sub-feature and the second sub-feature to obtain the second feature.

[0099] Optionally, the second encoder is an image encoder for processing information of the image type. The size of the scene canvas image is the same as that of the to-be-generated scene graph, and the position where the target commodity is located in the scene canvas image is the expected position of the target commodity in the to-be-generated scene graph. The scene canvas image can be understood as a preliminary and basic image framework of the to-be-generated scene. Generally speaking, this image is a "blueprint" or "sketch", which is used to guide the image generation model to understand the basic relationship between the commodity and the to-be-generated scene.

[0100] For example, Figure 3 is a schematic diagram of the target information and the scene graph provided in the first embodiment of the present application. Assuming the target commodity is a black sofa, the first image corresponding to the target commodity is as shown in the upper image in the Figure 3 image information. In this image, the target commodity is placed in the central position of the scene canvas. The second image corresponding to the target commodity is as shown in the lower image in the Figure 3 image information. The image generation model can encode the first image through the second encoder to obtain a first sub-feature, and the image generation model can perform block encoding (i.e., patch encoding) on the second image to obtain a second sub-feature. Among them, block encoding means dividing the second image into several image patches (patches), and then encoding the image patches, so that the second sub-feature is constituted by the features of the image patches. Among them, the resolution of the mask image is consistent with the resolution of the commodity image in the scene canvas image to reduce the blurriness of the commodity edge in the subsequently generated scene graph.

[0101] Optionally, after obtaining the first sub-feature and the second sub-feature, the first sub-feature and the second sub-feature are spliced to obtain the second feature.

[0102] It should be noted that through the above method, the model can effectively capture and retain the key information of the target commodity in the to-be-generated scene graph, including the appearance details of the commodity and its positional relationship in the scene. By introducing the mask image of the target commodity, this method can overcome the problem of blurred commodity edges in the traditional solution, thereby improving the accuracy of the generated scene graph.

[0103] In order to generate more accurate target features, in the image generation method provided in Embodiment 1 of this application, determining the target features according to the first feature, the second feature, and the third feature includes: determining the first position encoding corresponding to the first feature, the second position encoding corresponding to the second feature, and the third position encoding corresponding to the third feature through the position encoding module in the image generation model; determining the target first feature according to the first feature and the first position encoding, the target second feature according to the second feature and the second position encoding, and the target third feature according to the third feature and the third position encoding through the position encoding module; determining the target features according to the target first feature, the target second feature, and the target third feature.

[0104] Optionally, since text information and spatial structure information do not have specific position attributes, the position encoding module can determine the first position encoding based on a preset information (such as 0) to represent the global semantic features of the first feature. For the second feature, the position encoding module can calculate the second position encoding according to the position information of the target commodity in the scene canvas image to retain the spatial information. For the third feature, the third position encoding can be calculated according to the position information of the item in the scene canvas image to achieve perspective adaptability during the generation process.

[0105] Optionally, after obtaining the first position encoding, the second position encoding, and the third position encoding, the image generation model can add the first feature and the first position encoding to obtain the target first feature, add the second feature and the second position encoding to obtain the target second feature, and add the third feature and the third position encoding to obtain the target third feature.

[0106] After obtaining the target first feature, the target second feature, and the target third feature, the target first feature, the target second feature, and the target third feature can be concatenated to obtain the target features.

[0107] It should be noted that by introducing the position encodings corresponding to the first feature, the second feature, and the third feature respectively during the process of generating the target features, the richness of the information in the target features is further improved, thereby effectively improving the accuracy of the generated scene graph.

[0108] In order to generate more accurate positional encodings, in the image generation method provided in the first embodiment of the present application, the image information includes a first image, where the first image refers to a scene canvas image containing the target commodity, and the graphic and text information includes the position information of the item paired with the target commodity in the scene canvas image. Determining the first positional encoding corresponding to the first feature, the second positional encoding corresponding to the second feature, and the third positional encoding corresponding to the third feature through the positional encoding module in the image generation model includes: encoding the preset position information to obtain the first positional encoding; determining the target position of the target commodity in the scene canvas image according to the first image, and encoding the target position to obtain the second positional encoding; determining the central position of the item in the scene canvas image according to the position information of the item in the scene canvas image, and encoding the central position to obtain the third positional encoding.

[0109] Optionally, the input of the positional encoding module is the input of multiple encoders (i.e., the target information) and the outputs of multiple encoders (i.e., the first feature, the second feature, and the third feature).

[0110] Optionally, since the text information and the spatial structure information do not have specific position attributes and have a global influence on the generation of the entire scene graph, the preset position information can be set to 0, and the positional encoding module encodes the preset position information to obtain the first positional encoding.

[0111] In an alternative embodiment, the positional encoding module can identify the target position of the target commodity in the scene canvas image from the first image, and thus encode the target position to obtain the second positional encoding. Optionally, the target position can be the central position of the target commodity in the scene canvas image, that is, the coordinates of the center of the target commodity in the scene canvas image. The target position can also be the position of the corner points of the target rectangular frame in the scene canvas image, where the target rectangular frame refers to the rectangular frame surrounding the target commodity in the scene canvas image.

[0112] In an alternative embodiment, the positional encoding module can also identify the target position of the target commodity in the scene canvas image from the first image based on the second image, and thus encode the target position to obtain the second positional encoding.

[0113] Optionally, the graphic information includes the position information of the item paired with the target commodity in the scene canvas image. This position information refers to the position information of the corner points of the outer rectangular frame of the item image in the scene canvas image. Therefore, the position encoding module can calculate the center position of the item in the scene canvas image based on this position information, that is, calculate the coordinates of the center of the item in the scene canvas image. For example, by averaging the X-axis coordinates in the position information of the item and averaging the Y-axis coordinates, the center position of the item can be obtained. After obtaining the center position, the position encoding module can perform encoding processing on the center position to obtain the third position encoding.

[0114] It should be noted that through the above method, the position information of the commodity and the position information of the item are effectively integrated into the target feature, providing the model with the ability to deeply understand the scene layout and element positions, enabling the image generation model to better adjust the layout and perspective of the scene elements, improving the accuracy of the generated scene graph, and achieving a generation effect that is both in line with the description and natural and reasonable.

[0115] To generate more accurate position encoding, in the image generation method provided in Embodiment 1 of this application, generating a scene graph including the target commodity based on the target feature includes: updating the initial image feature by the multi-modal attention module in the image generation model according to the target feature and the initial image feature to obtain the updated image feature; and performing decoding processing on the updated image feature to obtain the scene graph.

[0116] Optionally, when the image generation model starts the image generation task, the image generation model randomly initializes the image feature (i.e., image token) in the latent space to obtain an initial image feature, which represents potential image information.

[0117] After obtaining the target feature, the image generation model concatenates the target feature and the initial image feature and injects them as input features into the multi-modal attention module. The multi-module attention module updates the initial image feature according to the target feature and the initial image feature to obtain the updated image feature. In this process, the multi-module attention module interacts the target feature with the initial image feature in the latent space, dynamically adjusting the weight of the initial image feature to make it better reflect the requirements of the scene description, the appearance of the commodity, the layout of the paired items, and the spatial structure, so that the updated image feature gradually approaches the scene graph to be generated. Among them, during the processing of the multi-modal attention module, the target feature will also be updated.

[0118] In an alternative embodiment, when there is one multi-modal attention module in the image generation model, the output result of the multi-modal attention module includes updated image features and updated target features, and the image generation model can directly determine the updated image features from the output result of the multi-modal attention module.

[0119] In an alternative embodiment, when there are multiple multi-modal attention modules in the image generation model, the multiple multi-modal attention modules are connected in series, and the output of the last multi-modal attention module among the multiple multi-modal attention modules is determined as the updated image features. For example, assuming there are N multi-modal attention modules, where N is a positive integer greater than 1, the input of the first multi-modal attention module among the N multi-modal attention modules is the feature obtained by concatenating the target features and the initial image features. For the multi-modal attention modules after the first multi-modal attention module, the input of this multi-modal attention module is the output of its corresponding previous multi-modal attention module, and the updated image features are determined from the output result of the Nth multi-modal attention module. The output result of the Nth multi-modal attention module includes updated image features and updated target features.

[0120] Optionally, after obtaining the updated image features, the updated image features can be decoded to obtain a scene graph. Optionally, the decoder for decoding can adopt the decoder structure in the aforementioned generative model that combines an autoencoder and probability modeling.

[0121] In an alternative embodiment, it can be implemented using Figure 4 the schematic diagram shown to generate the scene graph. Figure 4 is a working schematic diagram of the image generation model provided in the first embodiment of the present application. As Figure 4 shown, the text information and spatial structure information are encoded by a first encoder to obtain a first feature, the image information is encoded by a second encoder to obtain a second feature, and the text-image information is encoded by a third encoder to obtain a third feature. Then, after adding position encoding to the first feature, the second feature, and the third feature and concatenating them, the target features are obtained. After that, the target features and the initial image features are input into the multi-modal attention module, and the updated image features are output through the multi-modal attention module. The updated image features are decoded to obtain a scene graph.

[0122] It should be noted that by using the multi-modal attention module to process the target features, the image generation model can effectively integrate and utilize the information in the target features, dynamically adjust the initial image features, so as to accurately control the commodities, matching items and scene layouts during the generation process. The decoding process can convert the adjusted features into a visual scene graph, thus realizing the complete process from conditional input to final image generation and improving the accuracy of image generation.

[0123] To improve the processing effect of the initial image features, in the image generation method provided in Embodiment 1 of this application, the multi-modal attention module in the image generation model updates the initial image features according to the target features and the initial image features. The updated image features obtained include: normalizing the input features of the multi-modal attention module through a normalization layer to obtain a fourth feature, where the input features are determined based on the target features and the initial image features; processing the fourth feature through a linear layer to obtain a fifth feature, and processing the fourth feature through a multi-modal attention layer to obtain a sixth feature; processing the fifth feature and the sixth feature through a projection layer to obtain a seventh feature; performing a residual connection on the seventh feature and the input features to obtain the output features of the multi-modal attention module, and determining the updated image features according to the output features.

[0124] In an optional embodiment, the Figure 5 schematic diagram shown can be used to process the target features and the initial image features. Figure 5 is a schematic diagram of the multi-modal attention module provided in Embodiment 1 of this application. As Figure 5 shown, the input features of the multi-modal attention module (i.e., Figure 5 the (X h , X T , X I , X0) in h are determined based on the target features and the initial image features. Among them, X T is determined based on the initial image features, and X I , X T and X0 are determined based on the target features. For example, X I is determined based on the target first feature in the target features, X h is determined based on the target second feature in the target features, and X0 is determined based on the target third feature in the target features.

[0125] For example, when there is one multi-modal attention module in the image generation model, Figure 5 the input features of the multi-modal attention module in T, X I And X0 is equivalent to the target feature.

[0126] For another example, when there are multiple multi-modal attention modules in the image generation model, if Figure 5 the multi-modal attention module in Figure 5 is the first one among multiple multi-modal attention modules, its input feature is the feature obtained by concatenating the target feature and the initial image feature. If

[0127] Optionally, as Figure 5 shown, the input feature is input into the multi-modal attention module. First, the input feature is normalized by the normalization layer to obtain the fourth feature, so that different features can be compared and fused on the same scale. Then, the fourth feature is linearly transformed by the linear layer to obtain the fifth feature to enhance the expression ability of the model, and the fourth feature is processed by the multi-modal attention layer to obtain the sixth feature. Among them, the multi-modal attention layer is used to realize the interaction between the target feature and the initial image feature and dynamically adjust the generation process through the attention mechanism. Then, the fifth feature and the sixth feature are processed by the projection layer to obtain the seventh feature. Among them, the projection layer is used to project the processed feature back to the latent space to adapt to subsequent processing or output requirements. After obtaining the seventh feature, the seventh feature is subjected to residual connection with the input feature of the current multi-modal attention module to retain the original information and obtain the output feature of the multi-modal attention module.

[0128] Optionally, if the current multi-modal attention module is the last multi-modal attention module in the image generation model, the updated image feature is directly determined from the output feature. If the current multi-modal attention module is not the last multi-modal attention module in the image generation model, the output feature is input into the next multi-modal attention module for further processing until it is determined that the last multi-modal attention module in the image generation model has completed the processing of the image feature, and the updated image feature is determined from the output feature of the last multi-modal attention module.

[0129] It should be noted that through the processing of the normalization layer, the linear layer, the multi-modal attention layer and the projection layer, and the introduction of the residual connection, not only the visual quality and semantic consistency of the generated scene graph are improved, but also the response ability of the model to complex control conditions is enhanced, thereby improving the accuracy of image generation.

[0130] In order to improve the processing capability of the image generation model, in the image generation method provided in Example 1 of the present application, the image generation model is obtained in the following manner: obtaining a training sample set, wherein the training samples in the training sample set include sample information, the sample information includes image information of sample products, and at least one of the following information: text information used to describe the sample scene to be generated, graphic information of items matched with the sample product, spatial structure information of the sample scene to be generated, and the true label of the training sample is a sample scene graph containing the sample product; training the initial image generation model based on the training sample set to obtain the image generation model.

[0131] In an optional embodiment, the sample information includes image information of the sample product, text information used to describe the sample scene to be generated (or graphic information of items matching the sample product, or spatial structure information of the sample scene to be generated), and at least one of the remaining information.

[0132] In an optional embodiment, the sample information includes image information of the sample product, text information for describing the sample scene to be generated, graphic information of items matching the sample product, and spatial structure information of the sample scene to be generated.

[0133] Optionally, after obtaining the training sample set, the target processing system can train the initial image generation model according to the training sample set to obtain the image generation model. During the training process, the parameters of the multiple encoders in the image generation model can remain unchanged, while the parameters of the multimodal attention module and the position encoding module can be continuously updated as the model is trained.

[0134] In an optional embodiment, the Figure 6 , Figure 7 , Figure 8 , Figure 9 , Figure 10 as well as Figure 11 The schematic diagram shown determines the improvement effect of the scene graph generated by the method provided by the present application compared with the scene graph generated by the method in the related art. Figure 6 It is a schematic diagram of a scene graph provided according to the relevant technology Figure 1 , Figure 7 This is a schematic diagram of a scene diagram provided according to the first embodiment of the present application. Figure 1 ,like Figure 6 , Figure 7 As shown, Figure 6 , Figure 7 The target products in the above table are all shoe cabinets. Figure 7 , Figure 6 The shoe cabinet layout has changed, making it difficult to accurately represent the product structure. Figure 7 The expression of commodities is more accurate. Figure 8It is a schematic diagram of a scene graph provided according to related technologies Figure 2 , Figure 9 It is a schematic diagram of a scene graph provided according to the first embodiment of the present application Figure 2 , such as Figure 8 , Figure 9 shown. Compared with Figure 9 , Figure 8 the environmental light and shadow is dim Figure 9 while the environmental light and shadow of Figure 10 is more harmonious Figure 3 , Figure 11 It is a schematic diagram of a scene graph provided according to the first embodiment of the present application Figure 3 , such as Figure 10 , Figure 11 shown. Compared with Figure 11 , Figure 10 the tea table in Figure 11 is placed off-center from the sofa, and the environment of

[0135] It should be noted that through the above method, the effective training of the image generation model is achieved, so as to improve the processing ability of the trained image generation model

[0136] In the embodiments of the present application, by obtaining target information and processing the target information through an image generation model, a variety of key information related to the scene graph to be generated is injected into the image generation model as control conditions, so as to accurately control the generation process of the scene graph. By designing corresponding encoders for different information in the target information, the image generation model can extract more accurate features specifically, so as to better understand and restore the details and semantics of the scene graph, and improve the generation accuracy of the scene graph. The purpose of generating a scene graph based on multiple control conditions is achieved, thereby achieving the technical effect of improving the generation accuracy of the scene graph, and further solving the technical problem of low generation accuracy of the scene graph when generating a scene graph containing a product image for a product in related technologies

[0137] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application

[0138] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application.

[0139] Embodiment 2

[0140] According to an embodiment of the present application, there is also provided a method for generating an image, as Figure 12 shown, the method includes:

[0141] Step S1201, obtain target information uploaded by a client, where the target information includes image information of a target commodity, and at least one of the following information: text information for describing a scene to be generated, graphic and text information of an item to be paired with the target commodity, and spatial structure information of the scene to be generated.

[0142] Step S1202, perform encoding processing on the target information through multiple encoders in an image generation model in a cloud server to obtain target features, and generate a scene graph including the target commodity based on the target features, where the multiple encoders are used to process different information in the target information.

[0143] Step S1203, feedback the scene graph to the client.

[0144] In the embodiment of the present application, by obtaining target information and processing the target information through an image generation model, it is realized that a variety of key information related to the scene graph to be generated is injected into the image generation model as control conditions, so as to accurately control the generation process of the scene graph. By designing corresponding encoders for different information in the target information, the image generation model can extract more accurate features specifically, so as to better understand and restore the details and semantics of the scene graph, and improve the generation accuracy of the scene graph. The purpose of generating a scene graph based on multiple control conditions is achieved, thereby achieving the technical effect of improving the generation accuracy of the scene graph, and further solving the technical problem of low generation accuracy of the scene graph when generating a scene graph including a commodity image for a commodity in the related art.

[0145] In the cloud server, the specific method for generating an image is the same as the method in Embodiment 1, and will not be elaborated here.

[0146] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0147] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, the technical solution of this application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of the various embodiments of this application.

[0148] Embodiment 3

[0149] According to an embodiment of this application, there is also provided an image generation device for implementing the above image generation method, as Figure 13 shown. The device includes: a first acquisition unit 1301 and a generation unit 1302.

[0150] The first acquisition unit 1301 is used to acquire target information, where the target information includes image information of a target commodity and at least one of the following information: text information for describing a scene to be generated, graphic and text information of an item to be paired with the target commodity, and spatial structure information of the scene to be generated;

[0151] The generation unit 1302 is used to perform encoding processing on the target information through multiple encoders in the image generation model to obtain target features, and generate a scene graph including the target commodity based on the target features, where the multiple encoders are used to process different information in the target information.

[0152] In the image generation device provided in the third embodiment of the present application, the target information is obtained by the first acquisition unit 1301, where the target information includes the image information of the target commodity and at least one of the following information: the text information for describing the scene to be generated, the graphic and text information of the items paired with the target commodity, and the spatial structure information of the scene to be generated; the generation unit 1302 encodes the target information through multiple encoders in the image generation model to obtain target features, and generates a scene graph including the target commodity based on the target features, where the multiple encoders are used to process different information in the target information. In this solution, by obtaining the target information and processing the target information through the image generation model, a variety of key information related to the scene graph to be generated is injected into the image generation model as control conditions, so that the generation process of the scene graph can be accurately controlled. By designing corresponding encoders for different information in the target information, the image generation model can extract more accurate features specifically, so as to better understand and restore the details and semantics of the scene graph, and improve the generation accuracy of the scene graph. The purpose of generating a scene graph based on multiple control conditions is achieved, thus realizing the technical effect of improving the generation accuracy of the scene graph, and further solving the technical problem of low generation accuracy of the scene graph when generating a scene graph including the commodity image for the commodity in the related art.

[0153] Optionally, in the image generation device provided in the third embodiment of the present application, the generation unit includes: a first processing subunit, configured to encode the text information and the spatial structure information through a first encoder to obtain first features; a second processing subunit, configured to encode the image information through a second encoder to obtain second features; a third processing subunit, configured to encode the graphic and text information through a third encoder to obtain third features; and a determination subunit, configured to determine the target features according to the first features, the second features, and the third features.

[0154] Optionally, in the image generation device provided in the third embodiment of the present application, the image information includes a first image and a second image, the first image refers to a scene canvas image including the target commodity, the second image refers to a mask image of the target commodity, and the second processing subunit includes: a first processing module, configured to encode the first image through a second encoder to obtain first sub-features; a second processing module, configured to perform block encoding processing on the second image to obtain second sub-features; and a third processing module, configured to splice the first sub-features and the second sub-features to obtain second features.

[0155] Optionally, in the image generation device provided in Embodiment 3 of this application, the determination subunit includes: a first determination module, configured to determine a first position encoding corresponding to a first feature, a second position encoding corresponding to a second feature, and a third position encoding corresponding to a third feature through a position encoding module in the image generation model; a second determination module, configured to determine a target first feature according to the first feature and the first position encoding, determine a target second feature according to the second feature and the second position encoding, and determine a target third feature according to the third feature and the third position encoding through the position encoding module; and a third determination module, configured to determine a target feature according to the target first feature, the target second feature, and the target third feature.

[0156] Optionally, in the image generation device provided in Embodiment 3 of this application, the image information includes a first image, where the first image refers to a scene canvas image including a target commodity, and the graphic and text information includes the position information of an item paired with the target commodity in the scene canvas image. The first determination module includes: a first processing sub-module, configured to perform encoding processing on preset position information to obtain a first position encoding; a second processing sub-module, configured to determine the target position of the target commodity in the scene canvas image according to the first image, and perform encoding processing on the target position to obtain a second position encoding; and a third processing sub-module, configured to determine the central position of the item in the scene canvas image according to the position information of the item in the scene canvas image, and perform encoding processing on the central position to obtain a third position encoding.

[0157] Optionally, in the image generation device provided in Embodiment 3 of this application, the generation unit includes: a fourth processing sub-unit, configured to update the initial image feature according to the target feature and the initial image feature through a multi-modal attention module in the image generation model to obtain an updated image feature; and a fifth processing sub-unit, configured to perform decoding processing on the updated image feature to obtain a scene graph.

[0158] Optionally, in the image generation device provided in Embodiment 3 of this application, the fourth processing sub-unit includes: a fourth processing module, configured to normalize the input feature of the multi-modal attention module through a normalization layer to obtain a fourth feature, where the input feature is determined based on the target feature and the initial image feature; a fifth processing module, configured to process the fourth feature through a linear layer to obtain a fifth feature, and process the fourth feature through a multi-modal attention layer to obtain a sixth feature; a sixth processing module, configured to process the fifth feature and the sixth feature through a projection layer to obtain a seventh feature; and a fourth determination module, configured to perform a residual connection on the seventh feature and the input feature to obtain the output feature of the multi-modal attention module, and determine the updated image feature according to the output feature.

[0159] Optionally, in the image generation device provided in Embodiment 3 of the present application, the image generation device includes: a second acquisition unit configured to acquire a training sample set, where the training samples in the training sample set include sample information, and the sample information includes image information of a sample commodity and at least one of the following information: text information for describing a to-be-generated sample scene, graphic and text information of an item paired with the sample commodity, and spatial structure information of the to-be-generated sample scene. The true label of the training sample is a sample scene graph including the sample commodity; a training unit configured to train an initial image generation model according to the training sample set to obtain an image generation model.

[0160] It should be noted here that the above first acquisition unit 1301 and generation unit 1302 correspond to steps S201 to S202 in Embodiment 1. The above units and the corresponding steps have the same implemented examples and application scenarios, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.

[0161] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0162] Embodiment 4

[0163] The embodiments of the present application can provide an electronic device, and the electronic device can be any one of a group of electronic devices. Optionally, in this embodiment, the above electronic device can also be replaced with a terminal device such as a mobile terminal.

[0164] Optionally, in this embodiment, the above electronic device can be located in at least one of a plurality of network devices in a computer network.

[0165] In this embodiment, the above electronic device can execute program codes corresponding to the steps in the image generation method provided in any one of the above method embodiments.

[0166] Optionally, Figure 14 is a structural block diagram of an electronic device according to an embodiment of the present application. As Figure 14 shown, the electronic device 140 may include: one or more ( Figure 14 only one is shown in the figure) processors 1402, a memory 1404. The electronic device 140 may further include a memory controller for controlling and managing the memory 1404; the electronic device 140 may further include a peripheral interface for connecting a radio frequency module, an audio module, a display screen, etc.

[0167] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the image generation method and device in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above-mentioned image generation method. The memory can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory can further include a memory remotely set relative to the processor, and these remote memories can be connected to the terminal 10 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.

[0168] The processor can call the information and application programs stored in the memory through the transmission device to execute the program code corresponding to the steps in the image generation method provided in any one of the above method embodiments.

[0169] Those of ordinary skill in the art can understand that Figure 14 the structure shown is only schematic, and the electronic device can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 14 It does not limit the structure of the above electronic device. For example, the electronic device 140 may further include more or fewer components (such as a network interface, a display device, etc.) than those shown Figure 14 in the figure, or have a different configuration from that shown Figure 14 in the figure.

[0170] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and this program can be stored in a computer-readable storage medium. The storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0171] Embodiment 5

[0172] The embodiment of the present application also provides a computer-readable storage medium. Optionally, in this embodiment, the above storage medium can be used to save the program code executed by the image generation method provided in the first embodiment above.

[0173] Optionally, in this embodiment, the above storage medium may be located in any one of the electronic devices in the electronic device group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0174] Embodiment 6

[0175] The embodiments of the present application further provide a computer program product. Optionally, in this embodiment, the above computer program product may include a computer program, and when the computer program is executed by a processor, it implements the method for generating an image provided in the first embodiment above.

[0176] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.

[0177] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0178] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0179] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0180] In addition, the functional units in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0181] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.

[0182] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A method for generating an image, characterized in that, Including: Obtain target information, where the target information includes image information of a target commodity and at least one of the following information: text information for describing a scene to be generated, graphic and text information of an item paired with the target commodity, and spatial structure information of the scene to be generated; Encode the target information through multiple encoders in an image generation model to obtain target features, and generate a scene graph including the target commodity based on the target features, where the multiple encoders are used to process different information in the target information.

2. The method according to claim 1, wherein When the target information includes the image information, the text information, the graphic and text information, and the spatial structure information, encoding the target information through multiple encoders in an image generation model to obtain target features includes: Encode the text information and the spatial structure information through a first encoder to obtain a first feature; Encode the image information through a second encoder to obtain a second feature; Encode the graphic and text information through a third encoder to obtain a third feature; Determine the target features according to the first feature, the second feature, and the third feature.

3. The method according to claim 2, wherein The image information includes a first image and a second image. The first image refers to a scene canvas image including the target commodity, and the second image refers to a mask image of the target commodity. Encoding the image information through a second encoder to obtain a second feature includes: Encode the first image through the second encoder to obtain a first sub-feature; Perform block encoding on the second image to obtain a second sub-feature; Perform splicing processing on the first sub-feature and the second sub-feature to obtain the second feature.

4. The method according to claim 2, wherein Determining the target features according to the first feature, the second feature, and the third feature includes: Determine a first position encoding corresponding to the first feature, a second position encoding corresponding to the second feature, and a third position encoding corresponding to the third feature through a position encoding module in the image generation model; Determine a target first feature according to the first feature and the first position encoding, determine a target second feature according to the second feature and the second position encoding, and determine a target third feature according to the third feature and the third position encoding through the position encoding module; Determine the target features according to the target first feature, the target second feature, and the target third feature.

5. The method according to claim 4, characterized in that The image information includes a first image, the first image refers to a scene canvas image including the target commodity, and the graphic and text information includes position information of an item paired with the target commodity in the scene canvas image. Determining a first position encoding corresponding to the first feature, a second position encoding corresponding to the second feature, and a third position encoding corresponding to the third feature through a position encoding module in the image generation model includes: Encode preset position information to obtain the first position encoding; Determine the target position of the target commodity in the scene canvas image according to the first image, and perform encoding processing on the target position to obtain the second position encoding; Determine the central position of the item in the scene canvas image according to the position information of the item in the scene canvas image, and perform encoding processing on the central position to obtain the third position encoding.

6. The method according to any one of claims 1 to 5, characterized in that, Generating a scene graph including the target commodity based on the target feature includes: Updating the initial image feature by the multi-modal attention module in the image generation model according to the target feature and the initial image feature to obtain an updated image feature; Performing decoding processing on the updated image feature to obtain the scene graph.

7. The method according to claim 6, characterized in that, Updating the initial image feature by the multi-modal attention module in the image generation model according to the target feature and the initial image feature to obtain an updated image feature includes: Normalizing the input feature of the multi-modal attention module through a normalization layer to obtain a fourth feature, where the input feature is determined based on the target feature and the initial image feature; Processing the fourth feature through a linear layer to obtain a fifth feature, and processing the fourth feature through a multi-modal attention layer to obtain a sixth feature; Processing the fifth feature and the sixth feature through a projection layer to obtain a seventh feature; Performing a residual connection between the seventh feature and the input feature to obtain the output feature of the multi-modal attention module, and determining the updated image feature according to the output feature.

8. The method according to any one of claims 1 to 5, characterized in that, The image generation model is obtained by the following method: Obtain a training sample set, where the training samples in the training sample set include sample information, the sample information includes the image information of the sample commodity, and at least one of the following information: text information for describing the sample scene to be generated, graphic and text information of the item paired with the sample commodity, spatial structure information of the sample scene to be generated, and the true label of the training sample is a sample scene graph including the sample commodity; Train an initial image generation model according to the training sample set to obtain the image generation model.

9. A method for generating an image, characterized in that, Includes: Obtain the target information uploaded by the client, where the target information includes the image information of the target commodity, and at least one of the following information: text information for describing the scene to be generated, graphic and text information of the item paired with the target commodity, spatial structure information of the scene to be generated; Encoding the target information through multiple encoders in the image generation model in the cloud server to obtain a target feature, and generating a scene graph including the target commodity based on the target feature, where the multiple encoders are used to process different information in the target information; Feedback the scene graph to the client.

10. An image generation device, characterized in that, Includes: A first acquisition unit, configured to acquire target information, where the target information includes image information of a target commodity and at least one of the following information: text information for describing a scene to be generated, graphic and text information of an item paired with the target commodity, and spatial structure information of the scene to be generated; A generation unit, configured to perform encoding processing on the target information through multiple encoders in an image generation model to obtain target features, and generate a scene graph including the target commodity based on the target features, where the multiple encoders are used to process different information in the target information.

11. An electronic device, characterized in that, Comprising: A memory storing an executable program; A processor, configured to run the program, where when the program runs, it executes the image generation method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, where when the executable program runs, it controls the device where the storage medium is located to execute the image generation method according to any one of claims 1 to 9.

13. A computer program product, characterized in that, Including a computer program or instruction, where when the computer program or instruction is executed by a processor, it implements the image generation method according to any one of claims 1 to 9.