Image generation method and device, computing platform and computer storage medium
By obtaining product images and reference product images in the artificial intelligence AI interactive interface of the large language model, and using the large language model to process to generate target product images, the problem of insufficient image details in the prior art is solved, efficient and accurate image generation is achieved, meeting user needs and improving the practicality of the method.
Patent Information
- Application Number
- CN202510199322.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-13
AI Technical Summary
The images generated by the prior art have low accuracy and degree of detail expression and cannot meet the user's image generation needs.
In the artificial intelligence AI interactive interface based on the large language model, the product image corresponding to the preset product is obtained, the reference product image corresponding to the reference product is determined, and the product image and the reference product image are processed using the large language model to generate the target product image corresponding to the preset product.
The efficiency of image generation operations is improved, time and labor costs are reduced, and the generated target product images are similar to the reference product images and have certain differences, meeting the merchant's image generation needs and ensuring uniqueness.
Smart Images

Figure CN120147448A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to an image generation method, apparatus, computing platform, and computer storage medium. Background Art
[0002] In the field of e-commerce technologies, images play a crucial role. High-quality and well-optimized images can not only enhance the user experience but also directly affect the sales conversion rate of products and brand recognition.
[0003] Currently, the images uploaded by merchants on e-commerce platforms can be images generated through text-to-image technologies. Specifically, users can input image prompts into a preset image generation tool. The image prompts can include: background descriptions, image sizes, image elements, etc. Then, the image generation tool can generate corresponding images based on the input image prompts.
[0004] However, the images generated by the above implementation method completely depend on the description of the input image prompts, which easily leads to low accuracy and detail in the expression of details in the generated images, and thus easily fails to meet the user's image generation requirements. Summary of the Invention
[0005] Embodiments of this application provide an image generation method, apparatus, computing platform, and computer storage medium, which can ensure the accuracy and detail in the expression of details in the generated images and can meet the user's image generation requirements.
[0006] Embodiments of the present invention provide an image generation method, including:
[0007] In an artificial intelligence (AI) interaction interface based on a large language model, obtain a product image corresponding to a preset product;
[0008] Determine a reference product image corresponding to a reference product, where the number of main bodies of the preset product corresponding to the product image is the same as the number of main bodies of the reference product corresponding to the reference product image;
[0009] Use the large language model to process the product image and the reference product image to generate at least one target product image corresponding to the preset product, where the similarity between the target product image and the reference product image is greater than or equal to a preset threshold and there are differences between the target product image and the reference product image.
[0010] Embodiments of the present invention provide an image generation apparatus, including:
[0011] The first acquisition module is used to acquire a product image corresponding to a preset product in an artificial intelligence (AI) interaction interface based on a large language model;
[0012] The first determination module is used to determine a reference product image corresponding to a reference product, where the number of main bodies of the preset product is the same as that of the reference product;
[0013] The first processing module is used to process the product image and the reference product image by using the large language model to generate at least one target product image corresponding to the preset product, where the similarity between the target product image and the reference product image is greater than or equal to a preset threshold, and there are differences between the target product image and the reference product image.
[0014] An embodiment of the present invention provides a computing platform, including: a memory and a processor; wherein, the memory is used to store one or more computer instructions, and when the one or more computer instructions are executed by the processor, the image generation method in the above first aspect is implemented.
[0015] An embodiment of the present invention provides a computer storage medium for storing a computer program, and when the computer program is executed by a computer, the image generation method in the above first aspect is implemented.
[0016] An embodiment of the present invention provides a computer program product, including: a computer-readable storage medium storing computer instructions, and when the computer instructions are executed by one or more processors, the steps in the image generation method in the above first aspect are caused to be executed by the one or more processors.
[0017] The image generation method, device, computing platform and computer storage medium provided in this embodiment, in an artificial intelligence (AI) interaction interface based on a large language model, by acquiring a product image corresponding to a preset product, determining a reference product image corresponding to a reference product, and then processing the product image and the reference product image by using the large language model to generate at least one target product image corresponding to the preset product, effectively realizes the technical solution of using the "reference product image" and the "product image" to generate the target product image, that is, realizes the automatic generation solution of the target product image, thereby not only effectively improving the efficiency of the image generation operation, reducing the time cost and labor cost required for the image generation operation, and moreover, the generated target product image is relatively similar to the reference product image and there are certain differences, so that not only can the generated target product image meet the image generation effect of the merchant, but also it can be determined that the generated target product image has a certain uniqueness, further improving the practicability of this method. Description of the Drawings
[0018] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0019] Figure 1 It is a schematic diagram of the scenario of an image generation method provided by an exemplary embodiment of the present application;
[0020] Figure 2 It is a schematic flowchart of an image generation method provided by an exemplary embodiment of the present application;
[0021] Figure 3 It is a schematic flowchart of obtaining a product image corresponding to a preset product in an artificial intelligence (AI) interaction interface based on a large language model according to an exemplary embodiment of the present application;
[0022] Figure 4 It is a schematic diagram of the interface for obtaining a product image corresponding to a preset product provided by an exemplary embodiment of the present application;
[0023] Figure 5 It is a schematic flowchart of determining a reference product image corresponding to a reference product provided by an exemplary embodiment of the present application;
[0024] Figure 6 It is a schematic flowchart of another image generation method provided by an exemplary embodiment of the present application;
[0025] Figure 7 It is a schematic flowchart of using the large language model to process the product image and the reference product image to generate at least one target product image corresponding to the preset product according to an exemplary embodiment of the present application;
[0026] Figure 8 It is a schematic flowchart of an image generation method provided by an exemplary application embodiment of the present application;
[0027] Figure 9 It is a schematic diagram of the structure of an image generation device provided by an exemplary embodiment of the present application;
[0028] Figure 10 It is a schematic diagram of the structure of a computing platform provided by an exemplary embodiment of the present application. Detailed implementation manners
[0029] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0030] It should be noted that in the case where the embodiments of this application involve user information, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject. Additionally, various models (including but not limited to language models or large models) involved in this application comply with relevant standard regulations.
[0031] Furthermore, it should be noted that in the case where the embodiments of this application involve user interaction operations or trigger operations, the user interaction operations or trigger operations involved in the embodiments of this application include but are not limited to: interaction operations in various ways such as touch operations, gesture operations, voice operations, head movement operations, and eye movement operations; among them, touch operations include but are not limited to: click operations, double-click operations, long-press operations, swipe operations, pinch operations, or mouse hover operations, etc. Swipe operations include but are not limited to: straight-line swipes, curved swipes, etc.
[0032] To facilitate the understanding of the image generation method, device, computing platform, and computer storage medium provided by the embodiments of this application, the related technologies will be briefly described below:
[0033] In the application scenario of e-commerce, merchants have a deep tradition and reliance on the use of competitor product images. In the past, in the era without the assistance of artificial intelligence (AI) technology, merchants often manually referred to or even directly copied the product images of competitors. Small and medium-sized merchants' definition of "high-quality" creative images would also refer to the creativity of industry-leading merchants and understand which creative features such as visual elements, copywriting expressions, and color combinations in the market were favored by consumers by observing competitor product images. However, the above image processing operations greatly increased the image production time and graphic design costs.
[0034] With the rapid development of AI technology, image generation models have also been widely applied in the e-commerce field. Currently, the images uploaded by merchants on e-commerce platforms can be images generated through text-to-image technology. Specifically, users can input image prompts into a preset image generation tool. The image prompts can include: background description words, image size, image elements, etc. Then, the image generation tool can generate corresponding images based on the input image prompts.
[0035] However, although the above implementation method can save the art design cost, picture production time and trial-and-error cost of images to a certain extent. However, the generated images completely depend on the description of the input image prompts, which easily leads to low accuracy and detail in the expression of details of the generated images, and thus does not meet the user's image generation requirements.
[0036] To solve the problems existing in the related technology, this embodiment provides an image generation method, device, computing platform and computer storage medium. Specifically, refer to the appendix Figure 1 As shown, the execution subject of this image generation method can be the image generation device 200. The image generation device 200 can be implemented as a local server, a cloud server or a preset device. Among them, when the image generation device 200 is implemented as a cloud server, this image generation method can be executed in the cloud. A number of computing nodes (cloud servers) can be deployed in the cloud, and each computing node has processing resources such as computing and storage. In the cloud, multiple computing nodes can be organized to provide a certain service. Of course, a single computing node can also provide one or more services. The way for the cloud to provide this service can be to provide a service interface externally, and users call this service interface to use the corresponding service. The service interface includes forms such as Software Development Kit (SDK) and Application Programming Interface (API).
[0037] The image generation device 200 is communicatively connected to the client 100. Among them, the client 100 is used for users to perform applications so as to trigger an image generation operation. The above-mentioned client 100 can be any computing device with certain information interaction capabilities. Specifically, in implementation, the client 100 can be a mobile phone, a personal computer (PC), a tablet computer, a set application program, etc. In addition, the basic structure of the client 100 may include: at least one processor. The number of processors depends on the configuration and type of the client. The client 100 may also include a memory, which can be volatile, for example: Random Access Memory (RAM), or non-volatile, for example: Read-Only Memory (ROM), flash memory, etc., or may also include both types at the same time. Usually, an operating system (OS), one or more application programs, and program data, etc. are stored in the memory. In addition to the processing unit and the memory, the client 100 also includes some basic configurations, such as: a network card chip, an IO bus, a display component, and some peripheral devices, etc. Optionally, some peripheral devices may include, for example: a keyboard, a mouse, a stylus, a printer, etc. Other peripheral devices are well known in the art and will not be elaborated here.
[0038] The image generation device 200 refers to a device that can implement an image generation operation in a network virtual environment, usually referring to a device that uses a network for information planning and image generation operations. Physically, the image generation device 200 can be any device that can provide computing services and can perform corresponding image generation operations, such as: it can be a cluster server, a conventional server, a cloud server, a cloud host, a virtual center, etc. The composition of the image generation device 200 mainly includes a processor, a hard disk, a memory, a system bus, etc., which is similar to a general computer architecture.
[0039] In the above-mentioned embodiment of the present invention, the client 100 is network-connected to the image generation device 200, and this network connection can be a wireless or wired network connection. If the client 100 can be communicatively connected to the image generation device 200, the network mode of this mobile network can be any one of 2G (Global System for Mobile Communications GSM), 2.5G (General Packet Radio Service GPRS), 3G (Wideband Code Division Multiple Access WCDMA, Time Division-Synchronous Code Division Multiple Access TD-SCDMA), 4G (Long Term Evolution LTE), 4G+ (Enhanced Long Term Evolution LTE+), Worldwide Interoperability for Microwave Access WiMax, 5G, 6G, etc.
[0040] In the embodiments of the present application, the client 100 is for users to use, so as to be able to perform AI interaction operations with the large language model deployed in the image generation device 200. Specifically, an AI interaction interface based on the large language model in the image generation device 200 can be displayed on the client 100. The user can input an execution operation in the AI interaction interface, and the execution operation can be an interaction operation in natural language, so as to clarify the product image corresponding to the preset product for which image generation operation is required, where the number of preset products is one or more. For example: the preset product can be implemented as a liquid foundation of a preset brand, and the product image can at least include relevant information about the packaging bottle of the liquid foundation. As for Figure 1 other information in the product image is irrelevant to the image generation operation. For example, the brand information and foundation description information included on the packaging bottle of the liquid foundation are irrelevant to the processing operation of the image generation operation.
[0041] Image generation device 200: When the user inputs an execution operation in the AI interaction interface of the large language model, the image generation device 200 can obtain the product image corresponding to the preset product in the AI interaction interface based on the large language model. In order to implement the image generation operation, after obtaining the product image corresponding to the preset product, a reference product image corresponding to the reference product can be determined, where the category of the reference product can be the same as or different from the category of the preset product, and the number of main bodies of the preset product corresponding to the product image is consistent with the number of main bodies of the reference product corresponding to the reference product image. That is, when the number of main bodies of the preset product corresponding to the product image is 1, the number of main bodies of the reference product corresponding to the determined reference product image is also 1; when the number of main bodies of the preset product corresponding to the product image is 3, the number of main bodies of the reference product corresponding to the determined reference product image is also 3. This can ensure the quality and effect of the image generation operation for the preset product in the product image based on the reference product image to a certain extent.
[0042] After obtaining the product image and the reference product image, in order to ensure the quality and effect of the image generation to a certain extent, the pre-trained large language model can be used to process the product image and the reference product image, so as to generate at least one target product image corresponding to the preset product. For example, refer to the appendix Figure 1As shown, when a merchant user has an image generation requirement for a product image, the product image can be obtained in the AI interaction interface of the large language model. When the product image corresponds to Foundation A (i.e., the preset product), the reference product image corresponding to Foundation B (i.e., the reference product) needs to be determined at the same time. Then, the large language model can be used to analyze and process the reference product image and the product image, that is, the large language model can learn the image features of the reference product image and perform an image generation operation based on the learned image features, so as to generate at least one target product image corresponding to Foundation A, thus effectively completing the image generation operation.
[0043] For the target product image, the number can be one or more. Among them, the large language model is trained to perform an image generation operation. The similarity between the target product image and the reference product image is greater than or equal to a preset threshold, and there are differences between the target product image and the reference product image, that is, the similarity between the generated target product image and the reference product image is relatively high. Generally, the similarity between the target product image and the reference product image can be reflected by at least one of the following aspects: the similarity between the background of the target product image and the background of the reference product image, the similarity between the color tone of the target product image and the color tone of the reference product image, the similarity between the display effect features of the preset product in the target product image and the display effect features of the reference product in the reference product image, etc. Moreover, there must be certain differences between the target product image and the reference product image, so as to ensure to a certain extent that the generated target product image has a certain uniqueness.
[0044] In this embodiment, the technical solution of generating a "target product image" by using a "reference product image" and a "product image" is effectively realized, that is, the automatic generation scheme of image generating image is realized. Thus, it not only effectively improves the efficiency of image generation, reduces the time cost and labor cost required for image generation, but also the generated target product image is relatively similar to the reference product image and has certain differences. In this way, it can not only make the generated target product image meet the image generation effect of the merchant, but also ensure that the generated target product image has a certain uniqueness, further improving the practicability of this method.
[0045] The following will describe in detail the technical solutions provided by each embodiment of the present application with reference to the accompanying drawings.
[0046] Figure 2 It is a schematic flowchart of an image generation method provided by an exemplary embodiment of the present application; refer to the attached Figure 2As shown in the figure, this embodiment provides an image generation method, and the execution subject of this method is an image generation device. Among them, the image generation device can be implemented as software, or a combination of software and hardware. When the image generation device is implemented as hardware, it can specifically be various computing platforms capable of implementing image generation operations, including but not limited to personal computers, servers, etc. When the image generation device is implemented as software, it can be installed in the computing platforms exemplified above. Based on the above image generation device, image generation operations can be realized. Specifically, this image generation method can include:
[0047] Step S201: In an artificial intelligence AI interaction interface based on a large language model, obtain a product image corresponding to a preset product.
[0048] Step S202: Determine a reference product image corresponding to a reference product, and the number of main bodies of the preset product corresponding to the product image is the same as the number of main bodies of the reference product corresponding to the reference product image.
[0049] Step S203: Use the large language model to process the product image and the reference product image to generate at least one target product image corresponding to the preset product. The similarity between the target product image and the reference product image is greater than or equal to a preset threshold, and there are differences between the target product image and the reference product image.
[0050] The specific implementation methods and implementation principles of the above steps will be described in detail below:
[0051] Step S201: In an artificial intelligence AI interaction interface based on a large language model, obtain a product image corresponding to a preset product.
[0052] Among them, a large language model implemented based on artificial intelligence AI technology for realizing image generation operations can be deployed in the image generation device. In order to be able to realize image generation operations, an artificial intelligence AI interaction interface based on the large language model can be displayed in the image generation device. This AI interaction interface is used to interact with merchants in natural language to obtain a product image corresponding to a preset product that can trigger image generation operations.
[0053] In some instances, the image generation operation can be selectively triggered by the large language model deployed in the image generation device based on the detection results of the delivery effects for the preset products in the merchant stores. At this time, in the artificial intelligence (AI) interaction interface based on the large language model, obtaining the product image corresponding to the preset product may include: obtaining the delivery effect information of the product image corresponding to the preset product based on the large language model; when the delivery effect information does not meet the preset requirements, generating an image generation suggestion corresponding to the preset product based on the large language model; and displaying the product image corresponding to the preset product and the image generation suggestion in the AI interaction interface based on the large language model.
[0054] Specifically, for each merchant store in the e-commerce platform, in order to improve the intelligence level of the image generation operation to a certain extent, with the authorization of the merchant user obtained, the large language model deployed in the image generation device can automatically or periodically detect the delivery effects of at least some of the products in each merchant store. Among them, at least some of the products corresponding to the delivery effect detection operation can be specified by the merchant or automatically selected by the large language model. When at least some of the products corresponding to the delivery effect detection operation are automatically selected by the large language model, the large language model can automatically screen based on the advertising delivery costs of each product in the merchant store. In some instances, at least some of the products with higher advertising delivery costs can be selected as the preset products for which the delivery effect detection operation needs to be performed.
[0055] After the large language model detects the delivery effects of at least some of the products in the merchant store, each of the at least some of the detected products can be used as a preset product. Then, the delivery effect information of the product image corresponding to the preset product can be obtained. The delivery effect can be reflected by at least one of the following indicators: the click-through rate of the preset product, the click volume of the preset product, the conversion rate of the preset product, and so on. After obtaining the delivery effect, the delivery effect can be analyzed and processed to identify whether the delivery effect of the current preset product meets the preset requirements. When the identification result is that the delivery effect of the current preset product meets the preset requirements, it means that the delivery effect of the current preset product is good and meets the user's delivery needs. Therefore, there is no need to perform any modification operations on the release and promotion information of the preset product.
[0056] When the recognition result is that the delivery effect of the current preset product does not meet the preset demand, it means that the delivery effect of the current preset product is poor and does not meet the user's delivery demand. At this time, since the delivery effect of the preset product is likely to be related to the release image of the preset product, in order to ensure or even improve the delivery effect of the preset product, an image generation suggestion corresponding to the preset product can be generated based on the large language model. The image generation suggestion can be specifically expressed as "the delivery effect of the preset product is poor, and it is recommended to regenerate the product release image of the preset product to improve the delivery effect of the preset product". After generating the image generation suggestion corresponding to the preset product, in order to facilitate merchant users to promptly know the image generation suggestion for the preset product, the product image and image generation suggestion corresponding to the preset product can be displayed in the AI interactive interface based on the large language model, so that the product image corresponding to the preset product can be obtained under the active triggering of the large language model.
[0057] In other instances, when the delivery effect information does not meet the preset requirements, in order to further improve the accuracy and reliability of the method, before generating the image generation suggestion corresponding to the preset product based on the large language model, the large language model can be used to identify the correlation between the current delivery effect and the product image corresponding to the preset product, that is, to analyze the influence of the product image corresponding to the preset product on the delivery effect of the preset product; when the correlation between the delivery effect and the product image is greater than or equal to the preset threshold, it means that the product image has a greater influence on the delivery effect of the preset product, and then the large language model is allowed to be used to generate the image generation suggestion corresponding to the preset product; when the correlation between the delivery effect and the product image is less than the preset threshold, it means that the product image has a less influence on the delivery effect of the preset product, and then the large language model can be prohibited from being used to generate the image generation suggestion corresponding to the preset product, so that when the delivery effect is poor due to the product image, the image generation operation can be triggered; when the delivery effect is poor due to other non-product images, the image generation operation will not be triggered, thereby ensuring the quality and effectiveness of the image generation operation.
[0058] Step S202: determining a reference product image corresponding to the reference product, wherein the number of entities of the preset product corresponding to the product image is consistent with the number of entities of the reference product corresponding to the reference product image.
[0059] After obtaining the product image, in order to ensure the quality and effect of the image generation operation, a reference product image corresponding to the reference product can be determined. For the reference product image and the product image, the number of main bodies of the preset product corresponding to the product image can be made consistent with the number of main bodies of the reference product corresponding to the reference product image. In this way, when using the large language model for image generation operations, the large language model can learn the image features of the reference product image and can stably and effectively apply the image features to the preset product, thereby ensuring the generation quality and effect of the target product image to a certain extent.
[0060] In some instances, the reference product image can be the image corresponding to the reference product determined by the large language model to have a better placement effect (i.e., the placement effect index meets the preset requirements). Among them, the reference product can meet at least one of the following conditions: the leaf category of the reference product is the same as or similar to the leaf category of the preset product, the price range of the reference product is the same as or similar to the price range of the preset product, the target population of the reference product is the same as or similar to the target population of the preset product. At this time, determining the reference product image corresponding to the reference product can include: determining a reference product adapted to the preset product; obtaining the published image corresponding to the reference product through the e-commerce platform; based on the published image, determining the reference product image corresponding to the reference product, where all the published images can be directly determined as the reference product image corresponding to the reference product; or, the main view published image in the published images can be determined as the reference product image corresponding to the reference product, so as to accurately and effectively determine the reference product image corresponding to the reference product.
[0061] In other instances, the reference product image can not only be the image corresponding to the reference product determined by the large language model to have a better placement effect, but also be the image uploaded by the user. At this time, determining the reference product image corresponding to the reference product can include: displaying an image upload control in the artificial intelligence AI interaction interface based on the large language model; obtaining the uploaded image in response to the operation input by the user for the image upload control; determining the uploaded image as the reference product image corresponding to the reference product.
[0062] Among them, since the reference product image determined by the large language model may not meet the image generation requirements of merchant users. For example, the image background of the reference product image cannot meet the background generation requirements of merchant users, the color tone of the reference product image cannot meet the color tone generation requirements of merchant users, the display features of the reference product in the reference product image cannot meet the product display requirements of merchant users, etc. At this time, in order to enable the large language model to learn as much as possible the image features that can meet the needs of merchant users, merchant users can independently upload the reference product image corresponding to the reference product.
[0063] Specifically, in order to enable users to actively upload reference product images for image generation operations, an image upload control can be displayed in the artificial intelligence AI interaction interface based on the large language model. The image upload control can be implemented by clicking on the control or swiping the control, etc. When the user performs or inputs a selection operation on the image upload control, the image upload operation can be triggered, and the user can be reminded to select an image storage path. By accessing the image storage path and uploading the image stored in the preset area, the uploaded image can be obtained. Then, the uploaded image can be directly determined as the reference product image corresponding to the reference product, which also ensures the stability and reliability of determining the reference product image, and to a certain extent ensures that the determined reference product image can meet the image generation needs of merchant users.
[0064] It should be noted that when determining the reference product image corresponding to the reference product based on the user's upload operation, the user can perform the image upload work based on the preset image screening rules. The above image screening rules can include at least one of the following: the leaf category of the reference product corresponding to the uploaded image is the same as or similar to the leaf category of the preset product, the price range of the reference product corresponding to the uploaded image is the same as or similar to the price range of the preset product, and the target population of the reference product corresponding to the uploaded image is the same as or similar to the target population of the preset product.
[0065] Alternatively, the user can also perform the image upload work without relying on the image screening rules. At this time, for the uploaded image, the uploaded image can correspond to a reference product, and the number of main bodies of the reference product corresponding to the uploaded image should be consistent with the number of main bodies of the preset product corresponding to the product image. However, for the reference product corresponding to the uploaded image, the leaf category of the reference product may be the same as or different from the leaf category of the preset product, the price range of the reference product may be the same as or different from the price range of the preset product, and the target population of the reference product may be the same as or different from the target population of the preset product.
[0066] In addition, for the uploaded image, since it is an image obtained based on the user's upload operation, in order to avoid the situation where the generated target product image cannot be applied due to an image generation operation using a non-compliant uploaded image; after obtaining the uploaded image, the uploaded image can be first subjected to compliance detection, and then it can be determined whether the uploaded image can be determined as the reference product image based on the compliance detection result. At this time, after obtaining the uploaded image, the method in this embodiment can further include: obtaining a preset rule for detecting the uploaded image; using the preset rule to perform compliance detection on the uploaded image; and allowing the uploaded image to be determined as the reference product image corresponding to the reference product when the uploaded image passes the compliance detection.
[0067] After obtaining the uploaded image based on the user's upload operation, in order to ensure that the uploaded image meets the rules and requirements for image generation, a preset rule for detecting the uploaded image can be obtained. This preset rule is used to perform compliance detection on the uploaded image, and it can be a rule set by humans or a platform release rule determined in the image generation device. For example, the preset rule can include at least one of the following: whether there are violation marks in the image, whether there are illegal marks in the image, whether the number of subjects corresponding to the image is consistent with the number of subjects corresponding to the product image, etc. The specific content it includes can be flexibly configured and adjusted based on preset requirements.
[0068] After obtaining the preset rule, the preset rule can be used to perform a compliance detection operation on the uploaded image, that is, to determine whether the uploaded image meets the compliance requirements set by the preset rule. When the uploaded image meets all the compliance requirements set by the preset rule, it means that the uploaded image has passed the compliance detection. At this time, it is allowed to determine the uploaded image as the reference product image corresponding to the reference product, which ensures the reasonableness and legality of determining the reference product image to a certain extent. Correspondingly, when the uploaded image does not meet any of the compliance requirements set by the preset rule, it means that the uploaded image has not passed the compliance detection. At this time, in order to facilitate the user to timely understand the situation and reason why the uploaded image cannot be used as the reference product image, the uploaded image and the preset rule can be analyzed and processed by a large language model to determine the reason for non-passing, and the reason for non-passing can be associated and displayed with the uploaded image in the AI interaction interface based on the large language model. This can enable the user to quickly understand the unsuccessful situation and reason of the uploaded image, which is conducive to improving the user's good experience of performing image generation operations to a certain extent.
[0069] Step S203: Use a large language model to process the product image and the reference product image to generate at least one target product image corresponding to the preset product, where the similarity between the target product image and the reference product image is greater than or equal to a preset threshold, and there are differences between the target product image and the reference product image.
[0070] After obtaining the product image and the reference product image, in order to apply the image features of the reference product image to the preset product corresponding to the product image as much as possible, a large language model can be used to analyze and process the product image and the reference product image. Specifically, the product image and the reference product image are input into the large language model for analysis and processing to obtain at least one target product image output by the large language model and corresponding to the preset product, thus completing the generation operation of at least one target product image. And the number of generated target product images can be one or more.
[0071] Among them, for the generated target product image, it needs to meet the following requirements: the similarity between the target product image and the reference product image is greater than or equal to a preset threshold (i.e., the preset similarity threshold), that is, the target product image is similar to the reference product image. Specifically, the target product image and the reference product image may be similar in at least one of the following aspects: image background, image tone, image layout, etc. And, in order to ensure that the generated target product image has a certain uniqueness, there should be a difference between the generated target product image and the reference product image, that is, the difference degree between any target product image and the reference product image is greater than or equal to the preset difference degree threshold.
[0072] In addition, the specific implementation method for generating the target product image in this embodiment is not limited. In some examples, the target product image can be generated through the text-to-image link and image-to-image link included in the large language model. Among them, the large language model may include: a text-to-image sub-link and an image-to-image sub-link. At this time, using the large language model to process the product image and the reference product image to generate at least one target product image corresponding to the preset product may include: using the text-to-image sub-link in the large language model to analyze and process the product image and the reference product image to generate a first processing result corresponding to the preset product; using the image-to-image sub-link of the large language model to analyze and process the product image and the reference product image to generate a second processing result corresponding to the preset product; using the first processing result and the second processing result to generate at least one target product image corresponding to the preset product, which effectively ensures the accuracy and reliability of generating the target product image.
[0073] In other examples, when using the large language model to analyze and process the reference product image and the product image, for the reference product image, it can be classified according to whether it includes a human object. And, different types of reference product images may correspond to different processing strategies. Then, use the large language model and the processing strategy to perform processing operations on the product image and the reference product image. Specifically, using the large language model to process the product image and the reference product image to generate at least one target product image corresponding to the preset product may include: identifying whether the reference product image includes a human object; when the reference product image includes a human object, determining the image features irrelevant to the human object in the reference product image; using the large language model and the image features to process the product image and the reference product image to generate at least one target product image corresponding to the preset product main body.
[0074] Among them, for the reference product image, the reference product can correspond to different display methods. For example, the reference product can be directly displayed as the main subject in the reference product image; or, the reference product can be displayed as the main subject through a human object. For example, clothing products can be displayed through a model to ensure the display effect of the clothing products. Therefore, for the reference product image, the reference product image can be divided into different types according to whether there is a human object in the reference product image. For example, when there is a human object in the reference product image, the reference product image is determined as the first type of image; when there is no human object in the reference product image, the reference product image is determined as the second type of image. There are obvious differences in content, main subject, etc. between the above-mentioned first type of image and the second type of image. For example, there are obvious differences between the human features (such as facial features, body postures, etc.) included in the first type of image and the non-human features (such as clothing detail features, accessory features, etc.) included in the second type of image.
[0075] Since different types of reference product images can correspond to different image features, when determining the target product image, different processing strategies need to be adopted to analyze and process different types of reference product images. Specifically, in order to accurately generate at least one target product image corresponding to the preset product main subject, after obtaining the reference product image, the reference product image can be analyzed and processed to identify whether there is a human object in the reference product image. Specifically, an object recognition operation can be used to identify whether there is a human object in the reference product image.
[0076] In the case where the recognition result is that there is a human object in the reference product image, the reference product image includes both human object features and product features of the reference product. At this time, the product image may or may not include a human object. To prevent the large language model from learning the human object features in the reference product image and applying the human object features to generate at least one target product image corresponding to the preset product, which may lead to distortion of the human object in the target product image. When it is determined that there is a human object in the reference product image, the image features irrelevant to the human object in the reference product image can be determined. The image features can include at least one of the following: color feature, texture feature, shape feature, background feature, and the image positions of each image element in the reference product image.
[0077] After obtaining the image features irrelevant to the human object, the large language model and the image features irrelevant to the human object can be used to analyze and process the product image and the reference product image. That is, the large language model can directly learn the image features irrelevant to the human object in the reference product image, and perform an image generation operation based on the learned image features irrelevant to the human object and the product image, so as to generate at least one target product image corresponding to the preset product main body. This effectively ensures the quality and effect of the target product image generation.
[0078] In the case where the recognition result is that the reference product image does not include a human object, the large language model can be directly used to analyze and process the reference product image so that the large language model learns the image features of the reference product image. Then, the large language model can perform an image generation operation based on the learned image features of the reference product image and the preset product in the product image to obtain a target product image corresponding to the preset product. This also ensures the accuracy and reliability of the target product image generation.
[0079] Further, after generating at least one target product image corresponding to the preset product main body, a download operation can be performed on the target product image. At this time, the method in this embodiment may further include: displaying at least one target product image, where the target product image corresponds to an operation control for implementing image download; and in response to an operation input by the user for the operation control corresponding to the target product image, downloading the selected target product image to a preset area.
[0080] After generating at least one target product image corresponding to the preset product main body, the at least one target product image can be displayed. In order to enable the user to perform a download operation on any one of the target product images, the displayed target product image may correspond to an operation control for implementing the image download operation. The user can selectively input a selection operation for the operation control according to the need. For example, when the user needs to perform a download operation on multiple target product images, the user can input a selection operation for the operation controls corresponding to the multiple target product images, and then can perform a download operation on the selected target product images. Specifically, the selected target product images can be downloaded to a preset area to facilitate the user to view, apply, or perform secondary editing, etc. on the downloaded target product images. When the user does not need to perform a download operation on the target product image, that is, there is no need to input an execution operation for the operation control corresponding to the target product image. This effectively realizes the selective download operation of the target product image and further improves the flexibility and reliability of using the target product image.
[0081] The image generation method provided in this embodiment, in the artificial intelligence AI interaction interface based on a large language model, obtains a product image corresponding to a preset product, determines a reference product image corresponding to a reference product, and then uses the large language model to process the product image and the reference product image to generate at least one target product image corresponding to the preset product. This effectively implements the technical solution of using the "reference product image" and the "product image" to generate the target product image, that is, it realizes the automatic generation solution of the target product image. Thus, it not only effectively improves the efficiency of the image generation operation, reduces the time cost and labor cost required for the image generation operation, but also, the generated target product image is relatively similar to the reference product image and has certain differences. In this way, it can not only make the generated target product image meet the image generation effect of the merchant, but also ensure that the generated target product image has a certain uniqueness, further improving the practicality of this method.
[0082] Figure 3 It is a schematic flowchart of obtaining a product image corresponding to a preset product in an artificial intelligence AI interaction interface based on a large language model provided by an exemplary embodiment of this application; Figure 4 It is a schematic interface diagram of obtaining a product image corresponding to a preset product provided by an exemplary embodiment of this application; On the basis of the above embodiment, refer to the attached Figures 3 - 4 As shown, for the product image, it can not only obtain the product image through the placement effect of the preset product, but also obtain the product image through an image generation request input by the merchant user. At this time, in the artificial intelligence AI interaction interface based on a large language model, obtaining a product image corresponding to a preset product may include:
[0083] Step S301: In response to an image generation request input by a user in the AI interaction interface based on a large language model, display at least one recommended product information to be processed.
[0084] In order to obtain a product image corresponding to a preset product through human-computer interaction operations, the user can input an image generation request in the AI interaction interface based on a large language model. This image generation request is used to indicate that the merchant user has an image generation requirement. For example: The user can input an image generation request in the AI interaction interface, and this image generation request can be specifically implemented as "create a similar style". Then, after the image generation request input in the AI interaction interface based on a large language model, at least one recommended product information to be processed can be obtained. Refer to the attached Figure 4As shown, multiple recommended product lists can be displayed in the AI interaction interface. Each recommended product list can include multiple recommended products. For example, a recommended product list can include: a recommended product with ID xxxxxxxxxxxx1, a recommended product with ID xxxxxxxxxxxx2, a recommended product with ID xxxxxxxxxxxx3, and so on. The above-mentioned recommended products can be products included in the merchant stores determined by the large language model.
[0085] In some instances, at least one recommended product information can be determined by analyzing and comparing the advertising investment costs of each product in the merchant store. At this time, displaying at least one recommended product information to be processed can include: based on the image generation request, determining the corresponding merchant store, where the merchant store includes multiple products; determining the advertising investment costs corresponding to each product in the merchant store; and determining at least one recommended product information to be processed based on the advertising investment costs corresponding to each product.
[0086] Among them, determining at least one recommended product information to be processed based on the advertising investment costs corresponding to each product can include: sorting each product based on the level of the advertising investment costs corresponding to each product to obtain product sorting information; determining the number of products used to limit the recommended products, where the number of products can be one or more; and determining at least one recommended product that meets the number of products based on the product sorting information. The recommended product can be a product with a higher advertising investment cost.
[0087] Then, at least one recommended product information to be processed can be obtained based on at least one recommended product, and at least one recommended product information can be displayed in the AI interaction interface based on the large language model. For the obtained at least one recommended product information, at least one recommended product information can be displayed in the form of an advertising investment cost list, which effectively ensures the accuracy and reliability of obtaining and displaying at least one recommended product information.
[0088] In other instances, at least one recommended product information can be determined not only by analyzing and comparing the advertising investment costs of each product in the merchant store, but also by determining the estimated investment effects of each product in the merchant store. At this time, displaying at least one recommended product information to be processed can include: based on the image generation request, obtaining the corresponding merchant store, where the merchant store includes multiple products; determining the estimated investment effects corresponding to each product in the merchant store; and determining at least one recommended product information to be processed based on the estimated investment effects corresponding to each product.
[0089] Among them, determining at least one piece of recommended product information to be processed based on the estimated placement effects corresponding to each product may include: sorting each product according to the high or low of the estimated placement effects corresponding to each product to obtain product sorting information; determining the number of products used to limit the recommended products, where the number of products can be one or more; based on the product sorting information, determining at least one recommended product that meets the number of products, and the recommended product can be a product with a relatively high estimated placement effect.
[0090] Then, at least one piece of recommended product information to be processed can be obtained based on at least one recommended product, and at least one piece of recommended product information can be displayed in the AI interaction interface based on the large language model. For the obtained at least one piece of recommended product information, the at least one piece of recommended product information can be displayed in the form of an estimated placement effect list. The above-mentioned estimated placement effect list can include at least one of the following: click-through rate index list, click volume index list, conversion rate index list. In this way, it is effectively realized that at least one piece of recommended product information can be flexibly obtained and displayed based on different estimated placement effect lists.
[0091] Step S302: In response to the selection operation input by the user for the recommended product information, determine the information of the preset product, where the information of the preset product includes at least one product release image.
[0092] After displaying at least one piece of recommended product information to be processed, in order to implement the image generation operation, the user can input a selection operation for the recommended product information (which can be any one piece of recommended product information or any number of pieces of recommended product information), and then determine the preset product based on the selection operation input by the user for the recommended product information, and obtain the information of the preset product. The information of the preset product can include at least one of the following: product name, product description information, product specification parameters, at least one product release image, etc. In this way, it is realized that the information of the preset product can be stably determined through the user's product selection operation.
[0093] Step S303: Use the large language model and the information of the preset product to process at least one product release image of the preset product to obtain a product image corresponding to the preset product.
[0094] After obtaining the information of the preset product, the large language model and the information of the preset product can be used to analyze and process at least one product release image of the preset product, so as to obtain a product image corresponding to the preset product. In some instances, any one of at least one product release image of the preset product can be directly determined as the product image corresponding to the preset product. For example, the front view release image of the preset product can be determined as the product image corresponding to the preset product.
[0095] In some other examples, not only can the product image corresponding to the preset product be determined based on the front view release image included in at least one product release image of the preset product, but also the product image can be obtained by performing a matting operation on the preset product body. At this time, using the large language model and the information of the preset product to process at least one product release image of the preset product to obtain the product image corresponding to the preset product may include: using the large language model to perform a matting operation on the preset product body for at least one product release image to obtain at least one matted image corresponding to the preset product body; determining the product image corresponding to the preset product based on the at least one matted image.
[0096] Among them, for the product image, it can not only include the body of the preset product, but also include specific environmental information, product title information, product promotion copywriting, etc. To avoid the influence of the body information of other non-preset products on the image generation operation. After obtaining the information of the preset product, the large language model can be used to perform a matting operation on the preset product body for at least one product release image included in the information of the preset product, so as to obtain at least one matted image corresponding to the preset product body, and the matted image can only include the preset product body.
[0097] After obtaining at least one matted image, the at least one matted image can be analyzed and processed to determine the product image corresponding to the preset product. In some examples, any one of the at least one matted images is determined as the product image corresponding to the preset product, or the matted image corresponding to the main body release image is determined as the product image corresponding to the preset product.
[0098] In some other examples, after obtaining at least one matted image, the product image can also be determined based on a manual selection operation. At this time, determining the product image corresponding to the preset product based on the at least one matted image may include: in response to a selection operation input by the user for the matted image, determining an intermediate image; determining the product image corresponding to the preset product based on the intermediate image.
[0099] After obtaining at least one matted image, the at least one matted image can be displayed in the AI interaction interface of the large language model. The user can input a selection operation for the matted image, and then the intermediate image can be determined based on the selection operation input by the user for the matted image. The intermediate image is the matted image determined by the manual selection operation, and then the intermediate image can be analyzed and processed to determine the product image corresponding to the preset product.
[0100] In some instances, for the intermediate image and the product image, when the effect of the intermediate image meets the user's requirements, the intermediate image can be directly determined as the product image. Alternatively, when the effect of the intermediate image does not meet the user's requirements, a secondary editing operation can be performed on the intermediate image to determine the product image. At this time, based on the intermediate image, determining the product image corresponding to the preset product may include: displaying an operation interface for editing the intermediate image in response to a selection operation input by the user for the preset editing control corresponding to the intermediate image; and determining the product image corresponding to the preset product in response to the editing operation input by the user in the operation interface.
[0101] Among them, after obtaining the intermediate image, it can be recognized whether the intermediate image meets the user's requirements. When the intermediate image does not meet the user's requirements, the intermediate image can be displayed in the AI interaction interface of the large language model, and the preset editing control corresponding to the intermediate image can be displayed. The preset editing control is used to implement the preset editing operation for the intermediate image; then the user can input a selection operation for the preset editing control corresponding to the intermediate image. In response to the selection operation input by the user for the preset editing control corresponding to the intermediate image, an operation interface for editing the intermediate image can be displayed in the AI interaction interface of the large language model. Then the user can input an editing operation for the intermediate image in the operation interface. The editing operation may include at least one of the following: an operation for adjusting the image brightness, an operation for adjusting the image contrast, an operation for adjusting the image saturation, an operation for cropping the image, an operation for rotating the image, an operation for removing or adding elements in the image, etc. After obtaining the editing operation input by the user in the operation interface, the product image corresponding to the preset product after the editing operation can be determined, which effectively ensures that the determined product image can meet the user's requirements and further improves the quality and effect of the image generation operation based on the product image.
[0102] In this embodiment, in response to an image generation request input by the user in the AI interaction interface based on the large language model, at least one recommended product information to be processed is displayed, the information of the preset product is determined, and then at least one product release image of the preset product is processed by using the large language model and the information of the preset product to obtain the product image corresponding to the preset product, which effectively ensures that the obtained product image can meet the user's image generation requirements and further improves the accuracy and reliability of obtaining the product image corresponding to the preset product.
[0103] Figure 5 It is a schematic flowchart of determining the reference product image corresponding to the reference product provided by an exemplary embodiment of the present application; on the basis of the above embodiment, refer to the appendix Figure 5As shown, for the reference product image, it can be determined not only through the user's upload operation and the images corresponding to the reference products with better placement effects determined by the large language model, but also based on virtual products to determine the reference product image corresponding to the reference product. At this time, the method for determining the reference product image corresponding to the reference product in this embodiment may include:
[0104] Step S501: Display a plurality of recommended reference images in the artificial intelligence AI interaction interface based on the large language model, where the recommended reference images are generated based on virtual products.
[0105] To ensure that the generated reference product images comply with the preset legal usage rules, a plurality of recommended reference images can be generated based on virtual products. At this time, before displaying a plurality of recommended reference images in the artificial intelligence AI interaction interface based on the large language model, the method in this embodiment may further include: using the large language model to determine at least one virtual product and at least one recommended product that matches the reference product; determining the product release images corresponding to each of the at least one recommended product; using the large language model to process the product release images and the at least one virtual product to obtain the recommended reference images corresponding to each of the at least one virtual product, where the virtual product is different from the recommended product.
[0106] Specifically, in order to accurately obtain the recommended reference images, the large language model can be used to determine at least one virtual product and at least one recommended product that matches the reference product. Among them, the virtual product can be generated by the large language model based on any virtual information, and at least one recommended product that matches the reference product can be determined by at least one of the click-through rate, click volume, and conversion rate of the product. In some instances, the recommended product can at least meet at least one of the following: the similarity between the recommended product and the leaf category of the preset product is greater than or equal to the preset similarity threshold; the difference between the price range of the recommended product and the price range of the preset product is less than or equal to the preset difference threshold; the similarity between the user group portrait targeted by the recommended product and the user group portrait targeted by the preset product is greater than or equal to the preset similarity threshold. After obtaining a plurality of recommended reference images, a plurality of recommended reference images can be displayed in the AI interaction interface of the large language model, where any one of the recommended reference images can be generated based on a virtual product.
[0107] Step S502: Respond to the selection operation input by the user for the recommended reference image, and determine the reference product image corresponding to the reference product.
[0108] After displaying multiple recommended reference images in the AI interaction interface based on the large language model, the user can input a selection operation for one of the recommended reference images, and then determine the reference product image corresponding to the reference product in response to the selection operation input by the user for the recommended reference image, which effectively ensures the accuracy and reliability of determining the reference product image.
[0109] In this embodiment, multiple recommended reference images are displayed in the artificial intelligence (AI) interaction interface based on the large language model, and the reference product image corresponding to the reference product is determined in response to the selection operation input by the user for the recommended reference image, which effectively ensures the accuracy and reliability of determining the reference product image.
[0110] Figure 6 It is a schematic flowchart of another image generation method provided for an exemplary embodiment of the present application; based on any of the above embodiments, refer to the appendix Figure 6 As shown, after generating at least one target product image corresponding to a preset product, a product release operation can be performed on the target product image. At this time, the method in this embodiment may further include:
[0111] Step S601: In the artificial intelligence (AI) interaction interface based on the large language model, display a link association control for applying at least one target product image.
[0112] After generating at least one target product image corresponding to a preset product, a link association control for applying at least one target product image can be displayed in the AI interaction interface based on the large language model, where the link association control is used to identify information related to the product release link of at least one preset product.
[0113] Step S602: In response to the operation input by the user for the link association control, determine the product release link associated with the preset product.
[0114] When the user wants to apply the target product image to the product release link of at least one preset product, the user can input a selection operation for the link association control, and then determine the product release link associated with the preset product based on the operation input by the user for the link association control.
[0115] Step S603: Associate at least one target product image with the product release link and release at least one target product image.
[0116] After obtaining a product release link associated with a preset product, at least one target product image can be associated with the product release link, and a release operation can be performed on at least one target product image. In this way, it effectively realizes the application of the obtained target product image to the scenario of product release, which is conducive to ensuring the promotion quality and effect of the product.
[0117] In some other examples, after releasing at least one target product image, different target product images can be displayed for different users. At this time, the method in this embodiment may further include: obtaining the preference information of potential users targeted by the preset product; among at least one target product image, determining a product image to be displayed that matches the preference information, where the product image to be displayed is different from the target product image released by the preset product; and displaying the product image to be displayed to potential users.
[0118] Among them, after releasing the preset product based on at least one target product image, the potential users targeted by the preset product may have different effect preferences. For example: potential user A likes the target product image with a blue color tone, potential user B likes the target product image with a red color tone, etc. At this time, in order to improve the display effect of the preset product and display different target product images for different potential users, the preference information of potential users targeted by the preset product can be obtained first. The preference information may include at least one of the following: image background preference, image color tone preference, image layout preference, etc. The above preference information can be determined by analyzing and processing the user portraits of potential users.
[0119] After obtaining the preference information of potential users targeted by the preset product, at least one target product image and the preference information can be analyzed and processed to determine a product image to be displayed that matches the preference information among at least one target product image. The product image to be displayed may be the same as or different from the target product image released by the preset product. Then, the determined product image to be displayed that matches the preference information can be displayed to potential users. The product images to be displayed determined by different preference information may be different. In this way, for the same preset product, different product images to be displayed can be displayed for different potential users, thus realizing the image display effect of "one product for one person", and then the preset product can be displayed based on the image display preferences of different potential users, which is further conducive to increasing the success rate of potential users purchasing the preset product and the product conversion rate.
[0120] In some other examples, after publishing at least one target product image, an operation of tracking the effectiveness index data for the target product image can be performed, and based on the effectiveness index data obtained from the tracking operation, an optimization and update operation can be carried out on the large language model. At this time, the method in this embodiment may further include: obtaining the effectiveness index data of each target product image, where the effectiveness index data includes at least one of the following: click-through rate, number of clicks, conversion rate; and optimizing and updating the large language model based on the effectiveness index data.
[0121] After publishing at least one target product image, an operation of tracking the effectiveness can be performed on the target product image. Specifically, the effectiveness index data of each target product image can be obtained first. Among them, the effectiveness index data can include at least one of the following: click-through rate, number of clicks, conversion rate. The above-mentioned effectiveness index data can be obtained through statistical analysis processing operations on the publishing data in the e-commerce platform. After obtaining the effectiveness index data, an optimization and update operation can be carried out on the large language model based on the effectiveness index data, so as to obtain an optimized large language model, and then an image generation operation can be carried out based on the optimized large language model, which is beneficial to improving the quality and effect of the image generation operation.
[0122] In this embodiment, by displaying a link association control for applying at least one target product image in the artificial intelligence AI interaction interface based on the large language model, in response to the operation input by the user for the link association control, a product publishing link associated with the preset product is determined, and then at least one target product image is associated with the product publishing link and published, thus effectively realizing that the target product image can be applied to the product publishing link of the preset product, further improving the practicability of the method.
[0123] Figure 7 It is a schematic flowchart of using a large language model to process product images and reference product images to generate at least one target product image corresponding to a preset product provided by an exemplary embodiment of the present application; based on any of the above embodiments, refer to the appendix Figure 7 As shown, when the reference product image includes reference copy information; the target product image can be generated in combination with the display characteristics of the reference copy information included in the reference product image. Among them, the target product image can include product copy corresponding to the preset product, or the target product image does not include product copy corresponding to the preset product; at this time, using the large language model to process product images and reference product images to generate at least one target product image corresponding to a preset product in this embodiment may include:
[0124] Step S701: Obtain product information corresponding to the preset product.
[0125] When including reference copy information in the reference product image, when performing an image generation operation based on the reference product image and the product image, the generated target product image can include not only the main body of the preset product, but also the product copy corresponding to the preset product. In order to accurately generate the product copy that matches the preset product, the product information corresponding to the preset product can be obtained first. The product information can include at least one of the following: product name, product category, product description information, and so on.
[0126] Step S702: Use a large language model to process the product image, the reference product image, and the product information to generate at least one target product image corresponding to the preset product. The generated target product image includes the product copy of the preset product, and the gap between the position of the product copy in the target product image and the position of the reference copy information in the reference product image is less than or equal to a preset threshold.
[0127] After obtaining the product information corresponding to the preset product, a large language model can be used to analyze and process the product image, the reference product image, and the product information. That is, the product image, the reference product image, and the product information can be input into the large language model to obtain at least one target product image corresponding to the preset product output by the large language model. The generated target product image can include the product copy of the preset product and the main body of the preset product. Moreover, the display effect characteristics of the product copy in the target product image are similar to the display effect characteristics of the reference copy information in the reference product image. For example, the gap between the position of the product copy in the target product image and the position of the reference copy information in the reference product image is less than or equal to a preset threshold, or the gap between the font size of the product copy in the target product image and the font size of the reference copy information in the reference product image is less than or equal to a preset threshold, further ensuring the flexible reliability of generating at least one target product image.
[0128] In some other examples, after generating at least one target product image corresponding to the preset product, the user can perform a secondary editing operation on the at least one target product image according to the needs. At this time, the method in this embodiment can further include: displaying the at least one target product image in the artificial intelligence AI interaction interface based on the large language model, where the AI interaction interface includes adjustment controls for adjusting the target product image; in response to the operation input by the user for the adjustment control, displaying an editing configuration interface for adjusting the target product image; in response to the operation input by the user in the editing configuration interface, adjusting the product copy and / or the main body of the preset product in the target product image to obtain an adjusted target product image.
[0129] After obtaining the target product image including the product copywriting, it can be identified whether the obtained target product image can identify the user's image generation requirement. When the target product image meets the user's image generation requirement, the generated target product image can be directly applied or downloaded; when the target product image does not meet the user's image generation requirement, the user can perform a secondary editing operation on the generated target product image. Specifically, at least one target product image can be displayed in the artificial intelligence AI interaction interface based on the large language model. Among them, the AI interaction interface can include adjustment controls for adjusting the target product image. After displaying the adjustment controls for adjusting the target product image in the AI interaction interface, when the user needs to adjust the target product image, the user can input a selection operation for the adjustment controls, and then an editing configuration interface for adjusting the target product image can be displayed based on the operation input by the user for the adjustment controls. Among them, the editing configuration interface can include multiple editing controls for implementing image editing operations, and the editing controls can include at least one of the following: element editing controls for adjusting the position, size, and color of each display element in the image, brightness editing controls for adjusting the brightness of the image, contrast editing controls for adjusting the contrast of the image, saturation editing controls for adjusting the saturation of the image, size editing controls for cropping the size of the image, and so on.
[0130] After displaying the editing configuration interface for adjusting the target product image, the user can input corresponding editing operations in the editing configuration interface, and in response to the operations input by the user in the editing configuration interface, the product copywriting and / or the preset product main body in the target product image can be adjusted. Among them, the font, font size, display position, display effect, etc. of the product copywriting in the target product image are adjusted, and the color, display position, display effect, display size, etc. of the preset product main body are adjusted, so as to obtain the adjusted target product image and ensure that the adjusted target product image can meet the user's image display requirement.
[0131] In this embodiment, by obtaining the product information corresponding to the preset product, and then using the large language model to process the product image, the reference product image, and the product information, at least one target product image corresponding to the preset product can be stably generated, and the user can selectively perform a secondary editing operation on the generated at least one target product image, so as to ensure that the obtained target product image can meet the user's image generation requirement, and further improve the practicability of this method.
[0132] Specifically in application, refer to the appendix Figure 8As shown in the figure, an embodiment of the present application provides an image generation method implemented based on AI technology. The execution subject of the image generation method is an image generation device based on AI technology. A large language model for implementing image generation operations is deployed in the image generation device. Based on the large language model, data analysis operations, opportunity insight operations, creative production operations, business decision-making operations, etc. can be performed, which can help merchant users improve the release efficiency and delivery effect of goods. Specifically, the image generation method may include the following steps:
[0133] Step 1: Obtain high-quality creative images.
[0134] Among them, high-quality creative images (corresponding to the reference product images in the above embodiments) can be obtained through the user's image upload operation. At this time, obtaining high-quality creative images may include: in the AI interaction interface based on the large language model, an image upload interface can be displayed; obtain the image upload operation input by the user in the image upload; obtain the high-quality creative image uploaded by the user based on the image upload operation.
[0135] After obtaining the high-quality creative image uploaded by the user, an availability recognition operation can be performed on the high-quality creative image. Specifically, performing an availability recognition operation on the high-quality creative image may include: obtaining a detection rule for implementing the product availability recognition operation; based on the detection rule, performing a rationality and legality detection operation on the high-quality creative image, and detecting whether the number of main bodies included in the high-quality creative image is consistent with the number of preset product main bodies included in the product image; if the high-quality creative image meets the rationality and legality detection rules, and the number of main bodies included in the detected high-quality creative image is consistent with the number of preset product main bodies included in the product image, it is determined that the high-quality creative image passes the availability recognition operation. At this time, an image generation operation based on the high-quality creative image is allowed; otherwise, it is determined that the high-quality creative image fails the availability recognition operation. At this time, an image generation operation based on the high-quality creative image is prohibited, and a prompt message is fed back to the user.
[0136] In some other instances, high-quality creative images can be obtained from the recommended images generated by a large language model based on virtual goods. At this time, obtaining high-quality creative images may include: using the large language model to obtain a high-quality material library, where the high-quality material library includes high-quality material images corresponding to multiple high-quality goods. The above-mentioned high-quality goods can be determined by screening through a preset product screening rule. In some instances, the product screening rule may include at least one of the following: the high-quality goods correspond to recommended goods that belong to the same leaf category as the goods in the merchant's store; the price range of the high-quality goods is similar to the price range of the goods in the merchant's store; the brand / target audience of the high-quality goods is similar to the brand / target audience of the goods in the merchant's store, and so on. Then, the large language model can generate a recommended material list for the virtual goods based on the high-quality material images included in the high-quality material library. Among them, the recommended material list may include multiple recommended material images corresponding to the virtual goods, and display the recommended material list, obtain the selection operation input by the user for any one of the recommended material images in the recommended material list, and determine the high-quality creative image based on the selection operation, which effectively ensures the accuracy and reliability of obtaining the high-quality creative image.
[0137] Step 2: Obtain the product image of the preset product that needs to perform an image generation operation (or referred to as "image cloning operation").
[0138] Among them, the product image of the preset product that needs to perform an image generation operation can be obtained by the user uploading an image. At this time, obtaining the product image of the preset product that needs to perform an image generation operation may include: displaying an upload configuration interface; obtaining the image upload operation input by the user in the upload configuration interface; the user can upload an existing product image in the merchant's store, and then the product image of the preset product can be obtained based on the existing product image, which effectively ensures the accuracy and reliability of obtaining the product image.
[0139] In some other instances, the product image of the preset product that needs to perform an image generation operation can be obtained through the user's selection operation. At this time, obtaining the product image of the preset product that needs to perform an image generation operation may include: using the large language model to obtain multiple recommended product images that need to perform an image generation operation, where any recommended product image corresponds to a recommended product that is a published product with a relatively high advertising consumption or advertising investment cost in the merchant's store; obtaining the selection operation input by the user for any recommended product image; based on the selection operation input for any recommended product image, select the product image that needs an image generation operation (or creative operation), which effectively ensures the accuracy and reliability of obtaining the product image.
[0140] Step 3: Perform a matte extraction operation on the product image for the preset product body to obtain the image after matte extraction.
[0141] Among them, the number of obtained product images is one or more. When the number of product images is multiple, matte extraction operations for the preset product main body can be performed on each product image, so that the images after matte extraction can be obtained. After the images after matte extraction are obtained, it can be first determined whether the images after matte extraction meet the preset product matte extraction effect. When the obtained images after matte extraction meet the preset product matte extraction effect, image generation operations can be performed based on the images after matte extraction; when the obtained images after matte extraction do not meet the preset product matte extraction effect, the user can perform secondary editing operations on the images after matte extraction.
[0142] The secondary editing operations on the images after matte extraction can include: in response to the triggering operation of the user for the secondary editing operation, displaying a secondary editing interface for the images after matte extraction; obtaining the editing configuration operations input by the user for the images after matte extraction in the secondary editing interface; and performing corresponding editing configuration operations on the images after matte extraction based on the editing configuration operations to obtain the configured images after matte extraction, which effectively ensures that the obtained images after matte extraction can meet the user's product matte extraction effect.
[0143] Step 4: Use a large language model to perform image generation operations on the high-quality creative images and the images after matte extraction to obtain at least one creative image corresponding to the preset product main body (corresponding to the target product image in the above-mentioned embodiment), and the similarity between any creative image and the product image is greater than or equal to the preset threshold, and there are differences between any creative image and the product image.
[0144] Among them, the number of creative images can be one or more, and one or more creative images can be generated by the large language model through the analysis and processing of the high-quality creative images and the images after matte extraction. When the high-quality creative images include creative texts, the generated creative images can include product copy corresponding to the preset product main body, and the display characteristics of the product copy in the creative images are similar to the display characteristics of the creative texts in the high-quality creative images; when the high-quality creative images do not include creative texts, the generated creative images can include product copy corresponding to the preset product main body, or the generated creative images may not include product copy corresponding to the preset product main body.
[0145] Step 5: When the effect of the obtained creative images does not meet the user's requirements, secondary editing operations can be performed on the creative images to obtain the edited creative images.
[0146] After obtaining the creative image, it is possible to identify whether the creative image meets the user's image generation requirements. When the creative image meets the user's image generation requirements, no adjustment or editing operation needs to be performed on the generated creative image. When the creative image does not meet the user's image generation requirements, the user can trigger the secondary editing operation of the creative image. Specifically, a secondary editing interface for the creative image can be displayed; the editing configuration operations input by the user in the secondary editing interface can be obtained; and the secondary editing operation can be performed on the creative image based on the editing configuration operations. In some instances, performing the secondary editing operation on the creative image may include at least one of the following: performing a secondary editing operation on the display features of the preset commodity main body, performing a secondary editing operation on the display features of the commodity copywriting, etc., so as to obtain an edited creative image that meets the image generation requirements.
[0147] Step 6: Download / apply the edited creative image to the commodity / preview operation.
[0148] After obtaining the edited creative image, a flexible download operation / preview operation can be performed on the edited creative image according to the application requirements, which is convenient for the user to call, view or perform a secondary editing operation on the edited creative image; or, the edited creative image can be applied to the application scenario of commodity release according to the application requirements. Specifically, applying the edited creative image to the application scenario of commodity release may include: obtaining the commodity release link of the preset commodity; associating the edited creative image with the commodity release link of the preset commodity, and performing the commodity release operation of the preset commodity based on the edited creative image, thereby effectively realizing that the edited creative image can be applied to the application scenario of commodity release.
[0149] Further, after applying the edited creative image, a post-link effect collection operation can be performed on the applied creative image. For example, the effect indicators of the applied creative image can be detected, and the effect indicators may include at least one of the following: click-through rate, click volume or conversion rate. Then, the application effect corresponding to the creative image can be determined based on the detected click-through rate, click volume or conversion rate. Specifically, when the click-through rate, click volume or conversion rate meets the preset threshold, it is determined that the application effect corresponding to the creative image is good; when the click-through rate, click volume or conversion rate does not meet the preset threshold, it is determined that the application effect corresponding to the creative image is poor.
[0150] In some other instances, after performing the post-link effect collection operation on the applied creative image, an optimization operation can be performed on the large language model based on the collected effect indicators to obtain an optimized large language model, which is conducive to improving the quality and effect of the data processing operation of the large language model.
[0151] It should be noted that the large language model has the following capabilities: (1) Accurate generation and optimization of images: After obtaining high-quality creative images (e.g., competitive product images) uploaded by merchant users that they want to replicate, the large language model can automatically identify and extract high-quality elements (such as background, model posture, product placement, etc.), and generate and optimize images while maintaining the original creative style, making them more in line with their own brand characteristics and market demand. (2) Automated processing: The large language model can be implemented as a one-click operation tool, which can include: automatic image clipping, background replacement, copywriting, etc., greatly improving work efficiency. For example, merchants only need to enter a natural language description, such as "Please change the background of this competitive product image to my store style and adjust the copy position". After the above natural language description is obtained by the AI interactive interface based on the large language model, the large language model can automatically and instantly complete the image generation task. (3) “Differentiation” image generation operation: The large language model uses high-quality creative images as reference standards for image generation operations, but does not completely copy them, so that there is differentiation between the generated creative images and the high-quality creative images used as reference standards. Specifically, the background / color tone / layout is expected to be differentiated from the high-quality creative images so as to enhance the visual distinction in the search results page. In addition, merchants can perform secondary editing operations on the generated creative images based on the actual delivery effect, such as adjusting the generation parameters of the creative images, etc., so that the obtained creative images can meet the personalized needs of different users.
[0152] The technical solution provided by this application embodiment can generate the same picture based on the product image and the high-quality creative picture according to the high-quality creative picture that the merchant wants to replicate and the product image of the store's products that the merchant wants to apply. This capability can save the merchant's art costs and reduce the time required for the traditional design process; in addition, the above-mentioned image generation operation meets the AI demands of the advertising merchant in the early stage of a large number of creative material reserves, map candidate sets and other links, and can perform delivery and creative effect data (click-through rate, click volume, conversion rate) recovery operations for the generated creative pictures, and provide the delivery and creative effect data to the image generation device in a benign manner, so as to optimize and update the large language model with the effect data as a guide, which is conducive to improving the image generation quality and effect of the large language model. Specifically, this solution has the following advantages:
[0153] (1) The large language model has the ability to produce the same type of image. The ability to produce the same type of image can integrate operations such as advertising material image production, delivery, and effect analysis. It can also be linked to the large language model based on AI interactive operations to perform image generation operations. Through open dialogue, it can obtain more artificial intelligence AI capabilities, thereby better serving merchant users.
[0154] (2) The image generation technology of the large language model can also be integrated with other image processing capabilities. For example, it can integrate functions such as image matting, cropping, background layout, background replication, copywriting replication, and copywriting editing. Among them, the copywriting generation ability can directly understand the product based on the product selected by the customer, and combine the copywriting layout, font size, and copywriting style of the creative text in the high-quality creative image to perform AI writing, and can also perform secondary editing;
[0155] (3) Natural language dialogue operations can be performed in the AI interaction interface based on the large language model. Specifically, in the AI interaction interface, through natural language dialogue operations, merchants can express image generation requirements more naturally. The large language model can understand and execute complex creative instructions, providing a more user-friendly service. The advantage of the dialogue-based product is that it is detached from the fixed input content of the traditional GUI, which can help the technical team collect customer user feedback and innovative production requirements, and continuously optimize the AI large language model to ensure that the generated creative images meet the latest market trends and user preferences. For example, merchants only need to select the goods to be listed, and there is no need for merchants to re-upload white background images, etc. The large language model can automatically help merchants perform automatic image matting operations and replica operations of the same model, which can greatly reduce the operation costs of merchants and enhance the feedback service effect of interaction and experience at the same time;
[0156] (4) Creative generation operations oriented by effect indicators: In addition to the same model replication ability, after obtaining the creative image, the large language model can use the effect indicators of the creative image (such as click-through rate, click volume, or conversion rate, etc.) as the optimization orientation of the large language model, that is, the large language model can be optimized and updated based on the effect indicators of the creative image. In this way, when performing image generation operations based on the large language model, it can be determined that the generated creative image and the product copywriting placement in the creative image are available and have good placement effects, further improving the quality and effect of image generation.
[0157] In addition, in some of the processes described in the above embodiments and the accompanying drawings, multiple operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The operation numbers such as 11, 12, etc. are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.
[0158] Figure 9 This is a schematic structural diagram of an image generation device provided by an exemplary embodiment of the present application; refer to the appendixFigure 9 As shown, this embodiment provides an image generation device, which is used to execute Figure 2 the corresponding image generation method. Specifically, the image generation device may include:
[0159] A first acquisition module 11, configured to acquire a product image corresponding to a preset product in an artificial intelligence (AI) interaction interface based on a large language model;
[0160] A first determination module 12, configured to determine a reference product image corresponding to a reference product, where the number of main bodies of the preset product is the same as that of the reference product;
[0161] A first processing module 13, configured to use the large language model to process the product image and the reference product image to generate at least one target product image corresponding to the preset product, where the similarity between the target product image and the reference product image is greater than or equal to a preset threshold, and there are differences between the target product image and the reference product image.
[0162] Regarding the image generation device in this embodiment, it may also execute the description content of the embodiments shown in the foregoing embodiments Figures 1 - 8 For details, reference may be made to the detailed description in the foregoing embodiments, and details will not be elaborated herein.
[0163] Figure 10 It is a schematic structural diagram of a computing platform provided by an exemplary embodiment of the present application; as Figure 10 shown, this embodiment provides a computing platform, which is used to execute the above Figure 2 shown image generation method. In practice, the computing platform may include: a memory 24 and a processor 25.
[0164] The memory 24 is used to store computer programs and may be configured to store various other data to support operations on the computing platform. Examples of these data include instructions for any application or method operating on the computing platform, data structures, contact data, phone book data, messages, pictures, videos, etc.
[0165] The processor 25 is coupled to the memory 24 and is configured to execute the computer program in the memory 24 to: acquire a product image corresponding to a preset product in an artificial intelligence (AI) interaction interface based on a large language model; determine a reference product image corresponding to a reference product, where the number of main bodies of the preset product corresponding to the product image is the same as that of the reference product corresponding to the reference product image; use the large language model to process the product image and the reference product image to generate at least one target product image corresponding to the preset product, where the similarity between the target product image and the reference product image is greater than or equal to a preset threshold, and there are differences between the target product image and the reference product image.
[0166] In some instances, when the processor 25 obtains a product image corresponding to a preset product in an artificial intelligence (AI) interaction interface based on a large language model, the processor 25 is configured to perform: in response to an image generation request input by a user in the AI interaction interface based on the large language model, display at least one recommended product information to be processed; in response to a selection operation input by the user for the recommended product information, determine the information of the preset product, where the information of the preset product includes at least one product release image; use the large language model and the information of the preset product to process at least one product release image of the preset product to obtain a product image corresponding to the preset product.
[0167] In some instances, when the processor 25 uses the large language model and the information of the preset product to process at least one product release image of the preset product to obtain a product image corresponding to the preset product, the processor 25 is configured to perform: use the large language model to perform a matte operation on at least one product release image for the preset product main body to obtain at least one matte image corresponding to the preset product main body; based on the at least one matte image, determine a product image corresponding to the preset product.
[0168] In some instances, when the processor 25 determines a product image corresponding to the preset product based on at least one matte image, the processor 25 is configured to perform: in response to a selection operation input by the user for the matte image, determine an intermediate image; based on the intermediate image, determine a product image corresponding to the preset product.
[0169] In some instances, when the processor 25 determines a product image corresponding to the preset product based on the intermediate image, the processor 25 is configured to perform: in response to a selection operation input by the user for a preset editing control corresponding to the intermediate image, display an operation interface for editing the intermediate image; in response to an editing operation input by the user in the operation interface, determine a product image corresponding to the preset product.
[0170] In some instances, when the processor 25 determines a reference product image corresponding to a reference product, the processor 25 is configured to perform: display an image upload control in the artificial intelligence (AI) interaction interface based on the large language model; in response to an operation input by the user for the image upload control, obtain an uploaded image; determine the uploaded image as the reference product image corresponding to the reference product.
[0171] In some instances, after obtaining the uploaded image, the processor 25 is configured to perform: obtain a preset rule for detecting the uploaded image; use the preset rule to perform compliance detection on the uploaded image; in the case where the uploaded image passes the compliance detection, allow the uploaded image to be determined as the reference product image corresponding to the reference product.
[0172] In some instances, when the processor 25 determines a reference product image corresponding to a reference product, the processor 25 is configured to perform: displaying a plurality of recommended reference images in an artificial intelligence (AI) interaction interface based on a large language model, wherein the recommended reference images are generated based on virtual products; and determining a reference product image corresponding to the reference product in response to a selection operation input by a user for the recommended reference images.
[0173] In some instances, before displaying a plurality of recommended reference images in an artificial intelligence (AI) interaction interface based on a large language model, the processor 25 is configured to perform: determining at least one virtual product and at least one recommended product that matches the reference product by using a large language model; determining a product release image corresponding to each of the at least one recommended product; and processing the product release images and the at least one virtual product by using a large language model to obtain recommended reference images corresponding to the at least one virtual product respectively, wherein the virtual products are different from the recommended products.
[0174] In some instances, the recommended product satisfies at least one of the following: the similarity between the recommended product and the leaf category of a preset product is greater than or equal to a preset similarity threshold; the difference between the price range of the recommended product and the price range of the preset product is less than or equal to a preset difference threshold; the similarity between the user group portrait targeted by the recommended product and the user group portrait targeted by the preset product is greater than or equal to a preset similarity threshold.
[0175] In some instances, after generating at least one target product image corresponding to a preset product, the processor 25 is configured to perform: displaying, in an artificial intelligence (AI) interaction interface based on a large language model, a link association control for applying the at least one target product image; determining a product release link associated with the preset product in response to an operation input by a user for the link association control; associating the at least one target product image with the product release link; and releasing the at least one target product image.
[0176] In some instances, after releasing the at least one target product image, the processor 25 is configured to perform: obtaining preference information of potential users targeted by the preset product; determining, among the at least one target product image, a product image to be displayed that matches the preference information, wherein the product image to be displayed is different from the target product image released for the preset product; and displaying the product image to be displayed to the potential users.
[0177] In some instances, after releasing the at least one target product image, the processor 25 is configured to perform: obtaining effect metric data of each target product image, where the effect metric data includes at least one of the following: click-through rate, number of clicks, conversion rate; and optimizing and updating the large language model based on the effect metric data.
[0178] In some instances, when reference copywriting information is included in the reference product image; when the processor 25 uses a large language model to process the product image and the reference product image to generate at least one target product image corresponding to a preset product, the processor 25 is configured to perform: obtaining product information corresponding to the preset product; using the large language model to process the product image, the reference product image, and the product information to generate at least one target product image corresponding to the preset product, wherein the generated target product image includes the product copywriting of the preset product, and the gap between the position of the product copywriting in the target product image and the position of the reference copywriting information in the reference product image is less than or equal to a preset threshold value.
[0179] In some instances, after generating at least one target product image corresponding to a preset product, the processor 25 is configured to perform: displaying at least one target product image in an artificial intelligence AI interaction interface based on a large language model, wherein the AI interaction interface includes adjustment controls for adjusting the target product image; in response to an operation input by the user for the adjustment control, displaying an editing configuration interface for adjusting the target product image; and in response to an operation input by the user in the editing configuration interface, adjusting the product copywriting and / or the preset product main body in the target product image to obtain an adjusted target product image.
[0180] In some instances, when the processor 25 uses a large language model to process the product image and the reference product image to generate at least one target product image corresponding to a preset product, the processor 25 is configured to perform: identifying whether a human object is included in the reference product image; when a human object is included in the reference product image, determining image features irrelevant to the human object in the reference product image; and using the large language model and the image features to process the product image and the reference product image to generate at least one target product image corresponding to the preset product.
[0181] In some instances, when the processor 25 obtains a product image corresponding to a preset product in an artificial intelligence AI interaction interface based on a large language model, the processor 25 is configured to perform: obtaining placement effect information of the product image corresponding to the preset product based on the large language model; when the placement effect information does not meet the preset requirements, generating an image generation suggestion corresponding to the preset product based on the large language model; and displaying the product image corresponding to the preset product and the image generation suggestion in the AI interaction interface based on the large language model.
[0182] Further, as Figure 10 shown, the computing platform further includes: other components such as a communication component 26, a display 27, a power supply component 28, and an audio component 29. Figure 10Only some components are schematically shown, which does not mean that the computing platform only includes Figure 8 the components shown. Additionally, Figure 10 the components within the wireframe are optional components, rather than mandatory components, and can be determined according to the product form of the working node. The working node of this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, or an IOT device, or can also be a server device such as a conventional server, a cloud server, or a server array. If the working node of this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, or a smart phone, it may include Figure 10 the components within the wireframe; if the working node of this embodiment is implemented as a server device such as a conventional server, a cloud server, or a server array, it may not include Figure 10 the components within the wireframe.
[0183] The above-mentioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0184] The above-mentioned communication component is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on a communication standard, such as a mobile communication network such as 2G, 3G, 4G / LTE, 5G, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel.
[0185] The above-mentioned display includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from users. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation.
[0186] The above-mentioned power supply component provides power for various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.
[0187] The above-mentioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC). When the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in the memory or sent via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.
[0188] The above-mentioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC). When the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in the memory or sent via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.
[0189] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above method embodiments. Among them, the computer-readable storage medium can be implemented by volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices or any other non-transmission medium
[0190] Accordingly, an embodiment of the present application further provides a computer program product, which includes a computer program or instruction, which, when executed by a processor, enables the processor to implement the steps in the above method embodiments. It should be understood that each process or a combination of multiple processes in the above method flow can be implemented by the computer program or instruction. In addition, these computer programs or instructions can be applied to the processors of general-purpose computers, special-purpose computers, embedded processors or other programmable data processing devices, so that the processors of general-purpose computers, special-purpose computers, embedded processors or other programmable data processing devices can be used as devices to implement the corresponding functions in the above method embodiments.
[0191] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity object or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity object or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity object or device including the element.
[0192] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. An image generation method, characterized in that: include: In the artificial intelligence (AI) interactive interface based on a large language model, obtain the product image corresponding to the preset product; Determine a reference product image corresponding to the reference product, wherein the number of entities of the preset product corresponding to the product image is consistent with the number of entities of the reference product corresponding to the reference product image; The product image and the reference product image are processed by using the large language model to generate at least one target product image corresponding to the preset product, wherein the similarity between the target product image and the reference product image is greater than or equal to a preset threshold, and there are differences between the target product image and the reference product image.
2. The method according to claim 1, characterized in that In the artificial intelligence AI interactive interface based on the large language model, obtain the product image corresponding to the preset product, including: In response to an image generation request input by a user in an AI interaction interface based on a large language model, displaying at least one recommended product information to be processed; In response to a selection operation input by a user for the recommended product information, determining information of the preset product, wherein the information of the preset product includes at least one product release picture; At least one product release image of the preset product is processed using the large language model and the information of the preset product to obtain a product image corresponding to the preset product.
3. The method according to claim 2, characterized in that Processing at least one product release image of the preset product by using the large language model and the information of the preset product to obtain a product image corresponding to the preset product, including: Using the large language model, performing a cutout operation of a preset commodity subject on the at least one commodity release image to obtain at least one cutout image corresponding to the preset commodity subject; Based on the at least one cutout image, a product image corresponding to the preset product is determined.
4. The method according to claim 3, characterized in that Determining a product image corresponding to the preset product based on the at least one cutout image includes: In response to a selection operation input by a user with respect to the cutout image, determining an intermediate image; Based on the intermediate image, a product image corresponding to the preset product is determined.
5. The method according to claim 4, characterized in that Determining a product image corresponding to the preset product based on the intermediate image includes: In response to a selection operation input by a user for a preset editing control corresponding to the intermediate image, displaying an operation interface for editing the intermediate image; In response to an editing operation input by a user in the operation interface, a product image corresponding to the preset product is determined.
6. The method according to claim 1, characterized in that Determine the reference product image that corresponds to the reference product, including: Display image upload controls in the AI interactive interface based on a large language model; In response to an operation input by a user to the image upload control, obtaining an uploaded image; The uploaded image is determined as a reference product image corresponding to the reference product.
7. The method according to claim 6, characterized in that After obtaining the uploaded image, the method further includes: Obtaining a preset rule for detecting the uploaded image; Performing compliance detection on the uploaded image using preset rules; In the case where the uploaded image passes the compliance check, the uploaded image is allowed to be determined as a reference product image corresponding to the reference product.
8. The method according to claim 1, characterized in that Determine the reference product image that corresponds to the reference product, including: Displaying a plurality of recommended reference images in an artificial intelligence (AI) interactive interface based on a large language model, wherein the recommended reference images are generated based on virtual commodities; In response to a selection operation input by a user with respect to the recommended reference image, a reference product image corresponding to the reference product is determined.
9. The method according to claim 8, characterized in that Before displaying a plurality of recommended reference images in the artificial intelligence (AI) interaction interface based on the large language model, the method further includes: Determining at least one virtual commodity and at least one recommended commodity matching the reference commodity using the large language model; Determine a product release image corresponding to each of the at least one recommended product; The commodity release graph and the at least one virtual commodity are processed by using the large language model to obtain a recommendation reference graph corresponding to each of the at least one virtual commodity, wherein the virtual commodity is different from the recommended commodity.
10. The method according to claim 9, characterized in that The recommended products meet at least one of the following requirements: The similarity between the recommended product and the leaf category of the preset product is greater than or equal to a preset similarity threshold; The difference between the price range of the recommended product and the price range of the preset product is less than or equal to a preset difference threshold; The similarity between the user group portrait for the recommended product and the user group portrait for the preset product is greater than or equal to a preset similarity threshold.
11. The method according to any one of claims 1 to 10, characterized in that: After generating at least one target product image corresponding to the preset product, the method further includes: In an artificial intelligence (AI) interaction interface based on a large language model, a link association control for applying the at least one target product image is displayed; In response to an operation input by a user on a link association control, determining a product publishing link associated with the preset product; The at least one target product image is associated with the product publishing link, and the at least one target product image is published.
12. The method according to claim 11, characterized in that After publishing the at least one target product image, the method further includes: Obtaining preference information of potential users for preset products; Determining, from the at least one target product image, a product image to be displayed that matches the preference information, wherein the product image to be displayed is different from the target product image published by the preset product; The product image to be displayed is displayed to the potential user.
13. The method according to claim 11, characterized in that After publishing the at least one target product image, the method further includes: Acquire effect indicator data of each target product image, wherein the effect indicator data includes at least one of the following: click-through rate, click volume, and conversion rate; Based on the effect indicator data, the large language model is optimized and updated.
14. The method according to any one of claims 1 to 10, characterized in that: When the reference product image includes reference copy information; using the large language model to process the product image and the reference product image to generate at least one target product image corresponding to the preset product, including: Acquire commodity information corresponding to the preset commodity; The product image, the reference product picture and the product information are processed by using the large language model to generate at least one target product image corresponding to the preset product, wherein the generated target product image includes the product copy of the preset product, and the difference between the position of the product copy in the target product image and the position of the reference copy information in the reference product picture is less than or equal to a preset threshold.
15. The method according to claim 14, characterized in that After generating at least one target product image corresponding to the preset product, the method further includes: Displaying the at least one target product image in an artificial intelligence (AI) interaction interface based on a large language model, wherein the AI interaction interface includes an adjustment control for adjusting the target product image; In response to an operation input by a user on the adjustment control, displaying an editing configuration interface for adjusting the target product image; In response to an operation input by a user in the editing configuration interface, the product copy and / or the preset product body in the target product image is adjusted to obtain an adjusted target product image.
16. The method according to any one of claims 1 to 10, characterized in that: The method of processing the product image and the reference product image by using the large language model to generate at least one target product image corresponding to the preset product includes: Identifying whether the reference product image includes a person object; When the reference product image includes a person object, determining image features in the reference product image that are not related to the person object; The product image and the reference product image are processed using the large language model and the image features to generate at least one target product image corresponding to the preset product.
17. The method according to any one of claims 1 to 10, characterized in that: In the artificial intelligence AI interactive interface based on the large language model, obtain the product image corresponding to the preset product, including: Acquire, based on the large language model, delivery effect information of the product image corresponding to the preset product; When the delivery effect information does not meet the preset requirements, generating an image generation suggestion corresponding to the preset product based on the large language model; In an AI interactive interface based on a large language model, the product image corresponding to the preset product and the image generation suggestion are displayed.
18. A computing platform, characterized in that: include: A memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions, when executed by the processor, implement the method of any one of claims 1 to 17.
19. A computer storage medium, characterized in that: Used to store a computer program, which enables a computer to implement the method of any one of claims 1 to 17 when executed.
20. A computer program product, characterized in that include: A computer-readable storage medium storing computer instructions, which, when executed by one or more processors, causes the one or more processors to execute the steps of the method of any one of claims 1 to 17.