Image editing method, device and equipment and computer storage medium
By using a unified generative understanding model and deep thinking technology, the problem of lagging understanding and generation capabilities of virtual fitting systems when faced with rapidly changing fashion trends and segmented clothing styles has been solved. This enables edited images generated in AI fitting scenarios to accurately display the fitting effect, thus improving the practicality of virtual fitting systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-03-31
AI Technical Summary
Existing virtual fitting systems based on general generative models cannot meet users' expectations for accurate fitting results when faced with rapidly changing fashion trends and segmented clothing styles, and suffer from lagging understanding and generation capabilities.
By employing a unified generative understanding model and deep thinking technology, embedded representation information is generated by acquiring reference images and image editing instructions. The pre-trained generative understanding model is then used for image editing to ensure that the edited image accurately displays the try-on effect.
This enables edited images generated in AI-powered virtual try-on scenarios to clearly and precisely display the try-on effect, meeting users' expectations for accurate try-on and improving the practicality of the virtual try-on system.
Smart Images

Figure CN121767489A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an image editing method, apparatus, device, and computer storage medium. Background Technology
[0002] With the rapid development of science and technology, e-commerce platforms are becoming increasingly widespread. Currently, apparel is a key category for e-commerce platforms to develop and operate. Every day, a large number of users purchase clothing on e-commerce platforms. When a user is interested in a particular item of clothing, even with photos of models wearing it, it is difficult to see how it would look on them. Users often need to place an order and have it delivered to their home before trying it on. If the fit is not satisfactory, there is a certain probability of returning the item.
[0003] Currently, AI-based virtual try-on systems can provide users with a pre-purchase fitting experience. However, existing virtual try-on systems based on general generative models often have significant limitations. Specifically, when faced with rapidly changing fashion trends and specialized clothing style terminology, the understanding and generation capabilities of these large models lag significantly, thus preventing virtual try-on systems from meeting users' expectations for accurate fitting results. Summary of the Invention
[0004] This application provides an image editing method, apparatus, device, and computer storage medium that can guarantee the effect of image editing. When applied to AI virtual try-on scenarios, it can ensure that the generated try-on effect meets the user's expectations for accurate try-on results.
[0005] In a first aspect, embodiments of the present invention provide an image editing method, comprising: Acquire a reference image and corresponding image editing instructions; By understanding the reference image and the image editing instructions, a thought chain information for implementing image editing operations can be obtained; Based on the reference image, image editing instructions, and thought chain information, editing input information for implementing image editing operations is generated, wherein the editing input information is embedded representation information; The edited input information is processed using a pre-trained generative understanding model to obtain the edited image.
[0006] Secondly, embodiments of the present invention provide an image editing apparatus, comprising: The first acquisition module is used to acquire a reference image and image editing instructions corresponding to the reference image; The first processing module is used to understand the reference image and the image editing instructions to obtain the thought chain information for implementing the image editing operation; The first generation module is used to generate editing input information for implementing image editing operations based on the reference image, image editing instructions and the thought chain information, wherein the editing input information is embedded representation information; The first processing module is further configured to process the edit input information using a pre-trained generative understanding model to obtain the edited image.
[0007] Thirdly, embodiments of the present invention provide an electronic device, including: a memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions, when executed by the processor, implement the method in the first aspect described above.
[0008] Fourthly, embodiments of the present invention provide a computer storage medium for storing a computer program, which, when executed by a computer, implements the method described in the first aspect above.
[0009] Fifthly, embodiments of the present invention provide a computer program product, comprising: a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause one or more processors to perform the steps of the method described in the first aspect.
[0010] The image editing method, apparatus, device, and computer storage medium provided in this embodiment acquire a reference image and image editing instructions corresponding to the reference image. They then understand the reference image and the image editing instructions to obtain thought chain information for implementing image editing operations. Based on the reference image, image editing instructions, and the thought chain information, they generate editing input information for implementing image editing operations. The generative understanding model has high understanding and generation capabilities. The pre-trained generative understanding model then processes the editing input information. Because the editing input information incorporates thought chain information, the edited image can be accurately obtained, ensuring that the editing effect of the generated edited image meets the user's editing needs. When applied to AI virtual try-on scenarios, it ensures that the generated edited image clearly and explicitly displays the try-on effect, meeting the user's expectation for accurate try-on results and further guaranteeing the practicality of the method. Attached Figure Description
[0011] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1A schematic diagram of a scene for an exemplary embodiment of this application, illustrating an image editing method; Figure 2 A flowchart illustrating an image editing method provided for an exemplary embodiment of this application; Figure 3 A schematic diagram illustrating a process for generating editing input information to implement image editing operations, provided as an exemplary embodiment of this application; Figure 4 A schematic flowchart of another image editing method provided for an exemplary embodiment of this application; Figure 5 A flowchart illustrating the process of determining an edited image corresponding to the reference image, provided as an exemplary embodiment of this application; Figure 6 A schematic diagram illustrating the principle of an image editing method provided in an exemplary application embodiment of this application; Figure 7 A schematic diagram of the structure of an image editing device provided for an exemplary embodiment of this application; Figure 8 A schematic diagram of the structure of an electronic device provided in an exemplary embodiment of this application. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] It should be noted that, in the case of user information involved in the embodiments of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0014] The various models involved in this application (including but not limited to language models or large models) comply with relevant laws and standards. Furthermore, the technical solutions provided in the embodiments of this application can employ deep learning models with relatively large parameter scales. The large model is merely an example, and the embodiments of this application do not limit the number of model parameters supported by the deep learning model used, aiming to meet actual needs. The deep learning models involved in the embodiments of this application can be artificial intelligence-based language models (LM) or multimodal models (MM).
[0015] Additionally, it should be noted that when user interaction operations or triggering operations are involved in the embodiments of this application, these operations include, but are not limited to, various interaction methods such as touch operations, gesture operations, voice operations, head movement operations, and eye movement operations. Touch operations include, but are not limited to, click operations, double-click operations, long-press operations, swipe operations, pinch operations, or mouse hover operations. Swipe operations include, but are not limited to, straight-line swipes and curved-line swipes.
[0016] Terminology definition: The Unified Generative Understanding Model is a large-scale language model architecture that integrates multiple artificial intelligence (AI) capabilities. It unifies the traditionally separate tasks of understanding and generation into a single model framework, enabling it to handle multiple complex tasks such as text understanding, content generation, and reasoning analysis simultaneously.
[0017] Deep Thinking / Chain of Thought (CoT) is a prompting engineering technique that improves the performance of complex tasks by guiding large language models to demonstrate their reasoning processes. The idea behind CoT is to have the model progressively demonstrate the intermediate reasoning steps before providing the final answer, simulating the human thought process when solving problems.
[0018] AI-powered virtual try-on is an artificial intelligence-based solution that provides users with an immersive shopping experience. The system supports various try-on scenarios, including images of merchandise models and user-uploaded personal photos. AI-powered try-on deeply understands user needs through a unified understanding model, delivering a realistic try-on experience.
[0019] DIT architecture: short for "DiffusionTransformer", which replaces the U-Net architecture in the traditional diffusion model with the Transformer architecture. This design combines the powerful generative capabilities of the diffusion model with the advantages of the Transformer's sequence modeling and attention mechanisms, leveraging its multimodal unification, scalability, and conditional control capabilities in unified generative understanding.
[0020] Interleaved data: Interleaved data refers to a data format in which images and text are arranged alternately and organized in a mixed manner. This data structure is very important in multimodal AI training and applications. That is, the image and text content are not separate, but appear alternately in a logical order to form a coherent information flow.
[0021] To facilitate understanding of the image editing method, apparatus, device, and computer storage medium provided in the embodiments of this application, the relevant technologies are briefly described below: As a strategic priority category for e-commerce platforms, apparel carries a massive daily volume of user purchases. While product detail pages display numerous model images, they don't address the core concern of "how will it look on me?" In traditional shopping models, users must go through the entire process of ordering, receiving, and trying on clothes; if the fit doesn't meet expectations, returns and exchanges incur costs.
[0022] Currently, while virtual try-on systems based on artificial intelligence (AI) provide virtual clothing try-on experiences, they suffer from limitations due to their reliance on general, large-scale models. The lack of a deep-thinking-based understanding means these systems struggle to grasp rapidly changing fashion trends and specialized clothing terminology. This leads to difficulties in understanding complex text commands, ultimately preventing them from meeting users' expectations for precise try-on results.
[0023] To address the aforementioned technical problems, embodiments of this application provide an image editing method, apparatus, device, and computer storage medium, as detailed in the appendix. Figure 1As shown, the execution entity of this image editing method can be an image editing device 200, which can be implemented as a local server, a cloud server, or an edge server. When the image editing device 200 is implemented as a cloud server, the image editing method can be executed in the cloud. Several computing nodes (cloud servers) can be deployed in the cloud, each with computing, storage, and other processing resources. In the cloud, multiple computing nodes can be organized to provide a certain service; of course, a single computing node can also provide one or more services. The cloud can provide this service by providing a service interface, which users can call to use the corresponding service. Service interfaces include Software Development Kits (SDKs) and Application Programming Interfaces (APIs).
[0024] The image editing device 200 is communicatively connected to the client 100, which is used by the user to trigger image editing operations. The client 100 can be any computing device with a certain information interaction capability; specifically, it can be a mobile phone, a personal computer (PC), a tablet computer, a configuration application, etc. Furthermore, the basic structure of the client 100 may include at least one processor. The number of processors depends on the client's configuration and type. The client 100 may also include memory, which can be volatile, such as Random Access Memory (RAM), or non-volatile, such as Read-Only Memory (ROM), flash memory, etc., or both types. The memory typically stores the operating system (OS), one or more applications, and may also store program data. In addition to the processing unit and memory, the client 100 also includes some basic configurations, such as a network interface card (NIC) chip, an I / O bus, a display component, and some peripheral devices. Optionally, some peripheral devices may include, for example, a keyboard, a mouse, a stylus, a printer, etc. Other peripheral devices are well known in the art and will not be described in detail here.
[0025] Image editing device 200 refers to a device capable of performing image editing operations in a network virtual environment, typically referring to a device that utilizes a network for information planning and image editing operations. The image editing device 200 can be an image editing model used to implement image editing operations. Physically, the image editing device 200 can be any device capable of providing computing services and performing corresponding image editing operations, such as a processor, server, etc. The image editing device 200 mainly consists of a processor, hard disk, memory, system bus, etc., and its architecture is similar to that of a general-purpose computer.
[0026] In this embodiment described above, the image editing device 200 and the client 100 are connected via a network, which can be a wireless or wired network connection. If the image editing device 200 and the client 100 are connected via a communication connection, the network standard of the mobile network can be any one of 2G (Global System for Mobile Communications GSM), 2.5G (General Packet Radio Service GPRS), 3G (Wideband Code Division Multiple Access (WCDMA), Time Division Synchronous Code Division Multiple Access (TD-SCDMA), 4G (Long Term Evolution LTE), 4G+ (Enhanced Long Term Evolution LTE+), Global Microwave Access Interoperability (WiMax), 5G, 6G, etc.
[0027] In this embodiment, the client 100 is used by the user to generate an image editing request or image editing task to trigger image editing operations. The image editing task may include a reference image and image editing instructions corresponding to the reference image. Different scenarios may correspond to different image editing instructions. For example, in the AI virtual try-on application scenario, the image editing instruction can be implemented as a "virtual try-on instruction"; in the home application scenario, the image editing instruction can be implemented as a "home decoration background change instruction"; in the image generation application scenario, the image editing instruction can be implemented as a "pose change instruction," "effect change instruction," and so on. The image editing task can then be sent to the image editing device 200 so that the image editing device 200 can perform corresponding image editing operations based on the image editing task.
[0028] Image editing device 200 is used to acquire image editing tasks sent by client 100. The acquired image editing tasks may include a reference image and image editing instructions corresponding to the reference image. Different application scenarios may correspond to different image editing instructions. For example, in the home application scenario, the image editing instructions may be "change the sofa to a modern minimalist style" or "replace the dining table and chairs to a new Chinese style", etc. In the clothing application scenario, the image editing instructions may be "change the model's skirt to a high-IQ elite style", "change the model's overall style to Y2K style", or "change the model's overall style to wabi-sabi style", etc.
[0029] After acquiring the image editing task, the reference image and image editing instructions in the task can be understood to obtain the thought chain information used to implement the image editing operation. Then, the reference image, image editing instructions, and thought chain information can be analyzed and processed to generate editing input information for implementing the image editing operation. The editing input information is embedded representation information. Then, a pre-trained generative understanding model can be used to process the editing input information to obtain the edited image. This effectively ensures the accuracy and reliability of determining the edited image.
[0030] In this embodiment, when using the generative understanding model to analyze and process the edited input information, since the edited input information is generated based on the thought chain information and the generative understanding model has a high understanding and generation capability, the edited image can be obtained stably, and the editing effect of the generated edited image can meet the user's editing needs. When applied to the AI try-on scenario, it can ensure that the generated edited image clearly and explicitly displays the try-on effect, meeting the user's expectation for accurate try-on effect, and further ensuring the practicality of the method.
[0031] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0032] Figure 2 A schematic flowchart of an image editing method provided for an exemplary embodiment of this application; see attached document. Figure 2 As shown, this embodiment provides an image editing method. The execution subject of this method is an image editing device, which can be implemented as software or a combination of software and hardware. When the image editing device is implemented as hardware, it can be various electronic devices capable of performing image editing operations. In some instances, the image editing device can be implemented as an application client, server, cloud server, etc. When the image editing device is implemented as software, it can be installed in the electronic devices exemplified above. Specifically, the image editing method provided in this embodiment may include: Step S201: Obtain the reference image and the image editing instructions corresponding to the reference image.
[0033] Step S202: Understand the reference image and image editing instructions to obtain the thought chain information used to implement the image editing operation.
[0034] Step S203: Based on the reference image, image editing instructions, and thought chain information, generate editing input information for implementing image editing operations. The editing input information is embedded representation information.
[0035] Step S204: Use a pre-trained generative understanding model to process the editing input information to obtain the edited image.
[0036] The specific implementation methods and principles of each of the above steps are explained in detail below: Step S201: Obtain the reference image and the image editing instructions corresponding to the reference image.
[0037] When image editing is required, the image editing device can acquire a reference image and corresponding image editing instructions. The reference image and image editing instructions differ depending on the application scenario. For example, in a clothing try-on scenario, the reference image may include: a preset subject image (which can be any of the following: a person, animal, virtual object, etc.) and at least one clothing image; the image editing instructions are used to identify or characterize descriptive information for changing the clothing of the preset subject. Alternatively, the reference image may include: a preset subject image and a pose image; the image editing instructions are used to identify or characterize descriptive information for adjusting the pose of the preset subject. In a home furnishing scenario, the reference image may include: an original home scene image and a target home scene image; the image editing instructions are used to identify or characterize descriptive information for changing the background of the original home scene image based on the target home scene image. In the field of image generation applications, the reference image can be a single image of a preset scene; the image editing instructions are used to identify or characterize descriptive information for editing the reference image according to a specific effect.
[0038] For image editing instructions corresponding to the reference image, they can be implemented as either voice descriptions or text descriptions. When the image editing instruction is a text description, it can be a natural language description input by the user. For example, image editing instructions could be "Perform a fitting operation based on the provided second clothing image," "Adjust the posture of the first image based on the provided second pose image," or "Change the background of the first home furnishing image based on the background in the provided second home furnishing image." When the image editing instruction is a voice description, it can first undergo a voice-to-text conversion operation to obtain corresponding text descriptions, and then subsequent image editing operations can be performed based on the obtained text descriptions.
[0039] Furthermore, this embodiment does not limit the specific method of obtaining the reference image and the corresponding image editing instructions. In some instances, the reference image and the corresponding image editing instructions can be obtained through interactive operations with a client. In this case, obtaining the reference image and the corresponding image editing instructions may include: determining a client that is communicatively connected to the image editing device, which can generate and obtain the reference image and the corresponding image editing instructions; and then the reference image and the corresponding image editing instructions can be obtained actively or passively through the client, thus effectively ensuring the accuracy and reliability of obtaining the reference image and the corresponding image editing instructions.
[0040] In other instances, the reference image and its corresponding image editing instructions can be obtained not only through interactive operations with the client, but also through data upload operations via different upload interfaces in the image editing interface. In this case, obtaining the reference image and its corresponding image editing instructions may include: displaying a task upload interface for implementing image editing operations; identifying the image upload interface and the text upload interface in the task upload interface; performing an image upload operation via the image upload interface to obtain the reference image; and performing a text upload operation via the text upload interface to obtain the image editing instructions corresponding to the reference image. This also ensures the accuracy and reliability of obtaining the reference image and its corresponding image editing instructions.
[0041] Step S202: Understand the reference image and image editing instructions to obtain the thought chain information used to implement the image editing operation.
[0042] To ensure the quality and effectiveness of editing reference images, after obtaining the reference images and image editing instructions, the reference images and image editing instructions can be analyzed and understood to obtain the thought chain information used to implement image editing operations. This thought chain information can be obtained through in-depth analysis and understanding of the reference images and image editing instructions. In-depth analysis and understanding refers to the process of gaining profound insights by analyzing reference images and image editing instructions from multiple angles and levels, seeing through appearances to the essence. It goes beyond simply understanding and memorizing surface information, emphasizing a deeper understanding of information, critical thinking, and creative thinking.
[0043] In some instances, the understanding of reference images and image editing instructions can be achieved through a generative understanding model for implementing image editing operations. In this case, understanding the reference images and image editing instructions to obtain the thought chain information for implementing image editing operations can include: using the understanding sub-model in the generative understanding model to perform a deep understanding of the reference images and image editing instructions to obtain thought chain information.
[0044] The system includes a pre-configured generative understanding model for deep understanding operations. After obtaining the reference image and image editing instructions, these can be input into the generative understanding model. The understanding sub-model within the generative understanding model then performs deep understanding operations on the reference image and image editing instructions to obtain thought chain information. This effectively ensures the accuracy of obtaining thought chain information.
[0045] In other instances, the understanding of reference images and image editing instructions can be achieved through different understanding network models. In this case, understanding the reference image and image editing instructions to obtain the thought chain information for implementing image editing operations can include: inputting the reference image into an image understanding network model for deep understanding to obtain image understanding information; inputting the image editing instructions into a text understanding network model for deep understanding to obtain text understanding information; and generating thought chain information for implementing image editing operations based on the image understanding information and the text understanding information. This also ensures the flexibility and reliability of generating thought chain information.
[0046] Step S203: Based on the reference image, image editing instructions, and thought chain information, generate editing input information for implementing image editing operations. The editing input information is embedded representation information.
[0047] After obtaining the reference image, image editing instructions, and thought chain information, these can be analyzed and processed to generate editing input information for image editing operations. Since the thought chain information includes relevant information on a deep understanding of the reference image and image editing instructions, the generated editing input information can accurately identify or express the image editing operations.
[0048] In some instances, the editing input information used to implement image editing operations can be generated by analyzing and processing the reference image, image editing instructions, and thought chain information through multiple embedded models. In this case, generating the editing input information based on the reference image, image editing instructions, and thought chain information can include: inputting the reference image into a pre-trained image embedded model for analysis and processing to obtain the image editing input information output by the image embedded model; inputting the image editing instructions and thought chain information into a pre-trained text embedded model for analysis and processing to obtain the text editing input information output by the text embedded model; and determining the editing input information used to implement the image editing operation based on the image editing input information and the text editing input information. The generated editing input information is then used to implement the image editing operation, thus effectively ensuring the accuracy and reliability of the generated editing input information.
[0049] Step S204: Use a pre-trained generative understanding model to process the editing input information to obtain the edited image.
[0050] Since the editing input information is an embedded representation information generated based on the reference image, image editing instructions, and thought chain information, it can represent the relevant information of the reference image, image editing instructions, and thought chain information. Therefore, after obtaining the editing input information, it can be directly input into the pre-trained generative understanding model for analysis and processing, and then the generative understanding model can output the edited image.
[0051] In other instances, the generative understanding model includes a generative sub-model for image generation operations and an understanding sub-model for deep understanding operations. In this case, the image generation operations and deep understanding operations can be implemented through different sub-network models. Processing the edited input information using the pre-trained generative understanding model to obtain the edited image can include: processing the edited input information using the generative sub-model in the generative understanding model to obtain the edited image.
[0052] Specifically, after obtaining the editing input information, it can be input into the generative understanding model. The generative sub-model in the generative understanding model can then be used to perform image generation processing on the editing input information, thereby obtaining the edited image. This effectively ensures the accuracy and reliability of obtaining the edited image.
[0053] For edited images, different image editing commands can generate edited images with different editing effects. For example, in the AI virtual try-on application scenario, the edited image can be a try-on image, which can meet the user's try-on needs based on the reference image; in the home application scenario, the edited image can be a background-changed image, which can meet the user's background-changing needs based on the reference image; in the AI virtual try-on application scenario, the edited image can be a posture-adjusted image, which can meet the user's posture-adjusting needs based on the reference image; in the image generation application scenario, the edited image can be an adjusted effect image. For example, when the reference image is a toy doll image and the image editing command is "show a scene of a toy doll being hugged by a child", the edited image obtained after performing image editing operations on the reference image and the image editing command can be an image of "a child hugging a toy doll". When the reference image is a "fireworks display" and the image editing instruction is "show the scene after the fireworks display ends", the edited image can be a "single scene image after the fireworks end". This shows that the application scenarios and scope of this image editing method are relatively wide, thus effectively expanding the applicable scope and application scenarios of this image editing method.
[0054] The image editing method provided in this embodiment acquires a reference image and corresponding image editing instructions, understands the reference image and the image editing instructions to obtain thought chain information for implementing image editing operations, and generates editing input information for implementing image editing operations based on the reference image, image editing instructions, and thought chain information. The generative understanding model has high understanding and generation capabilities. Then, the pre-trained generative understanding model is used to process the editing input information. Since the editing input information incorporates thought chain information, the edited image can be accurately obtained, and the editing effect of the generated edited image can meet the user's editing needs. When applied to AI virtual try-on scenarios, it can ensure that the generated edited image clearly and explicitly displays the try-on effect, meeting the user's expectation for accurate try-on results and further ensuring the practicality of the method.
[0055] Figure 3 A flowchart illustrating the generation of editing input information for implementing image editing operations is provided as an exemplary embodiment of this application; based on the above embodiments, refer to the appendix. Figure 3As shown, for editing input information, it can be generated not only by multiple embedded models analyzing and processing the reference image, image editing instructions, and thought chain information separately, but also by a single generative understanding model analyzing and processing the reference image, image editing instructions, and thought chain information. In this case, the editing input information generated based on the reference image, image editing instructions, and thought chain information to implement image editing operations can include: Step S301: Use the understanding sub-model in the generative understanding model to process the reference image, image editing instructions and thought chain information to generate fused embedded information.
[0056] The generative understanding model includes an understanding sub-model for implementing deep understanding operations. When deep reasoning operations are required, the understanding sub-model in the generative understanding model can be called to analyze and process the reference image, image editing instructions, and thought chain information, thereby generating fused embedded information corresponding to the reference image, image editing instructions, and thought chain information.
[0057] In some instances, using the understanding sub-model in the generative understanding model to process the reference image, image editing instructions, and thought chain information to generate fused embedded information may include: using the understanding sub-model to process the reference image and image editing instructions to obtain image-text embedded information; using the understanding sub-model to process the thought chain information to generate thought-text embedded information corresponding to the thought chain information; and generating fused embedded information based on the image-text embedded information and the thought-text embedded information.
[0058] Specifically, after obtaining the reference image, image editing instructions, and thought chain information, the reference image and image editing instructions can be input into the generative understanding model. The understanding sub-model within the generative understanding model can then be used to transform the reference image and image editing instructions into embedded information, thereby obtaining image-text embedded information. Next, the thought chain information can be input into the generative understanding model. The understanding sub-model within the generative understanding model can then be used to transform the thought chain information into embedded information, thereby obtaining thought-text embedded information. After obtaining the image-text embedded information and the thought-text embedded information, they can be analyzed and processed. Specifically, the image-text embedded information and the thought-text embedded information can be spliced and fused to generate fused embedded information.
[0059] Step S302: Generate editing input information based on the fused embedded information.
[0060] After obtaining the fused embedded information, editing input information for image editing operations can be generated based on the fused embedded information. In some instances, the fused embedded information can be directly determined as the editing input information, which effectively ensures the accuracy and reliability of the generated editing input information.
[0061] In other instances, the editing input information can be generated based on the image embedded information corresponding to the reference image. In this case, generating the editing input information based on the fused embedded information may include: processing the reference image using an image processing sub-model to obtain the image embedded information; and generating the editing input information based on the fused embedded information and the image embedded information.
[0062] To ensure the effectiveness of image editing operations and minimize image distortion, an image processing sub-model can be used to process the reference image, thereby obtaining image embedded information. This image embedded information can be low-dimensional, dense, and semantically rich feature information representing the image information of the reference image. After obtaining the image embedded information, the fused embedded information and the image embedded information can be analyzed and processed. In some instances, the fused embedded information and the image embedded information can be spliced and fused to generate editing input information representing the image information.
[0063] In this embodiment, the understanding sub-model in the generative understanding model is used to process the reference image, image editing instructions, and thought chain information to generate fused embedded information. Then, the editing input information is generated based on the fused embedded information, which effectively ensures the accuracy and reliability of the generated editing input information.
[0064] Figure 4 A schematic flowchart of another image editing method provided for an exemplary embodiment of this application; based on any of the above embodiments, refer to the appendix. Figure 4 As shown, before processing the edit input information using a pre-trained generative understanding model, the generative understanding model can be obtained through model training. In this case, the method in this embodiment may further include: Step S401: Obtain training data for implementing image editing operations. The training data includes a reference image, thought chain information corresponding to the reference image, editing instructions, and the edited image. The reference image and the edited image satisfy the requirements defined by the thought chain information and editing instructions.
[0065] When there is a need to train a generative understanding model, training data for implementing image editing operations can be obtained first. This training data may include a reference image, thinking information corresponding to the reference image, editing instructions, and the edited image. The edited image included in the training data is not an ordinary edited image, but an edited image that meets the requirements of the thought chain information and editing quality constraints between the reference image and the edited image.
[0066] In some instances, training data for implementing image editing operations can be obtained through interactive operations with a client. In this case, obtaining training data for implementing image editing operations may include: identifying a client that is communicatively connected to the image editing device, which may store the training data; and then the training data can be actively or passively obtained through the client, thus effectively ensuring the accuracy and reliability of obtaining the training data.
[0067] In other instances, training data can be obtained not only through interactive operations with the client, but also through reference images, editing instructions, and task-generated images corresponding to historical image editing operations. In this case, obtaining training data for implementing image editing operations may include: obtaining reference images, editing instructions, and task-generated images corresponding to historical image editing operations; determining the thought chain information corresponding to historical image editing operations based on the reference images, editing instructions, and task-generated images; determining the edited image corresponding to the reference image based on the thought chain information, reference images, and task-generated images; and obtaining training data based on the reference images, editing instructions, thought chain information, and edited images.
[0068] To accurately obtain training data, the reference image, editing instructions, and task-generated image corresponding to historical image editing operations can be acquired first. The reference image can be the one corresponding to the historical image editing operation, and the task-generated image is the edited image corresponding to the historical image editing operation. In some instances, the reference image, editing instructions, and task-generated image corresponding to historical image editing operations can be stored in a preset area or a preset device. The reference image, editing instructions, and task-generated image corresponding to historical image editing operations can be obtained by accessing the preset area or preset device.
[0069] After obtaining the reference image, editing instructions, and task-generated image corresponding to historical image editing operations, these images can be analyzed to determine the thought chain information corresponding to the historical image editing operations. In some instances, the thought chain information can be obtained by analyzing the reference image, editing instructions, and task-generated image using a pre-trained large language model. In this case, determining the thought chain information corresponding to historical image editing operations based on the reference image, editing instructions, and task-generated image can include: inputting the reference image, editing instructions, and task-generated image into the pre-trained large language model for deep thinking operations, and obtaining the thought chain information output by the large language model corresponding to the historical image editing operations. This effectively ensures the accuracy and reliability of determining the thought chain information.
[0070] In other instances, thought chain information can be obtained by analyzing and processing reference images, editing instructions, and task-generated images using a pre-trained large language model. It can also be obtained by rewriting editing instructions. In this case, determining the thought chain information corresponding to historical image editing operations based on reference images, editing instructions, and task-generated images can include: rewriting editing instructions based on reference images and task-generated images to obtain rewritten execution instructions, where the rewritten execution instructions have more detailed information than the editing instructions; and determining the thought chain information based on reference images, task-generated images, and rewritten execution instructions.
[0071] Since image editing operations are closely related to reference images, editing instructions, and thought chain information, if the acquired training data only includes reference images, editing instructions, and task-generated images, it is necessary to generate corresponding thought chain information using these reference images, editing instructions, and task-generated images. Specifically, the editing instructions can be rewritten based on the reference images and task-generated images to obtain rewritten execution instructions. The rewritten execution instructions have more detailed information than the editing instructions, which can improve the accuracy and reliability of image editing operations to a certain extent when performing image editing operations based on the rewritten execution instructions.
[0072] In some instances, the rewritten execution instructions can be obtained through a pre-trained rewriting network model. In this case, rewriting the editing instructions based on the reference image and the task-generated image to obtain the rewritten execution instructions may involve inputting the reference image, the task-generated image, and the editing instructions into the pre-trained rewriting network model for processing, and obtaining the rewritten execution instructions output by the rewriting network model. This effectively ensures the accuracy and reliability of obtaining the rewritten execution instructions. In other instances, the rewritten execution instructions can not only be obtained through a pre-trained rewriting network model, but also through a preset rewriting algorithm to analyze and process the reference image, the task-generated image, and the editing instructions. As long as the rewritten execution instructions can be obtained stably, this will not be elaborated further.
[0073] After acquiring the reference image, the task-generated image, and the rewritten execution instructions, these elements can be analyzed to determine the thought chain information used to implement the image editing operation. In some instances, the thought chain information can be determined using a pre-trained multimodal large model. In this case, determining the thought chain information based on the reference image, the task-generated image, and the rewritten execution instructions can include: using the pre-trained multimodal large model to perform deep thinking analysis on the reference image, the task-generated image, and the rewritten execution instructions to obtain the thought chain information output by the multimodal large model. This effectively ensures the accuracy and reliability of determining the thought chain information.
[0074] It is worth noting that, for historical image editing operations, since the task-generated image is only generated by analyzing and processing the reference image and editing instructions, the image editing effect of the task-generated image may not meet the user's image editing needs. In this case, in order to ensure that the image editing effect between the edited image in the training data and the reference image meets the user's editing needs, after determining the thought chain information corresponding to the historical image editing operation, the thought chain information, the reference image, and the task-generated image can be analyzed and processed to stably determine the edited image corresponding to the reference image.
[0075] In some instances, the edited image can be determined using a pre-trained multimodal network model. In this case, determining the edited image corresponding to the reference image based on the thought chain information, the reference image, and the task-generated image can include: inputting the thought chain information, the reference image, and the task-generated image into the pre-trained multimodal network model for image generation processing, thereby obtaining the edited image corresponding to the reference image output by the multimodal network model. This effectively ensures the accuracy and reliability of determining the edited image.
[0076] In other instances, the edited image corresponding to the reference image can be determined not only by a pre-trained multimodal network model, but also by an image generation operation based on the rewritten execution instructions, thought chain information, and the reference image. In this case, determining the edited image corresponding to the reference image based on the thought chain information, the reference image, and the task-generated image may include: performing an image generation operation based on the rewritten execution instructions, thought chain information, and the reference image to obtain at least one generated image; and determining the edited image corresponding to the reference image based on the task-generated image and at least one generated image.
[0077] Specifically, to accurately obtain an edited image that meets the desired image editing effect, an image generation operation can be performed based on the rewritten execution instructions, thought chain information, and a reference image. This results in at least one generated image, with different generated images corresponding to different image editing effects. After obtaining the task-generated image and at least one generated image, these images can be analyzed to determine the edited image corresponding to the reference image. In some instances, the edited image can be either the task-generated image or the at least one generated image, or it can be an image whose image editing effect meets the user's image editing needs. This effectively ensures that the image editing effect between the edited image and the reference image satisfies the user's image editing requirements.
[0078] After obtaining the edited image corresponding to the reference image, training data can be generated based on the obtained reference image, editing instructions, and thought chain information. Specifically, the interleaved data triple consisting of the edited image, reference image, editing instructions, and thought chain information—instruction (i.e., editing instructions) - thought (i.e. thought chain information) - image (including: edited image and reference image)—can be determined as the training data. This effectively ensures the accuracy and reliability of the training data acquisition.
[0079] Step S402: Use the training data to perform cascade training on the understanding sub-model and the generative sub-model used to construct the generative understanding model to obtain the generative understanding model.
[0080] After obtaining the training data, the understanding sub-model and the generative sub-model used to construct the generative understanding model can be cascaded and trained to obtain the generative understanding model. In some instances, the understanding sub-model can correspond to the first set of model hyperparameters, and the generative sub-model can correspond to the second set of model hyperparameters. After obtaining the training data, the understanding sub-model and the generative sub-model used to construct the generative understanding model can be trained to obtain the model training loss function. Then, based on the model training loss function, the first set of model hyperparameters of the understanding sub-model and the second set of model hyperparameters of the generative sub-model can be adjusted and optimized to obtain the trained generative understanding model.
[0081] In other instances, cascaded training can be achieved by first freezing the hyperparameters of one sub-model and then adjusting the hyperparameters of another sub-model. In this case, cascading training of the understanding sub-model and the generative sub-model (which constitute the generative understanding model) using training data to obtain the generative understanding model can include: first, freezing the first set of model hyperparameters corresponding to the understanding sub-model; then, training the generative sub-model using training data to obtain the generative loss function corresponding to the generative sub-model; and finally, adjusting and optimizing the second set of model hyperparameters corresponding to the generative sub-model based on the generative loss function to obtain the trained generative sub-model; then, freezing the second set of model hyperparameters corresponding to the generative sub-model; then, training the understanding sub-model using training data to obtain the understanding loss function corresponding to the understanding sub-model; and finally, adjusting and optimizing the first set of model hyperparameters corresponding to the understanding sub-model based on the understanding loss function to obtain the trained understanding sub-model. The trained generative understanding model can then be obtained based on the trained generative and understanding sub-models, effectively ensuring the accuracy and reliability of obtaining the generative understanding model.
[0082] In this embodiment, by acquiring training data for implementing image editing operations, the understanding sub-model and the generative sub-model used to construct the generative understanding model can be cascaded and trained. This effectively ensures the accuracy and reliability of the training data for the generative understanding model, and facilitates the quality and effect of image editing operations based on the trained generative understanding model, further improving the practicality of the method.
[0083] Figure 5 A flowchart illustrating the determination of an edited image corresponding to a reference image is provided as an exemplary embodiment of this application; based on the above embodiments, refer to the appendix. Figure 5As shown, the edited image corresponding to the reference image can be determined not only by a pre-trained multimodal network model, but also by the image editing quality corresponding to the task-generated image and at least one generated image. Therefore, determining the edited image corresponding to the reference image based on the task-generated image and at least one generated image can include: Step S501: Based on the reference image, the rewritten execution instructions, and the thought chain information, determine the image editing quality corresponding to the task-generated image and at least one generated image.
[0084] For the reference image, since the task-generated image and at least one generated image have different image editing effects, in order to accurately obtain the edited image used for model training, after obtaining the reference image, the rewritten execution instructions, and the thought chain information, the reference image, the rewritten execution instructions, and the thought chain information can be analyzed and processed to determine the image editing quality corresponding to the task-generated image and at least one generated image.
[0085] In some instances, determining the image editing quality of the task-generated image and at least one generated image based on the reference image, the rewritten execution instructions, and the thought chain information can include: using a pre-trained quality detection model to analyze and process the reference image, the rewritten execution instructions, and the thought chain information, thereby determining the image editing quality of the task-generated image and at least one generated image output by the quality detection model. This effectively ensures the accuracy and reliability of determining the image editing quality of the task-generated image and at least one generated image.
[0086] In other instances, the image editing quality of the task-generated image and at least one generated image can be determined not only by a pre-trained quality detection model, but also by editing quality rules used to identify the image editing quality of the task-generated image and the generated image. In this case, determining the image editing quality of the task-generated image and at least one generated image based on the reference image, the rewritten execution instructions, and the thought chain information can include: determining the editing quality rules used to identify the image editing quality of the task-generated image and at least one generated image; and then using the editing quality rules to analyze and process the task-generated image and at least one generated image to obtain the image editing quality corresponding to each of the task-generated image and at least one generated image. This also ensures the accuracy and reliability of determining the image editing quality of the task-generated image and at least one generated image.
[0087] Step S502: Determine the edited image based on the task-generated image and the image editing quality corresponding to at least one generated image.
[0088] Since the task-generated images and different generated images can correspond to different image editing qualities, these qualities can reflect user satisfaction with the task-generated images and the image editing of each generated image to a certain extent. Generally, higher image editing quality leads to higher image editing satisfaction, and lower image editing quality leads to lower image editing satisfaction. Therefore, to ensure the training quality and effectiveness of the generative understanding model, it is necessary to obtain images with high satisfaction as edited images. Determining the edited image based on the task-generated images and the image editing quality of at least one generated image can include: sorting the task-generated images and at least one generated image based on their image editing quality to obtain image sorting information; and determining the edited image from among the task-generated images and at least one generated image based on the image sorting information. The edited image has a higher image editing quality than the other generated images, thus effectively ensuring the accuracy and reliability of determining the edited image.
[0089] In this embodiment, the image editing quality corresponding to the task-generated image and at least one generated image is determined based on the reference image, the rewritten execution instructions, and the thought chain information. The edited image is then determined based on the image editing quality corresponding to the task-generated image and at least one generated image. This effectively ensures the accuracy and reliability of determining the edited image and further guarantees the accuracy of generating training data based on the edited image.
[0090] In a specific application, taking a clothing fitting scenario as an example, this application embodiment provides an image editing method based on Unified Generative Understanding deep thinking. The main body executing this image editing method is the Unified Generative Understanding large model. (See attached diagram.) Figure 6As shown, the unified generative understanding model includes: an understanding sub-model {e.g., a Vision-Language Model (VLM)}, a generative sub-model {e.g., a Diffusion Transformer (DIT) based on the Transformer architecture}, and an image processing sub-model {e.g., an image embedding extraction model (SigMoidLossforLanguage-Image Pre-training (SigLIP))}. This method can inject the deep thought chains produced by the understanding sub-model into the editing input information of the generative sub-model, thereby achieving better image editing quality and effects. Specifically, this image editing method can include a model training phase and a model application phase of the unified generative understanding model. The model training phase can include the following steps: Step 11: Collect image data for common e-commerce vertical tasks (such as trying on clothes, changing the background of the smallest stock keeping unit (SKU), modifying the task pose, body shape, etc.).
[0091] The image data can include reference images, editing instructions, and task-generated images involved in historical image editing operations. The editing instructions are different for different application scenarios. For example, in the application scenario of trying on clothes, the editing instruction can be "perform a try-on operation on the subject in the original reference image"; in the application scenario of home decoration, the editing instruction can be "replace the background in the original reference image based on the editing reference image"; and in the application scenario of trying on clothes, the editing instruction can be "change the pose of the subject in the original reference image".
[0092] Step 12: Rewrite the editing instructions using a multimodal large model to obtain the rewritten execution instructions.
[0093] The multimodal large model can construct editing instructions based on reference images and task-generated images in the image data to obtain rewritten execution instructions. These rewritten execution instructions can be general editing instructions or task instructions that require reasoning capabilities (such as spatial reasoning, action / process reasoning, purpose / context reasoning, causal / temporal reasoning, etc.). Compared to the original editing instructions, these rewritten execution instructions contain more detailed information. For example, they can include detailed comparison information obtained by comparing the differences between the reference image and the task-generated image (including at least one of the following: pose differences, image style differences, background differences, theme clothing differences, etc.).
[0094] Step 13: Again, use the multimodal large model to perform in-depth thinking analysis on the reference image, the rewritten execution instructions, and the task-generated image to determine the thought chain information corresponding to the image data, which can be used as one of the training data for the understanding sub-model in the unified generation understanding large model.
[0095] In this process, reference images, rewritten execution instructions, and task-generated images are input into a multimodal large model, enabling the multimodal large model to perform in-depth thinking and analysis on the reference images, rewritten execution instructions, and task-generated images, and obtain the thought chain information corresponding to the image data output by the multimodal large model.
[0096] Step 14: Use the generating sub-model to perform image generation operations on the thought chain information, reference image, and rewritten execution instructions to obtain at least one generated image.
[0097] Step 15: Analyze and process at least one generated image and the task-generated image using a multimodal large model to verify the raw image effect corresponding to each of the at least one generated image and the task-generated image. Then, based on the raw image effect corresponding to each of the at least one generated image and the task-generated image, determine the edited image corresponding to the reference image.
[0098] After obtaining at least one generated image, a multimodal large model or large language model can be used to analyze and process the at least one generated image and the task-generated image based on a reference image, rewritten execution instructions, and thought chain information. Specifically, this analysis can include at least one of the following: whether the instructions are strictly met, whether the image is complete and consistent, whether visual errors occur, and whether it conforms to the laws of the physical world. This allows for the determination of the raw image effect corresponding to each of the at least one generated image and the task-generated image. Then, based on the raw image effects corresponding to each of the at least one generated image and the task-generated image, an edited image corresponding to the reference image can be determined. The number of edited images can be one or more images that have passed the above verification operations. In some instances, the edited images can be obtained by sorting and scoring the raw image effects, and the raw image effect must be greater than or equal to a preset lower limit value. The preset lower limit value is used to identify the minimum score information that satisfies the raw image effect requirement.
[0099] In other instances, the edited image corresponding to the reference image may include a positive sample image and a negative sample image. The positive sample image may be an image whose generated image effect after the above-mentioned large language model effect analysis operation is greater than or equal to a preset lower limit value, and the negative sample image may be an image whose generated image effect after the above-mentioned large language model effect analysis operation is less than a preset lower limit value.
[0100] It is worth noting that if the quality of the generated image and the task-generated image are both less than the preset lower limit, it means that the quality of the generated image and the task-generated image does not meet the requirements. In this case, in order to ensure the quality and effect of model training for the unified generation of the large model, all the above sample data can be ignored and training data for model training can be regenerated.
[0101] Step 16: Based on the verified edited image corresponding to the reference image, generate interleaved data triples of instruction (i.e., editing instruction) - thought (i.e. thought chain information) - image (including: reference image and the edited image corresponding to the reference image). Then, based on the interleaved data triples as training data, cascade training of the understanding sub-model and the generative sub-model based on the training data to obtain a unified generative understanding model after training.
[0102] Before cascading the understanding sub-model and the generative sub-model, the understanding sub-model can be trained using supervised fine-tuning (SFT) based on the constructed training data. After training, the hyperparameters of the understanding sub-model can be frozen. Then, the understanding sub-model and the generative sub-model can be cascaded and trained based on the constructed training data to obtain a unified generative understanding model after training. This effectively completes the training of the unified generative understanding model and ensures the image editing quality and effect of the unified generative understanding model.
[0103] After completing the training phase of the unified generative understanding model, image editing operations can be performed using the trained model to achieve end-to-end image editing. The image editing methods at this stage can include: Step 21: Obtain the image editing task.
[0104] In some instances, an image editing task may include an original reference image and image editing instructions. For example, the original reference image could be a "teddy bear image," and the image editing instructions could be "Can you show a scene of this crochet teddy bear being hugged by a child?"; or, the original reference image could be an "image of the Seattle Space Needle during a fireworks display," and the image editing instructions could be "Can you show the view around the Seattle Space Needle after the fireworks display?"; or, the original reference image could be an "image of several friends together," and the image editing instructions could be "Show a scene of their celebration dinner after winning an award," and so on.
[0105] In other instances, image editing tasks may include an original reference image, an editable reference image, and image editing instructions. For example, in a clothing try-on scenario, the original reference image could be a personal photo uploaded by the user, the editable reference image could be a "clothing image" used to perform the try-on operation, including "top images," "skirt images," etc., and the image editing instruction could be "perform a try-on operation on the subject in the original reference image." In a home decoration scenario, the original reference image could be a home decor image uploaded by the user, the editable reference image could be a home decor image used to change the background, and the image editing instruction could be "change the background in the original reference image based on the editable reference image." Again, in a clothing try-on scenario, the original reference image could be a personal photo uploaded by the user, the editable reference image could be a pose reference image or pose marker image used to change the pose, and the image editing instruction could be "change the pose of the subject in the original reference image."
[0106] Step 22: Input the image editing task into the understanding sub-model (VLM) for deep thinking operation, obtain the deep thought chain (COT) output by VLM corresponding to the image editing task, and use the understanding sub-model to analyze and process the image editing task and the deep thought chain to obtain the image and text embedded information corresponding to the image editing task and the thinking embedded information corresponding to the deep thought chain.
[0107] Example 1: In a virtual try-on scenario, the original reference image can be a personal photo uploaded by the user, and the editable reference image can be a "clothing image" used to perform the try-on operation, including "top images" and "skirt images," etc. The image editing instruction can be "perform a try-on operation on the subject in the original reference image." After deeply considering the above data, the resulting deep thought chain can include information such as "understanding information about the image task, specific operational information for image editing, and precautions for editing operations." For example, the deep thought chain could be "edit..." The goal is to replace the sleeveless white knit top (with three buttons and a collar) in the second image and the black skirt in the third image on the model in the first image, while keeping the background and model's pose unchanged in the first image. Specific steps include: removing the black down vest and black long-sleeved undershirt from the model, replacing the upper body with the white knit sleeveless top; replacing the brown-red sweatpants and white sneakers with a black skirt and shoes that match the skirt's style (such as heels or flats). After the replacement, the overall outfit will be a white knit top and black skirt, with the same style, pose, and background as the original. Figure 1 To.
[0108] Example 2: In a home decoration application scenario, the original reference image can be a home decoration image uploaded by the user, the editing reference image can be a home image used to change the home background, and the image editing instruction can be "to replace the background in the original reference image based on the editing reference image." After deep thinking about the above data, the resulting deep thinking chain can include information such as "understanding information about the image task, specific operational information about image editing, and precautions for editing operations." For example, the deep thinking chain can be: "The editing goal is to keep the sofa, coffee table, carpet, and related decorations (such as table ornaments, lamps, etc.) in the first image completely unchanged, and only replace the original background of the first image with a spacious and bright background in the second image—that is, the background wall, windows, green plants, marble floor, and carpet layout—with the original background of the first image, which has a landscape window, paintings, and original floor tiles. The specific operation is: to carefully extract the sofa, coffee table, carpet, and decorations from the first image, accurately fit them onto the second background, make appropriate projections and contour adjustments to the carpet and objects to ensure consistent lighting and shadows, and finally blend them into the new background, making the picture natural and coherent without any abrupt splicing marks."
[0109] Example 3: In the application scenario of trying on clothes, the original reference image can be a personal photo uploaded by the user, and the editing reference image can be a pose reference image or pose marker image used to change the pose, etc. The image editing instruction can be "to change the pose of the main subject in the original reference image". After deep thinking about the above data, the obtained deep thinking chain can include information such as "the understanding of the image task, the specific operation information of image editing, and the precautions for editing operations". For example, the deep thinking chain can be: "The goal is to adjust the body movements of the person in the first image according to the pose shown in the second image, and it is necessary to ensure that the background, clothing, and accessories of the first image remain unchanged. When implementing the transformation, the person's left arm should be adjusted from naturally hanging down (in pocket) to bent elbow and raised, with the forearm across the abdomen, and the left hand naturally resting on or grasping the front of the coat. The right shoulder should be slightly tucked in, so that the right hand is still carrying a bag but hangs more naturally at the side of the body, with a coordinated posture that conforms to the skeletal reference. The body center of gravity and overall standing posture can be slightly adjusted to make the movement natural, but the background (wallpaper, lines, dried flower basket, etc.) must remain consistent with the original." Figure 1 "No change required."
[0110] Example 4: In image generation applications, the original reference image can be a single photo of a model, the editable reference image can be "empty," and the image editing instruction can be "replace this outfit with clothing suitable for wearing in Australia in December." After deep consideration of the above data, the resulting thought chain can include: "Considering that December is summer in Australia, the weather is usually very hot. The original outfit is too warm. Therefore, the long skirt will be replaced with clothing more suitable for hot weather, such as light-colored shorts or a shorter, more breathable skirt. The T-shirt can be replaced with a sleeveless top or vest. Sneakers can be replaced with more open shoes, such as sandals or canvas shoes, and adding accessories such as sunglasses will complete this summer look," and so on.
[0111] Step 23: Input the original reference image included in the image editing task into the image processing sub-model for analysis and processing to obtain the image embedded information corresponding to the original reference image.
[0112] To avoid image distortion during image editing, an image processing sub-model can be used to extract features from the original reference image. This allows the acquisition of embedded image information corresponding to the original reference image, which identifies the pixel attribute information of the original reference image.
[0113] Step 24: Combine the embedded text and image information, the embedded thought information, and the embedded image information to obtain the editing input information used to perform image editing operations.
[0114] Step 25: Input the editing information into the generated sub-model DIT to perform image editing operations and obtain the edited image output by DIT.
[0115] The technical solution provided in this application embodiment ensures high-quality training data through quality inspection for the unified generative understanding model, and adopts cascaded training of the understanding sub-model and the generative sub-model. This realizes the model training operation through a four-stage training process of "instruction construction - thought generation - effect verification - joint training", which significantly improves the inference accuracy of image editing tasks based on the trained unified generative understanding model. Afterwards, the trained unified generative understanding model can be used for image editing. During image editing, the CoT deep thinking mechanism is systematically introduced into the unified generative understanding model, constructing a collaborative mechanism between deep reasoning on the understanding side and precise control on the generation side. This overcomes the technical bottleneck caused by the separation of the understanding model and the generation model in traditional solutions. While maintaining model unity, it significantly improves the ability to handle complex fitting scenarios, especially performing well in handling the professional needs of the e-commerce vertical field. Specifically, the understanding-side network model not only provides basic semantic encoding capabilities, but more importantly, it can output deep thinking analysis operations for image editing tasks, thereby obtaining the corresponding deep thought chain. Then, the deep thought chain with rich semantic information is injected as an enhancement condition into the generation-side network model, enabling it to more accurately meet the image editing needs of the edited image, thus ensuring the quality and effect of image editing operations.
[0116] Because the technical solution described in this application embodiment has a certain degree of generalization ability, it can be applied to various image editing scenarios, thus providing more possibilities for AI virtual try-on scenarios. When this technical solution is applied to a virtual try-on scenario, for the unified generative understanding model, an efficient fashion knowledge injection mechanism can be constructed. Through this mechanism, the latest fashion trends, professional clothing terminology, and style characteristics are integrated into the unified generative understanding model in real time, ensuring that the model's generation effect remains synchronized with the forefront of fashion. This enables the model to accurately understand and process the professional virtual try-on terminology and complex requirements of e-commerce platforms. Users then only need to upload a personal photo and their fitting requirements. The system then performs a deep analysis of these details. Specifically, for the unique needs of e-commerce fitting scenarios, a deep-thinking dataset containing professional knowledge such as clothing styles, matching rules, and body fit is pre-built. Based on this dataset, deep-thinking operations are performed on multiple dimensions, including clothing style matching, body fit assessment, and color matching rationality. This yields a deep thought chain. Combining this deep thought chain with the personal photo and fitting requirements, fitting images are generated, ensuring that the images have a realistic and expected fitting effect. This provides users with a more accurate virtual fitting experience, further guaranteeing the practicality of the method.
[0117] Furthermore, in some of the processes described in the above embodiments and accompanying drawings, multiple operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 11, 12, etc., are merely used to distinguish different operations and do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0118] Figure 7 A schematic diagram of an image editing apparatus provided for an exemplary embodiment of this application; see attached drawing. Figure 7 As shown, this embodiment provides an image editing device for performing the above-described... Figure 2 The image editing method shown, specifically, the image editing device may include: The first acquisition module 11 is used to acquire a reference image and image editing instructions corresponding to the reference image; The first processing module 12 is used to understand the reference image and image editing instructions to obtain the thought chain information used to implement image editing operations; The first generation module 13 is used to generate editing input information for implementing image editing operations based on the reference image, image editing instructions and thought chain information. The editing input information is embedded representation information. The first processing module 12 is also used to process the editing input information using a pre-trained generative understanding model to obtain the edited image.
[0119] The image editing device in this embodiment can also perform the above-described... Figures 1-6 The description of the embodiments shown is for reference only, and will not be elaborated upon here.
[0120] like Figure 8 As shown, this embodiment provides an electronic device for performing the above-described... Figure 2 The image editing method shown may include an electronic device that includes a memory 24 and a processor 25.
[0121] Memory 24 is used to store computer programs and can be configured to store various other data to support operation on the electronic device. Examples of this data include instructions for any application or method used to operate on the electronic device, data structures, contact data, phone book data, messages, pictures, videos, etc.
[0122] The processor 25, coupled to the memory 24, executes the computer program in the memory 24 for: acquiring a reference image and image editing instructions corresponding to the reference image; understanding the reference image and image editing instructions to obtain thought chain information for implementing image editing operations; generating editing input information for implementing image editing operations based on the reference image, image editing instructions, and thought chain information, wherein the editing input information is embedded representation information; and processing the editing input information using a pre-trained generative understanding model to obtain the edited image.
[0123] Furthermore, such as Figure 8 As shown, the electronic device also includes other components such as a communication component 26, a display 27, a power supply component 28, and an audio component 29. Figure 8 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 8 The components shown. Additionally... Figure 8 The components within the center frame are optional, not mandatory, and their specific requirements depend on the product form of the work node. In this embodiment, the work node can be a terminal device such as a desktop computer, laptop computer, smartphone, or IoT device, or a server-side device such as a conventional server, cloud server, or server array. If the work node in this embodiment is implemented as a terminal device such as a desktop computer, laptop computer, or smartphone, it may include... Figure 8 The components within the center frame; if the working node in this embodiment is implemented as a server-side device such as a conventional server, cloud server, or server array, it may not include... Figure 8 The component within the center frame.
[0124] The aforementioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0125] The aforementioned communication component is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel.
[0126] The aforementioned display includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a Touch Panel, the screen can be implemented as a touchscreen to receive input signals from the user. The Touch Panel includes one or more touch sensors to sense touches, swipes, and gestures on the Touch Panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.
[0127] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.
[0128] The aforementioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0129] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium includes volatile or non-volatile components, or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium. Accordingly, this application also provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a processor, the processor is able to implement the steps in the above method embodiments. It should be understood that each step or combination of steps in the above method flow can be implemented by the computer program or instructions. In addition, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device, so that the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device can be implemented as a means to implement the corresponding functions in the above method embodiments.
[0130] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0131] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. An image editing method, characterized in that, include: Acquire a reference image and corresponding image editing instructions; By understanding the reference image and the image editing instructions, a thought chain information for implementing image editing operations can be obtained; Based on the reference image, image editing instructions, and thought chain information, editing input information for implementing image editing operations is generated, wherein the editing input information is embedded representation information; The edited input information is processed using a pre-trained generative understanding model to obtain the edited image.
2. The method according to claim 1, characterized in that, Understanding the reference image and the image editing instructions yields the thought process information for implementing the image editing operation, including: The understanding sub-model in the generative understanding model is used to perform a deep understanding of the reference image and the image editing instructions to obtain the thought chain information.
3. The method according to claim 1, characterized in that, Based on the reference image, image editing instructions, and the thought chain information, editing input information for implementing image editing operations is generated, including: The understanding sub-model in the generative understanding model is used to process the reference image, image editing instructions, and thought chain information to generate fused embedded information; The editing input information is generated based on the fused embedded information.
4. The method according to claim 3, characterized in that, The understanding sub-model in the generative understanding model is used to process the reference image, image editing instructions, and thought chain information to generate fused embedded information, including: The reference image and the image editing instructions are processed using the understanding sub-model to obtain embedded image and text information; The thought chain information is processed using the understanding sub-model to generate thought-embedded information corresponding to the thought chain information; The fused embedded information is generated based on the embedded information of the graphics and the embedded information of the thinking.
5. The method according to claim 3, characterized in that, Based on the fused embedded information, the editing input information is generated, including: The reference image is processed using the image processing sub-model to obtain image embedded information; The editing input information is generated based on the fused embedded information and the image embedded information.
6. The method according to any one of claims 1-5, characterized in that, The edited input information is processed using a pre-trained generative understanding model to obtain the edited image, including: The edited input information is processed using the generative sub-model in the generative understanding model to obtain the edited image.
7. The method according to any one of claims 1-5, characterized in that, Before processing the edited input information using a pre-trained generative understanding model, the method further includes: Acquire training data for implementing image editing operations. The training data includes a reference image, thought chain information corresponding to the reference image, editing instructions, and an edited image. The reference image and the edited image satisfy the requirements defined by the thought chain information and the editing instructions. The understanding sub-model and the generative sub-model used to construct the generative understanding model are cascaded and trained using the training data to obtain the generative understanding model.
8. The method according to claim 7, characterized in that, Obtain training data for implementing image editing operations, including: Acquire the reference image, editing instructions, and task-generated image corresponding to the historical image editing operations; Based on the reference image, editing instructions, and task-generated image, determine the thought chain information corresponding to the historical image editing operation; Based on the thought chain information, the reference image, and the task-generated image, determine the edited image corresponding to the reference image; The training data is obtained based on the reference image, editing instructions, thought chain information, and the edited image.
9. The method according to claim 8, characterized in that, Based on the reference image, editing instructions, and the task-generated image, the thought chain information corresponding to the historical image editing operation is determined, including: The editing instructions are rewritten based on the reference image and the task-generated image to obtain rewritten execution instructions, wherein the rewritten execution instructions have more detailed information than the editing instructions. The thought chain information is determined based on the reference image, the task-generated image, and the rewritten execution instructions.
10. The method according to claim 9, characterized in that, Based on the thought chain information, the reference image, and the task-generated image, determining the edited image corresponding to the reference image includes: Based on the rewritten execution instructions, the thought chain information, and the reference image, an image generation operation is performed to obtain at least one generated image; Based on the task-generated image and the at least one generated image, an edited image corresponding to the reference image is determined.
11. The method according to claim 10, characterized in that, Based on the task-generated image and the at least one generated image, determining the edited image corresponding to the reference image includes: Based on the reference image, the rewritten execution instructions, and the thought chain information, determine the image editing quality corresponding to the task-generated image and each of the at least one generated image; The edited image is determined based on the image generated by the task and the image editing quality corresponding to each of the at least one generated image.
12. The method according to claim 11, characterized in that, Based on the image generated by the task and the image editing quality corresponding to each of the at least one generated image, the edited image is determined, including: Based on the image editing quality, the task-generated images and at least one generated image are sorted to obtain image sorting information; Based on the image sorting information, the edited image is determined from the task-generated image and the at least one generated image, wherein the image editing quality of the edited image is higher than that of the other generated images.
13. An image editing device, characterized in that, include: The first acquisition module is used to acquire a reference image and image editing instructions corresponding to the reference image; The first processing module is used to understand the reference image and the image editing instructions to obtain the thought chain information for implementing the image editing operation; The first generation module is used to generate editing input information for implementing image editing operations based on the reference image, image editing instructions and the thought chain information, wherein the editing input information is embedded representation information; The first processing module is further configured to process the edit input information using a pre-trained generative understanding model to obtain the edited image.
14. An electronic device, characterized in that, include: A memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions, when executed by the processor, implement the method of any one of claims 1-12.
15. A computer storage medium, characterized in that, Used to store a computer program that, when executed by a computer, implements the method of any one of claims 1-12.
16. A computer program product, characterized in that, include: A computer program, when executed by a processor of an electronic device, causes the processor to perform the steps of the method of any one of claims 1-12.