Training method of image processing model and image generation method and device

Through the modularly designed sample generation and model training components, the image prompt text set is generated by combining visual language and text processing models. The LoRA model fine-tuning is used to solve the problem of high training complexity of open source models in specific industries, and realize efficient and low-threshold image processing model training and generation.

CN120339756AInactive Publication Date: 2025-07-18ALIBABA CLOUD FEITIAN (HANGZHOU) CLOUD COMPUTING TECH CO LTD

Patent Information

Application Number
CN202510820510.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing open source large models cannot meet user needs in specific industry applications, resulting in poor image generation results, and complex training process and high cost, making it difficult to achieve efficient and low-threshold model optimization.

Method used

Through the modularly designed sample generation components and model training components, the model training task flow is constructed, the visual language model and text processing model are used to generate image prompt text sets, the target sample image set is generated by combining the text-generated graph model, and the initial image processing model is trained through the model training components, and the LoRA model is used for fine-tuning to reduce the training complexity.

Benefits of technology

It realizes efficient and low-threshold training of image processing models that meet the needs in specific industries, simplifies the model training process, reduces technical thresholds and costs, and improves the correlation and quality of generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339756A_ABST
    Figure CN120339756A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image processing model training method and device and an image generation method and device.The method comprises the steps that in response to selection operation for a sample generation assembly and a model training assembly, the sample generation assembly and the model training assembly are connected, and a model training task flow is constructed; obtaining an initial sample image, inputting the initial sample image into a sample generation component, generating an image prompt text set corresponding to the initial sample image through the sample generation component, and generating a target sample image set based on the image prompt text set; and inputting the target sample image set into a model training assembly, training the initial image processing model by using the target sample image set through the model training assembly, and obtaining a trained target image processing model. Through the modular design of the components, the user can use the model training function more conveniently, and the model training complexity is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of artificial intelligence technology, and in particular to a training method for an image processing model, an image generation method and a device. Background Art

[0002] With the development of deep learning technology, image generation has become an important research direction in the field of computer vision in the application scenario of AIGC (artificial intelligence generated content). At present, open source large models are often used for AIGC applications, but open source large models cannot meet the needs of specific industries, so the generated image resources cannot meet user requirements. Therefore, it is necessary to optimize the models of specific industries for open source large models. However, in the process of optimizing the models of specific industries, the time, manpower and material costs of the initial investment are high due to the difficulty of collecting materials and the complexity of the model training process. Therefore, how to train image processing models that meet specific needs more simply and efficiently is an urgent problem that needs to be solved. Summary of the invention

[0003] In view of this, an embodiment of this specification provides a training method for an image processing model. One or more embodiments of this specification also relate to an image generation method, a training device for an image processing model, an image generation device, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.

[0004] According to a first aspect of an embodiment of this specification, a method for training an image processing model is provided, comprising: In response to a selection operation on a sample generation component and a model training component, connecting the sample generation component and the model training component to construct a model training task flow; Acquire an initial sample image and input it into the sample generation component, generate an image prompt text set corresponding to the initial sample image through the sample generation component, and generate a target sample image set based on the image prompt text set; The target sample image set is input into the model training component, and the initial image processing model is trained by the model training component using the target sample image set to obtain a trained target image processing model.

[0005] According to a second aspect of the embodiments of this specification, there is provided an image generation method, including: Determine an image prompt word associated with the target industry and input it into a target image processing model, wherein the target image processing model is trained by the training method of the above-mentioned image processing model; A target image corresponding to the image prompt word output by the target image processing model is obtained, wherein the target image belongs to the target industry.

[0006] According to a third aspect of the embodiments of the present specification, a training device for an image processing model is provided, including: A connection module, configured to connect the sample generation component and the model training component in response to a selection operation for the sample generation component and the model training component, and construct a model training task flow; A generation module, configured to obtain an initial sample image and input it into the sample generation component, generate an image prompt text set corresponding to the initial sample image through the sample generation component, and generate a target sample image set based on the image prompt text set; A training module, configured to input the target sample image set into the model training component, and train an initial image processing model by using the target sample image set through the model training component to obtain a target image processing model that has completed training.

[0007] According to a fourth aspect of the embodiments of the present specification, an image generation device is provided, including: An input module, configured to determine an image prompt word associated with a target industry and input it into a target image processing model, where the target image processing model is obtained by training through the above-mentioned training method of the image processing model; An output module, configured to obtain a target image corresponding to the image prompt word output by the target image processing model, where the target image belongs to the target industry.

[0008] According to a fifth aspect of the embodiments of the present specification, a computing device is provided, including: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned training method of the image processing model and the image generation method are implemented.

[0009] According to a sixth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer-executable instructions. When the instructions are executed by a processor, the steps of the above-mentioned training method of the image processing model and the image generation method are implemented.

[0010] According to a seventh aspect of the embodiments of the present specification, a computer program product is provided, including a computer program or instructions. When the computer program or instructions are executed by a processor, the steps of the above-mentioned training method of the image processing model and the image generation method are implemented.

[0011] One embodiment of this specification realizes constructing a model training task flow by selecting a sample generation component and a model training component. Through component modular design, users can use the model training function more conveniently. The sample generation component generates an image prompt text set corresponding to the initial sample image, and generates a target sample image set based on the image prompt text set, achieving the purpose of automatic generalization and expansion based on a small number of samples, and obtaining a target sample image set for model training. Then, the model training component uses the target sample image set to train the initial image processing model, enabling users to more conveniently implement model training and obtaining a trained target image processing model. Description of the Drawings

[0012] Figure 1 Fig. shows a schematic diagram of a scenario of a method for training an image processing model provided by an embodiment of this specification; Figure 2 Fig. is a flowchart of a method for training an image processing model provided by an embodiment of this specification; Figure 3 Fig. is a flowchart of a processing procedure of a method for training an image processing model provided by an embodiment of this specification; Figure 4 Fig. is a schematic structural diagram of an apparatus for training an image processing model provided by an embodiment of this specification; Figure 5 Fig. is a flowchart of an image generation method provided by an embodiment of this specification; Figure 6 Fig. is a schematic structural diagram of an image generation apparatus provided by an embodiment of this specification; Figure 7 Fig. is a block diagram of the structure of a computing device provided by an embodiment of this specification. Detailed Embodiments

[0013] In the following description, numerous specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of this specification. Therefore, this specification is not limited by the specific embodiments disclosed below.

[0014] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used in one or more embodiments of this specification refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0015] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0016] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.

[0017] First, the noun terms involved in one or more embodiments of this specification are explained.

[0018] Component: In a computer user interface, a "component" refers to any element with which a user can directly interact. These elements allow users to input data, select options, or trigger certain actions. Components can be graphical (such as buttons, text boxes, checkboxes, etc.) or in more abstract forms (such as voice commands). Buttons are typically used to perform an operation or open a new interface; text boxes allow users to input or edit text information; checkboxes are used to represent a boolean value (yes / no), and users can select or deselect them; dropdown lists provide a list of options for users to choose from, usually only showing one selected value; list boxes display multiple options, and users can select one or more items; labels display static text, usually used to explain the functions of other components and can also be used as a button; scrollbars enable users to navigate through large datasets, such as browsing long documents or lists; images display static pictures and can sometimes also be used as buttons; comboboxes combine the functions of text boxes and dropdown lists, and users can both input text and select from the list. These components usually come with standard styles provided by the operating system or development toolkits, and developers can customize their appearance and behavior according to needs. Different operating systems and programming languages (such as Java, C#, Python, etc.) have their own set of component collections and provide corresponding APIs to create and manage these components.

[0019] In response to: "In response to" generally indicates a reaction or response to a certain situation, problem, request, or event. In the technical field, it can refer to the reaction of a system to input information or events.

[0020] Trigger operation: It refers to automatically performing a certain behavior or response when a specific condition or event occurs. Trigger operations on components mainly refer to the events and behaviors triggered after users interact with interface components (such as buttons, text boxes, dropdown menus, etc.). These trigger operations can include the following types: 1. Click event: The user clicks a button, link, or icon. 2. Mouse over / mouse enter event: The mouse pointer moves over and hovers above the component. 3. Text input event: The user inputs or modifies text in a text box. 4. Selection event: The user changes the options in a dropdown menu or selection box. 5. Focus / blur event: The component gains or loses input focus, etc.

[0021] Text-to-image: An AI-based generation technology that automatically generates an image that matches the described content by inputting a natural language description (such as a text prompt).

[0022] ComfyUI: A modular and visual AI workflow construction tool, mainly used for building and managing complex AI inference processes (such as text-to-image, image-to-image, etc.). Through a node-based operation interface, it enables users to flexibly combine different functional modules (such as model loading, parameter adjustment, post-processing, etc.), thereby achieving highly customized AI generation tasks.

[0023] LoRA: A lightweight method for fine-tuning large models, aiming to significantly reduce the number of parameters and computing resources required for fine-tuning through low-rank factorization technology.

[0024] AIGC has given rise to a large number of innovative applications. Although open-source large models have universality and can meet the needs of a wide range of scenarios, when focused on specific industries, the generated image effects are often not fine enough to fully meet the requirements of actual projects. However, optimizing the model effects for specific industries faces many challenges: Material collection requires a large amount of time and resources, the training environment setup is complex and has high hardware requirements, and the process of initiating training involves a series of technical problems such as hyperparameter tuning. These factors often trap us in a dilemma of high cost, long cycle, and high technical threshold when applying open-source large models to actual projects, and there is an urgent need for more efficient and user-friendly solutions to reduce the difficulty of industry implementation.

[0025] Based on this, in this specification, a training method for an image processing model is provided. This specification also relates to an image generation method, a training device for an image processing model, an image generation device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.

[0026] See Figure 1 , Figure 1 shows a schematic scenario diagram of a training method for an image processing model provided according to an embodiment of this specification, where the method can be applied to a visual workload construction tool, and users can build a model training task flow through this tool. Figure 1The sample generation component and the model training component can be understood as module components configured in the building tool. Users can directly call the module components to build a task flow. After building the model training task flow, users can directly use the model training task flow to train the model, so as to provide users with a process-based and simplified model training method and reduce the difficulty of use for users. Specifically, as a node in the model training task flow, the sample generation component is used to perform data enhancement and expansion based on a small amount of sample input data, solving the problem of difficult material collection in special industries. After generating the target sample image set through the sample generation component, the model training component can use the target sample image set to train the model, thereby obtaining an image processing model that meets the needs of special industries. Based on this, the training method of the image processing model provided in this specification can provide users with a model training method with convenient operation and low training threshold, and can perform data enhancement through the sample generation component to obtain training samples that conform to industry characteristics. Subsequently, the model training component automatically completes the model training, simplifies the model training process, and through modular component calls, subsequent plug-and-play and the addition of other modular components can also be realized according to different user needs to meet different scenario requirements.

[0027] See Figure 2 , Figure 2 which shows a flowchart of a method for training an image processing model according to an embodiment of this specification, specifically including the following steps.

[0028] Step 202: In response to a selection operation for the sample generation component and the model training component, connect the sample generation component and the model training component to construct a model training task flow.

[0029] Among them, the selection operation for the sample generation component and the model training component can be understood as the operation of a user selecting components in the component library of the building tool when using a visual workflow building tool. In order to solve the problem of reducing the difficulty of using the user's model training, modular components are provided for the user in the building tool. For example, the building tool can be ComfyUI, and the modular component can be a ComfyUI node component. The user can create a building project, select the modular components to be used in the project, and connect the selected components according to the user's selection operation, thereby constructing a model training task flow.

[0030] In practical applications, for the selection operations of the sample generation component and the model training component, after selecting the sample generation component and the model training component, the sample generation component and the model training component can be connected to construct a model training task flow. Among them, the sample generation component is used to expand sample data for users, so as to solve the problem of difficult material collection in special industries. The model training component is used to automatically train models for users, so as to solve the problem of complex model training processes and step operation procedures. Therefore, in the case where the user has the above two problem requirements, a model training task flow can be constructed by selecting the sample generation component and the model training component, so as to realize the execution of the model training task with one click. Correspondingly, in the case where the user already has a large amount of training materials, the sample generation component can also not be selected, but the model training component can be used alone for model training. The user can directly input the training materials prepared by himself into the model training component for model training, so as to achieve the purpose of automatic model training. It should be noted that the model training task flow can also include modular components that meet other user requirements, such as model evaluation components, visualization components, etc. The model evaluation component can be used to evaluate the training effect of the trained image processing model, and can also monitor the performance indicators of the model during the training process; the visualization component can be used to generate an intuitive training report to help users quickly locate training problems and optimize the model. In addition, users can also implement corresponding requirements through modular components with other functions provided in the building tool, or design modular components that meet their own requirements through the building tool. Therefore, based on the use of the building tool, a pluggable and selectable task flow building method is provided for users to meet the implementation of model training in different scenarios and different requirements.

[0031] In specific implementation, the model training task flow can be a task flow including sample generation ability and model training ability. Users can also continue to select and add components on the basis of this model training task flow to obtain a task flow with more perfect functions. Correspondingly, the model training task flow selected and constructed by the user this time can be saved in the task flow database, so that the user can directly call the model training task flow subsequently, thereby reducing the step operation of component selection, further improving the model training efficiency, and improving the user experience.

[0032] Step 204: Obtain the initial sample image and input it into the sample generation component, generate an image prompt text set corresponding to the initial sample image through the sample generation component, and generate a target sample image set based on the image prompt text set.

[0033] Among them, the initial sample image can be understood as a small number of sample images provided by the user. The initial sample image can be one or more material images related to a project or industry. Taking the initial sample image as the input of the sample generation component, an image prompt text set corresponding to the initial sample image is generated through the sample generation component. The image prompt texts in the image prompt text set can be understood as prompt word description texts generated by expanding based on the initial sample image. The image prompt text set can be used for subsequent input into the text-to-image large model to generate the corresponding target sample image set. The target sample image set can be understood as a set of target sample images. The target sample image is the sample image obtained after sample generation based on the initial sample image. The target sample image can be used for subsequent model training to train an image processing model that conforms to the characteristics of the project or industry.

[0034] In practical applications, the user can select material images related to the industry that meet the requirements as the initial sample image. For example, if the user wants to train an image processing model specifically for generating second-generation images, some existing material images of second-generation images can be selected as the initial sample image. After inputting the initial sample image into the sample generation component, more target sample images that meet the user's requirements can be generated through the sample generation component to form the target sample image set.

[0035] Furthermore, in order to be able to produce more target sample images that meet the user's requirements in the future, when inputting an image into the sample generation component, a related theme text can also be input together to clarify the image boundary. Specifically, obtaining the initial sample image and inputting it into the sample generation component includes: in response to a trigger operation for the sample generation component, determining a first image directory path and an image theme text; obtaining the initial sample image according to the first image directory path, and inputting the initial sample image and the image theme text into the sample generation component.

[0036] Among them, the triggering operation for the sample generation component can be understood as the user's operation of selecting input data for the sample generation component through clicking, selecting, etc. For example, by clicking on the image selection control in the sample generation component, the first image directory path is determined on the popped-up file selection page. The first image directory path can be understood as the storage path of the initial sample image. Through the first image directory path, the sample generation component can obtain the initial sample image. Correspondingly, the sample generation component may also include an input box for the theme description text. The user inputs the image theme text in the input box through the triggering operation for the sample generation component, such as a typing operation. The image theme text can be a specific and summary text for describing the initial sample image, and the image theme text is used to strengthen and clarify the boundary of the output image. After subsequently inputting the initial sample image and the image theme text into the sample generation component, the sample generation component can generate a training material image that better meets the user's needs based on the initial sample image and the associated image theme text.

[0037] In practical applications, the user can trigger the sample generation component through operations such as clicking and selecting in the user interface of the model training task flow. For example, the user can click on the image selection control to open the file browser and select the first image directory path containing the initial sample image from the file browser. At the same time, the user can also directly input the image theme text into the sample generation component through a specially designed input box to clarify the theme or boundary conditions of the image, such as a specific style, scene description, or object features to be emphasized, etc. Input the obtained initial sample image and the image theme text into the sample generation component together. This process can be automatically completed by a background script or designed to be manually started after the user confirms.

[0038] Based on this, through the user's triggering operation, the initial sample image and the image theme text are input into the sample generation component, providing an intuitive operation interface for the user to easily select images and input theme texts, simplifying the entire process from data preparation to model application, and reducing the technical threshold. Allowing the user to directly guide the direction of image generation by inputting the theme text enhances the control over the output content and helps ensure that the generated images are more in line with the actual needs. By combining the image with the theme text, the relevance between the generated image and the expected target can be effectively improved, and the probability of irrelevant or inappropriate images appearing can be reduced. Whether it is for artistic creation requiring a specific style or image generation tasks in professional fields, this method can flexibly adapt to different application scenarios.

[0039] Further, after processing the user input, sample generation operations can be performed based on the sample generation component, thereby realizing sample augmentation for the user based on a small number of samples. Specifically, an image prompt text set corresponding to the initial sample image is generated through the sample generation component, including: calling a vision-language model through the sample generation component, inputting the initial sample image and the image theme text into the vision-language model, and obtaining the image description text corresponding to the initial sample image output by the vision-language model; calling a text processing model through the sample generation component, inputting the image description text into the text processing model, and obtaining the image prompt text set corresponding to the image description text output by the text processing model, where the text processing model is used to perform augmentation processing on the image description text.

[0040] Among them, the vision-language model can be understood as a model that reversely converts image materials into image descriptions. For example, the vision-language model can be the JoyCaption large model. After inputting the initial sample image and the image theme text into the vision-language model, the vision-language model can reversely output the image description text for the initial sample image. The text processing model can be understood as a model used to supplement or expand based on the image description text. The text processing model can be the LLM (Large Language Model) large language model. Based on the text processing ability of the text processing model, the image description text is supplemented or expanded from multiple dimensions (such as perspective, scene, details), so as to obtain a certain number of image prompt texts, and an image prompt text set is generated according to the image prompt texts.

[0041] In practical applications, by combining the vision-language model and the text processing model, a high-quality image prompt text set is generated, which is further used to generate training material images. This method can not only extract meaningful description information from the initial sample image, but also enrich and expand these descriptions through the large language model (LLM) to better guide the subsequent image generation process.

[0042] In specific implementation, the initial sample image and the image theme text are passed as inputs to the vision-language model. The vision-language model analyzes the image content and generates a descriptive image description text in combination with the provided theme text. The image description text should accurately reflect the main elements, scenes, and the user-specified theme of the image as much as possible. The image description text is input into the text processing model, and the text processing model performs multi-dimensional supplementation or expansion on the input image description text, including but not limited to perspective changes, scene expansion, detail addition, etc. Subsequently, according to the requirements of different text-to-image models, the description text can be adjusted and optimized to generate multiple versions of the image prompt text, forming an image prompt text set. The description extracted by the vision-language model plus the expansion of the text processing model can ensure that the generated image prompt text is more specific and diverse, which helps to improve the quality and relevance of the finally generated images.

[0043] In a specific embodiment of this specification, the initial sample image shows a cat sunning itself by the window, and the theme text input by the user is "warm family life". After inputting the initial sample image and the image theme text into the vision-language model, the vision-language model can output the image description text "A lazy orange cat is lying comfortably on the windowsill in the living room, soaking up the sun as the sunlight filters through the curtains, filling the whole room with a warm atmosphere." After inputting the image description text into the text processing model, the text processing model can output prompt texts containing multiple different angles and details, such as "On a sunny afternoon, an orange cat is lying on a wooden windowsill, and outside the window is a lush garden", "Inside the family living room, a kitten is enjoying the afternoon sun, with a neatly arranged bookshelf and a soft sofa in the background", "In a warm family atmosphere, a cat is taking a nap by the window, surrounded by green plants and family photos". Subsequently, each image prompt text in the image prompt text set can be input into the corresponding text-to-image model according to the model requirements, so as to obtain multiple target sample images generated by the text-to-image model, realizing image data augmentation.

[0044] Based on this, the sample generation component provides a convenient and fast sample generation method for users through the vision-language model combined with the text processing model, enabling users to obtain training material images that meet specific requirements, facilitating better training of the model subsequently, and meeting user needs.

[0045] Further, after obtaining the image prompt text set, the corresponding target sample image set can be generated using the image prompt text set. Specifically, generating the target sample image set based on the image prompt text set includes: calling at least one text-to-image model through the sample generation component, inputting the image prompt text set into the at least one text-to-image model, and obtaining the first sample image set output by the at least one text-to-image model; calculating the image correlation values corresponding to the first sample images in the first sample image set, and screening the first sample images in the first sample image set according to the image correlation values to obtain the target sample image set.

[0046] Among them, after generating the image prompt text set, the sample generation component can call one or more text-to-image models to perform text-to-image generation (complementing each other) in a certain batch, so as to obtain the first sample image set output by at least one text-to-image model. The first sample images in the first sample image set may have a certain gap from the initial sample images or the image theme text, that is, they do not meet the user's requirements or the characteristics of the project or industry. Therefore, the image correlation values of the first sample images in the first sample image set can be calculated, and image screening can be performed using the image correlation values to obtain the screened target sample image set.

[0047] In practical applications, when inputting the image prompt text set into the text-to-image model, since the prompt words of different series of large models are different because their training methods and architectures are different, different prompt words will be generated for different series, that is, the image prompt texts in the image prompt text set are generated according to the model requirements of the text-to-image model. Therefore, when inputting, it can be input according to the corresponding relationship between the image prompt text and the text-to-image model. After obtaining the first sample image set output by the text-to-image model, perform correlation detection on the generated first sample images, remove the sample images that do not meet the requirements, use the remaining sample images as the target sample images, and construct the target sample image set. Correspondingly, the user can also manually select or remove from the first sample image set to retain the target sample images that better meet the user's needs. In addition, if the number of first sample images output by the text-to-image model is insufficient, the generation of the image prompt text set and the generation of the first sample image set can also be repeated to meet the corresponding image quantity conditions.

[0048] In a specific embodiment of this specification, the sample generation component invokes one or more text-to-image models (such as the Flux series, SD series, etc.), and according to the characteristics and requirements of each model, inputs the corresponding image prompt text into the corresponding model. Each text-to-image model will generate a batch of images based on the input prompt words to form a first sample image set. Calculate the image correlation value of each image in the first sample image set with the initial sample image or the image theme text, and use the image correlation value to screen the first sample image set. Images that do not meet the threshold conditions will be removed, and those images that are highly relevant to the original theme or the initial sample image will be retained, thereby obtaining the target sample image set.

[0049] Based on this, by using the generated image prompt text set, the text-to-image model is used to generate the first sample image set that meets the theme requirements, and the image correlation value is used for image screening to obtain the target sample image set, ensuring that the target sample images in the target sample image set are closer to the actual needs of the user or the requirements of a specific project scenario.

[0050] Correspondingly, in order to make the generated sample images closer to the actual needs of the user or the requirements of a specific project scenario, it is necessary to calculate the image correlation value of the first sample images for image screening. Specifically, calculating the image correlation value corresponding to the first sample image in the first sample image set includes: determining the first sample image in the first sample image set and the first image prompt text corresponding to the first sample image, where the first image prompt text is input into the image prompt text set; encoding the first sample image and the first image prompt text to obtain an image embedding and a text embedding, and calculating a first image correlation value based on the image embedding and the text embedding; inputting the first sample image and the image question text associated with the first sample image into a visual question answering model to obtain the image answer text corresponding to the first sample image output by the visual question answering model, and calculating a second image correlation value based on the image answer text; calculating the image correlation value corresponding to the first sample image according to the first image correlation value and the second image correlation value.

[0051] Among them, the first sample image can be understood as the image output by the text-to-image model, and the first image prompt text corresponding to the first sample image can be understood as the image prompt text used to generate the first sample image. After determining the first sample image and the corresponding first image prompt text, the first sample image and the first image prompt text can be encoded to obtain an image embedding and a text embedding, and a first image-related value can be calculated based on the image embedding and the text embedding. Subsequently, a second image-related value can also be calculated in combination with the image content. Specifically, the first sample image and the image question text can be input into a visual question answering model. The image question text can be the image question text generated based on the image theme, the initial sample image, or the image prompt word text. After inputting the first sample image and the image question text into the visual question answering model, an image answer text output by the visual question answering model for the first sample image can be obtained, and a second image-related value can be calculated based on the image answer text. Finally, the image-related value of the first sample image can be calculated according to the first image-related value and the second image-related value. Correspondingly, any first sample image in the first sample image set can be calculated for the image-related value in the above manner.

[0052] In practical applications, the image-related value can be calculated according to a set feedback function. The specific logic of the feedback function includes using a vision model to perform multi-dimensional inspection scoring, specifically including correlation score calculation and fine-grained score calculation. Use the vision-language model CLIP (Contrastive Language-Image Pre-Training) to calculate the embedding vectors of the image and the prompt word, and obtain the global alignment score as the first image-related value through cosine similarity; use the visual question answering model VQA (Visual Question Answering) to generate specific questions related to the first sample image, verify the accuracy of the image details, and calculate the corresponding second image-related value. Finally, the image-related value of the first sample image is calculated according to the first image-related value and the second image-related value.

[0053] In a specific embodiment of this specification, the first sample image and the corresponding first image prompt text "prompt" are input into the vision-language model. The vision-language model encodes them respectively to obtain an image embedding and a text embedding. The global alignment score is calculated through cosine similarity as the first image-related value, such as 0.8. Then, the first sample image and the image question text "Is there xxx in the image?" are input into the visual question-answering model. The visual question-answering model outputs the image answer text corresponding to the image question text "There is xxx in the image". If the model outputs an affirmative answer, the second image-related value can be a fixed bonus value, such as 0.1. Then, the first image-related value and the second image-related value are added together to obtain the final image-related value, such as 0.9. In addition, the first image-related value and the second image-related value can also be calculated according to a certain weight to obtain the image-related value. It should be noted that the image question text can be generated by the visual question-answering model based on the image prompt text of the first sample image, or can be determined by the user manually inputting. There can be multiple image question texts, and the second image-related value can be accumulated according to the image answer text corresponding to the image question text. If each image question text corresponds to an affirmative answer, the bonus values corresponding to each image answer text can be superimposed to obtain the final second image-related value, and the final image-related value is calculated with the first image-related value. Subsequently, the calculated image-related value can be compared with a preset image-related value threshold to realize the correlation detection and screening of the first sample image, and retain the target sample image that is more in line with the theme.

[0054] Based on this, through the above calculation of the image-related value and the use of the image-related value to perform image screening on the first sample image, it can be ensured that the target sample images in the target sample image set are closer to the actual needs of the user or the requirements of a specific project scenario.

[0055] Step 206: Input the target sample image set into the model training component. The model training component uses the target sample image set to train the initial image processing model to obtain the trained target image processing model, where the initial image processing model includes a target base model and a fine-tuning model.

[0056] Among them, after obtaining the target sample image, the target sample image can be input into the model training component, and the model training component starts to execute the model training task, thereby completing the training of the image processing model and obtaining the trained target image processing model. The initial image processing model can be understood as the image processing model before training. The initial image processing model can be a pre-trained model, that is, it has basic image processing capabilities. However, users need to train an image processing model that meets specific scenario or project requirements. Therefore, it is necessary to perform model training on the initial image processing model based on the target sample image set that conforms to the industry or project characteristics, and obtain the target image processing model with industry or project characteristics, which can meet the user's needs.

[0057] In practical applications, in order to reduce the complex processes, time, and material costs consumed during model training, the method of using a base model plus a fine-tuning model can be adopted for training. During model training, only the adjustment of the model parameters of the fine-tuning model needs to be concerned, reducing the number of parameters to be adjusted, thereby reducing the computational cost and storage requirements. Specifically, when implementing, the fine-tuning model can be a LoRA model. A low-rank matrix of the LoRA model is added to the target base model to capture specific task changes. During the model training process, only the model parameters of the LoRA model need to be optimized and fine-tuned, while the model parameters of the target base model remain unchanged.

[0058] Furthermore, in order to enable the model training component to perform model training based on the target sample image set, it is necessary to input the target sample image set into the model training component. Specifically, inputting the target sample image set into the model training component includes: in response to a trigger operation for the model training component, determining a second image directory path corresponding to the target sample image set; and loading the target sample image set by the model training component based on the second image directory path.

[0059] Among them, the trigger operation for the model training component can be understood as the operation by which the user determines the directory path of the target sample image set for the model training component. In practical applications, the user can click on the model training component to open a file browser for selecting an image directory path, and determine the second image directory path in the file browser. The second image directory path is the image storage path of the target sample image set. Subsequently, the model training component can construct and load the target sample image set based on the second image directory path, thereby realizing the input of the target sample image set into the model training component.

[0060] In practical applications, after the sample generation component generates the target sample image set, the target sample images can be archived in the training data format (while recording the corresponding image descriptions or image prompts), and then output to the second image directory path. When it is necessary to input the target sample image set into the model training component, the second image directory path corresponding to the target sample images can be output to the model training component, so that the model training component can load the sample image set based on the second image directory path, thereby realizing the input of the target sample image set into the model training component.

[0061] In a specific embodiment of this specification, the user can trigger an image selection operation by clicking an interface button or control in the model training component, and select the corresponding second image directory path in the displayed file browser. The model training component will read and load the target sample image set from the specified location based on the second image directory path determined by the user. At the same time, during the reading process, the read data can also be verified and preprocessed, such as standardization and normalization, to ensure data integrity and correctness, and the preprocessed data is cached to avoid repeated calculations. Once the data is loaded and passes the verification, this data can be used for the subsequent model training process.

[0062] Based on this, through effective data management, the use of data becomes more flexible, facilitating data interaction between components, and improving the training efficiency of subsequent model training based on training data.

[0063] Furthermore, to ensure the stability and consistency of the training process, the model training component has a complete environment setup process built into its internal implementation. Specifically, the model training component uses the target sample image set to train the initial image processing model, including: in response to the execution operation for the model training task, creating a training environment corresponding to the model training task through the model training component; loading the target base model and the fine-tuning model in the training environment, generating the initial image processing model according to the target base model and the fine-tuning model, and training the initial image processing model using the target sample image set by executing the model training task.

[0064] Among them, the execution operation for the model training task can be understood as that when the user creates a model training task, the object of this training, that is, the initial image processing model, is selected. After determining the initial image processing model, the user can click the execution control, so that the model training component starts to respond and execute the model training task. After the model training component starts to execute the model training task, an independent virtual environment will be created for the model training task to adapt to different model frameworks, and a unique directory path (installing different special dependencies) will be assigned to each task to store relevant model files, logs, configurations, etc., so as to avoid problems caused by various conflicts. It not only solves the common dependency conflict problems in the model training process, but also provides the user with an isolated and reproducible training environment.

[0065] In practical applications, the training environment can be understood as an independent virtual environment created for the model training task. The training environment will install corresponding dependency packages according to the requirements of the required base model and fine-tuning model to ensure no conflict with other projects or tasks. Load the target base model (such as Stable Diffusion or FLUX series models) and fine-tuning model (such as LoRA model) in the training environment, combine the target base model with the fine-tuning model to generate an initial image processing model available for training; specifically in implementation, the combination of the target base model and the fine-tuning model can be to directly insert the low-rank matrix of the fine-tuning model such as LoRA into the target base model. During the training process, the parameters of the fine-tuning model will be updated, while most of the parameters of the base model remain unchanged and are only adjusted when necessary. Subsequently, by executing the model training task, the initial image processing model can be trained using the target sample image set.

[0066] Based on this, through the above method, the user only needs to input the target sample image set into the model training component, can determine the target base model and fine-tuning model through simple operations, and then can realize automatically executing the model training task by the model training component, and realize the model training of the initial image processing model, reducing the technical threshold and operation complexity of the user in the model fine-tuning process.

[0067] Furthermore, in order to make the model fine-tuning more in line with the user's needs, the user can also select the training parameters by himself. Specifically, before training the initial image processing model using the target sample image set by executing the model training task, the method further includes: responding to the adjustment operation for the model training task, determining the training parameters of the model training task; and executing the model training task according to the training parameters to train the initial image processing model using the target sample image.

[0068] Among them, the adjustment operation for the model training task can be understood as an operation in which the user adjusts the relevant task settings of the model training task. The adjustment operation includes adjusting the execution information of the model training task, training parameter setting, etc. After the user clicks the setting control of the model training task, a corresponding setting page will be displayed. The setting page may include training parameters such as learning rate, batch size, number of training epochs, low-rank dimension, weight freezing strategy, etc.; and relevant execution information of the model training task such as execution start time setting, stop time setting, task interruption recovery setting, etc.

[0069] In practical applications, after the user completes the adjustment of the model training task through the corresponding interface, the training parameters corresponding to the model training task can be determined. The training parameters can be used subsequently when executing the model training task, that is, the model training task is executed according to the training parameters, and the initial image processing model is trained using the target sample images. This enables the user, through the model training component, to simply select the base model and set the training parameters, and then start the training process with one click. After the user completes the above configuration, the model training component will automatically execute the training process, enabling the user to obtain the trained image processing model.

[0070] Based on this, by setting the training parameters for the model training task by the user, the user can complete the model fine-tuning with simple operations, making the training process more flexible and controllable, especially suitable for scenarios where parameters need to be frequently adjusted to achieve good performance. In addition, by providing a friendly user interface, even users without a technical background can easily manage and optimize the model training task.

[0071] Further, training the initial image processing model using the target sample image set includes: inputting the target sample image set and the target prompt text set corresponding to the target sample image set into the initial image processing model to obtain a predicted image set output by the initial image processing model; calculating a model loss value based on the target sample image set and the predicted image set, and using the model loss value to adjust the model parameters of the fine-tuning model, and subsequently continuing to train the initial image processing model until a target image processing model that meets the training conditions is obtained.

[0072] Among them, the target sample image set can be understood as a sample label. After inputting the target prompt text set into the initial image processing model, the initial image processing model will output a predicted image set, which can be understood as a predicted value. The model loss value is calculated by comparing the predicted value with the sample label, that is, the true value, and the model parameters of the fine-tuning model are adjusted using the model loss value, thus completing one round of model training. Subsequently, the trained target image processing model is obtained through iterative training.

[0073] In practical applications, the target prompt texts in the target prompt text set can correspond one-to-one with the target sample images in the target sample image set. The target prompt text can be the image prompt text of the target sample image. After inputting the target prompt text into the initial image processing model, it can guide the model to generate corresponding predicted images. Calculate the model loss value according to the target sample image set and the predicted image set, that is, use the target sample images in the target sample image set as the ground truth values, and the predicted images in the corresponding predicted image set as the predicted values, and calculate the model loss value such as image reconstruction loss, similarity loss, perceptual loss, adversarial loss, etc. The model loss value is used to adjust the model parameters of the fine-tuning model, that is, the LoRA model, and through multiple rounds of iterative training, a target image processing model that meets the training stop conditions such as reaching the training rounds and the model parameters meeting the threshold is obtained.

[0074] Based on this, the above model training component automatically executes the model training task, performs model training on the initial image processing model, enabling users to achieve model fine-tuning with a simple operation process, providing users with a more convenient and fast model training function. It greatly simplifies the implementation process of complex tasks and reduces the technical threshold.

[0075] Further, after obtaining the target image processing model that has completed training, it further includes: determining the target fine-tuning model in the target image processing model, and the training log of the target image processing model; storing the target fine-tuning model and the training log in the model directory.

[0076] Among them, the target fine-tuning model can be understood as the optimized fine-tuning model obtained after the training task is completed. The training log of the target image processing model can be understood as the log information recorded during the execution of the model training task. The training log can include information such as training parameters, sample data information, loss value changes, and model structure. After storing the target fine-tuning model and the training log in the model directory, the trained target fine-tuning model can be directly loaded based on the model directory later, and the target fine-tuning model can be combined with the target base model, so that users can directly use the target image processing model that meets specific industry scenarios. And users can also read out the training log of the target fine-tuning model from the model directory to understand the training process of the target fine-tuning model, and can also use analysis tools to perform training analysis based on the training log to determine the training problems existing in the target fine-tuning model during the training process, which is convenient for users to locate problems and make quick adjustments.

[0077] In a specific embodiment of this specification, after obtaining the target image processing model that has completed training, the model weights of the target fine-tuning model, i.e., the optimized LoRA model, in the target image processing model and the corresponding training logs are saved to a specified model directory, so as to facilitate the use of the target image processing model based on the target fine-tuning model.

[0078] A training method for an image processing model provided in this specification includes, in response to a selection operation for a sample generation component and a model training component, connecting the sample generation component and the model training component to construct a model training task flow; obtaining an initial sample image and inputting it into the sample generation component, generating an image prompt text set corresponding to the initial sample image through the sample generation component, and generating a target sample image set based on the image prompt text set; inputting the target sample image set into the model training component, and training an initial image processing model using the target sample image set through the model training component to obtain a target image processing model that has completed training, where the initial image processing model includes a target base model and a fine-tuning model. It realizes constructing a model training task flow by selecting a sample generation component and a model training component, and enables users to use the model training function more conveniently through component modular design. By generating an image prompt text set corresponding to the initial sample image through the sample generation component and generating a target sample image set based on the image prompt text set, the purpose of automatic generalization and expansion based on a small number of samples is achieved, and a target sample image set for model training is obtained. Then, through the model training component, the initial image processing model is trained using the target sample image set, and the initial image processing model includes a target base model and a fine-tuning model. The model training can be completed only by optimizing the fine-tuning model, reducing the complexity of model training.

[0079] The following combines the attached Figure 3 , taking the application of the training method of the image processing model provided in this specification in an e-commerce scenario as an example, to further illustrate the training method of the image processing model. Among them, Figure 3 FIG. shows a process flow chart of a training method for an image processing model provided in an embodiment of this specification, specifically including the following steps.

[0080] Step 302: In response to a selection operation for a sample generation component and a model training component, connect the sample generation component and the model training component to construct a model training task flow.

[0081] In one embodiment, an e-commerce platform hopes to provide merchants with an automated tool that can quickly generate product display images (such as different angles, backgrounds, matching scenarios, etc.) that conform to the brand style based on a small number of provided product images and style descriptions, for enriching the content of product detail pages and improving conversion rates. However, the product images directly generated using open-source text-to-image models often do not meet the brand tone or product detail requirements. Therefore, it is necessary to fine-tune the model so that it can better restore product features and match specific brand styles (such as minimalist style, retro style, technological sense, etc.), and support multiple display scenarios (white background, life scenario, model wearing, etc.). Merchants can select a sample generation component and a model training component to construct a model training task flow.

[0082] Step 304: In response to a trigger operation for the sample generation component, determine a first image directory path and an image theme text, obtain initial sample images according to the first image directory path, and input the initial sample images and the image theme text into the sample generation component.

[0083] In one embodiment, a merchant can click on the sample generation component to determine the first image directory path of the initial sample images. The initial sample images can be flat lay images or hanging shot images of "white shirts" products. The image theme text can be descriptive texts such as "simple business, workplace dressing, paired with dress pants and leather shoes, background is an office", etc. Input the initial sample images and the image theme text into the sample generation component.

[0084] Step 306: Call a vision-language model through the sample generation component, input the initial sample images and the image theme text into the vision-language model, and obtain an image description text corresponding to the initial sample images output by the vision-language model.

[0085] In one embodiment, call a vision-language model through the sample generation component, and use the vision-language model for reverse description to generate an image description text.

[0086] Step 308: Call a text processing model through the sample generation component, input the image description text into the text processing model, and obtain an image prompt text set corresponding to the image description text output by the text processing model.

[0087] In one embodiment, call a text processing model through the sample generation component, and use the text processing model for text expansion to generate an image prompt text set.

[0088] Step 310: Call at least one text-to-image model through the sample generation component, input the image prompt text set into the at least one text-to-image model, and obtain a first sample image set output by the at least one text-to-image model.

[0089] In one embodiment, the image prompt texts in the image prompt text set are input into multiple text-to-image models, and the text-to-image models are used for image generation, so as to obtain a diversified first sample image set.

[0090] Calculate the image-related values corresponding to the first sample images in the first sample image set, screen the first sample images in the first sample image set according to the image-related values, and obtain a target sample image set. Determine the first sample images in the first sample image set and the first image prompt texts corresponding to the first sample images; encode the first sample images and the first image prompt texts to obtain image embeddings and text embeddings, and calculate the first image-related values based on the image embeddings and the text embeddings; input the first sample images and the image question texts associated with the first sample images into a visual question answering model, obtain the image answer texts corresponding to the first sample images output by the visual question answering model, and calculate the second image-related values based on the image answer texts; calculate the image-related values corresponding to the first sample images according to the first image-related values and the second image-related values.

[0091] In one embodiment, the CLIP model is used to evaluate the relevance of each generated image to the original theme, screen out high-quality images, and construct a target sample image set.

[0092] Step 312: In response to a trigger operation for the model training component, determine the second image directory path corresponding to the target sample image set, and load the target sample image set based on the second image directory path.

[0093] In one embodiment, the user clicks on the image selection control in the model training component to determine the second image directory path of the target sample image set, so that the model training component loads the target sample image set.

[0094] Step 314: Use the target sample image set to train the initial image processing model through the model training component.

[0095] In response to an adjustment operation for the model training task, determine the training parameters of the model training task; execute the model training task according to the training parameters, and use the target sample images to train the initial image processing model. In response to an execution operation for the model training task of the initial image processing model, create a training environment corresponding to the model training task through the model training component; load the target base model and the fine-tuning model in the training environment, generate an initial image processing model according to the target base model and the fine-tuning model, and train the initial image processing model by executing the model training task using the target sample image set.

[0096] Input the target sample image set and the target prompt text set corresponding to the target sample image set into the initial image processing model to obtain a predicted image set output by the initial image processing model; calculate the model loss value based on the target sample image set and the predicted image set, and use the model loss value to adjust the model parameters of the fine-tuning model, and continue with the initial image processing model to obtain a trained target image processing model.

[0097] In one embodiment, the user selects a base model and sets training parameters. The system responds to the corresponding operation and creates a training environment for the model training task, loads the target sample image set and the corresponding image prompt text set, inputs the image prompt text set into the initial image processing model to generate predicted images, and calculates the model loss value using the target samples and the predicted images for backpropagation and updates the LoRA module parameters.

[0098] Step 316: Determine the target fine-tuning model in the target image processing model and the training log of the target image processing model; store the target fine-tuning model and the training log in the model directory.

[0099] In one embodiment, after multiple rounds of model training iterations, save the finally trained LoRA model and its training log.

[0100] A training method for an image processing model provided in this specification realizes constructing a model training task flow by selecting a sample generation component and a model training component. Through component modular design, users can use the model training function more conveniently. The sample generation component generates an image prompt text set corresponding to the initial sample image, and based on the image prompt text set, a target sample image set is generated to achieve the purpose of automatic generalization and expansion based on a small number of samples, and a target sample image set for model training is obtained. Then, the model training component uses the target sample image set to train the initial image processing model, and the initial image processing model includes a target base model and a fine-tuning model. The model training can be completed only by optimizing the fine-tuning model, reducing the complexity of model training.

[0101] Corresponding to the above method embodiment, this specification also provides an embodiment of a training device for an image processing model. Figure 4 The structural schematic diagram of a training device for an image processing model provided by an embodiment of this specification is shown. As Figure 4 shown, the device includes: A connection module 402, configured to connect the sample generation component and the model training component in response to a selection operation for the sample generation component and the model training component, and construct a model training task flow; A generation module 404, configured to obtain an initial sample image and input it into the sample generation component, generate an image prompt text set corresponding to the initial sample image through the sample generation component, and generate a target sample image set based on the image prompt text set; A training module 406, configured to input the target sample image set into the model training component, and train an initial image processing model by using the target sample image set through the model training component to obtain a trained target image processing model.

[0102] Optionally, the generation module 404 is configured to, in response to a trigger operation for the sample generation component, determine a first image directory path and an image theme text; obtain an initial sample image according to the first image directory path, and input the initial sample image and the image theme text into the sample generation component.

[0103] Optionally, the generation module 404 is configured to call a vision-language model through the sample generation component, input the initial sample image and the image theme text into the vision-language model, and obtain an image description text corresponding to the initial sample image output by the vision-language model; call a text processing model through the sample generation component, input the image description text into the text processing model, and obtain an image prompt text set corresponding to the image description text output by the text processing model, where the text processing model is used to perform an expansion process on the image description text.

[0104] Optionally, the generation module 404 is configured to call at least one text-to-image model through the sample generation component, input the image prompt text set into the at least one text-to-image model, and obtain a first sample image set output by the at least one text-to-image model; calculate an image correlation value corresponding to a first sample image in the first sample image set, and screen the first sample images in the first sample image set according to the image correlation value to obtain a target sample image set.

[0105] Optionally, the generating module 404 is configured to determine a first sample image in the first sample image set and a first image prompt text corresponding to the first sample image, where the first image prompt text is input into the image prompt text set; encode the first sample image and the first image prompt text to obtain an image embedding and a text embedding, and calculate a first image correlation value based on the image embedding and the text embedding; input the first sample image and the image question text associated with the first sample image into a visual question answering model to obtain an image answer text corresponding to the first sample image output by the visual question answering model, and calculate a second image correlation value based on the image answer text; calculate an image correlation value corresponding to the first sample image according to the first image correlation value and the second image correlation value.

[0106] Optionally, the training module 406 is configured to, in response to a trigger operation for the model training component, determine a second image directory path corresponding to the target sample image set; load the target sample image set based on the second image directory path through the model training component.

[0107] Optionally, the training module 406 is configured to, in response to an execution operation for a model training task, create a training environment corresponding to the model training task through the model training component; load a target base model and a fine-tuning model in the training environment, generate an initial image processing model according to the target base model and the fine-tuning model, and train the initial image processing model by executing the model training task using the target sample image set.

[0108] Optionally, the apparatus further includes an adjustment module, configured to, in response to an adjustment operation for the model training task, determine training parameters of the model training task; execute the model training task according to the training parameters, and train the initial image processing model using the target sample image.

[0109] Optionally, the training module 406 is configured to input the target sample image set and a target prompt text set corresponding to the target sample image set into the initial image processing model to obtain a predicted image set output by the initial image processing model; calculate a model loss value according to the target sample image set and the predicted image set, adjust model parameters of the fine-tuning model using the model loss value, and continue to train the initial image processing model until a target image processing model that meets the training conditions is obtained.

[0110] Optionally, the apparatus further includes a storage module, configured to determine a target fine-tuning model in the target image processing model and a training log of the target image processing model; store the target fine-tuning model and the training log in a model directory.

[0111] A training device for an image processing model provided in this specification includes: a connection module configured to connect the sample generation component and the model training component in response to a selection operation for the sample generation component and the model training component, and construct a model training task flow; a generation module configured to obtain an initial sample image and input it into the sample generation component, generate an image prompt text set corresponding to the initial sample image through the sample generation component, and generate a target sample image set based on the image prompt text set; a training module configured to input the target sample image set into the model training component, and train an initial image processing model by using the target sample image set through the model training component to obtain a target image processing model that has completed training. It realizes constructing a model training task flow by selecting a sample generation component and a model training component, and enables users to use the model training function more conveniently through component modular design. An image prompt text set corresponding to the initial sample image is generated through the sample generation component, and a target sample image set is generated based on the image prompt text set, achieving the purpose of automatic generalization and expansion based on a small number of samples, and obtaining a target sample image set for model training. Then, the initial image processing model is trained by using the target sample image set through the model training component, and the initial image processing model includes a target base model and a fine-tuning model. The model training can be completed only by optimizing the fine-tuning model, reducing the model training complexity.

[0112] The above is a schematic solution of a training device for an image processing model in this embodiment. It should be noted that the technical solution of the training device for the image processing model belongs to the same concept as the technical solution of the above image processing model training method. For the details not described in the technical solution of the training device for the image processing model, reference can be made to the description of the technical solution of the above image processing model training method.

[0113] See Figure 5 , Figure 5 shows a flowchart of an image generation method provided according to an embodiment of this specification, which specifically includes the following steps.

[0114] Step 502: Determine image prompt words associated with the target industry and input them into the target image processing model, where the target image processing model is trained through the above image processing model training method.

[0115] Step 504: Obtain a target image corresponding to the image prompt words output by the target image processing model, where the target image belongs to the target industry.

[0116] In a specific embodiment of this specification, after obtaining the target image processing model through the above training method, the user can input image prompts for managing the target industry into the target image processing model. For example, in an e-commerce scenario, the user can input the prompt "a white dress shirt on a hanger, modern office background, subtle lighting, business casual style" into the target image processing model to generate the product images required by the merchant using the target image processing model. The target image belongs to the target industry, thus meeting the user's needs.

[0117] An image generation method provided in this specification enables the user to generate target images related to the target industry based on image prompts associated with the target industry through the target image processing model, allowing the user to use the target image processing model to generate more images that conform to industry characteristics, eliminating the need for the user to spend human and material resources on designing the target images, and improving the efficiency of target image generation.

[0118] Corresponding to the above method embodiment, this specification also provides an embodiment of an image generation device. Figure 6 The structural schematic diagram of an image generation device provided by an embodiment of this specification is shown. As Figure 6 shown, the device includes: An input module 602, configured to determine image prompts associated with the target industry and input them into the target image processing model, where the target image processing model is obtained by training through the training method of the above image processing model; An output module 604, configured to obtain the target image corresponding to the image prompt output by the target image processing model, where the target image belongs to the target industry.

[0119] An image generation device provided in this specification enables the user to generate target images related to the target industry based on image prompts associated with the target industry through the target image processing model, allowing the user to use the target image processing model to generate more images that conform to industry characteristics, eliminating the need for the user to spend human and material resources on designing the target images, and improving the efficiency of target image generation.

[0120] The above is a schematic solution of an image generation device in this embodiment. It should be noted that the technical solution of this image generation device and the technical solution of the above image generation method belong to the same concept. For the details not described in detail in the technical solution of the image generation device, reference can be made to the description of the technical solution of the above image generation method.

[0121] Figure 7FIG. 0 shows a block diagram of a computing device 700 provided according to an embodiment of the present specification. The components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 via a bus 730, and a database 750 is used to store data.

[0122] The computing device 700 further includes an access device 740, which enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 740 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC).

[0123] In an embodiment of the present specification, the above components of the computing device 700 and Figure 7 other components not shown therein may also be connected to each other, for example, via a bus. It should be understood that Figure 7 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art may add or replace other components as needed.

[0124] The computing device 700 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 700 can also be a mobile or stationary server.

[0125] Wherein, the processor 720 is configured to execute the following computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above-mentioned image processing model training method and image generation method are implemented.

[0126] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solutions of the above-mentioned image processing model training method and image generation method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the descriptions of the technical solutions of the above-mentioned image processing model training method and image generation method.

[0127] This specification also provides a computer-readable storage medium in one embodiment, which stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the steps of the above-mentioned image processing model training method and image generation method are implemented.

[0128] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solutions of the above-mentioned image processing model training method and image generation method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the descriptions of the technical solutions of the above-mentioned image processing model training method and image generation method.

[0129] This specification also provides a computer program product in one embodiment, including a computer program or instructions, and when the computer program or instructions are executed by a processor, the steps of the above-mentioned image processing model training method and image generation method are implemented.

[0130] The above is a schematic solution of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solutions of the above-mentioned image processing model training method and image generation method belong to the same concept. For the details not described in detail in the technical solution of the computer program product, reference can be made to the descriptions of the technical solutions of the above-mentioned image processing model training method and image generation method.

[0131] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0132] The computer instructions include computer program code, which may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0133] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described order of actions, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.

[0134] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0135] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not elaborate on all the details and do not limit the invention to only the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can understand and utilize this specification well.

Claims

1. A method for training an image processing model, comprising: In response to a selection operation for a sample generation component and a model training component, connecting the sample generation component and the model training component to construct a model training task flow; Obtain an initial sample image and input it into the sample generation component, generate an image prompt text set corresponding to the initial sample image through the sample generation component, and generate a target sample image set based on the image prompt text set; Input the target sample image set into the model training component, and train an initial image processing model by using the target sample image set through the model training component to obtain a trained target image processing model.

2. The method according to claim 1, wherein obtaining an initial sample image and inputting it into the sample generation component comprises: In response to a trigger operation for the sample generation component, determine a first image directory path and an image theme text; Obtain an initial sample image according to the first image directory path, and input the initial sample image and the image theme text into the sample generation component.

3. The method according to claim 2, wherein generating an image prompt text set corresponding to the initial sample image through the sample generation component comprises: Call a vision-language model through the sample generation component, input the initial sample image and the image theme text into the vision-language model, and obtain an image description text corresponding to the initial sample image output by the vision-language model; Call a text processing model through the sample generation component, input the image description text into the text processing model, and obtain an image prompt text set corresponding to the image description text output by the text processing model, wherein the text processing model is used to perform an expansion process on the image description text.

4. The method according to claim 1, wherein generating a target sample image set based on the image prompt text set comprises: Call at least one text-to-image model through the sample generation component, input the image prompt text set into the at least one text-to-image model, and obtain a first sample image set output by the at least one text-to-image model; Calculate an image correlation value corresponding to a first sample image in the first sample image set, and screen the first sample images in the first sample image set according to the image correlation value to obtain a target sample image set.

5. The method according to claim 4, wherein calculating an image correlation value corresponding to a first sample image in the first sample image set comprises: Determine a first sample image in the first sample image set and a first image prompt text corresponding to the first sample image, wherein the first image prompt text is input into the image prompt text set; Encode the first sample image and the first image prompt text to obtain an image embedding and a text embedding, and calculate a first image correlation value based on the image embedding and the text embedding; Input the first sample image and the image problem text associated with the first sample image into a visual question answering model, obtain the image answer text corresponding to the first sample image output by the visual question answering model, and calculate a second image-related value based on the image answer text; Calculate the image-related value corresponding to the first sample image according to the first image-related value and the second image-related value.

6. The method according to claim 1, wherein inputting the target sample image set into the model training component comprises: In response to a trigger operation for the model training component, determine a second image directory path corresponding to the target sample image set; Load the target sample image set by the model training component based on the second image directory path.

7. The method according to claim 1, wherein training the initial image processing model by the model training component using the target sample image set comprises: In response to an execution operation for a model training task, create a training environment corresponding to the model training task by the model training component; Load a target base model and a fine-tuning model in the training environment, and generate the initial image processing model according to the target base model and the fine-tuning model; Train the initial image processing model using the target sample image set by executing the model training task.

8. The method according to claim 7, before training the initial image processing model using the target sample image set by executing the model training task, the method further comprises: In response to an adjustment operation for the model training task, determine the training parameters of the model training task; Execute the model training task according to the training parameters, and train the initial image processing model using the target sample image.

9. The method according to claim 7, wherein training the initial image processing model using the target sample image set comprises: Input the target sample image set and the target prompt text set corresponding to the target sample image set into the initial image processing model, and obtain a predicted image set output by the initial image processing model; Calculate a model loss value according to the target sample image set and the predicted image set, adjust the model parameters of the fine-tuning model using the model loss value, and continue to train the initial image processing model until a target image processing model that meets the training conditions is obtained.

10. The method according to any one of claims 1-9, after obtaining the target image processing model that has completed training, further comprising: Determine the target fine-tuning model in the target image processing model and the training log of the target image processing model; Store the target fine-tuning model and the training log in a model directory.

11. An image generation method, comprising: Determine an image prompt word associated with a target industry and input it into a target image processing model, wherein the target image processing model is trained by the image processing model training method according to any one of claims 1-10; Obtain the target image corresponding to the image prompt word output by the target image processing model, where the target image belongs to the target industry.

12. An apparatus for training an image processing model, comprising: A connection module, configured to connect the sample generation component and the model training component in response to a selection operation for the sample generation component and the model training component, and construct a model training task flow; A generation module, configured to obtain an initial sample image and input it into the sample generation component, generate an image prompt text set corresponding to the initial sample image through the sample generation component, and generate a target sample image set based on the image prompt text set; A training module, configured to input the target sample image set into the model training component, and train an initial image processing model by using the target sample image set through the model training component to obtain a target image processing model that has completed training.

13. An image generation apparatus, comprising: An input module, configured to determine an image prompt word associated with the target industry and input it into the target image processing model, where the target image processing model is obtained by training using the image processing model training method according to any one of claims 1-10; An output module, configured to obtain the target image corresponding to the image prompt word output by the target image processing model, where the target image belongs to the target industry.

14. A computing device, comprising: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.

15. A computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

16. A computer program product, comprising a computer program or instruction. When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Image automatic generation method and device based on AIGC, equipment and medium

    CN117496302A

  • Image generation model training method and device and image processing method and device

    CN117745857A

  • Sample generation method and device, model training method and device and storage medium

    CN119625098A

  • Quality evaluation method and device based on text map and related product

    CN119850538A

Cited By

  • Label generation method of visual equipment, product, electronic equipment and medium

    CN121170364A

  • A label generation method, product, electronic device and medium of a visual device

    CN121170364B