Image processing model training methods, image processing methods

By employing a multi-path closed-loop data generation and multi-stage training strategy, the problems of overfitting and prompt word sensitivity in the image processing model training process were solved, improving the training data quality and editing effect of the model, and achieving more intelligent and unified image processing capabilities.

CN122134869APending Publication Date: 2026-06-02SHUXING TECH (BEIJING) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHUXING TECH (BEIJING) CO LTD
Filing Date
2026-02-28
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing image processing models suffer from overfitting during training, high sensitivity to prompts, inability to uniformly handle different image editing functions, and significant differences in generated objects, resulting in poor editing effects.

Method used

A multi-path closed-loop data generation approach is adopted, which generates high-quality training sample data through expert models, structural control paths, and rendering synthesis paths. Combined with quality cleaning operators for screening, a multi-stage training strategy and dynamic instruction fine-tuning are used to improve the quality of training sample data and the efficiency of model training.

Benefits of technology

It improves the quality and efficiency of training data for image processing models, enhances the model's adaptability to complex instructions and the similarity of object processing, and improves the uniformity and intelligence of editing effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134869A_ABST
    Figure CN122134869A_ABST
Patent Text Reader

Abstract

This specification provides a training method for an image processing model and an image processing method. The training method for the image processing model includes: generating a training sample dataset according to preset data generation rules, wherein the training sample dataset includes training sample images, processing prompt words, and training label images; training an initial image processing model based on the training sample dataset to obtain a reference image processing model; updating the processing prompt words in the training sample dataset to obtain a reference training sample set; and training the reference image processing model based on the reference training sample set to obtain a target image processing model. This method improves the quality of the training sample data in the training sample dataset, effectively improving the training efficiency of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to training methods for image processing models and image processing methods. This specification also relates to a computing device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] With the development of AI (Artificial Intelligence) technology, image editing using AI has become an important branch. Users can issue editing instructions to the image to be edited through a large AI model, which then processes the image to obtain a new one.

[0003] In training large-scale AI models for image editing, overfitting to the training set often relies on prompts. If the user-inputted prompts deviate from the training prompts, the AI ​​model's processing power significantly declines. Furthermore, the training data for these models varies in quality, and different image editing functions require different models, making a single model unsuitable. Additionally, when editing objects in an image (such as faces or animals), the model regenerates new objects, leading to significant differences between the edited and original images. Therefore, training a more intelligent image editing model with better results has become a pressing technical challenge for engineers. Summary of the Invention

[0004] In view of this, embodiments of this specification provide a training method for an image processing model and an image processing method. This specification also relates to a computing device, a computer-readable storage medium, and a computer program product to address the aforementioned problems in the prior art.

[0005] According to a first aspect of the embodiments of this specification, a method for training an image processing model is provided, comprising: A training sample data set is generated according to a preset data generation rule, wherein the training sample data set includes training sample images, processing prompt words, and training label images; An initial image processing model is trained based on the training sample data set to obtain a reference image processing model; Update the processing prompt words in the training sample data set to obtain a reference training sample set, and train the reference image processing model based on the reference training sample set to obtain the target image processing model.

[0006] According to a second aspect of the embodiments of this specification, an image processing method is provided, comprising: Obtain at least one image to be processed and an image processing prompt; The at least one image to be processed and the image processing prompt are input into the target image processing model, wherein the target image processing model is obtained by the training method of the image processing model described above; Obtain the target image output by the target image processing model.

[0007] According to a third aspect of the embodiments of this specification, an image processing method is provided, comprising: At least one image to be processed and an image processing prompt word sent by the receiving device; The at least one image to be processed and the image processing prompt are input into the target image processing model, wherein the target image processing model is obtained by the training method of the image processing model described above; Obtain the target image output by the target image processing model and send the target image to the end device.

[0008] According to a fourth aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the above method.

[0009] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.

[0010] According to a sixth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.

[0011] The image processing model training method provided in this specification includes generating a training sample data set according to preset data generation rules, wherein the training sample data set includes training sample images, processing prompt words, and training label images; training an initial image processing model based on the training sample data set to obtain a reference image processing model; updating the processing prompt words in the training sample data set to obtain a reference training sample set, and training the reference image processing model based on the reference training sample set to obtain a target image processing model.

[0012] The method provided in the embodiments of this specification employs a multi-path closed-loop data generation approach. It generates a sufficient number of training sample data through methods including expert model generation paths, structural control paths, rendering and compositing paths, and long-tail task completion. Then, a quality cleaning operator is used to filter and refine the generated training sample data, further improving the quality of the training sample dataset and providing a data foundation for subsequent image processing model training. This method enables the low-cost, large-scale generation of high-quality, high-difficulty editing data, solving the problem of training sample data shortage. A multi-stage training strategy is employed in the model training phase, effectively improving the training efficiency of the model. Attached Figure Description

[0013] Figure 1 This is a flowchart of a training method for an image processing model provided in one embodiment of this specification; Figure 2 This is a schematic diagram illustrating the grouping of various image processing tasks according to an embodiment of this specification; Figure 3 This is a schematic diagram illustrating the calculation of loss value based on object identification information provided in one embodiment of this specification; Figure 4 This is a flowchart illustrating an image processing method provided in one embodiment of this specification. Figure 5 This is a schematic diagram of the structure of a training device for an image processing model provided in one embodiment of this specification; Figure 6 This is a schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification; Figure 7 This is a flowchart illustrating an image processing method applied to a cloud-side device, provided in one embodiment of this specification. Figure 8 This is an architecture diagram of an image processing system provided in one embodiment of this specification; Figure 9 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation

[0014] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0015] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0016] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0017] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0018] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0019] Content Application Platform: This is an application platform that integrates various types of content. The content application platform provides users with a rich variety of multimedia content, including but not limited to live streaming, on-demand video, audio, text and image information, social interaction, shopping, etc.

[0020] Published Content: This refers to the content service provided to the content application platform. Published content is any content pre-published by the recommending account on the content platform. For example, the published content could be a product recommendation note published by the recommending account on the content application platform. This product recommendation note can include product images, product videos, product descriptions, product advantages and disadvantages, etc. This product recommendation note can be text and image content, text information, video, etc.

[0021] Token: The basic discrete unit for the model to process and measure text, which may correspond to a single character, subword, punctuation mark, or special symbol. The input text is segmented into a token sequence by a token segmenter, and the model encodes, predicts, and generates tokens at the token level. Context length, billing, and computational overhead are usually measured in terms of the number of tokens.

[0022] Attention mechanism: A computational mechanism used in sequence modeling to dynamically assign "attention weights," enabling the model to aggregate information from different positions in the sequence based on relevance when processing the representation of a particular position. It typically manifests as a weighted aggregation of a set of candidate contexts, thereby highlighting content more relevant to the current task and suppressing irrelevant content.

[0023] Transformer: A type of neural network model architecture with attention mechanism at its core, commonly used for sequence modeling tasks such as natural language processing. It models the dependencies between positions in a sequence in parallel through self-attention, and excels at capturing long-range context. A typical structure consists of multiple stacked multi-head self-attention and feedforward networks, combined with residual connections and layer normalization to stabilize training. It usually also introduces positional encoding to represent the order information of tokens.

[0024] Autoregressive Large Language Models (ALMs): A type of language model that generates text sequentially, commonly used for tasks such as dialogue, writing, and question answering. Its basic objective is to predict the next token given the generated content, outputting a sequence progressively from left to right. It typically employs a Transformer structure with causal masking constraints to ensure that the current position can only utilize preceding information, thus achieving continuous and controllable generation. The model acquires general language capabilities through pre-training on large-scale corpora and can be adapted to specific business scenarios through fine-tuning with instructions.

[0025] Diffusion-based large language models typically utilize a Transformer architecture with bidirectional attention, incorporating the "stepwise noise addition-stepwise noise reduction" generation approach of diffusion models into text modeling. This involves progressively refining the generation process using discrete token controls. During inference, the diffusion-based large language model predicts masked tokens, transforming them into tokens with genuine semantic information, and obtaining the final output. Compared to autoregressive large language models, its advantages lie in stronger global consistency and controllable generation capabilities.

[0026] With the development of internet technology, more and more users are using content publishing platforms to share content, and images are an important component of this content. Sometimes, users need to process images when publishing content on these platforms. For example, they might need to remove passersby from an image, add a jacket to someone wearing a short-sleeved shirt, or change the English text on a poster to Chinese.

[0027] With the development of AI (Artificial Intelligence) technology, more and more image processing tools rely on image processing models. However, current image processing models are usually general-purpose models. When training general-purpose models, there is a lack of high-quality training data for specific domains (such as "aging photos" or "changing clothing material to a specific material").

[0028] Image processing models are trained using fixed prompt word templates, which can lead to overfitting. If the user changes the word order of the prompts or inputs more complex commands, the image processing model may exhibit poor command compliance.

[0029] Image processing models require targeted training to perform specific tasks. For example, tasks such as image denoising and super-resolution cannot be completed within a single image processing model. When an image processing model processes images that require the generation of specific text or adjustments to specific objects (such as human or animal faces), problems such as character distortion, layout errors, and the regeneration of specific objects may occur. Therefore, how to construct rich training data for the model and, based on this data, train a unified image processing model to solve various problems in the image processing process has become a pressing technical problem for engineers.

[0030] This specification provides a training method for an image processing model and an image processing method. This specification also relates to a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0031] Figure 1 A flowchart illustrating a training method for an image processing model according to an embodiment of this specification is shown, specifically including the following steps: Step 102: Generate a training sample data set according to preset data generation rules, wherein the training sample data set includes training sample images, processing prompt words, and training label images.

[0032] In practical applications, the quality of training data is crucial during the training of image processing models. However, acquiring training data is costly, and its quality varies significantly, potentially failing to meet the requirements of the desired image processing task. Therefore, to address the problem of insufficient training data, the embodiments provided in this specification offer a method for generating a training sample dataset based on preset data generation rules. In this embodiment, corresponding training sample data is generated according to pre-set data generation rules, and subsequent image processing model training is performed based on this training sample data.

[0033] The training sample dataset can be understood as a collection of multiple generated training sample datasets. In this embodiment, the training sample dataset specifically includes training sample images, processing prompt words, and training label images. In the subsequent specific model training process, suitable training sample datasets will be selected for model training based on the actual training needs and training tasks.

[0034] Training sample images can be understood as the original images; processing prompts can be understood as information used to input into the image processing model, prompting the model to perform corresponding processing; training label images can be understood as the standard qualified images expected to be generated after processing based on the original images and processing prompts.

[0035] The preset data generation rule can be understood as the rule set for generating training sample data in the method provided in the embodiments of this specification. The corresponding training sample data can be generated through the preset data generation rule.

[0036] To improve the quality of the training sample dataset, in a specific embodiment provided in this specification, the training sample dataset is generated according to preset data generation rules, including: An initial training sample data set is generated according to preset data generation rules; The training sample data set is selected from the initial training sample data set according to the preset training data quality assessment rules.

[0037] The method provided in the embodiments of this specification generates a training sample data set by means of preset data generation rules. However, the quality of the generated training sample data varies during the process of generating training sample data by means of preset data generation rules. In order to ensure the high quality of the final generated training sample data set, in this embodiment, the generated training sample data can also be filtered according to preset training data quality evaluation rules.

[0038] Specifically, an initial training sample data set can be generated based on preset data generation rules. The initial training sample data set can be understood as the original set of training sample data generated based on preset data generation rules, in which the generated training sample data has not undergone any filtering process.

[0039] To improve the quality of training sample data and maintain its visual naturalness and content consistency, the generated initial training sample data set can be filtered based on preset training data quality evaluation rules. Specifically, deduplication mechanisms and quality cleaning operators can be used to filter the initial training sample data set.

[0040] In practical applications, a multi-level deduplication mechanism can be set up to filter the training sample data in the initial training sample dataset. This process can effectively prevent the model from overfitting and memorizing training data during the training process.

[0041] The purpose of quality cleaning operators is to ensure that data is clean, diverse, and of high value. These operators can include resolution and size filtering operators (filtering out images with resolutions below a preset threshold or aspect ratios exceeding a preset threshold), grayscale and monochrome image filtering operators (detecting the number of channels or color entropy of an image, removing images that are purely black and white, monochrome, or lacking in color), blur detection operators (removing out-of-focus or motion-blurred images), aesthetic scoring operators (scoring images based on an aesthetic scoring model, retaining images with scores greater than a preset score threshold, effectively improving the quality of images generated by image processing models), and image-text similarity filtering operators (filtering out images where the similarity between the text in the image and the recognized text is below a threshold), etc. The specific content of the quality cleaning operators provided in the embodiments of this specification is not limited; the actual application shall prevail.

[0042] This method uses a preset training data quality assessment rule to filter the initial training sample dataset, eliminating training samples that do not meet the requirements, thus obtaining the final training sample dataset. This improves the quality of the training sample data in the final dataset, providing a data foundation for the subsequent training of image processing models.

[0043] The initial training sample data set is generated according to preset data generation rules, including at least one of the following: Input the training sample images and processing prompts into a preset image generation model to generate training label images; or, Generate training label images based on training sample images and preset conditional control information; or, For training sample images containing text to be adjusted and preset template rules, generate training label images, where the training label images include the adjusted target text; or, The training data distribution information for long-tail image processing tasks is determined based on the training sample images, and target training sample data is generated for long-tail image processing tasks based on the training data distribution information.

[0044] The method provided in the embodiments of this specification provides a method for generating a multi-path synthesis initial training sample data set: Path 1: Expert Model Generation Path. This path utilizes an existing expert model (a pre-defined image generation model, such as a specific Inpainting model) to generate corresponding training label images based on training sample images and processing prompts. The training sample images can be pre-saved images or images pre-generated by the expert model. For example, a technician can use an expert model to generate corresponding training sample images based on text instructions, and then combine these training sample images with new text instructions to generate corresponding training label images.

[0045] Path 2: Structural control path, which introduces pre-defined control information such as segmentation map, key points, and depth map. Based on the pre-defined control information, corresponding training label images with higher difficulty are generated from the training sample images. For example, the pose and viewpoint of the objects in the training sample images are changed.

[0046] Path 3: Rendering and compositing path. For images containing text, the path can use preset template rules (such as 3D scenes generated by the rendering engine, graphic design templates, etc.) to render and generate training label images based on the content and layout of the text to be adjusted. The path also generates corresponding text content in the training label images, thus constructing more realistic training label images and solving the problem of OCR (Optical Character Recognition) data gaps.

[0047] After the above processing, a certain number of training sample data will be generated. At this time, the training sample data can be analyzed to determine the image processing task that each training sample data is suitable for. Furthermore, the distribution information of the current training sample data on long-tail image processing tasks (such as specific image processing styles, specific image processing methods, etc.) can be analyzed. Based on the training data distribution information, it can be determined whether there are any image processing long-tail tasks with shortcomings. When it is determined that there is a small amount of training sample data for a certain image processing long-tail task, targeted data supplementation can be performed for that image processing long-tail task to generate a corresponding number of training sample data.

[0048] The method provided in the embodiments of this specification employs a multi-path closed-loop data generation approach. It generates a sufficient number of training sample data through methods including expert model generation paths, structural control paths, rendering and compositing paths, and long-tail task completion. Then, a quality cleaning operator is used to filter and refine the generated training sample data, further improving the quality of the training sample dataset and providing a data foundation for subsequent image processing model training. This method enables the low-cost, large-scale production of high-quality, high-difficulty editing data, solving the problem of training sample data shortage.

[0049] Step 104: Train an initial image processing model based on the training sample data set to obtain a reference image processing model.

[0050] After the training sample dataset is generated, an initial image processing model can be trained based on the training sample data. This initial image processing model can be understood as an image processing model that has not yet been trained. In the method provided in the embodiments of this specification, a multi-stage progressive training strategy is adopted. This step is the first training stage, i.e., the pre-training stage of the model. After the model pre-training process, a corresponding reference image processing model for the first training stage is obtained. The reference image processing model can be understood as the image processing model generated after the pre-training process.

[0051] In one specific embodiment provided in this specification, training an initial image processing model based on the training sample data set to obtain a reference image processing model includes: Identify at least one image processing task; The training sample dataset is grouped according to each image processing task to obtain the training sample dataset subset corresponding to each image processing task. An initial image processing model is trained based on each subset of training sample data to obtain a reference image processing model.

[0052] Specifically, in this embodiment, multi-condition-aware bucket sampling is adopted. The training sample data set is grouped (i.e., bucketed) according to different types of image processing tasks (such as denoising, super-resolution, redrawing, text, etc.) to obtain the training sample data subset corresponding to each image processing task. In the model pre-training stage, training sample data is sampled from each training sample data subset in proportion and the initial image processing model is trained to balance the basic capabilities of the image processing model.

[0053] Furthermore, in a specific embodiment provided in this specification, the training sample data set is grouped according to each image processing task to obtain a subset of training sample data corresponding to each image processing task, including: Determine the target image processing task and the training sample selection information corresponding to the target image processing task, wherein the target image processing task is any one of the image processing tasks; The training sample dataset is filtered according to the training sample filtering information to obtain a subset of training samples corresponding to the target image processing task.

[0054] To better explain the grouping of training sample data, different image processing tasks in practical applications correspond to different training sample selection information. This selection information can be understood as filtering information used to select samples that meet the requirements of the corresponding image processing task. In this embodiment, we will take a specific image processing task as an example for explanation, that is, determining the target image processing task from among the various image processing tasks, and the corresponding training sample selection information for that target image processing task. The target image processing task is any one of the various image processing tasks.

[0055] Once the training sample selection information corresponding to the target image processing task is determined, the training sample data set can be selected based on this selection information, and the training sample data that meets the selection information can be added to the training sample subset corresponding to the target image processing task.

[0056] See Figure 2 , Figure 2 This specification illustrates a schematic diagram of grouping image processing tasks according to an embodiment, such as... Figure 2 As shown, the image processing model includes n image processing tasks, each with its corresponding ground truth information. In this embodiment, the image size is used as an example to illustrate the training sample selection information. Figure 2 As shown, the image size corresponding to image processing task 1 is 1:1. Therefore, training sample data whose difference between the ratio of the image size ratio and the ratio of the true value is less than a preset threshold can be added to the training sample data subset corresponding to the image processing task.

[0057] In practical applications, the training sample selection information also includes the specific task content of the image processing task, such as poster generation task, character adjustment task, etc. The training sample data can be selected according to the specific task content to ensure that the training sample data in the training sample data subset can meet the corresponding image processing task.

[0058] In one specific embodiment provided in this specification, the training sample dataset is filtered according to the training sample filtering information to obtain a subset of training samples corresponding to the target image processing task, including: The training sample dataset is filtered according to the training sample filtering information to obtain at least one target training sample data. Embedding processing is performed on the training sample data of each target to obtain the target training sample data feature information corresponding to each target training sample data; A subset of training samples corresponding to the target image processing task is generated based on the feature information of each target training sample data.

[0059] In practical applications, the image processing model provided in the embodiments of this specification adopts a diffusion model (DIT architecture) based on the Transformer architecture for explanation. To reduce the memory challenges brought by the DIT architecture during model training, the method provided in the embodiments of this specification can pre-embed the training sample data, converting it into offline tensor storage. This avoids redundant online computation during real-time model training, reduces the floating-point operation load per iteration, and allows the system's computing power to be concentrated on model training itself.

[0060] Specifically, taking the target image processing task as an example, during the process of filtering the training sample dataset according to its corresponding training sample selection information, at least one target training sample data can be obtained. Then, each target training sample data undergoes embedding processing, converting it into corresponding target training sample data feature information (i.e., tensor information), which is then stored. During model training, the target training sample data feature information generated after pre-embedding is used for training, effectively freeing up GPU memory space, reducing the degree of image processing model slicing, thereby reducing the amount of data communication during training and improving the training efficiency of the model.

[0061] In practical applications, image processing models reference the content of the original image when generating a new image. The original image may contain target objects that need to be retained in the new image. For example, an image processing task might be to "generate a scene with a red background, stage equipment, and choir members based on a person in an image, changing the person's posture to standing and holding a trumpet." Under this task, the faces of the people in the new image are usually regenerated, resulting in significant differences between the newly generated faces and those in the original image. This can lead users to perceive the newly generated image as unrealistic.

[0062] To address this issue, in another specific embodiment provided in this specification, an initial image processing model is trained based on each subset of training sample data, including: When the image processing task is an object processing task, obtain a subset of reference training sample data corresponding to the object processing task; The training sample data in the subset of reference training sample data are sorted according to object similarity to obtain the sorting result; Based on the sorting results, training sample data is selected from the subset of reference training sample data to train the initial image processing model.

[0063] In this embodiment, the object processing task in the image processing task is trained in a specific way. The object processing task can be understood as a task that needs to process the facial data of objects in the training sample images. The objects can include people, animals, dolls, etc. The object processing task can be to update the color of a puppy in the image, or to change the words on the chest of a doll in the image, or to generate a group photo based on the people in two images, etc.

[0064] For object processing tasks, this embodiment adopts a method of gradually increasing the object similarity in the training sample data to train the image processing model. Specifically, when the image processing task is an object processing task, a subset of reference training sample data corresponding to the object processing task is obtained. The subset of reference training sample data can be understood as a set of training sample data used for the object processing task.

[0065] In the subset of reference training sample data, the object similarity between the training sample images and the training label images in each training sample data is calculated, and each training sample data is sorted according to the object similarity to obtain the sorting result. The sorting result can be arranged in order of object similarity from low to high, or in order of high to low.

[0066] After obtaining the ranking results, training sample data can be selected to train the image processing model according to a pre-set training strategy. Specifically, the training strategy refers to gradually increasing the object similarity as the pre-training process progresses. That is, in the initial stage of pre-training, training sample data with low object similarity is selected for training. This helps the image processing model respond to complex poses and facial expressions. During the training process, the object similarity of the training data is gradually increased. This helps the image processing model respond to the same commands with relatively small variations, and can better preserve the facial features of the objects.

[0067] In addition, in order to improve the ability of the image processing model to process multimodal data, the method provided in the embodiments of this specification will use a mixture of text-to-image (T2I) training data and image-to-image (Image-to-Image) training data. The training ratio of the two types of training data is approximately 1:1. It should be noted that the ratio of the two is not strictly allocated according to 1:1, and can be set according to the actual application situation. In the method provided in the embodiments of this specification, this is not limited.

[0068] To further enhance the capabilities of the image processing model, the training data for Image-to-Image will be allocated according to the content of the images. For example, the ratio of natural images, design images, and people images is approximately 3:1:1. Design images can include posters, UI (User Interface) designs, and other images containing text content. Design images prepare the image processing model for editing the text content in the future.

[0069] For training data on text-based images, different proportions of training data will be designed. The proportion of training data for structural editing will be increased. This will train the image processing model to learn not only color or background changes, but also structural changes, enabling the image processing model to learn spatial reasoning rather than simple pixel replacement.

[0070] In practical applications, scenarios often arise where multiple images need to be processed, i.e., generating a single image from multiple reference images. For example, generating a group photo based on people in Image 2 and Image 3, with the background of Image 1 as the background. For this type of training data, the image processing model may overfit to a fixed order of reference images during training. Therefore, in another specific embodiment provided in this specification, an initial image processing model is trained based on subsets of training sample data, including: Obtain the training sample data to be processed, which includes training sample images, processing prompt words, and training label images; The current training sample data is constructed based on the training sample images, the processing prompt words, and the training label images, and an initial image processing model is trained based on the current training sample data.

[0071] In this embodiment, for this type of training sample data, the training sample images and processing prompts can be decoupled, thereby improving the understanding ability of the image processing model.

[0072] Specifically, the training sample data to be processed is obtained, which includes multiple training sample images, processing prompt words, and training label images.

[0073] Subsequently, the positional relationships between multiple training sample images are randomly arranged, and some unnecessary reference images are even randomly discarded. Processing prompts are then regenerated. After the above processing is completed, new current training sample data is constructed, and the initial image processing model is trained based on the current training sample data.

[0074] For example, taking the training sample data to be processed, which includes images 1, 2, 3, and 4, where images 1-3 are the current training sample data and image 4 is the training label image, and the processing prompt is "Generate image 4 based on the clothes in image 1, the person in image 2, and the background in image 3," for this training sample data, by analyzing the relationship between the images in the processing prompt and the labels of each image, a new processing prompt can be generated, such as "Make the person in image 2 wear the clothes in image 1 and stand in the background of image 3 to generate image 4," or "Make the person in image 2 wear the clothes in image 1 to generate image 4," and so on. Furthermore, to further improve the processing capability of the image processing model, the image labels of each image can be updated, for example, renaming the original image 1 as image 2 and the original image 2 as image 1, and then generating a new processing prompt, "Make the person in image 1 wear the clothes in image 2 and stand in the background of image 3 to generate image 4."

[0075] In the method provided in this embodiment, to address the issue of image processing models easily overfitting to the fixed order of images in scenarios with multiple training sample images, the editing effect of multiple training sample images is improved by decoupling the spatial order from the processing prompts. Specifically, the decoupling method involves organizing the training sample data to generate new permutations and combinations. During the data organization phase, the training sample images are randomly rearranged, unnecessary training sample images are randomly discarded, and the processing prompts are adjusted simultaneously. The identifiers of the training sample images in the processing prompts are adjusted according to the updated images to ensure the semantic accuracy of the processing prompts. Through this process, the image processing model can learn to understand the actual meaning of the image identifiers during the pre-training phase, rather than simply memorizing the order of the images, thereby improving the model's comprehension ability.

[0076] This step is the pre-training stage of the image processing model. In the pre-training stage, the content of the training sample images and processing prompts is dynamically reorganized, breaking the fixed word order, so that the image processing model truly has semantic understanding ability, rather than rigidly executing specific sentence patterns, making the image processing model more intelligent.

[0077] Step 106: Update the processing prompt words in the training sample data set to obtain a reference training sample set, and train the reference image processing model based on the reference training sample set to obtain the target image processing model.

[0078] After the first stage of training, a reference image processing model is obtained. This reference image processing model can be understood as the image processing model obtained after the initial image processing model has undergone pre-training. Subsequently, the second training stage, the dynamic instruction fine-tuning stage, begins. In this stage, to further improve the image processing model's understanding ability, a strategy is adopted that avoids using fixed prompt word templates. Instead, the processing prompt words are randomly shuffled, replaced with synonyms, and restructured to update the processing prompt words in the training sample dataset, thereby generating a new reference training sample set. This new set is then used to further train the reference image processing model. This forces the image processing model to learn the deep semantic correspondence between image content and processing prompt words, rather than memorizing specific word orders, thus significantly improving instruction compliance consistency.

[0079] To further explain the dynamic instruction fine-tuning stage, this specification provides a specific embodiment that further explains how to update the processing prompts. Specifically, updating the processing prompts in the training sample data set to obtain a reference training sample set includes: The following steps are taken: determine the processing prompt word to be processed, the corresponding training sample image to be processed, and the training label image to be processed, wherein the processing prompt word to be processed is any one of the processing prompt words; Adjust the processing prompt words to be processed based on the prompt word adjustment library to obtain the target processing prompt word to be processed, and determine the prompt word relationship information between the target processing prompt word and the processing prompt words to be processed; Based on the prompt word relationship information, a new reference training sample is constructed from the training sample image to be processed, the training label image to be processed, and the target prompt word to be processed.

[0080] In this embodiment, a specific processing prompt word is used as an example for explanation. The same processing method can be applied to all processing prompt words, and will not be repeated here. Specifically, a processing prompt word to be processed is determined from all processing prompt words, and the corresponding training sample image and training label image to be processed are determined simultaneously.

[0081] The prompts to be processed are adjusted based on the prompt adjustment library to generate new target prompts to be processed. The prompt adjustment library includes various update methods for adjusting and updating prompts, such as synonym replacement, word order adjustment, sentence restructuring, etc. After adjusting the prompts to be processed using the prompt adjustment library, new target prompts to be processed are obtained, and the prompt relationship information between the target prompts to be processed and the original prompts to be processed is further determined. The prompt relationship information can be understood as a representation of the difference between the updated target prompts to be processed and the original prompts to be processed. For example, if synonym replacement is used in the prompt adjustment library, the prompt relationship information is semantically similar; if sentence restructuring is used, it is necessary to determine whether the semantics are opposite or similar based on the semantic information of the target prompts to be processed and the original prompts to be processed.

[0082] The purpose of determining the relationship information of the prompt words is to re-determine the relationship between the training sample image to be processed and the training label image to be processed in the future. If they are similar in meaning, the input position relationship between the training sample image to be processed and the training label image to be processed is maintained. If they are opposite in meaning, it may be necessary to use the training label image to be processed as the training sample image and the training label image to be processed as the training label image.

[0083] After determining the relationship information of the prompt words, the association between the training sample image to be processed, the training label image to be processed, and the target prompt word to be processed is re-determined based on this relationship information, and a new reference training sample is constructed. This facilitates subsequent input into the image processing model for processing.

[0084] In another specific embodiment provided in this specification, in order to further enrich the processing capabilities of the image processing model, specific image processing tasks (such as image quality restoration, image colorization, image blurring, etc.) can be defined as specific instruction editing tasks and added to regular image processing tasks for mixed training to improve the processing capabilities of the image processing model.

[0085] In another specific embodiment provided in this specification, training the reference image processing model based on the reference training sample set includes: When the image processing task is an object processing task, obtain the training sample image, processing prompt words and training label image corresponding to the object processing task, wherein the training sample image includes object identification information; The dynamic weights of objects are determined based on the current training round, and the object label matrix corresponding to the object identification information is determined in the training sample image. The training sample image and the processing prompt words are input into the reference image processing model to obtain the predicted image generated by the reference image processing model based on the object dynamic weights, the training sample image and the processing prompt words, and the object prediction matrix is ​​determined in the predicted image based on the object identification information; A first loss value is calculated based on the predicted image and the training label image, and a second loss value is calculated based on the object label matrix and the object prediction matrix; The model parameters of the reference image processing model are adjusted based on the first loss value and the second loss value.

[0086] In this embodiment, the object processing task will be further explained. In practical applications, human image processing is an important branch, and human image processing is an important part of the object processing task. Therefore, in this embodiment, a targeted training strategy is adopted for the object processing task.

[0087] Specifically, in the method provided in the embodiments of this specification, object identification information is designed for object processing tasks to improve the consistency of objects in the image processing process. In the model pre-training stage, this embodiment assigns object identification information to each object in the image to construct a global semantic structure. In the second training stage, this object identification information is used to focus on refining fine-grained textures. The object identification information specifically refers to identification information used to characterize the identity information of objects in the image.

[0088] In practical applications, object identification information is determined in the early stages of denoising in the image processing model. If a large loss value of object identification information is used to constrain the object in the later stages of denoising, it will destroy the texture information of the object. Based on this, in this embodiment, a dynamic object weight is designed. The dynamic object weight is related to the magnitude of noise during the denoising process of the image processing model on the noisy image. If the noise is large, the dynamic object weight is high, which is used to strongly constrain the facial information of the object. If the noise is small, the dynamic object weight will be small, or even 0, allowing the image processing model to play freely and generate the corresponding image.

[0089] Furthermore, when the image processing task is an object processing task, the dynamic weights of the objects corresponding to the current round are determined, and the object label matrix corresponding to the object identification information is determined in the training sample images using an object detection algorithm. During training, the image processing model estimates the prediction image based on the currently predicted velocity field and determines the object prediction matrix in the prediction image based on the object identification information.

[0090] See Figure 3 , Figure 3This specification illustrates a schematic diagram of calculating the loss value using object identification information provided in an embodiment of this specification, such as... Figure 3 As shown, taking two input images as an example, a portrait image and a cat image are used as training sample images. The user's instruction is "Generate an image of a child holding a cat based on the input image, with the background being a living room." Simultaneously, the training label image is also input into the image encoder for vectorization processing. The encoder adds noise to the training label image for subsequent image denoising to generate the predicted image.

[0091] The feature information of the training label image, the feature information of the training sample image, and the feature information of the user command are input into each layer of the multimodal spread Transformer for processing. After denoising at multiple time steps, the final predicted image is generated. The first loss value can be calculated based on the predicted image and the training label image. In the predicted image, the predicted region of interest (ROI) of the cat and the predicted ROI of the face are extracted based on the object identification information. The ROI of the cat and the ROI of the face in the training sample image are compared to calculate the second loss value. Finally, the model parameters are adjusted based on the first and second loss values ​​to train the image processing model.

[0092] The first loss value is calculated using the predicted image and the training labeled image, and the second loss value is calculated using the object label matrix and the object prediction matrix. The model parameters of the reference image processing model are adjusted based on the first and second loss values. This method improves the consistency of object identification information by integrating a robust identity consistency constraint during training, thus addressing the issue where the facial features of the object in the model-generated image do not match the facial features of the object in the original image when the AI ​​executes instructions on the object (such as changing the background).

[0093] In another specific embodiment provided in this specification, training the reference image processing model based on the reference training sample set to obtain the target image processing model includes: Train the reference image processing model based on the reference training sample set to obtain the image processing model to be aligned; The image processing model to be aligned is updated based on a preset model alignment algorithm to obtain the target image processing model.

[0094] After the fine-tuning training in the second training phase described above, the final target image processing model can be obtained. In practical applications, the image processing model ultimately aims to generate image information that meets user needs. Therefore, more and more model training incorporates the RLHF (Reinforcement Learning from Human Feedback) stage. By combining reinforcement learning and human feedback techniques, the output of large models is optimized to better align with human preferences and values.

[0095] Based on this, in a specific embodiment provided in this specification, after training the reference image processing model with the reference training sample set to obtain the second training stage of the image processing model to be aligned, the image processing model to be aligned can also be obtained. The image processing model to be aligned can be understood as a model that needs further optimization and interactive learning through the RLHF algorithm; that is, the method provided in this specification can further include a third stage of feedback optimization based on reinforcement learning.

[0096] In this stage, the image processing model to be aligned can be further optimized and adjusted according to the preset model alignment algorithm, so that its output content is more in line with human preferences. After being updated by the preset model alignment algorithm, the final target image processing model can be obtained.

[0097] Specifically, updating the image processing model to be aligned based on a preset model alignment algorithm includes: The image processing model to be aligned is updated based on the direct preference algorithm and / or the group relative strategy algorithm.

[0098] In practical applications, preset model alignment algorithms may include Direct Preference Optimization (DPO) and Group Relative Policy Optimization (GRPO).

[0099] DPO is a training method specifically designed for large language models. It aims to optimize the model using human preference data without the need for complex reinforcement learning algorithms. The core of DPO is to directly adjust the model parameters through preference data, bypassing the fitting of explicit reward models and complex reinforcement learning optimization processes. This not only improves training efficiency but also avoids the instabilities common in traditional RLHF methods.

[0100] The method provided in the embodiments of this specification employs asymmetric gradient optimization and decoupled identity alignment. During the DPO optimization process, positive samples (i.e., data with good editing results) are given larger gradient weights to strengthen positive feedback and accelerate convergence. Furthermore, in the DPO stage, instead of calculating the loss value between the predicted image and the original image on the portrait data, the loss value of the object identification information region is calculated, maintaining facial similarity without affecting the model's ability in the reinforcement learning stage. This method improves the model's convergence speed and final performance on complex tasks.

[0101] GRPO is a policy optimization method in reinforcement learning, often used to solve decision problems in complex, high-dimensional, or non-stationary environments. Its objective function is designed to improve policy stability, robustness, and generalization while ensuring policy performance. The method provided in the embodiments of this specification employs a fine-grained score-weighted ensemble reward strategy, which weights the rewards according to the original scores of relevant numerical terms to obtain finer-grained scores. This results in more continuous reward information and more effective reward supervision.

[0102] This method also emphasizes a hard sample mining strategy. During the training process, it was found that the image processing model responds well to most of the training data. In the method provided in the embodiments of this specification, these samples are removed, allowing the image processing model to focus more on the more difficult training sample data. This is more conducive to improving training efficiency and the final model's capabilities.

[0103] For text editing tasks, the method provided in the embodiments of this specification also proposes a layout-aware OCR reward strategy. This strategy constructs a reward function that calculates not only the accuracy of text recognition but also the positional deviation of the recognition box and the matching degree of the font size and style of the generated text. If the generated text is crowded, exceeds the defined area, or has a sudden change in font style, a higher penalty is applied. This allows the image processing model to learn to generate text information that meets the requirements. This strategy solves the industry problem of existing AI's inability to accurately generate text, enabling the image processing model not only to write characters correctly but also to achieve accurate layout, thus enabling higher-level poster modification and generation tasks.

[0104] For object processing tasks, the method provided in the embodiments of this specification proposes an identity-based reward strategy for the target object. In practical applications, directly calculating the facial similarity between the reference image and the generated image can easily lead to the direct copying of faces from the reference image and pasting them into the generated image. In this method, the generated image is used as the reward for object identification information, thus preventing the model from obtaining false high scores through simple copying and pasting, and avoiding training failure problems such as strange poses and expressions caused by reward deception.

[0105] See Figure 4 , Figure 4 This specification illustrates a flowchart of an image processing method according to an embodiment, which specifically includes the following steps: Step 402: Obtain at least one image to be processed and an image processing prompt.

[0106] Step 404: Input the at least one image to be processed and the image processing prompt word into the target image processing model, wherein the target image processing model is obtained through the training method of the above-described image processing model.

[0107] Step 406: Obtain the target image output by the target image processing model.

[0108] Here, the image to be processed can be understood as the image that needs to be processed, and the image processing prompt can be understood as the response image processing instructions that the user wants to perform for the image to be processed.

[0109] For example, let's take a photo of a man wearing frameless glasses holding a bouquet of flowers as an example. The corresponding image processing prompts are: "Zoom out the camera view, replace the background with a stage scene with a red backdrop, stage equipment, and choir members; adjust the man's posture, change the posture to standing, with his hands placed on his abdomen and holding a trumpet, change his clothes to a black suit, remove the bouquet of flowers, add the action of holding a trumpet, adjust the overall lighting and color tone to match the stage environment, and replace the man's glasses with black-rimmed glasses."

[0110] The image to be processed and the image processing prompts are input into the target image processing model, which is an image processing model trained using the methods described above. The target image processing model will output the final target image.

[0111] In this embodiment, the target image processing model generated by the above-mentioned image processing model training method fully understands the semantic information of the image processing prompts and ensures that the generated target image is consistent with the face in the input image to be processed based on the consistency of object identification information, thereby improving the model processing effect of the target image processing model and enhancing the user experience.

[0112] Corresponding to the above method embodiments, this specification also provides embodiments of a training apparatus for an image processing model. Figure 5 A schematic diagram of a training apparatus for an image processing model according to an embodiment of this specification is shown. Figure 5 As shown, the device includes: The generation module 502 is configured to generate a training sample data set according to a preset data generation rule, wherein the training sample data set includes training sample images, processing prompt words, and training label images; The first training module 504 is configured to train an initial image processing model based on the training sample data set to obtain a reference image processing model. The second training module 506 is configured to update the processing prompt words in the training sample data set, obtain a reference training sample set, and train the reference image processing model based on the reference training sample set to obtain the target image processing model.

[0113] In one specific embodiment provided in this specification, the first training module 504 is further configured as follows: Identify at least one image processing task; The training sample dataset is grouped according to each image processing task to obtain the training sample dataset subset corresponding to each image processing task. An initial image processing model is trained based on each subset of training sample data to obtain a reference image processing model.

[0114] In one specific embodiment provided in this specification, the first training module 504 is further configured as follows: Determine the target image processing task and the training sample selection information corresponding to the target image processing task, wherein the target image processing task is any one of the image processing tasks; The training sample dataset is filtered according to the training sample filtering information to obtain a subset of training samples corresponding to the target image processing task.

[0115] In one specific embodiment provided in this specification, the first training module 504 is further configured as follows: The training sample dataset is filtered according to the training sample filtering information to obtain at least one target training sample data. Embedding processing is performed on the training sample data of each target to obtain the target training sample data feature information corresponding to each target training sample data; A subset of training samples corresponding to the target image processing task is generated based on the feature information of each target training sample data.

[0116] In one specific embodiment provided in this specification, the first training module 504 is further configured as follows: When the image processing task is an object processing task, obtain a subset of reference training sample data corresponding to the object processing task; The training sample data in the subset of reference training sample data are sorted according to object similarity to obtain the sorting result; Based on the sorting results, training sample data is selected from the subset of reference training sample data to train the initial image processing model.

[0117] In one specific embodiment provided in this specification, the first training module 504 is further configured as follows: Obtain the training sample data to be processed, which includes training sample images, processing prompt words, and training label images; The current training sample data is constructed based on the training sample images, the processing prompt words, and the training label images, and an initial image processing model is trained based on the current training sample data.

[0118] In one specific embodiment provided in this specification, the generation module 502 is further configured as follows: An initial training sample data set is generated according to preset data generation rules; The training sample data set is selected from the initial training sample data set according to the preset training data quality assessment rules.

[0119] In one specific embodiment provided in this specification, the generation module 502 is further configured to be at least one of the following: Input the training sample images and processing prompts into a preset image generation model to generate training label images; or, Generate training label images based on training sample images and preset conditional control information; or, For training sample images containing text to be adjusted and preset template rules, generate training label images, where the training label images include the adjusted target text; or, The training data distribution information for long-tail image processing tasks is determined based on the training sample images, and target training sample data is generated for long-tail image processing tasks based on the training data distribution information.

[0120] In one specific embodiment provided in this specification, the second training module 506 is further configured as follows: The following steps are taken: determine the processing prompt word to be processed, the corresponding training sample image to be processed, and the training label image to be processed, wherein the processing prompt word to be processed is any one of the processing prompt words; Adjust the processing prompt words to be processed based on the prompt word adjustment library to obtain the target processing prompt word to be processed, and determine the prompt word relationship information between the target processing prompt word and the processing prompt words to be processed; Based on the prompt word relationship information, a new reference training sample is constructed from the training sample image to be processed, the training label image to be processed, and the target prompt word to be processed.

[0121] In one specific embodiment provided in this specification, the second training module 506 is further configured as follows: When the image processing task is an object processing task, obtain the training sample image, processing prompt words and training label image corresponding to the object processing task, wherein the training sample image includes object identification information; The dynamic weights of objects are determined based on the current training round, and the object label matrix corresponding to the object identification information is determined in the training sample image. The training sample image and the processing prompt words are input into the reference image processing model to obtain the predicted image generated by the reference image processing model based on the object dynamic weights, the training sample image and the processing prompt words, and the object prediction matrix is ​​determined in the predicted image based on the object identification information; A first loss value is calculated based on the predicted image and the training label image, and a second loss value is calculated based on the object label matrix and the object prediction matrix; The model parameters of the reference image processing model are adjusted based on the first loss value and the second loss value.

[0122] In one specific embodiment provided in this specification, the second training module 506 is further configured as follows: Train the reference image processing model based on the reference training sample set to obtain the image processing model to be aligned; The image processing model to be aligned is updated based on a preset model alignment algorithm to obtain the target image processing model.

[0123] In one specific embodiment provided in this specification, the second training module 506 is further configured as follows: The image processing model to be aligned is updated based on the direct preference algorithm and / or the group relative strategy algorithm.

[0124] The apparatus provided in the embodiments of this specification employs a multi-path closed-loop data generation method. It generates a sufficient number of training sample data through methods including expert model generation path, structural control path, rendering and compositing path, and long-tail task completion. Then, a quality cleaning operator is used to filter and refine the generated training sample data, further improving the quality of the training sample dataset and providing a data foundation for subsequent image processing model training. This apparatus enables the low-cost, large-scale production of high-quality, high-difficulty editing data, solving the problem of training sample data shortage.

[0125] A multi-stage training strategy was adopted in the model training phase. During the training process, each image was pre-embedded to save offline tensor information. During model training, the feature information of the target training sample data generated after pre-embedding was used for training, which effectively freed up the GPU memory space, reduced the degree of image processing model slicing, thereby reducing the amount of data communication during the training process and improving the training efficiency of the model.

[0126] Secondly, in order to address the issue of image processing models easily overfitting to the fixed order of each image component in scenarios with multiple training sample images, the editing effect of multiple training sample images is improved by decoupling the spatial order from the processing prompts, thereby enhancing the understanding ability of the image processing model.

[0127] Finally, by updating the image processing model through a preset model alignment algorithm, the training efficiency is further improved, and the instability common in the traditional RLHF method is avoided, thereby improving the convergence speed and final effect of the model on complex tasks.

[0128] The above is a schematic scheme of an image processing model training device according to this embodiment. It should be noted that the technical solution of this image processing model training device and the technical solution of the image processing model training method described above belong to the same concept. For details not described in detail in the technical solution of the image processing model training device, please refer to the description of the technical solution of the image processing model training method described above.

[0129] Corresponding to the above method embodiments, this specification also provides embodiments of an image processing apparatus. Figure 6 A schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification is shown. For example... Figure 6 As shown, the device includes: The acquisition module 602 is configured to acquire at least one image to be processed and an image processing prompt word; The input module 604 is configured to input the at least one image to be processed and the image processing prompt words into the target image processing model, wherein the target image processing model is obtained by the training method of the image processing model described above; The generation module 606 is configured to obtain the target image output by the target image processing model.

[0130] In this embodiment, the target image processing model generated by the above-mentioned image processing model training method fully understands the semantic information of the image processing prompts and ensures that the generated target image is consistent with the face in the input image to be processed based on the consistency of object identification information, thereby improving the model processing effect of the target image processing model and enhancing the user experience.

[0131] The above is an illustrative scheme of an image processing apparatus according to this embodiment. It should be noted that the technical solution of this image processing apparatus and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the image processing apparatus, please refer to the description of the technical solution of the image processing method described above.

[0132] See Figure 7 , Figure 7 This specification illustrates a processing flowchart of an image processing method applied to a cloud-side device according to an embodiment, which specifically includes the following steps: Step 702: The receiving end device sends at least one image to be processed and an image processing prompt word.

[0133] Step 704: Input the at least one image to be processed and the image processing prompt word into the target image processing model, wherein the target image processing model is obtained through the training method of the above-described image processing model.

[0134] Step 706: Obtain the target image output by the target image processing model and send the target image to the end device.

[0135] The image processing method provided in this embodiment is applied to cloud-side devices and interacts with edge devices. Since image processing requires high computing resources, in order to ensure a good user experience for users using various types of edge devices, this method deploys the target image processing model on the cloud-side device and interacts with the edge device via the network. Users can transmit the image to be processed and image processing prompts to the cloud-side device through the edge device. The cloud-side device quickly generates the corresponding target image using its superior computing resources and returns the target image to the edge device. This ensures the efficient and rapid execution of image processing tasks, provides efficient and stable image processing services for different users, and improves the user experience.

[0136] See Figure 8 , Figure 8 This specification illustrates an architecture diagram of an image processing system according to one embodiment of the present specification. The image processing system may include a client 100 and a server 200. Client 100 is used to send at least one image to be processed and image processing prompts to server 200; Server 200 is used to input the at least one image to be processed and the image processing prompt words into the target image processing model, wherein the target image processing model is obtained through the training method of the above-mentioned image processing model; obtain the target image output by the target image processing model, and send the target image to client 100; Client 100 is also used to receive target images sent by server 200.

[0137] An image processing system may include multiple clients 100 and a server 200. Clients 100 can be referred to as edge devices, and the server 200 can be referred to as cloud devices. Multiple clients 100 can establish communication connections through the server 200. In an image processing scenario, the server 200 is used to provide image processing services between the multiple clients 100. Each client 100 can act as a sender or receiver, communicating through the server 200.

[0138] Users can interact with server 200 through client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In an image processing scenario, users can publish data streams to server 200 through client 100, server 200 can generate target images based on the data streams, and push the target images to other clients that have established communication.

[0139] In this system, client 100 and server 200 establish a connection via a network. The network provides the medium for communication between client 100 and server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. Data transmitted by client 100 may need to undergo encoding, transcoding, compression, or other processing before being published to server 200.

[0140] Client 100 can be a browser, an app (application), a web application such as an H5 (HyperText Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application. Client 100 can be developed based on the software development kit (SDK) of the corresponding service provided by server 200, such as a real-time communication (RTC) SDK. Client 100 can be deployed on a computing device and depends on the device or certain apps on the device to run. The computing device may have a display screen and support information browsing, such as a personal mobile terminal like a mobile phone, tablet, or personal computer. Various other types of applications can also be configured on the computing device, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platform software.

[0141] Server 200 may include servers providing various services, such as servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It should be noted that server 200 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server integrated with blockchain. The server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0142] It is worth noting that the image processing methods provided in the embodiments of this specification are generally executed by the server. However, in other embodiments of this specification, the client may also have similar functions to the server, thereby executing the image processing methods provided in the embodiments of this specification. In other embodiments, the image processing methods provided in the embodiments of this specification may also be executed jointly by the client and the server.

[0143] Figure 9A structural block diagram of a computing device 900 according to an embodiment of this specification is shown. The components of the computing device 900 include, but are not limited to, a memory 910 and a processor 920. The processor 920 is connected to the memory 910 via a bus 930, and a database 950 is used to store data.

[0144] The computing device 900 also includes an access device 940, which enables the computing device 900 to communicate via one or more networks 960. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 940 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0145] In one embodiment of this specification, the above-described components of the computing device 900 and Figure 9 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 9 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0146] The computing device 900 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 900 can also be a mobile or stationary server.

[0147] The processor 920 is used to execute the following computer program / instruction, which, when executed by the processor, implements the training method and the steps of the image processing method of the above-mentioned image processing model.

[0148] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device belongs to the same concept as the training method of the image processing model and the technical solution of the image processing method described above. For details not described in detail in the technical solution of the computing device, please refer to the description of the training method of the image processing model and the technical solution of the image processing method described above.

[0149] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the training method and image processing method steps of the above-described image processing model.

[0150] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solution of the image processing model training method and the image processing method described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the image processing model training method and the image processing method described above.

[0151] An embodiment of this specification also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the training method and image processing method steps of the above-described image processing model.

[0152] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product belongs to the same concept as the technical solution of the image processing model training method and the image processing method described above. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the image processing model training method and the image processing method described above.

[0153] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0154] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0155] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this specification is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this specification. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this specification.

[0156] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0157] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. These embodiments have been selected and specifically described in this specification to better explain the principles and practical applications of this specification, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A training method for an image processing model, characterized in that, include: A training sample data set is generated according to a preset data generation rule, wherein the training sample data set includes training sample images, processing prompt words, and training label images; An initial image processing model is trained based on the training sample data set to obtain a reference image processing model; Update the processing prompt words in the training sample data set to obtain a reference training sample set, and train the reference image processing model based on the reference training sample set to obtain the target image processing model.

2. The method as described in claim 1, characterized in that, An initial image processing model is trained based on the training sample data set to obtain a reference image processing model, including: Identify at least one image processing task; The training sample dataset is grouped according to each image processing task to obtain the training sample dataset subset corresponding to each image processing task. An initial image processing model is trained based on each subset of training sample data to obtain a reference image processing model.

3. The method as described in claim 2, characterized in that, The training sample dataset is grouped according to each image processing task to obtain a subset of training sample data corresponding to each image processing task, including: Determine the target image processing task and the training sample selection information corresponding to the target image processing task, wherein the target image processing task is any one of the image processing tasks; The training sample dataset is filtered according to the training sample filtering information to obtain a subset of training samples corresponding to the target image processing task.

4. The method as described in claim 3, characterized in that, The training sample dataset is filtered according to the training sample filtering information to obtain a subset of training samples corresponding to the target image processing task, including: The training sample dataset is filtered according to the training sample filtering information to obtain at least one target training sample data. Embedding processing is performed on the training sample data of each target to obtain the target training sample data feature information corresponding to each target training sample data; A subset of training samples corresponding to the target image processing task is generated based on the feature information of each target training sample data.

5. The method as described in claim 2, characterized in that, The initial image processing model is trained based on each subset of training sample data, including: When the image processing task is an object processing task, obtain a subset of reference training sample data corresponding to the object processing task; The training sample data in the subset of reference training sample data are sorted according to object similarity to obtain the sorting result; Based on the sorting results, training sample data is selected from the subset of reference training sample data to train the initial image processing model.

6. The method as described in claim 2, characterized in that, The initial image processing model is trained based on each subset of training sample data, including: Obtain the training sample data to be processed, which includes training sample images, processing prompt words, and training label images; The current training sample data is constructed based on the training sample images, the processing prompt words, and the training label images, and an initial image processing model is trained based on the current training sample data.

7. The method as described in claim 1, characterized in that, A training sample dataset is generated according to preset data generation rules, including: An initial training sample data set is generated according to preset data generation rules; The training sample data set is selected from the initial training sample data set according to the preset training data quality assessment rules.

8. The method as described in claim 7, characterized in that, An initial training sample dataset is generated according to preset data generation rules, including at least one of the following: Input the training sample images and processing prompts into a preset image generation model to generate training label images; or, Generate training label images based on training sample images and preset conditional control information; or, For training sample images containing text to be adjusted and preset template rules, generate training label images, where the training label images include the adjusted target text; or, The training data distribution information for long-tail image processing tasks is determined based on the training sample images, and target training sample data is generated for long-tail image processing tasks based on the training data distribution information.

9. The method as described in claim 1, characterized in that, Update the processing prompt words in the training sample dataset to obtain a reference training sample set, including: The following steps are taken: determine the processing prompt word to be processed, the corresponding training sample image to be processed, and the training label image to be processed, wherein the processing prompt word to be processed is any one of the processing prompt words; Adjust the processing prompt words to be processed based on the prompt word adjustment library to obtain the target processing prompt word to be processed, and determine the prompt word relationship information between the target processing prompt word and the processing prompt words to be processed; Based on the prompt word relationship information, a new reference training sample is constructed from the training sample image to be processed, the training label image to be processed, and the target prompt word to be processed.

10. The method as described in claim 1, characterized in that, Training the reference image processing model based on the reference training sample set includes: When the image processing task is an object processing task, obtain the training sample image, processing prompt words and training label image corresponding to the object processing task, wherein the training sample image includes object identification information; The dynamic weights of objects are determined based on the current training round, and the object label matrix corresponding to the object identification information is determined in the training sample image. The training sample image and the processing prompt words are input into the reference image processing model to obtain the predicted image generated by the reference image processing model based on the object dynamic weights, the training sample image and the processing prompt words, and the object prediction matrix is ​​determined in the predicted image based on the object identification information; A first loss value is calculated based on the predicted image and the training label image, and a second loss value is calculated based on the object label matrix and the object prediction matrix; The model parameters of the reference image processing model are adjusted based on the first loss value and the second loss value.

11. The method as described in claim 1, characterized in that, The reference image processing model is trained based on the reference training sample set to obtain the target image processing model, including: Train the reference image processing model based on the reference training sample set to obtain the image processing model to be aligned; The image processing model to be aligned is updated based on a preset model alignment algorithm to obtain the target image processing model.

12. The method as described in claim 11, characterized in that, Updating the image processing model to be aligned based on a preset model alignment algorithm includes: The image processing model to be aligned is updated based on the direct preference algorithm and / or the group relative strategy algorithm.

13. An image processing method, characterized in that, include: Obtain at least one image to be processed and an image processing prompt; The at least one image to be processed and the image processing prompt are input into the target image processing model, wherein the target image processing model is obtained by the training method of any one of the image processing models in claims 1-12 above; Obtain the target image output by the target image processing model.

14. An image processing method, characterized in that, Applications in cloud-side devices, including: At least one image to be processed and an image processing prompt word sent by the receiving device; The at least one image to be processed and the image processing prompt are input into the target image processing model, wherein the target image processing model is obtained by the training method of any one of the image processing models in claims 1-12 above; Obtain the target image output by the target image processing model and send the target image to the end device.

15. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 14.

16. A computer-readable storage medium storing a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 14.

17. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 14.