Image generation model training method and device, electronic equipment, storage medium and program product
During the training process of the image generation model, the image mask of the image label is generated and fused to the image label, which solves the problem of poor image generation quality caused by inaccurate image labels, and achieves higher quality image generation.
Patent Information
- Application Number
- CN202510264597.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, the training effect of the image generation model is directly determined by the image tag, and inaccurate image tags will lead to poor image generation quality.
By converting the image label of the text sample, the first text corresponding to the image label is obtained, the content difference between the first text and the text sample is obtained, and the image mask of the image label is generated based on the content difference, the image mask is fused to the image label, and the fused image label is obtained; and the image generation model is trained based on the image and the fused image label.
Effectively avoid the influence of image areas corresponding to content differences in image tags, and improve the image generation quality of image generation model.
Smart Images

Figure CN120107392A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a training method, device, electronic device, storage medium and program product for an image generation model. Background Art
[0002] Image generation model, also known as text-based image model, is a model based on artificial intelligence technology that can generate images corresponding to the input text description. By learning the association between text and image, the natural language description is converted into visual content. Image generation model is a deep learning model that can generate realistic images from given input (such as text, noise vector or other forms of conditions).
[0003] In the related art, the training of image generation models is usually based on text samples. The image generation model is called to generate images corresponding to the text samples, and the similarity between the image labels of the image and the text samples is used as the loss value to train the image generation model. Since the loss value is directly determined by the image label, the training effect of the image generation model is directly determined by the image label. Inaccurate image labels will lead to poor image generation quality of the trained image generation model. Summary of the invention
[0004] The embodiments of the present application provide a training method, device, electronic device, computer-readable storage medium and computer program product for an image generation model, which can effectively improve the image generation quality of the image generation model.
[0005] The technical solution of the embodiment of the present application is implemented as follows:
[0006] The present application embodiment provides a method for training an image generation model, including:
[0007] Based on the text sample, calling the image generation model to generate an image corresponding to the text sample;
[0008] Performing text conversion on the image tag of the text sample to obtain a first text corresponding to the image tag, where the first text is used to describe the content in the image tag;
[0009] Obtaining a content difference between the first text and the text sample, and generating an image mask of the image label based on the content difference;
[0010] The image mask is fused to the image label to obtain a fused image label; and the image generation model is trained based on the image and the fused image label.
[0011] The present application embodiment provides a training device for an image generation model, comprising:
[0012] A generation module, used to call an image generation model to generate an image corresponding to the text sample based on the text sample;
[0013] A conversion module, configured to perform text conversion on the image label of the text sample to obtain a first text corresponding to the image label, wherein the first text is used to describe the content in the image label;
[0014] A mask module, configured to obtain a content difference between the first text and the text sample, and generate an image mask of the image label based on the content difference;
[0015] A training module is used to fuse the image mask to the image label to obtain a fused image label; and train the image generation model based on the image and the fused image label.
[0016] In the above scheme, both the first text and the text sample include multiple words, and the above mask module is also used to respectively determine the similarity between each word in the first text and the text sample; determine the word whose similarity with the text sample is less than the similarity threshold as the first word; select the second word corresponding to the first word from the text sample, and use the first word and the second word as the content difference between the first text and the text sample.
[0017] In the above scheme, the above mask module is also used to perform the following processing for each word in the first text: when the word is the first word in the first text, the word and the text sample are encoded respectively to obtain a first vector and a second vector, and the similarity between the first vector and the second vector is determined as the similarity between the word and the text sample; when the word is not the first word in the first text, a second text is constructed based on the word and the text before the word in the first text, and the similarity between the second text and the text sample is determined as the similarity between the word and the text sample.
[0018] In the above scheme, the above mask module is also used to encode the first text to obtain a third vector, and encode the text sample to obtain a fourth vector; determine the distance between the third vector and the fourth vector, and when the value of the distance is not equal to zero, obtain the content difference between the first text and the text sample.
[0019] In the above scheme, the above mask module is also used to perform target segmentation on the image label based on the content difference to obtain the target segmentation result of the image label; wherein the target segmentation result includes a first image area and a second image area, the first image area corresponds to the content difference, and the second image area is different from the first image area; based on the target segmentation result, the pixel value of each pixel in the first image area is set to a first value, and the pixel value of each pixel in the second image area is set to a second value to obtain an image mask of the image label, wherein the first value is greater than the second value.
[0020] In the above scheme, the first text includes multiple words, the content difference includes the target word among the multiple words, and the above mask module is also used to determine the first pixel point corresponding to the word from the image label for each word in the first text; based on the first pixel point corresponding to each word, the first pixel point corresponding to the target word is expanded to obtain the first image area; and the image area in the image label that is different from the first image area is determined as the second image area.
[0021] In the above scheme, the mask module is also used to perform feature extraction on the image label based on the word to obtain the image feature corresponding to the word, wherein the image feature includes multiple feature elements; select pixel points corresponding to the feature elements from the image label; and determine the pixel point corresponding to the feature element with the largest value as the first pixel point corresponding to the word.
[0022] In the above scheme, the above mask module is also used to obtain the second pixel point corresponding to the target word, the second pixel point belongs to the first pixel point corresponding to each of the words, and the second pixel point is adjacent to the first pixel point corresponding to the target word; based on the first pixel points corresponding to each of the words, the relative position relationship between the first pixel point corresponding to the target word and the boundary of the image label is determined; based on the relative position relationship and the second pixel point, the first pixel point corresponding to the target word is expanded to obtain the first image area.
[0023] In the above scheme, the above-mentioned mask module is also used to determine the area boundary of the first image area based on the second pixel point of the target word when the relative position relationship indicates that there are no other first pixel points between the first pixel point corresponding to the target word and each of the boundaries; when the relative position relationship indicates that there are other first pixel points between the first pixel point corresponding to the target word and part of the boundaries, determine the area boundary of the first image area based on the boundary and the second pixel point of the target word; and determine the image area within the area boundary in the image label as the first image area.
[0024] In the above scheme, the image mask includes multiple mask values, each of which corresponds to a pixel point in the image label. The above training module is used to obtain the pixel value of each pixel point in the image label corresponding to the image mask; for each pixel point in the image label corresponding to the image mask, the pixel value of the pixel point is multiplied by the corresponding mask value to obtain the target pixel value of the pixel point, and the pixel value of the pixel point is updated based on the target pixel value to obtain the fused image label.
[0025] In the above scheme, the above training module is also used to fuse the image mask to the image to obtain a fused image; and use the pixel point in the fused image as the third pixel point; for the fourth pixel point in the fused image label, determine the pixel value difference between the fourth pixel point and the corresponding third pixel point; sum the determined pixel value differences to obtain a first loss value, and based on the first loss value, update the model parameters of the image generation model.
[0026] In the above scheme, the above training module is also used to call the image generation model to generate an image corresponding to the first text based on the first text; determine the second loss value based on the image corresponding to the first text and the image label, and determine the third loss value based on the image and the fused image label; perform weighted summation of the second loss value and the third loss value to obtain a fourth loss value, and update the model parameters of the image generation model based on the fourth loss value.
[0027] In the above scheme, the image label includes multiple fifth pixel points, the image corresponding to the first text includes sixth pixel points corresponding one-to-one to the fifth pixel points, and the image mask includes mask values corresponding one-to-one to the fifth pixel points; the above training module is also used to adjust each of the mask values in the image mask based on the content difference to obtain target mask values corresponding to each of the mask values; for each of the fifth pixel points, determine the pixel value difference between the fifth pixel point and the corresponding sixth pixel point, and multiply the pixel value difference and the target mask value corresponding to the fifth pixel point to obtain the loss value of the fifth pixel point; sum the loss values of each of the fifth pixel points to obtain the second loss value.
[0028] In the above scheme, the above training module is also used to perform target segmentation on the image label based on the content difference to obtain the target segmentation result of the image label, and the target segmentation result includes a first image area in the image label corresponding to the content difference, and a second image area in the image label that is different from the first image area; the target mask value corresponding to the pixel point in the first image area in the image mask is determined as the mask value corresponding to the pixel point in the second image area; the target mask value corresponding to the pixel point in the second image area in the image mask is determined as the mask value corresponding to the pixel point in the first image area.
[0029] In the above scheme, the above training module is also used to fuse the image mask to the image to obtain a fused image; and use the pixel point in the fused image as the seventh pixel point; for the eighth pixel point in the fused image label, determine the pixel value difference between the eighth pixel point and the corresponding seventh pixel point; sum the pixel point differences corresponding to each of the eighth pixels to obtain the third loss value.
[0030] The present application provides an image generating device, including:
[0031] An image generation module, used to obtain a third text, and call an image generation model to generate an image corresponding to the third text;
[0032] Wherein, the image generation model is trained using the above-mentioned image generation model training method.
[0033] An embodiment of the present application provides an electronic device, including:
[0034] A memory for storing computer executable instructions or computer programs;
[0035] The processor is used to implement the image generation model training method or image generation method provided in the embodiment of the present application when executing the computer executable instructions or computer program stored in the memory.
[0036] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions or a computer program for causing a processor to execute and implement the image generation model training method or image generation method provided in the embodiment of the present application.
[0037] The embodiment of the present application provides a computer program product, which includes computer executable instructions or computer programs, and the computer executable instructions or computer programs are stored in a computer-readable storage medium. The processor of the electronic device reads the computer executable instructions or computer programs from the computer-readable storage medium, and the processor executes the computer executable instructions or computer programs, so that the electronic device executes the training method or image generation method of the image generation model described in the embodiment of the present application.
[0038] The embodiments of the present application have the following beneficial effects:
[0039] Based on the text sample, the image generation model is called to generate the image corresponding to the text sample, the image label of the text sample is converted to text, the first text corresponding to the image label is obtained, the content difference between the first text and the text sample is obtained, and based on the content difference, an image mask of the image label is generated, and the image mask is fused to the image label to obtain a fused image label; and the image generation model is trained based on the image and the fused image label. In this way, by performing text conversion on the image label of the text sample, the first text corresponding to the image label is obtained, and the content difference between the first text and the text sample is obtained. Since there is a content difference between the first text and the text sample, it means that there is an image area corresponding to the content difference in the image label of the text sample, which makes it difficult for the image label to adapt to the text sample. It means that directly training the image generation model based on the image label will result in poor image generation quality of the trained image generation model. By generating an image mask of the image label based on the content difference, the image mask can accurately reflect the content difference between the first text and the text sample, and the image mask is fused into the image label to obtain a fused image label, so that the image area corresponding to the content difference in the image label is covered by the content difference reflected by the image mask, so that the obtained fused image label can cover the image area corresponding to the content difference in the image label. By training the image generation model based on the image and the fused image label, the training of the image generation model can effectively avoid the influence of the image area corresponding to the content difference in the image label, thereby effectively improving the image generation quality of the image generation model. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a schematic diagram of the architecture of the training system of the image generation model provided in the embodiment of the present application;
[0041] Figure 2 is a schematic diagram of the structure of an electronic device for training an image generation model provided in an embodiment of the present application;
[0042] Figure 3 is a schematic diagram of the structure of an electronic device for generating an image provided by an embodiment of the present application;
[0043] Figure 4 This is a flow chart of the training method of the image generation model provided in the embodiment of the present application. Figure 1 ;
[0044] Figure 5 This is a flow chart of the training method of the image generation model provided in the embodiment of the present application. Figure 2 ;
[0045] Figure 6 This is a flow chart of the training method of the image generation model provided in the embodiment of the present application. Figure 3 ;
[0046] Figure 7 This is a flow chart of the training method of the image generation model provided in the embodiment of the present application. Figure 4 ;
[0047] Figure 8 This is a flow chart of the training method of the image generation model provided in the embodiment of the present application. Figure 5 ;
[0048] Fig. 9 This is a flow chart of the training method of the image generation model provided in the embodiment of the present application. Figure 6 ;
[0049] Fig.10 This is a schematic diagram of the principle of the training method of the image generation model provided in the embodiment of the present application. Figure 1 ;
[0050] Fig.11 This is a schematic diagram of the principle of the training method of the image generation model provided in the embodiment of the present application. Figure 2 ;
[0051] Fig.12 It is a schematic diagram of the effect of the response point provided in the embodiment of the present application;
[0052] Fig.13 It is a structural schematic diagram of the image generation model provided in the embodiment of the present application;
[0053] Fig.14This is a schematic diagram of the principle of the training method of the image generation model provided in this application Figure 3 . DETAILED DESCRIPTION
[0054] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.
[0055] In the following description, reference is made to some embodiments, which describe subsets of all possible embodiments, but it can be understood that some embodiments may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0056] In the following description, the terms first\second\third involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that the first\second\third can be interchanged in a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0057] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those understood by a person skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0058] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0059] 1) Image Generation Model: Also known as Text-to-Image Model, it is a model based on artificial intelligence technology that can generate images corresponding to the input text description. By learning the association between text and image, the natural language description is converted into visual content. The image generation model is a deep learning model that can generate realistic images from given inputs (such as text, noise vectors, or other forms of conditions). By learning a large amount of image and text data, the model can understand the semantic information in the text description and map it into the image space to generate visual content that matches the description. Typical applications include artistic creation, image restoration, virtual scene generation, advertising design, etc.
[0060] 2) Image mask: An image mask is a binary image corresponding to the original image, where the value of each pixel indicates whether the pixel is covered (usually 0 for covered and 1 for uncovered). In deep learning, masks can also be used for self-supervised learning tasks, such as masked image modeling (MAE), where the model needs to reconstruct the original image based on the masked part.
[0061] 3) Object Segmentation: It is a key task in the field of computer vision and image processing, which involves assigning each pixel or region in an image or video to a specific object or background. The purpose of object segmentation is to accurately identify and separate the objects in the image so that each object has a clear boundary and can be analyzed and processed independently.
[0062] 4) Text Transformation: In the intersection of computer vision and natural language processing (NLP), it usually refers to the process of converting visual information (such as image labels) into text descriptions. This process involves converting visual elements such as objects, scenes, actions in the image into natural language descriptions, allowing computers to express image content in a way that humans can understand.
[0063] During the implementation of the embodiments of the present application, the applicant discovered that the related technology has the following problems:
[0064] In the related art, the training of image generation models is usually based on text samples. The image generation model is called to generate images corresponding to the text samples, and the similarity between the image labels of the image and the text samples is used as the loss value to train the image generation model. Since the loss value is directly determined by the image label, the training effect of the image generation model is directly determined by the image label. Inaccurate image labels will lead to poor image generation quality of the trained image generation model.
[0065] The embodiments of the present application provide a training method, device, electronic device, computer-readable storage medium and computer program product for an image generation model, which can effectively improve the image generation quality of the image generation model. The following describes an exemplary application of the training system for the image generation model provided by the embodiments of the present application.
[0066] See also Figure 1 , Figure 1 It is a schematic diagram of the architecture of the training system 100 for the image generation model provided in an embodiment of the present application. The terminal (terminal 400 is shown as an example) is connected to the server 200 via the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.
[0067] The terminal 400 is used for the user to use the client 410 to display images on a graphical interface 410-1 (graphic interface 410-1 is shown as an example). The terminal 400 and the server 200 are connected to each other via a wired or wireless network.
[0068] In some embodiments, the server 200 may be an independent physical server, or a server cluster or business system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. The terminal 400 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart TV, a smart watch, a car terminal, etc., but is not limited thereto. The electronic device provided in the embodiment of the present application may be implemented as a terminal or as a server. The terminal and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiment of the present application.
[0069] In some embodiments, the server 200 calls an image generation model based on a text sample to generate an image corresponding to the text sample, performs text conversion on the image label of the text sample to obtain a first text corresponding to the image label, generates a mask for the image label based on the content difference between the first text and the text sample, and fuses the image mask to the image label to obtain a fused image label; and trains the image generation model based on the image and the fused image label, and sends the trained image generation model to the terminal 400.
[0070] In other embodiments, the terminal 400 calls an image generation model based on the text sample to generate an image corresponding to the text sample, performs text conversion on the image label of the text sample to obtain a first text corresponding to the image label, generates a mask of the image label based on the content difference between the first text and the text sample, and fuses the image mask to the image label to obtain a fused image label; and trains the image generation model based on the image and the fused image label, and sends the trained image generation model to the server 200.
[0071] See also Figure 2 , Figure 2 is a schematic diagram of the structure of an electronic device for training an image generation model provided in an embodiment of the present application, wherein: Figure 2 The electronic device 500 shown may be Figure 1 The server 200 or the terminal 400 in Figure 2The electronic device 500 shown includes: at least one processor 430, a memory 450, and at least one network interface 420. The various components in the electronic device 500 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a drive bus, and a status signal bus. However, for the sake of clarity, the bus system 440 is not described in detail. Figure 2 Various buses are labeled as bus system 440 .
[0072] Processor 430 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0073] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 430.
[0074] The memory 450 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0075] In some embodiments, memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.
[0076] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0077] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB).
[0078] In some embodiments, the training device for the image generation model provided in the embodiments of the present application can be implemented in software. Figure 2 A training device 455 of an image generation model stored in a memory 450 is shown, which may be software in the form of a program or a plug-in, and includes the following software modules: a generation module 4551, a conversion module 4552, a mask module 4553, and a training module 4554. These modules are logical, and thus may be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.
[0079] See also Figure 3 , Figure 3 is a schematic diagram of the structure of an electronic device for generating an image provided by an embodiment of the present application, wherein: Figure 3 The electronic device 600 shown may be Figure 1 The server 200 or the terminal 400 in Figure 3 The electronic device 600 shown includes: at least one processor 530, a memory 550, and at least one network interface 520. The various components in the electronic device 600 are coupled together via a bus system 540. It is understood that the bus system 540 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 540 is not described in detail. Figure 3 Various buses are labeled as bus system 540 .
[0080] The processor 530 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0081] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 550 may optionally include one or more storage devices that are physically remote from the processor 530.
[0082] The memory 550 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.
[0083] In some embodiments, the memory 550 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.
[0084] Operating system 551, including system programs for processing various basic system services and performing hardware-related tasks, such as framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0085] The network communication module 552 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 520. Exemplary network interfaces 520 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB).
[0086] In some embodiments, the image generation device provided in the embodiments of the present application can be implemented in a software manner. Figure 3 The image generation device 555 stored in the memory 550 is shown, which can be software in the form of a program and a plug-in, etc., including the following software modules: an image generation module 5551, which is logical and can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be explained below.
[0087] In other embodiments, the training device of the image generation model provided in the embodiments of the present application can be implemented in hardware. As an example, the training device of the image generation model provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the training method of the image generation model provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs) or other electronic components.
[0088] In some embodiments, the terminal or server can implement the training method of the image generation model provided in the embodiment of the present application by running a computer executable instruction or a computer program. For example, the computer program can be a native program (e.g., a dedicated training program) or a software module in an operating system, for example, a training module that can be embedded in any program (such as an instant messaging client, an album program, an electronic map client, a navigation client); for example, it can be a local (Native) application (APP, Application), that is, a program that needs to be installed in the operating system to run. In short, the above-mentioned computer program can be an application, module or plug-in in any form.
[0089] The training method of the image generation model provided in the embodiment of the present application will be explained in combination with the exemplary application and implementation of the server or terminal provided in the embodiment of the present application.
[0090] See also Figure 4 , Figure 4 This is a flow chart of the training method of the image generation model provided in the embodiment of the present application. Figure 1 , will combine Figure 4 Steps 101 to 106 are shown for illustration. The training method of the image generation model provided in the embodiment of the present application can be implemented by a server or a terminal alone, or by a server and a terminal in collaboration. The following will be described using the server alone as an example.
[0091] In step 101, based on a text sample, an image generation model is called to generate an image corresponding to the text sample.
[0092] In some embodiments, the image generation model, also known as the text-generated graph model, is a model based on artificial intelligence technology that can generate an image corresponding to an input text description. By learning the association between text and image, the natural language description is converted into visual content. The image generation model is a deep learning model that can generate realistic images from given inputs (such as text, noise vectors, or other forms of conditions). By learning a large amount of image and text data, the image generation model can understand the semantic information in the text description and map it into the image space to generate visual content that matches the description. Typical applications include artistic creation, image restoration, virtual scene generation, advertising design, etc.
[0093] In some embodiments, in computer science and natural language processing (NLP), it usually refers to text data used to train, verify or test machine learning models. These text data can be in the form of sentences, paragraphs, articles, etc., which are used to train models to understand the structure, semantics and context of language. The quality and diversity of text samples are crucial to the learning effect of the model because the model will learn how to understand the text based on these samples.
[0094] In some embodiments, an image is a representation of visual information captured by digital technology or optical devices. In the field of computer vision and image processing, an image is represented as a collection of pixels, each of which has a specific color value. An image can be static (such as a photo) or dynamic (such as a frame in a video). Image data is widely used in various fields, including medical imaging, satellite imagery, self-driving cars, face recognition, artistic creation, etc. The image corresponding to the text sample is used to display the picture content expressed by the text sample.
[0095] As an example, in the application scenario of product display image generation on an e-commerce platform, an e-commerce platform hopes to provide merchants with a function to quickly generate product display images. Merchants only need to enter the text description of the product, and the platform can automatically generate high-quality product images for product detail pages or advertising. The merchant enters a product description, for example: a cute white cat doll wearing a red bow sitting on green grass with a blue sky and white clouds in the background. Call the text image model and take the text description as input. The model generates an image through the following steps: Text encoding: Convert the text description into a semantic vector. Image generation: Generate an image that matches the description based on the semantic vector. Image optimization: Optimize the details of the generated image to ensure high-quality output. Generate an image that highly matches the text description: The image contains a white cat doll wearing a red bow sitting on green grass with a blue sky and white clouds in the background.
[0096] As an example, in the application scenario of game scene design, a large number of scenes need to be designed for the game, such as forests, castles, deserts, etc. Designers hope to quickly generate scene concept maps through text descriptions. Input text sample: The designer enters a scene description, for example: a mysterious magic forest with tall and dense trees, fluorescent mushrooms on the ground, and a glowing ancient stone tower in the distance. The game development tool calls the Wensheng graph model to convert the text description into an image. The Wensheng graph model generates a magic forest scene image that meets the description. The image shows a dense forest, fluorescent mushrooms on the ground, and a glowing stone tower in the distance. The generated image can be used as a concept map of the game scene for designers to further optimize or directly used in game development.
[0097] As an example, see Fig.13 , Fig.13 is a structural diagram of an image generation model provided in an embodiment of the present application, Fig.13It is a structural diagram of the image generation model provided by the embodiment of the present application. The text sample first extracts high-dimensional semantic features through a text encoder (such as CLIP or BERT) to generate a Q matrix (Query matrix). The matrix carries the global semantic information of the text as the conditional input of the subsequent attention mechanism. At the same time, the text features are mapped to the image latent space through a fully connected layer and a potential projection to generate initial hidden variables. The Q matrix of the text interacts with the K matrix (Key) and the V matrix (Value) generated by the image features. The semantic relevance between the text and the image is calculated by dot multiplication, and the attention weight matrix is obtained after scaling and Softmax normalization. The weight distribution is adjusted in combination with specific parameters (such as the number of 4C channels) to enhance the feature response of the key area. The weighted V matrix is mapped through a dense layer to fuse the text semantics with the local features of the image and output context embedding. Context embedding extracts deep features through multiple residual blocks, and each residual block is followed by a spatial transformation layer (such as a deformable convolution) to dynamically adjust the spatial structure of the feature map to adapt to the complex geometric relationship described by the text. The resolution is compressed by downsampling, while retaining the jump connection of low-level feature details (such as edges and textures) for detail recovery in the subsequent upsampling stage. If the model is a diffusion model, the latent variables will be gradually denoised according to the number of time steps: Time step input: The noise prediction network at each time step receives the current latent variables and text conditions, and predicts the noise through convolutional layers and residual blocks. Conditional injection: Text features dynamically affect the denoising direction through cross attention (QKV mechanism) to ensure that the generated content is consistent with the text semantics. The denoised latent variables are dimensionalized through potential projections, and then upsampled through multi-level convolutional layers to restore the resolution. Low-level features are fused with high-level semantic features through jump connections to supplement details (such as object texture and color). The last layer of convolution generates an RGB image to complete the mapping from text to image.
[0098] In step 102, text conversion is performed on the image label of the text sample to obtain a first text corresponding to the image label.
[0099] In some embodiments, the first text is used to describe the content in the image label. A text sample refers to an input descriptive text, which usually contains a linguistic description of the image content. For example: a brown puppy sitting on the grass with a background of blue sky and white clouds. An image label refers to image data corresponding to a text sample, usually a picture. For example, a picture shows a brown puppy sitting on the grass with a background of blue sky and white clouds. Text conversion refers to converting an image label (in image form) into a text description through an image understanding model (such as an image annotation model or a visual language model). For example, converting an image label (in image form) into text: a brown puppy sitting on the grass with a background of blue sky and white clouds.
[0100] In some embodiments, the above text conversion can be implemented by an image understanding model, which is used to convert the input image into text. Assume that the image label is a picture showing a brown puppy sitting on the grass with a blue sky and white clouds in the background. Use a pre-trained visual language model (such as CLIP, BLIP or image annotation model), that is, an image understanding model to analyze the image and extract semantic information from the image. The image understanding model generates a detailed text description based on the image content, for example: a brown puppy sitting on the grass with a blue sky and white clouds in the background.
[0101] In some embodiments, the image label is an image-form supervisory signal used to guide the training of the image generation model, which is essentially a real image strictly paired with the text sample. During the conventional training process, the image generation model generates images with text samples as input and calculates the loss by the difference with the image label (real image), thereby optimizing the generation ability of the model.
[0102] In some embodiments, the image understanding model may include an image encoding layer and an image understanding layer. The above-mentioned text conversion of the image label of the text sample to obtain the first text corresponding to the image label can be achieved as follows: calling the image encoding layer, encoding the image label of the text sample to obtain an image label vector, calling the image understanding layer, decoding the image label vector, and obtaining the first text corresponding to the image label.
[0103] In some embodiments, the image understanding model is generally intended to enable a computer to recognize and understand the content in an image. The concepts of the image encoding layer and the image understanding layer mentioned here are key components for parsing and converting images. The main task of the image encoding layer is to convert the input image (in this case, the image form of the image label) into a vector representation that can be processed by machine learning and deep learning models. When an image label (for example, a picture containing a text label) is input into the image encoding layer, the layer extracts features from the image through a series of convolutional neural networks (CNNs) or other neural network structures, and finally obtains a vector of fixed dimensions, called an image label vector. This vector is an abstract representation of the image content. The image understanding layer receives the image label vector output by the image encoding layer and decodes it to generate a corresponding text description. The image understanding layer usually includes one or more sequence models such as recurrent neural networks (RNNs) or Transformers, which can process sequence data and convert image label vectors into text sequences. The specific steps are as follows: The image understanding layer maps each vector element to a specific character or word based on the sequence generated by the image label vector. The image label vector is decoded step by step to generate a text sequence, that is, the first text corresponding to the image label.
[0104] In step 103, the content difference between the first text and the text sample is obtained.
[0105] In some embodiments, content differences refer to differences in text, meaning, information, structure, etc. when comparing two texts (i.e., the first text and the text sample). For example, vocabulary differences: differences in words or phrases used in the two texts. Grammatical differences: differences in sentence structure or grammatical rules. Semantic differences: differences in meaning between the two texts. Information differences: differences in information or details contained in the two texts. Style or tone differences: differences in style, tone, or expression between the two texts.
[0106] In some embodiments, see Figure 5 , the first text and the text sample both include a plurality of words, Figure 5 This is a flow chart of the training method of the image generation model provided in the embodiment of the present application. Figure 2 , Figure 4 The illustrated step 103 may be performed by Figure 5 Steps 1031 to 1033 are shown to be implemented.
[0107] In step 1031, the similarity between each of the words in the first text and the text sample is determined respectively.
[0108] As an example, the first text includes word A, word B, word C and word D. The similarity between word A in the first text and the text sample is determined, the similarity between word B in the first text and the text sample is determined, the similarity between word C in the first text and the text sample is determined, and the similarity between word D in the first text and the text sample is determined.
[0109] In some embodiments, the above step 1031 can be implemented in the following manner: performing the following processing for each word in the first text: when the word is the first word in the first text, encoding the word and the text sample respectively to obtain a first vector and a second vector, and determining the similarity between the first vector and the second vector as the similarity between the word and the text sample; when the word is not the first word in the first text, constructing a second text based on the word and the text before the word in the first text, and determining the similarity between the second text and the text sample as the similarity between the word and the text sample.
[0110] In some embodiments, for each word in the first text, it is first necessary to encode it. Encoding generally refers to converting a word into a representation in the form of a vector. For the entire first text sample, it is also necessary to encode it to obtain a vector, called a second vector. This vector can be obtained by averaging the vectors of all words, such as text embedding technology, to capture the overall semantics of the entire text. When the word is the first word in the first text, the vector of the word (first vector) is directly calculated for similarity with the vector of the entire text sample (second vector). The similarity calculation can use cosine similarity, Euclidean distance or other similarity measurement methods to determine the similarity between the two vectors. This similarity value represents the similarity between the word and the entire text sample. When the word is not the first word in the first text, a new text sample needs to be constructed, called a second text. This second text is constructed based on the word and the text content before it. For example, the word and its first few words (or the previous sentence) can be taken as a new text sample. Encode this second text to obtain a new vector to represent this constructed text sample. Then, the new vector (second vector) is subjected to similarity calculation with the original text sample vector (second vector). The calculated similarity value represents the similarity between the word and the original text sample.
[0111] In some embodiments, the first vector refers to a vector representation of a single word in the text converted through word embedding technology or other encoding methods. This vector captures the semantic information of the word and can represent the characteristics of the word in a high-dimensional space. In the described process, the first vector specifically refers to the vector obtained by encoding the word currently being processed. The second vector refers to the vector representation of the entire text sample (such as a sentence or paragraph) converted through text embedding technology or other encoding methods. This vector represents the overall semantic information of the entire text sample. In the described process, the second vector specifically refers to the vector obtained after encoding the original text sample. Words are the basic language units in the text, usually referring to individual words. In the described process, each word is a single word in the first text sample, such as one, agile, brown, etc. in a sentence.
[0112] In some embodiments, encoding refers to the process of converting words or text samples in text into numerical vector representations. In natural language processing, encoding usually uses word embedding models, such as Word2Vec, GloVe, etc. These models can map words into continuous vector spaces so that semantically similar words are closer in the vector space. Encoding can be at the word level, or at the sentence or paragraph level. In the described process, encoding includes both encoding of individual words and encoding of entire text samples.
[0113] As an example, suppose there is a first text sample in Chinese as follows: First text sample: A quick brown fox jumped over a lazy dog. The server will process each word according to the following steps to process the first word: The first word is a. The server uses the word embedding model to encode a and obtain the first vector: [0.1, 0.2, 0.3]. The server encodes the entire first text sample A quick brown fox jumped over a lazy dog and obtains the second vector: [0.5, 0.6, 0.7]. The server calculates the similarity of the two vectors, such as using cosine similarity, assuming that the calculation result is 0.85. This similarity value of 0.85 is the similarity between a and the entire text sample. The second word is agile. The server uses the word embedding model to encode agile and obtains the first vector: [0.4, 0.5, 0.3]. The server constructs the second text based on the second word agile and the text one before it: agile. The server encodes this second text agile and obtains a new second vector: [0.55, 0.65, 0.45]. The server calculates the similarity between this new second vector and the second vector of the original text sample, assuming the result is 0.80. This similarity value of 0.80 is the similarity between agile and the entire text sample. Process the third word: The third word is brown. The server encodes brown and obtains the first vector: [0.3, 0.2, 0.8]. Based on the third word brown and the text before it, an agile brown, construct the second text: an agile brown. The server encodes this second text and obtains a new second vector: [0.6, 0.7, 0.5]. Calculate the similarity, assuming the result is 0.75.
[0114] In this way, the semantic association between each word in the text and the overall text can be captured in detail, thus providing a more refined perspective for text analysis. In particular, for understanding the importance of key words in the text and their contribution to the overall semantics, when processing words other than the first word, by constructing a second text containing the word and its previous text, the impact of contextual information on the meaning of the word can be considered. This processing method makes the similarity calculation more consistent with the contextual characteristics of natural language, and improves the accuracy and reliability of the similarity calculation. By analyzing the similarity between each word and the text sample, the main content and key information of the text can be more effectively identified, providing support for subsequent text understanding and application.
[0115] In step 1032, a word whose similarity with the text sample is less than a similarity threshold is determined as the first word.
[0116] In some embodiments, the similarity threshold is a preset value used to determine whether the similarity between two vectors is high enough. If the similarity between two vectors is lower than this threshold, they are considered to be semantically far apart. If the similarity between a word and the entire text sample is very low, it means that the word may not contribute much to understanding the semantics of the entire text sample, or it is not strongly related to the subject content of the text sample. The word whose similarity with the text sample is less than the similarity threshold is determined as the first word, thereby determining the key words in the text, i.e., those words that are highly relevant to the overall semantics of the text sample. Noise or irrelevant words are filtered out, thereby improving accuracy and efficiency in tasks such as text analysis, information retrieval or text generation.
[0117] In step 1033, a second word corresponding to the first word is selected from the text sample, and the first word and the second word are used as the content difference between the first text and the text sample.
[0118] In some embodiments, the second word selected from the text sample corresponding to the first word may be the second word in the text sample having the same position as the first word in the first text, and the second word corresponding to the first word may mean that the position of the first word in the first text is the same as the position of the second word in the text sample. The correspondence between the first word and the second word may refer to a vocabulary mapping relationship between the first word and the second word, and this vocabulary mapping relationship may be based on semantic similarity, context association, vocabulary replacement or other linguistic rules. The second word in the text sample corresponding to the first word in the first text may mean that there is a word in the second text sample that corresponds to a word in the first text sample in meaning or function.
[0119] As an example, assume that there are two text samples in the embodiment of the present application: the first text sample: "I like to eat apples." The second text sample: "He likes to eat bananas." In this example, "apple" and "banana" are two words with a corresponding relationship. They are both names of fruits, but are mentioned in different text samples. This correspondence may be based on the category of vocabulary (both are fruits), context (both are discussing favorite fruits), or semantic similarity (both are common fruit choices).
[0120] In some embodiments, the first word refers to a specific word that appears in the first text sample. This word may be a keyword, a subject word, or other words that are of great significance in the text. The selection of the first word is usually based on the content, purpose, or analysis needs of the text. The second word refers to a word in the second text sample that corresponds to the first word in the first text sample. This correspondence may be a direct synonym replacement or an indirect correspondence based on semantic similarity or contextual association. The selection of the second word is intended to reflect the equivalent expression or related concept of the first word in the second text sample.
[0121] In some embodiments, some words may be determined as first words, that is, their similarity is lower than a preset similarity threshold, indicating that these words have a weak semantic association with the entire text sample. A word is found in the text sample that is semantically corresponding or close to the word determined as the first word. This correspondence may be based on the semantic role, context or word meaning similarity of the words. For example, if the first word is a cat, and a dog appears in the text sample, and they play a similar semantic role in the text, then the dog can be regarded as a second word corresponding to the cat. Alternatively, the words in the text sample correspond one-to-one to the words in the first text, then there is a second word in the text sample that corresponds one-to-one to the first word, and the first word and the second word are used as the content difference between the first text and the text sample, which means that the two words are used to represent the main difference or difference between the two texts. This difference may be due to their semantic differences, or because of their different roles and importance in the text.
[0122] As an example, suppose there is a first text sample in Chinese as follows: A quick brown fox jumped over a lazy dog. The server performs the following operations: Determine the similarity between each word and the text sample: The server encodes each word, for example: one: [0.1, 0.2], only: [0.3, 0.4], quick: [0.5, 0.6], brown: [0.7, 0.8], fox: [0.9, 1.0], jumped: [1.1, 1.2], lazy: [1.3, 1.4], dog: [1.5, 1.6], At the same time, the server encodes the entire text sample to obtain a second vector: [1.7, 1.8]. The server calculates the similarity of each word vector with the second vector, for example using cosine similarity, and obtains the following results: one: 0.90, only: 0.85, agile: 0.75, brown: 0.65, fox: 0.95, skipped: 0.80, lazy: 0.70, dog: 0.60. The similarity threshold is set to 0.8. The server identifies words with a similarity less than 0.8, namely brown and lazy, which are determined to be the first words. The server selects the second word corresponding to brown from the text sample, assuming it is lazy, because they both describe the characteristics of animals in the text. The server selects the second word corresponding to lazy from the text sample, assuming it is agile, because they are semantically relative. The server uses brown and lazy as the content differences between the first text and the text sample, and records the corresponding second words lazy and agile.
[0123] In this way, by determining the similarity between each word in the first text and the text sample, and taking the words with similarity below the threshold as the first words, the server can accurately identify the key differences in the text. The refinement level of text analysis is improved, allowing the server to capture the contribution of each word in the text to the overall semantics, thereby gaining a deeper understanding of the text content; by determining the first word, the server can filter out words that are weakly related to the subject content of the text sample or semantically mismatched, which helps filter noise and enhance the robustness of text processing; selecting the second word corresponding to the first word and taking the two as content differences not only reveals the significant difference between the first text and the text sample, but also provides an important semantic basis for applications such as text comparative analysis, information extraction, and text generation, thereby optimizing the performance and effect of natural language processing-related tasks.
[0124] In some other embodiments, the above step 103 can also be implemented in the following manner: encoding the first text to obtain a third vector, and encoding the text sample to obtain a fourth vector; determining the distance between the third vector and the fourth vector, and when the value of the distance is not equal to zero, obtaining the content difference between the first text and the text sample.
[0125] In some embodiments, the third vector obtained by encoding the first text represents the overall semantic information of the first text. This is usually achieved by combining all word vectors in the text in some form (such as averaging, using an attention mechanism, etc.). The fourth vector obtained by encoding the text sample also represents the overall semantic information of the text sample. Calculating the distance between two vectors is a way to measure the similarity of the contents of the two texts. Common distance metrics include Euclidean distance, cosine distance, etc. This distance value reflects the degree of difference between the first text and the text sample in the semantic space. If the two vectors are very close, the distance between them will be small; if they differ greatly, the distance will be larger. When the value of the distance is not equal to zero, it means that there is a difference in content between the first text and the text sample. The larger the distance value, the more significant the content difference is.
[0126] In some embodiments, the third vector refers to the vector representation obtained by encoding the first text (i.e., the text to be analyzed) by word embedding technology or other encoding methods. This vector integrates the semantic information of all words in the first text, and is usually formed into a single vector by merging the vectors of individual words (e.g., taking the average, maximum value, or using a more complex model such as an attention mechanism). The third vector can be regarded as a point in the semantic space of the first text, which reflects the overall semantic content and theme of the text. The fourth vector refers to the vector representation obtained by encoding the text sample (i.e., the reference text for comparison) by the same encoding method. Similar to the third vector, the fourth vector also contains the semantic information of all words in the text sample and merges them into a vector. This vector represents the position of the text sample in the semantic space and can be used to compare with the vector of the first text to analyze the semantic relationship and difference between the two.
[0127] In some embodiments, the distance between the third vector and the fourth vector may refer to the Euclidean distance, Manhattan distance, or cosine similarity between the third vector and the fourth vector. If cosine similarity is selected, the dot product of the two vectors is first calculated, and then the modulus of each vector is calculated separately, and finally the dot product is divided by the product of the modulus of the two vectors. The value range of cosine similarity is between -1 and 1. The closer the value is to 1, the more similar the two vectors are. The distance can be defined as 1 minus the cosine similarity. If Manhattan distance is selected, the distance calculation formula is the sum of the absolute values of the differences between the corresponding elements of the two vectors. If Euclidean distance is selected, the distance calculation formula is the square root of the sum of the squares of the differences between the corresponding elements of the two vectors.
[0128] As an example, suppose there are two Chinese texts, the first text is a sentence, and the text sample is another sentence. The server will perform the following operations: First text: A fox is playing in the forest. Text sample: A cat is resting on the grass. The server uses the word embedding model to encode the first text A fox is playing in the forest. Each word is converted into a vector, and then these vectors are combined in some way (such as averaging) to obtain a third vector representing the entire first text. Assume that the third vector is: [0.1, 0.2, 0.3]. Similarly, the server uses the same word embedding model to encode the text sample A cat is resting on the grass., and combines the encoded vectors to obtain a fourth vector representing the text sample. Assume that the fourth vector is: [0.4, 0.5, 0.6]. The server calculates the distance between the third vector and the fourth vector, using the Euclidean distance as the metric. Because the calculated distance value is not equal to zero (actually 0.52), the server determines that there is a content difference between the first text and the text sample. This content difference indicates that although both texts are describing animal behavior, they refer to different animals (foxes and cats) and different locations of activity (forests and meadows), which leads to an overall semantic difference.
[0129] Thus, by encoding the first text to obtain the third vector and encoding the text sample to obtain the fourth vector, the server can effectively grasp the position of the two texts in the semantic space. Calculating the distance between the two vectors not only provides the server with a means to quantify the similarity between the first text and the text sample, but also directly reveals the content difference between the two when the distance value is not equal to zero. It provides an objective metric for text similarity assessment, which helps to improve the accuracy of text classification, retrieval and recommendation; the identification of content differences enables the server to better understand the deep meaning and nuances of the text, which is crucial for text analysis and information extraction tasks.
[0130] In step 104, an image mask of the image tag is generated based on the content difference.
[0131] In some embodiments, the image mask is a binary image of the same size as the original image, where the value of each pixel indicates whether the pixel is covered (usually 0 for covered and 1 for uncovered). In deep learning, masks can also be used for self-supervised learning tasks, such as masked image modeling (MAE), where the model needs to reconstruct the original image based on the masked part.
[0132] In some embodiments, see Figure 6 , the first text and the text sample both include a plurality of words, Figure 6 This is a flow chart of the training method of the image generation model provided in the embodiment of the present application. Figure 3 , Figure 4 Step 104 shown may be performed by Figure 6 Steps 1041 to 1042 are shown to be implemented.
[0133] In step 1041, target segmentation is performed on the image label based on the content difference to obtain a target segmentation result of the image label.
[0134] In some embodiments, the target segmentation result includes a first image region and a second image region, wherein the first image region corresponds to the content difference, and the second image region is different from the first image region. The image is segmented to obtain two different regions, the first image region and the second image region, wherein the first image region corresponds to regions with content differences, and the second image region is different from the first image region and may correspond to regions with no content differences or with other features.
[0135] In some embodiments, corresponding to content differences means that during the target segmentation process, those areas in the image that are visually related to the target object are identified. These areas usually have similar colors, textures or other features as the target object, and are significantly different from the background areas. The first image areas are these areas corresponding to content differences, which are considered to be part of the target object. The first image area is an area in the target segmentation result that is directly related to the content difference. The first image area is the part of the image that corresponds to the text content difference. In other words, it is those areas in the image that can reflect the text difference. The content difference may reflect some unique information or changes in the text, and the first image area is the embodiment of this information in the image.
[0136] As an example, suppose the first text describes a scene: there is a red apple in the picture, and the text sample describes it as: there is a green apple in the picture. The content difference is in red and green. In image segmentation, the first image region may be the region where the red apple is located in the image, because it corresponds to the content difference (different color) in the text.
[0137] As an example, the first text description: There is a fruit in the picture, and the text sample description: There is an apple in the picture. The content difference lies in the specificity of the fruit and the apple. The first image region may be a broader region of fruit in the image (such as a fruit basket), rather than a specific apple region.
[0138] As an example, the first text describes the details of a scene, while the text sample describes the whole. The content difference may be reflected in the segmentation of the local and the whole of the image. The first image region may be a local area in the image, such as an apple in a fruit basket, rather than the entire fruit basket.
[0139] In some embodiments, object segmentation is a key task in the field of computer vision and image processing, which involves assigning each pixel or region in an image or video to a specific object or background. The purpose of object segmentation is to accurately identify and separate the various objects in the image so that each object has a clear boundary and can be analyzed and processed independently.
[0140] In some embodiments, through the target segmentation process, the result obtained by the server is to segment the image into different areas, which are marked and classified according to the information of the image tag. The target segmentation result divides the image into at least two areas, referred to herein as the first image area and the second image area. The first image area corresponds to the content difference, and the first image area is associated with the text content difference because this area visually matches the features described by the text content difference. For example, if the content difference indicates that the first text describes a dynamic scene and the text sample describes a static scene, then the first image area may contain dynamic objects or scenes. The second image area is different from the first image area in visual content, and it may contain image content that is not directly related to the content difference between the first text and the text sample.
[0141] In some embodiments, the first text includes multiple words, and the content difference includes a target word among the multiple words. The above step 1041 can be implemented in the following manner: for each of the words in the first text, determine the first pixel point corresponding to the word from the image label; based on the first pixel point corresponding to each of the words, perform region expansion on the first pixel point corresponding to the target word to obtain the first image area; and determine the image area in the image label that is different from the first image area as the second image area.
[0142] In some embodiments, the first text is composed of multiple words, each of which may have a different contribution to the semantics of the text. Content difference analysis determines the differences between the first text and the text sample, and these differences include some specific target words, which play a key role in the generation of differences in semantics. The server uses the first pixel points corresponding to each word determined previously, especially the pixel points corresponding to the target word, to perform region expansion. This expansion may include morphological operations (such as dilation) to merge all related pixels into a continuous area. This area represents the visual part of the image corresponding to the target word, that is, the first image area. In the image, all areas except the first image area are classified as second image areas. These areas do not directly correspond to the target words and may represent the background, other objects in the image, or visual elements related to the text content difference.
[0143] As an example, the first text: a horse is running on a grassland, and there are birds flying in the sky. The image label is an image containing a horse running on a vast grassland, and a bird is flying in the sky. The multiple words in the first text include a horse, in, on the grassland, running, in the sky, and flying birds. Assume that the content difference analysis determines that horse and bird are target words. The server analyzes the image label and finds the first pixel points corresponding to the horse, which constitute the shape of the horse in the image. Similarly, the server finds the first pixel points corresponding to the bird, which constitute the shape of the bird in the image. The server uses morphological operations (such as dilation) to expand the first pixel point area corresponding to the target word horse. This merges all the pixels related to the horse to form a continuous image area, namely the first image area. For the target word bird, the server also performs area expansion to obtain the image area of the bird. The first image area includes all the pixels related to the horse and the bird, which are marked. The remaining part of the image label, that is, the pixels not including the horse and the bird, is determined as the second image area. This may include grassland, sky and other objects that may appear in the image.
[0144] In this way, by associating the words in the first text with the pixels in the image label and expanding the region based on these associations, the server can accurately identify and segment the parts of the image that are directly related to the text content. It greatly improves the accuracy of image understanding, not only identifying objects in the image, but also associating these objects with specific concepts in the text description, thereby providing richer semantic information for image annotation, classification and retrieval. It makes image analysis more flexible, and the importance of image regions can be adjusted according to the differences in different text content, thereby optimizing the results of image processing tasks. By distinguishing the first image region from the second image region, the server can effectively distinguish the main content and background information in the image.
[0145] In some embodiments, the above-mentioned determination of the first pixel point corresponding to the word from the image label can be achieved in the following manner: based on the word, feature extraction is performed on the image label to obtain image features corresponding to the word, and the image features include multiple feature elements; pixel points corresponding to the feature elements are selected from the image label; and the pixel point corresponding to the feature element with the largest value is determined as the first pixel point corresponding to the word.
[0146] In some embodiments, the feature elements in the image features may correspond to pixels one-to-one, or the feature elements correspond to at least one pixel. In the image tag, the embodiment of the present application selects pixels based on the feature elements. The feature elements here may be specific patterns, shapes, colors or other attributes in the image. The correspondence between feature elements and pixels may be one-to-one, that is, each feature element corresponds to one pixel; or one-to-many, that is, one feature element corresponds to multiple pixels.
[0147] In some embodiments, feature extraction of image tags can be achieved through a cross-attention mechanism, and the input text (words) and image (image tags) need to be encoded into feature representations respectively. For text, word embedding is usually used to convert it into a dense vector; for images, convolutional neural networks (CNNs) are usually used to extract image features. The cross-attention mechanism is used to calculate the attention weights between text and image features. This usually involves calculating the similarity between text features and image features, for example, using dot products or scaled dot products. The higher the similarity, the greater the attention weight, indicating that the association between text and image features is stronger. According to the calculated attention weights, the image features are weighted and summed to obtain image feature representations related to text features. This process can be regarded as the text features guiding the attention of image features, so that the model can focus on important image areas related to the text. Image features corresponding to the text are extracted from the weighted image features. These features may include multiple feature elements, such as color, texture, shape, etc., which together constitute a visual representation of the text description.
[0148] In some embodiments, the server performs feature extraction on the image tags based on the words in the first text. This means that the server will use image processing algorithms (such as convolutional neural networks, SIFT, HOG, etc.) to identify visual features in the image that are related to these words. The feature extraction process will produce a series of image features, which can be represented as multi-dimensional feature vectors, where each dimension is a feature element. These feature elements reflect the visual information related to specific words in the image. Once the features are extracted, the server needs to associate each feature element in the feature vector with a specific pixel in the image. The server maps the feature vector to the spatial structure of the image and determines the position of each feature element in the image. The server evaluates the value of each feature element in the feature vector and looks for the feature element with the largest value. This feature element with the largest value represents the visual feature in the image that is most strongly or most significantly related to the word. The server determines the pixel corresponding to the feature element with the largest value as the first pixel corresponding to the word. This pixel represents the position of the visual element in the image that is most related to the word.
[0149] As an example, suppose there is an image whose image label is a cat on the grass, the first text of the embodiment of the present application is a cat playing, and the target word of interest is cat. The server will perform feature extraction on the image, especially looking for features related to the word cat. This may involve using a pre-trained deep learning model, such as a convolutional neural network (CNN), to identify features such as the shape, texture, and color of the cat in the image. After feature extraction, the server will obtain a feature vector containing multiple feature elements. For example, the feature vector may have 100 elements, each of which represents the feature intensity at different positions in the image. These feature elements may be related to different parts of the cat (such as the head, body, tail) or the overall shape. Each feature element is associated with a specific pixel or pixel area in the image. The server maps these feature elements back to the image based on the spatial position information of the feature vector and finds the corresponding pixel points. The server calculates the value of each feature element in the feature vector and finds the maximum value. Assume that the value of the 45th feature element is the largest, indicating that it is most strongly associated with the concept of cat. The server then finds the pixel corresponding to this maximum feature element, which is located at a certain position in the image, such as the head of the cat. This pixel will be determined as the first pixel corresponding to the word "cat". The feature element values of the cat's head area in the image are [0.1, 0.2, 0.9, 0.3, ...], among which the value of the 45th element is 0.9, which is the highest among all feature elements. Therefore, the server will select the pixel corresponding to the 45th feature element as the first pixel. This point is located in the cat's head area and it best represents the position of the word "cat" in the image.
[0150] In this way, the visual elements in the image that are most relevant to the target word can be accurately located, thereby improving the accuracy and efficiency of understanding the image content. By analyzing the characteristic elements of the image, the server can distinguish the main content and minor details of the image, which is crucial for image recognition and classification tasks. Selecting the pixel with the largest eigenvalue as the first pixel ensures that the server can focus on the most significant and representative part of the image, which helps to improve the robustness of image segmentation and target detection.
[0151] In some embodiments, the above-mentioned area expansion of the first pixel point corresponding to the target word based on the first pixel point corresponding to each of the words to obtain the first image area can be achieved in the following manner: obtaining the second pixel point corresponding to the target word, the second pixel point belongs to the first pixel point corresponding to each of the words, and the second pixel point is adjacent to the first pixel point corresponding to the target word; based on the first pixel point corresponding to each of the words, determining the relative position relationship between the first pixel point corresponding to the target word and the boundary of the image label; based on the relative position relationship and the second pixel point, area expansion of the first pixel point corresponding to the target word to obtain the first image area.
[0152] In some embodiments, after determining the first pixel corresponding to the target word, the server needs to find the second pixel. The second pixel is those pixels that are adjacent to the first pixel and also belong to the target word. These pixels are usually located around the first pixel and may include a continuous area in the image related to the target word. The server then analyzes the relative position relationship between the first pixel corresponding to each word and the boundary of the image label. Check whether the first pixel is close to the edge, corner or other key position of the image, which is crucial for subsequent area expansion. The server uses the relative position relationship of the first pixel and the information of the second pixel to guide the area expansion process. For example, if the first pixel is located in the central area of the image, the server may use a wider neighborhood for expansion; if the first pixel is located at the edge of the image, the expansion may be more conservative to avoid exceeding the image boundary. Area expansion is usually achieved through morphological operations (such as dilation), which gradually merges adjacent pixels into the area of the first pixel until a certain condition is met, such as encountering a non-target pixel or image boundary.
[0153] In some embodiments, the relative position relationship refers to the spatial position of a pixel (such as the first pixel corresponding to the target word) relative to the image label boundary. This relationship describes the specific position of the pixel in the image, for example, it may be located at the center of the image, near the boundary, a corner point, or a specific area of the image. This position information is crucial for region expansion in image processing because it can help determine the direction and range of region expansion.
[0154] In some embodiments, the second pixel refers to a pixel adjacent to the first pixel in the image. These pixels are usually closely connected to the first pixel in space, and may share an edge or be within a certain neighborhood. In image processing, the second pixel is usually used for region expansion based on neighborhood information because they have a direct visual connection with the first pixel. In image processing, the boundary refers to the edge of the image or the dividing line between different areas in the image. The boundary specifically refers to the edge of the image label, that is, the physical boundary of the image or the boundary between a specific object in the image (such as the object represented by the target word) and the background. Boundaries are critical for image segmentation and object recognition because they mark important transition points in the image.
[0155] In some embodiments, region expansion is an operation in image processing that expands a specific image region by fusing adjacent pixels (such as a second pixel) around a determined pixel (such as a first pixel). This operation is usually implemented using morphological operations (such as dilation and erosion) to merge scattered pixels into a larger, continuous region to better represent objects or features in the image. Region expansion is based on the relative position relationship and the second pixel to accurately control the expansion of the first pixel corresponding to the target word so as to obtain a first image region related to the target object.
[0156] As an example, suppose that there is an image in the embodiment of the present application, the image label is a dog playing in the park, and the first text of the embodiment of the present application is a dog chasing a ball. In this example, the target word of the embodiment of the present application is dog. The server determines the first pixel corresponding to the word dog through feature extraction and matching process. Assume that this point is located at the head of the dog in the image. The server will find the second pixel adjacent to the first pixel. These second pixels are adjacent to the first pixel, for example, they may be pixels around the dog's head, such as ears, eyes or nose. The server will analyze the relative position relationship between the first pixel and the image label boundary. For example, if the first pixel is located on the dog's head, the server will determine the distance between the head pixel and the image boundary and whether it is close to the edge of the image. The server expands the region of the first pixel based on the relative position relationship between the first pixel and the second pixel: if the first pixel is close to the left boundary of the image, the server may be more inclined to expand to the right to avoid exceeding the image boundary. The server uses a morphological dilation operation, starting from the first pixel, and gradually includes the adjacent second pixel to form a larger area. Until all adjacent second pixel points are included in the region, or a preset condition is reached, such as the region size limit or the image boundary is encountered.
[0157] Continuing with the previous example, the first pixel (the pixel of the dog's head) is expanded to include the dog's ears, eyes, and nose pixels (the second pixel) into the region. Since the first pixel is close to the left edge of the image, the expansion process is mainly carried out to the right and downward to avoid the region exceeding the image boundary. The resulting region (the first image region) clearly outlines the dog's head.
[0158] In this way, the accuracy and efficiency of the region expansion process can be ensured, because it is expanded based on the spatial proximity of pixels, which helps maintain the continuity and structural integrity of objects in the image. By obtaining a second pixel adjacent to the first pixel, the server can more finely control the direction and range of region expansion, avoiding excessive or unnecessary expansion, thereby improving the quality of image segmentation. Considering the relative position relationship enables the server to adapt to the limitations of the image boundary and avoid exceeding the actual range of the image when expanding the region, which is crucial to maintaining the authenticity and accuracy of the image content.
[0159] In some embodiments, the above-mentioned regional expansion of the first pixel point corresponding to the target word based on the relative position relationship and the second pixel point to obtain the first image area can be achieved in the following manner: when the relative position relationship indicates that there are no other first pixel points between the first pixel point corresponding to the target word and each of the boundaries, the regional boundary of the first image area is determined based on the second pixel point of the target word; when the relative position relationship indicates that there are other first pixel points between the first pixel point corresponding to the target word and part of the boundaries, the regional boundary of the first image area is determined based on the boundary and the second pixel point of the target word; the image area within the regional boundary in the image label is determined as the first image area.
[0160] In some embodiments, when the relative position relationship between the first pixel corresponding to the target word and the image boundary indicates that there are no other first pixels between the first pixel and its adjacent boundary, the server will directly use the second pixel of the target word to determine the boundary of the first image area. This means that the space around the first pixel is empty and there are no other visual elements related to the target word. Therefore, the second pixel will be used to define a closed area that surrounds the first pixel to form the first image area. For example, if the first pixel is the head of a dog in the image and there are no other dog-related pixels around it, the server will determine the boundary of the dog's head based on the second pixel of the dog's head (such as pixels around the ears and eyes).
[0161] In some embodiments, when the relative position relationship indicates that there are other first pixel points between the first pixel point and the boundary, the server needs to consider these additional pixel points and the boundary to jointly determine the boundary of the first image area. In this case, the second pixel point and the other first pixel points together constitute a more complex area, and the server may need to use morphological operations or other image processing techniques to integrate these points to form a continuous area. For example, if the first pixel point is a dog's ear, and there are other dog-related pixel points (such as ears and body) next to it, then the server will determine the first image area containing the entire dog's head based on these pixel points and the image boundary.
[0162] In some embodiments, when there are no other first pixels between the first pixel and the boundary, the server identifies the first pixel corresponding to the target word, such as cat, and finds the second pixel adjacent to it. These second pixels may be pixels on the ears, body or tail of the cat. The server checks the relative position relationship between the first pixel and the image boundary. If it is determined that there are no other first pixels related to the target word around the first pixel, this means that the area is isolated. The server determines the boundary of the first image area based on the second pixel: using a morphological dilation operation, starting from the first pixel, the adjacent second pixel is gradually included to form a closed area. Determine the stopping condition of the dilation, such as reaching a certain area size or encountering an image boundary. Through the morphological closing operation, fill any holes that may exist inside the area to ensure the continuity of the boundary. The final closed area is the first image area, which contains the representation of the target word cat in the image.
[0163] In some embodiments, when there are other first pixels between the first pixel and the partial boundary, the server identifies the first pixel and finds the second pixel adjacent to it. The server finds that there are other first pixels between the first pixel and the partial boundary, which indicates that the target object extends in the image and overlaps with the boundary. The server uses a morphological dilation operation to merge the first pixel and its second pixel to form a preliminary area. The server needs to determine how to deal with the portion that overlaps with the image boundary: if other first pixels are close to the image boundary, the server may dilate along the boundary until it encounters the second pixel or the image boundary. If it is necessary to avoid the region from expanding beyond the image boundary, the server can use a boundary filling algorithm to expand the region only to the inside of the image boundary. The server may use an edge detection algorithm to optimize the region boundary to ensure the accuracy of the boundary. The server determines the area between the internal boundary and the optimized boundary as the first image area, which represents the overall range of the target object in the image.
[0164] As an example, there are no other first pixels between the first pixel and the boundary. The server determines a first pixel located on the cat's head through feature extraction, and identifies the surrounding second pixels, such as the cat's ears and back. The server checks and finds that there are no other first pixels related to the cat around the first pixel, and the first pixel is a distance away from the image boundary. Based on the second pixel of the cat's ears and back, the server performs a morphological dilation operation, starting from the first pixel and expanding until all second pixels are included. This process stops at the image boundary or reaches a certain area size. The final closed area, namely the cat's head and ear area, is determined as the first image area.
[0165] As an example, there are other first pixels between the first pixel and part of the boundary. The server also determines a first pixel located at the cat's head and identifies the surrounding second pixels. There are other first pixels related to the cat at the image boundary near the first pixel, which means that the cat's body extends to the image boundary. The server uses a morphological dilation operation, starting from the first pixel, to include the second pixel to form a preliminary area for the cat's head. Since the cat's body extends to the boundary, the server needs to handle this part specially. It may dilate along the boundary until it encounters the second pixel or the image boundary. If the cat's body is close to the boundary, the server may dilate along the inside of the boundary to ensure that the area expansion does not exceed the image range. The server determines the area between the internal boundary and the optimized boundary as the first image area, which will include the cat's head and part of the body, extending to the inside of the image boundary.
[0166] In some embodiments, the above-mentioned determination of the regional boundary of the first image area based on the second pixel point of the target word can be achieved in the following manner: based on the target word, the image label is specially extracted to obtain the image feature corresponding to the target word, the image feature includes multiple feature elements, and the feature elements correspond one-to-one to the pixel points in the image label; for the second pixel point in each direction of the first pixel point, the pixel point with the smallest value of the feature element between the second pixel point of the target word and the first pixel point of the target word is determined as the boundary pixel point corresponding to the direction, and the area formed by connecting the boundary pixels in each direction is determined as the regional boundary of the first image area.
[0167] In some embodiments, a specific extraction is performed on the label of the image based on the target word. This means that specific features related to the target word will be found in the label information of the image. Here, "specific extraction" may refer to a stable and discriminative feature extraction method that can extract features related to the target word from the image. The obtained image features are composed of multiple feature elements, each of which corresponds to a pixel in the image label. This shows that a feature value can be assigned to each pixel, which can be used to indicate the degree of association between the pixel and the target word. Check the second pixel in each direction of the first pixel to find the pixel with the smallest feature element value between the first pixel and the second pixel of the target word. This minimum value means that this pixel is in a sense the least related to the target word, so it may be a candidate point for the region boundary. Connect the boundary pixels in each direction to form a closed region boundary. This boundary defines the first image region, that is, the image region associated with the target word.
[0168] As an example, suppose that the embodiment of the present application has a target word "wheel," and the task of the embodiment of the present application is to use computer vision technology to determine the boundaries of the image area related to "wheel" in an image containing a variety of different objects. There is an image containing multiple objects, such as vehicles, pedestrians, and traffic signs. The image has been preprocessed and already has a label system that can identify each pixel in the image. For example, the label may be the color, texture, or some high-level features of the pixel. Using the target word wheel, the image features related to the wheel, such as shape, edge, color, etc., are extracted by a special constant extraction method. Assume that the feature extraction method of the embodiment of the present application generates a feature map, which is the same size as the original image, but the value of each pixel now represents the similarity between the point and the wheel feature. Each feature element in the feature map corresponds to a pixel in the original image. Consider a pixel in the original image (referred to as the first pixel in the embodiment of the present application), which is identified as part of the wheel. The pixels around this first pixel (the second pixel) will be checked. For each direction (up, down, left, right, and the angle between them), the pixel with the smallest feature element value will be found. This point represents the place that is least similar to the wheel feature in the current direction, which may be the boundary of the wheel area. A series of boundary pixel points are determined. These boundary pixel points are connected to form a closed contour, which defines the wheel area in the image. For example, if the first pixel point is at the edge of the wheel, then the value of the feature element will decrease rapidly in a certain outward direction, indicating that the wheel area has been left. The pixel point corresponding to this minimum value will be marked as a boundary pixel point. Finally, all these boundary pixel points are connected to clearly outline the outline of the wheel, that is, the regional boundary of the first image area.
[0169] In this way, by performing special extraction of image labels for target words, image features closely related to the target can be effectively obtained, and the correlation between the features and the target object can be enhanced, thereby improving the accuracy of boundary detection in subsequent processing. By searching for the pixel with the smallest feature element value around the first pixel to determine the boundary, not only the accuracy of the boundary is ensured, but also the robustness to image noise and complex scenes is improved. Finally, by connecting these boundary pixels, a clear and continuous area outline can be generated.
[0170] In some embodiments, the above-mentioned determination of the regional boundary of the first image area based on the boundary and the second pixel point of the target word can be achieved in the following manner: based on the target word, the image label is specially extracted to obtain the image feature corresponding to the target word, the image feature includes multiple feature elements, and the feature elements correspond one-to-one to the pixels in the image label; for the second pixel points in each direction of the first pixel point, the pixel point with the smallest value of the feature element between the second pixel point of the target word and the first pixel point of the target word is determined as the boundary pixel point corresponding to the direction; for the boundary where there are other first pixel points between the first pixel point corresponding to the target word, the pixel point with the smallest value of the feature element between the boundary and the first pixel point of the target word is determined as the boundary pixel point corresponding to the boundary; and the area formed by connecting the boundary pixel points is determined as the regional boundary of the first image area.
[0171] In some embodiments, the extracted image features are composed of a plurality of feature elements, which correspond one-to-one with each pixel in the image, indicating the degree to which each pixel is related to the target word. For the first pixel that has been identified as being associated with the target word, its second pixel in each direction will be checked. The feature element values between the first pixel and its second pixel in each direction will be calculated, and the minimum value among these values will be found. The position where this minimum value is located is determined as the boundary pixel in that direction, because it represents the transition from the target area to the non-target area in that direction. When there are a plurality of first pixels associated with the target word, it is necessary to determine the boundary between them. This may be to handle the situation where the target word appears multiple times in the image. For the boundary between these first pixels, the minimum value of the feature element between the boundary and the first pixel of the target word will be found. The position where this minimum value is located is considered to be the pixel of the boundary between these pixels. Once the boundary pixels in all directions are determined, they are connected to form a closed contour. This contour defines the region boundary of the first image region, i.e., the edge of the image region associated with the target word.
[0172] As an example, suppose that the embodiment of the present application has an image containing several different objects, such as a bicycle, a tree, and a dog. The goal of the embodiment of the present application is to use a specific target word "bicycle" to identify and determine the boundary of the bicycle area in the image. The image is preprocessed, and a model is trained using the target word "bicycle" or a trained model is used for feature extraction. The model analyzes the image label (which may include information such as color, shape, texture, etc.) and extracts features related to "bicycle". The resulting feature map is a two-dimensional matrix, each value of which represents the similarity between the corresponding pixel and the "bicycle" feature. Assume that the embodiment of the present application has marked a pixel point A as part of the bicycle, and this point is the first pixel point of the embodiment of the present application. Now check the neighboring pixels (second pixel points) of point A in the up, down, left, and right directions. For example, the pixel point B to the right of point A will be checked. On the line between point A and point B, the feature element value will be calculated and the minimum value will be found. Assuming that the pixel point corresponding to this minimum value is C, then point C is determined to be the boundary pixel point starting from point A in the right direction. If there are multiple pixels marked as bicycle parts in the image, such as point A and point D, the boundary between them needs to be determined. The minimum value of the feature element is searched in the area between point A and point D. Assuming that the minimum value corresponds to pixel E, point E is determined as the boundary pixel between point A and point D. Repeat the above steps until all boundary pixels are determined. Then, connect these boundary pixels to form a closed contour. This contour is the area boundary of the bicycle. On the right side of point A, find point C with the smallest eigenvalue and record it as the boundary point. On the left side of point A, find point F with the smallest eigenvalue and record it as the boundary point. Above point A, find point G with the smallest eigenvalue and record it as the boundary point. Below point A, find point H with the smallest eigenvalue and record it as the boundary point. On the boundary between point A and another bicycle part point D, find point E with the smallest eigenvalue and record it as the boundary point. Connect points C, F, G, H and E to form a closed contour, which is the area boundary of the bicycle in the image.
[0173] In this way, by processing the image label for the target word, the image features corresponding to the target word can be accurately obtained, which improves the correlation between the features and the target object, thereby reducing the possibility of misjudgment and missed judgment in subsequent boundary detection. Secondly, this method can effectively identify the target area by searching for the pixel point with the smallest feature element value in all directions of the first pixel point, and can maintain a high accuracy even in complex or noisy images. In addition, considering the boundary determination between multiple first pixel points, it ensures that the boundary of the entire target area is both continuous and closed, which helps to form a complete area outline.
[0174] In this way, by flexibly handling two different situations - whether there are other first pixels around the first pixel corresponding to the target word, and whether there is overlap with the image boundary, the region expansion can accurately capture the details of the target object and adapt to the limitations of the image boundary. In the first case, the region boundary is determined by the second pixel, which ensures the accuracy of region expansion and the clarity of the object boundary; in the second case, the combination of the boundary and the second pixel information ensures that even in complex image structures, the region can be effectively expanded without exceeding the actual range of the image. The quality of image segmentation is optimized and the ability to understand and analyze image content is improved.
[0175] In step 1042, based on the target segmentation result, the pixel value of each pixel in the first image area is set to a first value, and the pixel value of each pixel in the second image area is set to a second value to obtain an image mask of the image label.
[0176] In some embodiments, the first value is greater than the second value.
[0177] In some embodiments, the server has obtained a first image region through region expansion, which is related to the target word (e.g., cat) and contains a representation of the cat in the image. At the same time, the server also identifies a second image region, which contains all pixels except the first image region, which may be a background or other irrelevant objects. The server sets the pixel value of each pixel in the first image region to a first value. This first value is usually a specific value used to identify that these pixels belong to the target object. For example, this value can be set to 255 (white) in a grayscale image, or a unique color code in a color image. Correspondingly, the server sets the pixel value of each pixel in the second image region to a second value. The second value is used to identify that these pixels do not belong to the target object, and is usually different from the first value for easy distinction. For example, this value can be set to 0 (black) in a grayscale image. The server creates an image mask (also called a binary image or segmentation mask). This mask is an image of the same size as the original image, in which the target area is marked with the first value and the non-target area is marked with the second value. The image mask clearly divides the boundary between the target object and the non-target object.
[0178] As an example, suppose that the embodiment of the present application has an image that contains a clear target object - a red balloon. The task of the embodiment of the present application is to use image processing technology to segment the balloon and distinguish the rest of the background from the balloon. The server first analyzes the image label, which may include steps such as color recognition, edge detection, and texture analysis. In this example, the server may use color recognition technology to distinguish the red balloon from the background. During the analysis process, the server identifies the color characteristics of the balloon and distinguishes it from the color characteristics of the background. For example, it may find that the red pixel value of the balloon area is within a specific range, while the background pixel value is different. Based on the color difference, the server performs target segmentation and obtains a segmentation result. In this result, the balloon pixel points are marked as belonging to the target object, while the background pixel points are marked as not belonging to the target object. The segmentation result may be a binary image, in which the balloon area is marked as 1 and the background is marked as 0. The server sets the pixel values of all pixels marked as balloons to a first value, such as 1, which represents the balloon area in the grayscale image. The server sets the pixel values of all pixels marked as background to a second value, such as 0, which represents the non-balloon area in the grayscale image. Image mask example: Pixel value of balloon area (first image area): 1. Pixel value of background area (second image area): 0.
[0179] In this way, by accurately distinguishing the target object from the background, the efficiency and accuracy of image segmentation are ensured, which helps to extract key information from complex images. Secondly, by setting the pixel value of the first image area to the first value and the pixel value of the second image area to the second value, the image mask not only clearly identifies the target object, but also the binarization processing method greatly reduces the amount of data processing, improves the processing speed, and provides the possibility for real-time image processing. Finally, the generation of the image mask makes the visual contrast between the target object and the background more obvious, which is convenient for manual inspection and editing, thereby improving the visualization and interactivity of the entire image processing process.
[0180] In step 105, the image mask is fused to the image label to obtain a fused image label.
[0181] In some embodiments, the image mask includes multiple mask values, each of which corresponds to a pixel in the image label. Figure 7 , Figure 7 This is a flow chart of the training method of the image generation model provided in the embodiment of the present application. Figure 4 , Figure 4 The illustrated step 105 may be performed by Figure 7 Steps 1051 to 1052 are implemented as shown.
[0182] In step 1051, the pixel value of each pixel point in the image tag corresponding to the image mask is obtained.
[0183] In some embodiments, the pixels in the image tag correspond one-to-one to the mask values in the image mask, and the pixels in the image tag have pixel values, which refer to the numerical value of a single pixel in a digital image. This numerical value usually represents the color or brightness information of the pixel. In image processing and computer vision, the pixel value is the basic data unit for analyzing and processing images. The pixel value is usually composed of multiple components, such as the three components in the RGB (red, green, and blue) color model. The value of each component is also usually between 0 and 255, and they are combined to represent a specific color.
[0184] In step 1052, for each pixel point in the image label corresponding to the image mask, the pixel value of the pixel point is multiplied by the corresponding mask value to obtain the target pixel value of the pixel point, and the pixel value of the pixel point is updated based on the target pixel value to obtain the fused image label.
[0185] In some embodiments, each pixel in the image label corresponds to a mask value in the image mask. This means that each area in the image label has a corresponding mask value to indicate whether it belongs to the target area. For each pixel in the image label, its pixel value (such as 0 or 1) is multiplied by the corresponding mask value. For example, if the label value of a pixel is 1 (indicating that it is part of the target object) and the corresponding mask value is 255, the target pixel value will be 255, indicating that this pixel is enhanced as part of the target area. According to the calculated target pixel value, the pixel value of the corresponding pixel in the image label is updated. This process actually strengthens the target area while weakening or ignoring the background area. The effect of the mask is integrated with the original image label. The new image label not only contains the original category information, but also incorporates the regional importance information provided by the mask.
[0186] As an example, the expression of the above target pixel value can be:
[0187] X=x×y (1)
[0188] Among them, X is used to indicate the target pixel value, x is used to indicate the pixel value of the pixel point, and y is used to indicate the target pixel value.
[0189] In this way, not only the visual representation of the target area in the image label is strengthened, but also the emphasis on the target area is increased, making the image analysis and recognition tasks more accurate and efficient. Through the multiplication operation, the pixel values of the target area are enhanced, while the pixel values of the background or non-target area are weakened or ignored, thereby visually highlighting the key area. This processing not only improves the interpretability of the image, but also the image label fused with the image mask provides richer and more accurate information for image understanding.
[0190] In step 106, the image generation model is trained based on the image and the fused image label.
[0191] In some embodiments, the training of the image generation model based on the image and the fused image label can be achieved in the following manner: based on the image and the fused image label, the model parameters of the image generation model are updated to obtain a target image generation model.
[0192] In some embodiments, see Figure 8 , Figure 8 This is a flow chart of the training method of the image generation model provided in the embodiment of the present application. Figure 5 , Figure 4 Step 106 shown may be performed by Figure 8 Steps 1061A to 1063A are implemented as shown.
[0193] In step 1061A, the image mask is fused to the image to obtain a fused image, and a pixel point in the fused image is used as a third pixel point.
[0194] In some embodiments, a binary image (image mask) is overlaid or superimposed on the original image. The high-value (usually 255 or white) portion of the image mask will overlay the corresponding pixels in the original image, thereby highlighting these locations; while the low-value (usually 0 or black) portion of the mask will not change the pixels of the original image. In the image fused with the image mask, each pixel now contains information from two sources: the information of the original image and the information of the mask. The third pixel is used to refer to each pixel in the image after the mask information is fused. The third pixel is formed by merging the pixel of the original image (the first pixel) and the pixel of the mask (the second pixel).
[0195] In step 1062A, for a fourth pixel in the fused image label, a pixel value difference between the fourth pixel and the corresponding third pixel is determined.
[0196] In some embodiments, the fourth pixel refers to the pixel in the image label after the image mask is fused. The image label may be a classification map used to indicate the category or attribute of each pixel in the image. When the image mask is fused to the image label, the fourth pixel contains the original label information and the mask information. The third pixel refers to the pixel in the original image fused with the image mask. This step involves calculating the value difference between the fourth pixel (the pixel in the image label fused with the mask) and the corresponding third pixel (the pixel in the original image fused with the mask). This difference is usually determined by a simple subtraction operation, that is, difference = pixel value of the fourth pixel - pixel value of the third pixel.
[0197] In some embodiments, the fourth pixel refers to a pixel in a new image formed after the image mask is fused to the original image during image processing. This new image is usually obtained by superimposing or fusion of the original image and the mask image in some form. The value of the fourth pixel reflects the combination of the pixel value of the original image at that position and the mask value. In the context of image generation model training, the fourth pixel usually represents the predicted value of the model output, which is compared with the value of the corresponding position in the image label to evaluate the model performance. The pixel value difference refers to the difference between the pixel values of two pixels calculated in image processing. Specifically in the above description, the pixel value difference refers to the difference between the pixel value of the fourth pixel (the pixel value in the image label fused with the mask) and the pixel value of the corresponding third pixel (the pixel value in the original image fused with the mask). This difference is a quantitative indicator used to measure the deviation between the model prediction value and the actual target value. In the image generation model training process, the pixel value difference is an important component of the loss function, which is used to guide the update of the model parameters to reduce the gap between the model output and the target value. By minimizing these differences, the model can gradually improve the accuracy of its generated images.
[0198] In step 1063A, the determined pixel value differences are summed to obtain a first loss value, and based on the first loss value, the model parameters of the image generation model are updated.
[0199] In some embodiments, all calculated pixel value differences are summed. This sum represents a global error metric, namely the first loss value. The first loss value quantifies the overall difference between the image predicted by the model and the target image. The first loss value is usually the target to be minimized during the training process, because the smaller its value, the closer the image predicted by the model is to the target image. After obtaining the first loss value, the embodiment of the present application uses this value to update the parameters of the image generation model. This process involves using an optimization algorithm, such as gradient descent, to adjust the weights and biases of the model. The gradient of the loss function with respect to the model parameters is calculated, and this gradient indicates the direction of change of the parameters and can reduce the loss value. According to the gradient and the learning rate (a value that controls the update amplitude of the parameters), the model parameters will be updated. The purpose of the parameter update is to reduce the loss value so that the model's prediction is closer to the target value. Repeatedly, the model parameters will be updated according to the current loss value in each iteration. As the training proceeds, the model will gradually improve and the loss value will usually gradually decrease.
[0200] As an example, the expression of the first loss value may be:
[0201]
[0202] Among them, C i used to indicate the square value of the difference between the pixel values, It is used to indicate the pixel value difference, N is used to indicate the number of pixel value differences, mask is used to indicate the mask value, y i Used to indicate the pixel value of the third pixel point, Used to indicate the pixel value of the fourth pixel.
[0203] In this way, the image mask is fused to the original image, and the pixel after the fusion mask is defined as the third pixel. For the training and optimization of the image generation model, by directly superimposing the mask on the image, the specific area that the model needs to pay attention to and optimize, that is, the target area, is highlighted. For the fourth pixel in the image label of the fusion mask, the pixel value difference between it and the corresponding third pixel is calculated, which can quantify the difference between the model output and the real target, and provide direct feedback information for model training. The first loss value obtained by summing these pixel value differences is a key indicator for measuring model performance, which reveals the overall error between the model-generated image and the target image. Updating the model parameters of the image generation model based on the first loss value helps the model to more accurately capture the characteristics of the target area, reduce errors, and improve the quality and realism of the generated image.
[0204] In some embodiments, see Fig. 9 , Fig. 9 This is a flow chart of the training method of the image generation model provided in the embodiment of the present application. Figure 6 , Figure 4 Step 106 shown may be performed by Fig. 9 Steps 1061B to 1063B are implemented as shown.
[0205] In step 1061B, based on the first text, the image generation model is called to generate an image corresponding to the first text.
[0206] In some embodiments, the image generation model, also known as the text-generated graph model, is a model based on artificial intelligence technology that can generate an image corresponding to an input text description. By learning the association between text and image, the natural language description is converted into visual content. The image generation model is a deep learning model that can generate realistic images from given inputs (such as text, noise vectors, or other forms of conditions). By learning a large amount of image and text data, the image generation model can understand the semantic information in the text description and map it into the image space to generate visual content that matches the description. Typical applications include artistic creation, image restoration, virtual scene generation, advertising design, etc.
[0207] In step 1062B, a second loss value is determined based on the image corresponding to the first text and the image label, and a third loss value is determined based on the image and the fused image label.
[0208] In some embodiments, the second loss value is a quantitative indicator used to measure the overall difference between the image generated by the model and the given image label or target image in the image generation or recognition task. Specifically in the described scenario, the second loss value is calculated based on the image corresponding to the first text (i.e., the image generated by the model based on the text description) and the original image label (the label without the fusion mask). It measures the consistency between the global content of the generated image and the image label. This loss value can help the model learn how to generate images that match the label content based on the text description. Calculating the second loss value usually involves comparing each pixel of the generated image with the label image, or calculating some distance metric between them (such as the mean square error MSE), and aggregating these differences into a single value to guide the model training process.
[0209] In some embodiments, the third loss value is an error metric calculated specifically for the image label fused with the image mask during the training of the image generation model. It is used to measure the performance of the image generated by the model within the area specified by the mask. In the described scenario, the third loss value is calculated based on the generated image and the image label fused with the image mask. This means that only pixels within the mask area are considered when calculating the loss. The third loss value focuses on the details and accuracy of the generated image in a specific area, ensuring that the model can generate corresponding image features based on the area highlighted by the mask. This loss value helps the model pay attention to the generation quality of specific areas during training, which is particularly important for tasks that require attention to specific parts of the image (for example, focusing on people or objects in the image). When calculating the third loss value, the pixel differences within the mask area are usually weighted or only the errors in these areas are calculated.
[0210] In some embodiments, the image label includes multiple fifth pixel points, and the image corresponding to the first text includes sixth pixel points corresponding one-to-one to the fifth pixel points. The above-mentioned determination of the second loss value based on the image corresponding to the first text and the image label can be achieved in the following manner: based on the content difference, each of the mask values in the image mask is adjusted to obtain the target mask value corresponding to each of the mask values; for each of the fifth pixel points, the pixel value difference between the fifth pixel point and the corresponding sixth pixel point is determined, and the pixel value difference is multiplied by the target mask value corresponding to the fifth pixel point to obtain the loss value of the fifth pixel point; the loss values of each of the fifth pixel points are summed to obtain the second loss value.
[0211] In some embodiments, the fifth pixel in the image label refers to each pixel in the label image, which represents the category or attribute of each pixel in the image. The sixth pixel in the image corresponding to the first text refers to the corresponding pixel in the image generated according to the text description. These two pixels are one-to-one corresponding, that is, each pixel in the label image has a corresponding pixel in the generated image. Before calculating the second loss value, the mask value in the image mask is first adjusted based on the content difference. The content difference may refer to the content difference between the original image and the generated image in a specific area, or the difference between the image label and the generated image. By adjusting the mask value, a target mask value can be obtained. The purpose of this step is to make the mask better reflect the difference between the generated image and the label image, so as to give a more accurate weight when calculating the loss value. For each fifth pixel, calculate the pixel value difference between it and the corresponding sixth pixel. This difference represents the degree of mismatch between the generated image and the label image at this pixel. Multiply this pixel value difference by the target mask value corresponding to the fifth pixel. This step is to strengthen or weaken the influence of a specific pixel on the total loss value by weighted difference. If the target mask value is high, the error of that pixel contributes more to the total loss value; if the target mask value is low, the contribution is smaller. Sum the loss values of all fifth pixels, and the sum is the second loss value. This loss value is a measure of the overall difference between the generated image and the labeled image. It takes into account the errors of all pixels, and the error of each pixel is weighted according to its target mask value.
[0212] As an example, the expression of the second loss value may be:
[0213]
[0214] Wherein, L2 is used to indicate the second loss value, is used to indicate the pixel value difference between the fifth pixel and the corresponding sixth pixel, N is used to indicate the number of pixel value differences, (1-mask) is used to indicate the target mask value, yi is used to indicate the pixel value of the fifth pixel, Used to indicate the pixel value of the sixth pixel.
[0215] In this way, by adjusting the image mask and calculating the second loss value based on content differences, and by personalizing the mask value, the model can pay more attention to key areas, thereby strengthening the importance of these areas when generating images. The loss value calculation of each fifth pixel matches its position and role in the image, ensuring the pertinence and effectiveness of model training. By multiplying the pixel value difference between the fifth pixel and the corresponding sixth pixel by the target mask value, the model can more accurately quantify and optimize the error in a specific area, thereby improving the consistency between the generated image and the labeled image. The second loss value obtained by summing the loss values of all fifth pixels provides a comprehensive and detailed error metric for the model, allowing the model to more efficiently learn the mapping relationship from text to image during training, thereby generating more accurate and realistic images.
[0216] In some embodiments, the above-mentioned adjusting each of the mask values in the image mask based on the content difference to obtain the target mask values corresponding to each of the mask values can be achieved as follows: performing target segmentation on the image label based on the content difference to obtain a target segmentation result of the image label, the target segmentation result including a first image area in the image label corresponding to the content difference, and a second image area in the image label that is different from the first image area; determining the target mask value corresponding to the pixel point in the first image area in the image mask as the mask value corresponding to the pixel point in the second image area; determining the target mask value corresponding to the pixel point in the second image area in the image mask as the mask value corresponding to the pixel point in the first image area.
[0217] In some embodiments, the content difference refers to the difference between the original image and the image label or the generated image and the image label. This difference can be calculated in a variety of ways, for example, by comparing pixel values, feature vectors, or other image analysis techniques to quantify. Based on the content difference, the image label is segmented. This step involves segmenting the image label into two different regions: a first image region and a second image region. The first image region generally refers to the portion corresponding to the content difference, that is, the regions in the image label that show large differences, which may be the parts that the model needs to pay special attention to and optimize. The second image region refers to the portion of the image label that is different from the first image region, that is, the regions with small content differences or that do not need special attention. In the process of adjusting the mask value, the target mask value corresponding to the pixel point in the first image region will be set to the mask value corresponding to the pixel point in the second image region. This means that the mask value that may have a higher mask value in the first image region (for example, the mask value is close to 255, indicating an area that needs to be paid special attention to) will now be set to the mask value of the second image region. Conversely, the target mask value corresponding to the pixel point in the second image region will be set to the mask value corresponding to the pixel point in the first image region. This means that areas that may have had lower mask values in the second image (e.g., mask values close to 0, indicating areas of non-attention) will now be set to the mask values of the first image area. By adjusting the mask values in this way, the model is guided to focus on areas that are important in content differences, while reducing attention to unimportant areas. Such adjustments help the model learn more effectively how to generate image content that matches the key areas in the image labels, thereby improving the accuracy and quality of the generated images.
[0218] As an example, assume that the embodiment of the present application has an image label, which is an image containing two characters, each with a different background. The goal of the embodiment of the present application is to generate an image that matches this image label, but the model encounters difficulties in generating characters, resulting in a large difference between the generated characters and the characters in the label, while the background is relatively matched. The image label is segmented, and the embodiment of the present application obtains two areas: the first image area: the character area, with a large content difference. The second image area: the background area, with a small content difference. In the original image mask, the embodiment of the present application may mark the character area as a high mask value (for example, a mask value of 255) and the background area as a low mask value (for example, a mask value of 0). Now, the embodiment of the present application adjusts the mask value according to the content difference: the mask value (high value, such as 255) of the character area (first image area) in the original mask is adjusted to the mask value (low value, such as 0) of the background area (second image area). This means that the model no longer needs to pay special attention to the character area. The mask value (low value, such as 0) of the background area (second image area) in the original mask is adjusted to the mask value (high value, such as 255) of the character area (first image area). This means that the model needs to pay more attention to the background area. Through this adjustment, the embodiment of the present application actually changes the focus of the model. Originally, the model needed to focus on the character area (because the content difference there is large), but now the model is guided to pay attention to the background area, which helps the model to invest more learning resources in background generation, thereby improving the generation quality of the character area.
[0219] In this way, by adjusting the mask values in the image mask based on content differences and performing target segmentation on the image labels, the model can identify key areas with large content differences and relatively stable background areas, and then optimize the mask values according to the characteristics of these areas. This adjustment strategy enables the model to focus more on areas that need to be optimized during training, thereby improving the pertinence and efficiency of training. By exchanging the mask values of the first image area and the second image area, the model is guided to pay attention to the background areas that may have been ignored, while reducing the focus on the relatively matched character areas. This strategy not only helps to improve the model's performance in background generation, but also optimizes the model's ability to understand and reproduce the overall image structure.
[0220] In some embodiments, the above-mentioned determination of the third loss value based on the image and the fused image label can be achieved in the following manner: fusing the image mask to the image to obtain a fused image, and taking the pixel point in the fused image as the seventh pixel point; for the eighth pixel point in the fused image label, determining the pixel value difference between the eighth pixel point and the corresponding seventh pixel point; summing the pixel point difference values corresponding to each of the eighth pixels to obtain the third loss value.
[0221] In some embodiments, the seventh pixel refers to a pixel in the image after the image mask is fused to the original image. This means that for each pixel in the original image, the embodiment of the present application adjusts its value according to the information of the image mask. If the value of the mask on the pixel is 1 (or a positive value), the value of the pixel will be adjusted according to the influence of the mask; if the mask value is 0, the value of the pixel remains unchanged. The eighth pixel refers to a pixel in the image label fused with the image mask. This image label is a target image, which contains the value that the image generated by the model in the embodiment of the present application should have. The description proposes a method for calculating a third loss value, which is used to measure the generation effect of the image generation model in a specific mask area. The following is a detailed explanation of this process: the seventh pixel, the seventh pixel refers to a pixel in the image after the image mask is fused to the original image. This means that for each pixel in the original image, the embodiment of the present application adjusts its value according to the information of the image mask. If the value of the mask on the pixel is 1 (or a positive value), the value of the pixel will be adjusted according to the influence of the mask; if the mask value is 0, the value of the pixel remains unchanged. Eighth pixel: The eighth pixel refers to the pixel in the image label that is fused with the image mask. This image label is a target image, which contains the value that the image that the embodiment of the present application hopes the model should generate should have. In order to determine the third loss value, the pixel value difference between the eighth pixel and the corresponding seventh pixel is first calculated. This difference represents the degree of mismatch between the image generated by the model and the target image in the area fused with the mask. Next, the embodiment of the present application sums up the pixel value differences corresponding to all these eighth pixels. This sum is the third loss value, which is a scalar value used to measure the overall error of the model in generating the image in the entire mask area.
[0222] As an example, the expression of the third loss value may be:
[0223]
[0224] Among them, L 3 Used to indicate the third loss value, used to indicate the square value of the difference between the pixel values, is used to indicate the pixel value difference between the eighth pixel point and the corresponding seventh pixel point, N is used to indicate the number of pixel value differences, mask is used to indicate the mask value, yi is used to indicate the pixel value of the seventh pixel point, Used to indicate the pixel value of the eighth pixel.
[0225] This allows precise control of specific areas during model training. By applying masks to images, the model can focus on those parts that need to be optimized, thereby achieving more detailed and accurate image generation in key areas. Secondly, the pixel value difference between the seventh pixel and the eighth pixel is calculated, and the third loss value is determined by summing these differences, which effectively quantifies the generation error of the model in the masked area and provides a clear optimization target for the model. This targeted error metric helps the model better understand the local features and details in the training data, thereby improving the quality and realism of the generated images.
[0226] In step 1063B, the second loss value and the third loss value are weightedly summed to obtain a fourth loss value, and based on the fourth loss value, the model parameters of the image generation model are updated.
[0227] As an example, the expression of the fourth loss value may be:
[0228] L 4 =α 1 L 2 +α 2 L 3 (5)
[0229] Among them, L 4 Used to indicate the fourth loss value, α 1 and α 2 The weight used to indicate the second loss value and the third loss value, L 2 Used to indicate the second loss value, L 3 Used to indicate the third loss value.
[0230] In some embodiments, the end time of training the image generation model can be determined based on the following factors: 1) Convergence of loss function: Image generation models usually use loss function to measure the difference between the generated image and the target image. When the value of the loss function no longer decreases significantly or remains stable in multiple training cycles, the model can be considered to have converged and the training can be terminated. 2) Validation set performance: In addition to the training loss, the performance of the model on the validation set is also an important indicator. If the performance of the model on the validation set meets expectations or no longer improves significantly, then the training can be considered to end. Overfitting detection: If the performance of the model on the training set is getting better and better, but the performance on the validation set begins to decline, this may mean that the model is overfitting. In this case, training should be stopped to prevent the model from over-learning the training data and failing to generalize to new data. 3) Resource constraints: Training deep learning models usually requires a lot of computing resources and time. If training cannot continue due to resource constraints, then the training should also be considered to end. 4) Manual setting: The end time of training can also be set manually by the researcher. For example, a fixed number of training epochs can be set, or training can be stopped after a certain training time is reached. 5) Early Stopping: Used to prevent overfitting. When the performance of the model on the validation set no longer improves, the training will stop early. The training of the image generation model should end when the performance of the model reaches the expected level, or when continued training no longer brings significant performance improvement. At the same time, the limitations of computing resources and time can also be considered, as well as the risk of preventing model overfitting.
[0231] In this way, the model not only focuses on the overall visual quality when generating images, but also focuses on the accuracy of specific areas or details, which is crucial for achieving more refined image control. The fourth loss value obtained by weighted summation provides a comprehensive performance indicator that can more comprehensively reflect the performance of the model at the global and local levels, thereby guiding the update of model parameters. The model can better understand the relationship between the text description and the image content, and generate images that are more consistent with the text description and rich in details. By updating the model parameters based on the fourth loss value, the generation quality of the model can be effectively improved, the error can be reduced, and the model can be more robust and reliable in image generation tasks, thereby broadening its scope of application.
[0232] In this way, based on the text sample, the image generation model is called to generate the image corresponding to the text sample, the image label of the text sample is converted to text, the first text corresponding to the image label is obtained, the content difference between the first text and the text sample is obtained, and based on the content difference, the image mask of the image label is generated, and the image mask is fused to the image label to obtain the fused image label; and the image generation model is trained based on the image and the fused image label. In this way, by performing text conversion on the image label of the text sample, the first text corresponding to the image label is obtained, and the content difference between the first text and the text sample is obtained. Since there is a content difference between the first text and the text sample, it means that there is an image area corresponding to the content difference in the image label of the text sample, which makes it difficult for the image label to adapt to the text sample. It means that directly training the image generation model based on the image label will result in poor image generation quality of the trained image generation model. By generating an image mask of the image label based on the content difference, the image mask can accurately reflect the content difference between the first text and the text sample, and the image mask is fused into the image label to obtain a fused image label, so that the image area corresponding to the content difference in the image label is covered by the content difference reflected by the image mask, so that the obtained fused image label can cover the image area corresponding to the content difference in the image label. By training the image generation model based on the image and the fused image label, the training of the image generation model can effectively avoid the influence of the image area corresponding to the content difference in the image label, thereby effectively improving the image generation quality of the image generation model.
[0233] In some embodiments, the image generation method provided in the embodiments of the present application can be implemented by a server or a terminal alone, or by a server and a terminal in collaboration. The following will be described using the server alone as an example.
[0234] The image generation method provided in the embodiment of the present application can be implemented in the following manner: obtaining a third text, and calling an image generation model to generate an image corresponding to the third text.
[0235] In some embodiments, the above-mentioned image generation model is trained by the training device of the above-mentioned image generation model. Specifically, the image generation model is trained based on sample images and fused image labels, the sample images are generated based on text samples, the image mask is generated based on the difference between the text sample and a fourth text, and the fourth text is obtained by performing text conversion on the image labels.
[0236] In some embodiments, a large number of text samples and their corresponding image samples are required. These text samples describe the content of the image, while the image samples are the visual representation of these descriptions. In order to enhance the model's understanding of the correspondence between text and image, image masks are generated. These masks are generated based on the content difference between the text sample and the first text. The first text is obtained by text conversion of the image label of the text sample, which may include simplification, summary or other forms of rewriting of the original text. The generated image mask is fused with the image label to form a new input, which contains the semantic association information between the text and the image. The image generation model is trained using the fused image mask and image label, as well as the corresponding image samples. The goal of the training is to enable the model to learn how to generate corresponding images based on the text description. After training, when given a new text input, the image generation model can use the learned knowledge to generate an image that matches the text description.
[0237] As an example, suppose there is an application called AI Painter that allows users to generate corresponding art images by inputting text descriptions. The user enters a text description in the application, such as a landscape painting depicting a seaside at sunset, with several seagulls flying in the sky and a couple walking on the beach. After the application receives the user's text input, it passes it to a pre-trained image generation model. The image generation model processes the input text description, understands the semantics and visual elements in it, and then generates an image that matches the text description. In this example, the image generation model may generate an image of a sunset and seaside scenery, with elements such as seagulls, beaches, and couples in the image. The generated image is displayed on the application's interface, and users can enjoy the artwork generated based on their text description. Users can also adjust and edit the generated image to achieve their ideal effect.
[0238] In this way, by acquiring text and calling the image generation model to generate the corresponding image, a rapid conversion from text description to visual content can be achieved. By learning a large number of text samples and corresponding images, the image generation model can understand the semantic association between text and image and generate high-quality images that meet user expectations. The training method of integrating image labels can improve the image generation model's understanding of text content and the accuracy of image generation, making the generated image closer to the text description.
[0239] The following is an explanation of an exemplary application of the embodiments of the present application in an actual application scenario of an image generation model.
[0240] Image generation model, also known as text-to-image generation model, is an artificial intelligence technology that can generate corresponding images based on text descriptions. Image generation models can be used to automatically generate illustrations for novels, comics, games, etc., to provide inspiration for creators, and can even be used to generate personalized works of art. Designers can use image generation models to quickly generate creative design sketches for advertising, brochures, web design, etc., to improve design efficiency. In the field of education, image generation models can automatically generate illustrations for teaching materials, or generate images that match text content in presentations. Users can use image generation models on social media to generate personalized emoticons, cover images, etc. to increase interactivity and fun.
[0241] The image generation model based on the diffusion model (such as stable diffusion) has developed rapidly in recent years. However, due to the inaccurate text used to train the generated model (description errors or omissions), the generated image is prone to errors and cannot be aligned with the description text. Considering that when the image and text are inconsistent, there will be information difference between the image and the text, which can guide the model to erase the wrong information during learning. In order to effectively train the generative model when the text-generated image is not aligned with the text, this scheme introduces cycle consistency to optimize the supervised information in the generative model training process, generates auxiliary training data at the same time, and calculates the loss for the image part where the image and text are not aligned, so as to jointly train the model. The specific process is: in the conventional text-generated image training process, a cycle process is added to verify the consistency of the text information of the generated image with the input text, and a consistency error is generated. With the help of the image-text connection relationship of the image cross attention, the corresponding image inconsistent area is generated from the consistency information difference, and the image supervision mask is obtained, so as to optimize the generation loss, thereby achieving the erasure of wrong information during training. In this solution, the image-text inconsistent region mask can generate regional joint training data, so that the consistent region and the inconsistent region can be trained separately, thereby improving the utilization rate of training data.
[0242] The embodiments of the present application construct a cycle consistency verification process for generating images; obtain mismatching points on the image through text homomodal consistency and image-text cross-modal response under image cross-attention, and expand from points to regions through effective methods; use mismatching region information as auxiliary supervision information and generate new training samples to jointly train the model.
[0243] The embodiment of the present application optimizes the diffusion Wensheng graph generation model, which can be applied to products that use generation models, such as content creation, using the Wensheng graph large model to quickly generate images that match the text description, thereby improving the creative efficiency of artists and creators. Design and advertising marketing can use Wensheng graph technology to generate exquisite design drafts, and designers can further refine them on this basis, thereby providing support for product design and advertising marketing.
[0244] In some embodiments, see Fig.10 , Fig.10 This is a schematic diagram of the principle of the training method of the image generation model provided in the embodiment of the present application. Figure 1 The traditional learning process based on the diffusion model is to encode the input image into the latent space through the encoder, add noise, and generate the latent space features at time T through the diffusion process. Then, T steps of U-Net denoising are performed to generate noise predictions, and the loss between the predicted noise and the previously added noise is calculated to train the U-Net. Considering that the text description may be wrong and thus mismatched with the image, and the learning of wrong information will make the generation model's generation ability worse, it is necessary to focus on erasing the wrong information-inconsistency errors. This scheme introduces cycle consistency judgment to obtain mismatched error information, thereby generating more accurate supervision information, so that the generation model results are more consistent with the text description. The cycle consistency process of this scheme is as follows Fig.10 As shown, after the text-to-image direction is generated, the image-to-text reverse process is added, and the consistency errors of the previous and next texts are calculated, and the image mask is generated through the text consistency errors, and the mask is used to control the loss of the text-to-image generation result. The consistency error generation mask is to generate the image-text correspondence information with the help of the generative model attention back propagation, so that the errors in the text mined by the cycle consistency can be mapped to the image, thereby optimizing the image generation loss calculation process and finally obtaining a more accurate loss. In the cycle consistency process, the generated image is used to generate new text through the description model, and then the subsequent operations are continued: the new text and the supervised text generate consistency errors, and at the same time, an information difference mask is generated. The information difference mask removes the text mismatching parts when calculating the new image generation loss.
[0245] In some embodiments, see Fig.10 The training samples in the embodiment of the present application may be image-text pairs (i.e. Fig.10 The text samples and image labels shown in the figure contain images and corresponding text descriptions. Each sample is recorded as: (image, description text). Some images in the image-text pair samples do not match, such as a picture of a cat running and a text description of a dog running.
[0246] In some embodiments, see Fig.10 , Fig.10 The image generation model shown can be based on the stable diffusion series model (such as sd1.4, sd1.5, sdxl, etc.). Fig.10 The text generation model shown can be a text generation model for the BLIP series of images (BLIP, BLI P2 model). Since the task of this application is to improve the effect of the image generation model, Fig.10 Only the image generation model needs to update parameters, describing the model (i.e. Fig.10 The text generation model shown) does not need to be updated.
[0247] In some embodiments, see Fig.10 For the input text (i.e. the text sample described above), the image generation model is first used to generate an image, and then the original image (i.e. the image label described above) is generated by the text generation model. Fig.10 The text description is consistent with the input text, and the text inconsistency information is output to generate an image mask. A mask is generated for the parts of the supervised image that are inconsistent with the text. Finally, the generation loss is calculated. During the calculation, the parts covered by the mask are not subject to loss calculation.
[0248] In some embodiments, see Fig.10 , the input of the consistency calculation is the original supervised text 1 (that is, Fig.10 The text sample shown in the figure) (a cat in the grass) and the text generated based on the supervised image 2 (i.e. Fig.10The text shown) (a dog in the grass), the calculation process is: A. Initialize the optimal matching score s = 0, the optimal matching length (0, 0) indicates the matching length of the text sample and the text, and the optimal matching starting point (0, 0) indicates the match from the 0th position of the text sample to the 0th position of the text. B. Starting from the first word position, the text is matched with the text sample one by one to obtain the matching score. That is, the matching degree of the first word, the first and second words, the 123rd words... of the text with the text sample is calculated respectively to obtain the sequence matching score of the text. C. For the text sequence matching score, the position with the lowest matching score is selected and considered to be the wrong word position. D. The words at the wrong word position of the text sample and the text are constructed as an incorrect word pair (cat, dog), indicating that the supervised text contains cat, but the supervised image contains dog. Matching calculation of two texts: Input the two texts into the clip text representation model respectively, extract the text representation, and calculate the cosine similarity of the two text representations, and then divide it by the word length of the shortest word to obtain the cosine similarity of the average word, and the similarity ranges from 0 to 1. For example, for adog and a cat in grass, first extract the clip text representation, then calculate the cosine similarity of the two representations, and then divide the similarity by 2. For example, the text samples in the example, 2, get the similarities: 0.4, 0.1, 0.23, 0.3.
[0249] In some embodiments, for the generation of an image mask, the image position corresponding to each word of the input text is obtained according to the input text, the above-mentioned incorrect word pairs, and the image according to the reverse process of attention calculation. Obtain the response point corresponding to each word: calculate the image response points corresponding to all words - that is, according to the reverse process of attention, obtain the corresponding response points for each word in the text from left to right (each word has a maximum response point (the maximum response point in the figure, that is, the feature distance between the feature of the word and the feature of each response point, and the point with the largest feature distance is the maximum response point of the word)), and the maximum response point extends outward to a certain range and corresponds to the word, and the extension range is determined by the following response area. Obtain the response area corresponding to each word: use SAM to perform image segmentation, find the corresponding segmentation area on the image segmentation for the above-mentioned response points one by one, and obtain the response area corresponding to each word. Find the image response area corresponding to the incorrect word and make a mask: the pixels in this area on the mask are 0, and the other positions are 1. The embodiment of the present application can directly generate a regional style threshold based on global statistical information to obtain the target area. First, connect each response point (the edge of the image is also considered to be a response point, but the response value is very small) to obtain the size of the response value between the lines connecting the two response points, and take the minimum response value on the line as the boundary. When determining which word area a certain point m (coordinates x, y) in the image belongs to: find the nearest response point abcd to the left, right, top, and bottom of m respectively (if there is no response point in a certain direction, the response point is set to the edge of the image), and determine which response point the point belongs to based on the boundary position on the line connecting the two response points: (1) When the response point is on the side close to b of the ab boundary, there is no response point in the c and d directions - that is, there is no response point in the area above and below the point (x, y). When calculating, take the upper and lower parts of the area instead of the upper and lower parts of the point, that is, there is no response point above and below the line connecting the two points (xl, y)) - then the point Belongs to the area of the word corresponding to the b response point; (2) When there is only a response point and no response points in other directions, it belongs to area a; (3) When there are abc response points but no d direction response point, first determine the distance relationship between point m and the response point (assuming the closest is a, then b, and then c), and use the same method to first determine whether it belongs to a or b (assuming it belongs to a), and then determine whether it belongs to the first result or c (a or c); (4) When there are 4 direction response points, still follow the above method in order of distance, first determine which of the 2 response points with the closest distance it belongs to, and then determine which one it belongs to with the next response point and the last response point respectively.
[0250] In some embodiments, see Fig.11 , Fig.11 This is a schematic diagram of the principle of the training method of the image generation model provided in the embodiment of the present application. Figure 2, start with a text sample, such as a cat in the grass. This text sample describes the content of the image to be generated. Based on the text sample, call the image generation model to generate an image corresponding to the text description. The image generation model may be based on technologies such as generative adversarial networks (GAN) or variational autoencoders (VAE). The generated image will have a corresponding image label, such as a dog in the grass. This label may be different from the original text sample, reflecting the image content generated by the model. Perform text conversion on the image label to generate text describing the image label. For example, convert the image label a dog in the grass to the text a dog in the grass. Get the difference between the generated text and the original text sample. For example, the original text is a cat, and the generated text is a dog, and the difference is the difference between a cat and a dog. Generate an image mask based on the text difference. The image mask is used to identify the parts of the image label that are different from the original text sample. For example, the mask may identify the dog part of the image, while ignoring the same parts such as the grass. The generated image mask is fused into the image label to form an image label with a mask. This helps the model understand more clearly which parts need to be adjusted. Based on the masked image labels and the generated images, the parameters of the image generation model are updated. In this way, the image generation model can learn how to better generate images based on the text descriptions and reduce the difference between the generated images and the text descriptions.
[0251] In some embodiments, see Fig.11 , calling the image generation model, based on the text sample ( Fig.11 A cat in the grass is shown), generate an image; call the text generation model, based on the image label ( Fig.11 A dog is shown in the grass), generating text ( Fig.11 A dog in the grass is shown), determining the content difference between the text and the text sample ( Fig.11 differences shown - cat vs dog); determining differences in image labels from differences in content ( Fig.11 dog shown); generating an image mask based on the portion outside the difference in the image label and the portion corresponding to the difference; determining a loss based on the image, the image label and the image mask.
[0252] In some embodiments, for the process of determining the response point by backpropagation of attention, the high-dimensional tensor (cross-attention maps) of cross attention can be used to establish a good connection between the image and word features. This solution associates the generated sentences with the generated images. The specific process is: forward calculation of the text, and saving the features of each layer of cross attention (note that the cross attention feature in SD1.5 is the feature output after the text is injected into the cross attention. After the text is injected, a larger response value will be obtained in the place associated with the text in the image-it is precisely because of this that a picture-text connection can be established based on cross attention). For each text feature, calculate the mean feature of all layers of cross attention. For each text, draw a picture: draw the mean feature (as shown below Fig.12 As shown, Fig.12 Schematic diagram of the effect of the response point provided in the embodiment of the present application, with more obvious response in key actions and rankings).
[0253] In some embodiments, see Fig.12 , the response point of word 1 in the image is 11, the response point of word 2 in the image is 21, the response point of word 3 in the image is 31, the response point of word 4 in the image is 41, the response point of word 5 in the image is 51, the response point of word 6 in the image is 61, the response point of word 7 in the image is 71, and the response point of word 8 in the image is 81.
[0254] In some embodiments, see Fig.13 , Fig.13 is a schematic diagram of the structure of the image generation model provided in the embodiment of the present application, based on Fig.13 The principle of the image generation model shown is as follows: after adding noise to the original image encoding (embedding, in the application, VAE encoding is used to map to the latent feature space), the latent space representation at time T is obtained through the diffusion process, and the features of the target image (i.e., the original image features without noise) are predicted through T denoising U-Net (denoising process) operations - the features are subtracted from the original U-Net (the image generation model described above) input to obtain the noise prediction, and the features are decoded by VAE to obtain the target image. For text, the text embedding is obtained through the CLIP text branch and then controlled by the QKV of U-Net. Diffusion sampling is used to map the features of the noisy image VAE encoding to the latent space representation at time T. The subsequent denoising process of the image learns to produce a fitting of the noise representation so that the original image minus the noise representation obtains the real image representation, and the real image is obtained through the decoder D.
[0255] In some embodiments, see Fig.13 , Fig.13 The text generation model shown consists of multiple stacked residual blocks ( Fig.13 The residual blocks 31 and residual blocks 32 shown in FIG. 3 and the spatial transformation layer ( Fig.13 The spatial transformation layer 41 and the spatial transformation layer 42 shown in the figure) contain two spatial transformation layers in the image generation model. Each spatial transformation layer is a QKV process. In the first QKV matrix process, KV is the same as the input Q (Q is the output of the previous network structure), and in the second QKV matrix process, KV is a constraint used to control the generated text features.
[0256] In some embodiments, see Fig.13 , there are multiple residual blocks (such as residual block 31 and residual block 32) in the image generation model. The residual block is the core component of the deep residual network (ResNet). It solves the gradient vanishing problem in the deep network through skip connections, so that the network can be trained deeper. The spatial transformation layers (such as spatial transformation layer 41 and spatial transformation layer 42) in the image generation model may be used to perform spatial transformations on the input data, such as rotation, scaling, or translation. These layers can help the model better process the spatial information in the image data. The downsampling layer is used to reduce the spatial dimension of the data while increasing the depth of the features. This helps the model extract abstract features at a higher level. The Q matrix, K matrix, and V matrix are mentioned in the image generation model, which are key components in the attention mechanism. The attention mechanism weights the input data by calculating the relationship between the query (Q), key (K), and value (V) to capture important information in the data. Skip connections are used to pass low-level features directly to high-level layers, which helps information flow and gradient propagation. The fully connected layer is used to map features to the final output space. Time step embedding may be used to process time series data or dynamic systems. Hidden variables are used to represent the latent features of the data, while contextual embedding is used to capture the contextual information of the input data. Normalization and scaling operations are used to stabilize the training process and prevent gradients from exploding or disappearing.
[0257] In some embodiments, see Fig.13 , Fig.13 The residual block shown includes a fully connected layer and a convolutional layer. The convolutional layer is called to convolve the latent variable to obtain a first convolution result. The fully connected layer is called to process the time step embedding to obtain a first fully connected result. The first convolution result and the first fully connected result are added to obtain a first sum result. The latent variable jump connection and the first sum result are added to obtain a second sum result. The convolutional layer is called to convolve the second sum result to obtain an output.
[0258] In some embodiments, the spatial transformation layer includes a convolution layer, dense mapping, dot multiplication, scaling, and normalization. The convolution layer is called to perform potential projection on the latent variable to obtain a projection result; the projection result is densely mapped to obtain a first dense mapping result; different dense mappings are performed on the context mapping to obtain a K matrix and a V matrix; the K matrix and the Q matrix are dot multiplied to obtain a dot product result; the dot product result is scaled to obtain a scaled result; the scaled result is normalized to obtain a normalized result; the normalized result and the V matrix are constructed as an attention weight matrix, and the attention weight matrix is dot multiplied to obtain a dot product result; the convolution layer is called to convolve the dot product result to obtain an output.
[0259] In some embodiments, see Fig.10 , the conventional training set collected by the business may contain errors, so the cyclic correction of this method is performed. Fig.10The process organizes the input and output of the model. After obtaining the image mask generated by the cycle consistency, when calculating the generated MS E loss, the mask loss is calculated for the image, that is, the loss is not calculated for the parts of the mask that are 0. The model is trained with this loss. The process is as follows: a total of M rounds (such as 10) of iterations are performed on the full amount of training data (that is, N pairs of pictures and texts). In each round of iteration: each bs sample of the full amount of image data is a batch, and the model is updated once for each batch of data. When all batches (N / bs) have been trained once, that is, all samples have been trained once in the model, it is called a round of iteration. The training uses image-text pair samples. For a certain image-text pair sample, the original image is input and initial noise is added. The text is used to generate constraints. The model generates the predicted image latent space feature Z0, and the decoder restores it to an image. The fine-tuning loss of the generated model is calculated with the supervised image. During training: The first batch of the first round of pre-training parameters are initialized: the parameters of the open source trained model (stable-diffusion v1-5) are used for VAE, text_encoder, and U-Net. In this training, only the U-Net parameters need to be updated, and the others are not updated. The learning rate is initialized with 0.0004, and the learning rate is changed to 0.1 times the original after every 5 rounds of learning, and a total of 10 rounds of training are conducted. Extract bs image-text pair samples and input them into the model: For each image, the encoder generates a latent space representation Ei, randomly generates bs seeds, and generates corresponding bs noise maps (with the same dimension as Z0). This map is superimposed with the representation Ei as Z0, and then the diffusion process generates ZT (as the original input for subsequent U-Net denoising). The text information is clipped to obtain text representation, which is input into the generation model (the text representation is used as K and V input information), and T denoising U-Net forward calculations are performed on ZT under the KV constraint. After the first forward calculation, ZT-1 is obtained, and finally after T times, the U-Net output is predicted Z0, and the image is decoded to obtain the predicted image. Calculate batch loss: Calculate the generation loss (this step is the MSE loss of the predicted image and the supervision image under the mask), and count the total loss of the batch samples. Using the SGD random gradient descent method, the loss is reversed back to the model to obtain the gradient of the model parameters (U-Net) and update the parameters. Complete all N / bs batch training and end an iteration.
[0260] In some embodiments, see Fig.10 , calling the image generation model, based on the text sample ( Fig.10 A cat in the grass is shown), generate an image; call the text generation model, based on the image label ( Fig.10 A dog is shown in the grass), generating text ( Fig.10 A dog in the grass is shown), determining the content difference between the text and the text sample ( Fig.10 difference shown - cat vs dog); based on the difference, generate an image mask; determine a loss based on the image, image label, and image mask.
[0261] In some embodiments, see Fig.14 , Fig.14 This is a schematic diagram of the principle of the training method of the image generation model provided in this application Figure 3 In the above process, a new text description of the supervised image (such as a dog in the grass) can be obtained. In order to effectively use all data, data merging training is performed after the above cyclic correction training to avoid the problem that some masked concepts are not learned. The overall training process is similar to the above process: when training a batch of samples, text 11 and text 12 are respectively input into the generative model to be trained to generate images 21 and 22. Loss 1 and loss 2 are calculated based on the masks saved in the above process, where loss 1 uses the original mask (the mask value of the dog part is 0 at this time), and loss 2 uses 1-mask (the dog part is not 0 at this time). The two are weighted and merged to obtain the total loss, and the total loss is back-propagated and the generative model is updated.
[0262] In some embodiments, the above loss 2 may not be performed. For example, if the amount of training data is sufficient, and the number of samples that the model can learn will not decrease after masking, then the previous round of training on the masked part of the data will not cause the model to learn less concepts of the masked part (for example, the dog in a certain sample picture is masked, but there are still many dog photos to be trained in the training data. Even if a dog picture is masked, it does not cause the model to learn too much and miss the concept of the dog).
[0263] In some embodiments, to improve training efficiency, after completing the last round of cyclic correction training, the m ask of each sample in this round of training is saved, and then data merging training is performed (the saved mask is required). Another training method is to alternate the two, that is, cyclic correction and data merging training are performed on each batch of data in turn: first perform cyclic correction training on a batch of data, and then perform data merging training based on the cyclic correction output mask; then train the cyclic correction and data merging training process for the next batch of data.
[0264] In some embodiments, the MSE loss is used to calculate the mean square error of the output predicted image and the image of the image-text pair. The following y is the pixel value of each point in the image of the image-text pair. The p indicates the predicted pixel:
[0265]
[0266] In some embodiments, during the cyclic correction, a mask is used to remove the mismatched parts of the supervision image to avoid providing erroneous supervision information to the model. At this time, the MSE is calculated as follows, where the mask value for a specific part is 0 and the others are 1:
[0267]
[0268] When merging training, 1-mask training is used, where the loss is:
[0269]
[0270] When merging data for training, the total loss (that is, the fourth loss value described above) loss = a*MSE1+(1-a)*MSE2, where a is an adjustable weight.
[0271] The embodiment of the present application verifies the accuracy of the training data from the perspective of cycle consistency, and generates data inaccuracy information (image mask) on the image, and uses this inaccurate information to improve the accuracy of loss calculation during model training. The embodiment of the present application can be applied to generate image training (such as the training process mentioned above in this article), and can also be used to generate image data cleaning, which can find abnormal data. When inconsistencies are found in certain data, they are discarded, otherwise they are retained; they can also be used for other downstream applications, such as supplementing model training data based on inconsistent information. As in the above-mentioned merged training, inconsistent areas and consistent areas are trained separately, which improves the utilization rate of inconsistent data.
[0272] In some embodiments, a cycle consistency verification process for generating images is constructed; mismatching points on the image are obtained through text homomodal consistency and image-text cross-modal response under image cross-attention, and expanded from points to regions through effective methods; mismatching region information is used as auxiliary supervision information, and new training samples are generated to jointly train the model.
[0273] It is understandable that in the embodiments of the present application, related data such as text samples are involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0274] The following continues to describe the exemplary structure of the image generation model training device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2As shown, the software modules in the training device 455 of the image generation model stored in the memory 450 may include: a generation module 4551, which is used to call the image generation model to generate an image corresponding to the text sample based on the text sample; a conversion module 4552, which is used to perform text conversion on the image label of the text sample to obtain a first text corresponding to the image label, and the first text is used to describe the content in the image label; a mask module 4553, which is used to obtain the content difference between the first text and the text sample, and generate an image mask of the image label based on the content difference; a training module 4554, which is used to fuse the image mask to the image label to obtain a fused image label; and train the image generation model based on the image and the fused image label.
[0275] In some embodiments, the first text and the text sample both include multiple words, and the above-mentioned mask module 4553 is also used to respectively determine the similarity between each word in the first text and the text sample; determine the word whose similarity with the text sample is less than the similarity threshold as the first word; select the second word corresponding to the first word from the text sample, and use the first word and the second word as the content difference between the first text and the text sample.
[0276] In some embodiments, the mask module 4553 is also used to perform the following processing for each word in the first text: when the word is the first word in the first text, the word and the text sample are encoded respectively to obtain a first vector and a second vector, and the similarity between the first vector and the second vector is determined as the similarity between the word and the text sample; when the word is not the first word in the first text, a second text is constructed based on the word and the text before the word in the first text, and the similarity between the second text and the text sample is determined as the similarity between the word and the text sample.
[0277] In some embodiments, the above-mentioned mask module 4553 is also used to encode the first text to obtain a third vector, and encode the text sample to obtain a fourth vector; determine the distance between the third vector and the fourth vector, and when the value of the distance is not equal to zero, obtain the content difference between the first text and the text sample.
[0278] In some embodiments, the mask module 4553 is also used to perform target segmentation on the image label based on the content difference to obtain a target segmentation result of the image label; wherein the target segmentation result includes a first image area and a second image area, the first image area corresponds to the content difference, and the second image area is different from the first image area; based on the target segmentation result, the pixel value of each pixel in the first image area is set to a first value, and the pixel value of each pixel in the second image area is set to a second value to obtain an image mask of the image label, wherein the first value is greater than the second value.
[0279] In some embodiments, the first text includes multiple words, and the content difference includes a target word among the multiple words. The above-mentioned mask module 4553 is also used to determine, for each of the words in the first text, a first pixel point corresponding to the word from the image label; based on the first pixel point corresponding to each of the words, perform region expansion on the first pixel point corresponding to the target word to obtain the first image area; and determine the image area in the image label that is different from the first image area as the second image area.
[0280] In some embodiments, the mask module 4553 is also used to perform feature extraction on the image label based on the word to obtain the image feature corresponding to the word, wherein the image feature includes multiple feature elements; select pixel points corresponding to the feature elements from the image label; and determine the pixel point corresponding to the feature element with the largest value as the first pixel point corresponding to the word.
[0281] In some embodiments, the mask module 4553 is also used to obtain a second pixel point corresponding to the target word, wherein the second pixel point belongs to a first pixel point corresponding to each of the words, and the second pixel point is adjacent to the first pixel point corresponding to the target word; based on the first pixel points corresponding to each of the words, the relative position relationship between the first pixel point corresponding to the target word and the boundary of the image label is determined; based on the relative position relationship and the second pixel point, the first pixel point corresponding to the target word is expanded to obtain the first image area.
[0282] In some embodiments, the mask module 4553 is further used to determine the regional boundary of the first image area based on the second pixel point of the target word when the relative position relationship indicates that there are no other first pixel points between the first pixel point corresponding to the target word and each of the boundaries; determine the regional boundary of the first image area based on the boundary and the second pixel point of the target word when the relative position relationship indicates that there are other first pixel points between the first pixel point corresponding to the target word and part of the boundaries; and determine the image area within the regional boundary in the image label as the first image area.
[0283] In some embodiments, the image mask includes multiple mask values, each of which corresponds to a pixel point in the image label. The above-mentioned training module 4554 is used to obtain the pixel value of each pixel point in the image label corresponding to the image mask; for each pixel point in the image label corresponding to the image mask, the pixel value of the pixel point is multiplied by the corresponding mask value to obtain the target pixel value of the pixel point, and the pixel value of the pixel point is updated based on the target pixel value to obtain the fused image label.
[0284] In some embodiments, the above-mentioned training module 4554 is also used to fuse the image mask to the image to obtain a fused image; and use the pixel point in the fused image as the third pixel point; for the fourth pixel point in the fused image label, determine the pixel value difference between the fourth pixel point and the corresponding third pixel point; sum the determined pixel value differences to obtain a first loss value, and update the model parameters of the image generation model based on the first loss value.
[0285] In some embodiments, the above-mentioned training module 4554 is also used to call the image generation model to generate an image corresponding to the first text based on the first text; determine a second loss value based on the image corresponding to the first text and the image label, and determine a third loss value based on the image and the fused image label; perform weighted summation of the second loss value and the third loss value to obtain a fourth loss value, and update the model parameters of the image generation model based on the fourth loss value.
[0286] In some embodiments, the image label includes multiple fifth pixel points, the image corresponding to the first text includes sixth pixel points corresponding one-to-one to the fifth pixel points, and the image mask includes mask values corresponding one-to-one to the fifth pixel points; the above-mentioned training module 4554 is also used to adjust each of the mask values in the image mask based on the content difference to obtain target mask values corresponding to each of the mask values; for each of the fifth pixel points, determine the pixel value difference between the fifth pixel point and the corresponding sixth pixel point, and multiply the pixel value difference and the target mask value corresponding to the fifth pixel point to obtain the loss value of the fifth pixel point; sum the loss values of each of the fifth pixel points to obtain the second loss value.
[0287] In some embodiments, the above-mentioned training module 4554 is also used to perform target segmentation on the image label based on the content difference to obtain a target segmentation result of the image label, wherein the target segmentation result includes a first image area in the image label corresponding to the content difference, and a second image area in the image label that is different from the first image area; the target mask value corresponding to the pixel point in the first image area in the image mask is determined as the mask value corresponding to the pixel point in the second image area; the target mask value corresponding to the pixel point in the second image area in the image mask is determined as the mask value corresponding to the pixel point in the first image area.
[0288] In some embodiments, the above-mentioned training module 4554 is also used to fuse the image mask to the image to obtain a fused image; and use the pixel point in the fused image as the seventh pixel point; for the eighth pixel point in the fused image label, determine the pixel value difference between the eighth pixel point and the corresponding seventh pixel point; sum the pixel point difference values corresponding to each of the eighth pixels to obtain the third loss value.
[0289] The following is a description of an exemplary structure of the image generating device 555 provided in the embodiment of the present application implemented as a software module. In some embodiments, Figure 3 As shown, the software modules stored in the image generation device 555 of the memory 550 may include: an image generation module 5551, used to obtain a third text, and call an image generation model to generate an image corresponding to the third text; wherein the image generation model is trained using the training method of the above-mentioned image generation model.
[0290] The embodiment of the present application provides a computer program product, which includes computer executable instructions or computer programs, and the computer executable instructions or computer programs are stored in a computer-readable storage medium. The processor of the electronic device reads the computer executable instructions or computer programs from the computer-readable storage medium, and the processor executes the computer executable instructions or computer programs, so that the electronic device executes the training method or image generation method of the image generation model described in the embodiment of the present application.
[0291] The present application embodiment provides a computer-readable storage medium storing computer-executable instructions or computer programs, wherein the computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will be caused to execute the training method of the image generation model or the image generation method provided in the present application embodiment, for example, Figure 4 The training method of the image generation model is shown.
[0292] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or it may be various electronic devices including one or any combination of the above memories.
[0293] In some embodiments, computer executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.
[0294] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file storing other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).
[0295] As an example, computer executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed at multiple sites and interconnected by a communication network.
[0296] In summary, the embodiments of the present application have the following beneficial effects:
[0297] (1) Based on the text sample, the image generation model is called to generate the image corresponding to the text sample, the image label of the text sample is converted to text, the first text corresponding to the image label is obtained, the content difference between the first text and the text sample is obtained, and based on the content difference, an image mask of the image label is generated, and the image mask is fused to the image label to obtain a fused image label; and the image generation model is trained based on the image and the fused image label. In this way, by performing text conversion on the image label of the text sample, the first text corresponding to the image label is obtained, and the content difference between the first text and the text sample is obtained. Since there is a content difference between the first text and the text sample, it means that there is an image area corresponding to the content difference in the image label of the text sample, which makes it difficult for the image label to adapt to the text sample. Therefore, it means that directly training the image generation model based on the image label will result in poor image generation quality of the trained image generation model. By generating an image mask of the image label based on the content difference, the image mask can accurately reflect the content difference between the first text and the text sample, and the image mask is fused into the image label to obtain a fused image label, so that the image area corresponding to the content difference in the image label is covered by the content difference reflected by the image mask, so that the obtained fused image label can cover the image area corresponding to the content difference in the image label. By training the image generation model based on the image and the fused image label, the training of the image generation model can effectively avoid the influence of the image area corresponding to the content difference in the image label, thereby effectively improving the image generation quality of the image generation model.
[0298] (2) It can capture the semantic relationship between each word in the text and the overall text in detail, thus providing a more refined perspective for text analysis. In particular, for understanding the importance of key words in the text and their contribution to the overall semantics, when processing words other than the first word, by constructing a second text containing the word and its previous text, the impact of contextual information on the meaning of the word can be considered. This processing method makes the similarity calculation more consistent with the contextual characteristics of natural language, and improves the accuracy and reliability of the similarity calculation. By analyzing the similarity between each word and the text sample, the main content and key information of the text can be more effectively identified, providing support for subsequent text understanding and application.
[0299] (3) By determining the similarity between each word in the first text and the text sample, and taking the words with similarity below the threshold as the first words, the server can accurately identify the key differences in the text. This improves the level of refinement of text analysis, allowing the server to capture the contribution of each word in the text to the overall semantics, thereby gaining a deeper understanding of the text content; by determining the first word, the server can filter out words that are weakly related to the subject content of the text sample or semantically mismatched, which helps filter noise and enhance the robustness of text processing; selecting the second word corresponding to the first word and taking the two as content differences not only reveals the significant differences between the first text and the text sample, but also provides an important semantic basis for applications such as text comparative analysis, information extraction, and text generation, thereby optimizing the performance and effects of natural language processing-related tasks.
[0300] (4) By encoding the first text to obtain the third vector and encoding the text sample to obtain the fourth vector, the server can effectively grasp the position of the two texts in the semantic space. Calculating the distance between the two vectors not only provides the server with a means to quantify the similarity between the first text and the text sample, but also directly reveals the content difference between the two when the distance value is not equal to zero. It provides an objective metric for text similarity assessment, which helps to improve the accuracy of text classification, retrieval and recommendation; the identification of content differences enables the server to better understand the deep meaning and nuances of the text, which is crucial for text analysis and information extraction tasks.
[0301] (5) By associating the words in the first text with the pixels in the image label and performing region expansion based on these associations, the server can accurately identify and segment the parts of the image that are directly related to the text content. This greatly improves the accuracy of image understanding, not only identifying objects in the image, but also associating these objects with specific concepts in the text description, thereby providing richer semantic information for image annotation, classification and retrieval. This makes image analysis more flexible, and the importance of image regions can be adjusted according to the differences in different text content, thereby optimizing the results of image processing tasks. By distinguishing the first image region from the second image region, the server can effectively distinguish the main content and background information in the image.
[0302] (6) It can accurately locate the visual elements in the image that are most relevant to the target word, thereby improving the accuracy and efficiency of understanding the image content. By analyzing the characteristic elements of the image, the server can distinguish the main content and minor details of the image, which is crucial for image recognition and classification tasks. Selecting the pixel with the largest eigenvalue as the first pixel ensures that the server can focus on the most significant and representative part of the image, which helps to improve the robustness of image segmentation and target detection.
[0303] (7) It can ensure the accuracy and efficiency of the region expansion process because it expands based on the spatial proximity of pixels, which helps maintain the continuity and structural integrity of objects in the image. By obtaining a second pixel adjacent to the first pixel, the server can more finely control the direction and range of region expansion, avoid excessive or unnecessary expansion, and thus improve the quality of image segmentation. Considering the relative position relationship enables the server to adapt to the limitations of the image boundary and avoid exceeding the actual range of the image when expanding the region, which is crucial to maintaining the authenticity and accuracy of the image content.
[0304] (8) By flexibly handling two different situations - whether there are other first pixels around the first pixel corresponding to the target word and whether there is overlap with the image boundary, the region expansion can accurately capture the details of the target object and adapt to the limitations of the image boundary. In the first case, the region boundary is determined by the second pixel, which ensures the accuracy of region expansion and the clarity of the object boundary; in the second case, the combination of the boundary and the second pixel information ensures that even in complex image structures, the region can be effectively expanded without exceeding the actual range of the image. The quality of image segmentation is optimized and the ability to understand and analyze image content is improved.
[0305] (9) By accurately distinguishing the target object from the background, the efficiency and accuracy of image segmentation are ensured, which helps to extract key information from complex images. Secondly, by setting the pixel value of the first image area to the first value and the pixel value of the second image area to the second value, the image mask not only clearly identifies the target object, but also the binarization processing method greatly reduces the amount of data processing, improves the processing speed, and provides the possibility for real-time image processing. Finally, the generation of the image mask makes the visual contrast between the target object and the background more obvious, which is convenient for manual inspection and editing, thereby improving the visualization and interactivity of the entire image processing process.
[0306] (10) Not only does it strengthen the visual representation of the target area in the image label, but it also increases the emphasis on the target area, making image analysis and recognition tasks more accurate and efficient. Through the multiplication operation, the pixel values of the target area are enhanced, while the pixel values of the background or non-target area are weakened or ignored, thereby visually highlighting the key area. This processing not only improves the interpretability of the image, but the image label fused with the image mask provides richer and more accurate information for image understanding.
[0307] (11) The image mask is fused to the original image, and the pixel after the fusion mask is defined as the third pixel. For the training and optimization of the image generation model, by directly superimposing the mask on the image, the specific area that the model needs to focus on and optimize, that is, the target area, is highlighted. For the fourth pixel in the image label of the fusion mask, the pixel value difference between it and the corresponding third pixel is calculated, which can quantify the difference between the model output and the real target, providing direct feedback information for model training. The first loss value obtained by summing these pixel value differences is a key indicator for measuring model performance, which reveals the overall error between the model-generated image and the target image. Updating the model parameters of the image generation model based on the first loss value helps the model to more accurately capture the characteristics of the target area, reduce errors, and improve the quality and realism of the generated image.
[0308] (12) By adjusting the image mask and calculating the second loss value based on content differences, and by personalizing the mask value, the model can pay more attention to key areas, thereby strengthening the importance of these areas when generating images. The loss value calculation of each fifth pixel matches its position and role in the image, ensuring the pertinence and effectiveness of model training. By multiplying the pixel value difference between the fifth pixel and the corresponding sixth pixel with the target mask value, the model can more accurately quantify and optimize the error in a specific area, thereby improving the consistency between the generated image and the labeled image. The second loss value obtained by summing the loss values of all fifth pixels provides a comprehensive and detailed error metric for the model, allowing the model to more efficiently learn the mapping relationship from text to image during training, thereby generating more accurate and realistic images.
[0309] (13) By adjusting the mask values in the image mask based on content differences and performing target segmentation on the image labels, the model can identify key areas with large content differences and relatively stable background areas, and then optimize the mask values according to the characteristics of these areas. This adjustment strategy enables the model to focus more on those areas that need to be optimized during training, thereby improving the pertinence and efficiency of training. By exchanging the mask values of the first image area and the second image area, the model is guided to pay attention to the background area that may have been ignored, while reducing the attention to the relatively matched character area. This strategy not only helps to improve the performance of the model in background generation, but also optimizes the model's ability to understand and reproduce the overall image structure.
[0310] (14) Allows precise control of specific areas during model training. By applying masks to images, the model can focus on those parts that need to be optimized, thereby achieving more detailed and accurate image generation in key areas. Secondly, the pixel value difference between the seventh pixel and the eighth pixel is calculated, and the third loss value is determined by summing these differences, which effectively quantifies the generation error of the model in the masked area and provides a clear optimization target for the model. This targeted error metric helps the model better understand the local features and details in the training data, thereby improving the quality and realism of the generated images.
[0311] (15) This allows the model to focus not only on the overall visual quality when generating images, but also on the accuracy of specific areas or details, which is crucial for achieving more refined image control. The fourth loss value obtained by weighted summation provides a comprehensive performance indicator that can more comprehensively reflect the performance of the model at the global and local levels, thereby guiding the update of model parameters. This enables the model to better understand the relationship between the text description and the image content, and generate images that are more consistent with the text description and rich in details. By updating the model parameters based on the fourth loss value, the generation quality of the model can be effectively improved, the error can be reduced, and the model can be more robust and reliable in image generation tasks, thereby broadening its scope of application.
[0312] (16) By acquiring text and calling the image generation model to generate the corresponding image, a rapid conversion from text description to visual content can be achieved. The image generation model can understand the semantic relationship between text and image by learning a large number of text samples and corresponding images, and generate high-quality images that meet user expectations. The training method of integrating image labels can improve the image generation model's understanding of text content and the accuracy of image generation, making the generated image closer to the text description.
[0313] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent substitutions and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A training method for an image generation model, characterized in that: The method comprises: Based on the text sample, calling the image generation model to generate an image corresponding to the text sample; Performing text conversion on the image tag of the text sample to obtain a first text corresponding to the image tag, where the first text is used to describe the content in the image tag; Obtaining a content difference between the first text and the text sample, and generating an image mask of the image label based on the content difference; Fusion the image mask to the image label to obtain a fused image label; The image generation model is trained based on the image and the fused image label.
2. The method according to claim 1, characterized in that The first text and the text sample both include a plurality of words, and obtaining the content difference between the first text and the text sample includes: respectively determining the similarity between each of the words in the first text and the text sample; Determine a word whose similarity to the text sample is less than a similarity threshold as a first word; A second word corresponding to the first word is selected from the text sample, and the first word and the second word are used as the content difference between the first text and the text sample.
3. The method according to claim 2, characterized in that The respectively determining the similarity between each of the words in the first text and the text sample comprises: The following processing is performed for each word in the first text: When the word is the first word in the first text, respectively encode the word and the text sample to obtain a first vector and a second vector, and determine the similarity between the first vector and the second vector as the similarity between the word and the text sample; When the word is not the first word in the first text, a second text is constructed based on the word and the text preceding the word in the first text, and the similarity between the second text and the text sample is determined as the similarity between the word and the text sample.
4. The method according to claim 1, characterized in that: The obtaining of the content difference between the first text and the text sample includes: Encoding the first text to obtain a third vector, and encoding the text sample to obtain a fourth vector; The distance between the third vector and the fourth vector is determined, and when the value of the distance is not equal to zero, the content difference between the first text and the text sample is obtained.
5. The method according to any one of claims 1 to 4, characterized in that: The step of generating an image mask of the image tag based on the content difference comprises: Based on the content difference, performing target segmentation on the image label to obtain a target segmentation result of the image label; The target segmentation result includes a first image region and a second image region, the first image region corresponds to the content difference, and the second image region is different from the first image region; Based on the target segmentation result, the pixel value of each pixel in the first image area is set to a first value, and the pixel value of each pixel in the second image area is set to a second value to obtain an image mask of the image label, wherein the first value is greater than the second value.
6. The method according to claim 5, characterized in that The first text includes a plurality of words, the content difference includes a target word among the plurality of words, and the target segmentation of the image label based on the content difference is performed to obtain the target segmentation result of the image label, including: For each of the words in the first text, determining a first pixel point corresponding to the word from the image label; Based on the first pixel points corresponding to the words, performing region expansion on the first pixel points corresponding to the target words to obtain the first image region; An image region in the image tag that is different from the first image region is determined as the second image region.
7. The method according to claim 6, characterized in that The determining, from the image tag, a first pixel point corresponding to the word comprises: Based on the words, feature extraction is performed on the image tags to obtain image features corresponding to the words, wherein the image features include multiple feature elements; Selecting pixel points corresponding to the feature elements from the image labels; The pixel point corresponding to the feature element with the largest value is determined as the first pixel point corresponding to the word.
8. The method according to claim 6, characterized in that The step of performing region expansion on the first pixel points corresponding to the target words based on the first pixel points corresponding to the words to obtain the first image region includes: Acquire a second pixel point corresponding to the target word, where the second pixel point belongs to a first pixel point corresponding to each of the words, and the second pixel point is adjacent to the first pixel point corresponding to the target word; Based on the first pixel points respectively corresponding to the words, determining the relative position relationship between the first pixel point corresponding to the target word and the boundary of the image label; Based on the relative position relationship and the second pixel point, the first pixel point corresponding to the target word is region expanded to obtain the first image region.
9. The method according to claim 8, characterized in that The step of performing region expansion on the first pixel point corresponding to the target word based on the relative position relationship and the second pixel point to obtain the first image region includes: When the relative position relationship indicates that no other first pixel points exist between the first pixel point corresponding to the target word and each of the boundaries, determining the region boundary of the first image region based on the second pixel point of the target word; When the relative position relationship indicates that there are other first pixel points between the first pixel point corresponding to the target word and part of the boundary, determining the region boundary of the first image region based on the boundary and the second pixel point of the target word; An image region within the region boundary in the image tag is determined as the first image region.
10. The method according to any one of claims 1 to 9, characterized in that: The image mask includes a plurality of mask values, each of the mask values corresponds to a pixel point in the image label, and fusing the image mask to the image label to obtain a fused image label includes: Obtaining a pixel value of each pixel point in the image label corresponding to the image mask; For each pixel point in the image label corresponding to the image mask, multiply the pixel value of the pixel point by the corresponding mask value to obtain the target pixel value of the pixel point, and update the pixel value of the pixel point based on the target pixel value to obtain the fused image label.
11. The method according to any one of claims 1 to 10, characterized in that: The step of training the image generation model based on the image and the fused image label comprises: Fusing the image mask to the image to obtain a fused image, and using a pixel point in the fused image as a third pixel point; For a fourth pixel point in the fused image label, determining a pixel value difference between the fourth pixel point and a corresponding third pixel point; The determined pixel value differences are summed to obtain a first loss value, and based on the first loss value, the model parameters of the image generation model are updated.
12. The method according to any one of claims 1 to 10, characterized in that: The step of training the image generation model based on the image and the fused image label comprises: Based on the first text, calling the image generation model to generate an image corresponding to the first text; Determine a second loss value based on the image corresponding to the first text and the image label, and determine a third loss value based on the image and the fused image label; The second loss value and the third loss value are weightedly summed to obtain a fourth loss value, and based on the fourth loss value, the model parameters of the image generation model are updated.
13. The method according to claim 12, characterized in that The image label includes a plurality of fifth pixels, the image corresponding to the first text includes sixth pixels corresponding one-to-one to the fifth pixels, and the image mask includes mask values corresponding one-to-one to the fifth pixels; The determining a second loss value based on the image corresponding to the first text and the image label includes: Based on the content difference, each mask value in the image mask is adjusted to obtain a target mask value corresponding to each mask value; For each of the fifth pixel points, determine a pixel value difference between the fifth pixel point and the corresponding sixth pixel point, and multiply the pixel value difference by a target mask value corresponding to the fifth pixel point to obtain a loss value of the fifth pixel point; The loss values of each of the fifth pixels are summed to obtain the second loss value.
14. The method according to claim 13, characterized in that The adjusting each mask value in the image mask based on the content difference to obtain a target mask value corresponding to each mask value includes: Based on the content difference, performing target segmentation on the image label to obtain a target segmentation result of the image label, wherein the target segmentation result includes a first image region in the image label corresponding to the content difference and a second image region in the image label that is different from the first image region; Determining the mask value corresponding to the pixel point in the second image area as the target mask value corresponding to the pixel point in the first image area in the image mask; The mask value corresponding to the pixel point in the first image area is determined as the target mask value corresponding to the pixel point in the second image area in the image mask.
15. The method according to claim 12, characterized in that The determining of a third loss value based on the image and the fused image label comprises: Fusion the image mask to the image to obtain a fused image, and use a pixel point in the fused image as the seventh pixel point; For an eighth pixel point in the fused image label, determining a pixel value difference between the eighth pixel point and a corresponding seventh pixel point; The pixel point difference values corresponding to each of the eighth pixel points are summed to obtain the third loss value.
16. An image generation method, characterized in that: The method comprises: Obtaining a third text, and calling an image generation model to generate an image corresponding to the third text; Wherein, the image generation model is trained using the method described in any one of claims 1 to 15.
17. A training device for an image generation model, characterized in that: The device comprises: A generation module, used to call an image generation model to generate an image corresponding to the text sample based on the text sample; A conversion module, configured to perform text conversion on the image label of the text sample to obtain a first text corresponding to the image label, wherein the first text is used to describe the content in the image label; A mask module, configured to obtain a content difference between the first text and the text sample, and generate an image mask of the image label based on the content difference; A training module is used to fuse the image mask to the image label to obtain a fused image label; and train the image generation model based on the image and the fused image label.
18. An electronic device, characterized in that: The electronic device comprises: A memory for storing computer executable instructions or computer programs; A processor, configured to implement the method according to any one of claims 1 to 16 when executing computer executable instructions or computer programs stored in the memory.
19. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the method according to any one of claims 1 to 16 is implemented.
20. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the method according to any one of claims 1 to 16 is implemented.
Citation Information
Cited By
Focus area determination method and device, electronic equipment and storage medium
CN121074499A
Image-based screen defect detection method, system and device
CN121746387A