Image processing method and device, electronic equipment and storage medium

CN122153424APending Publication Date: 2026-06-05NETEASE (HANGZHOU) NETWORK CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NETEASE (HANGZHOU) NETWORK CO LTD
Filing Date
2024-12-03
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing open-source datasets ignore fine-grained details in text-to-image generation models, resulting in generated images that do not match the text descriptions and are unable to generate continuous story images.

Method used

By acquiring multiple images related to image keywords, segmenting them into image clusters, and generating role mask images and index codes for each image, a multi-role consistency dataset is constructed. A text generation model is then used to generate descriptive text, ensuring the matching and continuity between images and text.

Benefits of technology

It improves the accuracy of image-text matching and can generate continuous story images that maintain frame consistency, making it suitable for fields such as visual reasoning, image editing, video games, and computer-aided design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122153424A_ABST
    Figure CN122153424A_ABST
Patent Text Reader

Abstract

The application provides an image processing method and device, electronic equipment and storage medium. The method comprises: acquiring a plurality of images related to at least one image keyword; segmenting the plurality of images according to the at least one image keyword to obtain image clusters under different image keywords; for each image cluster, the following processing is performed: generating a description text of each image in the image cluster, and obtaining a role mask image of each image and an index code of the image cluster based on each image and / or the description text, wherein the index code comprises at least one position symbol, and each position symbol is used to indicate a character position in the description text which has the same description part for the same object. A multi-role consistency dataset can be constructed by the application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer application technology, and in particular to an image processing method, apparatus, electronic device and storage medium. Background Technology

[0002] Text-to-image conversion is the process of converting text into images using artificial intelligence technology. It can generate realistic images that match the description based on given text and has been widely used in fields such as visual reasoning, image editing, video games, animation production, and computer-aided design.

[0003] In related technologies, text-to-image generation models (e.g., text-to-image diffusion models) are typically used to generate images from text (i.e., natural language descriptions) through a diffusion process. Generally, open-source datasets available on the internet are used for training text-to-image generation models. These datasets contain hundreds of millions of text-image pairs (i.e., a text and a generated image form a text-image pair).

[0004] However, the text-image pairs provided by the aforementioned open-source datasets are relatively coarse. They focus more on representing the category information of the target objects in the image (such as people, animals, etc.) and ignore fine-grained details. As a result, the text-image generation model trained on this open-source dataset cannot generate images that accurately represent the fine-grained details in the text, making it impossible for the generated images to match the images described in the text accurately. Summary of the Invention

[0005] In view of the above, embodiments of this application provide at least one image processing method, apparatus, electronic device, and storage medium to overcome at least one of the above-mentioned defects.

[0006] In a first aspect, an exemplary embodiment of this application provides an image processing method, the method comprising: acquiring multiple images associated with at least one image keyword; segmenting the multiple images according to the at least one image keyword to obtain image clusters under different image keywords; and for each image cluster, performing the following processing: generating descriptive text for each image in the image cluster, and obtaining a role mask image for each image and an index code for the image cluster based on each image and / or the descriptive text, the index code including at least one position character, each position character indicating the position of a character in the descriptive text that has the same descriptive part for the same object.

[0007] Secondly, embodiments of this application also provide an image processing apparatus, the apparatus comprising: an acquisition module for acquiring multiple images associated with at least one image keyword; a segmentation module for segmenting the multiple images according to the at least one image keyword to obtain image clusters under different image keywords; and an identification module for performing the following processing for each image cluster: generating descriptive text for each image in the image cluster, and obtaining a role mask image for each image and an index code for the image cluster based on each image and / or the descriptive text, wherein the index code includes at least one position character, each position character indicating the position of a character in the descriptive text that has the same descriptive part for the same object.

[0008] Thirdly, embodiments of this application also provide an electronic device, a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the above-described image processing method.

[0009] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the above-described image processing method.

[0010] The image processing method, apparatus, electronic device, and storage medium provided in this application can quickly and effectively construct a multi-role consistent dataset to provide data support for subsequent processing.

[0011] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 A flowchart illustrating an exemplary embodiment of the image processing method provided in this application is shown.

[0014] Figure 2 This illustration shows a schematic diagram of the processing flow provided by an exemplary embodiment of this application;

[0015] Figure 3A flowchart illustrating the steps of acquiring multiple images provided in an exemplary embodiment of this application;

[0016] Figure 4 A schematic diagram illustrating the flowchart for calculating the aesthetic score of an image according to an exemplary embodiment of this application;

[0017] Figure 5 A flowchart illustrating the steps of filtering image clusters provided in an exemplary embodiment of this application;

[0018] Figure 6 A flowchart illustrating the steps for generating index codes provided in an exemplary embodiment of this application;

[0019] Figure 7 This invention provides a schematic diagram of the structure of an image processing apparatus according to an exemplary embodiment of the present application.

[0020] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an exemplary embodiment of this application. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0022] The terms “a,” “an,” “the,” and “the” are used in this specification to indicate the presence of one or more elements / components / etc.; the terms “including” and “having” are used to indicate an open-ended inclusion and to mean that there may be other elements / components / etc. in addition to the listed elements / components / etc.; the terms “first” and “second” are used only as markings and are not a limitation on the number of objects.

[0023] It should be understood that in the embodiments of this application, "at least one" means one or more, and "more than one" means two or more. "And / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. The character " / " generally indicates that the related objects before and after it are in an "or" relationship. "Contains A, B and / or C" means containing any one, two, or three of A, B, and C.

[0024] It should be understood that in the embodiments of this application, "B corresponding to A", "B corresponding to A", "A corresponds to B" or "B corresponds to A" means that B is associated with A, and B can be determined based on A. Determining B based on A does not mean that B is determined solely based on A; B can also be determined based on A and / or other information.

[0025] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0026] Text-to-image conversion is the process of converting text into images using artificial intelligence technology. It can generate realistic images that match the description based on given text and has been widely used in fields such as visual reasoning, image editing, video games, animation production, and computer-aided design.

[0027] In related technologies, text-to-image generation models (e.g., text-to-image diffusion models) are typically used to generate images from text (i.e., natural language descriptions) through a diffusion process. Generally, open-source datasets available on the internet are used for training these models, containing hundreds of millions of text-image pairs (i.e., a descriptive text and a generated image form a text-image pair).

[0028] However, the text-image pairs provided by the aforementioned open-source datasets are relatively coarse. They focus more on representing the category information of the target objects in the image (such as people, animals, etc.) and ignore fine-grained details. As a result, the text-image generation model trained on this open-source dataset cannot generate images that accurately represent the fine-grained details in the text, making it impossible for the generated images to match the images described in the text accurately.

[0029] The lack of fine-grained detail in open-source datasets makes it difficult for images generated by text-to-image generation models to fully express the semantics of the target text. Related techniques typically rely on manually labeling each image and / or text to obtain additional annotations, but this approach leads to a significant increase in manual labor costs.

[0030] Furthermore, with the continuous advancement of text-to-image technology, there has been great interest in generating stories or videos based on text-to-image generation models. However, open-source datasets in related technologies can only provide text-image pairs consisting of a single text and a single generated image, which cannot meet the needs of generating continuous story images.

[0031] In response to the problems in at least one of the above aspects, this application provides, in one aspect, a text-image dataset that can cover more fine-grained details, in order to help improve the accuracy of image-text matching.

[0032] Furthermore, existing open-source datasets often lack fixed feature attributes or rich background diversity in their text-image pairs, making it difficult for text-image generation models trained on these datasets to meet the needs of generating continuous story images. To address this, the dataset provided in this application uses index codes to indicate the same descriptive parts for the same object in different descriptive texts, ensuring natural and flexible character poses while maintaining consistency between frames. Moreover, it achieves clear separation of foreground and background by outputting a title for each image cluster and a separate description for each image, vividly representing the protagonist's pose and / or the background.

[0033] To facilitate understanding of this application, the image processing methods, apparatus, electronic devices, and storage media provided in the embodiments of this application will be described in detail below.

[0034] Please see Figure 1 This is a flowchart of an image processing method provided as an exemplary embodiment of this application.

[0035] Reference Figure 1 In step S101, multiple images related to at least one image keyword are acquired.

[0036] For example, various keyword image search methods can be used to obtain the above-mentioned multiple images. Here, each obtained image contains at least one image keyword. This application does not limit the method of obtaining images.

[0037] In an exemplary embodiment of this application, a multi-role consistency dataset can be constructed by selecting multiple roles and collecting images of each role set in different environments, layouts, and gestures. The constructed diverse and finely structured dataset allows roles to maintain identity consistency when performing different operations in different scenarios. The specific method of constructing this dataset will be described later.

[0038] The following is combined with Figure 2 The processing flow shown details the specific process of each step in the image processing method of this application. It should be understood that... Figure 2 The processing method shown is only an example, and this application is not limited to it.

[0039] exist Figure 2 In the example shown, process 11 is used to represent batch acquisition of multiple images. The following section will combine... Figure 3 This section introduces an exemplary method for acquiring multiple images.

[0040] Figure 3 A flowchart illustrating the steps for acquiring multiple images provided by an exemplary embodiment of this application is shown.

[0041] like Figure 3 As shown, in step S201, multiple initial images are acquired in batches based on at least one image keyword.

[0042] For example, multiple initial images can be collected from the internet and / or open-source datasets based on user-inputted image keywords or image keywords determined through other means to form an initial dataset. For instance, based on the principle of keyword image search, and with the help of image recognition technology and indexing mechanisms of search engines, images related to image keywords can be searched by matching image keywords with features in an image database.

[0043] In step S202, multiple images are selected from the multiple initial images based on the aesthetic score of each initial image.

[0044] For example, for each initial image in the initial dataset formed above, an aesthetic score can be calculated to help filter the dataset during batch downloading.

[0045] Figure 4 This is a schematic flowchart illustrating the calculation of an aesthetic score for an image provided by an exemplary embodiment of this application.

[0046] In this example, a preliminary image can be input into the CLIP model (Contrastive Language-Image Pre-training, a pre-trained model based on contrastive text-image pairs) to extract its image features. These features are then input into an MLP (Multi-Layer Perceptron) to obtain an aesthetic score for the preliminary image. This process can be repeated for each preliminary image in the dataset to obtain an aesthetic score for each image.

[0047] For example, the CLIP model achieves more accurate image feature extraction and classification by jointly processing images and text. MLP is a feedforward artificial neural network model that can consist of an input layer, several hidden layers, and an output layer, used in machine learning tasks such as classification and regression. In this application, the construction method of the CLIP model and MLP is not limited, as long as they can achieve the aforementioned image feature extraction and image aesthetic scoring. Alternatively, other methods can also be used to obtain the image aesthetic score.

[0048] Furthermore, it should be understood that the above-described method of filtering the dataset based on image aesthetic rating is only an example. Dataset filtering can also be performed using other dimensions besides evaluating the aesthetic appeal of images, and this application does not limit this to such methods.

[0049] return Figure 1 In step S102, multiple images are segmented according to at least one image keyword to obtain image clusters under different image keywords.

[0050] For example, each image keyword can be treated as a corresponding IP address, thereby enabling image clustering and data cleaning based on IP addresses, as described above. Figure 2 The processing procedure in the example is shown in step 22.

[0051] This application does not restrict the specific method of obtaining image clusters, as long as it can form an image cluster for each image keyword. That is, it is equivalent to using specific image keywords to segment multiple images. The images included in different image clusters can be completely different or partially repeated.

[0052] In an alternative example, K-means clustering can be used to cluster multiple images in the initial dataset to generate at least one smaller image cluster.

[0053] For example, in k-means clustering, k is the number of clusters required as input. In this embodiment, the number of cluster centers is determined by the number of image keywords crawled each time. For instance, k points are created as initial centroids. When the cluster assignment of any point changes, for each data point in the filtered initial dataset, for each centroid, the distance between the centroid and the data point is calculated, and the data point is assigned to the nearest cluster. For each cluster, the mean of all points in the cluster is calculated, and the mean is used as the centroid.

[0054] In a preferred example, after performing the above image clustering process, data cleaning can also be performed for each image cluster.

[0055] For example, CLIP can be used to evaluate each image in each image cluster from the text-image dimension and / or the image-image dimension, thereby filtering out unqualified samples from the image cluster based on the evaluation results.

[0056] The following is combined with Figure 5 This section will introduce the process of evaluating each image cluster from different dimensions. It should be understood that... Figure 5 The image cluster filtering method shown is only a preferred example. Image clusters can also be filtered in other ways, and this application does not limit this.

[0057] Figure 5 A flowchart illustrating the steps of filtering image clusters provided by an exemplary embodiment of this application is shown.

[0058] like Figure 5 As shown, in step S301, the first alignment index and / or the second alignment index corresponding to each image in the image cluster are calculated.

[0059] In this embodiment of the application, the first comparison index is used to characterize the degree of matching between the image and the reference text. Here, the reference text can be determined based on the image keywords corresponding to the image cluster.

[0060] For example, the image keywords of the image cluster can be directly used as the reference text, and various image-text comparison methods can be used to obtain a first comparison index to characterize the degree of matching between the image and the reference text. This application does not limit the image-text comparison method.

[0061] In this embodiment of the application, the second comparison index is used to characterize the degree of matching between the image and the reference image. Here, the reference image can be determined based on the image keywords corresponding to the image cluster.

[0062] For example, the reference image may refer to the image that best represents the corresponding image keyword. Optionally, the reference image may not be included in the image cluster. For example, an image may be assigned as the reference image in advance by manually annotating the image keyword. Alternatively, the reference image may be included in the image cluster, such as the image that best represents the image keyword (e.g., the one with the highest relevance) selected directly from the image cluster.

[0063] Based on this, various image-to-image comparison methods are used to obtain a second comparison index to characterize the degree of matching between the image and the reference image. This application does not limit the image-to-image comparison method.

[0064] In step S302, the images are filtered according to the first alignment index and / or the second alignment index corresponding to each image in the image cluster.

[0065] In the embodiments of this application, image clusters can be filtered based on one of the first alignment index and the second alignment index, or they can be filtered based on both the first alignment index and the second alignment index. This application does not limit this.

[0066] For example, a threshold can be set for different dimensions. For instance, a text-image threshold can be set for the text-image dimension. The first comparison index of each image in the image cluster is compared with the text-image threshold. Images with a first comparison index greater than the text-image threshold are retained, and images with a first comparison index not greater than (less than or equal to) the text-image threshold are removed from the image cluster.

[0067] And / or, a picture-to-picture threshold can also be set for the image-to-image dimension. For example, the second alignment index of each image in the image cluster is compared with the picture-to-picture threshold. Images with a second alignment index greater than the picture-to-picture threshold are retained, and images with a second alignment index not greater than the picture-to-picture threshold are removed from the image cluster.

[0068] by Figure 2 The example shown illustrates that the image cluster in process 22 represents an image cluster that has undergone IP clustering and data cleaning, and may include one or more such image clusters.

[0069] return Figure 1 In step S103, the role mask image of each image in each image cluster and the index code of the image cluster are obtained.

[0070] Through the above processing, multiple image clusters can be efficiently formed for large-scale images acquired in batches, and role mask images and index codes under each image cluster can be generated to construct a multi-role consistent dataset.

[0071] In the embodiments of this application, descriptive text for each image in each image cluster can be generated first, that is, one image corresponds to one descriptive text. For example, text for describing the image content can be generated in various ways, and this application does not limit this.

[0072] In a preferred embodiment, all images in the image cluster can be identified using a text generation model (e.g., GPT-4v) to generate descriptive text for each image.

[0073] For example, all images in an image cluster are collectively input into GPT-4v to obtain descriptive text for each image. Simultaneously, character alignment is performed on each descriptive text within GPT-4v. For instance, character alignment may include, but is not limited to, extracting the same character descriptions from each descriptive text to label the image set; for example, the same character descriptions could be used as image titles.

[0074] Furthermore, GPT-4v performs fine-grained cleaning on each image within an image cluster. For example, after performing the character alignment process described above, images that can be aligned are retained, while images that cannot be aligned (i.e., those whose descriptive text does not contain the same characters as the title) are removed. In addition, manual correction of incompatible images within the image cluster can be combined, such as manually removing images from the cluster that do not meet the requirements.

[0075] In this way, GPT-4v can not only output the descriptive text of each image in an image cluster, but also output more detailed descriptive text for each image. For example, it can output a title (i.e., common description, which refers to the same character description extracted from each descriptive text) for each image cluster, as well as a separate description for each image (i.e., individual description, which refers to the character description in the image's descriptive text other than the title).

[0076] In a preferred embodiment of this application, the descriptive text includes descriptive text for different objects. For example, different objects may include, but are not limited to, at least one of the following: characters in an image, background, and characters' actions. Here, characters may include, but are not limited to, at least one of the following: real humans, real animals, cartoon animals, and cartoon characters.

[0077] In an optional embodiment, the following configuration can be pre-executed for the text generation model to achieve the above character alignment processing under the constraints of the configuration. Exemplarily, the configuration processing for the text generation model includes at least one of the following: describing each image with the same sentence structure; using the same words to describe the same character; using the same words to describe the same modifier of the same character; describing the background with the same grammatical structure; and describing the action with the same grammatical structure.

[0078] Following the above Figure 2 The example of the image cluster shown, the configuration performed in GPT4V may include: all that is needed is to describe each image in the image cluster. Please describe each of the following images in approximately 50 words using the same sentence structure. The description includes the characters, background, surrounding decorations, and the characters' actions. If each character is dressed the same, then the description of the clothing in each description text will be the same. Treat each image to be described as a separate image, without a description linked to the previous one.

[0079] against Figure 2 In the image cluster shown, during processing step 33, the descriptive text generated for image 1 in the image cluster includes: a woman in a floral dress and a man in a white T-shirt and gray pants, on a sunny street with trees and white buildings in the background. The descriptive text generated for image 2 in the image cluster includes: a woman in a floral dress and a man in a white T-shirt and gray pants, kissing in front of a café. A complete list for each image is no longer provided. In the above processing steps of this example, the same words are used for describing the characters and clothing.

[0080] In this example, the output of the text generation model includes a title 301 for the image cluster, such as "A woman in a floral dress, a man in a white T-shirt and gray pants". Additionally, it outputs a separate description for each image, for example, Figure 2 302 to 305 in the image cluster are individual descriptions for each image, that is, different descriptions for the four images in the cluster. For example, the individual description 302 for image 1 in the cluster is "They are on a sunny street with trees and white buildings in the background", the individual description 303 for image 2 in the cluster is "They are kissing in front of a cafe", the individual description 304 for image 3 in the cluster is "They are standing side by side in front of a vibrant graffiti wall", and the individual description 305 for image 4 in the cluster is "The man is looking to the right, and they are standing next to a white wall".

[0081] In a preferred embodiment, the generated descriptive text can be returned in JSON format. For image descriptive text containing multiple objects, delimiters are used for separation. For example, {"Image 1":", "Image 2":", ..., "Image n":"}.

[0082] Furthermore, in the embodiments of this application, the character mask image of each image can be generated in the following way: based on each image and the description text corresponding to each image, the character mask image of each image is obtained.

[0083] In an optional example, a character mask image for each image can be obtained using a pre-trained segmentation model (e.g., Segment Anything) based on each image in the image cluster and the corresponding descriptive text (including the title and a separate description for each image). Various mask image generation methods can be used to obtain the character mask image, and this application does not impose any limitations on this.

[0084] Following the above Figure 2 In the example shown, in process 44, a character mask image for each image in the image cluster is obtained through a segmentation model. The character mask image distinguishes between the character and the background, and also distinguishes between different characters.

[0085] In this embodiment, an index code for an image cluster can be generated based on the descriptive text of each image. The following is a detailed explanation... Figure 6 This application presents an example of generating index codes, but is not limited to this.

[0086] Figure 6 A flowchart illustrating the steps for generating index codes provided in an exemplary embodiment of this application is shown.

[0087] like Figure 6 As shown, in step S401, the same description portion is extracted from the description text corresponding to all images in the image cluster.

[0088] For example, after obtaining the descriptive text corresponding to each image in an image cluster, the common descriptive parts can be extracted.

[0089] In the case where the text generation model outputs a title containing image clusters, the title of the image cluster can be directly obtained in step S401.

[0090] In step S402, the character position of the specified character in the extracted identical description portion in the description text is determined.

[0091] In the embodiments of this application, the specified character can be any character in the same description part, and the number of specified characters can be one or more. For example, the specified characters can include, but are not limited to, the start character and / or end character of the same description part.

[0092] In step S403, an index code is generated based on the determined character position. For example, the index code includes the character position.

[0093] Here, each image may include one or more objects. In a preferred embodiment, for the case where each image includes multiple objects, the same descriptive part can be extracted for each object to generate a location symbol for each object.

[0094] For example, for each object, the common descriptive portion for that object is extracted from all descriptive texts of the image cluster. In this case, the index code includes a positional character corresponding to each object.

[0095] Based on the above Figure 2 The title 301 of the image cluster shown includes, for example, a woman wearing a floral dress and a man wearing a white T-shirt and gray pants. The objects in the image include a woman and a man. In this case, a locator can be determined based on the description of "a woman wearing a floral dress" and another locator can be determined based on the description of "a man wearing a white T-shirt and gray pants". In this case, the index code of the image cluster includes the above two locators.

[0096] In other words, the index code of an image cluster includes at least one position character, each position character being used to indicate the position of a character in the descriptive text that has the same descriptive part for the same object.

[0097] In an optional example, the index code may also include at least one length identifier, with each length identifier corresponding to a location identifier, to locate the common descriptive portion of an object within an image cluster based on a pair of location identifiers and length identifiers. Alternatively, the length identifier may not be defined, and the descriptive portion between the nearest delimiter before the location identifier and the nearest delimiter after the location identifier can be directly extracted from the descriptive text (or a file stored in JSON format) as the common descriptive portion of an object.

[0098] by Figure 2 Taking the example shown, the image processing scheme of this application embodiment can output the role mask image of each image in the image cluster for each image cluster, and can also output the index code of the image cluster. In addition, it can also output the descriptive text of each image in the image cluster, which can be one descriptive text per image, or it can be output in the form of a title and a separate description.

[0099] It should be understood that the above-mentioned image segmentation processing steps (generating character mask images) and labeling processing steps (generating index codes) are not sequential. The segmentation processing steps can be performed first, followed by the labeling processing steps, or the labeling processing steps can be performed first, followed by the segmentation processing steps, or the labeling processing and segmentation processing steps can be performed simultaneously. This application does not restrict the order of the above steps.

[0100] The image processing scheme described in this application enables the construction of a multi-role consistency dataset, which maintains consistency between frames, ensures natural and flexible character poses, and achieves clear separation between foreground and background, thus facilitating the generation of continuous story images.

[0101] The dataset formed by the processing method described above in this application can be applied to different use cases, such as visual reasoning, image editing, video games, animation production, and computer-aided design, to provide data support.

[0102] In a preferred embodiment of this application, the constructed multi-role consistency dataset is intended to generate consistent character images across different backgrounds, and can be applied, for example, to generate continuous story images.

[0103] For example, a training dataset for training a text image generation model can be generated using the constructed multi-role consistency dataset. For instance, the training dataset includes multiple subsets, each subset including all images under an image cluster, the descriptive text corresponding to each image, and the index code of the image cluster. Then, based on the multiple subsets, the initial diffusion model is trained to obtain the text image generation model.

[0104] In a preferred example, the initial diffusion model can be trained using each subset by determining a first training text and multiple second training texts within the subset, and training the initial diffusion model based on all images within the subset, the first training text, and the multiple second training texts. Optionally, each subset may also include a character mask image corresponding to each image, so that model training is performed based on the image, the corresponding text, and the corresponding character mask image.

[0105] Here, the first training text may include the same descriptive portion extracted from the descriptive text corresponding to each image based on the index code corresponding to the subset. For example, the first training text may refer to the title of the aforementioned image cluster. Exemplarily, the first training text may include character portions used to describe the character's clothing features. Optionally, the character's clothing features may include, but are not limited to, at least one of the following: the character's top, bottom, shoes, hat, gloves, scarf, necklace, and other accessories.

[0106] Each second training text may include character descriptions from the image's corresponding descriptive text, excluding the first training text. For example, each second training text may refer to a separate description corresponding to each image within the aforementioned image cluster. For instance, the second training text is used to describe the character's pose features and / or background features. Optionally, the character's pose features may include, but are not limited to, at least one of the following: standing, squatting, lying flat, lying on one's side, sitting, bending over; and the character's background features may include, but are not limited to, at least one of the following: library, shopping mall, park, grassland, jungle, bedroom, swimming pool.

[0107] In a preferred embodiment, face sample images can be introduced on the basis of the above training dataset. The face sample images, the first training text and the second training text under the subset are used as conditional information, and a text image generation model is trained based on all images under the subset.

[0108] For example, the training process may include: inputting all images in the subset into the initial diffusion model, using the first training text, the second training text, and the face sample images as conditional information, and using the conditional information to guide the training of the initial diffusion model so that during the training process, the trained initial diffusion model can output a set of predicted images that conforms to the conditional information, that is, it can make the faces corresponding to the characters in the output predicted image set consistent between frames, the characters' clothing consistent with each of the first training texts, and the characters' poses and / or backgrounds consistent with each of the second training texts.

[0109] The text image generation model trained on the aforementioned multi-role consistency dataset not only maintains frame consistency across multiple consecutive story images (e.g., consistency in the faces corresponding to characters and / or consistency in character clothing) when generating consecutive story images, but also ensures natural and flexible character poses and diverse backgrounds within each consecutive story image. Compared to text image generation models trained on existing open-source datasets in related technologies, this model effectively improves the matching accuracy between generated images and images described in the text.

[0110] Based on the same application concept, this application also provides an image processing device corresponding to the method provided in the above embodiments. Since the principle of the device in this application to solve the problem is similar to the image processing method in the above embodiments of this application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0111] Figure 7 This is a schematic diagram of the structure of an image processing apparatus provided for an exemplary embodiment of this application. (See attached diagram.) Figure 7 As shown, the image processing apparatus 200 includes:

[0112] The acquisition module 210 acquires multiple images related to at least one image keyword;

[0113] The segmentation module 220 segments the multiple images according to the at least one image keyword to obtain image clusters under different image keywords;

[0114] The recognition module 230 performs the following processing for each image cluster: generating descriptive text for each image in the image cluster, and obtaining a role mask image for each image and an index code for the image cluster based on each image and / or descriptive text. The index code includes at least one position character, and each position character is used to indicate the position of characters in the descriptive text that have the same descriptive part for the same object.

[0115] In one possible implementation of this application, the descriptive text includes descriptive text for different objects, including characters in the image, background, and the characters' actions.

[0116] In one possible implementation of this application, the identification module 230 obtains the role mask image of each image in each image cluster by means of a segmentation model, based on each image in the image cluster and the corresponding descriptive text.

[0117] In one possible implementation of this application, the recognition module 230 generates descriptive text for each image in each image cluster by recognizing all images in the image cluster through a text generation model to generate descriptive text corresponding to each image, wherein character alignment is performed for the descriptive text corresponding to each image in the text generation model.

[0118] In one possible implementation of this application, the recognition module 230 generates descriptive text corresponding to each image in each image cluster by outputting a title for each image cluster and a separate description for each image, wherein the title includes the same character description extracted from each descriptive text, and the separate description includes the character description in the image's descriptive text other than the title.

[0119] In one possible implementation of this application, the recognition module 230 performs at least one of the following configuration processes for the text generation model: describing each image with the same sentence structure; using the same words to describe the same character; using the same words to describe the same modifier of the same character; describing the background with the same grammatical structure; and describing the action with the same grammatical structure.

[0120] In one possible implementation of this application, the recognition module 230 generates an index code for each image cluster by: extracting the same descriptive portion from the descriptive text corresponding to all images in the image cluster; determining the character position of a specified character in the extracted same descriptive portion in the descriptive text; and generating an index code based on the determined character position.

[0121] In one possible implementation of this application, each image includes multiple objects, and the recognition module 230 extracts the same descriptive parts in the following way: for each object, extract the same descriptive parts for that object from all descriptive texts of the image cluster, wherein the index code includes a positional character corresponding to each object respectively.

[0122] In one possible implementation of this application, the segmentation module 220 filters each image cluster by: calculating a first alignment index and / or a second alignment index corresponding to each image in the image cluster, wherein the first alignment index is used to characterize the degree of matching between the image and the reference text, and the second alignment index is used to characterize the degree of matching between the image and the reference image, wherein the reference text and the reference image are determined according to the image keywords corresponding to the image cluster; and filtering is performed according to the first alignment index and / or the second alignment index corresponding to each image in the image cluster.

[0123] In one possible implementation of this application, the acquisition module 210 is further configured to: acquire multiple preliminary images in batches based on the at least one image keyword; and select multiple images from the multiple preliminary images based on the aesthetic score of each preliminary image.

[0124] In one possible implementation of this application, the acquisition module 210 determines the aesthetic score of each preliminary image by: inputting the preliminary image into the CLIP model to extract the image features of the preliminary image; and inputting the extracted image features into a multilayer perceptron to obtain the aesthetic score of the preliminary image.

[0125] In one possible implementation of this application, a training module is further included, configured to: generate a training dataset, the training dataset comprising multiple subsets, each subset comprising all images under an image cluster, descriptive text corresponding to each image, and an index code of the image cluster; and train an initial diffusion model based on the multiple subsets to obtain a text image generation model.

[0126] In one possible implementation of this application, the training module trains the initial diffusion model using each subset in the following manner: determining a first training text and multiple second training texts under the subset, wherein the first training text includes the same descriptive portion extracted from the descriptive text corresponding to each image based on the index code corresponding to the subset, and each second training text includes character descriptions in the descriptive text corresponding to the image other than the first training text; and training the initial diffusion model based on all images under the subset, the first training text, and the multiple second training texts.

[0127] Based on the above-mentioned device, a multi-role consistency dataset can be constructed quickly and effectively.

[0128] Please see Figure 8 , Figure 8 A schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this application. For example... Figure 8 As shown, the electronic device 300 includes a processor 310, a memory 320, and a bus 330.

[0129] The memory 320 stores machine-readable instructions executable by the processor 310. When the electronic device 300 is running, the processor 310 communicates with the memory 320 via the bus 330. When the machine-readable instructions are executed by the processor 310, the steps of the image processing method in any of the above embodiments can be performed, as follows:

[0130] Obtain multiple images associated with at least one image keyword; segment the multiple images according to the at least one image keyword to obtain image clusters under different image keywords; for each image cluster, perform the following processing: generate descriptive text for each image in the image cluster, and obtain a role mask image for each image and an index code for the image cluster based on each image and / or the descriptive text, wherein the index code includes at least one position character, each position character being used to indicate the position of a character in the descriptive text that has the same descriptive part for the same object.

[0131] Based on the aforementioned electronic devices, multi-role consistency datasets can be constructed quickly and effectively.

[0132] This application also provides a computer-readable storage medium storing a computer program. When the computer program is run by a processor, it can execute the steps of the image processing method as described in any of the above embodiments, as follows:

[0133] Obtain multiple images associated with at least one image keyword; segment the multiple images according to the at least one image keyword to obtain image clusters under different image keywords; for each image cluster, perform the following processing: generate descriptive text for each image in the image cluster, and obtain a role mask image for each image and an index code for the image cluster based on each image and / or the descriptive text, wherein the index code includes at least one position character, each position character being used to indicate the position of a character in the descriptive text that has the same descriptive part for the same object.

[0134] Based on the aforementioned computer-readable storage media, multi-role consistency datasets can be constructed quickly and efficiently.

[0135] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0136] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0137] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0138] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0139] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An image processing method, characterized in that, The method includes: Retrieve multiple images associated with at least one image keyword; The multiple images are segmented according to at least one image keyword to obtain image clusters under different image keywords; For each image cluster, the following processing is performed: generating descriptive text for each image in the image cluster, and obtaining a character mask image for each image and an index code for the image cluster based on each image and / or descriptive text, wherein the index code includes at least one position character, each position character being used to indicate the position of a character in the descriptive text that has the same descriptive part for the same object.

2. The method according to claim 1, characterized in that, The descriptive text includes descriptive text for different objects, including characters in the image, the background, and the characters' actions.

3. The method according to claim 2, characterized in that, The role mask image for each image in each image cluster is obtained using the following method: By using a segmentation model, a role mask image for each image is obtained based on each image in the image cluster and the corresponding descriptive text.

4. The method according to claim 2, characterized in that, The descriptive text for each image in each image cluster is generated as follows: A text generation model is used to identify all images in an image cluster to generate descriptive text for each image, wherein character alignment is performed on the descriptive text for each image in the text generation model.

5. The method according to claim 4, characterized in that, In the text generation model, descriptive text for each image in each image cluster is generated in the following way: For each image cluster, output a title and a separate description for each image. The title includes the same character description extracted from each description text, and the separate description includes the character description from the image's description text other than the title.

6. The method according to claim 4, characterized in that, Perform at least one of the following configuration processes for the text generation model: Describe each image using the same sentence structure; Using the same words to describe the same character; The same word is used to describe the same modifier for the same character; The background is described using the same grammatical structure; Actions are described using the same grammatical structure.

7. The method according to claim 2, characterized in that, The index code for each image cluster is generated in the following way: Extract the common description parts from the description text corresponding to all images in the image cluster; Determine the character position in the description text of a specified character in the extracted identical description portion; Generate an index code based on the determined character position.

8. The method according to claim 7, characterized in that, Each image contains multiple objects. The steps for extracting parts with the same description include: For each object, extract the common descriptive portion for that object from all descriptive texts of the image cluster, where the index code includes a positional character corresponding to each object.

9. The method according to claim 1, characterized in that, This also includes filtering each image cluster in the following ways: Calculate a first alignment index and / or a second alignment index for each image in the image cluster, wherein the first alignment index is used to characterize the degree of matching between the image and the reference text, and the second alignment index is used to characterize the degree of matching between the image and the reference image, wherein the reference text and the reference image are determined based on the image keywords corresponding to the image cluster; The images are filtered according to the first alignment index and / or the second alignment index corresponding to each image in the image cluster.

10. The method according to claim 1, characterized in that, The steps to obtain multiple images associated with at least one image keyword include: Based on at least one image keyword, multiple initially selected images are acquired in batches; Based on the aesthetic score of each initial image, multiple images are selected from the multiple initial images.

11. The method according to claim 10, characterized in that, The aesthetic score for each initial image was determined using the following method: The initial selected images are input into the CLIP model to extract their image features; The extracted image features are input into a multilayer perceptron to obtain the aesthetic score of the initial image.

12. The method according to claim 1, characterized in that, Also includes: Generate a training dataset, which includes multiple subsets. Each subset includes all images under an image cluster, the descriptive text corresponding to each image, and the index code of the image cluster. Based on the aforementioned subsets, the initial diffusion model is trained to obtain a text image generation model.

13. The method according to claim 12, characterized in that, The initial diffusion model is trained using each subset in the following way: A first training text and multiple second training texts are determined under a subset. The first training text includes the same descriptive part extracted from the descriptive text corresponding to each image based on the index code corresponding to the subset. Each second training text includes character descriptions in the descriptive text corresponding to the image other than the first training text. The initial diffusion model is trained based on all images in the subset, the first training text, and multiple second training texts.

14. An image processing apparatus, characterized in that, The device includes: The acquisition module acquires multiple images associated with at least one image keyword; The segmentation module segments the multiple images according to the at least one image keyword to obtain image clusters under different image keywords; The recognition module performs the following processing for each image cluster: generating descriptive text for each image in the image cluster, and obtaining a role mask image for each image and an index code for the image cluster based on each image and / or descriptive text. The index code includes at least one position character, each position character being used to indicate the position of characters in the descriptive text that have the same descriptive part for the same object.

15. An electronic device, characterized in that, include: The device includes a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is in operation, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the method as described in any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the method as described in any one of claims 1 to 13.