Image processing method and device, equipment and storage medium
By dividing the line drawing image into sub-images and using causal sparse attention maps for attention calculation, the problem of increased computational complexity during image coloring is solved, and the inference efficiency is improved and different number of reference images are adapted.
Patent Information
- Application Number
- CN202510228056.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
During the image coloring process, as the number of reference images increases, the computational complexity increases quadratically, resulting in inferential efficiency.
By dividing the line drawing image into sub-images and obtaining the matching reference images separately, a dimensionality reduction transformation is performed to obtain a feature set. Based on the location encoding information and the noise latent space, a colored line draft image is generated, and attention calculation is performed using causal sparse attention maps to reduce the complexity of global attention calculations.
The inference efficiency of the coloring process of line drawing images is significantly improved, the complexity of attention calculation is reduced, and the adaptability to different number of reference images is enhanced.
Smart Images

Figure CN120147457A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and particularly to an image processing method, apparatus, device, and storage medium. Background Art
[0002] In recent years, the development and breakthrough of artificial intelligence technology have promoted the rapid development of the field of image processing. Artificial intelligence technology performs excellently in image generation and coloring tasks.
[0003] In the related art, the image coloring task is completed through a full attention mechanism. In the image coloring task, a line drawing image and multiple reference images are input into the full attention mechanism, and the full attention mechanism colors the line drawing image according to the multiple reference images. The full attention mechanism realizes context matching by calculating the pairwise relationships between all reference images, which helps to fuse the features of the context during the process of coloring the line drawing image.
[0004] In the related art, as the number of reference images increases, during the coloring process, the computational complexity increases quadratically, resulting in low inference efficiency. Summary of the Invention
[0005] The embodiments of the present application provide an image processing method, apparatus, device, and storage medium. The technical solutions provided by the embodiments of the present application are as follows:
[0006] According to one aspect of the embodiments of the present application, an image processing method is provided, and the method includes:
[0007] Dividing an N number of sub-line drawing images from a line drawing image, where N is an integer greater than 1;
[0008] Respectively obtaining reference images matching the N sub-line drawing images to obtain N reference sub-sets, and each reference sub-set corresponding to a sub-line drawing image includes at least one reference image matching the sub-line drawing image;
[0009] Respectively performing dimensionality reduction transformation on each of the reference images in the N reference sub-sets to obtain N reference feature sets, where the reference feature set includes the reference latent space of at least one reference image, and the reference latent space of the reference image is used to indicate the feature information of the reference image;
[0010] Based on the position encoding information, the noise latent space, and the N reference feature sets, a position-encoded noise latent space and at least one position-encoded reference latent space are obtained. The position encoding information includes N local position information and one central region information. Each local position information is used to indicate the position relationship between elements in any reference latent space in a reference feature set, and the central region information is used to indicate the position relationship between elements in the noise latent space. The noise latent space includes at least one real noise;
[0011] Perform a dimensionality reduction transformation on the line drawing image to obtain the line drawing latent space of the line drawing image, and the line drawing latent space is used to indicate the feature information of the line drawing image;
[0012] Based on the line drawing latent space, the position-encoded noise latent space, the causal sparse attention map, and the at least one position-encoded reference latent space, a colored line drawing image is generated. The causal sparse attention map is used to indicate the attention calculation rule between the line drawing latent space and the position-encoded reference latent space.
[0013] According to one aspect of the embodiments of the present application, a method for training an image processing model is provided. The method includes:
[0014] Obtain the training samples of the image processing model. The training samples include: a sample line drawing image, a reference coloring image, and N sample line drawing sub-images divided from the sample line drawing image. The reference coloring image is the colored image corresponding to the sample line drawing image;
[0015] Through the image processing model, obtain sample reference images matching the N sample line drawing sub-images respectively, and obtain N sample reference sub-sets. Each sample reference sub-set corresponding to a sample line drawing sub-image includes at least one sample reference image matching the sample line drawing sub-image;
[0016] Through the image processing model, perform a dimensionality reduction transformation on each sample reference image in the N sample reference sub-sets respectively to obtain N sample reference feature sets. The sample reference feature set includes the reference latent space of at least one sample reference image, and the reference latent space of the sample reference image is used to indicate the feature information of the sample reference image;
[0017] Through the image processing model, perform a dimensionality reduction transformation on the reference coloring image to obtain the coloring latent space of the reference coloring image, and the coloring latent space is used to indicate the feature information of the reference coloring image;
[0018] Performing T denoising operations on the colored latent space through the image processing model to obtain the noise latent space of the reference colored image, where the noise latent space includes at least one real noise, and T is an integer greater than or equal to 1;
[0019] Performing a dimensionality reduction transformation on the sample line drawing image through the image processing model to obtain the line drawing latent space of the sample line drawing image, where the line drawing latent space is used to indicate the feature information of the sample line drawing image;
[0020] Through the image processing model, according to the position encoding information, the noise latent space, and the N reference feature sets, obtaining the position-encoded noise latent space and at least one position-encoded reference latent space, where the position encoding information includes N local position information and one central region information, each local position information is used to indicate the position relationship between each element in any reference latent space in a sample reference feature set, and the central region information is used to indicate the position relationship between each element in the noise latent space;
[0021] Based on the line drawing latent space, the position-encoded noise latent space, the causal sparse attention map, and the at least one position-encoded reference latent space, adjusting the parameters of the image processing model through the image processing model to obtain the trained image processing model, where the causal sparse attention map is used to indicate the attention calculation rule between the line drawing latent space and the position-encoded reference latent space.
[0022] According to one aspect of the embodiments of the present application, an image processing device is provided, and the device includes:
[0023] A division module for dividing N line drawing sub-images from a line drawing image, where N is an integer greater than 1;
[0024] An acquisition module for respectively acquiring reference images matching the N line drawing sub-images to obtain N reference sub-sets, where each reference sub-set corresponding to a line drawing sub-image includes at least one reference image matching the line drawing sub-image;
[0025] A first transformation module for respectively performing a dimensionality reduction transformation on each reference image in the N reference sub-sets to obtain N reference feature sets, where the reference feature set includes the reference latent space of at least one reference image, and the reference latent space of the reference image is used to indicate the feature information of the reference image;
[0026] An encoding module, configured to obtain a position-encoded noise latent space and at least one position-encoded reference latent space according to position encoding information, a noise latent space, and the N reference feature sets. The position encoding information includes N local position information and one central region information. Each local position information is used to indicate the positional relationship between elements in any reference latent space in a reference feature set, and the central region information is used to indicate the positional relationship between elements in the noise latent space. The noise latent space includes at least one real noise;
[0027] A second transformation module, configured to perform a dimensionality reduction transformation on the line drawing image to obtain a line drawing latent space of the line drawing image, where the line drawing latent space is used to indicate the feature information of the line drawing image;
[0028] A generation module, configured to generate a colored line drawing image based on the line drawing latent space, the position-encoded noise latent space, a causal sparse attention map, and the at least one position-encoded reference latent space. The causal sparse attention map is used to indicate the attention calculation rule between the line drawing latent space and the position-encoded reference latent space.
[0029] According to one aspect of the embodiments of the present application, there is provided a training device for an image processing model. The device includes:
[0030] A first acquisition module, configured to acquire training samples of the image processing model. The training samples include: a sample line drawing image, a reference coloring image, and N sample line drawing sub-images divided from the sample line drawing image. The reference coloring image is the colored image corresponding to the sample line drawing image;
[0031] A second acquisition module, configured to respectively obtain sample reference images matching the N sample line drawing sub-images through the image processing model to obtain N sample reference sub-sets. Each sample reference sub-set corresponding to a sample line drawing sub-image includes at least one sample reference image matching the sample line drawing sub-image;
[0032] A first transformation module, configured to respectively perform a dimensionality reduction transformation on each sample reference image in the N sample reference sub-sets through the image processing model to obtain N sample reference feature sets. The sample reference feature sets include the reference latent spaces of at least one sample reference image, and the reference latent space of the sample reference image is used to indicate the feature information of the sample reference image;
[0033] A second transformation module, configured to perform a dimensionality reduction transformation on the reference coloring image through the image processing model to obtain a coloring latent space of the reference coloring image, where the coloring latent space is used to indicate the feature information of the reference coloring image;
[0034] A noise addition module, configured to perform T times of noise addition operations on the colored latent space through the image processing model to obtain a noise latent space of the reference colored image, where the noise latent space includes at least one real noise, and T is an integer greater than or equal to 1;
[0035] A third transformation module, configured to perform a dimensionality reduction transformation on the sample line drawing image through the image processing model to obtain a line drawing latent space of the sample line drawing image, where the line drawing latent space is used to indicate the feature information of the sample line drawing image;
[0036] An encoding module, configured to obtain a position-encoded noise latent space and at least one position-encoded reference latent space through the image processing model according to position encoding information, the noise latent space, and the N reference feature sets, where the position encoding information includes N local position information and one central region information, each local position information is used to indicate the position relationship between elements in any reference latent space in a sample reference feature set, and the central region information is used to indicate the position relationship between elements in the noise latent space;
[0037] An adjustment module, configured to adjust the parameters of the image processing model through the image processing model based on the line drawing latent space, the position-encoded noise latent space, the causal sparse attention map, and the at least one position-encoded reference latent space to obtain a trained image processing model, where the causal sparse attention map is used to indicate the attention calculation rule between the line drawing latent space and the position-encoded reference latent space.
[0038] According to one aspect of the embodiments of the present application, a computer device is provided. The computer device includes a processor and a memory. A computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above image processing method or to implement the training method of the above image processing model.
[0039] According to one aspect of the embodiments of the present application, a computer-readable storage medium is provided. A computer program is stored in the readable storage medium, and the computer program is loaded and executed by a processor to implement the above image processing method or to implement the training method of the above image processing model.
[0040] According to one aspect of the embodiments of the present application, a computer program product is provided. The computer program product includes a computer program, and the computer program is loaded and executed by a processor to implement the above image processing method or to implement the training method of the above image processing model.
[0041] The technical solutions provided by the embodiments of the present application at least include the following beneficial effects:
[0042] By performing attention calculation between the line drawing latent space of the line drawing image and the reference latent space of the reference image based on the attention calculation rule indicated by the causal sparse attention map, and generating a colored line drawing image based on the noise latent space. By performing attention calculation based on causal sparsity, only the attention calculation indicated by the causal sparse attention map needs to be carried out, avoiding global attention calculation, effectively reducing the complexity of attention calculation, and thus significantly improving the inference efficiency of the coloring process of the line drawing image. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a schematic diagram of a computer system provided by an embodiment of the present application;
[0044] Figure 2 is a flowchart of an image processing method provided by an embodiment of the present application;
[0045] Figure 3 is a schematic diagram of intercepting a sub-line drawing image provided by an embodiment of the present application;
[0046] Figure 4 is a schematic diagram of position encoding provided by an embodiment of the present application;
[0047] Figure 5 is a schematic diagram of a causal sparse attention map provided by an embodiment of the present application;
[0048] Figure 6 is a flowchart of generating a colored line drawing image provided by an embodiment of the present application;
[0049] Figure 7 is a schematic diagram of the architecture of an image processing model provided by an embodiment of the present application;
[0050] Figure 8 is a flowchart of generating a colored line drawing image provided by another embodiment of the present application;
[0051] Figure 9 is a flowchart of a training method of an image processing model provided by an embodiment of the present application;
[0052] Figure 10 is a block diagram of an image processing device provided by an embodiment of the present application;
[0053] Figure 11 is a block diagram of a training device of an image processing model provided by an embodiment of the present application;
[0054] Figure 12 is a block diagram of the structure of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.
[0056] Please refer to Figure 1 , which shows a schematic diagram of a computer system provided by an embodiment of this application. The computer system may include: a model training device 10 and a model using device 20.
[0057] The model training device 10 is an electronic device with data calculation, processing, and storage functions. The model training device 10 can be either a terminal device or a server. The model training device 10 may include, but is not limited to, electronic devices such as mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, game consoles, wearable devices, multimedia playback devices, augmented reality (AR) devices, virtual reality (VR) devices, etc. The model training device 10 is used to train an image processing model.
[0058] In this application, the image processing model can be a neural network model. For example, the above neural network model can be any one of the following: full connect neural network (FCNN), convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), etc., and can also be other neural network models, which are not limited in this application. Optionally, the model training device 10 can use machine learning to train the above image processing model so that the trained image processing model has the ability to quickly color line drawing images.
[0059] In some embodiments, such as Figure 1As shown below, the training process of the image processing model is as follows (only a brief description is provided here, and the specific training process is described below): Obtain training samples, including a sample line drawing image 301, a reference coloring image 302, and 4 sample line drawing sub-images; respectively obtain sample reference images that match the 4 sample line drawing sub-images 303 to obtain 4 sample reference sub-sets 304; perform dimensionality reduction transformation on each sample reference image in the above 4 sample reference sub-sets to obtain 4 sample reference feature sets 305; perform dimensionality reduction transformation and noise addition operation on the reference coloring image to obtain a noise latent space 306; according to the position encoding information 307, perform position encoding on the noise latent space 306 and the reference latent spaces of each sample reference image to obtain a position-encoded noise latent space 308 and at least one encoded reference latent space 309; perform dimensionality reduction transformation on the sample line drawing image 301 to obtain a line drawing latent space 310 of the sample line drawing image 301. For each prediction of noise and denoising operation, input the line drawing latent space 310 into the second neural network; and input the position-encoded noise latent space 308, at least one encoded reference latent space 309, and the output of the second neural network into the first neural network to obtain the predicted noise. In each prediction of noise and denoising operation, calculate the loss function value based on the predicted noise and the true noise, and adjust the parameters of the first neural network according to the loss function value to obtain an adjusted image processing model. Determine the image processing model after the last adjustment as the trained image processing model. Determine the noise latent space after the last denoising as the colored line drawing image 311.
[0060] The model usage device 20 is an electronic device with data calculation, processing, and storage functions. The model usage device 20 can be either a terminal device or a server. The model usage device 20 can include, but is not limited to, electronic devices such as mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, game consoles, wearable devices, multimedia playback devices, augmented reality devices, virtual reality devices, cloud technology platforms, intelligent robots, intelligent transportation terminal systems, and driving central control systems. The model usage device 20 uses the trained image processing model to color the line drawing image.
[0061] The model training device 10 and the model usage device 20 can be two independent devices or the same device.
[0062] Exemplarily, the server mentioned above can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms, but is not limited thereto.
[0063] Please refer to Figure 2 , which shows a flowchart of an image processing method provided by an embodiment of the present application. The execution subject of each step of this method can be a computer device. For example, this computer device can be Figure 1 the model usage device 20 in the computer system shown. This method can include at least one of the following steps (210-260):
[0064] Step 210, divide N line sub-images from the line drawing image, where N is an integer greater than 1.
[0065] A line drawing image is an image that only uses lines to outline the object's contour and structure. In some embodiments, the line drawing image can be stored in a bitmap format, a vector graphic format, or other image formats, and the embodiments of the present application do not limit this. An image in bitmap format is an image composed of pixels. In some embodiments, the bitmap format includes at least one of the following: PNG (Portable Network Graphics), PSD (Photoshop Document), JPEG (Joint Photographic Experts Group), TIFF (Tag Image File Format), etc., and can also include other formats, and the embodiments of the present application do not limit this. An image in vector graphic format is an image that describes lines with mathematical formulas. In some embodiments, the vector graphic format includes at least one of the following: SVG (Scalable Vector Graphics), DXF (Drawing Exchange Format), PDF (Portable Document Format), etc., and can also include other vector formats, and the embodiments of the present application do not limit this.
[0066] In some embodiments, the line drawing image can be derived from any of the following scenarios: comic drawing, animation coloring, game production, product design, illustration design, concept map for cultural relic restoration, black and white movies, etc., or from other scenarios, which are not limited in the embodiments of this application.
[0067] The sub-line drawing image is a part of the line drawing image. In some embodiments, the above N sub-line drawing images satisfy at least one of the following conditions: the contents included are not exactly the same, the sizes are the same, the areas are the same, and the shapes are the same. Other conditions may also be included, which are not limited in the embodiments of this application.
[0068] In some embodiments, obtain the line drawing image; extract N sub-images from the line drawing image as N sub-line drawing images, and there is an overlapping area between any two sub-line drawing images.
[0069] In some embodiments, obtain the line drawing image by any of the following methods: input by the user, input by the model, network download, etc. Other methods may also be available, which are not limited in the embodiments of this application.
[0070] In some embodiments, according to the extraction configuration information, extract N sub-images from the line drawing image as sub-line drawing images. The extraction configuration information is used to indicate the rules for extracting sub-images from the line drawing image. The extraction configuration information includes at least one of the following: extraction ratio, coordinate information, shape information, height information, width information, the number of sub-images N. Other information may also be included, which are not limited in the embodiments of this application.
[0071] The extraction ratio is used to indicate the ratio of the sub-image to the line drawing image. For example, the extraction ratio can be the ratio of the area of the sub-image to the area of the line drawing image. If the extraction ratio is 3 / 4, it means that the area of the sub-image accounts for 3 / 4 of the area of the line drawing image. For another example, the extraction ratio can be the ratio of the width of the sub-image to the width of the line drawing image, and the ratio of the length of the sub-image to the length of the line drawing image. If the extraction ratio is (0.2, 0.2), it means that the ratio of the width of the sub-image to the width of the line drawing image is 0.2, and the ratio of the length of the sub-image to the length of the line drawing image is 0.2.
[0072] The coordinate information is used to indicate the position of the key points of the sub-image on the line drawing image. In some embodiments, the key points of the sub-image can include at least one of the following: upper left vertex, lower left vertex, upper right vertex, lower right vertex, center point, etc. Other key points may also be included, which can be set according to the shape of the sub-image, and are not limited in the embodiments of this application.
[0073] The shape information is used to indicate the shape of the sub-image. In some embodiments, the shape of the sub-image can include at least one of the following: rectangle, square, circle, ellipse, regular polygon, etc., or it can be a free shape, which is not limited in the embodiments of this application.
[0074] The height information is used to indicate the height of the sub-image. For example, the height information indicates the pixel values occupied by the sub-image in terms of height. The width information is used to indicate the width of the sub-image. For example, the width information is used to indicate the pixel values occupied by the sub-image in terms of width.
[0075] In some embodiments, the cropping configuration information includes: a cropping ratio, coordinate information, shape information, and the number of sub-images N; according to the cropping configuration information, N sub-images are cropped from the line drawing image as N line drawing sub-images.
[0076] Exemplarily, assuming that the lower left vertex of the line drawing image is the origin of the coordinate axis, the size of the line drawing image is 256×256, the cropping ratio is 3 / 4, the coordinate information is used to indicate 4 coordinates, namely (0,0), (0,256), (256,0), and (256,256), the shape information indicates that the sub-image is a square, and the number of sub-images N is 4. Taking any one of the above coordinates as a vertex in the sub-image, 4 squares with a size of 192×192 are cropped from the line drawing image as 4 line drawing sub-images.
[0077] Based on the previous example, please refer to Figure 3 , which shows a schematic diagram of cropping line drawing sub-images provided by an embodiment of the present application. Based on the vertex (0,0) of the line drawing image, the cropped line drawing sub-image 302-1; based on the vertex (0,256) of the line drawing image, the cropped line drawing sub-image 302-2; based on the vertex (256,0) of the line drawing image, the cropped line drawing sub-image 302-3; based on the vertex (256,256) of the line drawing image, the cropped line drawing sub-image 302-4.
[0078] In the above manner, there are overlapping regions between the line drawing sub-images cropped from the line drawing image, which helps to perceive the context of each line drawing sub-image and improve the accuracy and consistency of color matching in the subsequent coloring process.
[0079] Step 220, respectively obtain reference images that match the N line drawing sub-images to obtain N reference sub-sets. Each reference sub-set corresponding to a line drawing sub-image includes at least one reference image that matches the line drawing sub-image.
[0080] The reference images are used to provide information such as color, texture, or style during the coloring process of the line drawing image. Matching the line drawing sub-image means conforming to the content of the line drawing sub-image. In some embodiments, any one of the above N line drawing sub-images corresponds to a reference sub-set.
[0081] In some embodiments, for any one of the above-mentioned N reference subsets, the reference images in the reference subset satisfy at least one of the following conditions: local semantic consistency with the corresponding line draft sub-image, color distribution matching, structural alignment, etc. There may also be other conditions, which are not limited in the embodiments of the present application. Local semantic consistency with the corresponding line draft sub-image means that the reference image and the corresponding line draft sub-image contain the same type of semantic objects. For example, if the line draft sub-image includes a character's face, the corresponding reference image should also include a character's face. Another example is that if the line draft sub-image includes clothing folds, the reference image should include clothing with similar textures. Color distribution matching means that the color distribution of the reference image is adapted to the gray-scale distribution of the line draft sub-image. Structural alignment means that the edge structure of the reference image is consistent with the edge structure of the line draft sub-image.
[0082] In some embodiments, the sources of the reference images include at least one of the following: user input, collection by relevant technicians, the Internet, public data sets, etc. There may also be other sources, which are not limited in the embodiments of the present application.
[0083] In some embodiments, the sizes of any two reference images are equal.
[0084] In some embodiments, in the reference image set, reference images that respectively match the N line draft sub-images are obtained to obtain N reference subsets, and the reference image set includes at least one reference image.
[0085] The reference image set is a set of pre-set reference images. In some embodiments, the reference image set is also referred to as a reference image pool. In some embodiments, each reference image in the reference image set satisfies at least one of the following conditions: belonging to the same style, coming from the same work, coming from the same author, containing the same semantic objects, etc. There may also be other conditions, which are not limited in the embodiments of the present application.
[0086] In some embodiments, the reference image set includes at least one of the following: original reference images, reference sub-images, global reference features, local reference features, label information, etc., and also includes other contents, which are not limited in the embodiments of the present application. The original reference image refers to a complete reference image. The reference sub-image is a sub-image cropped from the original reference image. The global reference feature is an image feature extracted from the original reference image. The local reference feature is an image feature extracted from the reference sub-image. The label information is information describing the original reference image or the reference sub-image. In some embodiments, the label information includes at least one of the following: category label, style label, color system label, light environment label, etc. There may also be other information, which are not limited in the embodiments of the present application.
[0087] In some embodiments, for any one of the N line draft sub-images, image features of the line draft sub-image are extracted; similarities between the image features of the line draft sub-image and the image features of each reference image in the reference image set are calculated respectively; the k reference images with the highest similarities in the reference image set are used as the corresponding reference subset of the line draft sub-image, where k is a positive integer.
[0088] In some embodiments, the image features of the line draft sub-image are extracted by a feature extraction model, and the feature extraction model is used to extract the image features of an image. In some embodiments, the feature extraction model is a pre-trained neural network model. Optionally, the feature extraction model can be any one of the following: fully-connected neural network, convolutional neural network, recurrent neural network, generative adversarial network, etc., and can also include other neural network models, which are not limited in the embodiments of the present application.
[0089] The similarity between the image features of the line draft sub-image and the image features of the reference image is used to measure whether the line draft sub-image and the reference image match. In some embodiments, the similarity between the image features of the line draft sub-image and the image features of the reference image can be the cosine similarity between the image features of the line draft sub-image and the image features of the reference image.
[0090] In some embodiments, the comprehensive similarities between the line draft sub-image and each reference image in the reference image set are calculated respectively, and the comprehensive similarity is calculated based on a second preset weight and different similarity metrics. The second preset weight is used to indicate the importance of different similarity metrics. In some embodiments, the second preset weight is preset by relevant technicians, which is not limited in the embodiments of the present application. In some embodiments, the similarity metrics can include at least one of the following: semantic similarity, color similarity, structural similarity, style similarity, etc., and can also include other similarity metrics, which are not limited in the embodiments of the present application. The semantic similarity is used to measure the matching degree of the sub-features of the line draft sub-image and the reference image. For example, the semantic similarity is the cosine similarity between the image features of the line draft sub-image and the image features of the reference image described above. The color similarity is used to measure the matching degree of the line draft sub-image and the reference image in terms of color distribution and lightness relationship. The structural similarity is used to measure the matching degree of the line draft sub-image and the reference image in terms of edges, shapes, textures, etc. The style similarity is used to measure the matching degree of the line draft sub-image and the reference image in terms of style.
[0091] Through the above method, the most matching reference image is effectively determined for each line draft sub-image according to the similarity between the line draft sub-image and the reference image, ensuring that the colored line draft image is more natural.
[0092] Step 230: Perform dimensionality reduction transformation on each reference image in the N reference subsets respectively to obtain N reference feature sets. The reference feature set includes the reference latent space of at least one reference image, and the reference latent space of the reference image is used to indicate the feature information of the reference image.
[0093] The latent space refers to the low-dimensional continuous representation space after the image is compressed or abstracted, and it is a set of latent and non-directly observable variables. The latent space includes the key features and semantic information of the image. The dimensionality reduction transformation refers to the process of mapping the high-dimensional image pixel space to the low-dimensional latent space (low-dimensional continuous representation space). In some embodiments, the dimensionality reduction transformation is performed on each reference image in the N reference subsets through a Variational Autoencoder (VAE) to obtain N reference feature sets. The reference latent space of the reference image is the latent space obtained after the dimensionality reduction transformation of the reference image. The reference latent spaces in each reference feature set correspond to the same line draft sub-image.
[0094] Step 240: According to the position encoding information, the noise latent space, and the N reference feature sets, obtain the position-encoded noise latent space and at least one position-encoded reference latent space. The position encoding information includes N local position information and one central region information. Each local position information is used to indicate the positional relationship between the elements in any reference latent space in a reference feature set, and the central region information is used to indicate the positional relationship between the elements in the noise latent space. The noise latent space includes at least one real noise.
[0095] The position encoding information is used to indicate the position encoding method of the noise latent space and each reference latent space in the N reference feature sets. The local position information is used to indicate the position encoding method of the reference latent spaces of each reference image in a reference feature set. The central region information is used to indicate the position encoding method of the noise latent space.
[0096] The noise latent space is a latent space filled with random real noise. Real noise, also known as Gaussian noise, is random noise that follows a Gaussian distribution. In some embodiments, the noise latent space can be directly generated or generated by adding real noise to a specified image. For example, the noise latent space can be directly generated by a first function according to the preset size of the noise latent space. The first function can be a custom function or an existing function, and the embodiments of the present application do not limit this. For example, at least one real noise can be randomly generated first, and then these noises can be added to a specified image. The specified image can be a line draft image, a reference image, or any image, and the embodiments of the present application do not limit this. In some embodiments, the size of the noise latent space is the same as the size of the reference latent space.
[0097] In some embodiments, based on the central region information, position encoding is performed on the noise latent space to obtain the noise latent space after position encoding; for any one of the N reference latent spaces in the N reference feature sets, based on the local position information corresponding to the reference latent space, position encoding is performed on the reference latent space to obtain the reference latent space after position encoding.
[0098] In some embodiments, the manner of performing position encoding based on the position encoding information is any one of the following: direct addition, concatenation, multiplicative scaling, etc., and other position encoding manners may also be included, which are not limited in the embodiments of the present application.
[0099] For any one of the above N line draft sub-images, the reference latent spaces of all the reference images in the reference feature set corresponding to the line draft sub-image are position-encoded using the same local position information. The noise latent space is position-encoded using the central position information. By performing position encoding on the reference latent spaces of each reference image, position information is explicitly injected, enabling the image processing model to distinguish features at different positions.
[0100] Integrating multiple reference image references into an image processing model is challenging because the two-dimensional position encoding in large-scale image processing models limits the ability to handle extreme aspect ratios and high resolutions when splicing references. Simply replacing two-dimensional position encoding with three-dimensional position encoding will cause the image processing model to be unstable due to changing the input domain. To overcome these challenges, local reusable position encoding is proposed in the embodiments of the present application, which can integrate any number of reference images without changing the existing two-dimensional encoding. This method reuses local region encoding, that is, performs position encoding using the same local position information, maintaining an appropriate aspect ratio and resolution. As shown in FIG. 4, which shows a schematic diagram of the position encoding information provided by an embodiment of the present application, N is 4, and the complete position encoding information 307 is divided into 5 regions: local region A, local region B, local region C, local region D, and central region E. The 4 local regions are the regions corresponding to the 4 line draft sub-images. The central region information of the central region corresponds to the noise latent space, while the other local regions share local position information with their respective reference feature sets.
[0101] In some embodiments, before performing position encoding, based on the relative position relationship of the N + 1 regions of the position encoding information, spatial concatenation is performed on the noise latent space and the N reference feature sets to obtain a concatenated latent space. Spatial concatenation refers to concatenating two or more tensors in the spatial dimension (height or width). In some embodiments, based on the position encoding information, position encoding is performed on the concatenated latent space to obtain the concatenated latent space after position encoding; from the concatenated latent space after position encoding, the noise latent space after position encoding and at least one of the reference latent spaces after position encoding are obtained.
[0102] Exemplarily, please refer to Figure 4 , which shows a schematic diagram of the position encoding provided by an embodiment of the present application. For the reference latent spaces in any one reference feature set, they are spliced along the height direction, and the noise latent space is spliced with 4 reference feature sets in the width direction to obtain the spliced latent space 401, as Figure 4 can be seen, the noise latent space in the spliced latent space 401 corresponds to 5 regions in the position encoding information 307 with the 4 reference feature sets. The reference feature set a corresponds to the local region A, the reference feature set b corresponds to the local region B, the reference feature set c corresponds to the local region C, the reference feature set d corresponds to the local region D, and the noise latent space e corresponds to the central region E.
[0103] In the above manner, the same two-dimensional position encoding (local position information) is used for the reference latent spaces in the same reference feature set, enabling the image processing model to flexibly process any number of reference images without changing the two-dimensional position encoding of the image processing model. This not only solves the efficiency bottleneck of the prior art in processing a large number of reference images, but also improves the coloring accuracy and identity consistency, and can be well applied to industrial-level comic line drawing coloring tasks that require high efficiency and high quality.
[0104] In some embodiments, the position encoding is to add any one reference latent space in the N reference feature sets to the corresponding local position information, or add the noise latent space to the central region information.
[0105] In some embodiments, the sizes of any two reference latent spaces in the above N reference feature sets are equal; the size of any one local position information is equal to the size of the reference latent space in the corresponding reference feature set; the size of the central region information is equal to the size of the noise latent space.
[0106] In some embodiments, for any one reference latent space in any one reference feature set, each element in the local position information corresponding to the reference feature set is added to the corresponding element of the reference latent space to obtain the position-encoded reference latent space.
[0107] In some embodiments, each element in the central region information is added to the corresponding element of the noise latent space to obtain the position-encoded noise latent space.
[0108] In the above manner, directly adding the corresponding position encoding to the reference latent space and the noise latent space helps to retain the original features while enhancing the position perception without introducing additional parameters, thereby improving the efficiency of the position encoding.
[0109] Step 250: Perform a dimensionality reduction transformation on the line drawing image to obtain the line drawing latent space of the line drawing image, where the line drawing latent space is used to indicate the feature information of the line drawing image.
[0110] The line drawing latent space of the line drawing image is the latent space obtained after the dimensionality reduction transformation of the line drawing image.
[0111] Step 260: Generate a colored line drawing image based on the line drawing latent space, the noise latent space after position encoding, the causal sparse attention map, and at least one reference latent space after position encoding, where the causal sparse attention map is used to indicate the attention calculation rule between the line drawing latent space and the reference latent space after position encoding.
[0112] The causal sparse attention map (Causal Sparse map) is used to implement the causal sparse attention mechanism during the coloring process of the line drawing image. The causal sparse attention mechanism simultaneously constrains the temporal dependence relationship (causality) and limits the interaction range (sparsity). The causal sparse attention map refers to the weight matrix in the causal self-attention mechanism and is used to represent the correlation between elements in the attention sequence. In some embodiments, the attention sequence is a sequence composed of the line drawing latent space and at least one reference latent space after position encoding.
[0113] Exemplarily, assume that at least one reference latent space after position encoding is R = {r1, r2, …, rn}, and the attention sequence S composed of at least one reference latent space R after position encoding and the line drawing latent space l is S = {l, r1, r2, …, rn}, where n is an integer greater than 1.
[0114] In some embodiments, the element S in the causal sparse attention map a,b reflects the degree of attention of latent space a to latent space b, where a and b are integers greater than 0 and less than or equal to n + 1. In some embodiments, the element S in the causal sparse attention map a,b can take a value greater than 0, or 0 or infinity or negative infinity. When the value of the element S a,b is a value greater than 0, it represents the attention weight of latent space a to latent space b. When the value of the element S a,b is 0 or infinity or negative infinity, it means that latent space a does not pay attention to latent space b, that is, there is no need to perform the attention calculation of latent space a to latent space b.
[0115] In the above manner, the full attention mechanism in the prior art is changed to a sparse attention mechanism, enhancing the efficiency and effectiveness of the coloring process. This not only reduces the computational complexity but also retains the necessary color information.
[0116] In some embodiments, the line drawing latent space and at least one position-encoded reference latent space are input into an image processing model. Based on the causal sparse attention map, the image processing model performs T noise prediction and denoising operations on the position-encoded noise latent space to obtain a colored line drawing image, where T is an integer greater than or equal to 1.
[0117] The image processing model is a neural network model that implements an image processing method. By performing noise prediction and denoising operations based on color constraints and line drawing constraints, the coloring process of the line drawing image is realized. The color constraints are provided by the above at least one position-encoded reference latent space. The line drawing constraints are provided by the line drawing latent space. In some embodiments, the image processing model is a diffusion model.
[0118] Noise prediction refers to the reverse process of the diffusion model, where the diffusion model predicts the real noise added to the current noise latent space. In some embodiments, based on at least one position-encoded reference latent space and the line drawing latent space, the image processing model predicts the real noise added to the current noise latent space.
[0119] The denoising operation refers to removing the predicted noise from the current noise latent space. By repeating the noise prediction and denoising operations T times, the real noise in the noise latent space can be gradually removed to generate a colored line drawing image.
[0120] In the above manner, during the process of the image processing model performing noise prediction, by combining the feature information provided by the line drawing latent space and the reference latent space, it is ensured that the color filling of the generated colored line drawing image conforms to the line drawing structure, and the coloring process of the line drawing image is performed more precisely.
[0121] In summary, the technical solution provided by the embodiments of the present application performs attention calculation between the line drawing latent space of the line drawing image and the reference latent space of the reference image based on the attention calculation rule indicated by the causal sparse attention map, and generates a colored line drawing image based on the noise latent space. By performing attention calculation based on causal sparsity, only the attention calculation indicated by the causal sparse attention map needs to be performed, avoiding global attention calculation, effectively reducing the complexity of attention calculation, and thus significantly improving the inference efficiency of the coloring process of the line drawing image. In addition, for each reference latent space in the same reference feature set, the same local position information is reused for position encoding, enhancing the adaptability to different numbers of reference images.
[0122] The process of generating a colored line drawing image is introduced below.
[0123] In some embodiments, for each noise prediction and denoising operation, attention calculation is performed based on a preset causal sparse attention map.
[0124] Exemplarily, please refer to Figure 5 , which shows a schematic diagram of the causal sparse attention map provided by an embodiment of the present application. The attention sequence S = {l, r1, r2, …, r12}, that is, the number of elements in the attention sequence is 13. It can be seen from Figure 5 that the causal sparse attention map is the degree of attention of the Q (Query) vector to the K (Key) vector. In the causal sparse attention map shown in Figure 5 , the elements filled with slashes are numerical values greater than 0, indicating the attention weights of the corresponding Q vectors to the K vectors, and the unfilled elements are 0 or infinitesimal, indicating that the corresponding Q vectors and K vectors do not perform attention calculations.
[0125] Assume that the sequence length of the line drawing latent space is Sl, and the sequence length of each reference latent space is Sr. The computational complexity of the full attention mechanism is: O(T × (Sl 2 + 2N × Sl × Sr + N 2 × Sr 2 )).
[0126] As the number of reference images increases, the computational complexity of this full attention mechanism may become quite large. However, the way of treating all reference images as complete images in attention calculation is inefficient. Instead, the reference images should mainly provide color information for the coloring process of the line drawing images, so pairwise calculations between them can be excluded. To alleviate this inefficiency problem, the full attention mechanism is replaced with a sparse attention mechanism by excluding calculations between reference images. This adjustment reduces the computational complexity to: O(T × (Sl 2 + 2N × Sl × Sr + N × Sr 2 )).
[0127] Figure 5 The meaning of the causal sparse attention map shown in
[0128] is as follows: Calculate the self-attention of the line drawing latent space, calculate the self-attention of each reference latent space r1, r2, …, r12 respectively, calculate the cross-attention of the line drawing latent space to each reference latent space r1, r2, …, r12, and do not calculate other attentions.
[0128] For the causal sparse attention map similar to the one shown in Figure 5 , perform T times of noise prediction and denoising operations. That is, for each of the following noise prediction and denoising operations, the attention calculation rule between the line drawing latent space indicated by the causal sparse attention map and the position-encoded reference latent space is: Calculate the self-attention of the line drawing latent space, calculate the self-attention of each reference latent space respectively, calculate the cross-attention of the line drawing latent space to each reference latent space, and do not calculate other attentions.
[0129] In some embodiments, please refer toFigure 6 , which shows a flowchart of generating a colored line drawing image provided by an embodiment of the present application. The process of generating the colored line drawing image includes the following steps (610-630).
[0130] Step 610, input at least one position-encoded reference latent space into the first neural network of the image processing model respectively. The first neural network performs self-attention calculation on each position-encoded reference latent space based on the causal sparse attention map, and obtains the self-attention calculation results of each position-encoded reference latent space respectively.
[0131] The first neural network is used to perform attention calculation and output predicted noise to perform denoising operation in the noise latent space. In some embodiments, the first neural network is a U-Net. In some embodiments, the first neural network includes a self-attention mechanism and / or a cross-attention mechanism. The self-attention mechanism refers to obtaining the dependencies within a latent space to understand the internal structure of the latent space. The cross-attention mechanism refers to obtaining the correlation information between two different latent spaces to understand the association relationship between different latent spaces.
[0132] Exemplarily, please refer to Figure 7 , which shows a schematic diagram of the architecture of the image processing model provided by an embodiment of the present application. Input each position-encoded reference latent space into the first neural network 701 in sequence, and perform self-attention calculation on each position-encoded reference latent space.
[0133] Step 620, save the self-attention calculation results of each position-encoded reference latent space in the buffer of the image processing model.
[0134] In some embodiments, the self-attention result of any position-encoded reference latent space includes the intermediate calculation results of self-attention calculation. The intermediate calculation results of self-attention calculation refer to the calculation results of each attention layer in the first neural network, including the K matrix and the V matrix.
[0135] The buffer of the image processing model, also known as the KV Cache (key-value buffer), is used to store the self-attention calculation results of each position-encoded reference latent space.
[0136] Exemplarily, as Figure 7 shown, for the self-attention result of any position-encoded reference latent space, the self-attention calculation in each causal sparse attention layer is saved in the KV Cache 702.
[0137] Since the reference images are clean and pre - existing, they do not need to go through the full diffusion denoising process together with the noise latent space. To maintain the independence between the reference latent spaces, the bidirectional attention between the reference latent space and the noise latent space is modified to unidirectional causal attention. The reference latent space only needs to perform one diffusion step, that is, self - attention calculation. The KV Cache is used to store the keys and values of the self - attention calculation layer - by - layer of the reference latent space, provide color - conditional guidance for the noise latent space during the denoising process, and ensure consistent retention of color information. By replacing sparse attention with causal sparse attention, the complexity of attention calculation is reduced to: O(T×(Sl 2 +N×Sl×Sr)+N×Sr 2 ).
[0138] Step 630, input the line - drawing latent space into the image - processing model. The image - processing model performs T times of noise prediction and denoising operations on the position - encoded noise latent space based on the causal sparse attention map and the self - attention calculation results of at least one position - encoded reference latent space respectively, to obtain the colored line - drawing image.
[0139] In some embodiments, as Figure 8 shown, step 630 includes the following steps (631 - 635).
[0140] Step 631, for the i - th noise prediction and denoising operation among the T times of noise prediction and denoising operations, in the buffer of the image - processing model, obtain the self - attention calculation results of at least one position - encoded reference latent space respectively, where i is a positive integer with an initial value of 1 and less than or equal to T.
[0141] In each noise prediction and denoising process, directly obtain the self - attention calculation results of at least one position - encoded reference latent space in the buffer of the image - processing model, without repeatedly calculating the self - attention calculation results of each position - encoded reference latent space.
[0142] In this way, it only needs to perform the self - attention calculation of each position - encoded reference latent space once before starting to execute the noise prediction and denoising operations, and save the self - attention calculation of each position - encoded reference latent space in the buffer of the image - processing model. When needed, obtain it from the buffer, effectively reducing the overhead of attention calculation and effectively reducing the computational complexity.
[0143] Step 632, input the line - drawing latent space into the second neural network of the image - processing model. The second neural network performs self - attention calculation on the line - drawing latent space based on the causal sparse attention map, to obtain the self - attention calculation result of the line - drawing latent space.
[0144] The second neural network is used to perform self-attention calculation on the line art latent space and integrate the line art latent space into the main branch layer by layer, that is, into the first neural network, so as to achieve precise control of the line art. The second neural network is also called the Line Art Guider.
[0145] In some embodiments, the second neural network includes a self-attention layer. Compared with the prior art, the second neural network removes the cross-attention layer and only retains the self-attention layer. This method reduces the number of parameters of the image processing model while not weakening its control effect.
[0146] Exemplarily, as Figure 7 shown, the line art latent space is input into the second neural network 703 of the image processing model, and the line art latent space is integrated into the first neural network 701 layer by layer.
[0147] In some embodiments, before step 632, it further includes: obtaining a color hint latent space and a color hint mask. The color hint latent space is used to indicate at least one color hint point, and the color hint point is used to indicate the color generated at the corresponding position in the line art image. The color hint mask is used to indicate the position of at least one color hint point in the line art image; performing channel concatenation on the color hint latent space, the color hint mask, and the line art latent space to obtain the concatenated line art latent space, and the concatenated line art latent space is input into the second neural network to perform self-attention calculation to obtain the self-attention calculation result of the line art latent space.
[0148] In some embodiments, a color hint image is obtained. The color hint image includes at least one color hint point; a dimensionality reduction transformation is performed on the color hint image to obtain a hint color latent space. The color hint point includes multiple pixels filled with the corresponding hint color. The color hint mask is used to distinguish the area where the color hint point is located in the color hint image and the area other than the color hint point. The pixel values in the color hint mask have various representation methods, including but not limited to the following methods: binary, multi-class, probability, and floating-point numbers, etc. Exemplarily, the color hint mask is represented in binary, 1 represents the area where the color hint point is located, and 0 represents the area other than the area where the color hint point is located.
[0149] By performing channel concatenation on the color hint latent space, the color hint mask, and the line art latent space, the color hint information and the line art information are fused together, which helps the image processing model understand the line art structure and the color hint.
[0150] In the above manner, different color hints are used to adjust specific areas in the line art image to meet diverse requirements.
[0151] Step 633: Input the line drawing latent space, the self-attention calculation result of the line drawing latent space, at least one position-encoded reference latent space, and the self-attention calculation results of at least one position-encoded reference latent space into the first neural network. Based on the causal sparse attention map, the first neural network outputs a fused latent space, which is used to indicate the geometric feature information of the line drawing image and the color information of at least one reference image.
[0152] In some embodiments, based on the causal sparse attention map, the first neural network performs cross-attention calculations of the line drawing latent space on each position-encoded reference latent space to obtain at least one cross-attention calculation result. The first neural network performs feature fusion on the self-attention calculation result of the line drawing latent space, the self-attention calculation results of at least one position-encoded reference latent space, and at least one cross-attention calculation result to obtain a fused latent space.
[0153] The geometric feature information of the line drawing image is used to describe the structural information, edge relationships, and shape characteristics in the line drawing image. The color information of the reference image is used to describe the color distribution, categories, and regional affiliations in the reference image. In some embodiments, the color information includes at least one of the following: color index, color label, and color region mask, and may also include other information, which is not limited in the embodiments of the present application. The color index is used to indicate a certain color, and each color corresponds to a unique color index. The color label is used to indicate the category of the target object. The color region mask is used to indicate the distribution positions of different colors.
[0154] In the above manner, by fusing the geometric feature information of the line drawing image and the color information of at least one reference image, the image processing model can better understand the line drawing structure, color distribution, and style, etc., ensuring that the filling of the line drawing image matches the line drawing structure.
[0155] Step 634: Based on the fused latent space, the first neural network performs the i-th noise prediction and denoising operation on the position-encoded noise latent space to obtain a denoised noise latent space.
[0156] In some embodiments, based on the fused latent space, the first neural network outputs the i-th predicted noise. The i-th predicted noise is removed from the noise latent space to obtain the i-th denoised noise latent space, and the i-th denoised noise latent space is used as the noise latent space for the (i + 1)-th predicted noise.
[0157] Performing the noise prediction and denoising operation T times means repeatedly executing Steps 631 to 634 T times.
[0158] Step 635: Determine the denoised noise latent space obtained from the last noise prediction and denoising operation as the colored line drawing image.
[0159] When i is equal to T, it is determined that this noise prediction and denoising operation is the last noise prediction and denoising operation.
[0160] In some embodiments, the denoised noise latent space and the line drawing image obtained from the last noise prediction and denoising operation are input into the decoder of the image processing model, and the decoder outputs a low-resolution colored line drawing image; the low-resolution colored line drawing image is input into the upsampling component, and the upsampling component outputs a high-resolution colored line drawing image. The upsampling component is used to restore the details of the colored line drawing image. In some embodiments, the upsampling component is a Guided Super-Resolution Pipeline (GSRP for short). The Guided Super-Resolution Pipeline is used to upsample the low-resolution colored output to generate a high-resolution color image, enhancing detail restoration and improving the output quality.
[0161] Exemplarily, as Figure 7 shown, the denoised noise latent space and the line drawing image obtained from the last noise prediction and denoising operation are input into the decoder 704 of the image processing model, and then the output of the decoder 704 and the line drawing image are input into the Guided Super-Resolution Pipeline 705, and the Guided Super-Resolution Pipeline 705 outputs a high-resolution colored line drawing image.
[0162] The usage process of the image processing model has been introduced above. Now, the training process of this image processing model is introduced. The method steps of the usage process and the training process of the image processing model correspond to each other. For the details not shown in the embodiments corresponding to the training process, reference can be made to the corresponding parts of the embodiments of the usage process.
[0163] Please refer to Figure 9 , which shows a flowchart of a training method for an image processing model provided by an embodiment of the present application. The execution subject of each step of this method can be a computer device. For example, this computer device can be the model training device 10 in the computer system as Figure 1 shown. This method may include at least one of the following steps (910-980):
[0164] Step 910, obtaining training samples for the image processing model. The training samples include: sample line drawing images, reference colored images, and N sample sub-line drawing images divided from the sample line drawing images. The reference colored image is the colored image corresponding to the sample line drawing image.
[0165] The training samples are used to train the coloring ability of the image processing model. The sample line drawing images include uncolored line drawing content. The reference colored image is the colored version corresponding to the sample line drawing image. The sample sub-line drawing images are a part of the sample line drawing images.
[0166] Step 920: Obtain sample reference images respectively matching N sample line draft sub-images through an image processing model, obtaining N sample reference sub-sets. Each sample reference sub-set corresponding to a sample line draft sub-image includes at least one sample reference image matching the sample line draft sub-image.
[0167] The implementation process and manner of step 920 are similar to those of step 220 above. For relevant descriptions, please refer to the corresponding part of step 220. This application will not elaborate here.
[0168] Step 930: Perform dimensionality reduction transformation on each sample reference image in the N sample reference sub-sets respectively through the image processing model, obtaining N sample reference feature sets. The sample reference feature set includes the reference latent space of at least one sample reference image, and the reference latent space of the sample reference image is used to indicate the feature information of the sample reference image.
[0169] The implementation process and manner of step 930 are similar to those of step 230 above. For relevant descriptions, please refer to the corresponding part of step 230. This application will not elaborate here.
[0170] In some embodiments, randomly obtain M sample reference images from the above N sample reference sub-sets, where M is a positive integer. In some embodiments, the value of M is preset by relevant technical personnel, and this application embodiment does not limit this. For example, randomly select reference images from each sample reference sub-set, keeping the total number of reference images cited constant at 3, 6, or 12. This strategy enhances the adaptability of the image processing model to different combinations of the number of reference images.
[0171] Step 940: Perform dimensionality reduction transformation on the reference colored image through the image processing model, obtaining the colored latent space of the reference colored image. The colored latent space is used to indicate the feature information of the reference colored image.
[0172] The colored latent space of the reference colored image is the latent space obtained after the reference colored image undergoes dimensionality reduction transformation.
[0173] Step 950: Perform T noise addition operations on the colored latent space through the image processing model, obtaining the noise latent space of the reference colored image. The noise latent space includes at least one real noise, where T is an integer greater than or equal to 1.
[0174] In some embodiments, during any noise addition, a true noise corresponding to a noise intensity is randomly obtained first, and the true noise of the noise intensity is randomly added to the coloring latent space. The noise intensity is a numerical value used to describe the noise level or intensity, indicating the degree of perturbation of the noise to the coloring latent space. The higher the noise intensity, the greater the impact of the noise on the coloring latent space. There are many ways to measure the noise intensity, which may include but are not limited to the following ways: root mean square error, signal-to-noise ratio, standard deviation, peak signal-to-noise ratio, etc. The embodiments of the present application do not limit this.
[0175] Step 960, perform a dimensionality reduction transformation on the sample line drawing image through the image processing model to obtain the line drawing latent space of the sample line drawing image, and the line drawing latent space is used to indicate the feature information of the sample line drawing image.
[0176] The implementation process and manner of step 960 are similar to those of step 250 above. For relevant descriptions, please refer to the corresponding part of step 250. The present application will not elaborate here.
[0177] Step 970, through the image processing model, according to the position encoding information, the noise latent space, and N reference feature sets, obtain the position-encoded noise latent space and at least one position-encoded reference latent space. The position encoding information includes N local position information and one central region information. Each local position information is used to indicate the position relationship between each element in any reference latent space in a sample reference feature set, and the central region information is used to indicate the position relationship between each element in the noise latent space.
[0178] The implementation process and manner of step 970 are similar to those of step 240 above. For relevant descriptions, please refer to the corresponding part of step 240. The present application will not elaborate here.
[0179] Step 980, through the image processing model, based on the line drawing latent space, the position-encoded noise latent space, the causal sparse attention map, and at least one position-encoded reference latent space, adjust the parameters of the image processing model to obtain the trained image processing model. The causal sparse attention map is used to indicate the attention calculation rule between the line drawing latent space and the position-encoded reference latent space.
[0180] In some embodiments, step 980 includes the following steps (981 to 985).
[0181] Step 981, input the line drawing latent space and at least one position-encoded reference latent space into the first neural network of the image processing model. The first neural network predicts the noise added to the noise latent space for the i-th time based on the causal sparse attention map to obtain the predicted noise for the i-th time, where i is a positive integer with an initial value of 1 and less than or equal to T.
[0182] In some embodiments, at least one position-encoded reference latent space is respectively input into a first neural network. The first neural network performs self-attention calculation on each position-encoded reference latent space based on a causal sparse attention map to obtain the self-attention calculation results of each of the at least one position-encoded reference latent spaces; the self-attention calculation results of each of the at least one position-encoded reference latent spaces are stored in a buffer of the image processing model; the line drawing latent space is input into the image processing model, and the first neural network predicts the noise added to the noise latent space at the i-th time based on the causal sparse attention map to obtain the predicted noise at the i-th time.
[0183] In some embodiments, the step of predicting the noise includes the following steps:
[0184] (a) In the buffer of the image processing model, obtain the self-attention calculation results of each of the at least one position-encoded reference latent spaces.
[0185] (b) Input the line drawing latent space into the third neural network and the fourth neural network of the image processing model. The third neural network and the fourth neural network respectively perform self-attention calculation on the line drawing latent space based on the causal sparse attention map to obtain a first calculation result and a second calculation result of the line drawing latent space. The first calculation result is obtained by the third neural network performing self-attention calculation on the line drawing latent space, and the second calculation result is obtained by the fourth neural network performing self-attention calculation on the line drawing latent space.
[0186] The third neural network and the fourth neural network are two line drawing guides with different styles. The third neural network and the fourth neural network independently perform self-attention calculation on the line drawing latent space.
[0187] In the above manner, the line drawing style is enhanced by randomly mixing two styles of line drawing guides, thereby improving the adaptability of the line drawing guide to different line drawing styles.
[0188] (c) Based on a first preset weight, the first calculation result, and the second calculation result, obtain the self-attention calculation result of the line drawing latent space.
[0189] The first preset weight is used to control the fusion ratio of the first calculation result and the second calculation result. In some embodiments, the first preset weight is preset by those skilled in the relevant art, and the embodiments of the present application do not limit this. In some embodiments, by adjusting the first preset weight, the first calculation result and the second calculation result of the line drawing latent space are randomly mixed.
[0190] In some embodiments, based on the first preset weight, perform weighted summation on the first calculation result and the second calculation result to obtain the self-attention calculation result of the line drawing latent space.
[0191] (d) The line draft latent space, the self-attention calculation result of the line draft latent space, at least one position-encoded reference latent space, and the self-attention calculation result of each of the at least one position-encoded reference latent space are input into a first neural network, and the first neural network outputs a fused latent space based on a causal sparse attention graph, where the fused latent space is used to indicate geometric feature information of the sample line draft image and color information of at least one sample reference image.
[0192] (e) The first neural network predicts the noise added to the noise latent space for the i-th time based on the fused latent space to obtain the i-th predicted noise.
[0193] In some embodiments, a color cue latent space and a color cue mask are obtained, the color cue latent space is used to indicate at least one color cue point, the color cue point is used to indicate the color generated at a position corresponding to the sample line draft image, and the color cue mask is used to indicate the position of at least one color cue point in the sample line draft image; channel stitching is performed on the color cue latent space, the color cue mask and the line draft latent space to obtain a stitched line draft latent space, and the stitched line draft latent space is input into the third neural network and the fourth neural network to perform self-attention calculation to obtain a self-attention calculation result of the line draft latent space.
[0194] In some embodiments, for any one of the at least one color hint point, based on the color hint mask, the variance of the color value of at least one pixel corresponding to the color hint point in the sample line draft image is calculated; when the variance of the color value of at least one pixel corresponding to the color hint point in the sample line draft image is greater than a first threshold, the color hint point is determined to be invalid.
[0195] The color value of a pixel represents the color information of the pixel. In some embodiments, the color value of a pixel can be represented by any of the following methods: RGB (Red Green Blue) model, RGBA (Red Green Blue Alpha) model, and HSV (Hue Saturation Value) model, etc. The color value of a pixel can also be represented by other methods, which are not limited in the embodiments of the present application. Unless otherwise specified, in the present application, the color value of a pixel is represented by the RGB model.
[0196] Exemplarily, the first threshold is 0.01, and the variance of the color value of at least one pixel corresponding to the color hint point in the sample line draft image is limited to no more than 0.01.
[0197] Through the above method, the color cue points sampled on the edge during the training stage are excluded, thereby avoiding ambiguity during the training stage.
[0198] Step 982: Calculate the loss function value for the i-th time based on the real noise added for the i-th time and the predicted noise for the i-th time in the noise latent space.
[0199] The loss function value is used to measure the difference between the added real noise and the predicted noise.
[0200] In some embodiments, calculate the mean square error between the real noise and the predicted noise based on the real noise added for the i-th time and the predicted noise for the i-th time in the noise latent space, and use it as the loss function value for the i-th time.
[0201] It should be noted that the loss function value for the i-th time can also be calculated by other means.
[0202] Step 983: Adjust the parameters of the first neural network based on the loss function value for the i-th time to obtain an adjusted image processing model.
[0203] In some embodiments, the first neural network includes a first fine-tuning module, and the first fine-tuning module includes at least one adjustable fine-tuning parameter; adjust the fine-tuning parameters of the first fine-tuning module based on the loss function value for the i-th time to obtain an adjusted image processing model. In some embodiments, the first fine-tuning module can be implemented as a LoRA (Low-Rank Adaptation) matrix.
[0204] Exemplarily, as Figure 7 shown, at least one LoRA matrix 706 is stacked on the first neural network 701, and the fine-tuning parameters (Tunning weight) on at least one LoRA matrix 706 are adjusted based on the loss function value obtained from each noise prediction and denoising operation. During the entire training process, the fixed parameters (Frozen weight) in the first neural network 701 are not adjusted.
[0205] In some embodiments, methods such as gradient descent, random search, and grid search can be used to adjust the parameters of the image processing model according to the total loss, and this application does not make any limitations in this regard.
[0206] Step 984: When i is less than T, assign the value obtained by adding 1 to i to i, remove the predicted noise for the i-th time in the noise latent space to obtain the noise latent space after the i-th denoising, and then start executing from step 981 again.
[0207] In some embodiments, the first neural network outputs the predicted noise for the i-th time based on the fusion latent space; remove the predicted noise for the i-th time in the noise latent space to obtain the noise latent space after the i-th denoising, and the noise latent space after the i-th denoising is used as the noise latent space for the (i + 1)-th predicted noise.
[0208] Performing T noise prediction and denoising operations means looping through steps 981 - 984 T times, that is, adjusting the image processing model T times.
[0209] Step 985, when i is equal to T, stop adjusting the image processing model and determine the last adjusted image processing model as the trained image processing model.
[0210] In some implementations, when i is equal to T, determine the noise latent space after the last denoising as the colored line drawing image; calculate the comprehensive loss function value based on the reference colored image and the colored line drawing image; when the comprehensive loss function value is greater than the first threshold value, restart from step 910. The first threshold value is used to control the accuracy of the image processing model. In some embodiments, the first threshold value is preset by those skilled in the relevant art, and the embodiments of the present application do not limit this.
[0211] In summary, the technical solution provided by the embodiments of the present application performs the attention calculation between the line drawing latent space of the sample line drawing image and the reference latent spaces of at least one reference image based on the attention calculation rule indicated by the causal sparse attention map, combines the noise latent space after position encoding, and adjusts the image processing model. By performing the attention calculation based on the causal sparse attention, the complexity of the attention calculation is effectively reduced, and the efficiency and accuracy of training the image processing model are significantly improved. In addition, for each reference latent space in the same reference feature set, the same local position information is reused to perform position encoding, enhancing the adaptability of the image processing model to different numbers of reference images.
[0212] The following is an embodiment of the apparatus of the present application, which can be used to execute the method embodiment of the present application. For details not disclosed in the embodiment of the apparatus of the present application, please refer to the method embodiment of the present application.
[0213] Please refer to Figure 10 , which shows a block diagram of an image processing apparatus provided by an embodiment of the present application. This apparatus has the functions implemented in the above examples, and these functions can be implemented by hardware or by hardware executing corresponding software. This apparatus can be the model usage device 20 introduced above, or can be set in the model usage device 20. As Figure 10 shown, the apparatus 1000 may include: a division module 1010, an acquisition module 1020, a first transformation module 1030, an encoding module 1040, a second transformation module 1050, and a generation module 1060.
[0214] The division module 1010 is used to divide an N number of sub line drawing images from the line drawing image, where N is an integer greater than 1.
[0215] An acquisition module 1020 is configured to respectively acquire reference images matching the N line draft sub-images, obtaining N reference sub-sets. Each reference sub-set corresponding to a line draft sub-image includes at least one reference image matching the line draft sub-image.
[0216] A first transformation module 1030 is configured to respectively perform dimensionality reduction transformation on each of the reference images in the N reference sub-sets, obtaining N reference feature sets. The reference feature set includes the reference latent space of at least one reference image, and the reference latent space of the reference image is used to indicate the feature information of the reference image.
[0217] An encoding module 1040 is configured to obtain a position-encoded noise latent space and at least one position-encoded reference latent space according to position encoding information, a noise latent space, and the N reference feature sets. The position encoding information includes N local position information and one central region information. Each local position information is used to indicate the position relationship between elements in any reference latent space in a reference feature set, and the central region information is used to indicate the position relationship between elements in the noise latent space. The noise latent space includes at least one real noise.
[0218] A second transformation module 1050 is configured to perform dimensionality reduction transformation on the line draft image, obtaining the line draft latent space of the line draft image, and the line draft latent space is used to indicate the feature information of the line draft image.
[0219] A generation module 1060 is configured to generate a colored line draft image based on the line draft latent space, the position-encoded noise latent space, a causal sparse attention map, and the at least one position-encoded reference latent space. The causal sparse attention map is used to indicate the attention calculation rule between the line draft latent space and the position-encoded reference latent space.
[0220] In some embodiments, the generation module 1060 is configured to input the line draft latent space and the at least one position-encoded reference latent space into an image processing model. The image processing model performs T times of noise prediction and denoising operations on the position-encoded noise latent space based on the causal sparse attention map to obtain the colored line draft image, where T is an integer greater than or equal to 1.
[0221] In some embodiments, the generation module 1060 includes: a first calculation sub-module, a storage sub-module, and a denoising sub-module (not shown in Figure 10 ).
[0222] The first computing sub-module is configured to input the at least one position-encoded reference latent space into the first neural network of the image processing model respectively. Based on the causal sparse attention map, the first neural network performs self-attention calculation on each of the position-encoded reference latent spaces to obtain the self-attention calculation results of each of the at least one position-encoded reference latent spaces.
[0223] The saving sub-module is configured to save the self-attention calculation results of each of the at least one position-encoded reference latent spaces in the buffer of the image processing model.
[0224] The denoising sub-module is configured to input the line drawing latent space into the image processing model. Based on the causal sparse attention map and the self-attention calculation results of each of the at least one position-encoded reference latent spaces, the image processing model performs T times of noise prediction and denoising operations on the position-encoded noise latent space to obtain the colored line drawing image.
[0225] In some embodiments, for the i-th noise prediction and denoising operation among the T times of noise prediction and denoising operations, the denoising sub-module is configured to obtain the self-attention calculation results of each of the at least one position-encoded reference latent spaces in the buffer of the image processing model, where i is a positive integer with an initial value of 1 and less than or equal to T; input the line drawing latent space into the second neural network of the image processing model, and the second neural network performs self-attention calculation on the line drawing latent space based on the causal sparse attention map to obtain the self-attention calculation result of the line drawing latent space; input the line drawing latent space, the self-attention calculation result of the line drawing latent space, the at least one position-encoded reference latent space, and the self-attention calculation results of each of the at least one position-encoded reference latent spaces into the first neural network, and the first neural network outputs a fused latent space based on the causal sparse attention map, where the fused latent space is used to indicate the geometric feature information of the line drawing image and the color information of the at least one reference image; perform the i-th noise prediction and denoising operation on the position-encoded noise latent space based on the fused latent space through the first neural network to obtain the denoised noise latent space; and determine the denoised noise latent space obtained from the last noise prediction and denoising operation as the colored line drawing image.
[0226] In some embodiments, the generation module 1060 further includes a prompt sub-module (in Figure 10(not shown in the figure) for obtaining a color hint latent space and a color hint mask, where the color hint latent space is used to indicate at least one color hint point, the color hint point is used to indicate the color generated at the corresponding position in the line drawing image, and the color hint mask is used to indicate the position of the at least one color hint point in the line drawing image; perform channel splicing on the color hint latent space, the color hint mask, and the line drawing latent space to obtain a spliced line drawing latent space, and input the spliced line drawing latent space into the second neural network to perform the self-attention calculation to obtain the self-attention calculation result of the line drawing latent space.
[0227] In some embodiments, the encoding module 1040 is configured to perform position encoding on the noise latent space based on the central region information to obtain the position-encoded noise latent space; for any one of the N reference latent spaces in the N reference feature sets, perform position encoding on the reference latent space based on the local position information corresponding to the reference latent space to obtain the position-encoded reference latent space.
[0228] In some embodiments, the position encoding is to add any one of the N reference latent spaces in the N reference feature sets to the corresponding local position information, or to add the noise latent space to the central region information.
[0229] In some embodiments, the obtaining module 1020 is configured to obtain, in the reference image set, reference images respectively matching the N line drawing sub-images to obtain the N reference sub-sets, where the reference image set includes at least one reference image.
[0230] In some embodiments, the obtaining module 1020 is further configured to obtain, in the reference image set, reference images respectively matching the N line drawing sub-images to obtain the N reference sub-sets, including: for any one of the N line drawing sub-images, extract the image features of the line drawing sub-image; calculate the similarity between the image features of the line drawing sub-image and the image features of each reference image in the reference image set respectively; use the k reference images with the highest similarity in the reference image set as the reference sub-set corresponding to the line drawing sub-image, where k is a positive integer.
[0231] In some embodiments, the partitioning module 1010 is configured to obtain the line drawing image; extract N sub-images from the line drawing image as the N line drawing sub-images, and there is an overlapping area between any two of the line drawing sub-images.
[0232] In summary, the technical solution provided by the embodiment of the present application performs attention calculation between the line drawing latent space of the line drawing image and the reference latent space of the reference image based on the attention calculation rule indicated by the causal sparse attention map, and generates a colored line drawing image based on the noise latent space. By performing attention calculation based on causal sparse attention, only the attention calculation indicated by the causal sparse attention map needs to be performed, avoiding global attention calculation, effectively reducing the complexity of attention calculation, and thus significantly improving the inference efficiency of the coloring process of the line drawing image. In addition, for each reference latent space in the same reference feature set, the same local position information is reused to perform position encoding, enhancing the adaptability to different numbers of reference images.
[0233] Please refer to Figure 11 , which shows a block diagram of a training device for an image processing model provided by an embodiment of the present application. This device has the functions implemented in the above examples, and these functions can be implemented by hardware or by software executed by the hardware. This device can be the model training device 10 introduced above, or can be provided in the model training device 10. As Figure 11 shown, the device 1100 may include: a first acquisition module 1110, a second acquisition module 1120, a first transformation module 1130, a second transformation module 1140, a noise addition module 1150, a third transformation module 1160, an encoding module 1170, and an adjustment module 1180.
[0234] The first acquisition module 1110 is configured to acquire training samples of the image processing model, where the training samples include: sample line drawing images, reference coloring images, and N sample line drawing sub-images divided from the sample line drawing images, and the reference coloring image is the colored image corresponding to the sample line drawing image.
[0235] The second acquisition module 1120 is configured to respectively obtain sample reference images matching the N sample line drawing sub-images through the image processing model, to obtain N sample reference sub-sets, and each sample reference sub-set corresponding to a sample line drawing sub-image includes at least one sample reference image matching the sample line drawing sub-image.
[0236] The first transformation module 1130 is configured to respectively perform dimensionality reduction transformation on each sample reference image in the N sample reference sub-sets through the image processing model, to obtain N sample reference feature sets, where the sample reference feature set includes the reference latent space of at least one sample reference image, and the reference latent space of the sample reference image is used to indicate the feature information of the sample reference image.
[0237] A second transformation module 1140, configured to perform a dimensionality reduction transformation on the reference colored image through the image processing model to obtain a colored latent space of the reference colored image, where the colored latent space is used to indicate the feature information of the reference colored image.
[0238] A noise adding module 1150, configured to perform T noise adding operations on the colored latent space through the image processing model to obtain a noise latent space of the reference colored image, where the noise latent space includes at least one real noise, and T is an integer greater than or equal to 1.
[0239] A third transformation module 1160, configured to perform a dimensionality reduction transformation on the sample line drawing image through the image processing model to obtain a line drawing latent space of the sample line drawing image, where the line drawing latent space is used to indicate the feature information of the sample line drawing image.
[0240] An encoding module 1170, configured to obtain a position-encoded noise latent space and at least one position-encoded reference latent space through the image processing model according to position encoding information, the noise latent space, and the N reference feature sets, where the position encoding information includes N local position information and one central region information, each local position information is used to indicate the position relationship between each element in any one reference latent space in a sample reference feature set, and the central region information is used to indicate the position relationship between each element in the noise latent space.
[0241] An adjustment module 1180, configured to adjust the parameters of the image processing model through the image processing model based on the line drawing latent space, the position-encoded noise latent space, the causal sparse attention map, and the at least one position-encoded reference latent space to obtain a trained image processing model, where the causal sparse attention map is used to indicate the attention calculation rule between the line drawing latent space and the position-encoded reference latent space.
[0242] In some embodiments, the adjustment module 1180 includes: a prediction sub-module, a first calculation sub-module, an adjustment sub-module, a loop sub-module, and a stop sub-module (not shown in Figure 11 ).
[0243] The prediction sub-module is configured to input the line drawing latent space and the at least one position-encoded reference latent space into a first neural network of the image processing model, and the first neural network predicts the noise added to the noise latent space for the i-th time based on the causal sparse attention map to obtain the predicted noise for the i-th time, where i is a positive integer with an initial value of 1 and less than or equal to T.
[0244] The first calculation sub-module is used to calculate the loss function value of the $i$-th time based on the real noise added for the $i$-th time and the predicted noise of the $i$-th time in the noise latent space.
[0245] The adjustment sub-module is used to adjust the parameters of the first neural network based on the loss function value of the $i$-th time to obtain an adjusted image processing model.
[0246] The loop sub-module is used to, when $i$ is less than $T$, assign the value obtained by adding 1 to $i$ to $i$, remove the predicted noise of the $i$-th time in the noise latent space to obtain the denoised noise latent space of the $i$-th time, and then input the line drawing latent space and the at least one position-encoded reference latent space into the first neural network of the image processing model again. The first neural network predicts the noise added to the noise latent space for the $i$-th time based on the causal sparse attention map to obtain the predicted noise of the $i$-th time, and then starts to execute the above steps.
[0247] The stop sub-module is used to, when $i$ is equal to $T$, stop adjusting the image processing model and determine the last adjusted image processing model as the trained image processing model.
[0248] In some embodiments, the prediction sub-module is used to input the at least one position-encoded reference latent space into the first neural network respectively. The first neural network performs self-attention calculation on each of the position-encoded reference latent spaces based on the causal sparse attention map to obtain the self-attention calculation results of the at least one position-encoded reference latent spaces respectively; save the self-attention calculation results of the at least one position-encoded reference latent spaces in the buffer of the image processing model; input the line drawing latent space into the image processing model, and the first neural network predicts the noise added to the noise latent space for the $i$-th time based on the causal sparse attention map to obtain the predicted noise of the $i$-th time.
[0249] In some embodiments, the prediction sub-module is further configured to obtain, in the buffer of the image processing model, the self-attention calculation results of the at least one position-encoded reference latent space respectively; input the line drawing latent space into the third neural network and the fourth neural network of the image processing model, and respectively perform self-attention calculation on the line drawing latent space by the third neural network and the fourth neural network based on the causal sparse attention map to obtain a first calculation result and a second calculation result of the line drawing latent space, where the first calculation result is obtained by the third neural network performing self-attention calculation on the line drawing latent space, and the second calculation result is obtained by the fourth neural network performing self-attention calculation on the line drawing latent space; obtain the self-attention calculation result of the line drawing latent space based on a first preset weight, the first calculation result, and the second calculation result; input the line drawing latent space, the self-attention calculation result of the line drawing latent space, the at least one position-encoded reference latent space, and the self-attention calculation results of the at least one position-encoded reference latent space respectively into the first neural network, and output, by the first neural network based on the causal sparse attention map, a fused latent space, where the fused latent space is used to indicate the geometric feature information of the sample line drawing image and the color information of the at least one sample reference image; predict the noise added to the noise latent space at the i-th time based on the fused latent space by the first neural network to obtain the predicted noise at the i-th time.
[0250] In some embodiments, the adjustment module 1180 further includes a hint sub-module (not shown in Figure 11 ), which is configured to obtain a color hint latent space and a color hint mask, where the color hint latent space is used to indicate at least one color hint point, the color hint point is used to indicate the color generated at the corresponding position of the sample line drawing image, and the color hint mask is used to indicate the position of the at least one color hint point in the sample line drawing image; perform channel splicing on the color hint latent space, the color hint mask, and the line drawing latent space to obtain a spliced line drawing latent space, and input the spliced line drawing latent space into the third neural network and the fourth neural network to perform the self-attention calculation to obtain the self-attention calculation result of the line drawing latent space.
[0251] In some embodiments, the device 1100 further includes a judgment module (in Figure 11(not shown in the figure) is used to calculate the variance of the color values of at least one pixel corresponding to any one of the at least one color hint points in the sample line drawing image based on the color hint mask; when the variance of the color values of at least one pixel corresponding to the color hint point in the sample line drawing image is greater than a first threshold, the color hint point is determined to be invalid.
[0252] In summary, the technical solution provided by the embodiments of the present application performs attention calculation between the line drawing latent space of the sample line drawing image and the reference latent spaces of at least one reference image based on the attention calculation rule indicated by the causal sparse attention map, combines the noise latent space after position encoding, and adjusts the image processing model. By performing attention calculation based on causal sparse attention, the complexity of attention calculation is effectively reduced, and the efficiency and accuracy of training the image processing model are significantly improved. In addition, for each reference latent space in the same reference feature set, the same local position information is reused to perform position encoding, enhancing the adaptability of the image processing model to different numbers of reference images.
[0253] It should be noted that when the device provided in the above embodiment implements its functions, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0254] Please refer to Figure 12 , which shows a structural block diagram of a computer device 1200 provided by an embodiment of the present application. The computer device 1200 can be Figure 1 the model training device 10 in the shown implementation environment, or Figure 1 the model using device 20 in the shown implementation environment, and is used to implement the image processing method or the training method of the image processing model provided in the above embodiment. Specifically:
[0255] Generally, the computer device 1200 includes a processor 1210 and a memory 1220.
[0256] The processor 1210 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1210 may be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), or programmable logic array (PLA). The processor 1210 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1210 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1210 may further include an AI processor, which is used to process computational operations related to machine learning.
[0257] The memory 1220 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1220 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1220 are used to store a computer program, and the computer program is configured to be executed by one or more processors to implement an image processing method or a training method for an image processing model.
[0258] Those skilled in the art can understand that Figure 12 the structure shown in
[0259] In an exemplary embodiment, a computer-readable storage medium is further provided. A computer program is stored in the storage medium, and when the computer program is executed by a processor, the above-mentioned image processing method or the training method of the image processing model is implemented. Optionally, the computer-readable storage medium may include: Read-Only Memory (ROM), Random Access Memory (RAM), Solid State Drives (SSD), or optical discs, etc. Among them, the random access memory may include Resistance Random Access Memory (ReRAM) and Dynamic Random Access Memory (DRAM).
[0260] In an exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the above-mentioned image processing method or the training method of the image processing model.
[0261] It should be noted that in the practical application of the collection and processing of relevant data (such as line drawing images, reference images, and sample line drawing images, etc.) in this application, the informed consent or separate consent of the personal information subject should be obtained strictly in accordance with the requirements of relevant national laws and regulations, and subsequent data use and processing behaviors should be carried out within the scope authorized by laws and regulations and the personal information subject.
[0262] It should be understood that "a plurality of" mentioned herein refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. In addition, the step numbers described in this article only exemplarily show a possible execution sequence between steps. In some other embodiments, the above steps may not be executed in the order of the numbers. For example, two steps with different numbers are executed simultaneously, or two steps with different numbers are executed in the reverse order of the illustration. The embodiments of this application do not make any limitations in this regard.
[0263] The above are only optional embodiments of this application, and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included in the protection scope of this application.
Claims
1. An image processing method, characterized in that: The method comprises: Divide the line draft image into N line draft sub-images, where N is an integer greater than 1; Respectively acquiring reference images matching the N line draft sub-images to obtain N reference subsets, wherein the reference subset corresponding to each line draft sub-image includes at least one reference image matching the line draft sub-image; Performing a dimensionality reduction transformation on each of the reference images in the N reference subsets to obtain N reference feature sets, wherein the reference feature sets include a reference latent space of at least one reference image, and the reference latent space of the reference image is used to indicate feature information of the reference image; According to the position coding information, the noise latent space and the N reference feature sets, a position-coded noise latent space and at least one position-coded reference latent space are obtained, wherein the position coding information includes N local position information and one central area information, each local position information is used to indicate the positional relationship between each element in any reference latent space in a reference feature set, and the central area information is used to indicate the positional relationship between each element in the noise latent space, and the noise latent space includes at least one real noise; Performing a dimensionality reduction transformation on the line draft image to obtain a line draft latent space of the line draft image, wherein the line draft latent space is used to indicate feature information of the line draft image; A colored line drawing image is generated based on the line drawing latent space, the position-encoded noise latent space, the causal sparse attention map and the at least one position-encoded reference latent space, wherein the causal sparse attention map is used to indicate an attention calculation rule between the line drawing latent space and the position-encoded reference latent space.
2. The method according to claim 1, characterized in that The generating a colored line drawing image based on the line drawing latent space, the position-encoded noise latent space, the causal sparse attention map and the at least one position-encoded reference latent space comprises: The line drawing latent space and the at least one position-encoded reference latent space are input into an image processing model, and the image processing model performs T noise prediction and denoising operations on the position-encoded noise latent space based on the causal sparse attention map to obtain the colored line drawing image, where T is an integer greater than or equal to 1.
3. The method according to claim 2, characterized in that The step of inputting the line drawing latent space and the at least one position-encoded reference latent space into an image processing model, and performing T noise prediction and denoising operations on the position-encoded noise latent space based on the causal sparse attention graph by the image processing model to obtain the colored line drawing image includes: Inputting the at least one position-encoded reference latent space into a first neural network of the image processing model respectively, and having the first neural network perform self-attention calculation on each of the position-encoded reference latent spaces based on the causal sparse attention graph to obtain a self-attention calculation result for each of the at least one position-encoded reference latent spaces; storing the self-attention calculation results of each of the at least one position-encoded reference latent spaces in a buffer of the image processing model; The line drawing latent space is input into the image processing model, and the image processing model performs T noise prediction and denoising operations on the position-encoded noise latent space based on the self-attention calculation results of the causal sparse attention map and the at least one position-encoded reference latent space to obtain the colored line drawing image.
4. The method according to claim 3, characterized in that The step of inputting the line drawing latent space into an image processing model, and performing T noise prediction and denoising operations on the position-encoded noise latent space based on the self-attention calculation results of the causal sparse attention map and the at least one position-encoded reference latent space to obtain the colored line drawing image, including: For the i-th noise prediction and denoising operation in the T noise prediction and denoising operations, obtaining, in a buffer of the image processing model, a self-attention calculation result of each of the at least one position-encoded reference latent spaces, where i is a positive integer having an initial value of 1 and being less than or equal to T; Inputting the line draft latent space into a second neural network of the image processing model, and having the second neural network perform self-attention calculation on the line draft latent space based on the causal sparse attention graph to obtain a self-attention calculation result of the line draft latent space; Inputting the line draft latent space, the self-attention calculation result of the line draft latent space, the at least one position-encoded reference latent space, and the self-attention calculation result of each of the at least one position-encoded reference latent space into the first neural network, and the first neural network outputs a fused latent space based on the causal sparse attention graph, wherein the fused latent space is used to indicate geometric feature information of the line draft image and color information of the at least one reference image; Based on the fused latent space, the first neural network performs the i-th noise prediction and denoising operation on the position-encoded noise latent space to obtain a denoised noise latent space; The denoised noise latent space obtained by the last noise prediction and denoising operation is determined as the colored line drawing image.
5. The method according to claim 4, characterized in that Before inputting the line draft latent space into the second neural network of the image processing model, and performing self-attention calculation on the line draft latent space by the second neural network based on the causal sparse attention graph to obtain the self-attention calculation result of the line draft latent space, the method further includes: Acquire a color hint latent space and a color hint mask, wherein the color hint latent space is used to indicate at least one color hint point, the color hint point is used to indicate a color generated at a position corresponding to the line draft image, and the color hint mask is used to indicate a position of the at least one color hint point in the line draft image; Channel stitching is performed on the color cue latent space, the color cue mask, and the line draft latent space to obtain a stitched line draft latent space, and the stitched line draft latent space is input into the second neural network to perform the self-attention calculation to obtain a self-attention calculation result of the line draft latent space.
6. The method according to any one of claims 1 to 5, characterized in that: The step of obtaining a position-encoded noise latent space and at least one position-encoded reference latent space according to the position encoding information, the noise latent space and the N reference feature sets includes: Based on the central area information, position encoding is performed on the noise latent space to obtain the position-encoded noise latent space; For any reference latent space in the N reference feature sets, position encoding is performed on the reference latent space based on local position information corresponding to the reference latent space to obtain the position-encoded reference latent space.
7. The method according to claim 6, characterized in that The position encoding is to add any one of the reference latent spaces in the N reference feature sets to the corresponding local position information, or to add the noise latent space to the central area information.
8. The method according to any one of claims 1 to 7, characterized in that: The obtaining of reference images matching the N line draft sub-images respectively to obtain N reference subsets includes: In the reference image set, reference images respectively matching the N line draft sub-images are obtained to obtain the N reference sub-sets, wherein the reference image set includes at least one reference image.
9. The method according to claim 8, characterized in that The step of acquiring reference images that respectively match the N line draft sub-images from the reference image set to obtain the N reference sub-sets includes: For any one of the N line draft sub-images, extract image features of the line draft sub-image; respectively calculating the similarity between the image features of the line draft sub-image and the image features of each of the reference images in the reference image set; The k reference images with the highest similarity in the reference image set are used as a reference subset corresponding to the line draft sub-image, where k is a positive integer.
10. A training method for an image processing model, characterized in that: The method comprises: Acquire a training sample of the image processing model, the training sample comprising: a sample line draft image, a reference colored image, and N sample line draft sub-images obtained by dividing the sample line draft image, the reference colored image being a colored image corresponding to the sample line draft image; Acquire sample reference images matching the N sample line draft sub-images respectively through the image processing model to obtain N sample reference subsets, wherein the sample reference subset corresponding to each sample line draft sub-image includes at least one sample reference image matching the sample line draft sub-image; Performing a dimensionality reduction transformation on each of the sample reference images in the N sample reference subsets by using the image processing model to obtain N sample reference feature sets, wherein the sample reference feature sets include a reference latent space of at least one sample reference image, and the reference latent space of the sample reference image is used to indicate feature information of the sample reference image; Performing a dimensionality reduction transformation on the reference colored image through the image processing model to obtain a colored latent space of the reference colored image, wherein the colored latent space is used to indicate feature information of the reference colored image; Performing T noise addition operations on the shading latent space through the image processing model to obtain a noise latent space of the reference shading image, wherein the noise latent space includes at least one real noise, and T is an integer greater than or equal to 1; Performing a dimensionality reduction transformation on the sample line draft image by using the image processing model to obtain a line draft latent space of the sample line draft image, wherein the line draft latent space is used to indicate feature information of the sample line draft image; Obtaining, by the image processing model, a position-encoded noise latent space and at least one position-encoded reference latent space according to the position encoding information, the noise latent space and the N reference feature sets, wherein the position encoding information includes N local position information and one central area information, each local position information is used to indicate a positional relationship between each element in any reference latent space in a sample reference feature set, and the central area information is used to indicate a positional relationship between each element in the noise latent space; The image processing model is used to adjust the parameters of the image processing model based on the line draft latent space, the position-encoded noise latent space, the causal sparse attention map and the at least one position-encoded reference latent space to obtain a trained image processing model, wherein the causal sparse attention map is used to indicate the attention calculation rules between the line draft latent space and the position-encoded reference latent space.
11. The method according to claim 10, characterized in that The method of adjusting the parameters of the image processing model based on the line drawing latent space, the position-encoded noise latent space, the causal sparse attention map and the at least one position-encoded reference latent space to obtain a trained image processing model includes: Inputting the line drawing latent space and the at least one position-encoded reference latent space into a first neural network of the image processing model, and having the first neural network predict the noise added to the noise latent space for the i-th time based on the causal sparse attention graph to obtain the i-th predicted noise, where i is a positive integer whose initial value is 1 and is less than or equal to T; Calculating the i-th loss function value based on the i-th added real noise and the i-th predicted noise in the noise latent space; Based on the i-th loss function value, adjusting the parameters of the first neural network to obtain an adjusted image processing model; When i is less than T, assign i a value obtained by adding 1 to i, remove the i-th predicted noise from the noise latent space to obtain the i-th denoised noise latent space, and again input the line drawing latent space and the reference latent space after the at least one position encoding into the first neural network of the image processing model, and predict the noise added to the noise latent space for the i-th time based on the causal sparse attention map by the first neural network to obtain the i-th predicted noise. When i is equal to T, stop adjusting the image processing model, and determine the image processing model after the last adjustment as the trained image processing model.
12. The method according to claim 11, characterized in that The step of inputting the line drawing latent space and the at least one position-encoded reference latent space into a first neural network of the image processing model, and predicting the noise added to the noise latent space for the i-th time based on the causal sparse attention graph by the first neural network to obtain the i-th predicted noise includes: Inputting the at least one position-encoded reference latent space into the first neural network respectively, and the first neural network performs self-attention calculation on each of the position-encoded reference latent spaces based on the causal sparse attention graph to obtain a self-attention calculation result for each of the at least one position-encoded reference latent spaces; storing the self-attention calculation results of each of the at least one position-encoded reference latent spaces in a buffer of the image processing model; The line drawing latent space is input into the image processing model, and the first neural network predicts the noise added to the noise latent space for the i-th time based on the causal sparse attention map to obtain the i-th predicted noise.
13. The method according to claim 12, characterized in that The inputting the line drawing latent space into the image processing model, and predicting the noise added to the noise latent space for the i-th time by the first neural network based on the causal sparse attention graph to obtain the i-th predicted noise, comprises: Obtaining, in a buffer of the image processing model, a self-attention calculation result of each of the at least one position-encoded reference latent spaces; Inputting the line draft latent space into a third neural network and a fourth neural network of the image processing model, wherein the third neural network and the fourth neural network respectively perform self-attention calculation on the line draft latent space based on the causal sparse attention graph to obtain a first calculation result and a second calculation result of the line draft latent space, wherein the first calculation result is obtained by the third neural network performing self-attention calculation on the line draft latent space, and the second calculation result is obtained by the fourth neural network performing self-attention calculation on the line draft latent space; Based on a first preset weight, the first calculation result, and the second calculation result, a self-attention calculation result of the line draft latent space is obtained; Inputting the line draft latent space, the self-attention calculation result of the line draft latent space, the at least one position-encoded reference latent space, and the self-attention calculation result of each of the at least one position-encoded reference latent space into the first neural network, and the first neural network outputs a fused latent space based on the causal sparse attention graph, wherein the fused latent space is used to indicate the geometric feature information of the sample line draft image and the color information of the at least one sample reference image; The first neural network predicts the noise added to the noise latent space for the i-th time based on the fused latent space to obtain the i-th predicted noise.
14. The method according to claim 13, characterized in that Before inputting the line draft latent space into the third neural network and the fourth neural network of the image processing model, and respectively performing self-attention calculation on the line draft latent space based on the causal sparse attention graph by the third neural network and the fourth neural network to obtain the first calculation result and the second calculation result of the line draft latent space, the method further includes: Acquire a color hint latent space and a color hint mask, wherein the color hint latent space is used to indicate at least one color hint point, the color hint point is used to indicate a color generated at a position corresponding to the sample line draft image, and the color hint mask is used to indicate a position of the at least one color hint point in the sample line draft image; Channel stitching is performed on the color cue latent space, the color cue mask, and the line draft latent space to obtain a stitched line draft latent space, and the stitched line draft latent space is input into the third neural network and the fourth neural network to perform the self-attention calculation to obtain a self-attention calculation result of the line draft latent space.
15. The method according to claim 14, characterized in that The method further comprises: For any one of the at least one color hint point, based on the color hint mask, calculating a variance of a color value of at least one pixel corresponding to the color hint point in the sample line draft image; When the variance of the color value of at least one pixel corresponding to the color hint point in the sample line draft image is greater than a first threshold, the color hint point is determined to be invalid.
16. An image processing device, characterized in that: The device comprises: A division module, used for dividing the line draft image into N line draft sub-images, where N is an integer greater than 1; An acquisition module, configured to respectively acquire reference images matching the N line draft sub-images to obtain N reference subsets, wherein the reference subset corresponding to each line draft sub-image includes at least one reference image matching the line draft sub-image; A first transformation module is configured to perform a dimensionality reduction transformation on each of the reference images in the N reference subsets to obtain N reference feature sets, wherein the reference feature sets include a reference latent space of at least one reference image, and the reference latent space of the reference image is used to indicate feature information of the reference image; an encoding module, configured to obtain a position-encoded noise latent space and at least one position-encoded reference latent space according to position encoding information, a noise latent space and the N reference feature sets, wherein the position encoding information includes N local position information and one central area information, each local position information is used to indicate a positional relationship between elements in any reference latent space in a reference feature set, and the central area information is used to indicate a positional relationship between elements in the noise latent space, and the noise latent space includes at least one real noise; A second transformation module, configured to perform a dimensionality reduction transformation on the line draft image to obtain a line draft latent space of the line draft image, wherein the line draft latent space is used to indicate feature information of the line draft image; A generation module is used to generate a colored line drawing image based on the line drawing latent space, the position-encoded noise latent space, the causal sparse attention map and the at least one position-encoded reference latent space, wherein the causal sparse attention map is used to indicate an attention calculation rule between the line drawing latent space and the position-encoded reference latent space.
17. A training device for an image processing model, characterized in that: The device comprises: A first acquisition module is used to acquire a training sample of the image processing model, wherein the training sample includes: a sample line draft image, a reference colored image, and N sample line draft sub-images obtained by dividing the sample line draft image, wherein the reference colored image is a colored image corresponding to the sample line draft image; a second acquisition module, configured to respectively acquire sample reference images matching the N sample line draft sub-images through the image processing model to obtain N sample reference subsets, wherein the sample reference subset corresponding to each sample line draft sub-image includes at least one sample reference image matching the sample line draft sub-image; A first transformation module, configured to perform a dimensionality reduction transformation on each of the sample reference images in the N sample reference subsets by using the image processing model to obtain N sample reference feature sets, wherein the sample reference feature sets include a reference latent space of at least one sample reference image, and the reference latent space of the sample reference image is used to indicate feature information of the sample reference image; A second transformation module, configured to perform a dimensionality reduction transformation on the reference colored image through the image processing model to obtain a colored latent space of the reference colored image, wherein the colored latent space is used to indicate feature information of the reference colored image; A noise adding module, configured to perform T noise adding operations on the shading latent space through the image processing model to obtain a noise latent space of the reference shading image, wherein the noise latent space includes at least one real noise, and T is an integer greater than or equal to 1; A third transformation module, configured to perform a dimensionality reduction transformation on the sample line draft image through the image processing model to obtain a line draft latent space of the sample line draft image, wherein the line draft latent space is used to indicate feature information of the sample line draft image; an encoding module, configured to obtain, by means of the image processing model, a position-encoded noise latent space and at least one position-encoded reference latent space according to the position encoding information, the noise latent space and the N reference feature sets, wherein the position encoding information includes N local position information and one central area information, each local position information is used to indicate a positional relationship between elements in any reference latent space in a sample reference feature set, and the central area information is used to indicate a positional relationship between elements in the noise latent space; An adjustment module is used to adjust the parameters of the image processing model based on the line draft latent space, the position-encoded noise latent space, the causal sparse attention map and the at least one position-encoded reference latent space through the image processing model to obtain a trained image processing model, wherein the causal sparse attention map is used to indicate an attention calculation rule between the line draft latent space and the position-encoded reference latent space.
18. A computer device, characterized in that: The computer device includes a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the method according to any one of claims 1 to 9, or to implement the method according to any one of claims 10 to 15.
19. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and the computer program is used to be executed by a processor to implement the method according to any one of claims 1 to 9, or to implement the method according to any one of claims 10 to 15.
20. A computer program product, characterized in that The computer program product comprises a computer program, which is loaded and executed by a processor to implement the method according to any one of claims 1 to 9, or to implement the method according to any one of claims 10 to 15.
Citation Information
Cited By
Green ammonia production hydrogen load prediction method based on sparse attention variational Bayes
CN121167295A