Model training method, image processing method and related equipment
By fusing the pixel category information and noise of the training image during the machine learning model training stage, non-uniformly distributed noise is generated, the problem of texture confusion in local areas in the generated image by machine learning model is solved, and the quality and accuracy of image generation is improved, which is suitable for image recovery, style conversion and text-guided generation.
Patent Information
- Application Number
- CN202410198064.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-21
- Publication Date
- 2025-08-22
AI Technical Summary
When generating images, the texture distribution of different local areas is chaotic, resulting in poor quality of generated images and inability to effectively utilize the data distribution differences in local areas.
By obtaining pixel category information and noise of the training image to fuse, a non-uniformly distributed second noise is generated, and a machine learning model is trained using a loss function to improve the similarity of the predicted noise. It combines position encoding and mask generation to assist the model in learning the spatial position and semantic information of the image.
It reduces the probability of confusion between textures in different local areas, improves the quality and accuracy of generated images, and is suitable for application scenarios such as image recovery, style conversion and text guidance generation.
Smart Images

Figure CN120525075A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to a model training method, an image processing method, and related equipment. Background Art
[0002] Artificial intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a branch of computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. Generating images using machine learning models is a common application of AI. This can be achieved by performing multiple denoising operations on pure noise.
[0003] Specifically, during the training phase of a machine learning model, noise can be sampled from a standard Gaussian distribution space. This noise can then be used to add noise to an image to obtain a noisy image. The noisy image is then fed into the machine learning model, which generates predicted noise. The model is then trained using a loss function that indicates the similarity between the predicted noise and the sampled noise. During the inference phase of the machine learning model, the trained model can be used to generate predicted noise multiple times. This predicted noise can then be used to perform denoising operations on the pure noise multiple times to generate an image.
[0004] However, the texture distribution of different local areas in an image may be different, resulting in different data distribution in different local areas of an image. Since in the training stage of the machine learning model, the loss function indicates the similarity between the predicted noise and the sampled noise, and the sampled noise is sampled from the standard Gaussian distribution space, that is, the sampled noise is uniformly distributed, then the predicted noise generated in the inference stage of the machine learning model is also uniformly distributed, resulting in confusion between the textures of different local areas in the image obtained based on the predicted noise generated by the machine learning model, that is, it is impossible to obtain a good quality image. Summary of the Invention
[0005] The present application provides a model training method, an image processing method, and related equipment, which are used to reduce the probability of confusion between textures in different local areas of a generated image, thereby facilitating the acquisition of images with better quality.
[0006] This application provides the following technical solutions:
[0007] In a first aspect, the present application provides a model training method that can be used in the field of image processing in the field of artificial intelligence. The method includes: a training device can obtain first information corresponding to a first training image, and the first information indicates the category of pixels in the first training image. For example, if the trained first machine learning model is used in an application scenario of image restoration, the first training image can be a low-resolution image. If the trained first machine learning model is used in an application scenario of image style conversion, the first training image can be an image before style conversion, and the "image before style conversion" can also be called a "content image". If the trained first machine learning model is used in an application scenario of text-guided image generation, the first training image can be an image that matches the descriptive text, etc.
[0008] The training device can also obtain the first noise. For example, the training device can sample the first noise from a preset sampling space. For example, the preset sampling space can be a sampling space that conforms to the Gaussian distribution, so that the sampled first noise can be noise that conforms to the Gaussian distribution; or, the preset sampling space can also be a sampling space with a non-Gaussian distribution, so that the sampled first noise can be noise that does not conform to the Gaussian distribution.
[0009] The training device can then fuse the first information and the first noise to obtain the second noise, and use the second noise to add noise to the second training image to obtain the third training image; wherein, the image content of the first training image and the second training image is the same, and the third training image is used to train the first machine learning model, and the first machine learning model is a model for predicting the noise carried in the image.
[0010] In this implementation, a second noise is used in the training stage of the first machine learning model. The second noise is obtained by updating the first noise using the first information. The first information indicates the category of each pixel in the first image. Since the second image and the first image have the same image content, the first information can also indicate the category of each pixel in the second image. Since the first information is introduced into the second noise, that is, the data distribution information of the second image is introduced into the second noise, the second noise is no longer uniformly distributed noise. The loss function used when training the first machine learning model indicates the similarity between the predicted noise generated by the first machine learning model and the second noise. Therefore, the predicted noise generated by the trained first machine learning model is no longer uniformly distributed. Different local areas in the generated predicted noise have different data distributions. Therefore, the probability of confusion between textures in different local areas in the image obtained based on the predicted noise generated by the machine learning model will be reduced, which is conducive to obtaining better quality images.
[0011] In one possible implementation, in this method, the training device can also input the first training image into the second machine learning model to obtain at least one mask generated by the second machine learning model that corresponds one-to-one to at least one category, the at least one mask including a first mask corresponding to the first category, the first category being any one of the at least one category, and the first mask indicating the pixels in the first training image that are of the first category; and then obtaining the first information based on the at least one mask, the first information including the first value of each pixel in the first training image, the first values corresponding to pixels of different categories in the first information are different, and the first values corresponding to pixels of the same category in the first information are the same. In this implementation, the second machine learning model is used to generate at least one mask, the at least one mask indicating pixels of at least one category in the first training image, and then obtaining the first information based on the at least one mask, that is, providing a solution for automatically obtaining the first information, which improves the implementation fluency of the solution.
[0012] In one possible implementation, the first mask includes a second value for each pixel in the first training image, where the second value corresponding to the first pixel in the first training image is 1, and the second value corresponding to the second pixel in the first training image is 0. The first pixel is a pixel in the first training image whose category is the first category, and the second pixel is a pixel other than the first pixel in the first training image, that is, the category of the second pixel is not the first category. The training device then obtains the first information based on the at least one mask, which may include: the training device position-encoding the first training image to obtain first feature information of the first training image, and then determining the first value of each pixel in the first training image based on the at least one mask and the first feature information.
[0013] Position encoding can also be called position embedding. "Position encoding" means that the spatial position information of each pixel in the first training image is utilized in the process of generating the first value of each pixel in the first feature information. "Utilizing the spatial position information of the pixel in the first training image" can also be understood as utilizing the two-dimensional position information of the pixel in the first training image, or can also be understood as utilizing the position information of the row and column of the pixel in the first training image. Exemplarily, the position encoding algorithm used can be rotation position encoding, sine-cosine position encoding, absolute position encoding, conditional position encoding, or other position encoding algorithms.
[0014] In this implementation, the first training image is position-encoded to obtain first feature information, and then the first information is determined based on at least one mask and the first feature information. The first information not only carries the semantic information of the pixel points, but also can reflect the spatial position information of the pixel points in the image, which is beneficial to assist the first machine learning model to learn richer information about the image, thereby further improving the accuracy of the final generated image, that is, it is beneficial to obtain a better image.
[0015] In one possible implementation, in one application scenario, if the trained first machine learning model is used in an image restoration application scenario, the first training image can be a low-resolution image, and the second training image can be a high-resolution image with the same image content as the first training image, or the second training image can be a noisy image obtained by performing at least one noise addition operation on the high-resolution image. This clarifies the relationship between the first and second training images when the trained first machine learning model is applied to image restoration, thereby improving the integration of this solution with actual application scenarios. In another application scenario, if the trained first machine learning model is used in an image style transfer application scenario, the first training image can be an image before style transfer, and the second training image can be an image with the same image content as the first training image after style transfer, or the second training image can be a noisy image obtained by performing at least one noise addition operation on the style transfer image. In another application scenario, if the trained first machine learning model is used in a text-guided image generation application scenario, the first training image can be an image that matches the descriptive text, and the second training image can be the first training image, or the second training image can be a noisy image obtained by performing at least one noise addition operation on the first training image.
[0016] In one possible implementation, the training device fuses the first information and the first noise to obtain the second noise, including: the training device adds the first information to the first noise to obtain the second noise. This implementation provides a specific implementation scheme for updating the first noise based on the first information, which is simple to operate and easy to implement.
[0017] In one possible implementation, in this method, the training device may also input the third training image into the first machine learning model, generate predicted noise through the first machine learning model, and train the first machine learning model according to the loss function to obtain the trained first machine learning model. The loss function indicates the similarity between the predicted noise and the second noise. The predicted noise indicates the noise carried in the third training image; further, since the goal of training the first machine learning model using the loss function is to improve the similarity between the predicted noise generated by the first machine learning model and the second noise, and the second noise is used to add noise to the second training image to obtain the third training image, that is, the second noise can be understood as the additional noise of the third training image relative to the second training image, then the predicted noise can also be understood as the additional noise of the third training image relative to the second training image.
[0018] In this implementation method, it is further clarified how to use the third training image to train the first machine learning model, so that a better first machine learning model can be obtained more smoothly, which is conducive to generating more user-satisfied images in the application stage with the help of the trained first machine learning model.
[0019] In one possible implementation, the training device inputs the third training image into the first machine learning model, and generates prediction noise using the first machine learning model. This includes: the training device inputs the third training image and the first information into the first machine learning model, and generates prediction noise using the first machine learning model. In this implementation, directly inputting the first information into the first machine learning model helps the first machine learning model more intuitively understand the image content of the desired image, thereby facilitating higher-precision image generation.
[0020] In one possible implementation, the training device inputs the third training image into the first machine learning model and generates predicted noise through the first machine learning model, including: inputting the third training image and the first training image into the first machine learning model and generating predicted noise through the first machine learning model, wherein the first training image is a low-resolution image; or inputting the third training image and text description information into the first machine learning model and generating predicted noise through the first machine learning model, wherein the first training image is a training image adapted to the text description information; or inputting the third training image, the first training image, and image style information into the first machine learning model and generating predicted noise through the first machine learning model, wherein the image style information is converted image style information corresponding to the first training image. This implementation provides three application scenarios of the solution, expands the application scenarios of the solution, and improves the implementation flexibility of the solution.
[0021] In a second aspect, an embodiment of the present application provides an image processing method that can be used in the field of image processing in the field of artificial intelligence. In this method, an execution device obtains first information corresponding to a first image, where the first information indicates the category of a pixel in the first image; the first image, the first information, and the second information are input into a first machine learning model to obtain predicted noise generated by the first machine learning model; the second information is denoised using the predicted noise to obtain a second image, wherein the second information is pure noise, or the second information is an image obtained by denoising the pure noise at least once using the predicted noise generated by the first machine learning model, and the image content of the second image is the same as that of the first image. Exemplarily, "pure noise" can be understood as noise that does not carry valid image information.
[0022] In the second aspect of this application, the specific implementation methods of the steps in the second aspect, the meanings of the terms and the beneficial effects brought about can all be referred to the first aspect. The difference is that the "first training image" in the first aspect is replaced by the "first image" in the second aspect, and the "third training image" in the first aspect is replaced by the "second information" in the second aspect. No further details will be given here.
[0023] In a third aspect, embodiments of the present application provide an image processing method that can be used in the field of image processing within the field of artificial intelligence. In this method, an execution device inputs second information into a trained first machine learning model to obtain first predicted noise generated by the trained first machine learning model, where the training method for the first machine learning model is the method of the first aspect; and denoises the second information using the first predicted noise to obtain a second image, where the second information is pure noise, or the second information is an image obtained by denoising the pure noise at least once using the predicted noise generated by the first machine learning model.
[0024] In one possible implementation, if the method is applied to an image restoration scenario, inputting the second information into a trained first machine learning model includes: inputting a first image and the second information into the trained first machine learning model, where the first image is a low-resolution image. The method also includes: inputting the first image and the second image into the trained first machine learning model to obtain a second predicted noise generated by the trained first machine learning model; and denoising the second image using the second predicted noise to obtain a third image, where the third image is a high-resolution image with the same image content as the first image.
[0025] In the third aspect of this application, the specific implementation methods of the steps in the third aspect, the meanings of the terms and the beneficial effects brought about can all be referred to the first aspect. The difference is that the "first training image" in the first aspect is replaced by the "first image" in the second aspect, and the "third training image" in the first aspect is replaced by the "second information" in the second aspect. No further details will be given here.
[0026] In a fourth aspect, the present application provides a model training device that can be used in the field of image processing in the field of artificial intelligence. The model training device includes: an acquisition module for acquiring first information and first noise, the first information indicating the category of pixel points in the first training image; a fusion module for fusing the first information and the first noise to obtain a second noise; a noise addition module for using the second noise to add noise to the second training image to obtain a third training image, wherein the image content of the first training image and the second training image is the same, and the third training image is used to train a first machine learning model, and the first machine learning model is a model for predicting the noise carried in the image.
[0027] In one possible implementation, the model training device also includes: a generation module, which is used to input the first training image into the second machine learning model to obtain at least one mask generated by the second machine learning model that corresponds one-to-one to at least one category, and the at least one mask includes a first mask corresponding to the first category, the first category is any one of the at least one category, and the first mask indicates the pixel points in the first training image that are classified as the first category; an acquisition module, which is also used to obtain first information based on at least one mask, wherein the first information includes a first value of each pixel point in the first training image, the first values corresponding to pixels of different categories in the first information are different, and the first values corresponding to pixels of the same category in the first information are the same.
[0028] In one possible implementation, the first mask includes a second value of each pixel in the first training image, the second value corresponding to the first pixel in the first training image is 1, and the second value corresponding to the second pixel in the first training image is 0, the first pixel is a pixel in the first training image whose category is the first category, and the second pixel is a pixel other than the first pixel in the first training image. The acquisition module is specifically used to perform position encoding on the first training image, obtain first feature information of the first training image, and determine the first value of each pixel in the first training image based on at least one mask and the first feature information.
[0029] In one possible implementation, if the trained first machine learning model is applied to an image restoration scenario, the first training image is a low-resolution image, the second training image is a high-resolution image with the same image content as the first training image, or the second training image is an image obtained by denoising the high-resolution image at least once.
[0030] In one possible implementation, the model training device also includes: a generation module, used to input the third training image into the first machine learning model, and generate predicted noise through the first machine learning model; a training module, used to train the first machine learning model according to the loss function to obtain the trained first machine learning model, and the loss function indicates the similarity between the predicted noise and the second noise.
[0031] In one possible implementation, the generation module is specifically used to input the third training image and the first information into the first machine learning model, and generate prediction noise through the first machine learning model.
[0032] In one possible implementation, the generation module is specifically used to: input the third training image and the first training image into the first machine learning model, and generate predicted noise through the first machine learning model, wherein the first training image is a low-resolution image; or, input the third training image and text description information into the first machine learning model, and generate predicted noise through the first machine learning model, wherein the first training image is a training image adapted to the text description information; or, input the third training image, the first training image and image style information into the first machine learning model, and generate predicted noise through the first machine learning model, and the image style information is the converted image style information corresponding to the first training image.
[0033] In the fourth aspect of this application, the specific implementation methods of the steps in the fourth aspect, the meanings of the terms and the beneficial effects brought about can all be referred to the first aspect and will not be repeated here.
[0034] In a fifth aspect, the present application provides an image processing device that can be used in the field of image processing in the field of artificial intelligence. The image processing device includes: an acquisition module for acquiring first information corresponding to a first image, where the first information indicates the category of pixels in the first image; a generation module for inputting the first image, the first information, and the second information into a first machine learning model to obtain predicted noise generated by the first machine learning model; a denoising module for denoising the second information using the predicted noise to obtain a second image, wherein the second information is pure noise, or the second information is an image obtained by denoising the pure noise at least once using the predicted noise generated by the first machine learning model, and the image content of the second image is the same as that of the first image.
[0035] In one possible implementation, the acquisition module is specifically used to: input the first image into the second machine learning model, obtain at least one mask generated by the second machine learning model that corresponds one-to-one to at least one category, the at least one mask includes a first mask corresponding to the first category, the first category is any one of the at least one category, and the first mask indicates the pixels in the first image that are classified as the first category; obtain first information based on the at least one mask, wherein the first information includes a first value of each pixel in the first image, the first values corresponding to pixels of different categories in the first information are different, and the first values corresponding to pixels of the same category in the first information are the same.
[0036] In the fifth aspect of this application, the specific implementation methods of the steps in the fifth aspect, the meanings of the terms and the beneficial effects brought about can all be referred to the second aspect and will not be repeated here.
[0037] In a sixth aspect, the present application provides an image processing device that can be used in the field of image processing in the field of artificial intelligence. The image processing device includes: an input module for inputting second information into a trained first machine learning model to obtain first predicted noise generated by the trained first machine learning model, and the training method of the first machine learning model is the method of the first aspect; a denoising module for denoising the second information using the first predicted noise to obtain a second image, wherein the second information is pure noise, or the second information is an image obtained by denoising the pure noise at least once using the predicted noise generated by the first machine learning model.
[0038] In one possible implementation, the input module is specifically configured to input a first image and second information into a trained first machine learning model, where the first image is a low-resolution image. The input module is further configured to input the first image and second image into the trained first machine learning model to obtain a second predicted noise generated by the trained first machine learning model. The denoising module is further configured to denoise the second image using the second predicted noise to obtain a third image, where the third image is a high-resolution image with the same image content as the first image.
[0039] In the sixth aspect of this application, the specific implementation methods of the steps in the sixth aspect, the meanings of the terms and the beneficial effects brought about can all be referred to the third aspect and will not be repeated here.
[0040] In the seventh aspect, the present application provides a training device, including a processor and a memory, the processor is coupled to the memory, the memory is used to store programs; the processor is used to execute the programs in the memory, so that the training device executes the method described in the first aspect above.
[0041] In an eighth aspect, the present application provides an execution device, including a processor and a memory, wherein the processor is coupled to the memory, the memory is used to store programs; the processor is used to execute the programs in the memory, so that the execution device executes the method described in the second or third aspect above.
[0042] In a ninth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes the method described in the first, second or third aspect above.
[0043] In a tenth aspect, the present application provides a computer program product, which includes a program. When the program runs on a computer, it enables the computer to execute the method described in the first aspect, the second aspect or the third aspect above.
[0044] In an eleventh aspect, the present application provides a chip system, which includes a processor for supporting the implementation of the functions involved in the above aspects, for example, sending or processing the data and / or information involved in the above methods. In one possible design, the chip system also includes a memory, which is used to store program instructions and data necessary for the terminal device or communication device. The chip system can be composed of a chip or can include a chip and other discrete devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 A schematic diagram of the structure of the artificial intelligence main framework provided in the embodiment of the present application;
[0046] Figure 2 A schematic diagram of the working principle of generating an image using a first machine learning model provided in an embodiment of the present application;
[0047] Figure 3 A system architecture diagram of an image processing system provided in an embodiment of the present application;
[0048] Figure 4 A flow chart of a method for training a model provided in an embodiment of the present application;
[0049] Figure 5 A schematic diagram of determining first information provided in an embodiment of the present application;
[0050] Figure 6 A schematic diagram of a training method for a model provided in an embodiment of the present application;
[0051] Figure 7 A schematic diagram of a flow chart of an image processing method provided in an embodiment of the present application;
[0052] Figure 8A schematic diagram of the beneficial effects provided by the embodiments of the present application;
[0053] Figure 9 A schematic diagram of the structure of a training device for a model provided in an embodiment of the present application;
[0054] Figure 10 A schematic diagram of the structure of an image processing device provided in an embodiment of the present application;
[0055] Figure 11 Another structural diagram of the image processing device provided in an embodiment of the present application;
[0056] Figure 12 A schematic diagram of the structure of an execution device provided in an embodiment of the present application;
[0057] Figure 13 A schematic diagram of the structure of the training device provided in an embodiment of the present application;
[0058] Figure 14 A schematic diagram of the structure of the chip provided in an embodiment of the present application. DETAILED DESCRIPTION
[0059] The following describes the embodiments of the present application in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present application, rather than all of the embodiments. It is known to those skilled in the art that with the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0060] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0061] In the embodiments of the present application, "sending" and "receiving" indicate the direction of signal transmission. For example, "sending information to XX device" can be understood as the destination of the information being XX device, which can include direct sending through the air interface, as well as indirect sending through the air interface by other units or modules. "Receiving information from YY device" can be understood as the source of the information being YY device, which can include direct receiving from YY device through the air interface, as well as indirect receiving from YY device through the air interface from other units or modules. "Sending" can also be understood as the "output" of the chip interface, and "receiving" can also be understood as the "input" of the chip interface. In other words, sending and receiving can be performed between devices or within a device, for example, between components, modules, chips, software modules or hardware modules within the device through a bus, trace or interface. It is understandable that information may undergo necessary processing, such as encoding, modulation, etc., between the source and destination of the information, but the destination can understand the valid information from the source. Similar expressions in this application can be understood similarly and will not be repeated.
[0062] In the embodiments of the present application, "indication" may include direct indication and indirect indication, and may also include explicit indication and implicit indication. The information indicated by a certain information (such as the indication information described below) is called information to be indicated. In the specific implementation process, there are many ways to indicate the information to be indicated, such as but not limited to, the information to be indicated can be directly indicated, such as the information to be indicated itself or the index of the information to be indicated. The information to be indicated can also be indirectly indicated by indicating other information, wherein there is an association between the other information and the information to be indicated; it is also possible to indicate only a part of the information to be indicated, while the other parts of the information to be indicated are known or agreed in advance, for example, the indication of specific information can be achieved with the help of the arrangement order of each information agreed in advance (such as predefined by the protocol), thereby reducing the indication overhead to a certain extent. The present application does not limit the specific method of indication. It is understandable that, for the sender of the indication information, the indication information can be used to indicate the information to be indicated, and for the receiver of the indication information, the indication information can be used to determine the information to be indicated.
[0063] First, the overall workflow of the artificial intelligence system is described. Figure 1 , Figure 1The following diagram illustrates a structural diagram of the AI framework. This framework is explained below from two perspectives: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). The "intelligent information chain" reflects the entire process from data acquisition to processing. For example, it encompasses the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. Throughout this process, data undergoes a condensed progression from "data-information-knowledge-wisdom." The "IT value chain," encompassing the entire process from the underlying infrastructure of human intelligence, information (provided and processed by technology), to the system's industrial ecosystem, reflects the value that AI brings to the information technology industry.
[0064] (1) Infrastructure
[0065] The infrastructure provides computing power support for artificial intelligence systems, enabling communication with the outside world and providing support through the basic platform. Communication with the outside world is achieved through sensors; computing power is provided by intelligent chips, which can specifically adopt hardware acceleration chips such as central processing units (CPUs), embedded neural network processing units (NPUs), graphics processing units (GPUs), application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs); the basic platform includes related platform guarantees and support such as distributed computing frameworks and networks, and can include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside world to obtain data, and this data is provided to the intelligent chips in the distributed computing system provided by the basic platform for calculation.
[0066] (2) Data
[0067] Data above the infrastructure layer represents data sources for AI. This data includes graphics, images, voice, and text, as well as IoT data from traditional devices. This includes business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.
[0068] (3) Data processing
[0069] Data processing generally includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.
[0070] Among them, machine learning and deep learning can symbolize and formalize data for intelligent information modeling, extraction, preprocessing, and training.
[0071] Reasoning refers to the process of simulating human intelligent reasoning in computers or intelligent systems, using formalized information to perform machine thinking and solve problems based on reasoning control strategies. Typical functions are search and matching.
[0072] Decision-making refers to the process of making decisions after intelligent information is reasoned, and usually provides functions such as classification, sorting, and prediction.
[0073] (4) General ability
[0074] After the data has undergone the data processing mentioned above, some general capabilities can be further formed based on the results of the data processing, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0075] (5) Smart products and industry applications
[0076] Smart products and industry applications refer to the products and applications of artificial intelligence systems in various fields. They are the encapsulation of the overall artificial intelligence solution, which productizes intelligent information decision-making and realizes practical application. Its application areas mainly include: smart terminals, smart manufacturing, smart transportation, smart homes, smart medical care, smart security, autonomous driving, smart cities, etc.
[0077] The method provided in this application can be applied to various application scenarios involving image processing, and optionally, can be applied to various application scenarios involving image generation using machine learning models. Since the method provided in this application can be used in fields such as smart terminals, smart healthcare, smart cities, and smart homes, the following provides examples of application scenarios in various application areas of this application.
[0078] Application Scenario 1: Image Restoration
[0079] In the embodiment of the present application, image restoration refers to restoring a high-resolution image based on a low-resolution image, that is, it is necessary to generate a high-resolution image based on a low-resolution image. The image content of the aforementioned low-resolution image and high-resolution image is consistent, and the difference is that the resolutions of the two images are different. "Low-resolution image" can also be called "low-precision image" and "high-resolution image" can also be called "high-precision image". For example, in the field of smart terminals, there may be a need to restore low-resolution images taken by mobile phones to high-resolution images. For another example, in the field of smart medical care, there may be a need to restore low-resolution medical images to high-resolution medical images. For another example, in the field of smart cities, remote sensing images of cities may be obtained, and there may also be a need to restore low-resolution remote sensing images to high-resolution remote sensing images.
[0080] Application Scenario 2: Image Style Transfer
[0081] For example, image style conversion can also be called image style transfer. In the field of smart terminals, there may be a need to convert the image style of image 1 from the original image style A to image style B, that is, it is necessary to generate image 2 with image style B based on image 1. The image content of image 1 and image 2 are the same but the image style is different. For example, image 1 is an image taken by a mobile phone, and you want to convert image 1 to image 2 in the style of Van Gogh; for another example, image 1 is taken on a sunny day, and you want to convert image 1 to image 2 with the image style of a cloudy day; for another example, image 1 uses filter A, and you want to convert image 1 to image 2 with the image style of filter B, and so on. The specific image style to which the original image is transferred can be determined based on the actual needs of the user and is not limited here.
[0082] Application Scenario 3: Text-guided Image Generation
[0083] For example, in the field of smart terminals or smart homes, there may be a demand to generate an image that matches a description text input by a user.
[0084] It should be noted that the above examples are only for facilitating understanding of the application scenarios of this solution and are not intended to limit this solution. Not all application scenarios of the embodiments of this application are listed here one by one.
[0085] In all the above scenarios, the idea of performing multiple denoising operations on pure noise can be used to generate images. In the aforementioned process of multiple denoising of pure noise, a machine learning model that has performed training operations (hereinafter referred to as the "first machine learning model" for the convenience of description) can be used to generate predicted noise. Correspondingly, in the training stage of the first machine learning model, the image can be denoised to obtain a noisy image, and the noisy image can be input into the first machine learning model. The predicted noise is generated by the first machine learning model, and the first machine learning model is trained using a loss function that indicates the similarity between the predicted noise and the correct noise. For example, "pure noise" in this application can be understood as noise that does not carry valid image information.
[0086] In order to understand this solution more intuitively, the following Figure 2 , the principle of generating images using the first machine learning model is introduced. Figure 2 A schematic diagram of the working principle of generating an image using a first machine learning model provided in an embodiment of the present application. Figure 2 Includes two sub-schematic diagrams, upper and lower. Figure 2 The upper diagram of represents the training phase of the first machine learning model. Figure 2 The following diagram represents the first stage of applying the machine learning model. Figure 2 In the above sub-schematic diagram, in the training stage of the first machine learning model, X0 represents the image acquired by the training device, Xt-1 represents the image obtained after performing t-1 noise addition operations on X0, Xt represents the image obtained after performing t noise addition operations on X0, that is, another noise addition operation is performed on Xt-1, and XT represents the image obtained after performing T noise addition operations on X0. Exemplarily, XT can be specifically expressed as pure noise.
[0087] Then, after the training device acquires an image, it can perform T trainings on the first machine learning model based on the aforementioned acquired image. Regarding the specific implementation method of the training device performing the tth training (i.e., any training) among the T trainings, the training device can input Xt into the first machine learning model, generate the predicted noise between Xt-1 and Xt through the first machine learning model, and then determine the function value of the loss function, wherein the loss function indicates the similarity between the predicted noise and the actual noise between Xt-1 and Xt; after determining the function value of the loss function, the training device adjusts the weight parameters of the first machine learning model based on the back propagation algorithm to implement one training of the first machine learning model. The training device can use the acquired multiple images to iteratively train the first machine learning model until the convergence condition is met, thereby obtaining the first machine learning model that has performed the training operation.
[0088] See also Figure 2 In the following diagram, during the application phase of the first machine learning model, the execution device can generate an image by performing multiple denoising operations on pure noise. For example, PT represents the pure noise acquired by the execution device. Denoising operations can be performed on PT T times to obtain P0, where P0 represents the desired image. Pt represents the noisy image obtained after performing Tt denoising operations on PT, and Pt-1 represents the noisy image obtained after performing one denoising operation on Pt.
[0089] In the application stage of the first machine learning model, after the execution device obtains a pure noise, it can use the first machine learning model that has performed the training operation to perform t prediction noise generation operations. For the specific implementation method of the execution device performing any one of the t prediction noise generation operations, the training device can input Pt into the first machine learning model, generate prediction noise through the first machine learning model, and use the above-mentioned prediction noise to perform a denoising operation on Pt to obtain Pt-1; then after the execution device performs t denoising operations, P0 can be obtained. It should be understood that Figure 2 The examples are only for facilitating understanding of this solution and are not intended to limit this solution.
[0090] Since the texture distribution of different local areas in the image obtained using the first machine learning model may be different, the data distribution of different local areas in an image may be different. However, in the related art, in the training stage of the first machine learning model, the loss function indicates the similarity between the predicted noise and the sampled noise, and the sampled noise is sampled from the standard Gaussian distribution space, that is, the sampled noise is uniformly distributed, then the predicted noise generated in the inference stage of the machine learning model is also uniformly distributed, resulting in confusion between the textures of different local areas in the image obtained based on the predicted noise generated by the machine learning model, that is, it is impossible to obtain a good quality image.
[0091] In order to solve the above problems, this application provides a model training method and an image processing method. Before introducing the method provided by this application in detail, please refer to Figure 3 , Figure 3 A system architecture diagram of the image processing system provided in the embodiment of the present application, Figure 3 In the figure, the image processing system 300 includes a training device 310, a database 320, an execution device 330 and a data storage system 340, and the execution device 330 includes a calculation module 331.
[0092] During the training phase of the first machine learning model 301, a training data set is stored in the database 320, the training device 310 generates the first machine learning model 301, and iteratively trains the first machine learning model 301 using the training data set to obtain a trained first machine learning model 301. The first machine learning model 301 can be specifically represented as a neural network or a non-neural network model. In the embodiments of the present application, only the first machine learning model 301 represented as a neural network is used as an example for description.
[0093] The first machine learning model 301 obtained by the training device 310 after executing the training operation can be deployed to the computing module 331 of the execution device 330. For example, the execution device 330 can be a desktop computer, a mobile phone, a tablet, an AR device, a virtual reality (VR) device, a robot, or other types of devices. The execution device 330 can call data, code, etc. in the data storage system 340, or it can store data, instructions, etc. in the data storage system 340. The data storage system 340 can be placed in the execution device 330, or the data storage system 340 can be an external memory relative to the execution device 330.
[0094] During the application phase of first machine learning model 301, execution device 330 inputs the second information into first machine learning model 301 in computing module 331, obtains predicted noise generated by first machine learning model 301, and uses the predicted noise to denoise the second information to obtain a second image. The second information is pure noise, or an image obtained by denoising pure noise at least once using the predicted noise generated by the first machine learning model. Execution device 330 repeats the aforementioned operations multiple times to obtain the desired final image.
[0095] In some embodiments of this application, please refer to Figure 2 , the execution device 330 and the client device can be integrated into the same device, so that the user can directly interact with the execution device 330. For example, when the client device is a mobile phone or a tablet, the execution device 330 can be a module in the host processor (Host CPU) of the mobile phone or tablet that uses the first machine learning model to process data. The execution device 330 can also be a graphics processing unit (GPU) or a neural network processor (NPU) in the mobile phone or tablet. The GPU or NPU is mounted on the host processor as a coprocessor, and the host processor assigns tasks.
[0096] It is worth noting that Figure 3This is merely a schematic diagram of the architecture of two data processing systems provided by an embodiment of the present invention, and the positional relationships between the devices, components, modules, etc. shown in the diagram do not constitute any limitation. For example, in other embodiments of the present application, the execution device 330 and the client device may be separate devices, and the execution device 330 may be configured with an input / output (I / O) interface to exchange data with the client device. After the execution device 330 generates the final desired image through the first machine learning model 301 in the computing module 331, the aforementioned image may be returned to the client device through the I / O interface.
[0097] In combination with the above description, the following describes the specific implementation process of the training phase and the application phase in the method provided by this application.
[0098] 1. Training Phase
[0099] In the embodiments of this application, please refer to Figure 4 , Figure 4 A flow chart of a method for training a model provided in an embodiment of the present application is provided. The method for training a model provided in an embodiment of the present application may include:
[0100] 401. Obtain first information, where the first information indicates categories of pixels in a first training image.
[0101] In an embodiment of the present application, the first information corresponding to the first training image can be used to indicate the category of the pixel points in the first training image. For example, the first training image includes at least one category of pixel points, and the first information can include the first value of each pixel point in the first training image. The first values corresponding to the pixels of different categories in the first information are different, and the first values corresponding to the pixels of the same category in the first information are the same. For example, pixel point 1 in the first training image is category 1, pixel point 2 in the first training image is category 2, and pixel point 3 in the first training image is category 1. The first values of pixel point 1 and pixel point 3 in the first information are both 1, that is, the first values corresponding to the pixels of the same category in the first information are the same; the first value of pixel point 2 in the first information is 2, that is, the first values corresponding to the pixels of different categories in the first information are different. It should be understood that the examples here are only for the convenience of understanding this solution and are not used to limit this solution.
[0102] For example, if the trained first machine learning model is used in an application scenario of image restoration, the first training image may be a low-resolution image, and the purpose of "image restoration" is to generate a high-resolution image with the same image content as the aforementioned low-resolution image. If the trained first machine learning model is used in an application scenario of image style conversion, the first training image may be an image before style conversion, and the "image before style conversion" may also be referred to as a "content image." If the trained first machine learning model is used in an application scenario of text-guided image generation, the first training image may be an image that matches the descriptive text, etc. When applied to other scenarios, the first training image may also be specifically expressed as other images, which is not limited in the embodiments of this application.
[0103] In one case, the first information corresponding to the first training image may be manually annotated information.
[0104] In another case, step 401 may include: the training device inputs the first training image into a second machine learning model that has performed a training operation, obtains N masks generated by the second machine learning model that has performed the training operation and corresponding to the N categories, and then obtains the first information based on the N masks, where N is an integer greater than or equal to 1. In the embodiment of the present application, at least one mask is generated by the second machine learning model, and the at least one mask indicates pixels of at least one category in the first training image, and then the first information is obtained based on the at least one mask, which provides a solution for automatically obtaining the first information, thereby improving the implementation fluency of the solution.
[0105] Optionally, the second machine learning model can be a machine learning model for performing image segmentation tasks. For example, the second machine learning model can be specifically expressed as a convolutional neural network (CNN), a fully connected neural network based on an attention mechanism, a residual neural network or other types of neural networks, etc. Exemplarily, the second machine learning model can adopt a segment anything model (SAM).
[0106] The granularity of the division of the N categories is related to the second machine learning model used. The N categories are divided at the granularity at which the second machine learning model supports division. For example, some machine learning models regard a person as a whole, so the category of the pixels included in the person in the image is all people; while some machine learning models regard a person's skin, clothes, eyes, and hair as different categories, so the category of the pixels of the person's skin in the image is skin, the category of the pixels of the person's clothes is clothes, the category of the pixels of the person's eyes is eyes, the category of the pixels of the person's hair is hair, etc. It should be understood that the examples here are only for the convenience of understanding this solution and are not used to limit this solution.
[0107] Among them, the N masks include a first mask corresponding to the first category, the first category is any one of the N categories, that is, the first mask is any one of the N masks, and the first mask among the N masks indicates the pixel points of the first category in the first training image, and each mask can be specifically expressed in the form of a matrix.
[0108] Exemplarily, in one case, each of the N masks may include a second value corresponding to each pixel in the first image, and the second value of each pixel in each of the N masks is either 0 or 1; then, in the first mask, a pixel with a value of 1 represents that the category of the pixel is the first category, and a pixel with a value of 0 represents that the category of the pixel is not the first category, that is, in the first mask, the second value corresponding to the first pixel in the first training image is 1, and the second value corresponding to the second pixel in the first training image is 0, the first pixel is a pixel in the first training image whose category is the first category, and the second pixel is a pixel other than the first pixel in the first training image, that is, the category of the second pixel is not the first category, then at least one mask generated by the second machine learning model can determine the category of each pixel in the first image.
[0109] Exemplarily, the training device obtains the first information based on at least one mask, which may include: the training device performs position encoding on the first training image to obtain first feature information of the first training image, and then determines the first value of each pixel in the first training image based on the at least one mask and the first feature information.
[0110] For example, the first feature information of the first training image may also be referred to as “initial position encoding of the first training image”. Exemplarily, the size of the first feature information may be 1xWxH, where W and H are the length and width of the first training image, respectively.
[0111] Among them, position encoding can also be called position embedding (PE), and the meaning of "position encoding" refers to the use of the spatial position information of the pixel in the first training image in the process of generating the first value of each pixel in the first feature information. "Using the spatial position information of the pixel in the first training image" can also be understood as using the two-dimensional position information of the pixel in the first training image, or it can also be understood as using the position information of the row and column of the pixel in the first training image. Exemplarily, the position encoding algorithm used can be rotary position embedding (RoPE), sine-cosine position encoding, absolute position encoding, conditional position encoding (conditional position encoding) or other types of position encoding algorithms, etc. The specific position encoding algorithm can be determined in combination with the actual application scenario, and is not limited in the embodiments of the present application.
[0112] To further understand this solution, a specific implementation method for determining a first value for each pixel of a first category (i.e., a first pixel) in a first training image based on a first mask and first feature information is first described herein. Optionally, the training device may multiply the first mask and the first feature information to obtain an intermediate result for each first pixel in at least one first pixel included in the first training image, and average the intermediate results for all first pixels to obtain the first value for each first pixel in the first information.
[0113] Exemplarily, the first training image includes at least one first pixel point, including first pixel point 1, first pixel point 2, first pixel point 3, first pixel point 4 and first pixel point 5. The intermediate results of each first pixel point obtained by the training device after multiplying the first mask and the first feature information include: the intermediate result of the first pixel point 1 is 1, the intermediate result of the first pixel point 2 is 0.8, the intermediate result of the first pixel point 3 is 1.4, the intermediate result of the first pixel point 4 is 1.6, and the intermediate result of the first pixel point 5 is 2. The value obtained by averaging 1, 0.8, 1.4, 1.6 and 2 is 1.36. Then, the first values of all the first pixel points in the first information are 1.36. It should be understood that the examples here are only for the convenience of understanding this scheme and are not used to limit this scheme.
[0114] Alternatively, the training device may multiply the first mask and the first feature information to obtain an intermediate result for each first pixel in at least one first pixel included in the first training image, obtain a maximum value from the intermediate results of all first pixels, and use the aforementioned maximum value as the first value of each first pixel in the first information. Alternatively, the training device may obtain a median value from the intermediate results of all first pixels, and use the aforementioned median value as the first value of each first pixel in the first information, etc., which is not limited in the embodiments of the present application.
[0115] The training device repeatedly performs the above operation on each of the N masks (that is, the operation performed on the first mask), thereby obtaining the first value of the pixel points of each category of the N categories of pixel points included in the first training image, that is, obtaining the first information.
[0116] For a more intuitive understanding of this solution, please refer to Figure 5 , Figure 5 A schematic diagram of determining first information provided in an embodiment of the present application. Figure 5 As shown, the training device generates N masks corresponding to the first training image, and obtains the first feature information after position encoding the first training image. The N masks can indicate N regions in the first training image, and each of the N regions includes pixels of a category in the first training image. Then, the first feature information can be divided into feature information of the N regions. The i-th mask is any one of the N masks, such as Figure 5 As shown, based on the i-th mask and the feature information of the i-th region included in the first feature information, the first value of all pixels in the i-th region can be obtained; for example, the i-th mask and the feature information of the i-th region included in the first feature information can be multiplied and then averaged to obtain the first value of all pixels in the i-th region. After obtaining the first value of the pixel points in each of the N regions, the first values of the pixel points in the N regions can be summarized to obtain the first information. It should be understood that Figure 5 The examples are only for facilitating understanding of this solution and are not intended to limit this solution.
[0117] In an embodiment of the present application, the first training image is position-encoded to obtain first feature information, and then the first information is determined based on at least one mask and the first feature information. The first information not only carries the semantic information of the pixel points, but also can reflect the spatial position information of the pixel points in the image, which is beneficial to assisting the first machine learning model to learn richer information about the image, thereby further improving the accuracy of the final generated image, that is, it is beneficial to obtain a better image.
[0118] In another case, in N masks, the second values corresponding to pixels of different categories are different. For example, N masks may include a total of N+1 values, and N values of the aforementioned N+1 values correspond one-to-one to N categories. For example, if the value of N is 5, then the N+1 values can include 0, 1, 2, 3, 4 and 5, and the N masks include 5 masks, where, in the first mask, the pixel point with a value of 1 represents that the category of the pixel point is category 1, and the pixel point with a value of 0 represents that the category of the pixel point is not category 1; in the second mask, the pixel point with a value of 2 represents that the category of the pixel point is category 2, and the pixel point with a value of 0 represents that the category of the pixel point is not category 2; in the third mask, the pixel point with a value of 3 represents that the category of the pixel point is category 3, and the pixel point with a value of 0 represents that the category of the pixel point is not category 3; in the fourth mask, the pixel point with a value of 4 represents that the category of the pixel point is category 4, and the pixel point with a value of 0 represents that the category of the pixel point is not category 4; in the fifth mask, the pixel point with a value of 5 represents that the category of the pixel point is category 5, and the pixel point with a value of 0 represents that the category of the pixel point is not category 5. The training device obtaining the first information based on the at least one mask may include: the training device superimposing the at least one mask together to obtain the first information.
[0119] 402. Obtain a first noise.
[0120] In an embodiment of the present application, the training device can sample the first noise from a preset sampling space. For example, the preset sampling space can be a sampling space that conforms to the Gaussian distribution, so that the sampled first noise can be noise that conforms to the Gaussian distribution; or, the preset sampling space can also be a sampling space with a non-Gaussian distribution, so that the sampled first noise can be noise that does not conform to the Gaussian distribution. It should be noted that since noise needs to be sampled from the sampling space multiple times during the multiple training processes of the first machine learning model, it is sufficient to ensure that the multiple noise sampling operations are sampled from the same preset sampling space.
[0121] 403. Fuse the first information and the first noise to obtain second noise.
[0122] In the embodiment of the present application, step 403 is an optional step. After acquiring the first information and the first noise, the training device may fuse the first information and the first noise to obtain the second noise.
[0123] Optionally, step 403 may include: the training device adds the first information to the first noise to obtain the second noise. In the embodiment of the present application, a specific implementation scheme for updating the first noise based on the first information is provided, which is simple to operate and easy to implement. Alternatively, the training device may multiply the first information by the first noise to obtain the second noise, or alternatively, the training device may subtract the first noise from the first information to obtain the second noise, etc. The specific method to be adopted can be determined in combination with the actual application scenario and is not limited in the embodiment of the present application.
[0124] 404. Use the second noise to add noise to the second training image to obtain a third training image, wherein the first training image and the second training image have the same image content, and the third training image is used to train the first machine learning model, which is a model for predicting noise carried in the image.
[0125] In an embodiment of the present application, after obtaining the second noise, the training device can use the second noise to add noise to the second training image to obtain a third training image. For example, in an application scenario, if the trained first machine learning model is used in an application scenario of image restoration, the first training image can be a low-resolution image, and the second training image can be a high-resolution image with the same image content as the first training image, or the second training image can be a noisy image obtained by performing at least one noise addition operation on a high-resolution image with the same image content as the first training image. In an embodiment of the present application, the relationship between the first training image and the second training image is clarified when the trained first machine learning model is applied to the scenario of image restoration, thereby improving the degree of integration between this solution and actual application scenarios.
[0126] In another application scenario, if the trained first machine learning model is used in an application scenario of image style conversion, the first training image may be an image before style conversion, and the second training image may be an image with the same image content as the first training image and after style conversion, or the second training image may be a noisy image obtained by performing at least one noise addition operation on the aforementioned style conversion image. In another application scenario, if the trained first machine learning model is used in an application scenario of text-guided image generation, the first training image may be an image that matches the descriptive text, and the second training image may be the first training image, or the second training image may be a noisy image obtained by performing at least one noise addition operation on the first training image.
[0127] It should be noted that step 403 is an optional step. If step 403 is not performed, the training device may also use the first noise to add noise to the second training image to obtain a third training image.
[0128] 405. Input the third training image into the first machine learning model, and generate prediction noise through the first machine learning model.
[0129] In the embodiment of the present application, step 405 is an optional step. For example, the predicted noise indicates the noise carried in the third training image. Furthermore, since the goal of training the first machine learning model using the loss function in the subsequent step 406 is to improve the similarity between the predicted noise generated by the first machine learning model and the second noise, and the second noise is used to add noise to the second training image to obtain the third training image, that is, the second noise can be understood as the additional noise in the third training image relative to the second training image, the predicted noise can also be understood as the additional noise in the third training image relative to the second training image.
[0130] Specifically, in an application scenario, if the trained first machine learning model is used in an application scenario of image restoration, step 405 may include: the training device inputs the third training image and the first training image into the first machine learning model, and generates predicted noise through the first machine learning model, wherein the first training image is a low-resolution image.
[0131] In another application scenario, if the trained first machine learning model is used in an application scenario of image style conversion, step 405 may include: the training device inputs the third training image, the first training image and the image style information into the first machine learning model, generates prediction noise through the first machine learning model, and the image style information is the converted image style information corresponding to the first training image.
[0132] Exemplarily, the image style information is the desired image style information. For example, the first training image can be understood as a "content image", and the image style information can specifically adopt a "style image", that is, the image content of the converted image is determined by the first training image, and the image style of the converted image is determined by the style image.
[0133] In another application scenario, if the trained first machine learning model is used in an application scenario where text is used to guide image generation, step 405 may include: the training device inputs the third training image and text description information into the first machine learning model, and generates prediction noise through the first machine learning model, wherein the first training image is a training image adapted to the text description information.
[0134] Optionally, in step 405, when the training device inputs the third training image into the first machine learning model, it may also input the first information into the first machine learning model, so that the first machine learning model uses more information to generate prediction noise.
[0135] Exemplarily, in an application scenario, if the trained first machine learning model is used in an application scenario of image restoration, step 405 may include: the training device inputs the third training image, the first training image and the first information into the first machine learning model, and generates predicted noise through the first machine learning model.
[0136] In another application scenario, if the trained first machine learning model is used in an application scenario of image style conversion, step 405 may include: the training device inputs the third training image, the first training image, the image style information and the first information into the first machine learning model, and generates predicted noise through the first machine learning model.
[0137] In another application scenario, if the trained first machine learning model is used in an application scenario where text is used to guide image generation, step 405 may include: the training device inputs the third training image, text description information and the first information into the first machine learning model, and generates prediction noise through the first machine learning model.
[0138] In the embodiments of the present application, three application scenarios of the present solution are provided, which expands the application scenarios of the present solution and improves the implementation flexibility of the present solution. In addition, the first information can be directly input into the first machine learning model to help the first machine learning model more intuitively understand the image content of the image to be generated, which is conducive to generating images with higher accuracy.
[0139] 406. Train the first machine learning model according to the loss function to obtain a trained first machine learning model, where the loss function indicates the similarity between the predicted noise and the second noise.
[0140] In this embodiment of the present application, step 406 is an optional step. After obtaining the predicted noise generated by the first machine learning model, the training device can train the first machine learning model according to a loss function to obtain a trained first machine learning model, where the loss function indicates the similarity between the predicted noise and the expected noise. Since step 403 is an optional step, if step 403 is performed, the expected noise can be the second noise generated in step 403; if step 403 is not performed, the expected noise can be the first noise.
[0141] Optionally, the training device generates a function value of a loss function based on the predicted noise and the second noise, and then updates the weight parameters of the first machine learning model based on the function value of the loss function and the backpropagation algorithm to achieve a single training of the first machine learning model, where the loss function indicates the similarity between the predicted noise and the second noise. In the embodiment of the present application, it is further clarified how to use the third training image to train the first machine learning model, so that a better first machine learning model can be obtained more smoothly, which is conducive to generating more user-satisfied images in the application stage with the help of the trained first machine learning model.
[0142] For example, in the process of training the first machine learning model, the predicted noise generated by the first machine learning model can be compared with the desired expected noise (i.e., the second noise obtained in step 403), and the weight vector of each layer of the machine learning model can be updated according to the difference between the two (of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the first machine learning model). For example, if the network's predicted noise is high, the weight vector is adjusted to make it predict lower, and the adjustment is continued until the first machine learning model can generate the desired expected noise or a value very close to the desired expected noise. Therefore, it is necessary to predefine "how to compare the difference between the predicted noise and the expected noise", which is the loss function or objective function, which are important equations for measuring the difference between the predicted noise and the expected noise. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference, and the training of the first machine learning model becomes a process of minimizing this loss as much as possible.
[0143] Machine learning models can use the backpropagation algorithm to correct the size of the initial machine learning model parameters during training, reducing the reconstruction error loss of the machine learning model. Specifically, forward propagation of the input signal to the output generates error loss. This error loss information is then backpropagated to update the initial machine learning model parameters, thereby converging the error loss. The backpropagation algorithm is a backward propagation movement dominated by error loss, aiming to obtain the optimal parameters of the machine learning model.
[0144] The training device can repeatedly execute steps 401 to 406 to implement iterative training of the first machine learning model until the convergence condition is met, and a first machine learning model that has performed the training operation (i.e., the first machine learning model after training) is obtained. The convergence condition can be that the number of training times for the first machine learning model reaches a preset number of times, and / or the convergence condition of the loss function is met.
[0145] It should be noted that the embodiment of the present application does not limit the number of executions between step 401 and steps 402 to 406. After executing step 401 once, steps 402 to 406 can be executed multiple times to achieve multiple training of the first machine learning model. After the training device obtains the next training image, the first information corresponding to the next training image is re-acquired.
[0146] For a more intuitive understanding of this solution, please refer to Figure 6 , Figure 6 A schematic diagram of a training method for a model provided in an embodiment of the present application. Figure 6 Taking the application scenario of image restoration provided by the method of the present application as an example, X0 represents a high-resolution image, Y0 represents a low-resolution image with the same content as the image X0, the training device inputs Y0 into the second machine learning model, obtains N masks corresponding to N categories generated by the second machine learning model, and determines the first information based on the N masks. In each noise addition operation of performing T noise addition operations on X0 to obtain XT, as shown in FIG. Figure 6 As shown, the training device samples the first noise from a preset sampling space, adds the first information to the first noise to obtain the second noise, and then performs a noise addition operation based on the second noise.
[0147] For example, in the t-th training operation among T training operations, the first noise can be obtained from the t-th sampling noise in the preset sampling space, and the sampled first noise is added to the first information to obtain the second noise; the second noise is used to perform a noise addition operation on Xt-1 to obtain Xt; Xt is input into the first machine learning model, and the predicted noise is generated by the first machine learning model; then, the function value of the loss function is generated according to the predicted noise and the aforementioned second noise, and the loss function indicates the similarity between the predicted noise and the second noise, and then based on the function value of the loss function and the back propagation algorithm, the weight parameters of the first machine learning model are updated to achieve one training of the first machine learning model. It should be understood that Figure 6 The examples are only for facilitating understanding of this solution and are not intended to limit this solution.
[0148] In this implementation, a second noise is used in the training stage of the first machine learning model. The second noise is obtained by updating the first noise using the first information. The first information indicates the category of the pixel points in the first image. Since the second image and the first image have the same image content, the first information can also indicate the category of the pixel points in the second image. Since the first information is introduced into the second noise, that is, the data distribution information of the second image is introduced into the second noise, the second noise is no longer uniformly distributed noise. The loss function used when training the first machine learning model indicates the similarity between the predicted noise generated by the first machine learning model and the second noise. Therefore, the predicted noise generated by the trained first machine learning model is no longer uniformly distributed. Different local areas in the generated predicted noise have different data distributions. Therefore, the probability of confusion between textures in different local areas in the image obtained based on the predicted noise generated by the machine learning model will be reduced, which is conducive to obtaining better quality images.
[0149] 2. Application Phase
[0150] In the embodiment of the present application, after executing the above Figure 3 After the training operation described in the corresponding embodiment, a first machine learning model that has been trained can be obtained, and then the first machine learning model that has been trained can be used to perform image processing in the application stage. For details, please refer to Figure 7 , Figure 7 This is a flow chart of an image processing method provided in an embodiment of the present application. The image processing method provided in an embodiment of the present application may include:
[0151] 701. Acquire first information corresponding to a first image, where the first information indicates a category of pixels in the first image.
[0152] In the embodiment of the present application, step 701 is an optional step. The specific implementation method of the execution device performing step 701 can refer to the above Figure 3 The description in the corresponding embodiment is different in that the “first training image” in the above description is replaced by the “first image” in this application, and the specific implementation method of step 701 is not repeated here.
[0153] 702. Input the second information into the first machine learning model to obtain prediction noise generated by the first machine learning model.
[0154] In the embodiment of this application, Figure 7 The first machine learning model in the corresponding embodiment can be Figure 3 The model after the training operation is performed according to the training method of the model shown in the corresponding embodiment.
[0155] Step 701 is an optional step. If step 701 is not performed, step 403 needs to be performed during the training phase of the first machine learning model. That is, the first information can be additionally introduced into the input of the first machine learning model, and / or, during the training phase of the first machine learning model, the first information can be fused with the first noise to guide the first machine learning model to learn the knowledge of data distribution in different local areas in the same image.
[0156] If step 701 is executed, step 702 may include: the execution device inputs the first image, the first information and the second information into the first machine learning model to obtain the predicted noise generated by the first machine learning model.
[0157] If step 701 is not performed, then if the first machine learning model is used in an application scenario of image restoration, then step 702 may include: the execution device inputs the second information and the first image into the first machine learning model, and generates predicted noise through the first machine learning model, where the first image is a low-resolution image. If the first machine learning model is used in an application scenario of image style conversion, then step 702 may include: the execution device inputs the second information, the first image, and the image style information into the first machine learning model, and generates predicted noise through the first machine learning model, where the image style information is the converted image style information corresponding to the first image. If the first machine learning model is used in an application scenario of text-guided image generation, then step 405 may include: the execution device inputs the second information and text description information into the first machine learning model, and generates predicted noise through the first machine learning model.
[0158] The specific implementation of the above steps can refer to the above Figure 3 The description in the corresponding embodiment is different in that the "first training image" in the above description is replaced by the "first image" in this application, and the "third training image" is replaced by the "second information". The specific implementation method of step 702 is not repeated here.
[0159] 703. De-noising the second information using the predicted noise to obtain a second image, wherein the second information is pure noise, or the second information is an image obtained by de-noising the pure noise at least once using the predicted noise generated by the first machine learning model.
[0160] In an embodiment of the present application, after obtaining the predicted noise generated by the first machine learning model, the execution device can use the predicted noise to denoise the second information and obtain a second image. After executing step 703, the execution device can enter step 702 again. If step 701 is executed, when entering step 702 again, the execution device can input the first image, the first information, and the second image into the first machine learning model to obtain the predicted noise generated by the first machine learning model, and use the predicted noise to denoise the second image to obtain a third image. The image content of the third image is the same as that of the first image. The execution device repeats the above operation at least once to obtain the final desired image.
[0161] For example, in one application scenario, if the trained first machine learning model is used in an image restoration application scenario, the desired image is a high-resolution image. In another application scenario, if the trained first machine learning model is used in an image style conversion application scenario, the desired image is a style-converted image.
[0162] If step 701 is not performed, then when the execution device enters step 702 again, the first image and the second image may be input into the first machine learning model to obtain the predicted noise generated by the first machine learning model. The predicted noise is then used to denoise the second image to obtain a third image. The third image has the same image content as the first image. The execution device repeats the above operation at least once to obtain the desired final image.
[0163] For example, in one application scenario, if the trained first machine learning model is used in an image restoration application scenario, the desired image is a high-resolution image. In another application scenario, if the trained first machine learning model is used in an image style conversion application scenario, the desired image is a style-converted image. In another application scenario, if the trained first machine learning model is used in a text-guided image generation application scenario, the desired image is an image that matches the descriptive text.
[0164] In order to have a more intuitive understanding of the beneficial effects of the method provided by this application, the beneficial effects of the method provided by this application are described here in combination with experimental data. Figure 8 , Figure 8A schematic diagram of the beneficial effects provided by the embodiments of the present application. Set14 and Urban100 are two different data sets. PSNR, SSIM, and FID are three indicators used to evaluate the accuracy of images obtained using machine learning models. A larger PSNR value represents a higher accuracy of the generated image, a larger SSIM value represents a higher accuracy of the generated image, and a smaller FID value represents a higher accuracy of the generated image. BebyGAN, USRGAN, HCFlow, HCFlow++, and SPSR are several machine learning models already available in the industry for generating images. From the above comparison, it can be seen that the images generated using the method provided in this application have the highest accuracy.
[0165] exist Figures 1 to 8 On the basis of the corresponding embodiments, in order to better implement the above solutions of the embodiments of the present application, the following also provides related devices for implementing the above solutions. Figure 9 , Figure 9 A structural schematic diagram of a training device for a model provided in an embodiment of the present application, wherein the neural network training device 900 includes: an acquisition module 901, for acquiring first information and first noise, the first information indicating the category of pixels in a first training image; a fusion module 902, for fusing the first information and the first noise to obtain second noise; and a noise addition module 903, for adding noise to the second training image using the second noise to obtain a third training image, wherein the first training image and the second training image have the same image content, and the third training image is used to train a first machine learning model, which is a model for predicting noise carried in an image.
[0166] Optionally, the model training device 900 also includes: a generation module 904, which is used to input the first training image into the second machine learning model to obtain at least one mask generated by the second machine learning model that corresponds one-to-one to at least one category, and the at least one mask includes a first mask corresponding to the first category, the first category is any one of the at least one category, and the first mask indicates the pixel points in the first training image that are classified as the first category; an acquisition module 901, which is also used to obtain first information based on at least one mask, wherein the first information includes a first value of each pixel point in the first training image, the first values corresponding to pixels of different categories in the first information are different, and the first values corresponding to pixels of the same category in the first information are the same.
[0167] Optionally, the first mask includes a second value of each pixel in the first training image, the second value corresponding to the first pixel in the first training image is 1, and the second value corresponding to the second pixel in the first training image is 0, the first pixel is a pixel in the first training image whose category is the first category, and the second pixel is a pixel other than the first pixel in the first training image. The acquisition module 901 is specifically used to perform position encoding on the first training image, obtain first feature information of the first training image, and determine the first value of each pixel in the first training image based on at least one mask and the first feature information.
[0168] Optionally, if the trained first machine learning model is applied to an image restoration scenario, the first training image is a low-resolution image, the second training image is a high-resolution image with the same image content as the first training image, or the second training image is an image obtained by denoising the high-resolution image at least once.
[0169] Optionally, the model training device also includes: a generation module 904, used to input the third training image into the first machine learning model, and generate predicted noise through the first machine learning model; a training module 905, used to train the first machine learning model according to the loss function to obtain the trained first machine learning model, and the loss function indicates the similarity between the predicted noise and the second noise.
[0170] Optionally, the generation module 904 is specifically used to input the third training image and the first information into the first machine learning model, and generate predicted noise through the first machine learning model.
[0171] Optionally, the generation module 904 is specifically used to: input the third training image and the first training image into the first machine learning model, and generate predicted noise through the first machine learning model, wherein the first training image is a low-resolution image; or, input the third training image and text description information into the first machine learning model, and generate predicted noise through the first machine learning model, wherein the first training image is a training image adapted to the text description information; or, input the third training image, the first training image and image style information into the first machine learning model, and generate predicted noise through the first machine learning model, and the image style information is the converted image style information corresponding to the first training image.
[0172] It should be noted that the information interaction, execution process, etc. between the modules / units in the model training device 900 are the same as those in the present application. Figures 1 to 8 The corresponding method embodiments are based on the same concept. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0173] See also Figure 10 , Figure 10 A structural schematic diagram of an image processing device provided in an embodiment of the present application, the image processing device 1000 includes: an acquisition module 1001, used to obtain first information corresponding to a first image, the first information indicating the category of pixels in the first image; a generation module 1002, used to input the first image, the first information, and the second information into a first machine learning model to obtain predicted noise generated by the first machine learning model; a denoising module 1003, used to denoise the second information using the predicted noise to obtain a second image, wherein the second information is pure noise, or the second information is an image obtained by denoising the pure noise at least once using the predicted noise generated by the first machine learning model, and the image content of the second image is the same as that of the first image.
[0174] Optionally, the acquisition module 1001 is specifically used to: input the first image into the second machine learning model to obtain at least one mask generated by the second machine learning model that corresponds one-to-one to at least one category, the at least one mask includes a first mask corresponding to the first category, the first category is any one of the at least one category, and the first mask indicates the pixels in the first image that are classified as the first category; obtain first information based on the at least one mask, wherein the first information includes a first value for each pixel in the first image, the first values corresponding to pixels of different categories in the first information are different, and the first values corresponding to pixels of the same category in the first information are the same.
[0175] It should be noted that the information interaction, execution process, etc. between the modules / units in the image processing device 1000 are the same as those in the present application. Figures 1 to 8 The corresponding method embodiments are based on the same concept. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0176] See also Figure 11 , Figure 11 Another structural schematic diagram of the image processing device provided in an embodiment of the present application, the image processing device 1100 includes: an input module 1101, used to input second information into the trained first machine learning model to obtain first predicted noise generated by the trained first machine learning model, and the training method of the first machine learning model is the method of the first aspect; a denoising module 1102, used to denoise the second information using the first predicted noise to obtain a second image, wherein the second information is pure noise, or the second information is an image obtained by denoising the pure noise at least once using the predicted noise generated by the first machine learning model.
[0177] Optionally, input module 1101 is specifically configured to input a first image and second information into a trained first machine learning model, where the first image is a low-resolution image. Input module 1101 is further configured to input the first image and second image into the trained first machine learning model to obtain a second predicted noise generated by the trained first machine learning model. Denoising module 1102 is further configured to denoise the second image using the second predicted noise to obtain a third image, where the third image is a high-resolution image with the same image content as the first image.
[0178] It should be noted that the information interaction, execution process, etc. between the modules / units in the image processing device 1100 are the same as those in the present application. Figures 1 to 8 The corresponding method embodiments are based on the same concept. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0179] Next, we will introduce an execution device provided by the embodiment of the present application. Figure 12 , Figure 12 A schematic diagram of the structure of an execution device provided in an embodiment of the present application. Specifically, the execution device 1200 includes: a receiver 1201, a transmitter 1202, a processor 1203 and a memory 1204 (wherein the number of processors 1203 in the execution device 1200 can be one or more, Figure 12 (taking one processor as an example), the processor 1203 may include an application processor 12031 and a communication processor 12032. In some embodiments of the present application, the receiver 1201, the transmitter 1202, the processor 1203 and the memory 1204 may be connected via a bus or other means.
[0180] The memory 1204 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1203. A portion of the memory 1204 may also include non-volatile random access memory (NVRAM). The memory 1204 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.
[0181] Processor 1203 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.
[0182] The methods disclosed in the above embodiments of the present application can be applied to or implemented by the processor 1203. The processor 1203 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the processor 1203. The above processor 1203 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and can further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 1203 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present application can be directly implemented as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in memory 1204, and processor 1203 reads the information in memory 1204 and, in conjunction with its hardware, completes the steps of the above method.
[0183] Receiver 1201 can be used to receive input digital or character information and generate signal input related to executing device-related settings and function control. Transmitter 1202 can be used to output digital or character information through the first interface. Transmitter 1202 can also be used to send instructions to the disk pack through the first interface to modify data in the disk pack. Transmitter 1202 can also include a display device such as a display screen.
[0184] In the embodiment of the present application, the processor 1203 is used to execute Figures 1 to 8 The method executed by the execution device in the corresponding embodiment. It should be noted that the specific manner in which the application processor 12031 in the processor 1203 executes the above steps is the same as that in the present application. Figures 1 to 8 The corresponding method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figures 1 to 8 The corresponding method embodiments are the same. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0185] The present application also provides a training device. Figure 13 , Figure 13 A structural diagram of a training device provided in an embodiment of the present application. Specifically, the training device 1300 is implemented by one or more servers. The training device 1300 may have relatively large differences due to different configurations or performances. It may include one or more central processing units (CPUs) 1322 (for example, one or more processors) and memory 1332, and one or more storage media 1330 (for example, one or more mass storage devices) storing application programs 1342 or data 1344. Among them, the memory 1332 and the storage medium 1330 can be temporary storage or permanent storage. The program stored in the storage medium 1330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the training device. Furthermore, the central processing unit 1322 can be configured to communicate with the storage medium 1330 to execute a series of instruction operations in the storage medium 1330 on the training device 1300.
[0186] The training device 1300 may also include one or more power supplies 1326, one or more wired or wireless network interfaces 1350, one or more input and output interfaces 1358, and / or one or more operating systems 1341, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0187] In the embodiment of the present application, the central processing unit 1322 is used to execute Figures 1 to 8 The method executed by the training device in the corresponding embodiment. It should be noted that the specific manner in which the central processing unit 1322 executes the above steps is the same as that in the present application. Figures 1 to 8 The corresponding method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figures 1 to 8 The corresponding method embodiments are the same. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0188] The present application also provides a computer-readable storage medium in which a program for signal processing is stored. When the program is run on a computer, the computer executes the above-mentioned Figures 1 to 8 The steps executed by the execution device in the method described in the embodiment shown, or the computer is caused to execute the steps as described above Figures 1 to 8 The illustrated embodiment describes the steps performed by the training device in the method.
[0189] The present application also provides a computer program product including a program, which, when the program is run on a computer, enables the computer to execute the above-mentioned Figures 1 to 8 The steps executed by the execution device in the method described in the embodiment shown, or the computer is caused to execute the steps as described above Figures 1 to 8 The illustrated embodiment describes the steps performed by the training device in the method.
[0190] The execution device and training device provided in the embodiment of the present application can be specifically a chip, which includes: a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin or a circuit. The processing unit can execute the computer execution instructions stored in the storage unit to enable the chip to execute the above Figures 1 to 8 Method 1 described in the illustrated embodiment. Optionally, the storage unit is a storage unit within the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM), etc.
[0191] For details, please refer to Figure 14 , Figure 14 A schematic diagram of the structure of a chip provided in an embodiment of the present application, which can be represented as a neural network processor NPU 140. NPU 140 is mounted on the host CPU as a coprocessor and is assigned tasks by the host CPU. The core of the NPU is arithmetic circuit 1403, which is controlled by controller 1404 to extract matrix data from memory and perform multiplication operations.
[0192] In some implementations, the arithmetic circuit 1403 includes multiple processing units (PEs). In some implementations, the arithmetic circuit 1403 is a two-dimensional systolic array. The arithmetic circuit 1403 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1403 is a general-purpose matrix processor.
[0193] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 1402 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 1401 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 1408.
[0194] Unified memory 1406 is used to store input and output data. Weight data is directly transferred to weight memory 1402 through the Direct Memory Access Controller (DMAC) 1405. Input data is also transferred to unified memory 1406 through the DMAC.
[0195] BIU stands for Bus Interface Unit, i.e., bus interface unit 1410 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 1409 .
[0196] The bus interface unit 1410 (BIU) is used for the instruction fetch memory 1409 to obtain instructions from the external memory, and is also used for the storage unit access controller 1405 to obtain the original data of the input matrix A or the weight matrix B from the external memory.
[0197] DMAC is mainly used to transfer input data in the external memory DDR to the unified memory 1406 or transfer weight data to the weight memory 1402 or transfer input data to the input memory 1401.
[0198] The vector calculation unit 1407 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.
[0199] In some implementations, the vector calculation unit 1407 can store the processed output vector to the unified memory 1406. For example, the vector calculation unit 1407 can apply a linear function and / or a nonlinear function to the output of the operation circuit 1403, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values to generate an activation value. In some implementations, the vector calculation unit 1407 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1403, for example, for use in a subsequent layer in a neural network.
[0200] An instruction fetch buffer 1409 connected to the controller 1404 is used to store instructions used by the controller 1404;
[0201] Unified memory 1406, input memory 1401, weight memory 1402, and instruction fetch memory 1409 are all on-chip memories. External memories are private to the NPU hardware architecture.
[0202] Among them, the operations of each layer in the machine learning model shown in the above embodiment can be performed by the operation circuit 1403 or the vector calculation unit 1407.
[0203] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the above-mentioned first aspect method.
[0204] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0205] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.
[0206] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0207] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
Claims
1. A model training method, characterized in that: The method comprises: Acquire first information and first noise, where the first information indicates a category of a pixel in a first training image; fusing the first information and the first noise to obtain second noise; The second training image is denoised using the second noise to obtain a third training image, wherein the first training image and the second training image have the same image content, and the third training image is used to train a first machine learning model, which is a model for predicting noise carried in an image.
2. The method according to claim 1, characterized in that The method further comprises: Inputting the first training image into a second machine learning model, obtaining at least one mask generated by the second machine learning model and corresponding to at least one category, wherein the at least one mask includes a first mask corresponding to a first category, the first category being any one of the at least one category, and the first mask indicating pixels in the first training image that are classified as the first category; The first information is obtained based on the at least one mask, wherein the first information includes a first value of each pixel in the first training image, the first values corresponding to pixels of different categories in the first information are different, and the first values corresponding to pixels of the same category in the first information are the same.
3. The method according to claim 2, characterized in that The first mask includes a second value of each pixel in the first training image, the second value corresponding to a first pixel in the first training image is 1, and the second value corresponding to a second pixel in the first training image is 0, the first pixel is a pixel in the first training image whose category is the first category, and the second pixel is a pixel other than the first pixel in the first training image, and obtaining the first information based on the at least one mask includes: performing position encoding on the first training image to obtain first feature information of the first training image; Determine the first value of each pixel in the first training image according to the at least one mask and the first feature information.
4. The method according to any one of claims 1 to 3, characterized in that If the trained first machine learning model is applied to an image restoration scenario, the first training image is a low-resolution image, the second training image is a high-resolution image with the same image content as the first training image, or the second training image is an image obtained by denoising the high-resolution image at least once.
5. The method according to any one of claims 1 to 3, characterized in that The method further comprises: Inputting the third training image into a first machine learning model to generate prediction noise through the first machine learning model; The first machine learning model is trained according to a loss function to obtain a trained first machine learning model, wherein the loss function indicates a similarity between the predicted noise and the second noise.
6. The method according to claim 5, characterized in that Inputting the third training image into the first machine learning model and generating predicted noise through the first machine learning model includes: The third training image and the first information are input into the first machine learning model, and the predicted noise is generated by the first machine learning model.
7. The method according to claim 5, characterized in that Inputting the third training image into the first machine learning model and generating predicted noise through the first machine learning model includes: Inputting the third training image and the first training image into the first machine learning model, and generating prediction noise through the first machine learning model, wherein the first training image is a low-resolution image; or Inputting the third training image and text description information into the first machine learning model, and generating prediction noise through the first machine learning model, wherein the first training image is an image adapted to the text description information; or The third training image, the first training image and the image style information are input into the first machine learning model, and prediction noise is generated by the first machine learning model, where the image style information is the converted image style information corresponding to the first training image.
8. An image processing method, characterized in that: The method comprises: Acquire first information corresponding to a first image, where the first information indicates categories of pixels in the first image; Inputting the first image, the first information, and the second information into a first machine learning model to obtain predicted noise generated by the first machine learning model; A second image is obtained by denoising the second information using the predicted noise, wherein the second information is pure noise, or the second information is an image obtained by denoising pure noise at least once using the predicted noise generated by the first machine learning model, and the second image has the same image content as the first image.
9. The method according to claim 8, characterized in that The acquiring first information corresponding to the first image includes: Inputting the first image into a second machine learning model to obtain at least one mask generated by the second machine learning model and corresponding to at least one category, wherein the at least one mask includes a first mask corresponding to a first category, the first category being any one of the at least one category, and the first mask indicating pixels in the first image that are classified as the first category; The first information is obtained based on the at least one mask, wherein the first information includes a first value of each pixel in the first image, the first values corresponding to pixels of different categories in the first information are different, and the first values corresponding to pixels of the same category in the first information are the same.
10. An image processing method, characterized in that: The method comprises: Inputting the second information into the trained first machine learning model to obtain a first prediction noise generated by the trained first machine learning model, wherein the training method of the first machine learning model is any one of the methods described in weights 1 to 7; A second image is obtained by denoising the second information using the first predicted noise, wherein the second information is pure noise, or the second information is an image obtained by denoising pure noise at least once using the predicted noise generated by the first machine learning model.
11. The method according to claim 10, characterized in that Inputting the second information into the trained first machine learning model includes: inputting a first image and the second information into the trained first machine learning model, wherein the first image is a low-resolution image; The method further comprises: Inputting the first image and the second image into the trained first machine learning model to obtain a second predicted noise generated by the trained first machine learning model; The second image is denoised using the second predicted noise to obtain a third image, where the third image is a high-resolution image having the same image content as the first image.
12. A model training device, characterized in that: The device comprises: an acquisition module, configured to acquire first information and first noise, wherein the first information indicates a category of pixels in a first training image; a fusion module, configured to fuse the first information and the first noise to obtain a second noise; A denoising module is configured to denoise the second training image using the second noise to obtain a third training image, wherein the first training image and the second training image have the same image content, and the third training image is used to train a first machine learning model, which is a model for predicting noise carried in an image.
13. An image processing device, characterized in that: The device comprises: an acquisition module, configured to acquire first information corresponding to a first image, where the first information indicates a category of pixels in the first image; an input module, configured to input the first image, the first information, and the second information into a first machine learning model to obtain predicted noise generated by the first machine learning model; A denoising module is configured to denoise the second information using the predicted noise to obtain a second image, wherein the second information is pure noise, or the second information is an image obtained by denoising pure noise at least once using the predicted noise generated by the first machine learning model, and the second image has the same image content as the first image.
14. An image processing device, characterized in that: The device comprises: An input module, configured to input the second information into a trained first machine learning model to obtain a first predicted noise generated by the trained first machine learning model, wherein the training method of the first machine learning model is any one of the methods described in any one of weights 1 to 7; A denoising module is used to denoise the second information using the first predicted noise to obtain a second image, wherein the second information is pure noise, or the second information is an image obtained by denoising pure noise at least once using the predicted noise generated by the first machine learning model.
15. A training device, characterized in that comprising a processor and a memory, the processor being coupled to the memory, The memory is used to store programs; The processor is configured to execute the program in the memory so that the training device performs the method according to any one of claims 1 to 7.
16. An execution device, characterized in that: comprising a processor and a memory, the processor being coupled to the memory, The memory is used to store programs; The processor is configured to execute the program in the memory, so that the execution device executes the method according to any one of claims 8 to 11.
17. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program, and when the program is run on a computer, the computer is caused to execute the method according to any one of claims 1 to 11.
18. A computer program product, characterized in that The computer program product comprises a program, which, when run on a computer, causes the computer to perform the method according to any one of claims 1 to 11.