An image processing method, a training method of a neural network, and related devices
By introducing LIF module into the pulsed neural network for feature extraction, the problem that pulsed neural network is difficult to deal with general visual tasks is solved, and more efficient and accurate feature extraction is achieved.
Patent Information
- Application Number
- CN202210302717.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-03-25
AI Technical Summary
Pulse neural networks are difficult to directly apply to perform mainstream general visual tasks because they are mainly suitable for processing sparse data.
Feature extraction of a single image through the LIF module, the feature information of the image block is divided into at least two groups, and the LIF module is input in sequence to realize the leakage and accumulation process, generate target data, and then update the feature information of the image block.
It realizes feature extraction of images through the LIF module, expands its application range to mainstream general visual tasks, and improves the efficiency and accuracy of feature extraction.
Smart Images

Figure CN114821096B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular, to an image processing method, a training method for a neural network, and related devices. Background Art
[0002] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. Using artificial intelligence for image processing is a common application method of artificial intelligence.
[0003] Currently, the spiking neural network (SNN), as a bionic neural network, has received extensive attention in recent years. The leaky integrate and fire (LIF) module in the spiking neural network has the advantages of fast and efficient calculation.
[0004] However, the spiking neural network is mainly applied to process sparse data. For example, multiple pictures collected by a dynamic vision sensor are processed through the spiking neural network, but the spiking neural network cannot be directly applied to perform mainstream general vision tasks. Summary of the Invention
[0005] Embodiments of this application provide an image processing method, a training method for a neural network, and related devices, which realize feature extraction of a single image through the LIF module, and further enable the LIF module to be applied to perform mainstream general vision tasks.
[0006] To solve the above technical problems, the embodiments of this application provide the following technical solutions:
[0007] In a first aspect, an embodiment of the present application provides an image processing method, which can apply artificial intelligence technology to the field of image processing. The method includes: an execution device inputs an image to be processed into a first neural network, and the first neural network extracts features from the image to be processed to obtain feature information of the image to be processed. Among them, the execution device extracts features from the image to be processed through the first neural network, including: the execution device obtains first feature information corresponding to the image to be processed. The image to be processed includes multiple image blocks, and the first feature information includes feature information of multiple image blocks in the image to be processed. The first feature information is also the feature information of the image to be processed; the execution device sequentially inputs the feature information of at least two groups of image blocks into the LIF module to obtain target data generated by the LIF module. The feature information of a group of image blocks includes the feature information of at least one image block; the execution device obtains second feature information corresponding to the image to be processed according to the target data. The second feature information includes the updated feature information of the image blocks, and the second feature information is the updated feature information of the image to be processed.
[0008] In this implementation, the feature information of the entire image to be processed is divided into the feature information of multiple image blocks in the image to be processed. The feature information of multiple image blocks can be divided into the feature information of at least two groups of image blocks. The feature information of at least two groups of image blocks is sequentially input into the LIF module to implement the leakage and accumulation process of the LIF module, obtain the target data generated by the LIF module, and then obtain the updated feature information of the image to be processed according to the target data; through the foregoing method, the LIF module can be used to extract features from a single image, and then the LIF module can be applied to perform mainstream general vision tasks, which is beneficial to improving the efficiency and accuracy of the feature extraction process.
[0009] In a possible implementation manner of the first aspect, the execution device sequentially inputs the feature information of at least two groups of image blocks into the LIF module to obtain the target data generated by the LIF module, including: the execution device sequentially inputs the feature information of at least two groups of image blocks into the LIF module, and when the excitation condition of the LIF module is satisfied, the activation function is used to generate the target data.
[0010] Among them, the target data is not binary data, that is, the output of the LIF module may not be pulse data, that is, the target data output by the LIF module may no longer be two fixed values, but higher-precision data; as an example, for example, the target data may be floating-point data. Optionally, the precision of the target data and the feature information of the image blocks may be the same, that is, the numerical bit levels of the target data and the feature information of the image blocks may be the same.
[0011] In this implementation manner, the data output by the LIF module is non-binarized data, that is, the accuracy of the target data output by the LIF module is improved, so that more abundant feature information of the image to be processed can be extracted. Therefore, during the process of feature extraction of the image to be processed, both the advantages of fast and effective calculation of the LIF module are retained, and more abundant feature information can be obtained.
[0012] In a possible implementation manner of the first aspect, the execution device sequentially inputs the feature information of at least two groups of image blocks into the LIF module, including: the execution device inputs the feature information of at least two groups of image blocks into the LIF module in multiple rounds; further, in each round, the execution device inputs the feature information of a group of image blocks into one LIF module. Optionally, the first neural network may include M parallel LIF modules. Then, in each round, the execution device can simultaneously input the feature information of M groups of image blocks into M parallel LIF modules respectively, and the input data is processed by the M parallel LIF modules respectively.
[0013] In a possible implementation manner of the first aspect, the feature information of at least two groups of image blocks includes the feature information of multiple rows of image blocks. The feature information of each row of image blocks includes the feature information of multiple image blocks in the same row, and the feature information of each group of image blocks includes the feature information of at least one row of image blocks. And / or, the feature information of at least two groups of image blocks includes the feature information of multiple columns of image blocks. The feature information of each column of image blocks includes the feature information of multiple image blocks in the same column, and the feature information of each group of image blocks includes the feature information of at least one column of image blocks.
[0014] In a possible implementation manner of the first aspect, the firing condition of the LIF module may include whether the value of a membrane potential in the LIF module is greater than or equal to a preset threshold. Further, since the feature information of the image block may include the feature information corresponding to the image block and at least one channel, correspondingly, the firing condition of the LIF module may include one or more thresholds, that is, the threshold values corresponding to different channels may be the same or different.
[0015] In a possible implementation manner of the first aspect, the first neural network is a multi-layer perceptron MLP, a convolutional neural network, or a neural network using a self-attention mechanism. The neural network using a self-attention mechanism can also be called a Transformer neural network.
[0016] In this implementation manner, regardless of whether the first neural network is an MLP, a convolutional neural network, or a residual Transformer neural network, it can be compatible with the LIF module through the image processing method provided in the embodiments of the present application. Since the MLP, convolutional neural network, and residual Transformer neural network can be applied to different application scenarios, the application scenarios and implementation flexibility of this solution are greatly expanded.
[0017] In a possible implementation manner of the first aspect, the method further includes: the execution device performs feature processing on the feature information of the image to be processed through a second neural network to obtain a prediction result corresponding to the image to be processed, where the first neural network and the second neural network are included in the same target neural network, and the task performed by the target neural network is any one of the following: image classification, image segmentation, object detection on an image, or super-resolution processing of an image. In the embodiments of the present application, multiple application scenarios of this solution are provided, greatly expanding the implementation flexibility of this solution.
[0018] In a second aspect, the embodiments of the present application provide a training method for a neural network, which can apply artificial intelligence technology to the field of image processing. The method includes: inputting an image to be processed into a first neural network, performing feature extraction on the image to be processed through the first neural network to obtain feature information of the image to be processed, performing feature processing on the feature information of the image to be processed through a second neural network to obtain a prediction result corresponding to the image to be processed; training the first neural network and the second neural network using a loss function according to the prediction result corresponding to the image to be processed and the correct result, where the loss function indicates the similarity between the prediction result and the correct result.
[0019] Among them, performing feature extraction on the image to be processed through the first neural network includes: obtaining first feature information corresponding to the image to be processed, where the image to be processed includes multiple image blocks, and the first feature information includes the feature information of the image blocks; sequentially inputting the feature information of at least two groups of image blocks into a leaky-integrate-and-fire (LIF) module to obtain target data generated by the LIF module, where the feature information of a group of image blocks includes the feature information of at least one image block; obtaining second feature information corresponding to the image to be processed according to the target data, where the second feature information includes the updated feature information of the image blocks, and both the first feature information and the second feature information are the feature information of the image to be processed.
[0020] In the second aspect of the present application, the training device is further configured to execute the steps performed by the execution device in each possible implementation manner of the first aspect. The specific implementation manners of the steps, the meanings of the terms, and the beneficial effects brought in each possible implementation manner of the second aspect of the present application can all be referred to in the first aspect and will not be elaborated here.
[0021] In a third aspect, an embodiment of the present application provides an image processing apparatus, which can apply artificial intelligence technology to the field of image processing. The image processing apparatus includes: an input unit for inputting an image to be processed into a first neural network; and a feature extraction unit for extracting features of the image to be processed through the first neural network to obtain feature information of the image to be processed.
[0022] Among them, the feature extraction unit includes: an acquisition subunit for acquiring first feature information corresponding to the image to be processed, the image to be processed includes a plurality of image blocks, and the first feature information includes feature information of the image blocks; a generation subunit for sequentially inputting feature information of at least two groups of image blocks into a leaky-integrate-and-fire (LIF) module to obtain target data generated by the LIF module, the feature information of a group of image blocks includes feature information of at least one image block; and an acquisition subunit for acquiring second feature information corresponding to the image to be processed according to the target data, the second feature information includes updated feature information of the image blocks, and both the first feature information and the second feature information are feature information of the image to be processed.
[0023] In the third aspect of the present application, the image processing apparatus is further configured to execute the steps performed by the execution device in each possible implementation manner of the first aspect. For the specific implementation manners of the steps, the meanings of the terms, and the beneficial effects brought about in each possible implementation manner of the third aspect of the present application, reference may be made to the first aspect and each possible implementation manner of the first aspect, which will not be elaborated herein.
[0024] In a fourth aspect, an embodiment of the present application provides a neural network training apparatus, which can apply artificial intelligence technology to the field of image processing. The neural network training apparatus includes: a feature extraction unit for inputting an image to be processed into a first neural network, and extracting features of the image to be processed through the first neural network to obtain feature information of the image to be processed; a feature processing unit for performing feature processing on the feature information of the image to be processed through a second neural network to obtain a prediction result corresponding to the image to be processed; and a training unit for training the first neural network and the second neural network according to the prediction result and the correct result corresponding to the image to be processed by using a loss function, where the loss function indicates the similarity between the prediction result and the correct result.
[0025] Among them, the feature extraction unit includes: an acquisition subunit, configured to acquire first feature information corresponding to an image to be processed, the image to be processed includes a plurality of image patches, and the first feature information includes the feature information of the image patches; a generation subunit, configured to sequentially input the feature information of at least two groups of image patches into a leaky-integrate-and-fire (LIF) module to obtain target data generated by the LIF module, the feature information of a group of image patches includes the feature information of at least one image patch; the acquisition subunit is further configured to acquire second feature information corresponding to the image to be processed according to the target data, the second feature information includes the updated feature information of the image patches, and both the first feature information and the second feature information are the feature information of the image to be processed.
[0026] In the fourth aspect of this application, the image processing device is further configured to execute the steps performed by the execution device in each possible implementation manner of the second aspect. For the specific implementation manners of the steps, the meanings of the terms, and the beneficial effects brought about in each possible implementation manner of the fourth aspect of this application, reference may be made to the second aspect and each possible implementation manner of the second aspect, which will not be elaborated here.
[0027] In a fifth aspect, an embodiment of this application provides a computer program product, which includes a program that, when running on a computer, causes the computer to execute the method described in the first aspect or the second aspect above.
[0028] In a sixth aspect, an embodiment of this application provides a computer-readable storage medium, in which a computer program is stored, and when the program runs on a computer, it causes the computer to execute the method described in the first aspect or the second aspect above.
[0029] In a seventh aspect, an embodiment of this application provides an execution device, including a processor and a memory, the processor is coupled to the memory, the memory is configured to store a program; the processor is configured to execute the program in the memory, so that the execution device executes the image processing method described in the first aspect above.
[0030] In an eighth aspect, an embodiment of this application provides a training device, including a processor and a memory, the processor is coupled to the memory, the memory is configured to store a program; the processor is configured to execute the program in the memory, so that the training device executes the training processing method of the neural network described in the second aspect above.
[0031] In a ninth aspect, an embodiment of the present application provides a chip system, which includes a processor for supporting a terminal device or a communication device to implement the functions involved in the above aspects. For example, it can send or process the data and / or information involved in the above methods. In a possible design, the chip system further includes a memory for storing necessary program instructions and data of the terminal device or the communication device. The chip system can be composed of chips or include chips and other discrete devices. Description of the Drawings
[0032] Figure 1a It is a schematic structural diagram of an artificial intelligence main body framework provided by an embodiment of the present application;
[0033] Figure 1b It is an application scenario diagram of an image processing method provided by an embodiment of the present application;
[0034] Figure 2a It is a system architecture diagram of an image processing system provided by an embodiment of the present application;
[0035] Figure 2b It is a schematic flow diagram of a graphics processing method provided by an embodiment of the present application;
[0036] Figure 3 It is a schematic flow diagram of an image processing method provided by an embodiment of the present application;
[0037] Figure 4 It is a schematic structural diagram of a first neural network in an image processing method provided by an embodiment of the present application;
[0038] Figure 5 It is a schematic diagram of the feature information of multiple image blocks in an image processing method provided by an embodiment of the present application;
[0039] Figure 6 It is a schematic diagram of a LIF unit in a first neural network in an image processing method provided by an embodiment of the present application;
[0040] Figure 7 It is a schematic diagram of the feature information of a group of image blocks in an image processing method provided by an embodiment of the present application;
[0041] Figure 8 It is a schematic diagram of the feature information of a group of image blocks in an image processing method provided by an embodiment of the present application;
[0042] Figure 9 It is a schematic diagram of sequentially inputting the feature information of at least two groups of image blocks into a LIF module in an image processing method provided by an embodiment of the present application;
[0043] Figure 10A schematic diagram of inputting the feature information of at least two groups of image blocks into the LIF module in the image processing method provided by the embodiments of the present application;
[0044] Figure 11 A schematic flowchart of the training method of the neural network provided by the embodiments of the present application;
[0045] Figure 12 A schematic structural diagram of the image processing device provided by the embodiments of the present application;
[0046] Figure 13 A schematic structural diagram of the training device of the neural network provided by the embodiments of the present application;
[0047] Figure 14 A schematic structural diagram of the execution device provided by the embodiments of the present application;
[0048] Figure 15 A schematic structural diagram of the training device provided by the embodiments of the present application;
[0049] Figure 16 A schematic structural diagram of the chip provided by the embodiments of the present application. Detailed implementation manners
[0050] The embodiments of the present application will be described below with reference to the accompanying drawings. Those of ordinary skill in the art can understand that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0051] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of the present application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.
[0052] First, the overall working process of the artificial intelligence system will be described. Please refer to Figure 1a , Figure 1aShown is a schematic structural diagram of an artificial intelligence subject framework. The above artificial intelligence subject framework will be elaborated from two dimensions: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). Among them, the "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be the general processes of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes the refinement process of "data - information - knowledge - wisdom". The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (provision and processing technology implementation) to the industrial ecological process of the system.
[0053] (1) Infrastructure
[0054] The infrastructure provides computing power support for the artificial intelligence system, enables communication with the external world, and is supported through the basic platform. It communicates with the external world through sensors; the computing power is provided by intelligent chips, which can specifically adopt hardware acceleration chips such as central processing unit (CPU), neural-network processing unit (NPU), graphics processing unit (GPU), application specific integrated circuit (ASIC), or field programmable gate array (FPGA); the basic platform includes relevant platform guarantees and supports such as distributed computing frameworks and networks, and can include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the external world to obtain data, and these data are provided to the intelligent chips in the distributed computing system provided by the basic platform for computing.
[0055] (2) Data
[0056] The data in the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voices, texts, and also involves the Internet of Things data of traditional devices, including the business data of existing systems and the sensed data such as force, displacement, liquid level, temperature, humidity, etc.
[0057] (3) Data Processing
[0058] Data processing usually includes data training, machine learning, deep learning, search, reasoning, decision-making and other methods.
[0059] Among them, machine learning and deep learning can perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on data.
[0060] Inference refers to the process of simulating the intelligent reasoning method of humans in a computer or intelligent system, and using formalized information for machine thinking and problem-solving according to the inference control strategy. The typical functions are search and matching.
[0061] Decision-making refers to the process of making decisions after intelligent information is inferred, and usually provides functions such as classification, sorting, prediction, etc.
[0062] (4) General capabilities
[0063] After the data is processed as mentioned above, some general capabilities can be further formed based on the results of the data processing, such as algorithms or a general system. For example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0064] (5) Intelligent products and industry applications
[0065] Intelligent products and industry applications refer to the products and applications of artificial intelligence systems in various fields, which are the encapsulation of the overall artificial intelligence solution, productize intelligent information decision-making, and achieve landing applications. Its application fields mainly include: intelligent terminals, intelligent manufacturing, intelligent transportation, smart homes, intelligent healthcare, intelligent security, autonomous driving, smart cities, etc.
[0066] The embodiments of this application can be applied to various application fields in the field of artificial intelligence. Specifically, it can be applied to tasks of performing image processing in various application fields. The foregoing tasks include, but are not limited to: feature extraction of images, image classification, image segmentation, object detection of images, super-resolution processing of images, or other types of tasks, etc., and will not be enumerated here.
[0067] As an example, in intelligent terminals, smart homes, intelligent security, or other application fields, there may be a need to use neural networks for image classification. For a more intuitive understanding of this solution, please refer to Figure 1b , Figure 1b which is an application scenario diagram of the image processing method provided by the embodiments of this application. Figure 1b Taking the application of the target neural network in the field of intelligent terminals as an example, in the training stage, the training device uses the training data set to iteratively train the neural network for image classification. In each training process, the training device backpropagates the gradient value to update the weight parameters of the foregoing neural network. After the training operation on the target neural network is completed, the trained neural network can be sent to the mobile device side to perform graphic classification through this neural network. It should be understood that Figure 1bThe examples herein are only for facilitating the understanding of this solution and are not used to limit this solution.
[0068] As another example, for instance, in the field of autonomous driving, there may be a need to perform object detection on the collected images using a neural network, that is, inputting the images collected by an autonomous driving vehicle into the neural network to obtain the category and location of at least one object in the to-be-processed image output by the neural network.
[0069] As another example, for instance, in the field of intelligent terminals, in an image retouching application, a function of performing image segmentation on the input image may be provided, that is, inputting the to-be-processed image into the neural network to obtain the category of each pixel in the to-be-processed image output by the neural network, and the category of each pixel is foreground or background.
[0070] As another example, for instance, in the fields of intelligent security, smart city, etc., there may be a need to perform super-resolution processing on the collected images, that is, inputting the images collected by a monitoring device into the neural network to obtain the processed image output by the neural network, and the resolution of the processed image is higher.
[0071] As another example, for instance, in the fields of intelligent terminals and intelligent security, there may be a need for face recognition based on the collected images, then it is necessary to use a neural network to extract features from the collected images of users, and then the extracted feature information can be matched with the pre-registered feature information to determine whether the current user is a registered user, etc.
[0072] It should be noted that the graphic processing method provided in the embodiments of this application can also be applied to other application scenarios, which are not enumerated here. In the above various application scenarios, during the process of using a neural network to process images, it is necessary to first extract features from the input images. In order to be able to apply the LIF module to the feature extraction process of a single image, the embodiments of this application provide an image processing method.
[0073] The following first combines Figure 2a to introduce the image processing system of the embodiments of this application. Figure 2a is a system architecture diagram of the image processing system provided in the embodiments of this application. In Figure 2a the image processing system 200 includes a training device 210, a database 220, an execution device 230, and a data storage system 240, and the execution device 230 includes a computing module 231.
[0074] Among them, a training data set is stored in the database 220. The training device 210 generates the first model / rule 201, and iteratively trains the first model / rule 201 by using the training data set to obtain the trained first model / rule 201, and deploys the trained first model / rule 201 to the computing module 231 of the execution device 230. The first model / rule 201 may specifically be embodied as a neural network or a non-neural network model. In the embodiments of the present application, only the case where the first model / rule 201 is embodied as a neural network is taken as an example for illustration.
[0075] The execution device 230 may specifically be embodied as different systems or devices, such as a mobile phone, a tablet computer, a laptop computer, a virtual reality (VR) device, a monitoring system, and so on. Among them, the execution device 230 may call data, code, etc. in the data storage system 240, or store data, instructions, etc. in the data storage system 240. The data storage system 240 may be disposed in the execution device 230, or the data storage system 240 may be an external memory relative to the execution device 230.
[0076] In some embodiments of the present application, please refer to Figure 2a , the execution device 230 can directly interact with the "user". It should be noted that Figure 2a is only a schematic diagram of the architectures of two image processing systems provided by the embodiments of the present invention. The positional relationships among the devices, components, modules, etc. shown in the figure do not constitute any limitation. For example, in some other embodiments of the present application, the execution device 230 and the client device may be separate independent devices. The execution device 230 is configured with an input / output (I / O) interface to perform data interaction with the client device. The "user" may input an image to be processed to the I / O interface through the client device, and the execution device 230 returns the processing result to the client device through the I / O interface for the user.
[0077] Combined with the above description, please refer to Figure 2b , Figure 2bA schematic flowchart of a graphic processing method provided by an embodiment of the present application. A1. The execution device inputs the image to be processed into the first neural network. A2. The execution device extracts features from the image to be processed through the first neural network to obtain the feature information of the image to be processed. Wherein, step A2 may include: 201. The execution device obtains the first feature information corresponding to the image to be processed. The image to be processed includes a plurality of image blocks, and the first feature information includes the feature information of the image blocks. 202. The execution device sequentially inputs the feature information of at least two groups of image blocks into the LIF module to obtain the target data generated by the LIF module. The feature information of a group of image blocks includes the feature information of at least one image block. 203. The execution device obtains the second feature information corresponding to the image to be processed according to the target data. The second feature information includes the updated feature information of the image blocks. Both the first feature information and the second feature information are the feature information of the image to be processed.
[0078] In this implementation, the feature information of at least two groups of image blocks is sequentially input into the LIF module to implement the leakage and accumulation process of the LIF module, obtain the target data generated by the LIF module, and then obtain the updated feature information of the image to be processed according to the target data. Through the foregoing method, feature extraction of a single image can be implemented through the LIF module, and then the LIF module can be applied to execute mainstream general vision tasks.
[0079] Combined with the above description, the specific implementation processes of the inference stage and the training stage of the image processing method provided by the embodiment of the present application will be described below.
[0080] I. Inference stage
[0081] In the embodiment of the present application, specifically, please refer to Figure 3 , Figure 3 A schematic flowchart of an image processing method provided by an embodiment of the present application. The image processing method provided by the embodiment of the present application may include:
[0082] 301. The execution device inputs the image to be processed into the first neural network.
[0083] In the embodiment of the present application, after obtaining the image to be processed, the execution device may input the image to be processed into the first neural network, and extract features from the image to be processed through the first neural network to obtain the feature information of the image to be processed.
[0084] Among them, the first neural network can be specifically expressed as a multilayer perceptron (MLP), a convolutional neural network (CNN), a neural network using a self-attention mechanism, or other types of neural networks. The neural network using a self-attention mechanism can also be called a Transformer neural network. The specific method can be flexibly determined in combination with the actual application scenario and is not limited here.
[0085] In order to understand this solution more intuitively, the following is combined with Figure 4 The overall architecture of the first neural network is explained. Figure 4 A schematic diagram of the structure of the first neural network in the image processing method provided in the embodiment of the present application is shown in FIG. Figure 4 As shown, the first neural network may include a segmentation unit and a LIF unit. Optionally, the first neural network may further include a channel mixing unit and an upsampling / downsampling unit.
[0086] The segmentation unit in the first neural network is used to extract features and segment the image to be processed to obtain the initial feature information (embedding) of the multiple image blocks (patches) included in the image to be processed. Since the image to be processed is composed of multiple image blocks, the feature information of the multiple image blocks is also the feature information of the image to be processed. Among them, the aforementioned segmentation operation is used to divide the image to be processed into multiple image blocks, and the execution order of the aforementioned feature extraction operation and segmentation operation can be flexibly determined according to the actual application scenario. It should be noted that Figure 4 The feature information of the multiple image blocks shown is only for the convenience of understanding the relationship between the multiple image blocks and the image to be processed. In actual situations, the feature information of the multiple image blocks is in the form of data.
[0087] The LIF unit in the first neural network is used to update the feature information of the image block. The aforementioned LIF unit at least includes the LIF module in the embodiment of the present application. The aforementioned LIF unit may also include other neural network layers. The specific implementation process of the LIF unit will be introduced in detail through the following steps 302 to 304.
[0088] The channel mixing unit in the first neural network is also used to update the feature information of the image block. The up-sampling unit and the down-sampling unit are both used to change the size of the feature information of the image to be processed; wherein the up-sampling unit is used to perform an up-sampling operation on the feature information of the image block to amplify the feature information of the image block; and the down-sampling unit is used to perform a down-sampling operation on the feature information of the image block to reduce the feature information of the image block.
[0089] It should be noted that in practical applications, the first neural network may include more or fewer units, the positions of the LIF units and the channel mixing units can be adjusted, and the numbers of the LIF units, the channel mixing units, and the upsampling / downsampling units can be the same or different, as long as there are LIF units in the first neural network. Figure 4 The first neural network shown is only an example for facilitating the understanding of this solution and is not used to limit this solution.
[0090] 302. The execution device obtains the first feature information corresponding to the image to be processed. The image to be processed includes multiple image blocks, and the first feature information includes the feature information of the image blocks.
[0091] In the embodiments of the present application, before the execution device updates the initial feature information of multiple image blocks through the LIF module, it may first obtain the first feature information corresponding to the image to be processed. Among them, the image to be processed includes multiple image blocks, and the first feature information includes the feature information of the foregoing multiple image blocks.
[0092] Furthermore, the first feature information may include the initial feature information of each image block or the updated feature information of each image block.
[0093] Optionally, after the execution device obtains the first feature information corresponding to the image to be processed, it may perform a convolution operation on the feature information of multiple image blocks again by using a convolutional neural network layer to update the feature information of the image blocks and obtain the updated first feature information. Among them, the convolutional neural network layer may specifically be a depthwise convolution layer or other types of convolutional neural network layers. When a depthwise convolution layer is selected, the computational amount of the foregoing convolution operation can be reduced.
[0094] 303. The execution device sequentially inputs the feature information of at least two groups of image blocks into the LIF module to obtain the target data generated by the LIF module.
[0095] In the embodiments of the present application, after the execution device obtains the feature information of multiple image blocks included in the image to be processed (that is, the first feature information or the updated first feature information), it may divide the feature information of multiple image blocks included in the image to be processed into the feature information of at least two groups of image blocks, and sequentially input the feature information of at least two groups of image blocks into the LIF module to implement the leakage and accumulation process of the LIF module and obtain the target data generated by the LIF module. The feature information of a group of image blocks includes the feature information of at least one image block.
[0096] Specifically, the execution device may sequentially input the feature information of at least two groups of image blocks into the LIF module, and when the excitation condition of the LIF module is satisfied, use the activation function to generate target data.
[0097] Among them, the target data may be binary data, that is, the output of the LIF module may be two preset values. Or, the target data may also be non-binary data, that is, the output of the LIF module may not be pulse data, that is, the target data output by the LIF module may no longer be two fixed values, but higher-precision data; as an example, for example, the target data may be floating-point data. Optionally, the precision of the target data and the feature information of the image block may be the same, that is, the numerical bit levels of the target data and the feature information of the image block may be the same.
[0098] In the embodiment of the present application, the output of the LIF module is non-binary data, that is, the precision of the target data output by the LIF module is improved, so that more abundant feature information of the image to be processed can be extracted. Then, in the process of feature extraction of the image to be processed, both the advantages of fast and effective calculation of the LIF module are retained, and more abundant feature information can be obtained.
[0099] Regarding the concept of the feature information of at least two groups of image blocks. The image to be processed may include two direction dimensions, horizontal and vertical. Correspondingly, the image to be processed may be sliced in these two direction dimensions of horizontal and vertical, that is, the feature information of multiple image blocks may include the feature information of multiple image blocks in the horizontal dimension and the feature information of multiple image blocks in the vertical dimension.
[0100] For a more intuitive understanding of this solution, please refer to Figure 5 , Figure 5 which is a schematic diagram of the feature information of multiple image blocks in the image processing method provided by the embodiment of the present application. Figure 5 One B1 in may represent the feature information of an image block, Figure 5 Taking the example that the feature information to be processed includes the feature information of 16 image blocks in, the feature information of an image block may include the feature information of at least one channel corresponding to an image block, Figure 5 Taking the example that the feature information of an image block includes the feature information of multiple channels corresponding to an image block in ; the information corresponding to different channels may be the same or different types of information. As an example, for example, one channel may be used to obtain any one of the following information: color, texture, brightness, or other information, etc., which is not limited here.
[0101] Such as Figure 5As shown, the feature information of the image to be processed may include the feature information of multiple image patches in the horizontal dimension and the feature information of multiple image patches in the vertical dimension, that is, it includes the feature information of multiple rows of image patches and the feature information of multiple columns of image patches. It should be understood that Figure 5 the examples in
[0102] are only for facilitating the understanding of the concept of the feature information of multiple image patches and are not used to limit this solution. The execution device can determine the feature information of one or more rows of image patches as the feature information of a group of image patches, or the execution device can also determine the feature information of one or more columns of image patches as the feature information of a group of image patches, so as to divide the feature information of multiple image patches into at least two groups of image patch feature information.
[0103] Regarding the process of sequentially inputting the feature information of at least two groups of image patches into the LIF module. There can be one or more LIF modules in a LIF unit of the first neural network. For a more intuitive understanding of this solution, please refer to Figures 6 to 8 Figure 6 which is a schematic diagram of a LIF unit in the first neural network in the image processing method provided by the embodiments of this application. Figure 6 Taking the first neural network specifically represented as an MLP as an example, as shown in the figure, a LIF unit can include multiple MLP layers, a depthwise separable convolutional layer, a vertical LIF module, and a horizontal LIF module.
[0104] Among them, the MLP layer refers to a neural network layer composed of at least one fully connected neuron; if the first neural network is specifically represented as a convolutional neural network, the MLP layer can be replaced by a convolutional neural network layer; if the first neural network is specifically represented as a Transformer neural network, the MLP layer can be replaced by a Transformer neural network layer. Further, the convolutional neural network layer refers to a neural network layer composed of at least one partially connected neuron, and the Transformer neural network layer refers to a neural network layer introducing an attention mechanism.
[0105] The feature information of each group of image patches obtained by the vertical LIF module includes the feature information of at least one row of image patches. For a more intuitive understanding of this solution, please refer to Figure 7 Figure 7 which is a schematic diagram of the feature information of a group of image patches in the image processing method provided by the embodiments of this application. As Figure 7 shown, C1, C2, C3, and C4 respectively represent the feature information of four groups of image patches, that is, the feature information of multiple image patches is divided into four groups in the vertical direction, and the feature information of each group of image patches includes the feature information of one row of image patches. It should be understood that Figure 7 the examples in
[0106] The feature information of each group of image patches obtained by the horizontal LIF module includes the feature information of at least one column of image patches, that is, the feature information of multiple image patches is grouped in the horizontal direction, and at least two groups obtained are sequentially input into the horizontal LIF module. For a more intuitive understanding of this solution, please refer to Figure 8 , Figure 8 FIG. Figure 8 As shown in the figure, D1, D2, D3, and D4 respectively represent the feature information of four groups of image patches, that is, the feature information of multiple image patches is divided into four groups in the horizontal direction, and the feature information of each group of image patches includes the feature information of one column of image patches. It should be understood that Figure 8 the example in
[0107] It should be noted that one LIF unit in the first neural network may include more or fewer neural network layers, Figure 6 the example in
[0108] Specifically, if the first neural network includes a vertical LIF module, the execution device can group the feature information of multiple image patches in the vertical direction, and sequentially input at least two groups of the obtained feature information of image patches into the vertical LIF module, that is, input the feature information of one group of image patches into the vertical LIF module each time.
[0109] After the execution device inputs at least one row of the feature information of image patches (that is, the feature information of one group of image patches) into the vertical LIF module each time, it will determine whether the firing condition of the vertical LIF module is satisfied. If the determination result is no, the vertical LIF module may not generate any value; if the determination result is yes, the vertical LIF module may generate target data using the activation function and reset a membrane potential of the vertical LIF module to 0. The execution device continues to input the next group of the feature information of image patches into the vertical LIF module to perform leakage and accumulation on the feature information of the two groups of image patches. The execution device repeats the foregoing operation to process all the feature information of the image patches through the vertical LIF module.
[0110] To further understand this solution, the following discloses a formula for the specific implementation manner of the LIF module:
[0111]
[0112] where τ represents the leakage parameter of the LIF module, which is a hyperparameter. When the value of t is greater than V h, The value of is less than or equal to V t When h is The value of is 0, V t h represents the excitation condition of the LIF module, represents the nth value in the feature information of a set of image patches input to the LIF module in the current round (i.e., the (t + 1)-th round), represents the membrane potential of the LIF module in the previous round (i.e., the t-th round), represents the membrane potential of the LIF module in the current round. When the excitation condition of the LIF module is satisfied, the LIF module generates The calculation formula of is as follows:
[0113]
[0114] where, and V t h can be understood with reference to the above description. ReLU is an example of an activation function. It should be understood that the examples in Equation (1) and Equation (2) are only for facilitating the understanding of this solution and are not used to limit this solution.
[0115] Furthermore, the data sizes of the feature information of the two sets of image patches are the same, that is, the values included in the feature information of the two sets of image patches can correspond one by one. Then, the LIF module can multiply the feature information of the image patches in the previous round by the leakage parameter and add it to the feature information of the image patches in the current round to obtain multiple target values included in the current round. represents the nth value among the multiple target values. When the value of is greater than the preset threshold, it is determined that satisfies the excitation condition of the LIF module, and the LIF module is excited to generate a target data using the activation function.
[0116] Even further, since the feature information of the image patches can include the feature information corresponding to the image patches and at least one channel, correspondingly, the excitation condition of the LIF module can include one or more thresholds; further, the threshold values corresponding to different channels can be the same or different.
[0117] For a more intuitive understanding of this solution, please refer to Figure 9 Figure 9 This is a schematic diagram of sequentially inputting the feature information of at least two sets of image patches into the LIF module in the image processing method provided by the embodiments of the present application. Figure 9 Taking the feature information of a group of image patches including the feature information of a row of image patches as an example. Among them, in the first round, the execution device can input the feature information of the first row of images (that is, the feature information of a group of image patches represented by C1) into the vertical LIF module; in the second round, the execution device can input the feature information of the first row of images (that is, the feature information of a group of image patches represented by C2) into the vertical LIF module; in the third round, the execution device can input the feature information of the first row of images (that is, the feature information of a group of image patches represented by C3) into the vertical LIF module; in the fourth round, the execution device can input the feature information of the first row of images (that is, the feature information of a group of image patches represented by C4) into the vertical LIF module, thus realizing the sequential input of the feature information of four groups of image patches into the LIF module. It should be understood that Figure 9 The examples in this are only for facilitating the understanding of this solution and are not used to limit this solution.
[0118] Optionally, one LIF unit of the first neural network may include M parallel vertical LIF modules. Then, in each round, the execution device can simultaneously input the feature information of M groups of image patches into M parallel vertical LIF modules respectively, and process the input data through the M parallel vertical LIF modules respectively.
[0119] Specifically, if the first neural network includes a horizontal LIF module, the execution device can group the feature information of multiple image patches in the horizontal direction, and sequentially input the obtained feature information of at least two groups of image patches into the horizontal LIF module, that is, input the feature information of one group of image patches into the horizontal LIF module each time.
[0120] After the execution device inputs the feature information of at least one column of image patches (that is, the feature information of one group of image patches) into the horizontal LIF module each time, it will judge whether the excitation condition of the horizontal LIF module is satisfied. If the judgment result is no, the horizontal LIF module may not generate any value; if the judgment result is yes, the horizontal LIF module can generate target data using the activation function and reset a membrane potential of the horizontal LIF module to 0. The execution device continues to input the next group of image patch feature information into the horizontal LIF module to perform leakage and accumulation on the feature information of two groups of image patches. The execution device repeats the foregoing operations to process all the feature information of the image patches through the horizontal LIF module.
[0121] It should be noted that the specific processing methods of the "vertical LIF module" and the "horizontal LIF module" for the input data are similar. For the specific implementation method of the horizontal LIF module, reference can be made to the above description, and details are not elaborated here.
[0122] Optionally, an LIF unit of the first neural network may include M horizontally arranged LIF modules in parallel. In each round, the execution device may simultaneously input the feature information of M groups of image patches into the M horizontally arranged LIF modules in parallel, and process the input data through the M horizontally arranged LIF modules respectively.
[0123] To understand this solution more intuitively, please refer to Figure 10 , Figure 10 which is a schematic diagram of sequentially inputting the feature information of at least two groups of image patches into the LIF module in the image processing method provided by the embodiment of the present application. Figure 10 In the example, the feature information of a group of image patches includes the feature information of a row of image patches, and there are two horizontally arranged LIF modules in an LIF unit. As shown in the figure, in the first round, the execution device inputs the feature information of the first column of image patches (i.e., the feature information of a group of image patches represented by E1) into the horizontally arranged LIF module, and inputs the feature information of the third column of image patches (i.e., the feature information of a group of image patches represented by F1) into the horizontally arranged LIF module. That is, in one round, the feature information of 2 groups of image patches is respectively input into 2 horizontally arranged LIF modules in parallel.
[0124] In the second round, the execution device inputs the feature information of the second column of image patches (i.e., the feature information of a group of image patches represented by E2) into the horizontally arranged LIF module, and inputs the feature information of the fourth column of image patches (i.e., the feature information of a group of image patches represented by F2) into the horizontally arranged LIF module, thereby realizing the input of the feature information of four groups of image patches into 2 horizontally arranged LIF modules in parallel. It should be understood that Figure 10 the example in
[0125] is only for facilitating the understanding of this solution and does not limit this solution.
[0126] 304. The execution device obtains second feature information corresponding to the image to be processed according to the target data, where the second feature information includes the updated feature information of the image patches.
[0127] In the embodiment of the present application, after the execution device obtains a plurality of target data generated by the LIF module, it may obtain second feature information corresponding to the image to be processed according to the foregoing plurality of target data. The second feature information includes the updated feature information of the image patches. Both the first feature information and the second feature information are the feature information of the image to be processed.
[0128] Specifically, if the first neural network only includes vertical LIF modules or horizontal LIF modules, the execution device may determine the multiple target data output by the vertical LIF modules or horizontal LIF modules as the second feature information corresponding to the image to be processed; alternatively, other neural network layers may be used to process the output multiple target data again, and the processed data may be determined as the second feature information corresponding to the image to be processed.
[0129] If the first neural network includes both vertical LIF modules and horizontal LIF modules, the execution device may fuse the target data output by the vertical LIF modules and horizontal LIF modules, and directly determine the fused data as the second feature information corresponding to the image to be processed. Alternatively, the execution device may perform an update operation using other neural network layers before or after performing the fusion operation.
[0130] Further, if the first neural network is specifically an MLP, the above-mentioned other neural network layers may be MLP layers; if the first neural network is specifically a convolutional neural network, the above-mentioned other neural network layers may be convolutional neural network layers; if the first neural network is specifically a Transformer neural network, the above-mentioned other neural network layers may be Transformer neural network layers, etc.; if the first neural network uses other types of neural networks, the above-mentioned other neural network layers may also be replaced with other types of neural network layers, etc., which will not be elaborated here.
[0131] In the embodiments of the present application, whether the first neural network is an MLP, a convolutional neural network, or a Transformer neural network, the LIF module can be compatible through the image processing method provided by the embodiments of the present application. Since the MLP, convolutional neural network, and Transformer neural network can be applied to different application scenarios, the application scenarios and implementation flexibility of the present solution are greatly expanded.
[0132] It should be noted that the steps described in steps 302 to 304 above are the steps executed by a LIF unit in the first neural network. After the execution device obtains the second feature information corresponding to the image to be processed, other neural network layers may be used to update the second feature information, that is, to update the feature information of the image to be processed again.
[0133] Further, in combination with the above Figure 4 For understanding, steps 302 to 304 may be executed multiple times in the process of the execution device obtaining the feature information of the image to be processed through the first neural network. In the embodiments of the present application, the number of executions between steps 302 to 304 and step 301 is not limited. After step 301 is executed once, steps 302 to 304 may be executed once or multiple times, and then step 305 may be entered.
[0134] 305. The execution device performs feature processing on the feature information of the image to be processed through a second neural network, and obtains a prediction result corresponding to the image to be processed.
[0135] In the embodiment of the present application, after the execution device generates the feature information of the image to be processed through a first neural network, it can perform feature processing on the feature information of the image to be processed through a second neural network, and obtain a prediction result corresponding to the image to be processed. Among them, the first neural network and the second neural network are included in the same target neural network, and the task performed by the target neural network is any one of the following: image classification, image segmentation, object detection on an image, super-resolution processing of an image, or other types of tasks, etc. Here, the specific implementation tasks of the target neural network are not exhaustively listed.
[0136] The specific meaning of the prediction result corresponding to the image to be processed depends on the type of task performed by the target neural network. As an example, for example, if the task performed by the target neural network is image classification, the prediction result corresponding to the image to be processed can be used to indicate the predicted category corresponding to the image to be processed; as another example, for example, if the task performed by the target neural network is object detection on an image, the prediction result corresponding to the image to be processed can be used to indicate the predicted category and predicted position of each object in the image to be processed; as another example, for example, if the task performed by the target neural network is image segmentation, the prediction result corresponding to the image to be processed can be used to indicate the predicted category of each pixel point in the image to be processed; as another example, for example, if the task performed by the target neural network is image segmentation, the prediction result corresponding to the image to be processed can include the processed image, etc. Here, it is not exhaustively listed.
[0137] In the embodiment of the present application, multiple application scenarios of the present solution are provided, greatly expanding the implementation flexibility of the present solution.
[0138] In the embodiment of the present application, the feature information of the entire image to be processed is divided into the feature information of multiple image blocks in the image to be processed. The feature information of the multiple image blocks can be divided into at least two groups of image block feature information. The at least two groups of image block feature information are sequentially input into the LIF module to obtain the target data generated by the LIF module, and then the updated feature information of the image to be processed is obtained according to the target data; through the foregoing method, feature extraction of a single image can be realized through the LIF module, and then the LIF module can be applied to perform mainstream general vision tasks.
[0139] II. Training stage
[0140] In the embodiment of the present application, specifically, please refer to Figure 11 , Figure 11It is a schematic flowchart of a method for training a neural network provided by an embodiment of the present application. The method for training a neural network provided by an embodiment of the present application may include:
[0141] 1101. The training device inputs the image to be processed into the first neural network.
[0142] 1102. The training device obtains the first feature information corresponding to the image to be processed. The image to be processed includes multiple image patches, and the first feature information includes the feature information of the image patches.
[0143] 1103. The training device sequentially inputs the feature information of at least two groups of image patches into the LIF module to implement the leakage and accumulation process of the LIF module, and obtains the target data generated by the LIF module.
[0144] 1104. The training device obtains the second feature information corresponding to the image to be processed according to the target data. The second feature information includes the updated feature information of the image patches.
[0145] 1105. The training device performs feature processing on the feature information of the image to be processed through the second neural network, and obtains the prediction result corresponding to the image to be processed.
[0146] In an embodiment of the present application, a training data set may be configured on the training device. The training data set is used to train the target neural network. The target neural network includes the first neural network and the second neural network. The task performed by the target neural network is any one of the following: image classification, object detection on an image, image segmentation, super-resolution processing of an image, or other types of tasks, etc., which are not enumerated here.
[0147] The training data set includes multiple training data. Each training data includes an image to be processed and the correct result corresponding to the image to be processed. The specific meaning of the correct result corresponding to the image to be processed depends on the type of task performed by the target neural network. The two concepts of "the correct result corresponding to the image to be processed" and "the prediction result corresponding to the image to be processed" are similar. The difference is that "the correct result corresponding to the image to be processed" includes correct information, and "the prediction result corresponding to the image to be processed" includes the information generated by the target neural network.
[0148] For the specific implementation manners of steps 1101 to 1105, reference may be made to Figure 3 the descriptions of steps 301 to 305 in the corresponding embodiment, which will not be elaborated here.
[0149] 1106. The training device trains the first neural network and the second neural network according to the prediction result corresponding to the image to be processed and the correct result corresponding to the image to be processed, using a loss function that indicates the similarity between the prediction result and the correct result.
[0150] In the embodiments of the present application, the training device can generate the function value of the loss function according to the prediction result corresponding to the image to be processed and the correct result corresponding to the image to be processed, perform gradient derivation on the function value of the loss function, and backpropagate the gradient value to update the weight parameters of the first neural network and the second neural network (i.e., the target neural network), so as to complete one training of the first neural network and the second neural network. The training device repeats steps 1101 to 1106 until the convergence condition is met.
[0151] Among them, the loss function indicates the similarity between the prediction result corresponding to the image to be processed and the correct result corresponding to the image to be processed. The type of the loss function can be flexibly selected in combination with the actual application scenario. As an example, if the task performed by the target neural network is image classification, the loss function can be selected as the cross-entropy loss function, 0-1 loss function or other types of loss functions, etc. The examples here are only for facilitating the understanding of this solution and are not used to limit this solution.
[0152] The convergence condition can be that the convergence condition of the loss function is met or the number of iterations reaches a preset number, etc., which is not limited here.
[0153] In the embodiments of the present application, not only the implementation steps of the first neural network in the execution stage are provided, but also the implementation steps of the first neural network in the training stage are provided, which expands the application scenario of this solution and improves the comprehensiveness of this solution.
[0154] To more intuitively understand the beneficial effects brought by this solution, the beneficial effects brought by this solution are described below in combination with experimental data. First, taking the target neural network performing the image classification task as an example, experiments were carried out on the ImageNet dataset, and the experimental results are shown in Table 1 below.
[0155] The neural network adopted Top-1 accuracy ResMLP-B24 81.0 DeiT-B 81.8 AS-MLP-B 83.3 Embodiments of this application 83.5
[0156] Table 1
[0157] Among them, ResMLP-B24, DeiT-B, and AS-MLP-B are three existing neural networks. The aforementioned three neural networks can be used to classify images. From the above data, it can be seen that the accuracy of the classification results obtained by using the model provided in the embodiments of the present application is the highest.
[0158] Next, taking the target neural network performing object detection on images as an example, the experimental results are shown in Table 2 below.
[0159] The neural network adopted mIoU DNL (the backbone network adopted is ResNet-101) 46.0 Swin-S 47.6 OCRNet (the backbone network adopted is HRNet-w48) 45.7 Embodiments of this application 49.0
[0160] Table 2
[0161] Among them, DNL, Swin-S, and OCRNet are all existing neural networks. mIoU is an index for evaluating the accuracy of the detection results of object detection on images. As can be seen from the above data, the accuracy of the object detection results obtained by using the model provided in the embodiments of the present application is the highest.
[0162] In Figures 1a to 11 On the basis of the corresponding embodiments, in order to better implement the above solutions of the embodiments of the present application, the following also provides related devices for implementing the above solutions. Specifically, refer to Figure 12 , Figure 12 is a schematic structural diagram of an image processing device provided in an embodiment of the present application. The image processing device 1200 includes: an input unit 1201, configured to input an image to be processed into a first neural network; a feature extraction unit 1202, configured to extract features of the image to be processed through the first neural network to obtain feature information of the image to be processed.
[0163] Among them, the feature extraction unit 1202 includes: an acquisition subunit 12021, configured to acquire first feature information corresponding to the image to be processed. The image to be processed includes multiple image blocks, and the first feature information includes feature information of the image blocks; a generation subunit 12022, configured to sequentially input feature information of at least two groups of image blocks into a leaky-integrate-and-fire (LIF) module to obtain target data generated by the LIF module. The feature information of a group of image blocks includes feature information of at least one image block; an acquisition subunit 12021, configured to acquire second feature information corresponding to the image to be processed according to the target data. The second feature information includes updated feature information of the image blocks. Both the first feature information and the second feature information are feature information of the image to be processed.
[0164] In a possible design, the generation subunit 12022 is specifically configured to sequentially input feature information of at least two groups of image blocks into the LIF module, and when the excitation condition of the LIF module is satisfied, use an activation function to generate target data, and the target data is not binary data.
[0165] In a possible design, the first neural network is a multi-layer perceptron (MLP), a convolutional neural network, or a neural network using a self-attention mechanism.
[0166] In a possible design, the image processing device 1200 further includes: a feature processing unit, configured to perform feature processing on the feature information of the image to be processed through a second neural network to obtain a prediction result corresponding to the image to be processed, where the first neural network and the second neural network are included in the same target neural network, and the task performed by the target neural network is any one of the following: classification, segmentation, object detection, or super-resolution.
[0167] It should be noted that the information interaction, execution process, etc. between the modules / units in the image processing device 1200 are based on the same concept as the corresponding method embodiments in this application. For specific content, reference can be made to the descriptions in the method embodiments shown above in this application, and details will not be elaborated here. Figures 3 to 10 Please refer to
[0168] Please refer to Figure 13 , Figure 13 FIG. is a schematic structural diagram of a training device for a neural network provided by an embodiment of the present application. The neural network training device 1300 includes: a feature extraction unit 1301, configured to input the image to be processed into a first neural network, and perform feature extraction on the image to be processed through the first neural network to obtain the feature information of the image to be processed; a feature processing unit 1302, configured to perform feature processing on the feature information of the image to be processed through a second neural network to obtain a prediction result corresponding to the image to be processed; a training unit 1303, configured to train the first neural network and the second neural network according to the prediction result corresponding to the image to be processed and the correct result by using a loss function, where the loss function indicates the similarity between the prediction result and the correct result.
[0169] Among them, the feature extraction unit 1301 includes: an acquisition subunit 13011, configured to acquire first feature information corresponding to the image to be processed, the image to be processed includes a plurality of image blocks, and the first feature information includes the feature information of the image blocks; a generation subunit 13012, configured to sequentially input the feature information of at least two groups of image blocks into a leaky-integrate-and-fire (LIF) module to obtain target data generated by the LIF module, the feature information of a group of image blocks includes the feature information of at least one image block; the acquisition subunit 13011 is further configured to acquire second feature information corresponding to the image to be processed according to the target data, the second feature information includes the updated feature information of the image blocks, and both the first feature information and the second feature information are the feature information of the image to be processed.
[0170] In a possible design, the generation subunit 13012 is specifically configured to sequentially input the feature information of at least two groups of image blocks into the LIF module, and when the excitation condition of the LIF module is satisfied, use an activation function to generate target data, and the target data is not binary data.
[0171] It should be noted that the information interaction, execution process, etc. among the modules / units in the neural network training device 1300 are based on the same concept as the corresponding method embodiments in this application. For specific content, please refer to the descriptions in the method embodiments shown above in this application, and details will not be repeated here. Figure 11 The corresponding method embodiments are based on the same concept. For specific content, please refer to the descriptions in the method embodiments shown above in this application, and details will not be repeated here.
[0172] Next, a kind of execution device provided by the embodiments of this application will be introduced. Please refer to Figure 14 , Figure 14 which is a schematic structural diagram of the execution device provided by the embodiments of this application. The execution device 1400 may specifically be a virtual reality (VR) device, a mobile phone, a tablet computer, a laptop computer, a smart wearable device, a monitoring data processing device, etc., and is not limited here. Specifically, the execution device 1400 includes: a receiver 1401, a transmitter 1402, a processor 1403, and a memory 1404 (where the number of processors 1403 in the execution device 1400 may be one or more, Figure 14 and one processor is taken as an example here). Among them, the processor 1403 may include an application processor 14031 and a communication processor 14032. In some embodiments of this application, the receiver 1401, the transmitter 1402, the processor 1403, and the memory 1404 may be connected through a bus or other means.
[0173] The memory 1404 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1403. A part of the memory 1404 may also include a non-volatile random access memory (NVRAM). The memory 1404 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof. Among them, the operation instructions may include various operation instructions for implementing various operations.
[0174] The processor 1403 controls the operation of the execution device. In a specific application, the various components of the execution device are coupled together through a bus system. Among them, the bus system may include a power bus, a control bus, a status signal bus, etc. in addition to a data bus. However, for the sake of clear illustration, all kinds of buses are referred to as the bus system in the figure.
[0175] The method disclosed in the embodiments of the present application can be applied to or implemented by the processor 1403. The processor 1403 can be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be completed by the integrated logic circuit in hardware or instructions in software form in the processor 1403. The above-mentioned processor 1403 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The processor 1403 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 1404, and the processor 1403 reads the information in the memory 1404 and combines its hardware to complete the steps of the above method.
[0176] The receiver 1401 can be used to receive input digital or character information, and generate signal inputs related to the relevant settings and function controls of the execution device. The transmitter 1402 can be used to output digital or character information through the first interface; the transmitter 1402 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 1402 can also include a display device such as a display screen.
[0177] In the embodiments of the present application, the processor 1403 is used to execute Figures 3 to 10 the image processing method executed by the execution device in the corresponding embodiment. Specifically, the application processor 14031 is used to execute the following steps:
[0178] Input the image to be processed into the first neural network, and perform feature extraction on the image to be processed through the first neural network to obtain the feature information of the image to be processed. Among them,
[0179] Performing feature extraction on the image to be processed through a first neural network, including: obtaining first feature information corresponding to the image to be processed, where the image to be processed includes multiple image blocks, and the first feature information includes the feature information of the image blocks; sequentially inputting the feature information of at least two groups of image blocks into a leaky-integrate-and-fire (LIF) module to obtain target data generated by the LIF module, where the feature information of a group of image blocks includes the feature information of at least one image block; and obtaining second feature information corresponding to the image to be processed according to the target data, where the second feature information includes the updated feature information of the image blocks, and both the first feature information and the second feature information are the feature information of the image to be processed.
[0180] It should be noted that the specific manner in which the application processor 14031 executes the above steps is based on the same concept as the corresponding method embodiments in this application, Figures 3 to 10 and the technical effects brought by it are the same as those of the corresponding method embodiments in this application. Figures 3 to 10 For specific content, reference can be made to the descriptions in the method embodiments shown above in this application, and details will not be elaborated here.
[0181] An embodiment of this application also provides a training device. Please refer to Figure 15 , Figure 15 which is a schematic structural diagram of a training device provided by an embodiment of this application. Specifically, the training device 1500 is implemented by one or more servers. The training device 1500 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 1522 (for example, one or more processors) and a memory 1532, and one or more storage media 1530 (for example, one or more mass storage devices) for storing application programs 1542 or data 1544. Among them, the memory 1532 and the storage media 1530 may be transient storage or persistent storage. The program stored in the storage media 1530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the training device. Further, the central processing unit 1522 may be configured to communicate with the storage media 1530 and execute a series of instruction operations in the storage media 1530 on the training device 1500.
[0182] The training device 1500 may further include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input / output interfaces 1558, and / or one or more operating systems 1541, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0183] In the embodiments of the present application, the central processing unit 1522 is configured to execute Figure 12 the image processing method executed by the training device in the corresponding embodiments. Specifically, the central processing unit 1522 is configured to execute the following steps:
[0184] Input the image to be processed into the first neural network, extract the feature information of the image to be processed through the first neural network, obtain the feature information of the image to be processed, and perform feature processing on the feature information of the image to be processed through the second neural network to obtain the prediction result corresponding to the image to be processed; according to the prediction result and the correct result corresponding to the image to be processed, use the loss function to train the first neural network and the second neural network, and the loss function indicates the similarity between the prediction result and the correct result.
[0185] Among them, extracting the feature information of the image to be processed through the first neural network includes: obtaining the first feature information corresponding to the image to be processed, the image to be processed includes multiple image blocks, and the first feature information includes the feature information of the image blocks; sequentially inputting the feature information of at least two groups of image blocks into the leaky-integrate-and-fire (LIF) module to obtain the target data generated by the LIF module, and the feature information of a group of image blocks includes the feature information of at least one image block; according to the target data, obtain the second feature information corresponding to the image to be processed, and the second feature information includes the updated feature information of the image blocks. Both the first feature information and the second feature information are the feature information of the image to be processed.
[0186] It should be noted that the specific manner in which the central processing unit 1522 executes the above steps is based on the same concept as the corresponding method embodiments in the present application, and the technical effects brought by it are the same as those of the corresponding method embodiments in the present application. For specific content, reference can be made to the descriptions in the method embodiments shown above in the present application, and details are not repeated here. Figure 11 Based on the same concept as the corresponding method embodiments in the present application, and the technical effects brought by it are the same as those of the corresponding method embodiments in the present application. Figure 11 For specific content, reference can be made to the descriptions in the method embodiments shown above in the present application, and details are not repeated here.
[0187] In the embodiments of the present application, there is also provided a computer program product, which when running on a computer, causes the computer to execute the steps executed by the execution device in the method described in the foregoing Figures 3 to 10 shown embodiments, or causes the computer to execute the steps executed by the training device in the method described in the foregoing Figure 11 shown embodiments.
[0188] In the embodiments of the present application, there is also provided a computer-readable storage medium, in which a program for signal processing is stored. When it runs on a computer, it causes the computer to execute the steps executed by the execution device in the method described in the foregoing Figures 3 to 10 shown embodiments, or causes the computer to execute the steps executed by the training device in the method described in the foregoing Figure 11 shown embodiments.
[0189] The image processing device, neural network training device, execution device, or training device provided by the embodiments of the present application may specifically be a chip. The chip includes a processing unit and a communication unit. The processing unit may be a processor, for example, and the communication unit may be an input / output interface, a pin, a circuit, or the like. The processing unit may execute computer-executable instructions stored in the storage unit to cause the chip to execute the image processing method described in the above Figures 3 to 10 illustrated embodiments, or to cause the chip in the training device to execute the neural network training method described in the above Figure 11 illustrated embodiments. Optionally, the storage unit is a storage unit within the chip, such as a register, a cache, etc. The storage unit may also be a storage unit outside the chip within the radio access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0190] Specifically, please refer to Figure 16 , Figure 16 which is a schematic structural diagram of the chip provided by the embodiments of the present application. The chip may be embodied as a neural network processor NPU 160. The NPU 160 is mounted on the main CPU (Host CPU) as a coprocessor, and tasks are allocated by the Host CPU. The core part of the NPU is the arithmetic circuit 160. The arithmetic circuit 1603 is controlled by the controller 1604 to extract matrix data from the memory and perform multiplication operations.
[0191] In some implementations, the arithmetic circuit 1603 includes multiple processing units (Process Engine, PE) internally. In some implementations, the arithmetic circuit 1603 is a two-dimensional systolic array. The arithmetic circuit 1603 may also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1603 is a general matrix processor.
[0192] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 1602 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 1601 and performs matrix operations with matrix B. The partial results or final results of the obtained matrix are saved in the accumulator 1608.
[0193] The unified memory 1606 is used to store input data and output data. The weight data is directly transferred through the Direct Memory Access Controller (DMAC) 1605, and the DMAC transfers it to the weight memory 1602. The input data is also transferred to the unified memory 1606 through the DMAC.
[0194] The BIU is the Bus Interface Unit, that is, the bus interface unit 1610, which is used for the interaction between the AXI bus, the DMAC, and the Instruction Fetch Buffer (IFB) 1609.
[0195] The bus interface unit 1610 (Bus Interface Unit, abbreviated as BIU) is used for the instruction fetch buffer 1609 to obtain instructions from the external memory, and is also used for the storage unit access controller 1605 to obtain the original data of the input matrix A or the weight matrix B from the external memory.
[0196] The DMAC is mainly used to transfer the input data in the external memory DDR to the unified memory 1606, or transfer the weight data to the weight memory 1602, or transfer the input data to the input memory 1601.
[0197] The vector calculation unit 1607 includes multiple arithmetic processing units, which, if necessary, further process the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolution / full connection layer network calculations in neural networks, such as Batch Normalization, pixel-level summation, upsampling of the feature plane, etc.
[0198] In some implementations, the vector calculation unit 1607 can store the processed output vector in the unified memory 1606. For example, the vector calculation unit 1607 can apply a linear function and / or a non-linear function to the output of the arithmetic circuit 1603, such as linear interpolation of the feature plane extracted by the convolutional layer, and another example is the vector of the accumulated value to generate the activation value. In some implementations, the vector calculation unit 1607 generates normalized values, pixel-level summation values, or both. In some implementations, the processed output vector can be used as the activation input to the arithmetic circuit 1603, such as for use in subsequent layers in the neural network.
[0199] The instruction fetch buffer 1609 connected to the controller 1604 is used to store the instructions used by the controller 1604;
[0200] The unified memory 1606, the input memory 1601, the weight memory 1602, and the fetch memory 1609 are all On-Chip memories. The external memory is private to the NPU hardware architecture.
[0201] Among them, Figures 3 to 11 The operations of each layer in the first neural network and the second neural network shown can be executed by the operation circuit 1603 or the vector calculation unit 1607.
[0202] Among them, the processor mentioned anywhere above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the method in the first aspect above.
[0203] In addition, it should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.
[0204] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits. However, for this application, software programs are more often the better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disc of a computer, and includes several instructions for causing a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in various embodiments of this application.
[0205] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0206] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
Claims
1. An image processing method, characterized in that, the method includes: inputting the image to be processed into a first neural network, and extracting features of the image to be processed through the first neural network to obtain feature information of the image to be processed; wherein, the extracting features of the image to be processed through the first neural network includes: obtaining first feature information corresponding to the image to be processed, the image to be processed includes a plurality of image blocks, and the first feature information includes feature information of the image blocks; sequentially inputting feature information of at least two groups of the image blocks into a Leaky-Integrate-and-Fire (LIF) module to obtain target data generated by the LIF module, and feature information of a group of the image blocks includes feature information of at least one image block; obtaining second feature information corresponding to the image to be processed according to the target data, the second feature information includes updated feature information of the image blocks, and both the first feature information and the second feature information are feature information of the image to be processed.
2. The method according to claim 1, characterized in that, the sequentially inputting feature information of at least two groups of the image blocks into the LIF module to obtain target data generated by the LIF module includes: sequentially inputting feature information of at least two groups of the image blocks into the LIF module, and when the firing condition of the LIF module is satisfied, generating the target data by using an activation function, and the target data is not binary data.
3. The method according to claim 1 or 2, characterized in that, the first neural network is a multi-layer perceptron (MLP), a convolutional neural network, or a neural network adopting a self-attention mechanism.
4. The method according to claim 1 or 2, characterized in that, the method further includes: performing feature processing on the feature information of the image to be processed through a second neural network to obtain a prediction result corresponding to the image to be processed, wherein the first neural network and the second neural network are included in the same target neural network, and the task performed by the target neural network is any one of the following: classification, segmentation, object detection, or super-resolution.
5. A training method of a neural network, characterized in that, the method includes: inputting the image to be processed into a first neural network, extracting features of the image to be processed through the first neural network to obtain feature information of the image to be processed, and performing feature processing on the feature information of the image to be processed through a second neural network to obtain a prediction result corresponding to the image to be processed; training the first neural network and the second neural network by using a loss function according to the prediction result and the correct result corresponding to the image to be processed, and the loss function indicates the similarity between the prediction result and the correct result; wherein, the extracting features of the image to be processed through the first neural network includes: obtaining first feature information corresponding to the image to be processed, the image to be processed includes a plurality of image blocks, and the first feature information includes feature information of the image blocks; Input the feature information of at least two groups of the image patches into a leaky-integrate-and-fire (LIF) module in sequence to obtain the target data generated by the LIF module. The feature information of one group of the image patches includes the feature information of at least one image patch. Obtain the second feature information corresponding to the image to be processed according to the target data. The second feature information includes the updated feature information of the image patches. Both the first feature information and the second feature information are the feature information of the image to be processed.
6. The method according to claim 5, wherein, the step of inputting the feature information of at least two groups of the image patches into the LIF module in sequence to obtain the target data generated by the LIF module includes: Input the feature information of at least two groups of the image patches into the LIF module in sequence. When the firing condition of the LIF module is satisfied, generate the target data by using an activation function. The target data is not binary data.
7. An image processing apparatus, wherein, the apparatus includes: an input unit configured to input an image to be processed into a first neural network; a feature extraction unit configured to extract the feature information of the image to be processed through the first neural network, and obtain the feature information of the image to be processed. Wherein, the feature extraction unit includes: an acquisition subunit configured to acquire the first feature information corresponding to the image to be processed. The image to be processed includes a plurality of image patches. The first feature information includes the feature information of the image patches; a generation subunit configured to input the feature information of at least two groups of the image patches into a leaky-integrate-and-fire (LIF) module in sequence to obtain the target data generated by the LIF module. The feature information of one group of the image patches includes the feature information of at least one image patch; the acquisition subunit is configured to obtain the second feature information corresponding to the image to be processed according to the target data. The second feature information includes the updated feature information of the image patches. Both the first feature information and the second feature information are the feature information of the image to be processed.
8. The apparatus according to claim 7, wherein, the generation subunit is specifically configured to input the feature information of at least two groups of the image patches into the LIF module in sequence. When the firing condition of the LIF module is satisfied, generate the target data by using an activation function. The target data is not binary data.
9. The apparatus according to claim 7 or 8, wherein, the first neural network is a multi-layer perceptron (MLP), a convolutional neural network, or a neural network adopting a self-attention mechanism.
10. The apparatus according to claim 7 or 8, wherein, the apparatus further includes: a feature processing unit configured to perform feature processing on the feature information of the image to be processed through a second neural network to obtain the prediction result corresponding to the image to be processed. Wherein, the first neural network and the second neural network are included in the same target neural network. The task performed by the target neural network is any one of the following: classification, segmentation, object detection, or super-resolution.
11. A training apparatus for a neural network, It is characterized in that the device includes: a feature extraction unit configured to input an image to be processed into a first neural network, and perform feature extraction on the image to be processed through the first neural network to obtain feature information of the image to be processed; a feature processing unit configured to perform feature processing on the feature information of the image to be processed through a second neural network to obtain a prediction result corresponding to the image to be processed; a training unit configured to train the first neural network and the second neural network according to the prediction result and the correct result corresponding to the image to be processed by using a loss function, where the loss function indicates the similarity between the prediction result and the correct result; wherein, the feature extraction unit includes: an acquisition subunit configured to acquire first feature information corresponding to the image to be processed, the image to be processed includes a plurality of image blocks, and the first feature information includes feature information of the image blocks; a generation subunit configured to sequentially input the feature information of at least two groups of the image blocks into a leaky-integrate-and-fire (LIF) module to obtain target data generated by the LIF module, and the feature information of a group of the image blocks includes feature information of at least one image block; the acquisition subunit is further configured to acquire second feature information corresponding to the image to be processed according to the target data, the second feature information includes updated feature information of the image blocks, and both the first feature information and the second feature information are feature information of the image to be processed.
12. The device according to claim 11, it is characterized in that the generation subunit is specifically configured to sequentially input the feature information of at least two groups of the image blocks into the LIF module, and when an excitation condition of the LIF module is satisfied, generate the target data by using an activation function, and the target data is not binary data.
13. A computer program product, it is characterized in that the computer program product includes a program, and when the program runs on a computer, it causes the computer to execute the method according to any one of claims 1 to 4, or causes the computer to execute the method according to claim 5 or 6.
14. A computer-readable storage medium, it is characterized in that the computer-readable storage medium stores a program, and when the program runs on a computer, it causes the computer to execute the method according to any one of claims 1 to 4, or causes the computer to execute the method according to claim 5 or 6.
15. An execution device, it is characterized in that it includes a processor and a memory, the processor is coupled to the memory, the memory is configured to store a program; the processor is configured to execute the program in the memory, so that the execution device executes the method according to any one of claims 1 to 4.
16. A training device, it is characterized in that it includes a processor and a memory, the processor is coupled to the memory, the memory is configured to store a program; the processor is configured to execute the program in the memory, so that the training device executes the method according to claim 5 or 6.
Citation Information
Patent Citations
Image processing method, neural network training method and related equipment
CN113065997A
Image enhancement model training method and device, image enhancement method and device and electronic equipment
CN113469897A