Data processing method and related device
By adjusting the number of fusion nodes according to their importance in the graph neural network, the allocation of computing resources is optimized, which solves the problem of wasted computing resources in image processing tasks and improves task efficiency and image quality.
Patent Information
- Application Number
- CN202410494329.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-23
- Publication Date
- 2025-10-24
AI Technical Summary
Existing graph neural networks, in image processing tasks, waste computational resources and incur excessive computing power because different nodes are treated equally, and cannot effectively utilize the 'uneven recovery' characteristic in super-resolution tasks.
By setting different fusion quantities and importance of nodes in the graph neural network, differentiated processing is performed based on the impact of nodes on the accuracy of image processing tasks, reducing the fusion quantity of unimportant nodes and optimizing the allocation of computing resources.
It reduces the computational overhead of graph neural networks, improving the efficiency and quality of image processing tasks, especially in super-resolution tasks where it better reconstructs high-frequency components.
Smart Images

Figure CN120833259A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence (AI), and particularly relates to a data processing method and related apparatus. BACKGROUND
[0002] The purpose of super-resolution (SR) is to construct a high-resolution image from a low-resolution image. SR work usually applies deep neural network knowledge learned from high-resolution training images to construct missing details in the low-resolution input.
[0003] When performing an image processing task (for example, an SR task) through a graph neural network, graph construction needs to be performed. In the graph construction process, different nodes in the image can be connected and fused. In the prior art, the same operation paradigm is performed on each node, that is, the importance of different nodes is considered to be the same. For example, the number of nodes fused on each node is the same when constructing the graph, which results in a large overall computing power overhead of the graph neural network. SUMMARY
[0004] In a first aspect, a data processing method is provided. The method includes: obtaining an image, the image including a first node and a second node, the first node and the second node being different pixel points or image blocks on the image; processing the image through a graph neural network to obtain a processing result; the processing result including a first fusion result and a second fusion result, the first fusion result being obtained by aggregating information of the first node and information of M nodes on the image, the second fusion result being obtained by aggregating information of the second node and information of N nodes on the image, wherein an influence degree of the first node on the accuracy of an image processing task is greater than an influence degree of the second node on the accuracy of the image processing task, and the M is greater than the N; and performing the image processing task according to the processing result.
[0005] In the existing technology, the same operation paradigm is performed on each node, that is, it is assumed that the importance of different nodes is the same. In this case, for different nodes, the number of nodes to be fused is the same (or the fusion number is unrelated to the importance of the node). However, for image processing tasks, the importance of different nodes on the image (that is, the degree of influence on the execution effect or execution accuracy of the image processing task) may be different. For example, in an embodiment of the present application, the image includes a first node and a second node. The fusion number used when fusing the node information of the first node is M, and the fusion number used when fusing the node information of the second node is N. The first node is more important to the image processing task than the second node, so M is set to be greater than N. In this case, for nodes that are not very important, setting the fusion number lower (compared to nodes with higher importance) will not have a great impact on the processing results of the image processing task, and at the same time can reduce the computing power overhead of the graph neural network.
[0006] In a possible implementation, the image processing task is super-resolution, image denoising, image segmentation, or image recognition.
[0007] In a possible implementation, the first node is a node in a high-frequency area relative to the second node in the image.
[0008] For example, the super-resolution task has the characteristic of "unbalanced restoration". In the process of super-resolution of low-definition images, the low-definition part that accounts for the vast majority of the image does not actually require a lot of changes; only a small amount of high-frequency details require a lot of reconstruction of the neural network. Therefore, for the nodes in the low-frequency area, there is no need to fuse a lot of nodes to ensure a good super-resolution effect. The existing super-resolution technical solutions have failed to fully utilize the characteristics of "unbalanced restoration" in super-resolution. The present application designs a graph of changes in node degree (in the embodiment of the present application, the node degree can also be referred to as the fusion number), so that the graph neural network used for super-resolution tasks can pay more attention to the parts of the picture that require high-frequency reconstruction.
[0009] In a possible implementation, the graph neural network includes multiple blocks, the first fusion result is obtained by aggregating the information of the first node and the information of M nodes on the image through a target block among the multiple blocks, and the second fusion result is obtained by aggregating the information of the second node and the information of N nodes on the image through the target block.
[0010] In a possible implementation, a total aggregation value related to the current computing power and a total influence degree of the nodes on the image on the accuracy of the image processing task can also be acquired, the value of M is determined according to the total aggregation value and a relationship between the influence degree of the first node on the accuracy of the image processing task and the total influence degree, and the value of N is determined according to the total aggregation value and a relationship between the influence degree of the second node on the accuracy of the image processing task and the total influence degree. That is, in the determination of the fusion quantity of the nodes, a total number is determined based on the current power requirement, and then the fusion quantity of each node is allocated based on the total number in the case of meeting the total number determination, so that the computing overhead of the graph neural network can be adapted to the power requirement.
[0011] In a possible implementation, the graph neural network includes a first block and a second block, and processing the image by using the graph neural network to obtain a processing result includes: performing first aggregation on the nodes in the image by using the first block, wherein the aggregation object of the nodes in the first aggregation is selected from the nodes included in one continuous image region; and performing aggregation on the nodes in the image by using the second block, wherein the aggregation object of the nodes in the second aggregation is selected from the nodes in a plurality of regions of the image, and there is a gap between adjacent regions in the plurality of regions.
[0012] The existing graph neural network image processing algorithm cannot be directly applied to the bottom vision field: previous graph algorithms often take small image patches as nodes of the graph; and the bottom vision has a higher requirement for the reconstruction of pixels, and the aggregation of patch nodes will affect the image quality. Therefore, a single pixel can be taken as a node of the graph to ensure the image quality of super-resolution reconstruction. The problem brought by taking a single pixel as a node of the graph is that, in the field of bottom vision, the resolution of the image to be processed is large, and if the pixels are directly taken as nodes of the graph in the process of constructing the graph, the search space of similar nodes is too large, which will cause a large power consumption. Therefore, an embodiment of the present application designs a sampling method, which can significantly reduce the search space of similar nodes to control the power consumption of graph construction. Specifically, for each node, the nodes can be collected in the local pixels and the global pixels and in a smaller sampling space, so as to significantly reduce the power consumption of graph construction.
[0013] In a second aspect, the present application provides a data processing apparatus, the apparatus comprising:
[0014] An acquisition module is configured to acquire an image, the image comprising a first node and a second node, the first node and the second node being different pixel points or image blocks on the image;
[0015] A processing module is configured to process the image by using a graph neural network to obtain a processing result, the processing result comprising a first fusion result and a second fusion result, the first fusion result being obtained by aggregating information of the first node and information of M nodes on the image, the second fusion result being obtained by aggregating information of the second node and information of N nodes on the image, wherein the first node has a greater influence on the accuracy of an image processing task than the second node, and M is greater than N; and the image processing task is performed according to the processing result.
[0016] In a possible implementation, the image processing task is super-resolution, image denoising, image segmentation, or image recognition.
[0017] In a possible implementation, the first node is a node in a high-frequency region of the image relative to the second node.
[0018] In a possible implementation, the graph neural network comprises a plurality of blocks, the first fusion result is obtained by aggregating information of the first node and information of M nodes on the image by using a target block in the plurality of blocks, and the second fusion result is obtained by aggregating information of the second node and information of N nodes on the image by using the target block.
[0019] In a possible implementation, the processing module is further configured to:
[0020] acquire a total aggregation value and a total influence degree of nodes on the image on the accuracy of the image processing task, the total aggregation value being related to a current computing power;
[0021] determine the value of M according to the total aggregation value and a relationship between the influence degree of the first node on the accuracy of the image processing task and the total influence degree;
[0022] determine the value of N according to the total aggregation value and a relationship between the influence degree of the second node on the accuracy of the image processing task and the total influence degree.
[0023] In a possible implementation, the graph neural network comprises a first block and a second block;
[0024] The processing module is specifically configured to:
[0025] performing the first aggregation, the aggregation object of a node is selected from nodes included in one continuous image region where the node is located;
[0026] performing the second aggregation, the aggregation object of a node is selected from nodes in a plurality of regions of the image, and there is a gap between adjacent regions in the plurality of regions.
[0027] In a third aspect, a chip is provided, which includes at least one processing unit and an interface circuit, the interface circuit is configured to provide program instructions or data for the at least one processing unit, the at least one processing unit is configured to execute the program instructions to implement the method in any one of the first aspect, the at least one processing unit includes a first hardware unit and a second hardware unit, the first hardware unit is configured to calculate a prefix sum in a channel dimension, and the second hardware unit is configured to calculate a prefix sum in a spatial dimension.
[0028] In a fourth aspect, an embodiment of the present application provides a data processing apparatus, which can include a memory, a processor and a bus system, wherein the memory is configured to store a program, and the processor is configured to execute the program in the memory to perform the method in the first aspect and any optional method thereof.
[0029] In a fifth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program, when the computer program is executed on a computer, the computer is caused to perform the method in the first aspect and any optional method thereof.
[0030] In a sixth aspect, an embodiment of the present application provides a computer program, when the computer program is executed on a computer, the computer is caused to perform the method in the first aspect and any optional method thereof.
[0031] In a seventh aspect, the present application provides a chip system, which includes a processor, and is configured to support the data processing apparatus to implement the functions involved in the above aspects, for example, to send or process the data involved in the above method; or, information. In a possible design, the chip system further includes a memory, the memory is configured to save necessary program instructions and data for the execution device or the training device. The chip system can be composed of a chip, or can include a chip and other discrete devices. BRIEF DESCRIPTION OF DRAWINGS
[0032] In order to more clearly illustrate the technical method of the embodiments of the present application, the drawings required to be used in the embodiments will be briefly introduced as follows.
[0033] Figure 1An application architecture schematic diagram provided for an embodiment of the present application;
[0034] Figures 2 to 7 An application architecture schematic diagram provided for an embodiment of the present application;
[0035] Figure 8 A data processing method schematic diagram provided for an embodiment of the present application;
[0036] Figure 9 A data processing method schematic diagram provided for an embodiment of the present application;
[0037] Figure 10 A search space schematic for an embodiment of the present application;
[0038] Figure 11 A model schematic for an embodiment of the present application;
[0039] Figure 12 A data processing device structure schematic provided for an embodiment of the present application;
[0040] Figure 13 An apparatus schematic provided for an embodiment of the present application;
[0041] Figure 14 An apparatus schematic provided for an embodiment of the present application;
[0042] Figure 15 A chip schematic provided for an embodiment of the present application. DETAILED DESCRIPTION
[0043] The embodiments of the present application will be described in detail with reference to the drawings. The terms used in the embodiments of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0044] The embodiments of the present application will be described in detail with reference to the drawings. The terms used in the embodiments of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0045] The terms "first", "second", "third", etc. in the specification and claims of this application and in the above figures are used to distinguish similar objects, and are not necessarily used to describe a particular sequential or chronological order. It should be understood that the terms so used are interchangeable under appropriate circumstances and are merely employed in the description of embodiments of this application for the purpose of differentiation among like objects. Furthermore, the terms "comprising", "having", "including", and any variations thereof in the specification and claims are intended to cover a non-exclusive inclusion, such that a process, method, system, product or apparatus that comprises, has, includes or contains a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, system, product or apparatus.
[0046] First, the overall workflow of the artificial intelligence system is described, please see Figure 1 , Figure 1 A structural diagram of an artificial intelligence subject framework is shown, and the above artificial intelligence subject framework is described below from two dimensions of "intelligent information chain" (horizontal axis) and "IT value chain" (vertical axis). Among them, the "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be a general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes a condensation process of "data-information-knowledge-wisdom". The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of human intelligence, information (provision and processing technology implementation) to the industrial ecological process of the system.
[0047] (1) Infrastructure
[0048] The infrastructure provides computing power support for the artificial intelligence system, realizes communication with the outside world, and realizes support through the underlying platform. Communication with the outside world through sensors; computing power is provided by intelligent chips (CPU, NPU, GPU, ASIC, FPGA, etc. Hardware acceleration chips); the underlying platform includes distributed computing framework and network related platform guarantee and support, which can include cloud storage and computing, interconnection network, etc. For example, sensors and external communication acquire data, which are provided to intelligent chips in the distributed computing system provided by the underlying platform for calculation.
[0049] (2) Data
[0050] The data on the upper layer of the infrastructure is used to represent the data source in the field of artificial intelligence. Data relates to graphics, images, speech, text, and also relates to Internet of Things data of traditional devices, including business data of existing systems and sensing data such as force, displacement, liquid level, temperature, humidity, etc.
[0051] (3) Data processing
[0052] Data processing generally includes data training, machine learning, deep learning, search, reasoning, decision, and the like.
[0053] Among them, machine learning and deep learning can model, extract, preprocess, train, and the like of symbolic and formalized intelligent information of data.
[0054] Reasoning refers to a process of simulating human intelligent reasoning methods in a computer or intelligent system, using formalized information to perform machine thinking and solve problems according to a reasoning control strategy, and a typical function is search and matching.
[0055] Decision refers to a process of decision-making after intelligent information is reasoned, and generally provides functions such as classification, sorting, and prediction.
[0056] (4) General capability
[0057] After data is processed by the above-mentioned data processing, some general capabilities can be formed based on the results of the data processing, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, and the like.
[0058] (5) Intelligent product and industry application
[0059] Intelligent product and industry application refers to products and applications of artificial intelligence systems in various fields, which is a packaging of the overall solution of artificial intelligence, and realizes application landing by productizing intelligent information decision-making. The application fields mainly include intelligent terminals, intelligent transportation, intelligent medical treatment, automatic driving, smart city, and the like.
[0060] The present application can be but is not limited to applied in the field of natural language processing in the field of artificial intelligence, and specifically can be applied in neural network search in the field of natural language processing and neural network reasoning in the field of natural language processing. The following will introduce multiple application scenarios landing to products.
[0061] In order to better understand the scheme of the embodiments of the present application, the following will first introduce the general technical scheme of the embodiments of the present application. Figures 2 to 5 The possible application scenarios of the embodiments of the present application are briefly introduced.
[0062] I. Image processing application program
[0063] The product form of the embodiments of the present application can be an image processing application program. The image processing application program can run on a terminal device or a server on the cloud side.
[0064] In a possible implementation, referring to Figure 2 , the image processing application program can implement an image processing task to obtain a processing result.
[0065] Image processing tasks can be image enhancement, image recognition, image segmentation, generation tasks, etc.
[0066] In one possible implementation, the user can open an image processing application installed on the terminal device and input an image. The image processing application can process the image using a model trained by the method provided in the embodiment of the present application, or using the method provided in the embodiment of the present application, and present the processing results to the user (the presentation method can be but is not limited to display, playback, saving, uploading to the cloud, etc.).
[0067] In one possible implementation, a user can open an image processing application installed on a terminal device and input an image. The image processing application can send the image to a server on the cloud side. The server on the cloud side processes the image using a model trained using the method provided in an embodiment of the present application, and transmits the processing result back to the terminal device. The terminal device can present the processing result to the user (the presentation method can be, but is not limited to, display, playback, saving, uploading to the cloud side, etc.).
[0068] Next, the image processing application in the embodiment of this application is introduced from the perspective of functional architecture and product architecture that implements the functions.
[0069] Reference Figure 2 , Figure 2 This is a schematic diagram of the functional architecture of the image processing application in the embodiment of this application:
[0070] In one possible implementation, Figure 2 As shown, an image processing application 102 can receive input parameters 101 (e.g., including an image) and generate a processing result 103. The image processing application 102 can be executed on (for example) at least one computer system and includes computer code that, when executed by one or more computers, causes the computers to execute a model trained by the method provided in the embodiments of the present application.
[0071] Reference Figure 3 , Figure 3 The following is a schematic diagram of the physical architecture for running image processing applications in the embodiment of the present application:
[0072] See also Figure 3 , Figure 3 Schematic diagram of a system architecture is shown. The system may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 3 (The description is made by taking one server as an example), the server 200 may provide image processing functions for one or more terminals.
[0073] The terminal 100 can install an image processing application or open a webpage related to the image processing function. The application and the webpage can provide an interface. The terminal 100 can receive parameters input by a user on the interface of the image processing function and send the parameters to the server 200. The server 200 can obtain a processing result based on the received parameters and return the processing result to the terminal 100.
[0074] It should be understood that, in some optional implementations, the terminal 100 can also complete the action of obtaining a processing result based on received parameters by itself without the cooperation of the server. The embodiments of the present application are not limited.
[0075] Next, the product form of the terminal 100 is described. Figure 3
[0076] The terminal 100 in the embodiments of the present application can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and the like. The embodiments of the present application are not limited in this regard.
[0077] Figure 4 An optional hardware structure schematic diagram of the terminal 100 is shown.
[0078] Referring to FIG. 1, Figure 4 As shown in FIG. 1, the terminal 100 can include a radio frequency unit 110, a memory 120, an input unit 130, a display unit 140, a camera 150 (optional), an audio circuit 160 (optional), a speaker 161 (optional), a microphone 162 (optional), a processor 170, an external interface 180, a power supply 190, and the like. Those skilled in the art can understand that Figure 4 The terminal or the multifunctional device is only an example and does not constitute a limitation on the terminal or the multifunctional device, and can include more or fewer components than those shown, or combine certain components, or different components.
[0079] The input unit 130 can be used to receive inputted digital or character information, and to generate key signal inputs related to user settings of the portable multifunctional device and control of functions. Specifically, the input unit 130 can include a touch screen 131 (optional) and / or other input devices 132. The touch screen 131 can collect touch operations of a user thereon or therearound (such as operations of the user using a finger, a joint, a stylus, or any suitable object on or near the touch screen), and drive corresponding connected devices according to pre-set programs. The touch screen can detect touch actions of the user on the touch screen, convert the touch actions into touch signals and send the touch signals to the processor 170, and can receive commands from the processor 170 and execute the commands; the touch signals at least include touch point coordinate information. The touch screen 131 can provide an input interface and an output interface between the terminal 100 and the user. In addition, the touch screen can be implemented in various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch screen 131, the input unit 130 can also include other input devices. Specifically, the other input devices 132 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, on-off keys, etc.), trackballs, mice, joysticks, etc.
[0080] Among them, the other input devices 132 can receive an inputted image.
[0081] The display unit 140 can be used to display information inputted by the user or provided to the user, various menus of the terminal 100, interactive interfaces, file display, and / or playing of any kind of multimedia files. In the embodiments of the present application, the display unit 140 can be used to display interfaces of image processing type application programs, processing results, etc.
[0082] The storage 120 can be used to store instructions and data. The storage 120 can mainly include a storage instruction area and a storage data area. The storage data area can store various data such as multimedia files, texts, etc.; the storage instruction area can store software units such as operating systems, applications, instructions required by at least one function, etc., or their subsets, expanded sets. It can also include a non-volatile random access memory; provide the processor 170 with software and applications that include management of hardware, software, and data resources in the computing processing device, support control. It is also used for storage of multimedia files, and storage of running programs and applications.
[0083] The processor 170 is the control center of the terminal 100, connects each part of the whole terminal 100 by various interfaces and lines, executes various functions of the terminal 100 and processes data by running or executing the instructions stored in the memory 120 and calling the data stored in the memory 120, thereby overall controlling the terminal device. Optionally, the processor 170 can include one or more processing units; preferably, the processor 170 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface and application program, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 170. In some embodiments, the processor, the memory, can be implemented on a single chip, and in some embodiments, they can also be implemented on separate chips respectively. The processor 170 can also be used to generate corresponding operation control signals to send to the corresponding components of the computing processing device, read and process the data in the software, especially read and process the data and programs in the memory 120, so that each functional module therein executes the corresponding function, thereby controlling the corresponding components to act according to the requirements of the instructions.
[0084] The memory 120 can be used to store software codes related to the data processing method, and the processor 170 can execute the steps of the data processing method of the chip, or can also schedule other units (such as the above-mentioned input unit 130 and display unit 140) to realize the corresponding functions.
[0085] The RF unit 110 (optional) can be used to receive and send signals in the process of information or communication, for example, after receiving the downlink information of the base station, the processor 170 processes it; in addition, the uplink data is sent to the base station. Generally, the RF circuit includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF unit 110 can also communicate with network devices and other devices through wireless communication. The wireless communication can use any communication standard or protocol, including but not limited to global system for mobile communication (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), long term evolution (LTE), email, short messaging service (SMS), etc.
[0086] In the embodiments of the present application, the RF unit 110 can send images to the server 200 and receive the processing results sent by the server 200.
[0087] It should be understood that the RF unit 110 is optional, which can be replaced by other communication interfaces, for example, it can be a network interface.
[0088] The terminal 100 also includes a power supply 190 (such as a battery) for supplying power to each component. Preferably, the power supply can be logically connected to the processor 170 through a power management system, so as to realize the functions of managing charging, discharging, and power consumption management through the power management system.
[0089] The terminal 100 also includes an external interface 180, which can be a standard Micro USB interface, or a multi-pin connector, and can be used to connect the terminal 100 and other devices for communication, or can be used to connect a charger to charge the terminal 100.
[0090] Although not shown, the terminal 100 can also include a flash, a wireless fidelity (WiFi) module, a Bluetooth module, sensors with different functions, etc., which will not be described here. Some or all of the methods described below can be applied in the terminal 100 as shown. Figure 4
[0091] Next, the product form of the server 200 is described. Figure 4 The product form of the server 200 is described.
[0092] Figure 5 A structural diagram of the server 200 is provided, as shown in the figure. Figure 5 The server 200 includes a bus 201, a processor 202, a communication interface 203, and a memory 204. The processor 202, the memory 204, and the communication interface 203 communicate through the bus 201.
[0093] The bus 201 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 In the figure, only one thick line is used, but it does not mean that there is only one bus or one type of bus.
[0094] The processor 202 can be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), etc.
[0095] The memory 204 can include a volatile memory, such as a random access memory (RAM). The memory 204 can also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a mechanical hard drive (HDD), or a solid state drive (SSD).
[0096] The memory 204 can be used to store software code related to the data processing method, and the processor 202 can execute the steps of the chip data processing method or schedule other units to realize the corresponding functions.
[0097] It should be understood that the terminal 100 and the server 200 described above can be centralized or distributed devices, and the processors (for example, the processor 170 and the processor 202) in the terminal 100 and the server 200 can be hardware circuits (for example, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a general-purpose processor, a digital signal processing (DSP), a microprocessor, a microcontroller, or the like) or a combination of the hardware circuits. For example, the processor can be a hardware system with an instruction execution function, such as a CPU, a DSP, or the like, or a hardware system without an instruction execution function, such as an ASIC, an FPGA, or the like, or a combination of the hardware system without an instruction execution function and the hardware system with an instruction execution function.
[0098] It should be understood that the steps related to the model inference process in the embodiments of the present application involve AI-related operations. When performing AI operations, the instruction execution architecture of the terminal device and the server is not limited to the architecture of the processor combined with the memory described above. The following will be described in combination with Figure 6 The system architecture provided by the embodiments of the present application will be described in detail.
[0099] Figure 6 The system architecture provided by the embodiments of the present application will be described in detail. Figure 6 As shown in the figure, the system architecture 500 includes an execution device 510, a training device 520, a database 530, a client device 540, a data storage system 550, and a data acquisition system 560.
[0100] The execution device 510 includes a computing module 511, an I / O interface 512, a preprocessing module 513, and a preprocessing module 514. The target model / rule 501 can be included in the computing module 511, and the preprocessing module 513 and the preprocessing module 514 are optional.
[0101] The execution device 510 can be a terminal device or a server running an image processing application.
[0102] The data acquisition device 560 is used to acquire training samples. The training samples can be images, etc. After the training samples are acquired, the data acquisition device 560 stores the training samples in the database 530.
[0103] The training device 520 can obtain the target model / rule 501 by training a neural network (for example, a graph neural network in the embodiments of the present application) based on the training samples maintained in the database 530.
[0104] It should be understood that the training device 520 can perform a pre-training process on the neural network to be trained based on the training samples maintained in the database 530, or fine-tune the model based on the pre-training.
[0105] It should be noted that in actual application, the training samples maintained in the database 530 can not all be collected from the data collection device 560, but can also be received from other devices. In addition, it should be noted that the training device 520 can not completely train the target model / rule 501 based on the training samples maintained in the database 530, but can also obtain training samples from the cloud or other places for model training, and the above description should not be regarded as a limitation of the embodiments of the present application.
[0106] The target model / rule 501 trained by the training device 520 can be applied to different systems or devices, such as the execution device 510 shown in the figure. Figure 6 The execution device 510 can be a terminal such as a mobile phone terminal, a tablet computer, a notebook computer, an augmented reality (AR) / virtual reality (VR) device, a vehicle-mounted terminal, etc., and can also be a server, etc.
[0107] Specifically, the training device 520 can deliver the trained model to the execution device 510.
[0108] In the Figure 6 , the execution device 510 is configured with an input / output (I / O) interface 512 for data interaction with external devices, and a user can input data (such as images in the embodiments of the present application, etc.) to the I / O interface 512 through the client device 540.
[0109] The pre-processing module 513 and the pre-processing module 514 are used for pre-processing the input data received by the I / O interface 512. It should be understood that there can be no pre-processing module 513 and pre-processing module 514 or only one pre-processing module. When there is no pre-processing module 513 and pre-processing module 514, the input data can be directly processed by the calculation module 511.
[0110] During the pre-processing of the input data by the execution device 510, or during the calculation and other related processing of the calculation module 511 of the execution device 510, the execution device 510 can call the data, code, etc. in the data storage system 550 for corresponding processing, or store the data, instructions, etc. obtained by the corresponding processing in the data storage system 550.
[0111] Finally, the I / O interface 512 provides the processed results to the client device 540 and thus to the user.
[0112] exist Figure 6 In the illustrated case, the user can manually input data, and this "manual input data" can be operated through the interface provided by I / O interface 512. In another case, client device 540 can automatically send input data to I / O interface 512. If the automatic transmission of input data by client device 540 requires user authorization, the user can set the corresponding permissions in client device 540. The user can view the results output by execution device 510 on client device 540, and the specific presentation form can be a display, sound, action, etc. Client device 540 can also serve as a data acquisition terminal, collecting input data input into I / O interface 512 and output results from I / O interface 512 as new sample data and storing them in database 530. Of course, collection can also be performed without client device 540, and instead the I / O interface 512 directly stores the input data input into I / O interface 512 and output results from I / O interface 512 as new sample data in database 530.
[0113] It is worth noting that Figure 6 This is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, Figure 6 In the embodiment, the data storage system 550 is an external memory relative to the execution device 510. In other cases, the data storage system 550 can also be placed in the execution device 510. It should be understood that the execution device 510 can be deployed in the client device 540.
[0114] Next, a more detailed architecture of the execution entity that executes the data processing method in the embodiment of the present application is introduced.
[0115] The following combination Figure 6 The system architecture provided in the embodiments of the present application is introduced in detail. Figure 6 This is a schematic diagram of the system architecture provided in the embodiment of this application. Figure 6 As shown, the system architecture 500 includes an execution device 510 , a training device 520 , a database 530 , a client device 540 , a data storage system 550 , and a data collection system 560 .
[0116] The execution device 510 includes a calculation module 511, an I / O interface 512, a pre-processing module 513, and a post-processing module 514. The calculation module 511 may include the target model / rule 501, and the pre-processing module 513 and the post-processing module 514 are optional.
[0117] Data acquisition device 560 is used to collect training samples. Training samples can be images, text data, audio data, etc. In the embodiment of the present application, training samples are the data used to train multiple candidate neural networks. After collecting the training samples, data acquisition device 560 stores them in database 530.
[0118] It should be understood that a search space may also be maintained in the database 530 .
[0119] The training device 520 can construct multiple candidate neural networks based on the search space maintained in the database 530, and train the multiple candidate neural networks based on the training samples to search for the target model / rule 501. In the embodiment of the present application, the target model / rule 501 can be a target neural network.
[0120] It should be noted that, in actual applications, the training samples maintained in the database 530 may not all be collected by the data acquisition device 560, but may also be received from other devices. It should also be noted that the training device 520 may not train the target model / rule 501 entirely based on the training samples maintained in the database 530, but may also obtain training samples from the cloud or other places for model training. The above description should not be used as a limitation on the embodiments of the present application.
[0121] The target model / rule 501 obtained by training the training device 520 can be applied to different systems or devices, such as Figure 6 The execution device 510 shown may be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) / virtual reality (VR) device, a vehicle terminal, etc., or a server or cloud, etc.
[0122] Specifically, the training device 520 may transmit the target neural network to the execution device 510 .
[0123] exist Figure 6 In the embodiment, the execution device 510 is configured with an input / output (I / O) interface 512 for data interaction with an external device, and the user can input data (such as the data to be processed in the embodiment of the present application) to the I / O interface 512 through the client device 540.
[0124] The preprocessing modules 513 and 514 are configured to perform preprocessing on the input data received by the I / O interface 512. It should be understood that there can be no preprocessing modules 513 and 514 or only one preprocessing module. When there is no preprocessing module 513 and 514, the input data can be directly processed by the computing module 511.
[0125] During the preprocessing of the input data by the execution device 510 or the processing performed by the computing module 511 of the execution device 510, the execution device 510 can call data, code, etc. in the data storage system 550 for the corresponding processing, or store the data, instructions, etc. obtained by the corresponding processing in the data storage system 550.
[0126] Finally, the I / O interface 512 presents the processing result (for example, the data processing result in the embodiments of the present application) to the client device 540, thereby providing the user.
[0127] From the inference side of the model:
[0128] In the embodiments of the present application, the computing module 511 of the execution device 510 can obtain the code stored in the data storage system 550 to implement the data processing method in the embodiments of the present application.
[0129] In the embodiments of the present application, the computing module 511 of the execution device 510 can include hardware circuits (such as application specific integrated circuit (ASIC), field-programmable gate array (FPGA), general-purpose processor, digital signal processing (DSP), microprocessor or microcontroller, etc.) or a combination of these hardware circuits. For example, the training device 520 can be a hardware system with an execution instruction function, such as a CPU, a DSP, etc., or a hardware system without an execution instruction function, such as an ASIC, an FPGA, etc., or a combination of the hardware system without an execution instruction function and the hardware system with an execution instruction function.
[0130] Specifically, the computing module 511 of the execution device 510 can be a hardware system with an execution instruction function. The data processing method provided in the embodiments of the present application can be software code stored in a memory. The computing module 511 of the execution device 510 can obtain the software code from the memory and execute the obtained software code to implement the data processing method provided in the embodiments of the present application.
[0131] It should be understood that the computing module 511 of the execution device 510 can be a combination of a hardware system without an execution instruction function and a hardware system with an execution instruction function, and part of the steps of the data processing method provided in the embodiments of the present application can also be implemented by the hardware system without an execution instruction function in the computing module 511 of the execution device 510, which is not limited here.
[0132] From the training side of the model:
[0133] In the embodiments of the present application, the training device 520 can obtain the code stored in the storage (not shown in the figure) to implement the neural network search method in the embodiments of the present application. Figure 7
[0134] In the embodiments of the present application, the training device 520 can include a hardware circuit (such as an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, etc.), or a combination of these hardware circuits. For example, the training device 520 can be a hardware system with an execution instruction function, such as a CPU, a DSP, etc., or a hardware system without an execution instruction function, such as an ASIC, an FPGA, etc., or a combination of the hardware system without an execution instruction function and the hardware system with an execution instruction function.
[0135] Specifically, the training device 520 can be a hardware system with an execution instruction function, and the data processing method provided in the embodiments of the present application can be software code stored in the storage. The training device 520 can obtain the software code from the storage and execute the obtained software code to implement the neural network search method provided in the embodiments of the present application.
[0136] It should be understood that the training device 520 can be a combination of a hardware system without an execution instruction function and a hardware system with an execution instruction function, and part of the steps of the neural network search method provided in the embodiments of the present application can also be implemented by the hardware system without an execution instruction function in the training device 520, which is not limited here.
[0137] II. Server-provided image processing function type cloud service:
[0138] In a possible implementation, the server can provide the end side with a service of the image processing function through an application programming interface (API).
[0139] The terminal device can send relevant parameters (for example, image data) to the server through an API provided by the cloud, and the server can obtain a processing result based on the received parameters and return the processing result to the terminal.
[0140] The description of the terminal and the server can refer to the description of the above embodiments, which will not be repeated here.
[0141] As Figure 8 A flow of using a cloud service of the image processing function provided by a cloud platform is shown.
[0142] 1. Open and purchase the image processing service.
[0143] 2. The user can download a software development kit (SDK) corresponding to the image processing service. The cloud platform usually provides multiple development versions of the SDK for the user to select according to the needs of the development environment, for example, a JAVA version of the SDK, a python version of the SDK, a PHP version of the SDK, an Android version of the SDK, and the like.
[0144] 3. The user downloads the SDK of the corresponding version to the local according to the needs, imports the SDK project to the local development environment, configures and debugs in the local development environment, and can also develop other functions in the local development environment, so as to form an application that integrates the image processing function.
[0145] 4. During the use of the image processing function application, when the image processing function is needed, the API call of the image processing function can be triggered. When the application triggers the image processing function, an API request is initiated to the running instance of the image processing function service in the cloud environment, wherein the API request carries an image. The running instance in the cloud environment processes the image and obtains a processing result.
[0146] 5. The cloud environment returns the processing result to the application, thereby completing one image processing function call.
[0147] Since the embodiments of the present application involve the application of a large number of neural networks, in order to facilitate understanding, the related terms and concepts related to neural networks involved in the embodiments of the present application will be introduced first.
[0148] (1) Neural network
[0149] The neural network can be composed of neural units, which can be operation units taking xs and intercept 1 as inputs, and the output of the operation units can be:
[0150]
[0151] where s = 1, 2, … n, n is a natural number greater than 1, Ws is the weight of xs, b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of the activation function can be used as the input of the next layer of convolutional layer, and the activation function can be a sigmoid function. The neural network is a network formed by connecting a plurality of the above single neural units, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of a plurality of neural units.
[0152] (2) Loss function
[0153] In the process of training a deep neural network, because it is desired that the output of the deep neural network is as close as possible to the value that is truly intended to be predicted, the weight vector of each layer of the neural network can be updated according to the difference between the predicted value of the current network and the truly intended target value by comparing the predicted value of the current network with the truly intended target value (of course, there is usually an initialization process before the first update, that is, the parameters of each layer in the deep neural network are pre-configured), for example, if the predicted value of the network is too high, the weight vector is adjusted to make it predict lower, and the adjustment is continuously made until the deep neural network can predict the truly intended target value or a value very close to the truly intended target value. Therefore, it is necessary to define in advance "how to compare the difference between the predicted value and the target value", which is the loss function or the objective function, which is an important equation for measuring the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference, and then the training of the deep neural network becomes a process of trying to minimize the loss.
[0154] (3) Backpropagation algorithm
[0155] The convolutional neural network can adopt a back propagation (BP) algorithm to correct the size of the parameters in the initial super-resolution model in the training process, so that the reconstruction error loss of the super-resolution model becomes smaller and smaller. Specifically, the forward transmission of the input signal until the output generates an error loss, and the error loss information is propagated backward to update the parameters in the initial super-resolution model, so as to make the error loss converge. The back propagation algorithm is a back propagation movement dominated by error loss, aiming to obtain the optimal parameters of the super-resolution model, such as the weight matrix.
[0156] (4) Deep neural network
[0157] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with many layers of hidden layers, where "many" has no special measurement standard. From the position of different layers of the DNN, the neural network inside the DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the number of layers in between is the hidden layer. The layers are fully connected, that is, any neuron in the i-th layer is connected to any neuron in the i+1-th layer. Although the DNN looks very complex, in terms of the work of each layer, it is actually not complex. In simple terms, it is expressed as the following linear relationship expression: wherein, is an input vector, is an output vector, is a bias vector, W is a weight matrix (also called a coefficient), and a() is an activation function. Each layer only performs the following simple operation on the input vector to obtain the output vector Since the DNN has many layers, the number of coefficients W and bias vectors is also large. These parameters in the DNN are defined as follows: taking the coefficient W as an example: assuming that in a three-layer DNN, the linear coefficient of the fourth neuron in the second layer to the second neuron in the third layer is defined as The superscript 3 represents the layer number of the coefficient W, and the subscripts correspond to the output third layer index 2 and the input second layer index 4. In summary, the coefficient of the k-th neuron in the L-1-th layer to the j-th neuron in the L-th layer is defined as It is noted that the input layer is without W parameters. In deep neural networks, more hidden layers allow the network to better capture complex situations in the real world. In theory, the more parameters a model has, the higher its complexity and the greater its "capacity" to perform more complex learning tasks. Training a deep neural network is a process of learning weight matrices, and the ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network (weight matrices formed by vectors W of many layers).
[0158] The purpose of super-resolution (SR) is to construct a high-resolution image from a low-resolution image. SR work usually applies deep neural network knowledge learned from high-resolution training images to construct missing details in low-resolution inputs.
[0159] When performing an image processing task (e.g., an SR task) through a graph neural network, graph construction is required. During graph construction, different nodes in an image can be connected and fused. In the prior art, the same operation paradigm is performed on each node, that is, the importance of different nodes is considered to be the same. For example, the number of fused nodes is the same for each node during graph construction, which results in a large overall computing power expenditure of the graph neural network.
[0160] To solve the above problems, an embodiment of a data processing method is provided in the present application. Referring to Figure 8 , Figure 8 An embodiment of a data processing method provided in the present application is shown as follows: Figure 11 As shown in the figure, the data processing method provided in the present application can include:
[0161] 901, an image is obtained, the image including a first node and a second node, the first node and the second node being different pixel points or image blocks on the image;
[0162] The image can be an image that needs to perform an image processing task.
[0163] The image can include a plurality of nodes. The nodes herein can correspond to pixel points or image blocks (patches) composed of a plurality of pixel points (for example, a 2*2 pixel patch). The present application is not limited thereto.
[0164] 902、through the graph neural network, processing the image to obtain a processing result; the processing result includes a first fusion result and a second fusion result, the first fusion result is obtained by aggregating the information of the first node and the information of M nodes on the image, and the second fusion result is obtained by aggregating the information of the second node and the information of N nodes on the image, wherein the influence degree of the first node on the accuracy of the image processing task is greater than the influence degree of the second node on the accuracy of the image processing task, and M is greater than N.
[0165] Among them, the graph neural network can perform feature extraction on the image to obtain the feature representation of each node. Specifically, the graph neural network can include multiple blocks, and each block corresponds to a stage. For example, referring to the MGB module (or GAL module) in Figure 11 , Figure 9 , each block can obtain the information (such as feature representation) of each node, and then connect and fuse the nodes. For example, for a node, the information of a certain number of other nodes in the image that are similar to the node can be fused into the node to update the information of the node, thereby completing the information interaction between the nodes.
[0166] For example, in a possible implementation, the graph neural network includes multiple blocks, the first fusion result is obtained by aggregating the information of the first node and the information of M nodes on the image through a target block in the multiple blocks, and the second fusion result is obtained by aggregating the information of the second node and the information of N nodes on the image through the target block. That is, the same block of the graph neural network uses different aggregation quantities when aggregating different nodes.
[0167] In the prior art, the same operation paradigm is performed on each node, that is, the importance of different nodes is considered to be the same. In this case, the number of nodes for fusion is the same for different nodes (or the fusion quantity is independent of the importance of the node).
[0168] For example, for convolution, the same convolution kernel scans all pixels of the feature map; for attention mechanism, each pixel needs to aggregate information from a fixed number of pixels in a fixed size neighborhood; for graph neural network, any pixel needs to aggregate K neighboring nodes around it. The previous super-resolution scheme is too rigid and does not fully utilize the characteristics of super-resolution "restoring imbalance".
[0169] However, the importance of different nodes on the image (i.e., the degree of influence on the execution effect or execution accuracy of the image processing task) may be different for image processing tasks. For example, in an embodiment of the present application, the image includes a first node and a second node, the fusion quantity used for fusion of node information of the first node is M, the fusion quantity used for fusion of node information of the second node is N, the importance of the first node for the image processing task is greater than that of the second node, and therefore M is set to be greater than N. In this case, for nodes that are not very important, the fusion quantity is set to be lower (compared to nodes with higher importance), which does not greatly affect the processing result of the image processing task, and at the same time, the computational power consumption of the graph neural network is reduced.
[0170] For example, the super-resolution task has the feature of "restoring imbalance". In the process of super-resolution of a low-resolution image, the low-resolution part, which accounts for a large part of the image, does not need to be changed a lot; only a small amount of high-frequency detail part needs to be reconstructed by the neural network. Therefore, for nodes in the low-frequency region, a large number of nodes do not need to be fused, and a good super-resolution effect can be ensured. The existing super-resolution technical solution fails to fully utilize the feature of "restoring imbalance" in super-resolution. The present application designs a graph with varying node degrees (in an embodiment of the present application, the node degree can also be referred to as the fusion quantity), so that the graph neural network used for the super-resolution task can pay more attention to the part of the image that needs high-frequency reconstruction.
[0171] Next, the construction of the graph in an embodiment of the present application is specifically introduced. In the embodiment, each pixel on the feature map can be taken as a node of the constructed graph. It is assumed that the dimension of the feature map is R H*W*C , and there are a total of H*W nodes; each node is an R C feature. It should be noted that the embodiment is also applicable to the case where an image patch (for example, a patch composed of 2*2 pixels) is used as a node.
[0172] In a possible implementation, taking the super-resolution task as an example of the image processing task, it is necessary to accurately determine which positions in the image belong to the high-frequency region and which positions belong to the low-frequency region. For example, referring to Figure 10 , the difference between the effect of down-sampling and up-sampling of the feature map and the original feature map can be used as an index for judging the high-frequency and low-frequency regions:
[0173]
[0174] In addition, the standard deviation of the feature can also be used as a standard for judging the importance of the node, and the embodiment of the present application is not limited thereto.
[0175] It should be understood that the feature map can also be processed using traditional edge detection operators, such as Sobel, Laplacian, Canny, Prewitt, etc.
[0176] In a possible implementation, a total aggregation value related to the current computing power and a total influence degree of the nodes on the image on the accuracy of the image processing task can also be obtained, the value of M is determined according to the total aggregation value and a relationship between the influence degree of the first node on the accuracy of the image processing task and the total influence degree, and the value of N is determined according to the total aggregation value and a relationship between the influence degree of the second node on the accuracy of the image processing task and the total influence degree. That is, when determining the fusion quantity of the nodes, a total number can be determined based on the current computing power requirement, and then the fusion quantity of each node is allocated based on the total number under the condition that the total number is determined, so that the computing overhead of the graph neural network can be adapted to the computing power requirement.
[0177] For example, the sum of the degrees (that is, the fusion quantity) of each node of the graph (budget, that is, the total aggregation value) can be set as Deg based on the current computing power, and the degree Deg(v) of each node v can be obtained as follows:
[0178] deg(v) = D F (v) Deg / ∑D F ;
[0179] Wherein, Df(v) is the influence degree of the node, Deg is the total aggregation value, the accumulation of DF represents the total influence degree, and deg(v) is the degree of the node.
[0180] It should be noted that the allocation of the node degree can not be allocated according to the above formula, for example, the importance index can be normalized by softmax before the node degree is allocated.
[0181] In a possible implementation, the graph neural network includes a first block and a second block; the nodes in the image can be aggregated by the first block; wherein the aggregation object of the node when the first aggregation is performed is selected from the nodes included in one continuous image region; the nodes in the image are aggregated by the second block; wherein the aggregation object of the node when the second aggregation is performed is selected from the nodes in the multiple regions of the image, and there is a gap between adjacent regions in the multiple regions.
[0182] The existing graph neural network image processing algorithm cannot be directly applied to the bottom vision field: previous graph algorithms often take small image patches as nodes of the graph; and the bottom vision has a higher requirement for the reconstruction of pixels, and the aggregation of patch nodes will affect the image quality. Therefore, a single pixel can be taken as a node of the graph to ensure the image quality of super-resolution reconstruction. The problem brought by taking a single pixel as a node of the graph is that, in the field of bottom vision, the resolution of the image to be processed is large, and if the pixel is directly taken as the node of the graph in the process of constructing the graph, the search space of similar nodes is too large, which will cause a large algorithmic cost. Therefore, the embodiment of the present application designs a sampling method, which can significantly reduce the search space of similar nodes, so as to control the algorithmic cost of graph construction. Specifically, for each node, the nodes can be collected in the local pixels and the global pixels and in a smaller sampling space, so as to significantly reduce the algorithmic cost of graph construction.
[0183] For each node, it is necessary to find and connect similar nodes in the image to complete the construction of the graph. However, since the super-resolution task often needs to process images with large resolution, when the pixel is taken as the node of the graph, it often leads to too many graph nodes, the space of searching similar nodes is too large, and a large amount of algorithmic cost is consumed. To solve this problem, the embodiment of the present application performs global and local sampling on all nodes for each node, so that each node searches for similar nodes in a smaller sampling space (rather than globally). Then in the sampling space, according to the node degree Deg(v) set in the previous step, the Deg(v) most similar nodes are selected and connected with the current node v.
[0184] For example, local sampling can be to collect all nodes around the current node (for example, refer to the middle graph of Figure 10 ).
[0185] For example, global sampling can be to collect every N pixels in the entire image range (for example, refer to the right graph of Figure 11 ).
[0186] When performing node aggregation, according to the constructed graph, each node can perform weighted aggregation on the nodes connected thereto. Let v be the current node; u be the critical node of v; f k (u,v) be a learnable neural network for measuring the similarity of u and v; be the K-th layer of the IPG network, and the feature of the u node, then the node aggregation can be represented as:
[0187]
[0188] Refer to Figure 11 , Figure 12An example of a structure of a model is shown with the bottom visual network SwinIR as an example.
[0189] 903. Perform the image processing task according to the processing result.
[0190] Referring to Table 1, Table 1 is an example of an experimental result based on an embodiment of the present application.
[0191] Table 1
[0192]
[0193] Referring to Figure 12 , Figure 12 An example of a data processing apparatus provided by an embodiment of the present application is shown as follows. Figure 13 As shown, the data processing apparatus 1200 provided by an embodiment of the present application can include:
[0194] The acquisition module 1201 is configured to acquire an image, the image including a first node and a second node, the first node and the second node being different pixel points or image blocks on the image.
[0195] The description of the acquisition module 1201 can refer to the description of step 901 in the above embodiments, which will not be repeated here.
[0196] The processing module 1202 is configured to process the image by a graph neural network to obtain a processing result, the processing result including a first fusion result and a second fusion result, the first fusion result being obtained by aggregating information of the first node and information of M nodes on the image, the second fusion result being obtained by aggregating information of the second node and information of N nodes on the image, wherein the influence degree of the first node on the accuracy of the image processing task is greater than the influence degree of the second node on the accuracy of the image processing task, and M is greater than N; and the image processing task is performed according to the processing result.
[0197] The description of the processing module 1202 can refer to the description of steps 902 and 903 in the above embodiments, which will not be repeated here.
[0198] In a possible implementation, the image processing task is super-resolution, image denoising, image segmentation, or image recognition.
[0199] In a possible implementation, the first node is a node in a high-frequency region of the image relative to the second node.
[0200] In a possible implementation, the graph neural network comprises a plurality of blocks, the first fusion result is obtained by aggregating information of the first node and information of M nodes on the image through a target block in the plurality of blocks, and the second fusion result is obtained by aggregating information of the second node and information of N nodes on the image through the target block.
[0201] In a possible implementation, the processing module 1202 is further configured to:
[0202] obtain a total aggregation value and a total influence degree of the nodes on the image on the accuracy of the image processing task, the total aggregation value being related to the current computing power;
[0203] determine the value of M according to the total aggregation value and a relationship between the influence degree of the first node on the accuracy of the image processing task and the total influence degree;
[0204] determine the value of N according to the total aggregation value and a relationship between the influence degree of the second node on the accuracy of the image processing task and the total influence degree.
[0205] In a possible implementation, the graph neural network comprises a first block and a second block.
[0206] The processing module 1202 is specifically configured to:
[0207] perform first aggregation on the nodes in the image through the first block, wherein the aggregation object of the nodes in the first aggregation is selected from the nodes included in one continuous image region;
[0208] perform aggregation on the nodes in the image through the second block, wherein the aggregation object of the nodes in the second aggregation is selected from the nodes in a plurality of regions of the image, and there is a gap between adjacent regions in the plurality of regions.
[0209] Next, an execution device provided by an embodiment of the present application is introduced. Please refer to Figure 13 , Figure 13 A structural schematic diagram of an execution device provided by an embodiment of the present application, the execution device 1300 can be a virtual reality (VR) device, a mobile phone, a tablet computer, a notebook computer, a smart wearable device, a monitoring data processing device, a server, or the like, which is not limited here. Specifically, the execution device 1300 comprises a receiver 1301, a transmitter 1302, a processor 1303, and a memory 1304 (wherein the number of processors 1303 in the execution device 1300 can be one or more, Figure 14The processor 1303 can include an application processor 13031 and a communication processor 13032, for example. In some embodiments of the present application, the receiver 1301, the transmitter 1302, the processor 1303 and the memory 1304 can be connected through a bus or other means.
[0210] The memory 1304 can include read-only memory and random access memory, and provide the processor 1303 with instructions and data. A portion of the memory 1304 can also include non-volatile random access memory (NVRAM). The memory 1304 stores processor and operating instructions, executable modules or data structures, or a subset thereof, or an expanded set thereof, wherein the operating instructions can include various operating instructions for implementing various operations.
[0211] The processor 1303 controls the operation of the execution device. In a specific application, various components of the execution device are coupled together through a bus system, which can include a data bus in addition to a power bus, a control bus, and a status signal bus, etc. However, for the sake of clarity, all the buses are referred to as a bus system in the figure.
[0212] The method disclosed in the embodiments of the present application can be applied to the processor 1303 or implemented by the processor 1303. The processor 1303 can be an integrated circuit chip having a signal processing capability. In the implementation process, the steps of the above method can be completed by the integrated logic electric circuit or the instruction of the software form in the processor 1303. The processor 1303 described above can be a general processor, a digital signal processor (DSP), a microprocessor or a microcontroller. Further, the processor 1303 can further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The processor 1303 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general processor can be a microprocessor or the processor can also be any conventional processor or the like. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware code processor to execute, or be executed by a combination of hardware and software modules in the code processor. The software module can be located in a random access memory, a flash memory, a read only memory, a programmable read only memory or an electrically erasable programmable memory, a register or other mature storage medium in the art. The storage medium is located in the storage 1304, and the processor 1303 reads the information in the storage 1304 and combines the hardware to complete the steps of the above method.
[0213] The receiver 1301 can be used to receive input digital or character information, and generate signal input related to the relevant settings and function control of the execution device. The transmitter 1302 can be used to output digital or character information; the transmitter 1302 can also be used to send instructions to the disk set to modify the data in the disk set.
[0214] In the embodiments of the present application, in one case, the processor 1303 is used to execute the data processing method executed by the execution device in the above embodiments.
[0215] The embodiments of the present application also provide a training device, please refer to Figure 14 , Figure 12 is a structural schematic diagram of the training device provided by the embodiments of the present application, and the training device 1400 can be deployed with Figure 8The apparatus described in the corresponding embodiments, in particular, the training device 1400 is implemented by one or more servers, and the training device 1400 can be different in configuration or performance, and can include one or more central processing units (CPUs) 1414 (for example, one or more processors) and a memory 1432, one or more storage media 1430 (for example, one or more mass storage devices) storing an application 1442 or data 1444. The memory 1432 and the storage medium 1430 can be temporary storage or persistent storage. The program stored in the storage medium 1430 can include one or more modules (not shown in the figure), and each module can include a series of instruction operations in the training device. Further, the central processing unit 1414 can be configured to communicate with the storage medium 1430 and execute the series of instruction operations in the storage medium 1430 on the training device 1400.
[0216] The training device 1400 can further include one or more power supplies 1426, one or more wired or wireless network interfaces 1450, one or more input / output interfaces 1458, or one or more operating systems 1441, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and the like.
[0217] In the embodiments of the present application, the central processing unit 1414 is configured to execute the method in the corresponding embodiments. Figure 15
[0218] In the embodiments of the present application, a computer program product is also provided, which, when running on a computer, causes the computer to perform the steps performed by the data processing apparatus described above, or causes the computer to perform the steps performed by the data processing apparatus described above.
[0219] In the embodiments of the present application, a computer readable storage medium is also provided, which stores a program for signal processing, and when running on a computer, causes the computer to perform the steps performed by the data processing apparatus described above, or causes the computer to perform the steps performed by the data processing apparatus described above.
[0220] The execution device, training device or terminal device provided in the embodiments of the present application can specifically be a chip, and the chip includes: a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin or a circuit. The processing unit can execute the computer execution instructions stored in the storage unit to enable the chip in the execution device to execute the data processing method described in the above embodiment, or to enable the chip in the training device to execute the data processing method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit can also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0221] For details, please refer to Figure 15 , This is a schematic diagram of the structure of a chip provided in an embodiment of the present application. This chip can be represented as a neural network processor NPU 1500. NPU 1500 is mounted on the host CPU as a coprocessor and is assigned tasks by the host CPU. The core of the NPU is arithmetic circuit 1503, which is controlled by controller 1504 to extract matrix data from memory and perform multiplication operations.
[0222] In some implementations, arithmetic circuit 1503 includes multiple processing units (PEs). In some implementations, arithmetic circuit 1503 is a two-dimensional systolic array. Arithmetic circuit 1503 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, arithmetic circuit 1503 is a general-purpose matrix processor.
[0223] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 1502 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 1501 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 1508.
[0224] The unified memory 1506 is used to store input data and output data. The weight data is transferred to the weight memory 1502 through a Direct Memory Access Controller (DMAC) 1505. The input data is also transferred to the unified memory 1506 through the DMAC.
[0225] The BIU is a Bus Interface Unit 1510 for the interaction between the AXI bus and the DMAC and the Instruction Fetch Buffer (IFB) 1509.
[0226] The BIU 1510 is used for the instruction fetch buffer 1509 to fetch instructions from the external memory, and is also used for the DMAC 1505 to fetch the original data of the input matrix A or the weight matrix B from the external memory.
[0227] The DMAC is mainly used to transfer the input data in the external memory DDR to the unified memory 1506, or to transfer the weight data to the weight memory 1502, or to transfer the input data to the input memory 1501.
[0228] The vector calculation unit 1507 includes a plurality of operation processing units, and further processes the output of the operation circuit as needed, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / full connection layer network calculation in neural networks, such as Batch Normalization, pixel-level summation, upsampling of feature planes, etc.
[0229] In some implementations, the vector calculation unit 1507 can store the processed output vector to the unified memory 1506. For example, the vector calculation unit 1507 can apply a linear function; or, a nonlinear function to the output of the operation circuit 1503, such as linear interpolation on the feature planes extracted by the convolutional layer, and further, for example, a vector of accumulated values to generate activation values. In some implementations, the vector calculation unit 1507 generates normalized values, pixel-level summation values, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1503, for example, for use in subsequent layers in the neural network.
[0230] The controller 1504 is connected to the instruction fetch buffer 1509, which is used to store instructions used by the controller 1504;
[0231] The unified memory 1506, the input memory 1501, the weight memory 1502, and the instruction memory 1509 are on-chip memories. The external memory is private to the NPU hardware architecture.
[0232] Any processor mentioned in the above can be a general central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling execution of the above programs.
[0233] It should be noted that the apparatus embodiments described above are merely illustrative, and the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Some or all of the modules can be selected to achieve the purpose of the embodiment according to actual needs. In addition, the connection relationship between the modules in the apparatus embodiment provided in the present application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.
[0234] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and necessary general hardware, and of course can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure for implementing the same function can also be various, such as analog circuits, digital circuits, or special circuits. However, for the present application, software program implementation is a better embodiment. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., including a plurality of instructions for making a computer device (which can be a personal computer, a training device, or a network device, etc.) execute the methods described in various embodiments of the present application.
[0235] In the above embodiments, all or part can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, it can be implemented in the form of a computer program product in whole or in part.
[0236] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, training device or data center to another website, computer, training device or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be stored by the computer or a data storage device such as a training device, a data center, etc. integrated with one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
Claims
1. A data processing method, characterized by, The method comprises: obtaining an image, the image comprising a first node and a second node, the first node and the second node being different pixel points or image blocks on the image; processing the image by a graph neural network to obtain a processing result, the processing result comprising a first fusion result and a second fusion result, the first fusion result being obtained by aggregating information of the first node and information of M nodes on the image, the second fusion result being obtained by aggregating information of the second node and information of N nodes on the image, wherein the first node has a greater influence on the accuracy of an image processing task than the second node, and M is greater than N; performing the image processing task according to the processing result.
2. The method of claim 1, wherein, The image processing task is super-resolution, image denoising, image segmentation, or image recognition.
3. The method according to claim 1 or 2, characterized in that, The first node is a node in a high-frequency region of the image relative to the second node.
4. The method according to any one of claims 1 to 3, characterized in that, The graph neural network comprises a plurality of blocks, the first fusion result being obtained by aggregating information of the first node and information of M nodes on the image by a target block in the plurality of blocks, and the second fusion result being obtained by aggregating information of the second node and information of N nodes on the image by the target block.
5. The method according to any one of claims 1 to 4, characterized in that, The method further comprises: obtaining a total aggregation value and a total influence of nodes on the image on the accuracy of the image processing task, the total aggregation value being related to a current computing power; determining the value of M according to the total aggregation value and a relationship between the influence of the first node on the accuracy of the image processing task and the total influence; determining the value of N according to the total aggregation value and a relationship between the influence of the second node on the accuracy of the image processing task and the total influence.
6. The method according to any one of claims 1 to 5, characterized in that, The graph neural network comprises a first block and a second block. The processing of the image by the graph neural network to obtain a processing result comprises: performing first aggregation on nodes in the image by the first block, wherein the aggregation objects of the nodes during the first aggregation are selected from nodes included in a continuous image region; performing aggregation on nodes in the image by the second block, wherein the aggregation objects of the nodes during the second aggregation are selected from nodes in a plurality of regions of the image, and there is a gap between adjacent regions in the plurality of regions.
7. A data processing apparatus, characterized by The apparatus comprises: an obtaining module configured to obtain an image, the image comprising a first node and a second node, the first node and the second node being different pixel points or image blocks on the image; A processing module is used to process the image through a graph neural network to obtain a processing result; the processing result includes a first fusion result and a second fusion result, the first fusion result is obtained by aggregating the information of the first node and the information of M nodes on the image, and the second fusion result is obtained by aggregating the information of the second node and the information of N nodes on the image, wherein the influence of the first node on the accuracy of the image processing task is greater than the influence of the second node on the accuracy of the image processing task, and M is greater than N; according to the processing result, the image processing task is executed.
8. The apparatus of claim 7, wherein, The image processing task is super-resolution, image denoising, image segmentation or image recognition.
9. The apparatus of claim 7 or 8, wherein, The first node is a node in a high-frequency area relative to the second node in the image.
10. The apparatus of any one of claims 7 to 9, wherein, The graph neural network includes multiple blocks, the first fusion result is obtained by aggregating the information of the first node and the information of M nodes on the image through a target block among the multiple blocks, and the second fusion result is obtained by aggregating the information of the second node and the information of N nodes on the image through the target block.
11. The apparatus of any one of claims 7 to 10, wherein, The processing module is further configured to: Obtaining a total aggregate value and a total influence of the nodes on the image on the accuracy of the image processing task, wherein the total aggregate value is related to the current computing power; Determining a value of M according to the total aggregate value and a relationship between the degree of influence of the first node on the accuracy of the image processing task and the total degree of influence; The value of N is determined according to the total aggregate value and the relationship between the degree of influence of the second node on the accuracy of the image processing task and the total degree of influence.
12. The apparatus of any one of claims 7 to 11, wherein, The graph neural network includes a first block and a second block; The processing module is specifically used to: Performing a first aggregation on the nodes in the image through the first block; wherein, when performing the first aggregation, the nodes to be aggregated are selected from nodes included in a continuous image region; The nodes in the image are aggregated through the second block; wherein, when performing the second aggregation, the aggregation objects of the nodes are selected from nodes in multiple areas of the image, and there are gaps between adjacent areas in the multiple areas.
13. A computing device, comprising: The computing device comprises at least one processor and at least one memory, wherein the at least one memory stores computer-readable instructions; the at least one processor executes the computer-readable instructions to enable the computing device to perform the method according to any one of claims 1 to 6.
14. A computer-readable storage medium, characterized in that, The method comprises computer-readable instructions; the computer-readable instructions are used to implement the method according to any one of claims 1 to 6.
15. A computer program product, characterised in that, The method comprises computer-readable instructions; the computer-readable instructions are used to implement the method according to any one of claims 1 to 6.
16. A chip, characterized by comprising at least one processing unit and an interface circuit for providing program instructions or data to the at least one processing unit, the at least one processing unit being configured to execute the program instructions to implement the method of any one of claims 1 to 6.