Data processing chips, modules, terminals, data processing methods and devices
By introducing a mailbox control unit to monitor the amount of intermediate processing data and calling the neural network processing subunit for calculation when the target amount of data is reached, the problem of the NPU waiting for a frame of image data to be processed is solved, resulting in faster data processing speed and lower power consumption.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
- Filing Date
- 2021-11-10
- Publication Date
- 2026-05-26
AI Technical Summary
In existing technologies, the neural network processing unit (NPU) needs to wait for a frame of image data to be processed before it can continue processing, resulting in long data processing delays and slow speed.
A mailbox control unit is introduced to monitor the amount of data processed in the intermediate stage. When the target amount of data is reached, the neural network processing subunit is invoked to perform calculations, thereby reducing waiting time.
It improves data processing speed, reduces neural network processing waiting time, and lowers power consumption.
Smart Images

Figure CN116128031B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a data processing chip, module, terminal, and data processing method. Background Technology
[0002] In the field of artificial intelligence, neural networks are typically deployed on a Network Processing Unit (NPU). The NPU then processes the data to be processed. The process of processing image data using an NPU is generally as follows: Image data is received through a data transmission interface. After the data detection unit detects the received image data, it controls the Power Management Unit (PMU) to supply power and timing to the data processing system. The received image data is stored in memory. The image data processing subunit in the data processing system, after being powered on, reads the image data from memory, processes the read image data, writes the intermediate processed data back into memory, and then the NPU reads and processes the intermediate processed data written by the image data processing subunit.
[0003] In related technologies, the NPU is controlled by a microcontroller unit (MCU) to read data. Specifically, the MCU monitors the processing progress of the image data processing subunit. Whenever the image data processing subunit finishes processing a frame, the MCU calls the NPU to continue processing that frame.
[0004] In the aforementioned related technologies, the NPU can only be called for subsequent processing after the image data processing subunit has finished processing one frame of image. This causes the NPU to wait for a period of time before it can continue processing, resulting in a long delay in NPU data processing and a slow data processing speed. Summary of the Invention
[0005] This application provides a data processing chip, module, terminal, and data processing method, which can improve data processing speed. The technical solution is as follows:
[0006] On the one hand, a data processing chip is provided, the data processing chip comprising: a data processing subunit, a neural network processing subunit, and a mailbox control unit;
[0007] The mailbox control unit is used to record the first data volume of intermediate processing data generated by the data processing subunit, and based on the first data volume, control the neural network processing subunit to perform data processing.
[0008] On the other hand, a data processing module is provided, the data processing module including the data processing chip as described above.
[0009] On the other hand, a terminal is provided, the terminal including the data processing module as described above.
[0010] On the other hand, a data processing method is provided, which is applied to a data processing chip. The data processing chip includes a data processing subunit, a neural network processing subunit, and a mailbox control unit. The mailbox control unit is used to record a first data volume of intermediate processing data generated by the data processing subunit, and based on the first data volume, controls the neural network processing subunit to perform data processing. The method includes:
[0011] For any given frame of input image, obtain the image data of the input image;
[0012] Record the first data volume of the first intermediate processing data corresponding to the image data, wherein the first intermediate processing data is the data obtained by the image data processing subunit processing a portion of the image data;
[0013] In response to the first data volume matching the target data volume, the neural network processing subunit is invoked to perform neural network calculations on the first intermediate processed data.
[0014] On the other hand, a data processing apparatus is provided, the apparatus comprising:
[0015] The acquisition module is used to acquire image data of any input image frame;
[0016] The recording module is used to record the first data volume of the first intermediate processing data corresponding to the image data, wherein the first intermediate processing data is the data obtained by the image data processing subunit processing a portion of the image data in the image data;
[0017] The calling module is used to call the neural network processing subunit to perform neural network calculations on the first intermediate processing data in response to the matching of the first data volume with the target data volume.
[0018] On the other hand, a computer-readable storage medium is provided that stores at least one piece of program code, the at least one piece of program code being executed by a processor to implement the data processing method as described above.
[0019] On the other hand, a computer program product is also provided, which stores at least one piece of program code, said at least one piece of program code being loaded and executed by a processor to implement the data processing method described above.
[0020] In this embodiment of the application, the first intermediate processing data generated is monitored. When the first data volume of the first intermediate processing data meets the required target data volume, the neural network processing subunit can be called to perform neural network calculation. Thus, the neural network processing subunit does not need to wait for all frames of images to be processed by the image data processing subunit, and can start neural network calculation in advance, reducing the waiting time and thereby improving the data processing speed. Attached Figure Description
[0021] Figure 1 A schematic diagram of the structure of an image data processing chip provided in an exemplary embodiment of this application is shown;
[0022] Figure 2 This illustration shows a schematic diagram of the structure of a data packet provided in an exemplary embodiment of this application;
[0023] Figure 3 A schematic diagram of the structure of an image data processing chip provided in an exemplary embodiment of this application is shown;
[0024] Figure 4 A schematic diagram of the structure of an image data processing chip provided in an exemplary embodiment of this application is shown;
[0025] Figure 5 A flowchart illustrating a data processing method provided in an exemplary embodiment of this application is shown;
[0026] Figure 6 A flowchart illustrating a data processing method provided in an exemplary embodiment of this application is shown;
[0027] Figure 7 A flowchart illustrating a data processing method provided in an exemplary embodiment of this application is shown;
[0028] Figure 8 A flowchart illustrating a data processing method provided in an exemplary embodiment of this application is shown;
[0029] Figure 9 This invention provides a structural block diagram of a data processing apparatus according to an exemplary embodiment of the present application.
[0030] Figure 10 A structural block diagram of a terminal provided in one embodiment of this application is shown. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0032] In this document, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the image data and other data involved in this application are data authorized by the user or fully authorized by all parties.
[0033] To extend the terminal's standby time, the terminal can operate in a low-power mode when the user is not using it. In low-power mode, the terminal powers down unused modules while maintaining power to modules that require Always-On (AON) mode. For example, in low-power mode, the terminal powers the Wireless Fidelity (Wi-Fi) module and authentication module, while powering down the video playback module and screen brightness module.
[0034] For example, if the terminal is a mobile phone, the user can lock the screen when not using the phone; this locked state is a low-power state. With the development of terminal technology, mobile phones can now collect facial images or fingerprint information while the screen is locked. The identity verification module then verifies the collected facial images or fingerprint information to unlock the phone. Therefore, the information collection module and the identity verification module need to be powered on while the phone is locked; that is, they must be in AON mode.
[0035] In this embodiment of the application, in order to further save the power consumption of the terminal in AON mode, the terminal monitors the first data volume of intermediate processing data generated by the image data processing subunit. In response to the first data volume matching the target data volume, the neural network processing subunit is called to perform neural network calculation on the intermediate processing data. In this way, the neural network processing subunit can start processing the intermediate processing data in advance without waiting for a whole frame of image processing to be completed, thereby reducing the waiting time of the neural network and improving the data processing speed.
[0036] Please refer to Figure 1 This illustrates a schematic diagram of the structure of a data processing chip provided in an exemplary embodiment of this application. See also... Figure 1 The data processing chip includes a data processing subunit 11, a neural network processing subunit 12, and a mailbox control unit 13. The mailbox control unit 13 is used to record a first data volume of intermediate processing data generated by the data processing subunit 11, and control the neural network processing subunit 12 to perform data processing based on the first data volume.
[0037] The mailbox control unit 13 is connected to the neural network processing subunit 12 and the data processing subunit 11, respectively.
[0038] Here, "data volume" refers to the amount of intermediate processing data. The target data volume represents the minimum amount of data that the neural network processing subunit 12 reads each time. In this embodiment, the data volume of the generated intermediate processing data is monitored by the mailbox control unit 13, so that when the data volume of the intermediate processing data reaches the target data volume, the neural network processing subunit 12 can be invoked to process the intermediate processing data. In this way, the neural network processing subunit 12 does not need to wait for the data processing subunit 11 to finish processing a frame of image data before it can start neural network calculation, thereby reducing the waiting time of the neural network processing subunit 12, thus improving the processing speed of the image data processing system, reducing the image processing time of the image data processing system, and reducing the power consumption of the image data processing system.
[0039] The data processing subunit 11 is a unit capable of processing image data. For example, this data processing subunit 11 is an image signal processor (ISP). When processing image data, the data processing subunit 11 reads the image data and processes it. The neural network processing subunit 12 is a processing unit capable of deploying a neural network. For example, this neural network processing subunit 12 is a neural network processing unit (NPU). When processing image data, the neural network processing subunit 12 can call network parameters and intermediate processing data, and process the intermediate processing data based on the network parameters.
[0040] It should be noted that the number of neural network processing sub-units 12 can be one or more, and this is not specifically limited in this embodiment. For example, the number of neural network processing sub-units 12 can be one, two, or four. When there are multiple neural network processing sub-units 12, different neural networks are deployed in each of the multiple neural network processing sub-units 12, and the network parameters of the multiple neural networks are stored in the same first storage unit 12. Alternatively, each neural network processing sub-unit 12 corresponds to one first storage unit 12, and the network parameters of the neural networks deployed in each neural network processing sub-unit 12 are stored in the corresponding first storage unit 12.
[0041] Another point to note is that the image data read by the data processing subunit 11 is the original image data, or the image data read by the data processing subunit 11 is the image data processed by the neural network processing subunit 12.
[0042] Accordingly, the neural network processing subunit 12 is used to implement an intermediate process in the entire image data processing process. The data processing subunit 11 can continue to process the data processing results processed by the neural network processing subunit 12. For example, the neural network processing subunit 12 is used to denoise the image data. Accordingly, the data processing subunit 11 preprocesses the image data to obtain intermediate processing data, the neural network processing subunit 12 denoises the intermediate processing data, and the image processing result obtained from the denoising process is rewritten to the buffer unit, and the data processing subunit 11 continues to process the denoised image processing result.
[0043] The mailbox control unit 13 is a mailbox mechanism (MB) control unit. Accordingly, data is transmitted between the neural network processing subunit 12 and the data processing subunit 11 through the MB control unit in a data packet mechanism, thereby realizing the scheduling of the neural network processing subunit 12 and the data processing subunit 11.
[0044] When the MB control unit controls the neural network processing subunit 12 and the data processing subunit 11, it can schedule them in the form of data packets. Correspondingly, the MB control unit monitors the amount of intermediate processing data generated by the data processing subunit 11, and generates data packets based on the data information of this intermediate processing data. This data information includes the amount of intermediate processing data, its storage address, etc. Based on this data information, the target fields in the data packets are encoded to obtain data packets representing the data information of the intermediate processing data.
[0045] See Figure 2 The data packet includes multiple fields, such as the data source field (ISP_NPU), data size field (Date Size), data address field (Data Addr), and data attribute field (CMD). ISP_NPU indicates whether the data originates from an ISP or an NPU. For example, if the ISP_NPU encoding is 0001, it indicates that the intermediate processing data comes from an ISP; conversely, if the encoding is 0010, it indicates that the intermediate processing data comes from an NPU. It should be noted that the correspondence between the encoded content and the data source can be set as needed, but this embodiment does not impose specific limitations on it.
[0046] Date Size indicates the amount of intermediate processing data. For example, if Date Size is encoded as 0010, it means the amount of intermediate processing data is 8 bits. It should be noted that the correspondence between the encoded content and the data amount can be set as needed; in this embodiment, no specific limitation is made. Data Addr indicates the storage address of the intermediate processing data. For example, if Data Addr is encoded as 0010, it means the intermediate processing data is stored in the first sub-storage unit. CMD indicates the data attribute of the intermediate processing data. For example, if CMD is encoded as 011, it means the intermediate processing data is readable and writable. The data packet may also include other fields, such as an End Of File (EOF) field to indicate the end of the data packet content.
[0047] In this embodiment, the mailbox control unit 13 records intermediate processing data information through data packets. This allows the data volume and other data information of the intermediate processing data to be represented simply by setting different fields in the data packets. Therefore, when the data processing subunit 11 and the neural network processing subunit 12 interact with the mailbox control unit 13, they can interact based on data packets, thereby reducing the resources occupied by data transmission between the mailbox control unit 13 and the data processing subunit 11 and the neural network processing subunit 12, and thus reducing the energy consumption of the mailbox control unit 13 in interacting with the data processing subunit 11 and the neural network processing subunit 12.
[0048] In some embodiments, the mailbox control unit 13 is used to monitor the processing speed of the neural network processing subunit 12; in response to the processing speed of the neural network processing subunit 12 being less than the processing speed of the data processing subunit 11, it controls the data processing subunit 11 to reduce its processing speed; the data processing subunit 11 is used to process the image data based on the reduced processing speed. Accordingly, the mailbox control unit is used to determine the amount of data consumed by the neural network processing subunit 12 when processing the first intermediate processing data; based on the amount of data consumed, it determines a first processing speed of the neural network processing subunit 12; in response to the first processing speed being less than a second processing speed, it controls the image data processing subunit 11 to reduce the second processing speed, the second processing speed being the processing speed at which the image data processing subunit 11 processes the image data.
[0049] The mailbox control unit 13 determines the processing speed of the neural network processing subunit 12 and the data processing subunit 11 based on the amount of data read by each subunit; alternatively, the mailbox control unit 13 determines the processing speed of the neural network processing subunit 12 and the data processing subunit 11 based on the amount of data generated by each subunit. In some embodiments, the mailbox control unit 13 monitors the amount of data read by the neural network processing subunit 12 each time it processes data, and determines the processing speed of the neural network processing subunit 12 based on the amount of data read by the neural network processing subunit 12 per unit time. The mailbox control unit 13 monitors the amount of intermediate processing data generated by the data processing subunit 11 per unit time, and determines the processing speed of the data processing subunit 11 based on the amount of intermediate processing data generated by the data processing subunit 11.
[0050] In this embodiment, the mailbox control unit 13 controls the processing speed of the data processing subunit 11 in reverse according to the processing speed of the neural network processing subunit 12. This indirectly realizes the control of the second storage unit 13, thereby preventing the second storage unit 13 from having insufficient storage space, which would cause intermediate processing data to be unable to be stored, and ensuring that the image data processing process can proceed normally.
[0051] Please refer to Figure 3 This illustrates a schematic diagram of the structure of a data processing chip provided in an exemplary embodiment of this application. See also... Figure 3 The data processing chip further includes: a power management unit 14, a first cache unit 15, and a second cache unit 16; the power management unit 14 is used to control the power supply to the first cache unit 15 and the second cache unit 16; the first cache unit 15 and the second cache unit 16 are used to cache data; the power management unit 14 is used to control the power supply to the first cache unit 15 and not to the second cache unit 16 in response to a first state of the data processing subunit 11, and to control the power supply to the first cache unit 15 and the second cache unit 16 in response to a second state of the data processing subunit 11.
[0052] See Figure 3 The power management unit 14 remains operational in AON mode, enabling it to control the power supply to other working units as needed. The power management unit 14 controls the power supply to the first buffer unit 15 and the second buffer unit 16 in different states of the data processing unit 14.
[0053] The power management unit 14 is connected to the first cache unit 15 and the second cache unit 16, respectively. The first cache unit 15 and the second cache unit 16 are connected to the data processing subunit 11 and the neural network processing subunit 12, respectively. Accordingly, the power management unit 14 is configured to de-supply the data processing subunit 11 and the neural network processing subunit 12 in response to not performing data processing, and to supply power to the data processing subunit 11 and the neural network processing subunit 12 in response to performing data processing. Alternatively, it may maintain power supply to the data processing subunit 11 and the neural network processing subunit 12.
[0054] In some embodiments, the power management unit 14 is also connected to the power supply module of the terminal, which is connected to the first cache unit 15 and the second cache unit 16 respectively. Accordingly, the power management unit 14 controls the power supply module to supply power to the first cache unit 15 or the second cache unit 16.
[0055] In some embodiments, the first state includes a state where no data is processed, and the second state includes a state where data is processed. For example, the data processing chip is a chip that performs image data processing in an authentication module. Accordingly, the data processing chip is used to process image data of the acquired face image or fingerprint. Then, the first state is the state when the terminal is in a locked screen state and no unlock operation is detected, and the second state is the state when the terminal is in a locked screen state and an unlock operation is detected.
[0056] In some embodiments, the first state includes a state of training a neural network model. The second state includes a state of processing data based on the neural network model.
[0057] For example, the first cache unit 15 is used to cache the model parameters of the neural network model. In the first state, the neural network model is trained, and the model parameters of the trained neural network model are stored in the first cache unit 15. In the second state, the terminal processes data using the cached model parameters in the first cache unit 15.
[0058] Based on the characteristics of the first cache unit 15 and the second cache unit 16, the data stored in the first cache unit 15 and the second cache unit 16 will be cleared after power failure. In order to ensure that the data processing chip can call the data in the first cache unit 15 for data processing at any time, it is necessary to ensure that the data stored in the first cache unit 15 is not cleared. In this embodiment of the application, in both the first state and the second state, the power management unit 14 maintains power supply to the first cache unit 15. In this way, even if other working units are powered off in the second state in AON mode, it can be ensured that the data in the first cache unit 15 will not be cleared due to power failure. Thus, in AON mode, only some cache units are powered, reducing the number of working units powered and reducing the power consumption in AON mode.
[0059] The first cache unit 15 and the second cache unit 16 are cache units in the memory used to store data. The storage space size of the first cache unit 15 and the second cache unit 16 can be set as needed, and the storage space sizes of the first cache unit 15 and the second cache unit 16 may be the same or different. In this embodiment, no specific limitation is made. The memory can be any type of volatile memory. For example, the memory can be static random access memory (SRAM).
[0060] In some embodiments, the first cache unit 15 and the second cache unit 16 are different units of the same volatile memory. Accordingly, the cache units of the volatile memory are divided, and the power management unit 14 supplies power to the different units of the volatile memory respectively.
[0061] In some embodiments, the first cache unit 15 and the second cache unit 16 are different units of different volatile memories. In some embodiments, the first cache unit 15 is a cache unit of the volatile memory of the data processing module where the data processing subunit 11 and the neural network processing subunit 12 are located. Alternatively, the second cache unit 16 is a cache unit in volatile memory outside the data processing module where the data processing subunit 11 or the neural network processing subunit 12 is located.
[0062] In this embodiment, a volatile memory is provided inside the data processing module where the data processing subunit 11 or the neural network processing subunit 12 is located. The storage unit of the volatile memory inside the data processing module where the data processing subunit 11 and the neural network processing subunit 12 are located is determined as the first cache unit 15. In this way, when the data processing subunit 11 or the neural network processing subunit 12 needs to read the data in the first cache unit 15, it only needs to read it from the internal volatile memory, thereby saving the time of internal data transfer and reducing the power consumption of reading data.
[0063] In some embodiments, the data processing subunit 11 and the neural network processing subunit 12 are used to process data using a neural network model in a second state. The first cache unit 15 is used to store the data of the neural network model, and the second cache unit 16 is used to store the data to be processed, intermediate processing data, and processing result data of the data processing subunit 11 and the neural network processing subunit 12.
[0064] The data of the neural network model includes the model parameters of the neural network model. The data processing subunit 11 or the neural network processing subunit 12 is used to perform neural network calculations based on the data to be processed and intermediate processing data stored in the first cache unit 15 and the second cache unit 16, to obtain the processing result data, and write the processing result data to the second cache unit 16.
[0065] In this embodiment, to ensure that the data processing chip can access the neural network model data in the first cache unit 15 for data processing at any time, it is necessary to ensure that the data stored in the first cache unit 15 is not cleared. In this embodiment, in both the first and second states, the power management unit 14 maintains power supply to the first cache unit 15. Thus, in AON mode, even if power is cut off to other working units in the second state, it can be ensured that the data in the first cache unit 15 will not be cleared due to power failure. In this way, in AON mode, only some cache units are powered, reducing the number of working units powered and lowering the power consumption in AON mode.
[0066] In some embodiments, the second cache unit 16 includes a first sub-cache unit and a second sub-cache unit, the data processing sub-unit 11 is used to write the intermediate processing data into the first sub-cache unit, and the neural network processing sub-unit 12 is used to write the processing result data into the second sub-cache unit.
[0067] That is, the first sub-cache unit is used to store the intermediate processing data written by the data processing sub-unit 11, and the second sub-cache unit is used to store the processing result data written by the neural network processing sub-unit 12. In this embodiment, by further dividing the second cache unit 16, the division of the second cache unit 16 is made more refined, thereby enabling partitioned power supply to the second cache unit 16 and further saving the energy consumption of the second cache unit 16.
[0068] It should be noted that when the first cache unit 15 and the second cache unit 16 are storage units in different static memories, the first cache unit 15 is a storage unit in the static memory integrated in the neural network processing subunit 141.
[0069] In some embodiments, in response to the processing speed of the neural network processing subunit 12 being less than the processing speed of the data processing subunit 11, the mailbox control unit 13 controls the data processing subunit 11 to reduce its processing speed. In some embodiments, in response to the processing speed of the neural network processing subunit 12 being less than the processing speed of the data processing subunit 11, the mailbox control unit 13 determines unprocessed intermediate processing data in the second cache unit 16; in response to the amount of intermediate processing data in the second cache unit 16 being greater than a first preset data amount, the mailbox control unit 13 controls the data processing subunit 11 to reduce its processing speed. Alternatively, in response to the processing speed of the neural network processing subunit 12 being less than the processing speed of the data processing subunit 11, the mailbox control unit 13 determines the remaining storage space in the second cache unit 16; in response to the amount of data that can be stored in the remaining storage space not exceeding a second preset data amount, the control unit controls the data processing subunit to reduce its processing speed.
[0070] It should be noted that the power management unit 11 controls the supply voltage to the data processing unit 14 based on its data processing speed; the supply voltage is positively correlated with the data processing speed. Accordingly, the power management unit 11 controls the supply voltage to the data processing unit 14 based on Dynamic Voltage and Frequency Scaling (DVFS) technology. This achieves NPU design reuse in non-AON mode, while employing DVFS technology in AON mode to formulate different voltage strategies based on the NPU hardware configuration and the computing power of AON mode, ensuring that different AON modes operate at appropriate voltage levels and achieving optimal power consumption in the given scenario.
[0071] In some embodiments, see Figure 4The data processing chip further includes: a data transmission interface 17; the data transmission interface 17 is used to receive data; the data processing subunit 11 is used to switch from the first state to the second state in response to the data received by the data transmission interface 17 containing specific information. The specific information includes at least one of the following: motion detection information and human detection information.
[0072] The data transmission interface 17 remains operational in AON mode, ensuring timely transmission of image data to the data processing subunit 11. For example, in an authentication scenario, this interface 17 transmits image data acquired by the image acquisition module to the data processing subunit 11. It can also transmit data acquired by other modules in the terminal to the data processing subunit 11. For instance, this interface 17 may be a Mobile Industry Processor Interface (MIPI).
[0073] In the embodiments of this application, the power management unit maintains power supply to the first buffer unit in both the first and second states of the data processing unit, and only supplies power to the second buffer unit in the second state of the data processing unit. In this way, in AON mode, when in the first state, only the first buffer unit needs to be powered, thereby reducing the number of units that need to be powered and reducing the power consumption in AON mode.
[0074] Please refer to Figure 5 This illustrates a data processing method provided by an exemplary embodiment of this application, see [link to example]. Figure 5 The method includes:
[0075] Step S501: For any frame of input image, the terminal acquires the image data of that input image.
[0076] The input image data can be either the original image data or the image data processed by a neural network processing subunit. This application does not impose specific limitations on this aspect in its embodiments.
[0077] In some embodiments, the image data of the input image is the raw image data. Accordingly, the terminal caches the image data to a cache medium via a data transmission interface, and then reads the image data from the cache medium for processing. The neural network-based data processing system further includes a cache medium and a data transmission interface. The data transmission interface is connected to the cache medium, and the cache medium is connected to the image data processing subunit. The data transmission interface is used to receive the image data and write it to the cache medium. The cache medium is used to cache the image data. The image data processing subunit is used to read the image data from the cache medium.
[0078] In AON mode, the buffer medium and data transmission interface remain operational to ensure timely transmission of image data to the image data processing subunit. For example, in an authentication scenario, this data transmission interface transmits the image data corresponding to the face image captured by the image acquisition module to the image data processing subunit in the image data processing system. The buffer medium is a buffer within a buffer, which can be of any type, such as a buffer register. The data transmission interface also transmits data acquired by other modules in the terminal to the image data processing subunit. For example, this data transmission interface could be a Mobile Industry Processor Interface (MIPI). For instance, if the other module is a camera module, this data transmission interface connects the camera module and the image data processing subunit, transmitting images acquired by the camera module to the buffer medium so that the image data processing subunit can read image data from the buffer medium.
[0079] In some embodiments, the image data is image data processed by the neural network processing subunit. Accordingly, after the terminal processes the intermediate processing data through the neural network processing subunit, it stores the processing result in the storage unit. The image data processing subunit reads the image data from the storage unit and continues to process the image data.
[0080] For example, the neural network processing subunit is used to denoise image data. Correspondingly, the image data processing subunit preprocesses the image data to obtain intermediate processing data, denoises the intermediate processing data through the neural network processing subunit, and rewrites the denoised image processing result back into the storage unit, whereby the image data processing subunit continues to process the denoised image processing result.
[0081] Step S502: The terminal records the first data volume of the first intermediate processing data corresponding to the image data. The first intermediate processing data is the data obtained by the image data processing subunit processing a portion of the image data.
[0082] The first intermediate processing data is the intermediate processing data generated by the terminal image data processing subunit after processing the image data. The first data volume indicates the amount of intermediate processing data generated.
[0083] After processing the image data, the image data processing subunit stores the generated intermediate processing data in the storage unit. In this embodiment, the intermediate processing data is recorded by monitoring the image data processing subunit. Specifically, the terminal records the intermediate processing data via data packets. See also... Figure 6 This process is implemented through the following steps S5021-S5022, including:
[0084] Step S5021: The terminal monitors the first intermediate processing data generated by the image data processing subunit and obtains the first data volume of the first intermediate processing data.
[0085] After processing the image data, the image data processing subunit generates intermediate processing data and writes it into the storage unit. In this step, the terminal monitors the data information of the intermediate processing data generated by the image data processing subunit through the MB control unit. This data information includes the data volume, data source, and storage address of the first intermediate processing data.
[0086] Step S5022: The terminal encodes the first field of the data packet to be encoded based on the first data volume to obtain a data packet used to record the first data volume.
[0087] The first field is a field representing the amount of data in the data packet. The length and position of this first field in the data packet can be set as needed, but are not specifically limited in this embodiment. For example, see... Figure 2 The first field is the data size field. It should be noted that the correspondence between the encoded content of the field and the data size can be set as needed; however, this embodiment does not impose specific limitations on this. For example, see... Figure 2 If the Date Size is encoded as 0010, it means that the amount of data in the intermediate processing data is 8 bits.
[0088] In this embodiment, the control unit records intermediate processing data information through data packets. This allows the amount of intermediate processing data and other data information to be represented simply by setting different fields in the data packets. Therefore, when the image data processing subunit and the neural network processing subunit interact with the control unit, they can interact based on data packets, thereby reducing the resources occupied by data transmission between the control unit and the image data processing subunit and the neural network processing subunit, and thus reducing the energy consumption of data interaction between the control unit and the image data processing subunit and the neural network processing subunit.
[0089] Step S503: In response to the first data volume matching the target data volume, the terminal calls the neural network processing subunit to perform neural network calculations on the first intermediate processing data.
[0090] The target data volume represents the minimum amount of data read by the neural network processing subunit each time. This target data volume can be set as needed; however, in this embodiment, it is not specifically limited. The terminal executes step S203 in response to a match between the first data volume and the target data volume by configuring the MB control unit. Accordingly, the terminal configures the MB control unit before this step. In some embodiments, the terminal displays a setting interface, which includes a target setting option for setting the target data volume. The user inputs the target data volume based on this target setting option, and the terminal configures the MB control unit based on the input target data volume.
[0091] In this step, in response to the first data volume matching the target data volume, the terminal generates a data processing instruction, and calls the neural network processing subunit based on the data processing instruction. See also Figure 7 This process is implemented through the following steps S5031-S5032, including:
[0092] Step S5031: The terminal generates a data processing instruction based on the storage address of the first intermediate processing data.
[0093] This storage address is the address where the first intermediate processing data is stored in the storage unit. The data processing instructions are used to instruct the neural network processing subunit to begin processing the intermediate processing data.
[0094] In some embodiments, the storage address is a storage address recorded in a data packet. Accordingly, the terminal generates a data processing instruction based on the data packet used to record the storage address. Specifically, the terminal determines the storage address of the first intermediate processing data; and encodes the second field of the data packet to be encoded based on the storage address to obtain the data packet used to record the storage address.
[0095] The data packet may be the same as or different from the data packet in step S5022; this embodiment does not specifically limit this. If the data packet is the same as the data packet in step S5022, the data packet includes a first field and a second field, representing the data volume and storage address, respectively. The positions of the first and second fields in the data packet can be set as needed; this embodiment does not specifically limit this.
[0096] Step S5032: The terminal sends the data processing instruction to the neural network processing subunit, which triggers the neural network processing subunit to perform neural network calculations on the first intermediate processing data corresponding to the storage address.
[0097] In this step, the terminal sends a data processing instruction to the neural network processing subunit through the MB control unit, so that the neural network processing subunit can read the first intermediate processing data from the storage unit based on the data processing instruction, and then perform neural network calculations on the first intermediate processing data.
[0098] In the application embodiment, the storage address of intermediate processing data is recorded through data packets. This allows the storage address and other data information of the first intermediate processing data to be represented simply by setting different fields in the data packets. Therefore, when the image data processing subunit and the neural network processing subunit interact with the MB control unit, they can interact based on data packets, thereby reducing the resources occupied by data transmission between the MB control unit and the image data processing subunit and the neural network processing subunit, and further reducing the energy consumption of data interaction between the MB control unit and the image data processing subunit and the neural network processing subunit.
[0099] It should be noted that after the terminal records the first data volume of the first intermediate processing data corresponding to the image data, the image data processing subunit continues to process other data in the image data, see [link to relevant documentation]. Figure 8 The process includes:
[0100] (1) The terminal continues to process other image data in the image data through the image data processing subunit to obtain the second intermediate processing data.
[0101] This step is the same as step S502, in which the image data processing subunit processes the image data to obtain the first intermediate processing data, and will not be repeated here.
[0102] (2) Based on the second intermediate processing data, the terminal continues to execute the second data volume of recording the second intermediate processing data corresponding to the image data.
[0103] The principle of recording the amount of data for the first intermediate processing data in this step is the same as that in step S502, and will not be repeated here.
[0104] In this implementation, while the terminal calls the neural network processing subunit to perform neural network calculations on the first terminal data, it continues to process the image data through the image data processing subunit. Thus, after the neural network processing subunit finishes processing the first intermediate processing data, if the amount of the second intermediate processing data matches the target data amount, it can directly read the second intermediate processing data for processing. This keeps the image data processing subunit and the neural network processing subunit processing data simultaneously, reducing the waiting time of the neural network unit, further improving the data processing speed, and thus reducing the terminal's power consumption in AON mode.
[0105] In some embodiments, the terminal adjusts the second processing speed of the image data processing subunit based on the first processing speed of the neural network processing subunit, so that the first processing speed and the second processing speed are matched. This process is achieved through the following steps (i)-(iii), including:
[0106] (i) The terminal determines the amount of data consumed when the neural network processing subunit processes the first intermediate processing data.
[0107] The terminal monitors the amount of first intermediate processing data consumed by the neural network processing subunit through the MB control module. The terminal records the amount of intermediate processing data read by the neural network processing subunit from the storage unit through the MB control module, and uses this amount as the amount of data consumed by the neural network unit. Specifically, the MB module records the amount of first intermediate processing data read by the neural network processing subunit from the storage unit each time, or the amount of first intermediate processing data read by the neural network processing subunit from the storage unit per unit time. In this embodiment, no specific limitation is made.
[0108] (ii) The terminal determines the first processing speed of the neural network processing subunit based on the amount of data consumed.
[0109] The terminal also records the time taken to consume the first intermediate processing data, and uses the ratio of data volume to time as the first processing speed of the neural network processing subunit.
[0110] (iii) In response to the first processing speed being less than the second processing speed, the image data processing subunit is controlled to reduce the second processing speed, the second processing speed being the processing speed at which the image data processing subunit processes the image data.
[0111] Before this step, the terminal also determines a second processing speed. The principle by which the terminal determines the second processing speed is the same as the principle by which it determines the first processing speed, and will not be repeated here.
[0112] In some embodiments, in response to a first processing speed being less than a second processing speed, the terminal controls the image data processing subunit to reduce its processing speed. In some embodiments, the terminal controls the second processing speed in conjunction with the amount of intermediate processing data that can be stored in the storage unit. This process involves: the terminal determining a third data amount, which is the maximum amount of intermediate processing data that the storage unit can store; and reducing the second processing speed based on the third data amount and the first processing speed. Correspondingly, in response to a first processing speed being less than a second processing speed, the terminal determines unprocessed intermediate processing data in the storage unit. In response to the amount of intermediate processing data in the storage unit being greater than a first preset data amount, the control unit controls the image data processing subunit to reduce its processing speed. Alternatively, in response to a first processing speed being less than a second processing speed, the terminal determines the amount of data that can still be stored in the unoccupied storage units of the storage unit. In response to the amount of data that can still be stored in the unoccupied storage units of the storage unit being no greater than a second preset data amount, the control unit controls the image data processing subunit to reduce its processing speed.
[0113] In this embodiment, the terminal controls the processing speed of the image data processing subunit in reverse based on the first processing speed of the neural network processing subunit. This indirectly controls the storage unit, thereby preventing insufficient storage space in the storage unit from causing intermediate processing data to be unable to be stored, and ensuring that the image data processing process can proceed normally.
[0114] It should be noted that the power management unit controls the supply voltage to the image data processing subunit based on its data processing speed; this supply voltage is positively correlated with the data processing speed. Accordingly, the power management unit uses Dynamic Voltage and Frequency Scaling (DVFS) technology to control the supply voltage to the image data processing subunit. This achieves NPU design reuse in non-AON mode, while employing DVFS technology in AON mode to formulate different voltage strategies based on the NPU hardware configuration and the computing power of AON mode, ensuring that different AON modes operate at appropriate voltage levels and achieving optimal power consumption in various scenarios.
[0115] In this embodiment of the application, the first intermediate processing data generated is monitored. When the first data volume of the first intermediate processing data meets the required target data volume, the neural network processing subunit can be called to perform neural network calculation. Thus, the neural network processing subunit does not need to wait for all frames of images to be processed by the image data processing subunit, and can start neural network calculation in advance, reducing the waiting time and thereby improving the data processing speed.
[0116] Please refer to Figure 9 This diagram illustrates a structural block diagram of a neural network-based data processing apparatus according to an embodiment of this application. This neural network-based data processing apparatus can be implemented as all or part of a processor through software, hardware, or a combination of both. The apparatus includes:
[0117] The acquisition module 901 is used to acquire the image data of any input image frame;
[0118] Recording module 902 is used to record the first data volume of the first intermediate processing data corresponding to the image data, wherein the first intermediate processing data is the data obtained by the image data processing subunit processing a portion of the image data in the image data;
[0119] Module 903 is invoked in response to the first data volume matching the target data volume, and the neural network processing subunit is invoked to perform neural network calculations on the first intermediate processed data.
[0120] In some embodiments, the recording module 902 includes:
[0121] The monitoring unit is used to monitor the first intermediate processing data generated by the image data processing subunit and obtain the first data volume of the first intermediate processing data.
[0122] An encoding unit is used to encode the first field of the data packet to be encoded based on the first data amount, so as to obtain a data packet for recording the first data amount.
[0123] In some embodiments, the calling module 903 includes:
[0124] The generation unit is used to generate data processing instructions based on the storage address of the first intermediate processing data;
[0125] The sending unit is used to send the data processing instruction to the neural network processing subunit, which triggers the neural network processing subunit to perform neural network calculations on the first intermediate processing data corresponding to the storage address.
[0126] In some embodiments, the device further includes:
[0127] The first determining module is used to determine the storage address of the first intermediate processing data;
[0128] The encoding module is used to encode the second field of the data packet to be encoded based on the storage address, so as to obtain the data packet used to record the storage address.
[0129] In some embodiments, the device further includes:
[0130] The data processing module is used to further process other image data in the image data through the image data processing subunit to obtain the second intermediate processing data;
[0131] The recording module 902 is also used to continue recording a second amount of data based on the second intermediate processing data corresponding to the image data.
[0132] In some embodiments, the device further includes:
[0133] The second determining module is used to determine the amount of data consumed when the neural network processing subunit processes the first intermediate processing data;
[0134] The third determining module is used to determine the first processing speed of the neural network processing subunit based on the amount of data consumed.
[0135] The control module is configured to control the image data processing subunit to reduce the second processing speed in response to the first processing speed being less than the second processing speed, wherein the second processing speed is the processing speed at which the image data processing subunit processes the image data.
[0136] In some embodiments, the control module includes:
[0137] A determining unit is used to determine a third data quantity, which is the maximum data quantity that the storage unit can store for the intermediate processing data.
[0138] The reduction unit is used to reduce the second processing speed based on the third data volume and the first processing speed.
[0139] In this embodiment of the application, the first intermediate processing data generated is monitored. When the first data volume of the first intermediate processing data meets the required target data volume, the neural network processing subunit can be called to perform neural network calculation. Thus, the neural network processing subunit does not need to wait for all frames of images to be processed by the image data processing subunit, and can start neural network calculation in advance, reducing the waiting time and thereby improving the data processing speed.
[0140] Please refer to Figure 10 This diagram illustrates a structural block diagram of a terminal 1000 provided in an exemplary embodiment of this application. The terminal 1000 can be a smartphone, computer, tablet computer, or other terminal with neural network deployment capabilities. The terminal 1000 in this application includes a data processing module 1010 as described in the above embodiments, which includes the data processing chip shown in the embodiments of this application.
[0141] In some embodiments, the terminal 1000 further includes one or more of the following components: a processor 1020 and a memory 1030.
[0142] The processor 1020 may include one or more processing cores. The processor 1020 connects to various parts within the terminal 1000 using various interfaces and lines, and performs various functions and processes data of the terminal 1000 by running or executing program code, programs, code sets, or program code sets stored in the memory 1030, and by calling data stored in the memory 1030. Optionally, the processor 1020 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 1020 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), Neural-network Processing Unit (NPU), and modem. Specifically, the CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display on the screen; the NPU is used to implement Artificial Intelligence (AI) functions; and the modem is used for wireless communication. It is understandable that the aforementioned modem may not be integrated into the processor 1020, but may be implemented as a separate chip.
[0143] The memory 1030 may include random access memory (RAM) or read-only memory. Optionally, the memory 1030 may include a non-transitory computer-readable storage medium. The memory 1030 may be used to store program code, programs, code, code sets, or program code sets. The memory 1030 may include a program storage area and a data storage area, wherein the program storage area may store program code for implementing an operating system, program code for at least one function (such as touch function, sound playback function, image playback function, etc.), program code for implementing the various method embodiments described below, etc.; the data storage area may store data created according to the use of the terminal 1000 (such as audio data, phone book, etc.).
[0144] In some embodiments, the terminal 1000 further includes a display screen. The display screen is a display component used to display a user interface. Optionally, the display screen is a touch-enabled display screen, through which users can perform touch operations on the display screen using their fingers, styluses, or any suitable object. The display screen is typically located on the front panel of the terminal 1000. The display screen can be designed as a full-screen, curved screen, irregularly shaped screen, dual-sided screen, or foldable screen. The display screen can also be designed as a combination of a full-screen and a curved screen, or a combination of an irregularly shaped screen and a curved screen, etc., which are not limited in this embodiment.
[0145] In addition, those skilled in the art will understand that the structure of the terminal 1000 shown in the above figures does not constitute a limitation on the terminal 1000. The terminal 1000 may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the terminal 1000 may also include a microphone, speaker, radio frequency circuit, input unit, sensor, audio circuit, Wireless Fidelity (Wi-Fi) module, power supply, Bluetooth module, etc., which will not be described in detail here.
[0146] This application also provides a computer-readable storage medium storing at least one piece of program code, which is loaded and executed by the processor to implement the data processing methods shown in the above embodiments.
[0147] This application also provides a computer program product that stores at least one piece of program code, which is loaded and executed by the processor to implement the data processing methods shown in the above embodiments.
[0148] In some embodiments, the computer program involved in the present application embodiments may be deployed and executed on a computer device, or executed on multiple computer devices located in one location, or executed on multiple computer devices distributed in multiple locations and interconnected through a communication network. Multiple computer devices distributed in multiple locations and interconnected through a communication network may constitute a blockchain system.
[0149] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more program codes or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of computer programs from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0150] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A data processing chip, characterized in that, The data processing chip includes: an image data processing subunit, a neural network processing subunit, and a mailbox control unit; The mailbox control unit is used to record the first data volume of intermediate processing data generated by the image data processing subunit. The intermediate processing data is the data obtained by the image data processing subunit from processing a portion of the image data. In response to the first data volume matching the target data volume, the neural network processing subunit is invoked to perform neural network calculations on the intermediate processing data. The target data volume represents the minimum data volume read by the neural network processing subunit each time.
2. The data processing chip according to claim 1, characterized in that, The mailbox control unit is used to determine the amount of data consumed by the neural network processing subunit when processing the intermediate processing data; based on the amount of data consumed, determine a first processing speed of the neural network processing subunit; and in response to the first processing speed being less than a second processing speed, control the image data processing subunit to reduce the second processing speed, where the second processing speed is the processing speed of the image data processing subunit in processing the image data.
3. A data processing module, characterized in that, The data processing module includes the data processing chip as described in any one of claims 1-2.
4. A terminal, characterized in that, The terminal includes the data processing module as described in claim 3.
5. A data processing method, characterized in that, The method is applied to a data processing chip, which includes an image data processing subunit, a neural network processing subunit, and a mailbox control unit. The mailbox control unit records a first data volume of first intermediate processing data generated by the image data processing subunit, and controls the neural network processing subunit to perform data processing based on the first data volume. The method includes: For any given frame of input image, obtain the image data of the input image; Record the first data volume of the first intermediate processing data corresponding to the image data, wherein the first intermediate processing data is the data obtained by the image data processing subunit from processing a portion of the image data; In response to the first data volume matching the target data volume, the neural network processing subunit is invoked to perform neural network calculations on the first intermediate processed data, where the target data volume represents the minimum amount of data read by the neural network processing subunit each time.
6. The method according to claim 5, characterized in that, The first data volume of the first intermediate processing data corresponding to the image data includes: Monitor the first intermediate processing data generated by the image data processing subunit to obtain the first data volume of the first intermediate processing data; Encode the first field of the data packet to be encoded based on the first data volume to obtain a data packet for recording the first data volume.
7. The method according to claim 5, characterized in that, The step of invoking the neural network processing subunit to perform neural network calculations on the first intermediate processed data includes: Based on the storage address of the first intermediate processing data, generate data processing instructions; The data processing instruction is sent to the neural network processing subunit, and the data processing instruction is used to trigger the neural network processing subunit to perform neural network calculations on the first intermediate processing data corresponding to the storage address.
8. The method according to claim 7, characterized in that, The method further includes: Determine the storage address of the first intermediate processing data; The second field of the data packet to be encoded is encoded based on the storage address to obtain a data packet for recording the storage address.
9. The method according to claim 5, characterized in that, After recording the first data volume of the first intermediate processing data corresponding to the image data, the method further includes: The image data processing subunit continues to process other image data in the image data to obtain second intermediate processing data. Based on the second intermediate processing data, the second data volume corresponding to the second intermediate processing data of the image data is recorded.
10. The method according to claim 5, characterized in that, The method further includes: Determine the amount of data consumed by the neural network processing subunit when processing the first intermediate processing data; Based on the amount of data consumed, the first processing speed of the neural network processing subunit is determined; In response to the first processing speed being less than the second processing speed, the image data processing subunit is controlled to reduce the second processing speed, where the second processing speed is the processing speed at which the image data processing subunit processes the image data.
11. The method according to claim 10, characterized in that, The control of the image data processing subunit to reduce the second processing speed includes: Determine a third data volume, which is the maximum amount of data that the storage unit can store for the first intermediate processing data; Based on the third data volume and the first processing speed, the second processing speed is reduced.
12. A data processing apparatus, characterized in that, The device includes: The acquisition module is used to acquire image data of any input image frame; The recording module is used to record the first data volume of the first intermediate processing data corresponding to the image data, wherein the first intermediate processing data is the data obtained by the image data processing subunit processing a portion of the image data in the image data; The calling module is used to call the neural network processing subunit to perform neural network calculations on the first intermediate processing data in response to the first data volume matching the target data volume, wherein the target data volume represents the minimum amount of data read by the neural network processing subunit each time.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one piece of program code, which is executed by a processor to implement the data processing method as described in any one of claims 5 to 11.
14. A computer program product, characterized in that, The computer program product stores at least one piece of program code, which is loaded and executed by a processor to implement the data processing method as described in any one of claims 5 to 11.