A data processing device and method for image data

By introducing data processing methods of decoding, channel filling and extraction modules into neural network accelerators, the problem of insufficient computing flexibility when processing multiple data forms is solved, and unified processing of data forms and flexible processing of channel data is realized.

CN114554218BActive Publication Date: 2025-06-06AXERA TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210149411.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-18
Publication Date
2025-06-06
Estimated Expiration
2042-02-18

AI Technical Summary

Technical Problem

Neural network accelerators have poor computing flexibility when processing multiple data forms and cannot effectively process data in different data forms.

Method used

A data processing device and method for image data are designed, including a decoding module, a channel filling module and an extraction module. The decoding module converts the data into the form required by the accelerator hardware, the channel filling module adjusts the number of channels, and the extraction module extracts the data of the target channel.

Benefits of technology

Through the processing of these modules, the data form before data input into the convolution operation unit is compliant, which improves the data processing flexibility of the neural network accelerator.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114554218B_ABST
    Figure CN114554218B_ABST
Patent Text Reader

Abstract

The present application provides a data processing device and method for image data. Image data may exist in different bit formats in memory, such as 8-bit, 10-bit, 12-bit or 14-bit formats, etc., and accelerator hardware is designed based on a fixed data format, and a decoding module is used to uniformly process the data into the data format required by the accelerator hardware. The data output by the image sensor may have different numbers of channels, and the neural network model uniformly adopts the format of A2 channels. Using the data processing device for image data provided by the present application, image data in a first format with A1 channels can be processed into third data in a second format with A2 channels. For some image data, when the amount of data contained in one of its channels is large, the data of the target channel is usually first processed by de-mosaicing and raw domain noise reduction, and the data of the target channel is extracted by an extraction module so that the accelerator hardware can process individual channels separately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data processing device and method for image data. Background Art

[0002] As more and more applications such as consumer electronics, automotive electronics, and industrial control are introduced into artificial intelligence, artificial intelligence is facing unprecedented rapid development, and technologies such as deep learning and neural networks have reached a climax. At the same time, in order to meet the increasing demand for neural network computing power, research institutions and companies have begun to continuously develop neural network accelerators for use in security, autonomous driving and other fields. Neural network accelerators are mainly responsible for convolution operations, but before performing convolution operations, the data format needs to be processed uniformly.

[0003] However, neural network accelerators have poor computational flexibility. Since data storage comes in a variety of forms, neural network accelerators are unable to process data in a variety of data forms. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a data processing device and method for image data, which is used to add a level of data pre-processing device before performing neural network processing so that the data format before inputting into the convolution operation unit is compliant.

[0005] An embodiment of the present application provides a data processing device for image data, comprising:

[0006] A decoding module, used for acquiring first data about the image to be input in the memory, and converting the first data into second data according to the data format required by the accelerator hardware;

[0007] A channel filling module is used to fill the second data with the third data of the second format of A2 channels when the second data is the image data of the first format of A1 channels; wherein A1 and A2 are both positive integers greater than 0, and A1 <A2;

[0008] The extraction module is used to extract the data of the target channel from the A2 channels of the third data to form the fourth data.

[0009] In the above technical solution, image data is usually in the form of a tightly arranged 8-bit in memory, that is, each pixel is in the form of 8 bits. However, considering that the input image may be in the form of 10 bits, 12 bits, 14 bits, etc., and when the neural network model is processing, the hardware is designed based on a fixed data form. The decoding module can uniformly process the data into the data form required by the accelerator hardware. Since the data output by some image sensors has different numbers of channels, for example, the neural network model uniformly adopts the format of A2 channels, the channel filling module can process the image data in the first format with A1 channels into the third data in the second format with A2 channels. For some image data, when the amount of data contained in a certain channel is large, usually the data of the target channel is first processed such as demosaic (removing mosaic) and noise reduction in the raw domain. The extraction module can extract the data of the target channel so that the neural network accelerator hardware can separately process individual channels.

[0010] In some alternative embodiments, during the process of converting the first data into the second data, the decoding module is specifically configured to:

[0011] Convert the received B1-bit first data into B2-bit second data, where both B1 and B2 are positive integers greater than 0, and B1 < B2; and

[0012] Fill the blank bits in the B2-bit second data with null values.

[0013] In the above technical solution, when the decoding module converts the first data into the second data, it adopts the method of filling null values into the first data to obtain the second data.

[0014] In some alternative embodiments, the decoding module includes a decoding cache module, a null value supplement module, and a decoding memory;

[0015] When performing the data conversion from B1 bit to B2 bit, the decoding cache module is configured to cache the data of C pixel points input, and at each clock unit, output the data of D / B2 pixel points to the null value supplement module; where C, D, and D / B2 are all positive integers greater than 0, D is the total bit width of the data bus, and it satisfies the relationship of B1×C < D < B1×(C + 1);

[0016] The null value supplement module is configured to fill the blank bits in the B2-bit second data with null values;

[0017] The decoding memory is configured to cache the B2-bit data and then output the second data.

[0018] In the above technical solution, when the decoding module performs data conversion from B1 bit to B2 bit, since the total bit width of the data bus is D bit, the total bit width of the data of C pixels cached by the decoding cache module each time cannot exceed D bit, and the total bit width of the data of D / B2 pixels processed by the null value filling module is also D bit. The second data output by the decoding memory meets the total bit width of the data bus, that is, the second data can be directly input into the neural network accelerator hardware for processing.

[0019] In some optional implementations, the decoding cache module includes: a 14-bit cache module, a 12-bit cache module, a 10-bit cache module and an 8-bit cache module, and the null value supplement module includes: a 14-bit null value supplement module, a 12-bit null value supplement module, a 10-bit null value supplement module and an 8-bit null value supplement module;

[0020] When performing 14-bit to 16-bit data conversion, the 14-bit cache module is configured to cache the input data of 9 pixels, and output the data of 8 pixels to the 14-bit null value supplement module in each clock unit; the 14-bit null value supplement module is configured to fill the high-bit or low-bit data of each pixel with a null value, for example, fill the blank high-bit or low-bit data with 0 to convert it into 16-bit data;

[0021] When performing 12-bit to 16-bit data conversion, the 12-bit cache module is configured to cache the input data of 9 pixels, and output the data of 8 pixels to the 12-bit null value supplement module in each clock unit; the 12-bit null value supplement module is configured to fill the high-bit or low-bit data of each pixel with a null value, for example, fill the blank high-bit or low-bit data with 0 to convert it into 16-bit data;

[0022] When performing 10-bit to 16-bit data conversion, the 10-bit cache module is configured to cache the input data of 9 pixels, and output the data of 8 pixels to the 10-bit null value supplement module in each clock unit; the 10-bit null value supplement module is configured to fill the high-bit or low-bit data of each pixel with a null value, for example, fill the blank high-bit or low-bit data with 0 to convert it into 16-bit data;

[0023] When performing 8-bit to 16-bit data conversion, the 8-bit cache module is configured to cache the input data of 9 pixels, and, in each clock unit, output the data of 8 pixels to the 8-bit null value supplement module; the 8-bit null value supplement module is configured to fill the high-order or low-order data of each pixel with null values, for example, filling the blank high-order or low-order data with 0 to convert it into 16-bit data.

[0024] In the above technical solution, image data is usually stored in memory in a tightly packed 8-bit format, that is, each pixel is in the 8-bit format. However, considering that the input image may exist in the 10-bit, 12-bit or 14-bit format (there are two reasons: 1. Different image sensors may produce different numbers of pixel bits, and 2. In actual use, the number of bits of image pixels may be compressed according to demand. For example, in dark light conditions, 14-bit storage is used to store more useful information, while in normal light conditions, 8-bit storage is used to save storage space), and when the neural network model is processing, the hardware is designed based on the 16-bit data format, so it is necessary to use a decoding module to restore the data to a unified 16-bit format.

[0025] In some optional embodiments, the channel filling module includes a filling cache module, a filling module and a filling memory;

[0026] The filling module is configured to fill the input A1 / A2 data into the data form of A2 channels in each clock unit when the second data is the image data of the first format of A1 channels; the filling memory is configured to output the third data after caching the data of A2 channels; the filling cache module is configured to cache the remaining (1-A1 / A2) data, which is used to be pieced together with the second data of the next clock unit.

[0027] In the above technical solution, the image data of the first format of A1 channels is, for example, image data in the format of rgb, and the data of A2 channels is, for example, data in the format of rgbIr. In order to support these two data forms, the neural network uniformly uses data in the format of rgbIr for training. If the input data is in the format of rgbIr, the channel filling module directly skips the processing. If the input data is in the format of rgb, a data channel will be added through the filling module (the channel data of the data channel is set to 0) to simulate the third data in the format of rgbIr, that is, the Ir channel is all 0. Since adding a data channel increases 1 / 4 of the data, 3 / 4 of the input data is filled in each clock unit, and the remaining 1 / 4 of the data is pieced together with the second data of the next clock unit.

[0028] In some optional embodiments, the extraction module includes an extraction data module, a channel integration module and an extraction memory;

[0029] The data extraction module is configured to extract the data of the target channel of the third data; the channel integration module is configured to integrate the data of the target channel extracted every A2 clock units; the extraction memory caches the integrated data of the target channel and then outputs the fourth data; wherein the fourth data is used to input the first neural network of the accelerator hardware for processing to obtain the processed data of the target channel.

[0030] In the above technical solution, in the target channel, for example, in the RGBIr format image data under extremely dark conditions, the data of the g channel usually contains a large amount of data, and the data of the g channel needs to be processed separately. Therefore, in each clock unit, the data extraction module extracts the data of the g channel as 1 / 4 of 128-bit data, and it takes 4 clock units to extract the data of the g channel to integrate it into 128-bit data. The extraction memory caches the integrated 128-bit data and outputs it as the fourth data.

[0031] In some optional embodiments, the method further comprises:

[0032] The fusion module is used to fuse the processed data of the target channel with the data of the other three channels to obtain a preprocessed whole image, and the preprocessed whole image is used to input the second neural network of the accelerator hardware.

[0033] An embodiment of the present application provides a data processing method for image data, comprising:

[0034] Convert first data about the image to be input in the memory into second data according to the data format required by the accelerator hardware;

[0035] When the second data is image data of the first format with A1 channels, the second data is padded with third data of the second format with A2 channels; wherein A1 and A2 are both positive integers greater than 0, and A1 <A2;

[0036] The fourth data of the target channel among the A2 channels of the third data is extracted.

[0037] In the above technical solution, image data is usually in the form of a tightly arranged 8-bit in memory, that is, each pixel is in the form of 8 bits. However, considering that the input image may be in the form of 10-bit, 12-bit or 14-bit, etc., and when the neural network model is processing, the hardware is designed based on a fixed data form. The decoding module can uniformly process the data into the data form required by the accelerator hardware. Since the data output by some image sensors has different numbers of channels, and the neural network model uniformly adopts the format of A2 channels, the channel filling module can process the image data in the first format with A1 channels into the third data in the second format with A2 channels. For some image data, when the amount of data contained in a certain channel is large, usually the data of the target channel is first subjected to demosaic (removing mosaic) and noise reduction in the raw domain, etc. The extraction module can extract the data of the target channel so that the neural network accelerator hardware can process individual channels separately.

[0038] In some alternative embodiments, converting the first data regarding the image to be input in the memory into the second data according to the data form required by the accelerator hardware includes:

[0039] converting the received first data of B1 bits into the second data of B2 bits, where both B1 and B2 are positive integers greater than 0, and B1 < B2; and

[0040] filling the blank bits in the second data of B2 bits with null values.

[0041] In the above technical solution, when the decoding module converts the first data into the second data, it adopts the method of filling the first data with null values to obtain the second data.

[0042] In some alternative embodiments, when the second data is the image data in the first format with A1 channels, filling the second data into the third data in the second format with A2 channels includes:

[0043] When the second data is the image data in the first format with A1 channels, the data of A1 / A2 is processed into the data form of A2 channels and output per clock unit, and the remaining (1 - A1 / A2) data is pieced together with the second data of the next clock unit.

[0044] In the above technical solution, the image data of the first format of A1 channels is, for example, image data in the format of rgb, and the data of A2 channels is, for example, data in the format of rgbIr. In order to support these two data forms, the neural network uniformly uses data in the format of rgbIr for training. If the input data is in the format of rgbIr, the channel filling module directly skips the processing. If the input data is in the format of rgb, a data channel will be added through the filling module (the channel data of the data channel is set to 0) to simulate the third data in the format of rgbIr, that is, the Ir channel is all 0. Since adding a data channel increases 1 / 4 of the data, 3 / 4 of the input data is filled in each clock unit, and the remaining 1 / 4 of the data is pieced together with the second data of the next clock unit.

[0045] In some optional embodiments, the method comprises:

[0046] The fourth data is input into the first neural network of the accelerator hardware for processing to obtain processed data of the target channel.

[0047] In the above technical solution, the fourth data, for example, the data of the g channel in the RGBIr format image data under extremely dark conditions, usually contains a large amount of data. Therefore, the data of the g channel needs to be input into the first neural network of the accelerator hardware for separate processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0049] Figure 1 A functional module diagram of a data processing device for image data provided in an embodiment of the present application;

[0050] Figure 2 It is a schematic diagram of the functional module structure of the decoding module;

[0051] Figure 3 A functional module diagram of a decoding module according to another embodiment of the present application;

[0052] Figure 4 A functional module diagram of a channel filling module provided in an embodiment of the present application;

[0053] Figure 5 A functional module diagram of an extraction module provided in an embodiment of the present application; and

[0054] Figure 6 The flowchart of steps of a data processing method for image data provided by an embodiment of the present application.

[0055] Icons: 1 - Decoding module, 11 - Decoding cache module, 12 - Null value filling module, 13 - Decoding memory, 111 - 14-bit cache module, 112 - 12-bit cache module, 113 - 10-bit cache module, 114 - 8-bit cache module, 121 - 14-bit null value filling module, 122 - 12-bit null value filling module, 123 - 10-bit null value filling module, 124 - 8-bit null value filling module, 2 - Channel filling module, 3 - Extraction module, 21 - Filling cache module, 22 - Filling module, 23 - Filling memory, 31 - Extracted data module, 32 - Channel integration module, 33 - Extraction memory. Specific embodiments

[0056] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application.

[0057] Please refer to Figure 1 , Figure 1 which is a functional module diagram of a data processing device for image data provided by an embodiment of the present application, specifically including a decoding module 1, a channel filling module 2, and an extraction module 3.

[0058] Among them, the decoding module 1 is used to obtain the first data of the to-be-input image in the memory and convert the first data into the second data according to the data form required by the accelerator hardware. The channel filling module 2 is used to fill the second data into the third data of the second format with A2 channels when the second data is the image data of the first format with A1 channels; where A1 and A2 are both positive integers greater than 0, and A1 < A2. The extraction module 3 is used to extract the data of the target channels from the A2 channels of the third data and form the fourth data.

[0059] In the embodiments of the present application, image data is usually in a tightly arranged form of 8 bits in memory, that is, each pixel is in the form of 8 bits. However, considering that the input image may be in the form of 10 bits, 12 bits, 14 bits, etc., and when the neural network model is processing, the hardware is designed based on a fixed data form. The decoding module 1 can be used to uniformly process the data into the data form required by the accelerator hardware. Since the data output by some image sensors has different numbers of channels, and the neural network model uniformly adopts a format of A2 channels, the channel filling module 2 can process the image data in the first format with A1 channels into the third data in the second format with A2 channels. For some image data, when the data volume contained in a certain channel is large, usually the data of the target channel is first subjected to demosaic (de-mosaicing) and noise reduction in the raw domain, etc. The extraction module 3 can extract the data of the target channel so that the neural network accelerator hardware can process individual channels separately.

[0060] In some alternative embodiments, during the process of converting the first data into the second data, the decoding module 1 is specifically configured to: convert the received B1-bit first data into B2-bit second data, where both B1 and B2 are positive integers greater than 0, and B1 < B2; and fill the blank bits in the B2-bit second data with null values.

[0061] In the embodiments of the present application, when the decoding module 1 converts the first data into the second data, it adopts a method of filling null values in the first data to obtain the second data.

[0062] Please refer to Figure 2 , Figure 2 FIG. is a schematic structural diagram of the functional modules of the decoding module 1. The decoding module 1 includes a decoding cache module 11, a null value supplement module 12, and a decoding memory 13.

[0063] Among them, when performing the data conversion from B1 bits to B2 bits, the decoding cache module 11 is configured to cache the data of C pixel points input, and output D / B2 pixel points of data to the null value supplement module 12 per clock unit; where C, D, and D / B2 are all positive integers greater than 0, and further satisfy the relationship of B1×C < D < B1×(C + 1), and D is the total bit width of the data bus. The null value supplement module 12 is configured to fill the blank bits in the B2-bit second data with null values. The decoding memory 13 is configured to cache the B2-bit data and then output the second data.

[0064] In the embodiment of the present application, when the decoding module 1 performs data conversion from B1 bit to B2 bit, since the total bit width of the data bus is D bit, the total bit width of the data of C pixel points cached by the decoding cache module 11 each time cannot exceed D bit, and the total bit width of the data of D / B2 pixel points processed by the null value filling module 22 is also D bit. The second data output by the decoding memory 13 meets the total bit width of the data bus, that is, the second data can be directly input into the neural network accelerator hardware for processing.

[0065] Please refer to Figure 3 , Figure 3 This is a functional module diagram of the decoding module 1 of another embodiment of the present application. In this embodiment, the decoding cache module 11 includes: a 14-bit cache module 111, a 12-bit cache module 112, a 10-bit cache module 113 and an 8-bit cache module 114, and the null value supplement module 12 includes: a 14-bit null value supplement module 121, a 12-bit null value supplement module 122, a 10-bit null value supplement module 123 and an 8-bit null value supplement module 124.

[0066] Among them, when performing 14-bit to 16-bit data conversion, the 14-bit cache module 111 is configured to cache the input data of 9 pixel points, and, in each clock unit, output the data of 8 pixel points to the 14-bit null value supplement module 121; the 14-bit null value supplement module 121 is configured to fill the high-order or low-order data of each pixel point with null values, for example, filling the blank high-order or low-order data with 0 to convert it into 16-bit data.

[0067] When performing 12-bit to 16-bit data conversion, the 12-bit cache module 112 is configured to cache the input data of 9 pixels, and, in each clock unit, output the data of 8 pixels to the 12-bit null value supplement module 122; the 12-bit null value supplement module 122 is configured to fill the high-order or low-order data of each pixel with null values, for example, filling the blank high-order or low-order data with 0 to convert it into 16-bit data.

[0068] When performing 10-bit to 16-bit data conversion, the 10-bit cache module 113 is configured to cache the input data of 9 pixels, and, in each clock unit, output the data of 8 pixels to the 10-bit null value supplement module 123; the 10-bit null value supplement module 123 is configured to fill the high-order or low-order data of each pixel with null values, for example, filling the blank high-order or low-order data with 0 to convert it into 16-bit data.

[0069] When performing 8-bit to 16-bit data conversion, the 8-bit cache module 114 is configured to cache the input data of 9 pixels, and, in each clock unit, output the data of 8 pixels to the 8-bit null value supplement module 124; the 8-bit null value supplement module 124 is configured to fill the high-order or low-order data of each pixel with null values, for example, filling the blank high-order or low-order data with 0 to convert it into 16-bit data.

[0070] In the embodiment of the present application, the image data is usually stored in the memory in a tightly packed 8-bit format, that is, each pixel is in the 8-bit format. However, considering that the input image may exist in the 10-bit, 12-bit or 14-bit format (there are two reasons: 1. Different image sensors may generate different numbers of pixel bits, and 2. In actual use, the number of bits of image pixels may be compressed as needed. For example, in dark light conditions, 14-bit storage is used to store more useful information, while in normal light conditions, 8-bit storage is used to save storage space), and when the neural network model is processing, the hardware is designed based on the 16-bit data format, so it is necessary to restore the data to the 16-bit format through the decoding module 1.

[0071] Please refer to Figure 4 , Figure 4 is a functional module diagram of the channel filling module 2 , which specifically includes a filling cache module 21 , a filling module 22 and a filling memory 23 .

[0072] Among them, the filling module 22 is configured to fill the input A1 / A2 data into the data form of A2 channels in each clock unit when the second data is the image data of the first format of A1 channels; the filling memory 23 is configured to output the third data after caching the data of A2 channels; the filling cache module 21 is configured to cache the remaining (1-A1 / A2) data, which is used to be pieced together with the second data of the next clock unit.

[0073] In the embodiment of the present application, the image data of the first format of A1 channels is, for example, image data in the format of rgb, and the data of A2 channels is, for example, data in the format of rgbIr. In order to support these two data forms, the neural network uniformly uses data in the format of rgbIr for training. If the input data is in the format of rgbIr, the channel filling module 2 directly skips the processing. If the input data is in the format of rgb, a data channel will be added through the filling module 22 (the channel data of the data channel is set to 0) to simulate the third data in the format of rgbIr, that is, the Ir channel is all 0. Since the addition of the data channel increases 1 / 4 of the data, 3 / 4 of the input data is filled in each clock unit, and the remaining 1 / 4 of the data is pieced together with the second data of the next clock unit.

[0074] Please refer to Figure 5 , Figure 5 3 is a functional module diagram of the extraction module 3 . The extraction module 3 specifically includes an extraction data module 31 , a channel integration module 32 and an extraction memory 33 .

[0075] Among them, the data extraction module 31 is configured to extract the data of the target channel of the third data; the channel integration module 32 is configured to integrate the data of the target channel extracted every A2 clock units; the extraction memory 33 caches the integrated data of the target channel and then outputs the fourth data; wherein the fourth data is used to input the first neural network of the accelerator hardware for processing to obtain the processed data of the target channel.

[0076] In an embodiment of the present application, in the target channel, for example, in the RGBIr format image data under extremely dark conditions, the data of the g channel usually contains a large amount of data, and the data of the g channel needs to be processed separately. Therefore, in each clock unit, the data extraction module 31 extracts the data of the g channel as 1 / 4 of 128-bit data, and it takes 4 clock units to extract the data of the g channel to integrate it into 128-bit data. The extraction memory 33 caches the integrated 128-bit data and outputs it as the fourth data.

[0077] In some optional embodiments, it also includes: a fusion module, which is used to fuse the processed data of the target channel with the data of the other three channels to obtain a preprocessed whole image, and the preprocessed whole image is used to input the second neural network of the accelerator hardware.

[0078] Please refer to Figure 6 , Figure 6 A flowchart of a method for processing image data provided in an embodiment of the present application specifically includes:

[0079] Step 101: Convert the first data regarding the to-be-input image in the memory into second data according to the data format required by the accelerator hardware;

[0080] Step 102: When the second data is the image data of the first format with A1 channels, fill the second data into the third data of the second format with A2 channels; where both A1 and A2 are positive integers greater than 0, and A1 < A2;

[0081] Step 103: Extract the fourth data of the target channel from the A2 channels of the third data.

[0082] In the embodiments of the present application, the image data in the memory usually adopts a form of 8-bit tightly arranged, that is, each pixel point is in the form of 8 bits. However, considering that the input image may be in the form of 10-bit, 12-bit or 14-bit, etc., and when the neural network model is processing, the hardware is designed based on a fixed data form. The decoding module 1 can uniformly process the data into the data form required by the accelerator hardware. Since the data output by some image sensors has different numbers of channels, and the neural network model uniformly adopts the format of A2 channels, the channel filling module 2 can process the image data of the first format with A1 channels into the third data of the second format with A2 channels. For some image data, when the amount of data contained in a certain channel is large, usually the data of the target channel is first subjected to demosaic (de-mosaicing) and raw domain noise reduction, etc. The extraction module 3 can extract the data of the target channel so that the neural network accelerator hardware can perform separate processing on individual channels.

[0083] In some optional embodiments, converting the first data regarding the to-be-input image in the memory into second data according to the data format required by the accelerator hardware includes: converting the received first data of B1 bits into second data of B2 bits, where both B1 and B2 are positive integers greater than 0, and B1 < B2; and filling the blank bits in the second data of B2 bits with null values.

[0084] In the embodiments of the present application, when the decoding module 1 converts the first data into the second data, it adopts the method of filling the first data with null values to obtain the second data.

[0085] In some optional embodiments, when the second data is image data of the first format of A1 channels, the second data is padded with third data of the second format of A2 channels, including: when the second data is image data of the first format of A1 channels, the data of A1 / A2 is processed into the data form of A2 channels and output in each clock unit, and the remaining (1-A1 / A2) data is pieced together with the second data of the next clock unit.

[0086] In the embodiment of the present application, the image data of the first format of A1 channels is, for example, image data in the format of rgb, and the data of A2 channels is, for example, data in the format of rgbIr. In order to support these two data forms, the neural network uniformly uses data in the format of rgbIr for training. If the input data is in the format of rgbIr, the channel filling module 2 directly skips the processing. If the input data is in the format of rgb, a data channel will be added through the filling module 22 (the channel data of the data channel is set to 0) to simulate the third data in the format of rgbIr, that is, the Ir channel is all 0. Since the addition of the data channel increases 1 / 4 of the data, 3 / 4 of the input data is filled in each clock unit, and the remaining 1 / 4 of the data is pieced together with the second data of the next clock unit.

[0087] In some optional embodiments, the method further includes: inputting the fourth data into the first neural network of the accelerator hardware for processing to obtain processed data of the target channel.

[0088] In an embodiment of the present application, the fourth data is, for example, the data of the g channel in the RGBIr format image data under extremely dark conditions. Since the data of the g channel usually contains a large amount of data, the data of the g channel needs to be input into the first neural network of the accelerator hardware for separate processing.

[0089] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0090] In addition, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0091] Furthermore, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0092] In this document, relational terms such as first and second, etc. are used merely to distinguish one entity or operation from another entity or operation, but do not necessarily require or imply any such actual relationship or order between these entities or operations.

[0093] The above description is only an embodiment of the present application and is not intended to limit the protection scope of the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A data processing device for image data, characterized in that, it comprises: a decoding module, configured to obtain first data about an image to be input in a memory, and convert the first data into second data according to the data format required by accelerator hardware; a channel filling module, configured to fill the second data into third data of a second format with A2 channels when the second data is image data of a first format with A1 channels; wherein, both A1 and A2 are positive integers greater than 0, and A1 < A2; and an extraction module, configured to extract data of target channels from the A2 channels of the third data and form fourth data; the channel filling module includes a filling cache module, a filling module and a filling memory; the filling module is configured to, when the second data is image data of a first format with A1 channels, fill the input data of A1 / A2 into a data form with A2 channels at each clock unit; the filling memory is configured to cache the data of A2 channels and then output the third data; the filling cache module is configured to cache the remaining (1 - A1 / A2) data, and this data is used to piece together with the second data of the next clock unit.

2. The device according to claim 1, characterized in that, during the process of converting the first data into the second data, the decoding module specifically is configured to: convert the received first data of B1 bits into second data of B2 bits, wherein both B1 and B2 are positive integers greater than 0, and B1 < B2; and fill the blank bits in the second data of B2 bits with null values.

3. The device according to claim 2, characterized in that, the decoding module includes a decoding cache module, a null value supplement module and a decoding memory; when performing the data conversion from B1 bits to B2 bits, the decoding cache module is configured to cache data of C pixel points input, and output D / B2 pixel points of data to the null value supplement module at each clock unit; wherein, C, D and D / B2 are all positive integers greater than 0, D is the total bit width of the data bus, and it satisfies the relationship of B1×C < D < B1×(C + 1); the null value supplement module is configured to fill the blank bits in the second data of B2 bits with null values; and the decoding memory is configured to cache the data of B2 bits and then output the second data.

4. The device according to claim 3, characterized in that, the decoding cache module includes: a 14-bit cache module, a 12-bit cache module, a 10-bit cache module and an 8-bit cache module, and the null value supplement module includes: a 14-bit null value supplement module, a 12-bit null value supplement module, a 10-bit null value supplement module and an 8-bit null value supplement module; When performing 14-bit to 16-bit data conversion, the 14-bit cache module is configured to cache the input data of 9 pixels, and output the data of 8 pixels to the 14-bit null value supplement module in each clock unit; the 14-bit null value supplement module is configured to fill the high-bit or low-bit data of each pixel with null values ​​to convert it into 16-bit data; When performing 12-bit to 16-bit data conversion, the 12-bit cache module is configured to cache the input data of 9 pixels, and output the data of 8 pixels to the 12-bit null value supplement module in each clock unit; the 12-bit null value supplement module is configured to fill the high-bit or low-bit data of each pixel with null values ​​to convert it into 16-bit data; When performing 10-bit to 16-bit data conversion, the 10-bit buffer module is configured to buffer the input data of 9 pixels, and output the data of 8 pixels to the 10-bit null value supplement module in each clock unit; the 10-bit null value supplement module is configured to fill the high-bit or low-bit data of each pixel with null values ​​to convert it into 16-bit data; When performing 8-bit to 16-bit data conversion, the 8-bit cache module is configured to cache the input data of 9 pixels, and, in each clock unit, output the data of 8 pixels to the 8-bit null value supplement module; the 8-bit null value supplement module is configured to fill the high-bit or low-bit data of each pixel with null values ​​to convert it into 16-bit data.

5. The device according to claim 1, It is characterized in that The extraction module includes an extraction data module, a channel integration module and an extraction memory; The data extraction module is configured to extract data of a target channel of the third data; The channel integration module is configured to integrate the data of the target channel extracted every A2 clock units; the extraction memory caches the integrated data of the target channel and outputs the fourth data; wherein the fourth data is used to input the first neural network of the accelerator hardware for processing to obtain the processed data of the target channel.

6. The device as claimed in claim 5, It is characterized in that Also includes: A fusion module is used to fuse the processed data of the target channel with the data of the other three channels to obtain a preprocessed whole image, and the preprocessed whole image is used to input the second neural network of the accelerator hardware.

7. A data processing method for image data, It is characterized in that include: Convert first data about the image to be input in the memory into second data according to the data format required by the accelerator hardware; When the second data is image data of the first format with A1 channels, the second data is padded with third data of the second format with A2 channels; wherein A1 and A2 are both positive integers greater than 0, and A1 <A2; Extracting fourth data of a target channel from A2 channels of the third data; When the second data is image data in a first format with A1 channels, padding the second data to third data in a second format with A2 channels includes: When the second data is image data in a first format with A1 channels, processing A1 / A2 of the data into a data form with A2 channels and outputting it per clock unit, and piecing together the remaining (1 - A1 / A2) of the data with the second data of the next clock unit.

8. The method according to claim 7, wherein, converting the first data regarding the image to be input in the memory into second data according to the data form required by the accelerator hardware includes: converting the received first data of B1 bits into second data of B2 bits, where both B1 and B2 are positive integers greater than 0, and B1 < B2; and padding the blank bits in the second data of B2 bits with null values.

9. The method according to claim 7, wherein, it includes: inputting the fourth data into a first neural network of the accelerator hardware for processing to obtain the data of the target channel after processing.

Citation Information

Patent Citations

  • Channel adjustment method, device and equipment for convolutional neural network model

    CN112766276A

  • Image data processing system and method

    CN113038269A