Row neighborhood image memory bank, computing platform and image data pipeline processing method
By stacking memory chips in the image memory bank and using bit-segment fission technology to form row neighborhood image memory bank, the insufficient memory bank bandwidth and memory wall problems are solved, efficient and low-power image data processing is achieved, and multi-algorithm parallel applications are supported, reducing system costs.
Patent Information
- Application Number
- CN202411906201.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-13
AI Technical Summary
The existing technology has insufficient memory bank bandwidth, memory wall problems, and high power consumption and high cost problems in image processing, which limits the development of image processing and artificial intelligence technology.
Using a row neighborhood image memory bank, multiple data bits are formed by stacking N memory chips of the same model horizontally and making data bits independent, combining bit bit segment fission technology, each data bit segment corresponds to one image pixel, achieving efficient data processing.
It significantly improves the bandwidth of the memory bank, solves the memory wall problem, reduces system power consumption, and supports parallel applications of multiple algorithms, reduces dependence on dedicated hardware, and reduces the overall cost of the system.
Smart Images

Figure CN119991409A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of high-speed image processing, and in particular to a row neighborhood image storage body, a computing platform and an image data pipeline processing method. Background Art
[0002] With the advent of the era of artificial intelligence and big data, the demand for computing power has increased dramatically, especially in the fields of deep learning and image processing. However, the improvement of computing power is not only limited by the capabilities of computing devices, but also by storage bandwidth and algorithm optimization. In practical applications, although hardware such as GPUs provide powerful computing power, insufficient storage bandwidth and memory wall problems have become performance bottlenecks. The memory wall problem is mainly manifested in the bottleneck of the data channel between the processor and the storage, which makes data transmission more time-consuming than the calculation itself. Especially in a multi-processor environment, it is difficult for the storage to supply data in a timely manner, resulting in a decrease in processing speed.
[0003] Existing storage technologies, such as general-purpose storage, can only process one piece of data per access operation, resulting in low storage bandwidth, which cannot meet the needs of the big data era. To solve these problems, the industry generally adopts high-bandwidth storage chips (such as GDDR6, HBM) and storage chip stacking technology. Although these technologies have improved storage bandwidth to a certain extent, they do not solve the problem of high effective storage bandwidth, and the number of storage chip stacks is limited by factors such as cost, space and power consumption.
[0004] In the field of image processing, especially face recognition algorithms and YOLO algorithms in deep learning, a large number of convolution calculations are required, which usually rely on GPU boards. Although GPUs have high computing power, they have disadvantages such as high power consumption, high noise, and high construction and operation costs. Therefore, it is urgent to develop high-performance lightweight pipeline convolution technology.
[0005] In summary, the problems existing in the existing technologies mainly include insufficient storage bandwidth, memory wall problem, and high power consumption and high cost in image processing. These problems limit the development of image processing and artificial intelligence technology, especially in scenarios that require large-scale data processing and high-speed computing. Summary of the invention
[0006] The present invention provides a row neighborhood image storage body, a computing platform and an image data pipeline processing method, which are used to solve the problems of insufficient storage body bandwidth, memory wall problem and high power consumption and high cost in image processing in the prior art, and realize lightweight and efficient data processing.
[0007] The present invention provides a row neighborhood image storage body, comprising: N memory chips of the same model are stacked horizontally, and the address bits and timing bits of the N memory chips are connected together one by one, while the data bits are independent of each other; Performing bit segmentation fission on the data bits of each storage chip to form a plurality of data bit segments, each data bit segment corresponding to an image pixel, so as to form the row neighborhood image storage body; Among them, a single row neighborhood image storage body can store at least one W×H 8-bit grayscale image, wherein W is the number of 8-bit grayscale image pixels in a row in the horizontal direction, and H is the number of 8-bit grayscale image pixels in a column in the vertical direction; W≥512, H≥512.
[0008] According to a row neighborhood image storage body provided by the present invention, the data bits of each storage chip are 8×k bits, wherein k≥2.
[0009] According to a row neighborhood image storage body provided by the present invention, the data bits of each storage chip are segmented and split according to 8 bits / pixel.
[0010] According to a row neighborhood image storage provided by the present invention, N≥2 and N≤8.
[0011] The present invention provides a computing platform, comprising two row neighborhood image storage bodies and a processing chip as described above, wherein the two row neighborhood image storage bodies are respectively connected to the processing chip.
[0012] The present invention provides an image data pipeline processing method, which is implemented based on a computing platform, and the method comprises: Determine the specific location of the image data to be read in the first row of the neighborhood image storage body by the processing chip; Sending a read instruction to the first row neighborhood image storage body through a processing chip to locate a first source data bit segment of a storage chip of the first row neighborhood image storage body, and performing a parallel read operation on the first source data bit segment to obtain first source image data; Processing the first source image data by a processing chip to obtain first target image data; The first destination data bit segment of the storage chip of the second row neighborhood image storage body that needs to be stored is determined by the processing chip, and the first target image data is sent to the first destination data bit segment for storage.
[0013] According to the image data pipeline processing method provided by the present invention, after the first target image data is sent to the first destination data bit segment for storage by the processing chip, the method further includes: Determine the specific location of the image data to be read in the second row neighborhood image storage body by the processing chip; Sending a read instruction to the second row neighborhood image storage body through the processing chip to locate the second source data bit segment of the storage chip of the second row neighborhood image storage body, and performing a parallel read operation on the second source data bit segment to obtain second source image data; Processing the second source image data by a processing chip to obtain second target image data; The second destination data bit segment of the storage chip of the first row neighborhood image storage body that needs to be stored is determined by the processing chip, and the second target image data is sent to the second destination data bit segment for storage.
[0014] According to the image data pipeline processing method provided by the present invention, the image data to be read includes W×H image pixels; The processing chip performs a parallel read operation on the first source data bit segment to obtain the first source image data, including: By adopting the memory bank continuous reading mode, W image pixels of each row of the first source data bit segment are read out successively in row order through the processing chip until the image pixels of H rows are read out, thereby obtaining the first source image data.
[0015] According to the image data pipeline processing method provided by the present invention, the first source image data is processed by a processing chip to obtain first target image data, which specifically includes: the first source image data is processed by a processing chip based on a preset algorithm to obtain the first target image data.
[0016] According to the image data pipeline processing method provided by the present invention, there are multiple preset algorithms; The first source image data is processed by a processing chip based on a preset algorithm to obtain the first target image data, specifically including: processing the first source image data in the order of the 1st row to the Hth row, using a processing pipeline to process each row of source image data in the order of multiple algorithms to obtain row target image data corresponding to each algorithm, until all rows of source image data are processed to obtain first target image data corresponding to each algorithm.
[0017] The row neighborhood image storage provided by the present invention stacks N memory chips of the same model horizontally and makes the data bits independent. This structure significantly improves the bandwidth of the storage body and effectively solves the problems of insufficient storage bandwidth and memory wall. This design allows parallel access to multiple data bit segments, each data bit segment corresponding to an image pixel, thereby improving the data transmission rate and reducing the data transmission time.
[0018] In addition, the bit segmentation fission technology further optimizes the utilization of storage chips, allowing each storage chip to store more image pixel data and improve storage efficiency. This improvement not only alleviates the urgent demand for storage bandwidth in the big data era, but also provides strong support for high-speed image processing technology, especially in scenarios that require large-scale data processing and high-speed computing.
[0019] The image data pipeline processing method provided by the present invention realizes efficient processing of image data through the coordinated work of the processing chip and the row neighborhood image storage body. It realizes pipeline operation of data processing by reading out and processing the data in the first row neighborhood image storage body in parallel, and then storing the processing result in the second row neighborhood image storage body. This method not only improves the speed of data processing, but also reduces the power consumption of the system by optimizing the data flow and reducing unnecessary data movement.
[0020] Secondly, this method supports the parallel application of multiple algorithms, allowing a single hardware platform to flexibly handle multiple image processing tasks, reducing the reliance on dedicated hardware, thereby reducing the overall cost of the system. This efficient and flexible image data processing method provides new possibilities for lightweight and efficient data processing, especially in the fields of deep learning and artificial intelligence, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0022] Figure 1 It is a schematic diagram of the structure of the row neighborhood image storage body provided by the present invention.
[0023] Figure 2 This is one of the schematic diagrams of the row neighborhood image storage provided by the present invention.
[0024] Figure 3 This is the second schematic diagram of the row neighborhood image storage provided by the present invention.
[0025] Figure 4 It is a structural schematic diagram of the computing platform provided by the present invention.
[0026] Figure 5 It is one of the schematic diagrams of the image data pipeline processing method provided by the present invention.
[0027] Figure 6It is a schematic diagram of realizing pipeline processing of image data provided by the present invention.
[0028] Figure 7 This is the second schematic diagram of the image data pipeline processing method provided by the present invention.
[0029] Figure 8 It is a schematic diagram of the bidirectional pipeline processing provided by the present invention.
[0030] Fig. 9 It is a schematic diagram of implementing image data pipeline processing using multiple algorithms provided by the present invention. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0032] The embodiment of the present invention discloses a row neighborhood image storage body, see Figure 1 ,include: N memory chips of the same type are stacked in a horizontal direction, and the address bits and timing bits of the N memory chips are connected together one by one, while the data bits are independent of each other; wherein N≥2 and N≤8.
[0033] Among them, the memory chip can be a DDR4 memory chip.
[0034] The data bits of each storage chip are segmented and fissioned to form multiple data bit segments, each data bit segment corresponds to an image pixel, so as to form a row neighborhood image storage body.
[0035] For example, each 32-bit data bit is split into four 8-bit data bit segments, corresponding to four image pixels. One parallel read and write operation can read and write 4×4=16 image pixels at the same time.
[0036] The data bits of each memory chip are 8×k bits, where k≥2.
[0037] Specifically, the data bits of each memory chip are greater than or equal to 16-bit, 32-bit, 64-bit, 128-bit or higher bit DDR memory chips.
[0038] The data bits of each memory chip are segmented and split into 8 bits per pixel.
[0039] For example, if the data bit length of the memory chip is 16 bits, the data bit segment is split into two segments, that is, bit0-bit7 is the first data bit segment, and bit8-bit15 is the second data bit segment; If the data bit length of the memory chip is 32 bits, the data bit segment is split into 4 segments, bit0-bit7 is the first data bit segment, bit8-bit15 is the second data bit segment, bit16-bit23 is the third data bit segment, and bit24-bit31 is the fourth data bit segment.
[0040] A similar method is used to perform segmented fission on the data of the storage chip with a word length of 64 bits and 128 bits.
[0041] The row neighborhood image storage body includes N memory chips, and the data bits of the memory chips are 32 bits, corresponding to 4 image pixels, so the row neighborhood storage body can read out or write 4×N image pixels in parallel at one time.
[0042] See also Figure 2 and Figure 3 , schematic diagrams of two types of row neighborhood image storage are shown respectively.
[0043] Figure 2 In the figure, “·” represents an image pixel, and its data bit is 8 bits. The data bit word length of a single DDR memory chip used is 32 bits, and the 32 bits are segmented into four 8-bit segments. In a read or write operation of the row neighbor memory bank, four 8-bit image pixels can be read or written at the same time.
[0044] Figure 3 In the figure, “·” represents an image pixel, and its data bit is 8 bits. The data bit length of the two DDR memory chips used is 32 bits, and the 32 bits are segmented into four 8-bit segments. In a read or write operation of the row neighborhood memory bank, eight 8-bit image pixels can be read or written at the same time.
[0045] Specifically, storage bandwidth usually refers to the rate of data transmission, that is, the amount of data that can be transmitted per unit time. It can be measured in bits per second (bps). The higher the bandwidth, the faster the data transmission speed.
[0046] The mathematical expression of the memory bank bandwidth of the row neighborhood image memory bank is: C = N × n1 × f0 / 8bit (1) Wherein, N is the number of memory chips of the row neighborhood image memory bank stacked in the horizontal direction; n1 is the number of data bits of a single memory chip; and f0 is the read and write clock frequency of the row neighborhood image memory bank.
[0047] Formula (1) clearly states the factors related to improving the memory bandwidth of the horizontally stacked row neighborhood image memory, namely, the number of stacked memory chips, the number of data bits of a single memory chip, and the memory read and write clock frequency.
[0048] The row neighborhood image storage body provided by the embodiment of the present invention stacks N memory chips of the same model in a horizontal direction, and connects the address bits and timing bits of these memory chips one by one, while the data bits are independent of each other; the stacked memory chips can work in parallel, thereby improving the reading and writing speed of data. Through the stacking technology and the bit segmentation fission technology, the effective bandwidth of the row neighborhood image storage body is improved, and the memory wall problem is solved.
[0049] Take a single row neighborhood image storage body storing a W×H 8-bit grayscale image as an example. W is the number of 8-bit grayscale image pixels in a row in the horizontal direction, and H is the number of 8-bit grayscale image pixels in a column in the vertical direction. W≥512, H≥512. Then for the row neighborhood image storage body, N×k data bit segments are stored, where N is the number of storage chips included in the row neighborhood image storage body, and k is the number of data bit segments formed by each storage chip.
[0050] The row neighborhood image storage provided by the embodiment of the present invention can stack N storage chips of the same type horizontally and make data bits independent. The present invention significantly improves the bandwidth of the storage and solves the memory wall problem.
[0051] The embodiment of the present invention also provides a computing platform, see Figure 4 , including two row neighborhood image storage bodies and a processing chip as described above, and the two row neighborhood image storage bodies are respectively connected to the processing chip.
[0052] In this embodiment, the processing chip may be a field programmable gate array (FPGA) processor. FPGA can perform multiple processing tasks simultaneously, which makes it very suitable for parallel processing of image data, especially in scenarios where multiple algorithms need to be applied simultaneously for image processing.
[0053] Of course, it can also be other parallel processing chips, such as application-specific integrated circuits (ASICs), graphics processing units (GPUs), or other AI chips.
[0054] The image data pipeline processing method provided by the present invention is described below. The image data pipeline processing method described below and the row neighborhood image storage body and computing platform described above can be referenced to each other.
[0055] The embodiment of the present invention also provides an image data pipeline processing method, which is implemented based on a computing platform, see Figure 5 , the method comprising: 501. Determine, by means of a processing chip, a specific location of image data to be read in a first row neighborhood image storage body.
[0056] Using image data management software or hardware logic, the processing chip determines the starting and ending addresses of specific image data in the first row of the neighborhood image storage bank. The specific storage location is calculated based on the image resolution, storage bank layout, and data storage strategy (such as row priority or column priority).
[0057] 502. Send a read instruction to the first row neighborhood image storage body through the processing chip to locate the first source data bit segment of the storage chip of the first row neighborhood image storage body, and implement a parallel read operation on the first source data bit segment to obtain the first source image data.
[0058] According to the position determined in step 501, an accurate read instruction including row address and column address information is generated by the processing chip. The read instruction is sent to the first row neighborhood image storage body through the control interface to activate the corresponding storage chip and prepare for data reading.
[0059] Utilize the parallel read capability of the memory bank to read data from multiple data segments at the same time, reducing the read time. Use a high-speed data bus or interface, such as a DDR interface, to support high-throughput data processing.
[0060] 503. Process the first source image data by using a processing chip to obtain first target image data.
[0061] The first source image data read out in parallel is sent to the FPGA processing chip through the first row neighborhood image storage body. The image data is processed according to the preset image processing algorithm (such as filtering, edge detection, color conversion, etc.). The calculations that may be involved in the processing process include mathematical operations, logical operations, and data conversion.
[0062] The processed first target image data is verified by the processing chip to ensure the integrity and correctness of the data. Error detection and correction mechanisms such as ECC (error correction code) are applied to improve the reliability of the data.
[0063] 504. Determine, by means of a processing chip, a first destination data bit segment of a storage chip of a second row neighborhood image storage body that needs to be stored, and send the first target image data to the first destination data bit segment for storage.
[0064] The processing chip determines the storage location based on the data processing results and the layout of the second row neighborhood image storage volume. Considering the data writing order and storage strategy, the storage path is optimized to reduce the writing time and potential conflicts.
[0065] The processed first target image data is written into a determined storage location through the processing chip. The parallel writing capability of the storage body is utilized to improve the data writing efficiency; the atomicity and consistency of the writing operation are ensured to prevent data damage.
[0066] The image data pipeline processing method provided by the embodiment of the present invention realizes efficient processing of image data through the cooperation of the processing chip and the row neighborhood image storage body. It realizes pipeline operation of data processing by reading out and processing the data in the first row neighborhood image storage body in parallel, and then storing the processing result in the second row neighborhood image storage body. This method not only improves the speed of data processing, but also reduces the power consumption of the system by optimizing the data flow and reducing unnecessary data movement.
[0067] Taking the processing chip as an FPGA processor as an example, through the above steps 501 to 504, the FPGA processes a grayscale image in the order of the 1st row to the Hth row using a processing pipeline method, such as Figure 6 shown.
[0068] Specifically, the image is divided into multiple blocks, each block is processed in parallel by N memory chips, and the width of the image is W and the height is H. Assuming that each memory chip can process a smaller block, the block size can be determined according to the capacity of the memory chip.
[0069] In step 501, the processing chip determines the specific position of the image data to be read in the first row of the neighborhood image storage body. For example, if the processing starts from the upper left corner of the image, the initial position is (0, 0).
[0070] In step 502, a read instruction is issued to the first row neighborhood image storage body through the processing chip to locate the first source data bit segment of the first row. Assuming that each storage chip stores 8 pixels (determined according to the data bit width of the storage chip), and the first row neighborhood image storage body includes 8 storage chips, then each read can read 8×8=64 image pixels.
[0071] W image pixels of each row of the first source data bit segment are read out successively in row order until all H rows of image pixels are read out, thereby obtaining the first source image data.
[0072] In step 503, the first source image data is processed by the processing chip to obtain the first target image data. The processing may include image enhancement, filtering or other image processing algorithms. For example, if the algorithm is edge detection, the FPGA will apply the edge detection algorithm to the read pixel data.
[0073] In step 504, the first destination data bit segment of the storage chip of the second row neighborhood image storage bank to be stored is determined by the processing chip, and the first target image data is sent to the first destination data bit segment for storage.
[0074] In this embodiment, in addition to implementing single-algorithm pipeline processing, bidirectional pipeline processing can also be implemented.
[0075] Specifically, see Figure 7 After the processing chip sends the first target image data to the first destination data segment for storage, the method further includes: 701. Determine, by means of a processing chip, a specific location of image data to be read in a second row neighborhood image storage body.
[0076] 702. Send a read instruction to the second row neighborhood image storage body through the processing chip to locate the second source data bit segment of the storage chip of the second row neighborhood image storage body, and implement a parallel read operation on the second source data bit segment to obtain second source image data.
[0077] 703. Process the second source image data by using a processing chip to obtain second target image data.
[0078] 704. Determine, by means of the processing chip, a second destination data bit segment of a storage chip of the first row neighborhood image storage body that needs to be stored, and send the second target image data to the second destination data bit segment for storage.
[0079] Figure 8 The figure shows a schematic diagram of bidirectional pipeline processing, wherein the first row of neighboring image storage bodies and the second row of neighboring image storage bodies implement bidirectional pipeline processing through a processing chip.
[0080] In this embodiment, bidirectional pipeline operation is used to allow image data to be read and stored alternately between two row neighborhood image storage bodies. This design allows data processing to be performed in parallel between the two storage bodies, thereby improving the overall efficiency of data processing.
[0081] Secondly, by performing round-trip data reading and writing between the two storage bodies, the system can read new data from one storage body while processing data in another storage body, so that data processing can be performed almost uninterruptedly, significantly improving the system throughput.
[0082] Thirdly, bidirectional pipelining makes better use of the resources of the memory bank and the processor. While processing one data set, another data set is being read or written, reducing the idle time of the processor and the memory bank.
[0083] Further, step 503 includes: processing the first source image data based on a preset algorithm by a processing chip to obtain the first target image data.
[0084] like Fig. 9 As shown, when there are multiple preset algorithms, the first source image data is processed in the order of the 1st row to the Hth row through the processing chip using a processing pipeline in the order of multiple algorithms to obtain row target image data corresponding to each algorithm, until all rows of source image data are processed to obtain the first target image data corresponding to each algorithm.
[0085] In this embodiment, a comprehensive computing method of row neighborhood image storage, computing platform and image data pipeline processing is applied to perform Gaussian filtering, Sobel operator, median filtering and binarization on a 512×512 grayscale image, which takes only 0.0196ms and achieves ultra-high-speed image processing.
[0086] Through the pipeline approach, image data can be processed continuously through multiple algorithms, reducing waiting time and improving the overall processing speed.
[0087] Secondly, multi-algorithm processing enables image data to be analyzed by multiple algorithms simultaneously, which helps to extract richer image features and enhances the depth and breadth of data processing.
[0088] Again, pipeline processing allows the processing chip to prepare data for the next algorithm while processing one algorithm, thereby maximizing processor utilization.
[0089] In addition, the multi-algorithm pipeline can add or remove algorithms as needed, providing a high degree of flexibility. In addition, the system can be easily expanded to include more algorithms or process more data.
[0090] The multi-algorithm pipeline processing method significantly improves the efficiency, quality and overall performance of image processing by applying multiple image processing algorithms in parallel. This processing method is crucial for modern image processing systems, especially in scenarios that require fast, efficient and high-quality image analysis.
[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A row neighborhood image storage body, characterized in that: include: N memory chips of the same model are stacked horizontally, and the address bits and timing bits of the N memory chips are connected together one by one, while the data bits are independent of each other; Performing bit segmentation fission on the data bits of each storage chip to form a plurality of data bit segments, each data bit segment corresponding to an image pixel, so as to form the row neighborhood image storage body; Among them, a single row neighborhood image storage body can store at least one W×H 8-bit grayscale image, wherein W is the number of 8-bit grayscale image pixels in a row in the horizontal direction, and H is the number of 8-bit grayscale image pixels in a column in the vertical direction; W≥512, H≥512.
2. The row neighborhood image storage according to claim 1, characterized in that: The data bits of each memory chip are 8×k bits, where k≥2.
3. The row neighborhood image storage according to claim 1 or 2, characterized in that: The data bits of each memory chip are segmented and split into 8 bits per pixel.
4. The row neighborhood image storage according to claim 1, characterized in that: N≥2 and N≤8.
5. A computing platform, characterized in that: It comprises two row neighborhood image storage bodies and a processing chip as claimed in any one of claims 1 to 4, wherein the two row neighborhood image storage bodies are respectively connected to the processing chip.
6. A pipeline processing method for image data, characterized in that: Based on the computing platform according to claim 5, the method comprises: Determine the specific location of the image data to be read in the first row of the neighborhood image storage body by the processing chip; Sending a read instruction to the first row neighborhood image storage body through a processing chip to locate a first source data bit segment of a storage chip of the first row neighborhood image storage body, and performing a parallel read operation on the first source data bit segment to obtain first source image data; Processing the first source image data by a processing chip to obtain first target image data; The first destination data bit segment of the storage chip of the second row neighborhood image storage body that needs to be stored is determined by the processing chip, and the first target image data is sent to the first destination data bit segment for storage.
7. The image data pipeline processing method according to claim 6, characterized in that: After sending the first target image data to the first destination data bit segment for storage through the processing chip, the method further includes: Determine the specific location of the image data to be read in the second row neighborhood image storage body by the processing chip; Sending a read instruction to the second row neighborhood image storage body through the processing chip to locate the second source data bit segment of the storage chip of the second row neighborhood image storage body, and performing a parallel read operation on the second source data bit segment to obtain second source image data; Processing the second source image data by a processing chip to obtain second target image data; The processing chip determines the second destination data bit segment of the storage chip of the first row neighborhood image storage body that needs to be stored, and sends the second target image data to the second destination data bit segment for storage.
8. The image data pipeline processing method according to claim 6, characterized in that: The image data to be read includes W×H image pixels; The processing chip performs a parallel read operation on the first source data bit segment to obtain the first source image data, including: By adopting the memory bank continuous reading mode, W image pixels of each row of the first source data bit segment are read out successively in row order through the processing chip until the image pixels of H rows are read out, thereby obtaining the first source image data.
9. The image data pipeline processing method according to claim 8, characterized in that: Processing the first source image data by a processing chip to obtain first target image data specifically includes: The first source image data is processed by a processing chip based on a preset algorithm to obtain the first target image data.
10. The image data pipeline processing method according to claim 9, characterized in that: There are multiple preset algorithms; Processing the first source image data based on a preset algorithm by a processing chip to obtain the first target image data specifically includes: The first source image data is processed in the order of the 1st row to the Hth row through the processing chip using a processing pipeline in the order of multiple algorithms to obtain row target image data corresponding to each algorithm, until all rows of source image data are processed to obtain first target image data corresponding to each algorithm.