FPGA (Field Programmable Gate Array) parallel computing unit scheduling method and device for image processing
By establishing a PCIe interface data transmission channel, allocating physical memory buffers, and building descriptor queues in the FPGA parallel computing unit, the data block size and cache management are optimized, which solves the data transmission and status monitoring problems in the existing technology and realizes efficient and secure image processing.
Patent Information
- Application Number
- CN202511136110.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-14
AI Technical Summary
Existing FPGA parallel computing unit scheduling methods have shortcomings in data block processing, data transmission and cache management, and status monitoring, which affect processing efficiency and performance.
A data transmission channel is established through the PCIe interface, a physically continuous memory area is allocated as a direct memory access buffer, a descriptor queue is created to record data transmission information, the optimal size of the data block is calculated according to the image resolution and the overlapping area is set, a two-dimensional cache array is constructed and the processing delay is monitored, and dynamic buffer adjustment and status tracking are achieved.
It significantly improves the security and reliability of FPGA parallel computing, solves deficiencies in data security, access control, and privacy protection, and improves processing efficiency and system performance.
Smart Images

Figure CN120744980A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular to a method and device for scheduling FPGA parallel computing units for image processing. Background Art
[0002] Existing FPGA parallel computing unit scheduling methods have obvious shortcomings. Traditional systems lack optimization in data block processing and fail to dynamically adjust data block size based on actual image resolution, affecting processing efficiency.
[0003] Furthermore, existing technologies face bottlenecks in data transmission and cache management. Most systems lack a comprehensive PCIe transmission mechanism and dynamic buffer adjustment strategies, resulting in insufficient parallel computing efficiency.
[0004] Existing systems have technical shortcomings in state monitoring. The lack of real-time monitoring capabilities for processing delays makes it difficult to achieve efficient resource scheduling through state tracking, impacting overall system performance. Addressing these issues is crucial for improving FPGA parallel computing performance. Summary of the Invention
[0005] In response to the problems in the existing technology, the present application provides an FPGA parallel computing unit scheduling method and device for image processing, which can effectively solve the shortcomings of traditional technologies in data security, access control and privacy protection, and significantly improve the security and reliability of FPGA parallel computing.
[0006] In order to solve at least one of the above problems, the present application provides the following technical solutions: In a first aspect, the present application provides an FPGA parallel computing unit scheduling method for image processing, comprising: A data transmission channel is established between a host and a field programmable gate array via a PCIe interface, a physically continuous memory area is allocated as a direct memory access buffer, a descriptor queue is created in the shared memory to record source and destination address information of the data transmission, the descriptor queue is mapped to the user space, the burst transmission length and address alignment are configured, the optimal size of the data block is calculated according to the image resolution, and an overlapping area is set; converting the image data into grayscale values and dividing the data into a plurality of data blocks according to the optimal size, adding boundary pixels to each of the data blocks and setting a timestamp including the data block number and processing time, writing the data blocks into the direct memory access buffer, triggering a doorbell register to send a task signal to a field programmable gate array, and causing the field programmable gate array to pre-fetch descriptors from the descriptor queue and write the data blocks into a dynamic random access memory; A two-dimensional cache array is constructed and the processed data blocks are written to corresponding locations according to the timestamp, the processing delay is monitored and the buffer size is dynamically adjusted, the field programmable gate array writes the data block status into the status register and triggers an interrupt signal, the missing and duplicate conditions of the data block are detected by the status tracker, the host reads the data block from the dynamic random access memory according to the content of the status register, and the host sends a data read completion signal to the field programmable gate array.
[0007] Furthermore, the method further includes: loading a PCIe device driver on the host side and enabling an interrupt response mechanism, allocating a physical continuous memory area through a kernel module and registering the area as a direct memory access buffer, mapping the direct memory access buffer to a user space address, configuring the PCIe link to the third generation specification and setting the link width, and starting a bus master to connect the PCIe interface to an advanced extensible interface bus; The field programmable gate array loads the PCIe intellectual property core configuration parameters, configures the Advanced Extensible Interface Memory Mapping Protocol and the Advanced Extensible Interface Reduced Interface Protocol, creates a descriptor queue in the shared memory to record the source and destination address information of the data transmission, sets the size of the direct memory access receive buffer and the transmit buffer according to system requirements, configures the burst transmission length and address alignment, and establishes an interrupt vector table to bind the task identifier to the interrupt handling function.
[0008] Furthermore, the method further includes: mapping the physical address space of the descriptor queue to the process virtual address space through a memory mapping function, configuring the base address register space to implement access to the control register and the status register, setting the direct memory access burst transfer length to an integer multiple of the basic transfer unit, configuring the address boundary alignment mode to ensure data access efficiency, and writing the read and write pointer update information of the descriptor queue into the doorbell register; The optimal size of a single data block is calculated based on the width and height of the input image. The data block size is set to an integer multiple of the computing power of the processing unit. An overlapping area mapping relationship is established between adjacent data blocks. The pixel width of the overlapping area is calculated to ensure the continuity of the boundary pixels. The data block size and overlapping area parameters are written into the configuration register, and a data block index table is generated to record the association information between blocks.
[0009] Furthermore, the method further includes: converting the RGB color space of the input image into a grayscale image, dividing the grayscale image data into a plurality of data blocks based on the optimal size, adding boundary pixels around each of the data blocks to form an extended data block, calculating position information of the extended data block in a global coordinate system, generating a data block descriptor including a start address and an end address, and assigning a globally unique data block number to each of the data blocks; Create a timestamp management unit to record the processing status of the data block, the timestamp includes two time attribute fields, the data block number and the data block processing time, establish a timestamp mapping table to store the timing information of the data block, write the data block and its corresponding timestamp information into the specified location of the direct memory access buffer, and update the data block status register to record the transmission completion flag.
[0010] Furthermore, the method further includes: writing task information into a specified offset address of a base address register space, triggering a doorbell register to generate an interrupt request signal, configuring an interrupt mask bit of an interrupt status register, activating an interrupt controller to enable an interrupt response, reading an interrupt vector table to obtain an entry address of an interrupt service routine, sending a task start signal to a field programmable gate array, and updating a task status register to record a task execution status; The field programmable gate array reads descriptor information from the descriptor queue to a prefetch buffer, parses the descriptor to obtain a source address and a destination address of a data block, configures transmission parameters of a direct memory access controller, establishes a burst transmission channel to write the data block into a designated storage unit of a dynamic random access memory, updates a read pointer of the descriptor queue, and writes a data block transmission completion flag into a status register.
[0011] Furthermore, the method further includes: calculating the number of rows and columns of the two-dimensional cache array according to the image resolution and the size of the data block, allocating a fixed-size storage space to each cache unit, reading the timestamp information to determine the write location of the data block, establishing a cache address mapping table to record the storage location of the data block in the two-dimensional cache array, writing the processed data block to the corresponding cache unit, and updating the cache status table to record the storage status of the data block; The processing time field in the timestamp is read to calculate the processing delay of the data block, the buffer size is adjusted according to the statistical value of the processing delay, the buffer threshold is configured to control the read and write rate of the data block, the field programmable gate array writes the processing status of the data block into a specified bit field of the status register, generates an interrupt request signal to notify the processing completion event, and updates the interrupt flag bit of the interrupt status register.
[0012] Furthermore, the method further includes: a state tracker reading a cache state table of the two-dimensional cache array, checking the continuity of data blocks according to data block numbers, establishing a data block state mapping table to record missing and duplicate data block information, writing the data block state mapping table to a specific bit field of a status register, updating a detection counter of the state tracker, and generating a data block state detection report to record location information of abnormal data blocks; The host reads the status register to obtain the processing status of the data block, configures the read parameters of the direct memory access controller, reads the processed data block from the dynamic random access memory to the system memory, updates the read status flag of the data block, writes the read completion flag to the field programmable gate array, clears the corresponding interrupt status bit, and releases the storage space occupied by the read data block.
[0013] In a second aspect, the present application provides an FPGA parallel computing unit scheduling device for image processing, comprising: an image processing module for establishing a data transmission channel between a host and a field programmable gate array through a PCIe interface, allocating a physically continuous memory area as a direct memory access buffer, creating a descriptor queue in the shared memory to record source and destination address information of the data transmission, mapping the descriptor queue to the user space, configuring the burst transmission length and address alignment, calculating the optimal size of the data block according to the image resolution, and setting the overlapping area; a data writing module, configured to convert the image data into grayscale values and divide the data into a plurality of data blocks according to the optimal size, add boundary pixels to each data block and set a timestamp including the data block number and processing time, write the data block into the direct memory access buffer, trigger a doorbell register to send a task signal to a field programmable gate array, and cause the field programmable gate array to prefetch descriptors from the descriptor queue and write the data block into a dynamic random access memory; In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, the steps of the FPGA parallel computing unit scheduling method for image processing are implemented.
[0014] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the FPGA parallel computing unit scheduling method for image processing.
[0015] In a fifth aspect, the present application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the FPGA parallel computing unit scheduling method for image processing.
[0016] It can be seen from the above technical solution that the present application provides an FPGA parallel computing unit scheduling method and device for image processing, which realizes data security isolation and access control through innovatively designed data security transmission mechanism, physical continuous memory allocation and descriptor queue management. A timestamp-based data block encryption strategy is constructed, combined with dynamic buffer adjustment and status monitoring to establish a reliable data protection system. A processing delay monitoring mechanism is introduced to ensure the privacy of the data processing process through real-time performance evaluation and resource scheduling. Based on data encryption transmission, this method effectively solves the shortcomings of traditional technologies in data security, access control and privacy protection, and significantly improves the security and reliability of FPGA parallel computing. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 Schematic diagram of the flow of the FPGA parallel computing unit scheduling method for image processing in an embodiment of the present application; Figure 2 This is a structural diagram of an FPGA parallel computing unit scheduling device for image processing in an embodiment of the present application; Figure 3 Schematic diagram of the structure of the electronic device in the embodiment of the present application.
[0019] Reference numerals: Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver program storage unit 9144, antenna 9111, speaker 9131, microphone 9132. DETAILED DESCRIPTION
[0020] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0021] The acquisition, storage, use, and processing of data in this application's technical solution comply with relevant national laws and regulations.
[0022] Taking into account the problems existing in the prior art, the present application provides an FPGA parallel computing unit scheduling method and device for image processing, which realizes data security isolation and access control through innovatively designed data security transmission mechanism, physical continuous memory allocation and descriptor queue management. A timestamp-based data block encryption strategy is constructed, combined with dynamic buffer adjustment and status monitoring to establish a reliable data protection system. A processing delay monitoring mechanism is introduced to ensure the privacy of the data processing process through real-time performance evaluation and resource scheduling. Based on data encryption transmission, this method effectively solves the shortcomings of traditional technologies in data security, access control and privacy protection, and significantly improves the security and reliability of FPGA parallel computing.
[0023] In order to effectively solve the deficiencies of traditional technologies in data security, access control, and privacy protection, and significantly improve the security and reliability of FPGA parallel computing, this application provides an embodiment of an FPGA parallel computing unit scheduling method for image processing, see Figure 1 The FPGA parallel computing unit scheduling method for image processing specifically includes the following contents: Step S101: establishing a data transmission channel between a host and a field programmable gate array via a PCIe interface, allocating a physically continuous memory area as a direct memory access buffer, creating a descriptor queue in the shared memory to record the source and destination address information of the data transmission, mapping the descriptor queue to the user space, configuring the burst transmission length and address alignment, calculating the optimal size of the data block based on the image resolution, and setting the overlap area; Optionally, this embodiment addresses the problems of low PCIe transmission bandwidth utilization, poor memory management efficiency, and unreasonable data block partitioning in traditional FPGA image processing systems by innovatively designing a data transmission optimization solution based on intelligent scheduling. Regarding PCIe transmission efficiency evaluation, this embodiment designs a transmission efficiency scoring formula: Transfer_Score = α×(Bandwidth_Utilization) + β×(DMA_Efficiency) + γ×(Memory_Coherency), where α, β, and γ are weight coefficients representing bandwidth utilization, DMA transmission efficiency, and memory coherence, respectively. At the same time, a data block optimization scoring formula is introduced: Block_Score = (Processing_Units × Block_Size) / (Overlap_Size + Transfer_Overhead) to evaluate the rationality of data block partitioning.
[0024] This embodiment first achieves efficient data transmission through a deeply optimized PCIe interface configuration. During the PCIe channel establishment process, the system adopts a hierarchical configuration strategy to complete the parameter settings of the physical layer, data link layer, and transaction layer in sequence. In the physical layer configuration, the system selects the appropriate PCIe generation and channel width based on the actual bandwidth requirements. For example, for high-resolution image processing, the PCIe Gen3 x8 configuration is preferred to provide sufficient transmission bandwidth. In the data link layer configuration, the system improves the reliability of data transmission by optimizing flow control parameters and retransmission strategies. In the transaction layer configuration, the system implements request priority management to ensure that key data transmission requests can be processed first. In order to improve transmission efficiency, the system sets the maximum load size and maximum read request size in the PCIe configuration space to align them with the processor cache line size, reducing the fragmentation overhead during data transmission. This multi-level PCIe optimization mechanism provides efficient and reliable channel guarantees for subsequent data transmission.
[0025] This embodiment innovatively implements a continuous physical memory allocation mechanism. When allocating a DMA buffer, the system uses a reserved memory area to ensure the continuity of the allocated physical memory space. In the specific implementation, the system reduces the number of page table entries and improves address conversion efficiency through the large page table mechanism of the memory management unit. At the same time, the system establishes a buffer management table to record the usage status and access rights of each memory block. In order to improve memory access efficiency, the system considers the NUMA architecture characteristics when allocating memory, and gives priority to memory areas local to the processor node. In the process of creating a descriptor queue, the system adopts a ring buffer design and realizes efficient descriptor access through read-write pointer management. The queue size is dynamically adjusted according to actual processing needs to avoid waste of resources. This optimized memory management mechanism significantly improves data transmission efficiency.
[0026] This embodiment achieves efficient user space access through an optimized mapping mechanism. When mapping the descriptor queue to the user space, the system uses zero-copy technology to avoid the data copy overhead between the kernel space and the user space. During the mapping process, the system implements fine-grained access control through page table entry settings to ensure the security of data access. At the same time, the system implements a mapping address cache mechanism to reduce the overhead of repeated mapping operations. In terms of burst transmission configuration, the system selects the optimal burst transmission length based on the PCIe channel characteristics and memory access mode. Address alignment requirements are implemented through hardware masks to ensure that all transmission operations are performed on legal address boundaries. This sophisticated configuration mechanism provides stable performance guarantees for data transmission.
[0027] This embodiment establishes an intelligent data block partitioning mechanism. When calculating the optimal data block size, the system comprehensively considers multiple factors, including the parallelism of the FPGA processing unit, PCIe transmission efficiency, and memory access characteristics. By establishing a performance model, the system can predict the processing efficiency under different block sizes and thus select the optimal configuration. When setting the overlapping area, the system determines the appropriate number of overlapping pixels based on the characteristics of the image processing algorithm to ensure the accuracy of boundary processing. In order to improve processing efficiency, the system implements a data block prefetching mechanism to preload the next data block while the current block is being processed. This optimized partitioning mechanism significantly improves parallel processing efficiency.
[0028] This embodiment provides a complete data management solution. Through multi-layered optimization design, the system optimizes the entire process from PCIe transmission to data block processing, significantly improving the overall performance of the FPGA image processing system. This deeply optimized management mechanism provides reliable technical support for practical applications.
[0029] Step S102: converting the image data into grayscale values and dividing the data into a plurality of data blocks according to the optimal size, adding boundary pixels to each data block and setting a timestamp including the data block number and processing time, writing the data block into the direct memory access buffer, triggering a doorbell register to send a task signal to a field programmable gate array, and causing the field programmable gate array to pre-fetch descriptors from the descriptor queue and write the data block into a dynamic random access memory; Optionally, this embodiment addresses the issues of low data preprocessing efficiency, discontinuous block boundary processing, and complex data transmission synchronization in traditional FPGA image processing by designing a timestamp-based data preprocessing and transmission solution. Regarding data preprocessing evaluation, this embodiment devises a preprocessing efficiency scoring formula: Process_Score = α × (Color_Convert / Base_Time) + β × (Block_Efficiency) + γ × (Border_Quality), where α, β, and γ are weight coefficients representing the color conversion speed ratio, data block partitioning efficiency, and boundary processing quality, respectively. A transmission efficiency scoring formula: Transfer_Score = (Block_Size × Queue_Depth) / (Setup_Time + Transfer_Delay) is also introduced to evaluate data transmission performance. These two scoring metrics comprehensively reflect the performance characteristics of the preprocessing and transmission processes.
[0030] This embodiment first achieves high-quality grayscale conversion through a deeply optimized color conversion mechanism. For input images of varying formats, an adaptive channel extraction algorithm is used to separate RGB data. During grayscale conversion, the system dynamically adjusts conversion weights based on image content characteristics. For example, for areas containing significant detail, the green channel weight is increased to preserve more texture information; for areas with highlights, the weight is appropriately reduced to avoid information loss. The converted grayscale data is quickly accessed through a multi-level cache mechanism, and continuous pixel data is processed in a pipelined manner. Especially when processing high-resolution images, the system simultaneously converts multiple pixel data points through a parallel processing mechanism, significantly improving processing efficiency. During data block partitioning, the system employs an adaptive block partitioning strategy based on previously determined optimal size parameters. Considering the parallel processing capabilities and storage resource limitations of the FPGA, the selection of block size requires a balance between processing efficiency and resource utilization. By establishing a data block management table, the system implements efficient block-level management, including block creation, location, and status tracking.
[0031] This embodiment adopts an innovative boundary processing mechanism to ensure data continuity. When adding boundary pixels to data blocks, the system adopts different filling strategies according to the spatial position of the block. For data blocks at the edge of the image, boundary pixels are generated by mirror filling or boundary extension; for internal data blocks, real overlapping pixels are extracted from adjacent blocks. This intelligent boundary processing mechanism effectively avoids the discontinuity of inter-block processing. The timestamp design adopts a dual-field structure. The data block number field ensures global uniqueness, and the processing time field is used for task scheduling and status tracking. The system maintains the temporal relationship of data blocks through the timestamp mapping table and supports block-level synchronization and sequence control based on timestamps. This complete preprocessing mechanism significantly improves the quality and efficiency of data preparation.
[0032] This embodiment implements an efficient data transmission mechanism. When writing data blocks to the DMA buffer, the system adopts a batch transmission strategy and implements continuous data storage through a pre-allocated buffer. The doorbell register is triggered using an edge trigger method to ensure reliable signal transmission. The FPGA side reads descriptor information in advance through a pre-fetch mechanism, establishes a transmission request queue, and implements pipeline operation of the transmission process. When data is written to the dynamic random access memory, the system quickly locates the storage location through the address mapping table and adopts a burst transmission mode to improve data throughput. This optimized transmission mechanism significantly reduces data transmission overhead.
[0033] This embodiment provides a reliable preprocessing and transmission solution. Through optimized color conversion, data segmentation, and transmission control, the system achieves an efficient data preparation process. This reliable processing mechanism provides high-quality data support for subsequent FPGA parallel computing. The innovation of this embodiment lies in the deep optimization of the data preprocessing and transmission process. Through reasonable resource scheduling and precise process control, the system's processing efficiency is significantly improved. This optimization effect is particularly evident when processing large-scale image data. At the same time, the design of the solution fully considers the actual application needs and has good scalability and adaptability.
[0034] Step S103: construct a two-dimensional cache array and write the processed data block to the corresponding position according to the timestamp, monitor the processing delay and dynamically adjust the buffer size, the field programmable gate array writes the data block status into the status register and triggers an interrupt signal, detects the missing and duplication of the data block through the status tracker, the host reads the data block from the dynamic random access memory according to the content of the status register, and the host sends a data reading completion signal to the field programmable gate array.
[0035] Optionally, this embodiment addresses issues such as disorganized two-dimensional data, inefficient cache management, and unreliable state synchronization in FPGA image processing by innovatively designing a data management solution based on adaptive caching. Regarding cache performance evaluation, this embodiment devises a cache efficiency scoring formula: Cache_Score = α × (Access_Speed / Base_Speed) + β × (Space_Utilization) + γ × (Hit_Rate), where α, β, and γ are weight coefficients representing access speed ratio, space utilization, and hit rate, respectively. A delay evaluation formula: Delay_Score = (Processing_Time × Block_Count) / (Buffer_Size + Sync_Overhead) is also introduced to evaluate the relationship between processing delay and buffer configuration. These two scoring metrics comprehensively reflect data management performance.
[0036] This embodiment first achieves efficient data management through a deeply optimized cache organization mechanism. The construction of the two-dimensional cache array adopts a hierarchical design, including a data storage layer, an address mapping layer, and a state management layer. In the data storage layer, the system organizes data using a block storage structure based on the characteristics of the internal storage resources of the FPGA. The size of each storage unit is optimized according to the data block size to ensure efficient use of storage space. The address mapping layer realizes rapid conversion from timestamp to storage location and supports efficient data positioning by establishing a multi-level mapping table. The state management layer maintains the usage status of each storage unit, including idle, writing, occupied and other status information. Especially when processing large-scale data, this hierarchical management mechanism significantly improves data access efficiency. In order to handle data write conflicts, the system implements a timestamp-based arbitration mechanism to ensure that data is written to the cache in the correct order. At the same time, by establishing a write queue, the system can cache multiple data blocks to be written, reducing write waiting time.
[0037] This embodiment innovatively implements an adaptive buffer management mechanism. When monitoring processing delays, the system uses a sliding window method to count the processing time of data blocks, and predicts future processing loads by analyzing the changing trend of processing time. The buffer size is adjusted using a gradual strategy to avoid frequent and large adjustments that lead to system instability. When it is detected that the processing delay continues to increase, the system will appropriately increase the buffer size to provide more data cache space; when the processing delay decreases, the buffer size will be reduced accordingly to avoid resource waste. This dynamic adjustment mechanism can optimize resource allocation according to the actual processing load and improve the adaptability of the system. At the same time, the system implements a buffer data pre-fetching mechanism to load the data to be processed in advance according to the timestamp of the data block, thereby reducing data waiting time.
[0038] This embodiment achieves reliable data tracking through an optimized state synchronization mechanism. The status register adopts a bit field design, and different bit fields record the processing status, error flag and control information of the data block respectively. The interrupt signal is generated using an edge trigger method to ensure the reliable transmission of the signal. The state tracker detects the missing and duplication of data blocks by scanning the timestamp sequence. When an anomaly is found, the error flag is immediately updated and the corresponding recovery mechanism is activated. For example, for missing data blocks, the system will try to re-request the data; for duplicate data blocks, the version with the newer timestamp is retained. This complete state management mechanism significantly improves the reliability of data processing.
[0039] This embodiment establishes an efficient data readback mechanism. After receiving the interrupt signal, the host quickly locates the data block to be processed by reading the status register. The data reading process adopts a batch transfer method, and efficient data handling is achieved through the DMA controller. After the reading is completed, the system notifies the FPGA through a special completion signal, triggering the release of related resources. This optimized readback mechanism significantly reduces the data transmission overhead. The innovation of this embodiment is that it realizes the all-round optimization of the data management process, and significantly improves the processing efficiency of the system through adaptive resource scheduling and reliable state synchronization. This optimization effect is particularly obvious when processing complex image data. At the same time, the design of the solution fully considers the actual application needs and has good scalability and robustness.
[0040] From the above description, it can be seen that the FPGA parallel computing unit scheduling method for image processing provided by the embodiment of the present application can realize data security isolation and access control through innovative design of data security transmission mechanism, physical continuous memory allocation and descriptor queue management. Build a data block encryption strategy based on timestamp, combine dynamic buffer adjustment and status monitoring, and establish a reliable data protection system. Introduce a processing delay monitoring mechanism to ensure the privacy of the data processing process through real-time performance evaluation and resource scheduling. Based on data encryption transmission, this method effectively solves the shortcomings of traditional technologies in data security, access control and privacy protection, and significantly improves the security and reliability of FPGA parallel computing.
[0041] In one embodiment of the FPGA parallel computing unit scheduling method for image processing of the present application, the following contents may also be specifically included: Step S201: The host side loads a PCIe device driver and enables an interrupt response mechanism, allocates a physically continuous memory area through a kernel module and registers it as a direct memory access buffer, maps the direct memory access buffer to a user space address, configures the PCIe link to the third generation specification and sets the link width, and starts a bus master to connect the PCIe interface to an Advanced Extensible Interface bus. Step S202: The field programmable gate array loads the PCIe intellectual property core configuration parameters, configures the Advanced Extensible Interface Memory Mapping Protocol and the Advanced Extensible Interface Reduced Interface Protocol, creates a descriptor queue in the shared memory to record the source address and destination address information of the data transmission, sets the size of the direct memory access receive buffer and the send buffer according to system requirements, configures the burst transmission length and address alignment, and establishes an interrupt vector table to bind the task identifier to the interrupt processing function.
[0042] Optionally, this embodiment addresses issues such as complex PCIe interface configuration, inefficient memory management, and difficulty synchronizing data transmission in FPGA systems by innovatively designing an interface initialization solution based on dual-end collaboration. Regarding PCIe performance evaluation, this embodiment employs a link efficiency scoring formula: Link_Score = α × (Link_Width × Link_Speed) + β × (DMA_Efficiency) + γ × (Interrupt_Response), where α, β, and γ are weight coefficients representing link bandwidth, DMA transmission efficiency, and interrupt response speed, respectively. A protocol adaptation scoring formula: Protocol_Score = (Transaction_Rate × Buffer_Utilization) / (Setup_Overhead + Sync_Delay) is also introduced to evaluate the interface protocol's adaptation.
[0043] This embodiment first achieves reliable device initialization through a deeply optimized driver loading mechanism. During the host-side driver loading process, the system adopts a hierarchical initialization strategy to complete device identification, resource allocation and function configuration in sequence. The driver reads the device identification information through the PCI configuration space to ensure correct matching with the FPGA device. The interrupt response mechanism adopts the MSI-X method, supports multi-queue interrupt distribution, and significantly improves interrupt processing efficiency. Especially when processing high-concurrency data transmission, the system binds interrupt requests to specific processor cores through interrupt affinity configuration to reduce interrupt processing delays. During the memory allocation process, the system applies for physically continuous memory pages through the kernel memory management interface, avoiding memory fragmentation problems. In order to improve memory access efficiency, the system takes into account the processor cache line size when allocating DMA buffers to ensure data access alignment. This optimized initialization mechanism provides a stable hardware foundation for subsequent data transmission.
[0044] This embodiment innovatively implements a PCIe link configuration mechanism. The PCIe link configuration adopts an adaptive adjustment strategy to select the optimal link parameters based on the actual bandwidth requirements. The system is preferentially configured to the PCIe Gen3 specification to support higher transmission rates. The selection of link width takes into account FPGA resource constraints and actual bandwidth requirements, and the optimal configuration is determined through a performance evaluation model. During the link training process, the system ensures signal quality through fine-grained electrical parameter adjustments. The startup of the bus master controller adopts a progressive strategy, first completing basic initialization and then gradually enabling advanced functions. The connection with the advanced extensible interface bus is achieved through a bridge controller, which supports address space mapping and transmission protocol conversion. This sophisticated link configuration mechanism significantly improves the reliability of data transmission.
[0045] This embodiment achieves efficient protocol adaptation through optimized IP core configuration. The configuration of the PCIe IP core on the FPGA side adopts a parameterized design to support flexible function customization. When configuring the advanced extensible interface protocol, the system supports both memory mapping and streamlined interface modes to meet the transmission requirements of different scenarios. The creation of the descriptor queue adopts a ring buffer structure, and efficient descriptor access is achieved through read and write pointer management. The buffer size is set based on actual processing requirements, and the optimal configuration is predicted through the performance model. The selection of burst transmission parameters takes into account the memory access characteristics to ensure transmission efficiency. The establishment of the interrupt vector table adopts a disperse-gather method to support a flexible interrupt handling mechanism. This complete protocol configuration mechanism provides reliable protocol support for data transmission.
[0046] This embodiment establishes a comprehensive resource management framework. From driver loading to protocol configuration, every step has been thoroughly optimized to ensure efficient resource utilization. DMA buffer management employs a multi-level caching strategy to support fast data access. By establishing a complete interrupt handling framework, the system implements an efficient task synchronization mechanism. This optimized resource management mechanism significantly improves overall system performance.
[0047] This embodiment provides an innovative solution for PCIe interface configuration. Through a dual-end collaborative configuration mechanism, the system achieves full-stack optimization from hardware to software. This reliable configuration mechanism provides an efficient data transmission channel for FPGA image processing systems. The technical solution of this embodiment offers excellent scalability, adapting to data processing needs of varying scales and providing reliable technical support for FPGA application development.
[0048] In one embodiment of the FPGA parallel computing unit scheduling method for image processing of the present application, the following contents may also be specifically included: Step S301: Mapping the physical address space of the descriptor queue to the process virtual address space through a memory mapping function, configuring the base address register space to implement access to the control register and status register, setting the direct memory access burst transfer length to an integer multiple of the basic transfer unit, configuring the address boundary alignment mode to ensure data access efficiency, and writing the read and write pointer update information of the descriptor queue into the doorbell register; Step S302: Calculate the optimal size of a single data block based on the width and height of the input image, set the data block size to an integer multiple of the computing power of the processing unit, establish an overlapping area mapping relationship between adjacent data blocks, calculate the pixel width of the overlapping area to ensure the continuity of the boundary pixels, write the data block size and overlapping area parameters into the configuration register, and generate a data block index table to record the association information between blocks.
[0049] Optionally, this embodiment addresses issues such as complex address mapping, irrational data block partitioning, and inefficient boundary processing in FPGA image processing by innovatively designing a data organization solution based on intelligent mapping. Regarding address mapping evaluation, this embodiment devises a mapping efficiency scoring formula: Map_Score = α × (Translation_Speed / Base_Speed) + β × (Space_Efficiency) + γ × (Access_Latency), where α, β, and γ are weight coefficients representing the address translation speed ratio, space utilization, and access latency, respectively. A data block optimization scoring formula: Block_Score = (Processing_Units × Block_Size) / (Overlap_Cost + Communication_Overhead) is also introduced to evaluate the rationality of data block partitioning.
[0050] This embodiment first achieves efficient storage access through a deeply optimized address mapping mechanism. In the address mapping process of the descriptor queue, the system adopts a hierarchical mapping strategy to map the physical address space into segments of the virtual address space. The implementation of the mapping function takes into account the memory access mode and adopts a page table hierarchical mechanism to reduce the address conversion overhead. In order to improve access efficiency, the system pre-calculates page table entries during the mapping process and caches hot page table entries in the TLB. In particular, for frequently accessed descriptor areas, the system reduces the page table hierarchy through a large page table mechanism, significantly reducing address conversion delays. In the configuration of the base address register space, the system adopts a regional partitioning strategy to map the control register and status register to different address intervals for easy independent access and management. This optimized mapping mechanism provides efficient address conversion support for data access.
[0051] This embodiment innovatively implements a transmission parameter configuration mechanism. The DMA burst transfer length is set based on the data block characteristics and bus characteristics, and the optimal transfer unit is calculated through a performance model. The system configures the burst length as an integer multiple of the basic transfer unit to ensure transmission efficiency. Address alignment requirements are implemented through hardware masks, and all transfer operations are performed on legal address boundaries. The read and write pointers of the descriptor queue are updated using atomic operations, and the processing unit is notified in real time through the doorbell register. This sophisticated parameter configuration mechanism significantly improves data transmission efficiency. The doorbell mechanism is implemented using edge triggering to ensure reliable transmission of update signals. At the same time, the system tracks the usage of the queue in real time by establishing a pointer status table, supporting dynamic queue management.
[0052] This embodiment achieves efficient image processing through an optimized data block partitioning mechanism. When calculating the optimal size of the data block, the system comprehensively considers multiple factors: first, based on the resolution characteristics of the input image, the basic data block size is calculated; then, according to the parallelism and computing power of the FPGA processing unit, the block size is adjusted to meet the hardware processing efficiency requirements; finally, storage resource limitations are considered to ensure that the data block can effectively utilize the on-chip storage. The final size of the data block is set to an integer multiple of the computing power of the processing unit. This alignment mechanism ensures that the processing unit can be fully utilized. When establishing the overlapping area, the system determines the minimum necessary number of overlapping pixels by analyzing the characteristics of the image processing algorithm. The mapping of the overlapping area uses shared memory to support data sharing between adjacent blocks. This optimized partitioning mechanism significantly improves processing efficiency.
[0053] This embodiment establishes a complete data management framework. The configuration register design utilizes a multi-field structure, storing not only data block parameters but also processing control information. The data block index table utilizes a multi-level index structure, supporting fast inter-block relationship queries. By establishing a comprehensive parameter management mechanism, the system achieves efficient organization and access of data blocks. This optimized management mechanism provides reliable data support for image processing.
[0054] This embodiment provides an innovative solution for data organization. Through optimized address mapping and data block management, the system optimizes the entire process, from storage access to data organization. This reliable organization mechanism provides efficient data support for FPGA image processing. The technical solution of this embodiment has good scalability and can adapt to image processing requirements of various scales and types, providing reliable technical support for FPGA application development. This optimization effect is particularly evident when processing large-scale image data.
[0055] In one embodiment of the FPGA parallel computing unit scheduling method for image processing of the present application, the following contents may also be specifically included: Step S401: Convert the RGB color space of the input image into a grayscale image, divide the grayscale image data into multiple data blocks based on the optimal size, add boundary pixels around each data block to form an extended data block, calculate the position information of the extended data block in the global coordinate system, generate a data block descriptor including a start address and an end address, and assign a globally unique data block number to each data block; Step S402: Create a timestamp management unit to record the processing status of the data block, the timestamp contains two time attribute fields, the data block number and the data block processing time, establish a timestamp mapping table to store the timing information of the data block, write the data block and its corresponding timestamp information into the specified location of the direct memory access buffer, and update the data block status register to record the transmission completion flag.
[0056] Optionally, this embodiment addresses issues such as inaccurate color conversion, chaotic data block management, and unreliable timing control in FPGA image processing by innovatively designing a timestamp-based data preprocessing solution. In terms of color conversion evaluation, this embodiment designs a conversion quality scoring formula: Color_Score = α×(Detail_Preserve) + β×(Edge_Clarity) + γ×(Noise_Suppress), where α, β, and γ are weight coefficients representing detail preservation, edge clarity, and noise suppression, respectively. At the same time, a timing management scoring formula is introduced: Timing_Score = (Processing_Speed × Block_Count) / (Timestamp_Overhead + Sync_Delay) to evaluate the effectiveness of timing control.
[0057] This embodiment first achieves high-quality grayscale conversion through a deeply optimized color conversion mechanism. When processing RGB images, the system uses an adaptive weighting method for channel synthesis, dynamically adjusting the contribution of each channel by analyzing image content features. For highlight areas, the system appropriately reduces the weight of the green channel to avoid information oversaturation; for dark areas, the red channel weight is increased to retain more detailed information. When calculating weights, the system takes into account the visual characteristics of the human eye, making the converted grayscale image more consistent with human perception. In particular, when processing images containing rich textures, the system improves detail performance through a local contrast enhancement algorithm. The conversion process adopts a pipeline design, supporting parallel processing of multiple pixels, significantly improving conversion efficiency. During the data block division process, the system adopts an adaptive block strategy based on the previously determined optimal size. The size of each data block not only takes into account the processing power of the FPGA, but also takes into account the principle of local data access to ensure processing efficiency. This optimized conversion mechanism provides high-quality input data for subsequent processing.
[0058] This embodiment innovatively implements a boundary extension mechanism. When adding boundary pixels to a data block, the system adopts different filling strategies according to the position of the block. For data blocks inside the image, the system extracts actual pixel values from adjacent blocks as boundary pixels; for data blocks at the edge of the image, an intelligent filling algorithm is used. For example, for areas containing obvious edge features, the system generates boundary pixels by edge extension; for texture areas, a texture synthesis algorithm is used to fill the boundary. The position calculation of the extended data block adopts a hierarchical mapping method, first determining the base position of the block in the global coordinate system, and then calculating the actual range after the extended boundary. During the descriptor generation process, the system not only records the address information, but also includes the attribute mark of the block to facilitate subsequent processing control. The allocation of data block numbers adopts a partition coding method to ensure the uniqueness and continuity of the numbers. This complete data organization mechanism significantly improves the reliability of data management.
[0059] This embodiment achieves reliable timing control through optimized timestamp management. The timestamp management unit adopts a hierarchical design, including three functional modules: timestamp generation, mapping management, and status tracking. The timestamp design adopts a dual-field structure, the data block number field is used to uniquely identify the data block, and the processing time field records the expected processing time window. This design not only supports timestamp-based data synchronization, but can also be used for load balancing and anomaly detection. The timestamp mapping table adopts a multi-level index structure to support fast timing information query and update. During the data writing process, the system ensures the consistency of timestamp information through atomic operations. The status register is updated using a bit field design, and different bit fields record the transmission status, processing status, and error flag respectively. This sophisticated timing management mechanism significantly improves the controllability of data processing.
[0060] This embodiment establishes a complete data management framework. From color conversion to timing control, each link has been deeply optimized to ensure the quality and efficiency of data processing. Especially when processing large-scale image data, the system significantly improves the processing speed through parallel processing and pipeline operation. The introduction of the timestamp mechanism not only solves the data synchronization problem, but also provides the system with reliable status tracking capabilities. The innovation of this embodiment is that it realizes the all-round optimization of the data preprocessing process, and significantly improves the processing efficiency of the system through reasonable resource scheduling and fine process control. This optimization effect is outstanding in practical applications, especially when processing image processing tasks that require precise timing control, the system exhibits superior performance and reliability.
[0061] In one embodiment of the FPGA parallel computing unit scheduling method for image processing of the present application, the following contents may also be specifically included: Step S501: writing task information to a specified offset address in the base address register space, triggering the doorbell register to generate an interrupt request signal, configuring the interrupt mask bit in the interrupt status register, activating the interrupt controller to enable interrupt response, reading the interrupt vector table to obtain the interrupt service routine entry address, sending a task start signal to the field programmable gate array, and updating the task status register to record the task execution status; Step S502: The field programmable gate array reads descriptor information from the descriptor queue into a prefetch buffer, parses the descriptor to obtain the source address and destination address of the data block, configures the transmission parameters of the direct memory access controller, establishes a burst transmission channel to write the data block into a designated storage unit of the dynamic random access memory, updates the read pointer of the descriptor queue, and writes a data block transmission completion flag into a status register.
[0062] Optionally, this embodiment addresses issues such as unreliable task triggering, interrupt response latency, and low data transmission efficiency in FPGA systems by innovatively designing a task control solution based on intelligent interrupts. Regarding task control evaluation, this embodiment employs the interrupt efficiency scoring formula: Interrupt_Score = α × (Response_Time / Base_Time) + β × (Queue_Efficiency) + γ × (Signal_Stability), where α, β, and γ are weight coefficients representing response time ratio, queue processing efficiency, and signal stability, respectively. Furthermore, a transfer performance scoring formula is introduced: Transfer_Score = (DMA_Speed × Burst_Length) / (Setup_Time + Switching_Overhead) to evaluate data transmission efficiency.
[0063] This embodiment first implements reliable interrupt control through a deeply optimized task triggering mechanism. During the task information writing process, the system adopts a hierarchical write strategy, first writing the task parameters to the specified register space, and then determining the specific register location through offset address calculation. The base address register space is organized in a segmented management manner, and different types of control information are mapped to independent address intervals to avoid access conflicts. The doorbell register is triggered in an edge-sensitive manner, and the reliability of the trigger signal is ensured by a hardware level detection mechanism. During the configuration of the interrupt status register, the system implements fine-grained interrupt control by setting the interrupt mask bit, and can selectively enable or disable specific interrupt sources based on task priority. Especially when dealing with multi-tasking concurrent scenarios, this sophisticated interrupt control mechanism can effectively avoid interrupt storm problems and ensure stable system operation. The interrupt vector table management adopts a dynamic update strategy to support flexible scheduling of interrupt service routines. This optimized triggering mechanism provides reliable interrupt support for task control.
[0064] This embodiment innovatively implements an interrupt response mechanism. The activation of the interrupt controller adopts a step-by-step startup strategy, first completing the basic configuration, and then gradually enabling various functions. The entry address of the interrupt service program is obtained by a fast table lookup method, and efficient program jumps are achieved through a pre-established vector mapping table. The sending of the task start signal adopts a synchronization mechanism to ensure that the FPGA can correctly receive the task start instruction. The task status register is updated using atomic operations to avoid state confusion caused by concurrent access. During the status recording process, the system adopts a multi-field design, which not only includes the basic execution status, but also records detailed progress information. This complete response mechanism significantly improves the reliability of task control.
[0065] This embodiment achieves efficient storage access through an optimized data transmission mechanism. After receiving the task signal, the FPGA first pre-fetches a certain amount of descriptor information from the descriptor queue, and realizes fast access to data through the pre-fetch buffer area. The descriptor is parsed in a parallel processing manner, and the source address and destination address information are extracted at the same time. During the configuration of the DMA controller, the system selects the optimal transmission parameters according to the characteristics of the data block, including burst length, address alignment, etc. The transmission channel is established in a dedicated channel manner to avoid interference with other transmission operations. During the data writing process, the system ensures that the data can be stored continuously through the pre-allocation mechanism of the storage unit, reducing access delays. This optimized transmission mechanism significantly improves data throughput.
[0066] This embodiment establishes a complete state management framework. The read pointer of the descriptor queue is updated using a secure update mechanism, ensuring the atomicity of queue operations. The transfer completion flag is written using a status register bit field design, supporting the parallel recording of multiple states. By establishing a complete state tracking mechanism, the system achieves reliable monitoring of the entire task execution process. This optimized management mechanism provides reliable state support for task control.
[0067] This embodiment provides an innovative solution for task control. Through optimized interrupt handling and data transmission mechanisms, the system optimizes the entire process from task triggering to data transmission. This reliable control mechanism provides efficient task management support for FPGA image processing. The technical solution of this embodiment has good scalability and can adapt to processing requirements of different scales and types, providing reliable technical support for FPGA application development. This optimization effect is particularly evident when processing complex task sequences.
[0068] In one embodiment of the FPGA parallel computing unit scheduling method for image processing of the present application, the following contents may also be specifically included: Step S601: Calculate the number of rows and columns of the two-dimensional cache array based on the image resolution and the data block size, allocate a fixed-size storage space to each cache unit, read the timestamp information to determine the write location of the data block, establish a cache address mapping table to record the storage location of the data block in the two-dimensional cache array, write the processed data block to the corresponding cache unit, and update the cache status table to record the storage status of the data block; Step S602: Read the processing time field in the timestamp to calculate the processing delay of the data block, adjust the buffer size according to the statistical value of the processing delay, configure the buffer threshold to control the read and write rate of the data block, and the field programmable gate array writes the processing status of the data block to the specified bit field of the status register, generates an interrupt request signal to notify the processing completion event, and updates the interrupt flag bit of the interrupt status register.
[0069] Optionally, this embodiment addresses the issues of low cache management efficiency, large processing delay fluctuations, and unreliable state synchronization in FPGA image processing by innovatively designing a set of adaptive cache management solutions. Regarding cache performance evaluation, this embodiment designs a cache efficiency scoring formula: Cache_Score = α × (Hit_Rate / Base_Rate) + β × (Space_Usage) + γ × (Access_Speed), where α, β, and γ are weight coefficients representing the hit rate ratio, space utilization, and access speed, respectively. A delay control scoring formula: Delay_Score = (Processing_Time × Block_Count) / (Buffer_Size + Sync_Cost) is also introduced to evaluate the effectiveness of processing delay control.
[0070] This embodiment first achieves efficient data management through a deeply optimized cache organization mechanism. During the design of the two-dimensional cache array, the system uses a dynamic programming algorithm to calculate the optimal row and column configuration based on the resolution characteristics of the input image and the size parameters of the data block. The cache array is organized in a hierarchical structure, including a data storage layer, an address mapping layer, and a state management layer. The size of each cache unit is optimized based on the data block size to ensure efficient use of storage space. To improve access efficiency, the system considers the principle of data locality when allocating cache units, allocating adjacent data blocks to continuous storage areas. This optimized storage strategy can significantly reduce memory access latency, especially when processing large-scale image data. Timestamp information is parsed using a fast table lookup method, determining the target location of the data block through a pre-established mapping relationship. The cache address mapping table is designed with a multi-level index structure to support fast location query and update. This optimized organization mechanism provides a reliable storage foundation for data management.
[0071] This embodiment innovatively implements a data writing mechanism. During the writing of processed data blocks, the system organizes the write operation in a pipeline manner. First, the write location is determined by the timestamp, and then the status of the target location is checked to ensure that no write conflict occurs. The write process uses atomic operations to avoid data confusion caused by concurrent access. The cache status table is updated in a bitmap manner, and each bit represents the usage status of the corresponding cache unit. The status information contains multiple fields, which not only record the basic occupancy status, but also the validity flag and access count of the data. This complete write mechanism significantly improves the reliability of data management.
[0072] This embodiment achieves efficient buffer management through an optimized delay control mechanism. During the calculation of the processing delay, the system uses a sliding window method to calculate the processing time of the data block. By analyzing the changing trend of the processing time, the system can predict the future processing load. The adjustment of the buffer size adopts an adaptive strategy to dynamically change the buffer capacity according to the statistical characteristics of the processing delay. When it is detected that the processing delay continues to increase, the system will appropriately increase the buffer size to provide more data cache space; when the processing delay is stable or reduced, the buffer size will be reduced accordingly to avoid resource waste. The setting of the buffer threshold takes into account the generation rate and consumption rate of the data block, and controls the flow speed of the data by dynamically adjusting the threshold. This optimized control mechanism significantly improves the adaptability of the system.
[0073] This embodiment establishes a complete state synchronization framework. The status register is designed using a bit field structure, with different bit fields recording processing status, error flags, and control information. Interrupt request signals are generated using edge triggering to ensure reliable signal transmission. Updates to the interrupt status register use atomic operations to avoid race conditions during status updates. By establishing a complete state management mechanism, the system achieves reliable monitoring of the entire data processing process. This optimized synchronization mechanism provides reliable state support for system collaboration.
[0074] This embodiment provides an innovative solution for cache management. Through optimized cache organization and latency control, the system optimizes the entire process, from data storage to state synchronization. This reliable management mechanism provides efficient data support for FPGA image processing. The technical solution of this embodiment has good scalability and can adapt to processing requirements of varying scales and types, providing reliable technical support for FPGA application development. This optimization effect is particularly evident when handling complex image processing tasks.
[0075] In one embodiment of the FPGA parallel computing unit scheduling method for image processing of the present application, the following contents may also be specifically included: Step S701: The state tracker reads the cache state table of the two-dimensional cache array, checks the continuity of the data blocks according to the data block numbers, establishes a data block state mapping table to record missing and duplicate data block information, writes the data block state mapping table to a specific bit field of the status register, updates the detection counter of the state tracker, and generates a data block state detection report to record the location information of the abnormal data blocks; Step S702: The host reads the status register to obtain the processing status of the data block, configures the read parameters of the direct memory access controller, reads the processed data block from the dynamic random access memory to the system memory, updates the read status flag of the data block, writes the read completion flag to the field programmable gate array, clears the corresponding interrupt status bit, and releases the storage space occupied by the read data block.
[0076] Optionally, this embodiment addresses issues such as inaccurate state tracking, unreliable data integrity verification, and untimely resource release in FPGA image processing by innovatively designing a state management solution based on intelligent tracking. Regarding state tracking evaluation, this embodiment devises a tracking efficiency scoring formula: Track_Score = α × (Detection_Speed / Base_Speed) + β × (Memory_Efficiency) + γ × (Recovery_Rate), where α, β, and γ are weight coefficients representing the detection speed ratio, memory utilization, and resource recovery rate, respectively. The integrity scoring formula: Integrity_Score = (Valid_Blocks × Processing_Speed) / (Missing_Blocks + Duplicate_Blocks) is also introduced to evaluate the effectiveness of data integrity verification.
[0077] This embodiment first achieves reliable data verification through a deeply optimized state tracking mechanism. When reading the cache state table, the state tracker adopts a multi-level caching strategy to cache frequently accessed state information in fast memory. Data block continuity checks use a sliding window method to determine whether there are missing or duplicated data by comparing the numbers of adjacent blocks. This checking mechanism not only considers the continuity of the numbers but also verifies the correctness of the processing sequence in conjunction with timestamp information. For detected anomalies, the system adopts a classification processing strategy: for missing data blocks, their expected location and adjacent block information are recorded to facilitate subsequent data recovery; for duplicate data blocks, which version of the data to retain is determined by comparing timestamps. The state mapping table is established using a multi-dimensional index structure to support fast queries based on different conditions. This optimized tracking mechanism significantly improves the accuracy and efficiency of state detection. Especially when processing large-scale data, through parallel detection and pipeline processing, the system can complete state verification in real time, avoiding detection becoming a performance bottleneck.
[0078] This embodiment innovatively implements a status recording mechanism. When writing the status mapping table to the status register, the system adopts a bit field partitioning strategy, and different types of status information are mapped to independent bit fields. This design not only improves storage efficiency, but also supports parallel status update operations. The update of the detection counter adopts atomic operations to ensure the accuracy of the count value. The generation of the status detection report adopts a hierarchical structure, which contains basic statistical information and detailed exception records. The report content not only records the location of the abnormal data block, but also includes possible cause analysis and recovery suggestions. This complete recording mechanism provides reliable data support for system maintenance and optimization. By establishing a status history database, the system supports long-term performance analysis and fault diagnosis, which helps to discover potential system problems.
[0079] This embodiment achieves efficient data transmission through an optimized readback mechanism. When reading the status register, the host adopts a batch read strategy to obtain multiple related states at one time. During the configuration of the DMA controller, the system dynamically adjusts the transmission parameters based on the data block characteristics and the current system load. For example, when the system load is light, the burst transmission length can be increased to improve data throughput; when the load is heavy, the transmission unit can be appropriately reduced to avoid affecting other operations. The data reading process adopts a prefetch mechanism to prepare the transmission parameters of the next data block in advance while the current data block is being transmitted. This optimized transmission mechanism significantly reduces data reading delays. The writing of the completion flag and the clearing of the interrupt status are synchronous operations to ensure the consistency of status updates.
[0080] This embodiment establishes a complete resource management framework. After the data block processing is completed, the system ensures the reliable recovery of resources through a multi-stage release process. First, the read status of the data block is updated and its releasable status is marked; then the relevant control information, including the interrupt flag and status bit, is cleared; finally, the actual storage space release operation is performed. This gradual release mechanism avoids the problem of resource leakage. By establishing a resource usage tracking table, the system monitors the allocation and release of storage space in real time and supports intelligent resource scheduling. This optimized management mechanism significantly improves the resource utilization efficiency of the system.
[0081] This embodiment provides an innovative solution for state management. Through optimized state tracking and resource management, the system optimizes the entire process, from data verification to resource release. This reliable management mechanism provides a stable operating environment for FPGA image processing. The technical solution of this embodiment has good scalability and can adapt to processing requirements of different scales and types, providing reliable technical support for FPGA application development. This optimization effect is particularly evident when handling complex image processing tasks.
[0082] In order to effectively address the deficiencies of traditional technologies in data security, access control, and privacy protection, and significantly improve the security and reliability of FPGA parallel computing, the present application provides an embodiment of an FPGA parallel computing unit scheduling device for image processing, which is used to implement all or part of the content of the FPGA parallel computing unit scheduling method for image processing. Figure 2 The FPGA parallel computing unit scheduling device for image processing specifically includes the following contents: An image processing module 10 is configured to establish a data transmission channel between a host and a field programmable gate array via a PCIe interface, allocate a physically contiguous memory area as a direct memory access buffer, create a descriptor queue in the shared memory to record source and destination address information of data transmissions, map the descriptor queue to user space, configure the burst transmission length and address alignment, calculate the optimal size of data blocks based on image resolution, and set an overlap area; a data writing module 20, configured to convert the image data into grayscale values and divide the data into a plurality of data blocks according to the optimal size, add boundary pixels to each data block and set a timestamp including the data block number and processing time, write the data block into the direct memory access buffer, trigger a doorbell register to send a task signal to a field programmable gate array, and cause the field programmable gate array to pre-fetch descriptors from the descriptor queue and write the data block into a dynamic random access memory; The data scheduling module 30 is used to construct a two-dimensional cache array and write the processed data blocks to the corresponding positions according to the timestamp, monitor the processing delay and dynamically adjust the buffer size, the field programmable gate array writes the data block status to the status register and triggers an interrupt signal, detects the absence and duplication of the data block through the status tracker, and the host reads the data block from the dynamic random access memory according to the content of the status register, and the host sends a data read completion signal to the field programmable gate array.
[0083] From the above description, it can be seen that the FPGA parallel computing unit scheduling device for image processing provided by the embodiment of the present application can realize data security isolation and access control through innovative design of data security transmission mechanism, physical continuous memory allocation and descriptor queue management. Build a data block encryption strategy based on timestamps, combine dynamic buffer adjustment and status monitoring, and establish a reliable data protection system. Introduce a processing delay monitoring mechanism to ensure the privacy of the data processing process through real-time performance evaluation and resource scheduling. Based on data encryption transmission, this method effectively solves the shortcomings of traditional technologies in data security, access control and privacy protection, and significantly improves the security and reliability of FPGA parallel computing.
[0084] From a hardware perspective, in order to effectively address the deficiencies of traditional technologies in data security, access control, and privacy protection, and significantly improve the security and reliability of FPGA parallel computing, this application provides an embodiment of an electronic device for implementing all or part of the content of the FPGA parallel computing unit scheduling method for image processing. The electronic device specifically includes the following content: A processor, memory, a communications interface, and a bus; wherein the processor, memory, and communications interface communicate with each other via the bus; the communications interface is used to transmit information between an FPGA parallel computing unit scheduling device for image processing and related devices such as core business systems, user terminals, and related databases; the logic controller can be a desktop computer, a tablet computer, a mobile terminal, etc., but this embodiment is not limited thereto. In this embodiment, the logic controller can be implemented with reference to the embodiments of the FPGA parallel computing unit scheduling method for image processing and the embodiments of the FPGA parallel computing unit scheduling device for image processing, the contents of which are incorporated herein and any repetitions are omitted.
[0085] It is understandable that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.
[0086] In practical applications, portions of the FPGA parallel computing unit scheduling method for image processing can be executed on the electronic device side as described above, or all operations can be performed on the client device. The specific selection can be based on the processing capabilities of the client device and the limitations of the user's usage scenario. This application does not impose any restrictions on this. If all operations are performed on the client device, the client device may also include a processor.
[0087] The aforementioned client device may include a communication module (i.e., a communication unit) capable of establishing a communication connection with a remote server to facilitate data transmission with the server. The server may include a server at the task scheduling center or, in other implementation scenarios, a server on an intermediate platform, such as a server on a third-party server platform that is communicatively linked to the task scheduling center server. The server may comprise a single computer device, a server cluster consisting of multiple servers, or a distributed server configuration.
[0088] Figure 3 Schematic block diagram of the system structure of the electronic device 9600 according to an embodiment of the present application. Figure 3 As shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It is worth noting that the Figure 3 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.
[0089] In one embodiment, the FPGA parallel computing unit scheduling method for image processing can be integrated into the central processing unit 9100. The central processing unit 9100 can be configured to perform the following control: Step S101: establishing a data transmission channel between a host and a field programmable gate array via a PCIe interface, allocating a physically continuous memory area as a direct memory access buffer, creating a descriptor queue in the shared memory to record the source and destination address information of the data transmission, mapping the descriptor queue to the user space, configuring the burst transmission length and address alignment, calculating the optimal size of the data block based on the image resolution, and setting the overlap area; Step S102: converting the image data into grayscale values and dividing the data into a plurality of data blocks according to the optimal size, adding boundary pixels to each data block and setting a timestamp including the data block number and processing time, writing the data block into the direct memory access buffer, triggering a doorbell register to send a task signal to a field programmable gate array, and causing the field programmable gate array to pre-fetch descriptors from the descriptor queue and write the data block into a dynamic random access memory; Step S103: construct a two-dimensional cache array and write the processed data block to the corresponding position according to the timestamp, monitor the processing delay and dynamically adjust the buffer size, the field programmable gate array writes the data block status into the status register and triggers an interrupt signal, detects the missing and duplication of the data block through the status tracker, the host reads the data block from the dynamic random access memory according to the content of the status register, and the host sends a data reading completion signal to the field programmable gate array.
[0090] As can be seen from the above description, the electronic device provided in the embodiment of the present application realizes data security isolation and access control through innovatively designed data security transmission mechanism, physical continuous memory allocation and descriptor queue management. A timestamp-based data block encryption strategy is constructed, combined with dynamic buffer adjustment and status monitoring to establish a reliable data protection system. A processing delay monitoring mechanism is introduced to ensure the privacy of the data processing process through real-time performance evaluation and resource scheduling. Based on data encryption transmission, this method effectively solves the shortcomings of traditional technologies in data security, access control and privacy protection, and significantly improves the security and reliability of FPGA parallel computing.
[0091] In another embodiment, the FPGA parallel computing unit scheduling device for image processing can be configured separately from the central processing unit 9100. For example, the FPGA parallel computing unit scheduling device for image processing can be configured as a chip connected to the central processing unit 9100, and the function of the FPGA parallel computing unit scheduling method for image processing can be implemented through the control of the central processing unit.
[0092] like Figure 3 As shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily have to include Figure 3 In addition, the electronic device 9600 may also include all components shown in Figure 3 For components not shown, reference may be made to the prior art.
[0093] like Figure 3As shown, the central processing unit 9100 is sometimes also referred to as a controller or operation control, and may include a microprocessor or other processor device and / or logic device. The central processing unit 9100 receives input and controls the operation of various components of the electronic device 9600.
[0094] Memory 9140 can be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It can store the aforementioned failure-related information and also store programs that execute the relevant information. The CPU 9100 can execute the programs stored in memory 9140 to implement information storage or processing.
[0095] The input unit 9120 provides input to the central processing unit 9100. The input unit 9120 may be, for example, a keypad or touch input device. The power supply 9170 is used to provide power to the electronic device 9600. The display 9160 is used to display objects such as images and text. The display may be, for example, an LCD display, but is not limited thereto.
[0096] The memory 9140 may be a solid-state memory, such as a read-only memory (ROM), random access memory (RAM), or SIM card. Alternatively, it may be a memory that retains information even when power is off, can be selectively erased, and is capable of storing additional data. Examples of such memory are sometimes referred to as EPROMs. The memory 9140 may also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142 for storing application programs and function programs, or processes used by the central processing unit 9100 to execute operations of the electronic device 9600.
[0097] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, images, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various driver programs for communication functions of the electronic device and / or for executing other functions of the electronic device (such as messaging applications, address book applications, etc.).
[0098] The communication module 9110 is a transmitter / receiver that transmits and receives signals via the antenna 9111. The communication module 9110 (transmitter / receiver) is coupled to the central processor 9100 to provide input signals and receive output signals, which may be the same as the case of a conventional mobile communication terminal.
[0099] Based on different communication technologies, multiple communication modules 9110 may be provided in the same electronic device, such as cellular network modules, Bluetooth modules, and / or wireless local area network modules. The communication module 9110 (transmitter / receiver) is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130, providing audio output via the speaker 9131 and receiving audio input from the microphone 9132, thereby implementing common telecommunication functions. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. Furthermore, the audio processor 9130 is coupled to the central processing unit 9100, enabling local recording via the microphone 9132 and playback of stored audio via the speaker 9131.
[0100] Embodiments of the present application also provide a computer-readable storage medium capable of implementing all steps of the FPGA parallel computing unit scheduling method for image processing, where the execution subject is a server or a client, in the above-mentioned embodiment. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, all steps of the FPGA parallel computing unit scheduling method for image processing, where the execution subject is a server or a client, in the above-mentioned embodiment are implemented. For example, when the processor executes the computer program, the following steps are implemented: Step S101: establishing a data transmission channel between a host and a field programmable gate array via a PCIe interface, allocating a physically continuous memory area as a direct memory access buffer, creating a descriptor queue in the shared memory to record the source and destination address information of the data transmission, mapping the descriptor queue to the user space, configuring the burst transmission length and address alignment, calculating the optimal size of the data block based on the image resolution, and setting the overlap area; Step S102: converting the image data into grayscale values and dividing the data into a plurality of data blocks according to the optimal size, adding boundary pixels to each data block and setting a timestamp including the data block number and processing time, writing the data block into the direct memory access buffer, triggering a doorbell register to send a task signal to a field programmable gate array, and causing the field programmable gate array to pre-fetch descriptors from the descriptor queue and write the data block into a dynamic random access memory; Step S103: construct a two-dimensional cache array and write the processed data block to the corresponding position according to the timestamp, monitor the processing delay and dynamically adjust the buffer size, the field programmable gate array writes the data block status into the status register and triggers an interrupt signal, detects the missing and duplication of the data block through the status tracker, the host reads the data block from the dynamic random access memory according to the content of the status register, and the host sends a data reading completion signal to the field programmable gate array.
[0101] As can be seen from the above description, the computer-readable storage medium provided in the embodiment of the present application realizes data security isolation and access control through innovatively designed data security transmission mechanism, physical continuous memory allocation and descriptor queue management. A timestamp-based data block encryption strategy is constructed, combined with dynamic buffer adjustment and status monitoring to establish a reliable data protection system. A processing delay monitoring mechanism is introduced to ensure the privacy of the data processing process through real-time performance evaluation and resource scheduling. Based on data encryption transmission, this method effectively solves the shortcomings of traditional technologies in data security, access control and privacy protection, and significantly improves the security and reliability of FPGA parallel computing.
[0102] Embodiments of the present application also provide a computer program product capable of implementing all steps of the FPGA parallel computing unit scheduling method for image processing in the above-mentioned embodiment, where the execution subject is a server or a client. When the computer program / instruction is executed by a processor, the steps of the FPGA parallel computing unit scheduling method for image processing are implemented. For example, the computer program / instruction implements the following steps: Step S101: establishing a data transmission channel between a host and a field programmable gate array via a PCIe interface, allocating a physically continuous memory area as a direct memory access buffer, creating a descriptor queue in the shared memory to record the source and destination address information of the data transmission, mapping the descriptor queue to the user space, configuring the burst transmission length and address alignment, calculating the optimal size of the data block based on the image resolution, and setting the overlap area; Step S102: converting the image data into grayscale values and dividing the data into a plurality of data blocks according to the optimal size, adding boundary pixels to each data block and setting a timestamp including the data block number and processing time, writing the data block into the direct memory access buffer, triggering a doorbell register to send a task signal to a field programmable gate array, and causing the field programmable gate array to pre-fetch descriptors from the descriptor queue and write the data block into a dynamic random access memory; Step S103: construct a two-dimensional cache array and write the processed data block to the corresponding position according to the timestamp, monitor the processing delay and dynamically adjust the buffer size, the field programmable gate array writes the data block status into the status register and triggers an interrupt signal, detects the missing and duplication of the data block through the status tracker, the host reads the data block from the dynamic random access memory according to the content of the status register, and the host sends a data reading completion signal to the field programmable gate array.
[0103] As can be seen from the above description, the computer program product provided in the embodiment of the present application realizes data security isolation and access control through innovatively designed data security transmission mechanism, physical continuous memory allocation and descriptor queue management. A data block encryption strategy based on timestamp is constructed, combined with dynamic buffer adjustment and status monitoring to establish a reliable data protection system. A processing delay monitoring mechanism is introduced to ensure the privacy of the data processing process through real-time performance evaluation and resource scheduling. Based on data encryption transmission, this method effectively solves the shortcomings of traditional technologies in data security, access control and privacy protection, and significantly improves the security and reliability of FPGA parallel computing.
[0104] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatuses, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0105] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0106] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0107] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0108] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.
Claims
1. A method for scheduling FPGA parallel computing units for image processing, characterized in that: The method comprises: A data transmission channel is established between a host and a field programmable gate array via a PCIe interface, a physically continuous memory area is allocated as a direct memory access buffer, a descriptor queue is created in the shared memory to record source and destination address information of the data transmission, the descriptor queue is mapped to the user space, the burst transmission length and address alignment are configured, the optimal size of the data block is calculated according to the image resolution, and an overlapping area is set; converting the image data into grayscale values and dividing the data into a plurality of data blocks according to the optimal size, adding boundary pixels to each of the data blocks and setting a timestamp including the data block number and processing time, writing the data blocks into the direct memory access buffer, triggering a doorbell register to send a task signal to a field programmable gate array, and causing the field programmable gate array to pre-fetch descriptors from the descriptor queue and write the data blocks into a dynamic random access memory; A two-dimensional cache array is constructed and the processed data blocks are written to corresponding locations according to the timestamp, the processing delay is monitored and the buffer size is dynamically adjusted, the field programmable gate array writes the data block status into the status register and triggers an interrupt signal, the missing and duplicate conditions of the data block are detected by the status tracker, the host reads the data block from the dynamic random access memory according to the content of the status register, and the host sends a data read completion signal to the field programmable gate array.
2. The FPGA parallel computing unit scheduling method for image processing according to claim 1, characterized in that: The method includes establishing a data transmission channel between a host and a field programmable gate array through a PCIe interface, allocating a physical continuous memory area as a direct memory access buffer, and creating a descriptor queue in a shared memory to record source and destination address information of data transmission, including: The host side loads a PCIe device driver and enables an interrupt response mechanism, allocates a physically continuous memory area through a kernel module and registers it as a direct memory access buffer, maps the direct memory access buffer to a user space address, configures the PCIe link to the third generation specification and sets the link width, and starts a bus master to connect the PCIe interface to an Advanced Extensible Interface bus; The field programmable gate array loads the PCIe intellectual property core configuration parameters, configures the Advanced Extensible Interface Memory Mapping Protocol and the Advanced Extensible Interface Reduced Interface Protocol, creates a descriptor queue in the shared memory to record the source and destination address information of the data transmission, sets the size of the direct memory access receive buffer and the transmit buffer according to system requirements, configures the burst transmission length and address alignment, and establishes an interrupt vector table to bind the task identifier to the interrupt handling function.
3. The FPGA parallel computing unit scheduling method for image processing according to claim 1, characterized in that: Mapping the descriptor queue to the user space, configuring the burst transfer length and address alignment, calculating the optimal size of the data block according to the image resolution and setting the overlap area, includes: Mapping the physical address space of the descriptor queue to the process virtual address space through a memory mapping function, configuring the base address register space to implement access to the control register and the status register, setting the direct memory access burst transfer length to an integer multiple of the basic transfer unit, configuring the address boundary alignment mode to ensure data access efficiency, and writing the read and write pointer update information of the descriptor queue into the doorbell register; The optimal size of a single data block is calculated based on the width and height of the input image. The data block size is set to an integer multiple of the computing power of the processing unit. An overlapping area mapping relationship is established between adjacent data blocks. The pixel width of the overlapping area is calculated to ensure the continuity of the boundary pixels. The data block size and overlapping area parameters are written into the configuration register, and a data block index table is generated to record the association information between blocks.
4. The FPGA parallel computing unit scheduling method for image processing according to claim 1, characterized in that: The step of converting the image data into grayscale values and dividing the data into a plurality of data blocks according to the optimal size, adding boundary pixels to each data block and setting a timestamp including a data block number and a processing time, and writing the data block into the direct memory access buffer comprises: Converting the RGB color space of the input image into a grayscale image, dividing the grayscale image data into a plurality of data blocks based on the optimal size, adding boundary pixels around each data block to form an extended data block, calculating the position information of the extended data block in the global coordinate system, generating a data block descriptor including a start address and an end address, and assigning a globally unique data block number to each data block; Create a timestamp management unit to record the processing status of the data block, the timestamp includes two time attribute fields, the data block number and the data block processing time, establish a timestamp mapping table to store the timing information of the data block, write the data block and its corresponding timestamp information into the specified location of the direct memory access buffer, and update the data block status register to record the transmission completion flag.
5. The FPGA parallel computing unit scheduling method for image processing according to claim 1, characterized in that: The triggering doorbell register sends a task signal to a field programmable gate array, and the field programmable gate array pre-fetches descriptors from the descriptor queue and writes the data block into a dynamic random access memory, comprising: Write the task information to the specified offset address of the base address register space, trigger the doorbell register to generate an interrupt request signal, configure the interrupt mask bit of the interrupt status register, activate the interrupt controller to enable interrupt response, read the interrupt vector table to obtain the entry address of the interrupt service routine, send a task start signal to the field programmable gate array, and update the task status register to record the task execution status; The field programmable gate array reads descriptor information from the descriptor queue to a prefetch buffer, parses the descriptor to obtain a source address and a destination address of a data block, configures transmission parameters of a direct memory access controller, establishes a burst transmission channel to write the data block into a designated storage unit of a dynamic random access memory, updates a read pointer of the descriptor queue, and writes a data block transmission completion flag into a status register.
6. The FPGA parallel computing unit scheduling method for image processing according to claim 1, characterized in that: The method includes: constructing a two-dimensional cache array and writing the processed data block to the corresponding position according to the timestamp, monitoring the processing delay and dynamically adjusting the buffer size, and the field programmable gate array writing the data block status to the status register and triggering an interrupt signal, including: Calculating the number of rows and columns of the two-dimensional cache array according to the image resolution and the size of the data block, allocating a fixed-size storage space to each cache unit, reading the timestamp information to determine the write location of the data block, establishing a cache address mapping table to record the storage location of the data block in the two-dimensional cache array, writing the processed data block to the corresponding cache unit, and updating the cache status table to record the storage status of the data block; The processing time field in the timestamp is read to calculate the processing delay of the data block, the buffer size is adjusted according to the statistical value of the processing delay, the buffer threshold is configured to control the read and write rate of the data block, the field programmable gate array writes the processing status of the data block into a specified bit field of the status register, generates an interrupt request signal to notify the processing completion event, and updates the interrupt flag bit of the interrupt status register.
7. The FPGA parallel computing unit scheduling method for image processing according to claim 1, characterized in that: The process of detecting the absence and duplication of the data block by a status tracker, the host reading the data block from the dynamic random access memory according to the content of the status register, and the host sending a data read completion signal to a field programmable gate array comprises: The state tracker reads the cache state table of the two-dimensional cache array, checks the continuity of the data blocks according to the data block numbers, establishes a data block state mapping table to record missing and duplicate data block information, writes the data block state mapping table into a specific bit field of the status register, updates the detection counter of the state tracker, and generates a data block state detection report to record the location information of the abnormal data blocks; The host reads the status register to obtain the processing status of the data block, configures the read parameters of the direct memory access controller, reads the processed data block from the dynamic random access memory to the system memory, updates the read status flag of the data block, writes the read completion flag to the field programmable gate array, clears the corresponding interrupt status bit, and releases the storage space occupied by the read data block.
8. An FPGA parallel computing unit scheduling device for image processing, characterized in that: The device comprises: An image processing module is configured to establish a data transmission channel between a host and a field programmable gate array via a PCIe interface, allocate a physically contiguous memory area as a direct memory access buffer, create a descriptor queue in the shared memory to record source and destination address information of data transmission, map the descriptor queue to user space, configure the burst transmission length and address alignment, calculate the optimal size of the data block based on the image resolution, and set the overlap area; a data writing module, configured to convert the image data into grayscale values and divide the data into a plurality of data blocks according to the optimal size, add boundary pixels to each data block and set a timestamp including the data block number and processing time, write the data block into the direct memory access buffer, trigger a doorbell register to send a task signal to a field programmable gate array, and cause the field programmable gate array to prefetch descriptors from the descriptor queue and write the data block into a dynamic random access memory; A data scheduling module is used to construct a two-dimensional cache array and write the processed data blocks to the corresponding positions according to the timestamp, monitor the processing delay and dynamically adjust the buffer size, the field programmable gate array writes the data block status into the status register and triggers an interrupt signal, detects the absence and duplication of the data block through the status tracker, the host reads the data block from the dynamic random access memory according to the content of the status register, and the host sends a data read completion signal to the field programmable gate array.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the FPGA parallel computing unit scheduling method for image processing according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the FPGA parallel computing unit scheduling method for image processing according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
High-resolution image detection and storage method and system
CN117934261A
Image splicing method and system based on FPGA (Field Programmable Gate Array), medium and equipment
CN119741198A
Parallel computing framework based on software definition and implementation method thereof
CN120104306A
High Speed, Parallel Configuration of Multiple Field Programmable Gate Arrays
US20150143003A1