A hardware-software integrated parallel reconfigurable real-time image processing system
Through the integrated parallel reconstructible image real-time processing system of software and hardware, the problems of resource competition and dynamic algorithm loading in multi-core DSP systems are solved, and efficient image processing and system performance improvement are achieved.
Patent Information
- Application Number
- CN202510413757.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing image real-time processing system developed based on DSP lacks hardware-level isolation when resource competition in parallel with multi-core tasks, resulting in performance fluctuations, and it is difficult to dynamically load or replace algorithm components at runtime, and cannot effectively utilize software and hardware computing resources.
The real-time image processing system of integrated software and hardware is adopted to reconstruct the DSP node unit architecture through the main control unit, and the routing data exchange is realized using SRIO high-speed interconnection unit. Combining the task allocation of the master core and slave core processing units, a serial parallel structure is formed to realize efficient resource scheduling and dynamic algorithm reconstruction.
It improves the communication efficiency between resources, enhances the flexibility of the system and the software operation capabilities of the hardware platform, adapts to a variety of image processing applications, and improves the efficiency and reliability of image processing.
Smart Images

Figure CN119917295B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of embedded real-time image processing, and particularly relates to a software and hardware integrated parallel reconfigurable real-time image processing system. Background Art
[0002] Image real-time processing systems mainly utilize parallel processing resources such as DSPs to complete image processing tasks such as enhancement, detection, and recognition, and the processing efficiency is equivalent to the frame rate of the acquired images. Existing image real-time processing systems developed based on DSPs often rely on software development tools provided by hardware manufacturers, use static compilation and linking, have a high degree of coupling of multi-core tasks, and cannot dynamically load or replace algorithm components during operation. When multiple tasks are parallel, there is a lack of hardware-level isolation for resource competition, such as cache pollution and memory bandwidth preemption, resulting in performance fluctuations.
[0003] Patent application with publication number CN109710399A: A DSP communication task scheduling system and method. This method classifies the communication tasks of each sub-module of the DSP communication task scheduling system according to the functions of the communication tasks, sets the priorities of each type of communication task according to the system function requirements, sets the operation state machine of the task scheduling system, and schedules tasks according to the state of the state machine, message permissions, and classification. This solution focuses on the collaborative work between sub-modules, does not consider the comprehensive scheduling of multiple DSPs and multi-cores, and the integrated optimization of software algorithms on this hardware platform, and it is difficult to synergistically utilize the advantages of the software and hardware computing resources of the image real-time processing system. Summary of the Invention
[0004] To solve the deficiencies in the prior art, the purpose of the present invention is to provide a software and hardware integrated parallel reconfigurable real-time image processing system, which overcomes the problems of high difficulty in process reconstruction, difficult efficient allocation of computing resources, and difficult effective utilization of DSP multi-core resources under traditional multi-core DSP hardware platforms.
[0005] The above technical objective of the present invention is achieved through the following technical solutions:
[0006] A software and hardware integrated parallel reconfigurable real-time image processing system includes a system input module, a resource scheduling module, and a multi-core task allocation module;
[0007] The system input module is used to receive the image to be processed and a preset process configuration file and send them to the resource scheduling module;
[0008] The resource scheduling module includes a main control unit, a DSP node unit, and an SRIO high-speed interconnection unit:
[0009] The main control unit reconstructs the architectures of each DSP node unit according to the process configuration file to form the processing processes of each DSP node unit, and at the same time configures the image processing algorithms in each DSP node unit based on the process configuration file; the SRIO high-speed interconnection unit is a routed data exchange channel, which realizes the exchange of input and output data when each DSP node unit uses the image processing algorithm to process the image to be processed.
[0010] The multi-core task allocation module is deployed in each DSP node unit and includes a main core processing unit and a slave core processing unit:
[0011] The main core processing unit is responsible for the main process processing of the image processing algorithm in the DSP node unit, uses the main core of the DSP node unit to allocate the image processing tasks and forms a task queue, and then outputs it to the slave core processing unit; the slave core processing unit uses the slave core of the DSP node unit to execute each image processing task in the task queue, and after all the image processing tasks are completed, it summarizes and outputs them to the main core processing unit, and the main core processing unit outputs the image processing results through the SRIO high-speed interconnection unit.
[0012] Further, the configuration information in the process configuration file stores the serial-parallel structures of each DSP node unit and the parameters of the image processing algorithms deployed in each DSP node unit in a depth-first traversal manner; among them, the serial-parallel structure includes a serial architecture, a parallel architecture, or a serial-parallel combined architecture.
[0013] Further, the image to be processed is sent to the SRIO high-speed interconnection unit, and then, according to the serial-parallel structure of the DSP node unit set by the process configuration file, the image is re-forwarded to the DSP node unit based on routing.
[0014] Further, the DSP node unit is a computing node, and serves as different process nodes according to its processing process and executes the corresponding image processing algorithms.
[0015] Further, the main control unit extracts the image processing algorithms in each DSP node unit, obtains the parameters of the image processing algorithms and the serial-parallel structures of each DSP node unit through the process configuration file, and reconstructs the architectures of the DSP node units based on the serial-parallel structures.
[0016] After receiving the data packet of the image processing algorithm configured with parameters, the DSP node unit performs verification. After the verification is passed, it first erases the original data in the FLASH, then writes the image processing algorithm into the FLASH. After that, the DSP node unit runs according to the written image processing algorithm and the serial-parallel order of each DSP node unit in the serial-parallel structure of the DSP node unit.
[0017] Further, in the SRIO high-speed interconnection unit, a routing and switching circuit method is adopted. Each DSP node unit is modularly designed as an end node unit, and each DSP node unit is interconnected through a high-speed interface; during the execution of the image processing task, data and results of each DSP node unit are mutually transmitted between each DSP node unit through the routing and switching circuit.
[0018] Further, the processing result of the last DSP node unit in the serial-parallel structure of the DSP node unit is output through the SRIO high-speed interconnection unit as the final image processing result.
[0019] Further, the call relationship and process of the main core processing unit and the slave core processing unit are as follows:
[0020] When the main core processing unit receives the image task execution instruction issued by the user, it enters the system state and starts to execute the main process of the image processing algorithm. At this time, the slave core processing unit is in a waiting state;
[0021] When the main process runs to the stage that requires parallel processing, the main core processing unit divides the image to be processed into blocks, forms different image processing tasks based on the blocks and adds them to the task queue, and stores the resources required for the image processing tasks in the handle; the slave core processing unit wakes up the slave core of the corresponding DSP node unit according to the handle to perform parallel processing of the image processing tasks; whenever there is an idle slave core, a new image processing task is taken out from the task queue for processing; when the slave core finishes processing all the image processing tasks, the main core processing unit switches from the system state to the user state to obtain control, then performs the summary processing of the data obtained from the image processing tasks, and sends the processing result to the SRIO high-speed interconnection unit.
[0022] A large field of view target segmentation and counting device, which uses the above-mentioned software and hardware integrated parallel reconfigurable image real-time processing system to perform spot counting on the large field of view starry sky image.
[0023] Compared with the prior art, the present invention has the following technical characteristics:
[0024] 1. The present invention provides a resource-efficient scheduling method based on routing. By adopting the routing and switching circuit method, each DSP node unit is designed as an independent and decoupled computing module. Image data, parameters, and processing results can all be sent to the DSP node unit at a specific address through the routing and switching circuit, effectively improving the communication efficiency between various resources.
[0025] 2. The present invention can form any serial architecture, parallel architecture or serial-parallel combined architecture of DSP node units according to actual application requirements through a configuration file and a main control unit; if a certain DSP node unit fails, only the process configuration file needs to be modified to reconstruct the processing relationship of the DSP node unit, avoid the faulty DSP node unit, and the system function can be restored.
[0026] 3. By updating the image processing algorithm application program data and parameters in the main control unit, the present invention can adapt to a variety of image processing applications and different forms of image input. Description of the Drawings
[0027] Figure 1 It is the system block diagram in the embodiment of the present invention. Detailed Embodiment
[0028] The present invention provides a software and hardware integrated parallel reconfigurable image real-time processing system, including a system input module, a resource scheduling module, and a multi-core task allocation module; the resource scheduling module includes a main control unit, a DSP node unit, and an SRIO high-speed interconnection unit; the multi-core task allocation module is deployed in each DSP node unit and includes a main core processing unit and a slave core processing unit, where:
[0029] 1. System input module.
[0030] The system input module is used to receive the image to be processed and a preset process configuration file and send them to the resource scheduling module for the allocation and processing of image processing tasks; the process configuration file presets the serial-parallel structure of the DSP node unit and the parameters of the image processing algorithm.
[0031] Among them, the configuration information in the process configuration file stores the serial-parallel structure of each DSP node unit and the parameters of the image processing algorithm deployed in each DSP node unit in a depth-first traversal manner; among them, the serial-parallel structure is set according to actual application requirements, including a serial architecture, a parallel architecture or a serial-parallel combined architecture; in the parallel architecture, the image can be processed in blocks, and the computing power of each DSP node unit is M. In the framework of independent operation of each DSP node unit, the computing power of N DSP node units running simultaneously can reach M*N, for the superposition of computing power. In the serial architecture, multiple DSP node units can jointly complete a set of image processing algorithm data streams. Through pipeline function division, each DSP node unit can be modularized, facilitated and specialized, improving the computing efficiency of each DSP, simplifying the operation of a single DSP, and improving the reliability of the design.
[0032] Among them, the image to be processed is sent to the SRIO high-speed interconnection unit of the resource scheduling module, and then, according to the serial-parallel structure of the DSP node units set in the process configuration file, the image is re-forwarded to the DSP node units based on routing.
[0033] 2. Resource scheduling module.
[0034] (2.1) Main control unit and DSP node units.
[0035] After receiving the process configuration file, the main control unit in the resource scheduling module reconstructs the serial architecture, parallel architecture, or serial-parallel combined architecture of each DSP node unit according to the configuration information in the process configuration file to form the processing flow of each DSP node unit; at the same time, configures the image processing algorithms in each DSP node unit according to the parameters in the configuration information; the DSP node unit is a computing node, and serves as different process nodes according to its processing flow and executes the corresponding image processing algorithms.
[0036] The main control unit extracts the image processing algorithms in each DSP node unit, and uses the configuration information in the process configuration file to obtain the parameters corresponding to the image processing algorithms and the serial-parallel structure of the DSP node units, and updates the DSP node units through the SPI / UART low-speed bus; after receiving the data packet of the image processing algorithm configured with parameters, the DSP node unit performs verification. After the verification passes, the original data in the FLASH is first erased, and then the image processing algorithm is written into the FLASH. After that, the DSP node unit runs according to the written image processing algorithm and the serial-parallel order of each DSP node unit in the serial-parallel structure of the DSP node unit.
[0037] (2.2) SRIO high-speed interconnection unit.
[0038] The SRIO high-speed interconnection unit in the resource scheduling module is a routing-based data exchange channel, which realizes the high-speed exchange of input and output data when each DSP node unit processes the image to be processed using the image processing algorithm.
[0039] The specific settings of the SRIO high-speed interconnection unit are as follows:
[0040] In the SRIO high-speed interconnection unit, the routing exchange circuit method is adopted, and each DSP node unit is modularly designed as an end node unit. Each DSP node unit is interconnected through a 4x SRIO Gen2 high-speed interface, and the single-link bandwidth is 5 Gbps / channel; the data and results of each DSP node unit during the execution of the image processing task are transmitted between each DSP node unit through the routing exchange circuit.
[0041] 3. Multi-core task allocation module.
[0042] The multi-core task allocation module is deployed in each DSP node unit, including a main core processing unit and slave core processing units, where:
[0043] The main core processing unit is responsible for the main process of the image processing algorithm in the DSP node unit. It uses the main core of the DSP node unit to allocate and recycle and integrate the image processing tasks to form a task queue, and outputs it to the slave core processing unit through the SRIO bus.
[0044] The slave core processing unit receives the task queue output by the main core processing unit, uses the slave cores of the DSP node unit to execute each image processing task in the task queue, and summarizes and outputs it to the main core processing unit of the DSP node unit after the image processing task is completed. The main core processing unit outputs the image processing result through the SRIO high-speed interconnection unit; among them, the processing result of the last DSP node unit in the serial-parallel structure of the DSP node unit is output through the SRIO high-speed interconnection unit as the final image processing result.
[0045] Among them, the call relationship and process of the main core processing unit and the slave core processing units are as follows:
[0046] When the main core processing unit receives the image task execution instruction issued by the user, it enters the system state and starts to execute the main process of the image processing algorithm. At this time, the slave core processing units are in a waiting state; when the main process runs to the stage that needs parallel processing, the main core processing unit divides the image to be processed into blocks, forms different image processing tasks based on the blocks and adds them to the task queue, and stores the resources required for the image processing tasks in the handle; the slave core processing units wake up the slave cores of the corresponding DSP node units according to the handle to perform parallel processing of the image processing tasks; whenever there is an idle slave core, it takes out a new image processing task from the task queue for processing; when the slave cores complete all the image processing tasks, the main core processing unit switches from the system state to the user state to obtain control, and then performs the summary processing of the data obtained from the image processing tasks, and sends the processing result to the SRIO high-speed interconnection unit.
[0047] In an embodiment of the present invention, a processing system is used to implement the image processing function of star image spot counting, specifically as follows:
[0048] Each image processing algorithm is written in advance in the main control unit, including image segmentation threshold calculation application, spot binary segmentation application (including main core and slave cores), spot connected region calculation and counting application; the parameters of each image processing algorithm and the set serial-parallel structure of the DSP node unit are written in advance in the configuration information of the process configuration file; after receiving the process configuration file, the resource scheduling module deploys the corresponding image processing algorithm to DSP node units DSP1 to DSP4; the image processing algorithms of the corresponding DSP node units are:
[0049] DSP1: Image segmentation threshold calculation application;
[0050] DSP2, DSP3: Spot binary segmentation application, including main core control code and slave core calculation code;
[0051] DSP4: Spot connected region calculation and counting application.
[0052] Among them, the serial-parallel structure is set as: DSP2 and DSP3 are in parallel relationship, and they are in serial relationship with DSP1 and DSP4.
[0053] The image to be processed is a 16-bit grayscale image with a size of 4608×4608, which is sent to DSP1 through the SRIO high-speed interconnection unit; after receiving the image, DSP1 traverses the image pixels, statistically calculates the image histogram, calculates the image segmentation threshold through the OTSU algorithm, and sends the upper half (size 2304×4608) of the image to DSP2 through the SRIO high-speed interconnection unit, and sends the lower half (size 2304×4608) of the image data to DSP3.
[0054] The main core processing unit of DSP2 divides the image with a size of 2304×4608 into data with a size of 144×144, a total of 512 blocks; forms image processing tasks for all image blocks and adds them to the task queue, traverses the slave cores, and if there are idle slave cores, takes out the image processing tasks from the task queue for allocation and processing.
[0055] After receiving the image processing task, the slave core processing unit of DSP2 segments the 144×144 image block according to the image segmentation threshold, sets the points greater than the threshold to 1, and the points less than the threshold to 0, and returns the processed image block result to the main core processing unit after completion.
[0056] After receiving the results of all image processing tasks being processed, the main core processing unit in DSP2 aggregates them to form a binary image with a size of 2304×4608. The non-zero points in the image represent the target, and the rest represent the background; the binary image result is sent to DSP4 through the SRIO high-speed interconnection channel.
[0057] The processing process of DSP3 is the same as that of DSP2.
[0058] After receiving the processing results of DSP2 and DSP3, DSP4 first stitches the results into a whole image, performs connected region labeling through the region growing method, and the number of regions is the total number of spots. Finally, the spot counting result can also be output through the SRIO high-speed interconnection unit.
[0059] In summary, the resource-efficient scheduling method based on routing in the present invention adopts a routing and switching circuit method, designs each DSP node unit as an independent and decoupled computing module, and image data, parameters, and processing results can all be sent to the DSP node unit at a specific address through the routing and switching circuit, effectively improving the communication efficiency between various resources; the DSP node units are combined into any serial architecture, parallel architecture, or serial-parallel combination architecture according to actual application requirements through a process configuration file and a main control unit; by updating the image processing algorithms and parameters in the main control unit, it can adapt to various image processing applications and different forms of image input.
[0060] This system design scheme improves the flexibility of software operation on the hardware platform and the overall optimization of the performance of software and hardware, and improves the operation efficiency of the software and hardware as a whole. Furthermore, it achieves performance improvement in real-time image processing scenarios such as satellites and drones; when applying the method of the present invention to the task of large field-of-view target segmentation and counting with a size of 4608×4608, the average time consumption is improved from 4400 ms to 620 ms, and the system operation efficiency is improved.
[0061] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A software and hardware integrated parallel reconfigurable real-time image processing system, characterized in that It includes a system input module, a resource scheduling module, and a multi-core task allocation module; The system input module is used to receive the image to be processed and the preset process configuration file and send them to the resource scheduling module; The resource scheduling module includes a main control unit, a DSP node unit, and an SRIO high-speed interconnection unit: The main control unit reconstructs the architectures of the DSP node units according to the process configuration file to form the processing processes of the DSP node units, and at the same time configures the image processing algorithms in the DSP node units based on the process configuration file; The SRIO high-speed interconnection unit is a routing-based data exchange channel, which realizes the exchange of input and output data when the DSP node units use the image processing algorithms to process the image to be processed; The multi-core task allocation module is deployed in each DSP node unit and includes a main core processing unit and a slave core processing unit: The main core processing unit is responsible for the main process processing of the image processing algorithm in the DSP node unit, allocates the image processing tasks using the main core of the DSP node unit and forms a task queue and then outputs it to the slave core processing unit; The slave core processing unit uses the slave core of the DSP node unit to execute each image processing task in the task queue, and after all the image processing tasks are executed, it summarizes and outputs them to the main core processing unit, and the main core processing unit outputs the image processing result through the SRIO high-speed interconnection unit; The configuration information in the process configuration file stores the serial-parallel structures of the DSP node units and the parameters of the image processing algorithms deployed in the DSP node units in a depth-first traversal manner; Among them, the serial-parallel structure includes a serial architecture, a parallel architecture, or a serial-parallel combined architecture; The image to be processed is sent to the SRIO high-speed interconnection unit, and then, according to the serial-parallel structure of the DSP node unit set by the process configuration file, the image is re-forwarded to the DSP node unit based on routing; The calling relationship and process of the main core processing unit and the slave core processing unit are as follows: When the main core processing unit receives the image task execution instruction issued by the user, it enters the system state and starts to execute the main process of the image processing algorithm. At this time, the slave core processing unit is in a waiting state; When the main process runs to the stage that needs parallel processing, the main core processing unit divides the image to be processed into blocks, forms different image processing tasks based on the blocks and adds them to the task queue, and stores the resources required for the image processing tasks in the handle; The slave core processing unit wakes up the slave core of the corresponding DSP node unit according to the handle to perform parallel processing of the image processing tasks; Whenever there is an idle slave core, it takes out a new image processing task from the task queue for processing; When the slave core finishes processing all the image processing tasks, the main core processing unit switches from the system state to the user state to obtain control, then performs the summary processing of the data obtained from the image processing tasks, and sends the processing result to the SRIO high-speed interconnection unit.
2. The software and hardware integrated parallel reconfigurable image real-time processing system according to claim 1, wherein The DSP node unit is a computing node, and serves as different process nodes according to its processing process and executes the corresponding image processing algorithms.
3. The software and hardware integrated parallel reconfigurable image real-time processing system according to claim 1, wherein, The master control unit extracts the image processing algorithms in each DSP node unit, obtains the parameters of the image processing algorithms and the serial-parallel structure of each DSP node unit through the process configuration file, and reconstructs the architecture of the DSP node unit based on the serial-parallel structure. After receiving the data packet of the image processing algorithm configured with parameters, the DSP node unit performs verification. After passing the verification, it first erases the original data in the FLASH, then writes the image processing algorithm into the FLASH. After that, the DSP node unit runs according to the written image processing algorithm and the serial-parallel order of each DSP node unit in the serial-parallel structure of the DSP node unit.
4. The software and hardware integrated parallel reconfigurable image real-time processing system according to claim 1, characterized in that In the SRIO high-speed interconnection unit, a routing and switching circuit method is adopted. Each DSP node unit is modularly designed as an end node unit, and each DSP node unit is interconnected through a high-speed interface; the data and results during the execution of the image processing task by each DSP node unit are transmitted to each other between each DSP node unit through the routing and switching circuit.
5. The software and hardware integrated parallel reconfigurable image real-time processing system according to claim 1, characterized in that The processing result of the last DSP node unit in the serial-parallel structure of the DSP node unit is output through the SRIO high-speed interconnection unit as the final image processing result.
6. A large field of view target segmentation and counting device, characterized in that, This device uses the software and hardware integrated parallel reconfigurable image real-time processing system according to any one of claims 1-5 to perform spot counting on the large field-of-view starry sky image.
Citation Information
Patent Citations
A DSP communication task scheduling system and method
CN109710399A
Dynamic reconfigurable intelligent computing cluster and configuration method thereof
CN108628800A
multi-core parallel signal processing system and method based on an SRIO bus
CN109656861A