Image convolution optimization method based on FT-M6678 chip
By designing a dedicated image convolution optimization method for the FT-M6678 DSP chip, using hardware instructions and mask data to construct, the problem that traditional computing methods are difficult to meet real-time image convolution processing is solved, and efficient and low-latency convolution operations are achieved, suitable for resource-constrained environments.
Patent Information
- Application Number
- CN202510230517.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
AI Technical Summary
Traditional general computing methods are difficult to meet the requirements of real-time image convolution processing of high-resolution images or multi-channel data, especially in environments where computing resources are constrained.
An image convolution optimization method specially designed for the FT-M6678 DSP chip is designed, and the dot product calculation is performed by constructing mask data, using hardware instructions such as _ddotp4, and processing image data column by column to achieve efficient convolution operation.
It significantly improves the efficiency of convolutional operations, especially suitable for scenarios that require high real-time and low latency, such as real-time image processing and embedded vision systems, with a performance improvement of about 40% compared to traditional methods.
Smart Images

Figure CN120147660A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital signal processing, and relates to an image convolution optimization method based on the FT-M6678 chip. Background Art
[0002] The FT-M6678 chip is a multi-core DSP chip completely independently developed by the National University of Defense Technology. It adopts the advanced 28nm process technology and has significant advantages of high performance and low power consumption. The chip supports a processing clock frequency of up to 1.25 GHz and has an 8-core high-speed processing architecture, which can significantly improve the signal processing ability. This makes the FT-M6678 perform excellently in processing complex computing tasks and is particularly suitable for application scenarios that require high performance and low power consumption. Based on the hardware characteristics of the FT-M6678, optimizing the algorithm design can further unleash its parallel computing potential and provide strong support for embedded real-time systems.
[0003] Image convolution is an important basic operation in digital image processing and is commonly used in tasks such as edge detection, blurring, and feature extraction. The computational complexity of the convolution operation is relatively high, especially when processing high-resolution images or multi-channel data, the amount of calculation will increase significantly. Traditional general computing methods are difficult to meet the requirements of real-time processing. Therefore, optimizing the algorithm for a specific hardware platform has become an important means to improve the efficiency of the convolution operation.
[0004] In practical applications, optimizing the performance of the DSP chip requires making full use of its hardware characteristics, such as multi-core parallel computing capabilities, dedicated hardware instruction sets (such as dot product operation instructions), and efficient memory access mechanisms. However, due to the large differences in the architecture characteristics of different DSP platforms, how to design and optimize efficient algorithms to achieve the optimal performance on a specific DSP platform is still an issue worthy of research. Summary of the Invention
[0005] The purpose of the present invention is to provide an image convolution optimization method based on the FT-M6678 chip, which is beneficial to efficiently execute image processing tasks with limited hardware resources and is particularly suitable for real-time image processing and environments with limited computing resources.
[0006] To achieve the above purpose, the technical solution of the present invention is as follows:
[0007] An image convolution optimization method based on the FT-M6678 chip includes the following steps:
[0008] Step S100: Construct a set of mask data according to the convolution kernel;
[0009] Step S200: Input the image data pointer, convolution kernel, and the length of each column of image data, and output the convolution result;
[0010] Step S300, initialize the memory pointer and load the image data and convolution kernels.
[0011] Step S400, perform dot product calculation on the image data and convolution kernels using the _ddotp4 hardware instruction.
[0012] Step S500, process the image data column by column and store the convolution results into the output pointer.
[0013] Step S600, repeat the above steps until all convolution calculations are completed.
[0014] The convolution kernel in Step S100 is an 8-bit signed number, and when stored, 4 8-bit numbers are combined into a 32-bit number.
[0015] The image data in Step S200 is a 16-bit signed number, and when stored, 2 16-bit numbers are combined into a 32-bit number.
[0016] In Step S300, initialize the memory pointer to align the input and output data in 4-byte units, and use the hardware optimization instructions _mem8_const and _amem4_const to load the image data.
[0017] In Step S400, use the _ddotp4 hardware instruction to perform parallel dot product calculation on the loaded image data and convolution kernels. This instruction requires two 32-bit variables, src1 and src2, as inputs, and two 32-bit variables, dst_o and dst_e, as outputs. Where dst_o is equal to the high 8 bits of the high 16-bit word of src1 multiplied by the high 16-bit word of src2, plus the low 8 bits of the high 16-bit word of src1 multiplied by the low 16-bit word of src2. Where dst_e is equal to the high 8 bits of the high 16-bit word of src1 multiplied by the low 16-bit word of src2, plus the low 8 bits of the low 16-bit word of src1 multiplied by the low 16-bit word of src2.
[0018] In Step S500, process the image data column by column, complete the convolution calculation of two pixels in each loop, and store the calculation results into the output pointer.
[0019] The advantages of the present invention are as follows: 1. The present invention proposes a 5×5 convolution optimization method specifically designed for the FT-M6678 DSP chip, which makes full use of the hardware characteristics of the chip to achieve overall optimization of memory access and calculation operations, and significantly improves the efficiency of convolution operations. 2. It is particularly suitable for scenarios requiring high real-time performance and low latency, such as real-time image processing, embedded vision systems, and tasks in other resource-constrained environments. 3. Under the same conditions, the time taken for this method to run 1000 times and calculate 1024 data points on the DSP platform is approximately 4 ms, while the running time of the traditional DSP platform without using this method is approximately 15 ms, showing a significant performance improvement. Brief Description of the Drawings
[0020] Figure 1 is the flowchart of the present invention;
[0021] Figure 2 is the illustration of the _ddotp4 function;
[0022] Figure 3 is the schematic diagram of the convolution kernel. Detailed Description of the Preferred Embodiment
[0023] The present invention will be further described below with reference to the accompanying drawings. The accompanying drawings are only for illustrative purposes and should not be construed as a limitation of this patent.
[0024] In order to describe this embodiment more concisely, some components that are well known to those skilled in the art but are not relevant to the main content of this creation will be omitted in the accompanying drawings or description. In addition, for the convenience of expression, some components in the accompanying drawings will be omitted, enlarged, or reduced, but this does not represent the size or all structures of the actual product.
[0025] The present invention discloses an image convolution optimization method based on the FT-M6678 chip. This solution is mainly used in environments with real-time image processing and limited computing resources. By making full use of the multi-core architecture and hardware instruction optimization of the FT-M6678 chip, the efficiency and performance of convolution calculation are significantly improved. The present invention will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0026] As Figure 1 shown, the image convolution optimization method based on the FT-M6678 chip includes the following steps
[0027] Step S100, constructing a set of mask data according to the convolution kernel;
[0028] As Figure 3 shown, taking M0, M1, M2, M3, M4 as examples, construct MASK01, MASK23, MASK45:
[0029] Adding a '0' at the front or end of the convolution data sequence to form a new data sequence, which includes 5 original data M0, M1, M2, M3, M4 and the added '0'.
[0030] Constructing the high 16-bit word of the first pair of convolution data MASK01, which is composed of 0 and M0 in the original data sequence, where 0 is the low byte and M0 is the high byte.
[0031] Constructing the low 16-bit word of the first pair of convolution data MASK01, which is composed of M0 and M1 in the original data sequence, where M0 is the low byte and M1 is the high byte.
[0032] Construct the high 16-bit word of the second pair of convolutional data MASK23, where the high 16-bit word is formed by concatenating M1 and M2 in the original data sequence, with M1 being the low byte and M2 being the high byte.
[0033] Construct the low 16-bit word of the second pair of convolutional data MASK23, where the low 16-bit word is formed by concatenating M2 and M3 in the original data sequence, with M2 being the low byte and M3 being the high byte.
[0034] Construct the high 16-bit word of the third pair of convolutional data MASK45, where the high 16-bit word is formed by concatenating M3 and M4 in the original data sequence, with M3 being the low byte and M4 being the high byte.
[0035] Construct the low 16-bit word of the third pair of convolutional data MASK45, where the low 16-bit word is formed by concatenating M4 and the added '0' in the original data sequence, with M4 being the low byte and '0' being the high byte.
[0036] Process and extract different data blocks of the convolutional data, such as MASK01, MASK23, MASK45. In a specific implementation, use the method of incrementing the pointer to sequentially extract each 16-bit data block and store it in a predetermined variable, such as mask1_01, mask1_23, etc. Each time data is extracted, the pointer will automatically increment to ensure that the data blocks are stored in order.
[0037] Step S200, input the image data pointer, convolutional kernel, and the length of each column of image data, and output the convolutional result; where the image data is a 16-bit signed number, usually stored as two 16-bit numbers combined into a 32-bit number, and the convolutional kernel is an 8-bit signed number, usually stored as four 8-bit numbers combined into a 32-bit number.
[0038] Step S300, initialize the memory pointer, and load the image data and the convolutional kernel;
[0039] Initialize the memory pointer to align the input and output data to 4 bytes, and use the hardware optimization instructions _mem8_const and _amem4_const to load the image data.
[0040] Verify the memory addresses of multiple pointer variables (such as data_in, data_out, and data_length) to ensure that their addresses are 4-byte aligned. If the address of a certain variable does not meet this condition, an error will be triggered through an assertion operation.
[0041] Next, perform the convolution operation. First, calculate the number of convolution loops count, and initialize the accumulative variables sum1 and sum2. Then, use the _mem8_const and _amem4_const functions to load 5 consecutive columns of data from memory, with 5 data in each column, ensuring that the data can be efficiently read and used for convolution calculation.
[0042] Through the _loll and _hill instructions, each column of data, such as col1_0123, col2_0123, etc., is first decomposed into two 16-bit data blocks, such as col1_01, col1_23, col2_01, col2_23, etc. Then these data are organized into a 6x5 data matrix.
[0043] Step S400, use the _ddotp4 hardware instruction to perform dot product calculation on the image data and the convolution kernel;
[0044] As Figure 2 shown, use the _ddotp4 hardware instruction to perform parallel dot product calculation on the loaded image data and the convolution kernel. This instruction requires two 32-bit variables as inputs, src1 and src2, and the outputs are two 32-bit variables, dst_o and dst_e. Among them, dst_o is equal to the high 8 bits of the high 16-bit word of src1 multiplied by the high 16-bit word of src2, plus the low 8 bits of the low 16-bit word of src1 multiplied by the high 16-bit word of src2. Among them, dst_e is equal to the high 8 bits of the high 16-bit word of src1 multiplied by the low 16-bit word of src2, plus the low 8 bits of the low 16-bit word of src1 multiplied by the low 16-bit word of src2.
[0045] The description is as Figure 2 shown. The result returns two 32-bit data, where dst_e is the result of the first point and dst_o is the result of the second point.
[0046] Step S500, process the image data column by column and store the convolution result into the output pointer.
[0047] After one round of loop, the sum1 obtained by summing dst_e is the calculation result of the first point, and the sum2 obtained by summing dst_o is the calculation result of the second point. The input image data pointer is incremented by 2, and data_out is incremented by 2. Perform the next round of loop until the end.
[0048] Step S600, repeat the above steps until all convolution calculations are completed.
[0049] The above is only the preferred embodiment of the present invention, and is not used to limit the scope of implementation of the present invention. That is, all equivalent changes and modifications made to the content of the scope of the patent application of the present invention should fall within the technical scope of the present invention.
Claims
1. An image convolution optimization method based on FT-M6678 chip, characterized in that: The following steps are included: Step S100, constructing a set of mask data according to the convolution kernel; Step S200, input image data pointer, convolution kernel, length of each column of image data, and output convolution result; Step S300, initializing the memory pointer, loading the image data and convolution kernel; Step S400, using the _ddotp4 hardware instruction to perform dot product calculation on the image data and the convolution kernel; Step S500, processing the image data column by column, and storing the convolution result into the output pointer; Step S600, repeat the above steps until all convolution calculations are completed.
2. The image convolution optimization method according to claim 1, characterized in that: The convolution kernel in step S100 is an 8-bit signed number, which is stored as 4 8-bit numbers combined into a 32-bit number.
3. The image convolution optimization method according to claim 1, characterized in that: In step S200, the image data is a 16-bit signed number, and when stored, two 16-bit numbers are combined into a 32-bit number.
4. The image convolution optimization method according to claim 1, characterized in that: In step S300, the memory pointer is initialized so that the input and output data are aligned to 4 bytes, and the image data is loaded using the hardware optimized instructions _mem8_const and _amem4_const.
5. The image convolution optimization method according to claim 1, characterized in that: In step S400, the _ddotp4 hardware instruction is used to perform parallel dot product calculation on the loaded image data and the convolution kernel. The instruction requires the input of two 32-bit variables, src1 and src2, and the output is 32-bit variables dst_o and dst_e; wherein dst_o is equal to the high 16-bit word of src1 multiplied by the high 8-bit word of the high 16-bit word of src2, plus the low 16-bit word of src1 multiplied by the low 8-bit word of the high 16-bit word of src2; wherein dst_e is equal to the high 16-bit word of src1 multiplied by the high 8-bit word of the low 16-bit word of src2, plus the low 16-bit word of src1 multiplied by the low 8-bit word of the low 16-bit word of src2.
6. The image convolution optimization method according to claim 1, characterized in that: In step S500, the image data is processed column by column, and the convolution calculation of two pixels is completed in each cycle, and the calculation result is stored in the output pointer.