Class increment target detection system based on heterogeneous embedded mode

By adopting a heterogeneous embedded system on small instruments, decomposing neural networks and combining the computing advantages of FPGA and ARM, low-power and efficient object detection are achieved, solving the problem of category changes in real-time and complex scenarios on small instruments, and improving computing speed and resource utilization efficiency.

CN120259738APending Publication Date: 2025-07-04TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510316168.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing object detection technologies are difficult to achieve efficient and real-time object detection on small instruments, especially in terms of computing resources and energy efficiency ratios, and traditional methods cannot effectively respond to category changes and real-time update requirements in complex scenarios.

Method used

Using a heterogeneous embedded class-increment object detection system, the embedded microprocessor ARM and field programmable gate array FPGA are used to decompose the neural network module into the full-connection layer and the non-full-connection layer parts, which are deployed on the ARM and FPGA respectively, and incremental learning is performed through knowledge distillation and quantization technology, and combined with 8-bit integer operation optimization calculation to achieve object detection.

Benefits of technology

Realizing low-power and efficient object detection on a resource-constrained platform can adapt to multiple goals in complex scenarios and perform incremental learning, improve computing speed and real-timeness, and reduce resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259738A_ABST
    Figure CN120259738A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of visual detection and machine learning, and aims to perform a target detection task on a real-time scene on a small instrument. Therefore, the technical scheme adopted by the invention is as follows: the heterogeneous embedded class increment target detection system comprises an embedded microprocessor ARM, a field programmable gate array FPGA, a server and a neural network module, the neural network module H (x) is decomposed into H (x) = G (F (x)), F (x) is a full-connection layer part of the neural network module, and F (x) is a full-connection layer part of the neural network module. The G (.) is a part of the neural network module except a full connection layer, the F (.) is deployed in the ARM, and the G (.) is deployed in the FPGA; the neural network module is a teacher and student network model, and knowledge distillation is carried out between the student network and the teacher network; after training is completed, the trained neural network module is used for target detection. The method is mainly applied to design and manufacturing occasions of small visual detection equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of visual detection and machine learning technology, and particularly relates to a heterogeneous embedded-based class incremental object detection system. Background Art

[0002] With the rapid development of object detection technology, the introduction of convolutional neural networks has increased the average accuracy of object detection technology from 20% to 85%, and object detection algorithms based on convolutional neural networks have become mainstream. By integrating the core capabilities of the network into a distributed platform, edge intelligence services can be easily provided. In this way, it can meet the fundamental needs of industry digitization in real-time services, data optimization, and application intelligence. Instruments such as drones can capture real-time images, and then object detection can be performed on the real-time images. By integrating the capabilities of the object detection network into such small instruments to form an object detection system, it can be applied to many object detection scenarios.

[0003] However, since the data to be processed in the real world may come from a large number of data streams containing complex categories and is very different from the data available during the initial training, the performance of the network may be affected to some extent. In this case, the performance of the network in early learning tasks drops rapidly, which is usually described as catastrophic forgetting or catastrophic interference. These situations affect the performance of existing models. The object detection model cannot change tasks during the test process. If you want to change the task of the model, such as adding new detectable categories, the model needs to be retrained on both the new and old category datasets at the same time. And the old category dataset may be very large, making the training time long and resource consumption high. More seriously, since traditional update methods need to save historical data and jointly train with new class data, they cannot handle the problem of loss of original data or inability to store historical data due to privacy reasons. Therefore, there are great disadvantages in adding new categories in this way. The introduction of continuous learning technology emerged as the times require to improve existing models to handle different data dynamic scenarios. Incremental learning, continuous learning, and lifelong learning and other technologies all refer to learning from continuous data streams, processing new tasks, and preserving the knowledge learned from previous learning tasks. At present, most of the incremental learning in the field of object detection is based on knowledge distillation. When a new task appears, first keep a copy of the existing detector, and then use it as a teacher network to guide the student network. By introducing distillation loss into the network loss function, the student network can imitate the teacher network to maintain the ability to detect old object classes, and at the same time can add new classes to achieve the ability to detect new classes.

[0004] In addition, the introduction of convolutional neural networks has brought problems such as high computational complexity and high spatial complexity while improving the accuracy. Due to the lightweight characteristics of small instruments and devices, the on-board energy and computing power are extremely limited, and it is difficult to deploy accurate and efficient detection algorithms. At the same time, the application scenarios have high requirements for real-time performance and computational performance. The traditional CPU architecture can no longer meet the high-real-time application scenarios. Especially on platforms with limited computing power, storage space, and operating power consumption such as embedded systems, it is difficult to deploy object detection technologies based on convolutional neural networks. Even if it is deployed on the CPU architecture through optimization or pruning, the recognition frame rate of the object detection system cannot meet the real-time requirements in the monitoring system.

[0005] Currently, embedded GPUs can be used as a solution. Their powerful parallel computing capabilities provide computing power support for network deployment, and at the same time, the rich software libraries make the deployment more convenient. However, due to the special computing architecture of embedded GPUs, they have significant disadvantages in terms of power consumption, cost, size, heat dissipation, and latency. Especially in terms of the energy efficiency ratio, which is the most concerned in object detection systems, they perform poorly. Therefore, when choosing an embedded GPU for edge computing applications, careful consideration and trade-offs must be made.

[0006] As a semi-custom integrated circuit, FPGA has natural advantages in terms of power consumption and computing power. Therefore, the design of a low-power object detection system based on FPGA has become one of the important research topics currently. FPGA is a hardware reconfigurable device based on look-up tables. Compared with traditional CPUs, FPGA has a large computing bandwidth and data transmission rate due to its unique architecture that does not depend on instruction sets, non-shared memory, and parallel computing methods. At the same time, due to the hardware reconfigurability of FPGA, users can freely define the hardware structure and data interface, thus achieving the purpose of deploying the required algorithms on FPGA. Therefore, using FPGA as an accelerator has become the best choice to improve the real-time performance and energy efficiency ratio of object detection systems. Summary of the Invention

[0007] To overcome the deficiencies of the prior art, the present invention aims to propose a visual detection technical solution capable of completing the target detection task for real-time scenes on small instruments. To this end, the technical solution adopted by the present invention is a heterogeneous embedded-based class incremental target detection system, including an embedded microprocessor ARM, a field programmable gate array FPGA, a server, and a neural network module. The neural network module H(·) is decomposed into H(x) = G(F(x)), where F(·) is the fully connected layer part of the neural network module, and G(·) is the part of the neural network module other than the fully connected layer. F(·) is deployed on the ARM, and G(·) is deployed on the FPGA. First, the basic initialization phase is used to initialize G(·), that is, offline training is performed on half of the classes in the dataset. After the basic initialization phase, the layers in G(·) are frozen. Then, during incremental learning, only F(·) is plastic and is updated on new data: the original pixel-level samples are stored in the replay buffer, and the compressed representation of the stored feature map tensor is stored. Specifically, for the input image x, the output of G(x) is a feature map z of size p*q*d, where p*q is the spatial grid size and d is the feature dimension. After initializing G(x) on the basic initialization dataset, all basic initialization samples are pushed through G(x) to obtain these feature maps for training the quantization PQ (Product Quantization) model. The PQ model set in the server encodes each feature map tensor into a p*q*s integer array, where s is the number of indices required for storage, that is, the number of codebooks used by PQ. After training the PQ model, the compressed representation of all basic initialization samples is obtained, and the compressed samples are added to the memory replay buffer. Then, the new samples are input into the model H(·) one by one. The new samples are compressed using the PQ model, a random subset of the samples is reconstructed from the memory buffer, and F(·) is updated on this mixed sample set for one iteration. The replay buffer is set to the upper limit of the memory. If the memory buffer is full, a new compressed example is added and an existing example is selected for deletion. Otherwise, only the new compressed sample is directly added. Among them, each iteration forms a student network, that is, a new H(x). At the same time, the student network also serves as the teacher network for the student network formed in the next iteration, and knowledge distillation is performed between the student and teacher networks. After training is completed, the trained neural network module is used for target detection.

[0008] Among them, 8-bit integers are used in the FPGA to replace floating-point numbers for the deployment of the partial neural network F(·). The specific process is as follows:

[0009] (1) Neural network module training:

[0010] First, the neural network module is trained, usually using standard floating-point parameters for training;

[0011] (2) Parameter quantization:

[0012] After training, the weight parameters of the neural network module are quantized into 8-bit integers. First, calculate the parameter range. For each weight matrix, find the minimum and maximum values to determine the quantization range. Then, calculate the quantization factor by dividing the quantization range into 256 evenly spaced intervals and computing the width of each interval, which will be the quantization factor. Finally, obtain the quantized parameters by mapping the floating-point parameter values to the closest quantization factor to get 8-bit integer values.

[0013] The specific formula is as follows:

[0014] V q = Q * (V x - min(V x ))

[0015] V' x = V q / Q + min(V x )

[0016] where Q = S / R, R = max(V x ) - min(V x ), S = 1 << $bits - 1, V x represents the original floating-point input, V q represents the quantized fixed-point value, V' x is the floating-point number restored according to the quantization parameters, and bits represents the quantization bit width.

[0017] (3) Inference process:

[0018] During the inference process, the input data and the quantized model parameters are used for forward propagation with 8-bit integers.

[0019] In the calculation, 8-bit integer operations, namely integer addition and multiplication, are used to perform convolution and fully connected operations.

[0020] In the FPGA, a multiplier and an accumulator are used for convolution calculation. The image pixel data stream is read serially. In each clock cycle, the data read first flows in sequentially from left to right and top to bottom. In each clock cycle, the convolution window moves one pixel from left to right and top to bottom to implement the convolution operation under each window.

[0021] Design a rectified linear unit (ReLU) module to determine the sign bit of the convolution output.

[0022] For the pooling layer, the output result is written back to the corresponding position of the output feature map. In each block, the maximum or average pooling operation is performed, and the pooling result is written back to the corresponding position of the output feature map.

[0023] The features and beneficial effects of the present invention are as follows:

[0024] In the field of visual detection, the present invention pays attention to the resource occupancy and running speed of the network model. Compared with traditional methods, it occupies less computing resources and is suitable for more complex real-world scenarios. It can achieve target detection on a low-power platform with limited resources. In addition, this method has flexibility and real-time performance, and can perform incremental learning according to multiple targets in complex scenarios to increase the recognizable categories to cope with various complex environments. Brief Description of the Drawings

[0025] Figure 1 Implementation structure flowchart.

[0026] Figure 2 Deployment flowchart.

[0027] Figure 3 Multilayer perceptron.

[0028] Figure 4 (a) Full-precision matrix multiplication operation; (b) Quantized low-precision matrix multiplication operation.

[0029] Figure 5 Vitis HLS design process. Detailed Implementation Manner

[0030] The present invention designs a heterogeneous embedded system to improve the computing speed and reduce resource utilization while meeting the algorithm performance requirements. This system aims to achieve efficient system deployment of the entire network from three levels: task memory access, computing resources, and computing tasks by dividing the tasks between the FPGA and the ARM, enabling the target detection network to run on a small instrument and process data in real time.

[0031] The technical solution adopted by the present invention is a class incremental object detection system based on heterogeneous embedding, including an embedded microprocessor ARM, a field programmable gate array FPGA, a server, and a neural network module. The neural network module H(·) is decomposed into H(x) = G(F(x)), where F(·) is the fully connected layer part of the neural network module, and G(·) is the part of the neural network module other than the fully connected layer. F(·) is deployed on the ARM, and G(x) is deployed on the FPGA. First, use the basic initialization stage to initialize G(·), that is, perform offline training on half of the classes in the dataset. After the basic initialization stage, the layers in G(·) are frozen. Then, during incremental learning, only F(·) is plastic and is updated on new data: the original pixel-level samples are stored in the replay buffer, and the compressed representation of the stored feature map tensor is stored. Specifically, for the input image x, the output of G(x) is a feature map z of size p*q*d, where p*q is the spatial grid size and d is the feature dimension. After initializing G(x) on the basic initialization dataset, all basic initialization samples are pushed through G(x) to obtain these feature maps, which are used to train the quantization PQ (Product Quantization) model. The PQ model set in the server encodes each feature map tensor into a p*q*s integer array, where s is the number of indices required for storage, that is, the number of codebooks used by PQ. After training the PQ model, the compressed representation of all basic initialization samples is obtained, and the compressed samples are added to the memory replay buffer. Then, the new samples are input into the model H(·) one by one; use the PQ model to compress the new samples, reconstruct a random subset of the samples from the memory buffer, and update F(·) on this mixed sample set for one iteration; the replay buffer is set to the upper limit of the memory. If the memory buffer is full, a new compressed example is added and an existing example is selected for deletion. Otherwise, only the new compressed sample is directly added. Among them, each iteration forms a student network, that is, a new H(x). At the same time, the student network also serves as the teacher network of the student network formed in the next iteration, and knowledge distillation is performed between the student and teacher networks; after training is completed, object detection is performed using the trained neural network module.

[0032] Among them, 8-bit integers are used in the FPGA to replace floating-point numbers for the deployment of the partial neural network F(·). The specific process is as follows:

[0033] (1) Neural network module training:

[0034] First, train the neural network module, usually using standard floating-point parameters for training;

[0035] (2) Parameter quantization:

[0036] After training, the weight parameters of the neural network module are quantized into 8-bit integers. First, calculate the parameter range. For each weight matrix, find the minimum and maximum values to determine the quantization range. Then, calculate the quantization factor by dividing the quantization range into 256 evenly spaced intervals and calculating the width of each interval, which will be the quantization factor. Finally, obtain the quantization parameters by mapping the floating-point parameter values to the closest quantization factor to get 8-bit integer values.

[0037] The specific formulas are as follows:

[0038] V q = Q * (V x - min(V x ))

[0039] V' x = V q / Q + min(V x )

[0040] where Q = S / R, R = max(V x ) - min(V x ), S = 1 << $bits - 1, V x represents the original floating-point input, V q represents the quantized fixed-point value, V' x is the floating-point number restored according to the quantization parameters, and bits represents the quantization bits;

[0041] (3) Inference process:

[0042] During the inference process, the input data and the quantized model parameters are used for forward propagation with 8-bit integers.

[0043] In the calculation, 8-bit integer operations, namely integer addition and multiplication, are used to perform convolution and fully connected operations.

[0044] In the FPGA, a multiplier and an accumulator are used for convolution calculation. The image pixel data stream is read serially. In each clock cycle, the data read first flows in from left to right and top to bottom in sequence. In each clock cycle, the convolution window moves one pixel from left to right and top to bottom to implement the convolution operation under each window.

[0045] Design a rectified linear unit (ReLU) module to determine the sign bit of the convolution output.

[0046] For the pooling layer, the output result is written back to the corresponding position of the output feature map. In each block, the maximum or average pooling operation is performed, and the pooling result is written back to the corresponding position of the output feature map.

[0047] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0048] Online incremental (lifelong) learning for object detection is employed in the present invention, where the agent has to learn one example at a time under severe memory and computational constraints. In object detection, the system has to output all the bounding boxes of an image with the correct labels. Different from previous work, the system described in the present invention can learn this task online, introducing new classes over time. The present invention achieves this function by using a novel memory replay mechanism, which can effectively replay the entire scene, and proposes the "Replay for the Online Detection of Objects (RODEO)" model, which replays the compressed representations stored in a fixed-capacity memory buffer to perform object detection step by step in a streaming manner. This algorithm can achieve class-incremental learning on the PASCAL VOC 2007 dataset and is computationally efficient, enabling deployment on embedded systems.

[0049] For the input image x, the object detection network H in RODEO can be decomposed into H(x) = G(F(x)), where G consists of a partial network of ResNet101, and F consists of the remaining part of ResNet101. We first initialize G(·) using the basic initialization phase, where RODEO first performs offline training on half of the classes in the dataset. After this basic initialization phase, the layers in G are frozen because the CNN layers of the partial ResNet101 learn general and transferable representations. Then, during incremental learning, only F is plastic and is updated on new data. Different from previous incremental image recognition methods, it stores the compressed representation of the feature map tensor instead of the original (pixel-level) samples in the replay buffer. Specifically, for the input image x, the output of G(x) is a feature map z of size p*q*d, where p*q is the spatial grid size and d is the feature dimension. After initializing G on the basic initialization dataset, all basic initialization samples are pushed through G to obtain these feature maps for training the quantization (Product Quantization, PQ) model. The PQ model encodes each feature map tensor into a p*q*s integer array, where s is the number of indices required for storage, i.e., the number of codebooks used by PQ. After training the PQ model, the compressed representations of all basic initialization samples are obtained, and the compressed samples are added to the memory replay buffer of the present invention. Then, new samples are input into the model H, one by one. The new samples are compressed using the PQ model, a random subset of samples is reconstructed from the memory buffer, and F is updated on this mixed sample set for one iteration. RODEO sets the replay buffer to the upper limit of the memory. If the memory buffer is full, new compressed examples are added and an existing example is selected for deletion. Otherwise, the new compressed samples are simply added. For all experiments, 8 bits or 1 byte equivalent is used to store the codebook indices, i.e., the size of each codebook is 256. COCO uses 64 codebooks and VOC uses 32 codebooks. For PQ calculations, RODEO uses the publicly available Faiss library. Algorithm 1 gives a description of the entire training process. For lifelong learning that needs to learn from a potentially infinite data stream, it is impossible to store all previous examples in the memory replay buffer. Since the capacity of the memory buffer of the present invention is fixed, it is important to replace less useful examples over time. The present invention uses a replacement strategy to replace the images with the fewest unique labels in the replay buffer.

[0050] PQ (Product Quantization), which is also a technique in the rodeo algorithm, is a step in the process of class incremental learning for an initial model. It is carried out on the server. Through PQ, the intermediate layer feature map of the CNN (with a size of p×q×d) is compressed into a more compact representation (an integer array of p×q×s) by PQ. For example, for the COCO dataset with a size of 25×30×2048 (width 25, height 30, and 2048 channels), it is compressed into a compact representation (such as 25×30×64) through Product Quantization (PQ), reducing memory occupancy and facilitating the efficient progress of incremental learning. PQ compresses the feature map generated by G, and after compression, it is placed in the buffer, facilitating F to update the weight parameters according to the feature maps replayed from the buffer. Every time a new class is learned, there will be new feature maps in the buffer, and F continuously replays from the buffer to update, so that it can avoid being biased towards new classes and forgetting old classes.

[0051] Since the present invention needs to be deployed on an embedded system, and the neural network model needs to be deployed on the FPGA and ARM respectively. The neural network on the FPGA requires a frozen network model. The ARM can deploy a changing network model, that is, a smaller incremental network model. Here, it is expected to deploy all the networks before the last layer of the backbone network ResNet101 as frozen layers on the FPGA, and deploy the remaining networks as incremental networks on the ARM to achieve lifelong learning.

[0052]

[0053] The splitting of tasks on the heterogeneous processor architecture mainly includes the following aspects: First, for the preprocessing of image data, such as denoising, scaling, enhancement, etc., it can be processed on the ARM side to utilize its strong general computing ability; Second, in terms of hardware network acceleration, a knowledge distillation mechanism is introduced in the network design stage to minimize the difference in the output of the old classes between the original network and the new network. The same parts of the new and old models are deployed on the FPGA side, and the changing fully connected layers are deployed on the ARM side, so as to deploy the class incremental object detection network on the heterogeneous embedded processor; Finally, the result is output to the display through the HP interface. In short, by reasonably allocating computing tasks between the ARM and the FPGA, the advantages of the heterogeneous processor architecture can be fully utilized to achieve an efficient and accurate object detection system.

[0054] The first training is for the teacher network, and after one incremental training, the student network is obtained. Among them, the same part of the two networks is the frozen layer G (the non-fully connected layer part). After the initial training (before class incremental learning), the weights of G are fixed (frozen) to ensure that the learned general features are not damaged during subsequent incremental learning. The part that changes is the fully connected layer, which is the unfrozen layer F and is plastic. It is updated on new data, adapts to new classes through incremental learning, and dynamically adjusts the weights to integrate new and old knowledge. When F is updated, it contacts both new and old data simultaneously to avoid over-biasing towards new classes, so that the so-called forgetting of past learning content can be achieved by this algorithm, resulting in a high accuracy for newly learned classes and a low accuracy for old classes. After one training, the teacher network changes to obtain a student network, and the student network learns new classes; when conducting the next learning, the student network in the previous stage becomes the teacher network for this learning, and so on in a cycle. The training code is set to ten processes, and when all ten processes are completed, the code stops.

[0055] Specifically, based on the class incremental object detection algorithm, the present invention performs object detection tasks on real-time scenes on a small instrument through the division of labor and cooperation between FPGA and ARM. Through this process, the effects of meeting the algorithm performance requirements while improving the computing speed and reducing resource utilization can be achieved.

[0056] The present invention uses the FPGA+ARM architecture to carry the object detection network and addresses the problem of diverse object types in complex environments based on the class incremental algorithm, deploying the network in a low-power environment, so as to achieve the purpose of real-time object detection on small devices such as drones.

[0057] The implementation structure of the present invention is as Figure 1 shown, applicable to scenarios that require real-time detection and have diverse categories. This network consists of an external camera and a heterogeneous embedded system with an FPGA+ARM architecture. The convolutional neural network is pre-designed on a computer, and the designed network is deployed. The deployment process of this system can be divided into four parts: model quantization, hardware logic generation, design of the data processing pipeline, and generation of hardware descriptions, as Figure 2 shown.

[0058] The problem overview of the present invention can be divided into two processes: model quantization and model deployment.

[0059] 1. Model quantization:

[0060] Most models trained on CPUs or GPUs use 32-bit floating-point (FP32) weights. Edge computing devices have low data bit widths and limited storage space, making them suitable for operations using low-bit-width data. If a model is to be deployed on an embedded device, the floating-point numbers need to be quantized into fixed-point numbers, which not only facilitates data reading, writing, and processing but also reduces the network size.

[0061] The most common frameworks for neural networks are Deep Neural Network (DNN), Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), etc. Taking DNN as an example, DNN consists of multiple non-linear transformations that form multiple processing layers and is an algorithm for high-level abstraction of data. As Figure 3 shown.

[0062] Among them, the multi-layer non-linear transformation is the dot product of the output of the previous layer and the weights corresponding to the neurons in the current layer (x*W), plus the bias (x*W+b) and activation processing (f(x*W+b)). Each neuron receives some inputs and performs some dot product calculations, which is very time-consuming, especially on devices that do not support hard floating-point operations, where the calculation of DNN becomes even more sluggish.

[0063] Therefore, quantization is introduced to reduce the computational amount to adapt to the deployment of the model on FPGAs. There is a close relationship between 8-bit quantization technology and FPGA deployment. FPGA is a programmable hardware accelerator commonly used in applications such as high-performance computing, embedded systems, and accelerating deep learning inference. The purpose of 8-bit uniform quantization is to use 8-bit integers to replace floating-point numbers, attempting to use fixed-point dot products to replace floating-point dot products, reducing the precision of the parameters to 8 bits, thereby reducing the storage and computational overhead of the model. This technology is usually used for the deployment of neural networks, especially on embedded devices or mobile devices, to reduce the memory footprint of the model and improve the inference speed. This greatly reduces the computational overhead of neural networks on devices without hard floating-point capabilities. Additionally, this method can also reduce the memory and storage footprint of the model.

[0064] The following are the steps of a basic 8-bit uniform quantization technology solution:

[0065] (1) Model training:

[0066] First, train the neural network model, usually using standard floating-point parameters for training.

[0067] (2) Parameter quantization:

[0068] After training, the weight parameters of the model are quantized to 8-bit integers. First, calculate the parameter range. For each weight matrix, find the minimum and maximum values to determine the quantization range. Here, the weight matrix is the calculation parameter for the input image during the calculation of the convolutional neural network, that is, an internal parameter of a model. Each parameter determines how the convolutional kernel weights the input data for calculation. The weight matrix is the multi-dimensional arrangement form of the weight parameters and directly participates in the convolutional operation. Then calculate the quantization factor. Divide the quantization range into 256 evenly spaced intervals and calculate the width of each interval, which will become the quantization factor. Finally, obtain the quantization parameter. Map the floating-point parameter value to the closest quantization factor to get an 8-bit integer value.

[0069] The specific formula is as follows:

[0070] V q = Q * (V x - min(V x ))

[0071] V' x = V q / Q + min(V x )

[0072] where Q = S / R, R = max(V x ) - min(V x ), S = 1 << $bits - 1. V x represents the original floating-point input, V q represents the quantized fixed-point value, V' x is the floating-point number restored according to the quantization parameter, and bits represents the quantization bit number.

[0073] (3) Inference process:

[0074] During the inference process, the input data and the quantized model parameters are used for forward propagation with 8-bit integers.

[0075] In the calculation, 8-bit integer operations (usually integer addition and multiplication) are used to perform operations such as convolution and fully connected.

[0076] 8-bit uniform quantization is a technique with precision loss because it reduces the representation of parameters to a lower precision level. Therefore, it may have a certain impact on the performance of the model.

[0077] The schematic diagrams of full-precision (32-bit floating-point number, FP32) matrix operations and quantized operations (assuming 8-bit integer quantization, INT8) are as Figure 4 shown.

[0078] The combination of 8-bit quantization technology and FPGA deployment can significantly improve the performance of deep learning models. On FPGAs, 8-bit integer calculations are generally faster than floating-point calculations, enabling higher inference speeds. Additionally, since 8-bit quantization reduces storage requirements, model parameters can be more easily loaded into the FPGA's memory, reducing data transfer overhead. The resources of FPGAs are limited, so reducing the storage requirements of the model is crucial for deploying deep learning models on FPGAs. 8-bit quantization technology makes the storage resources required by the model smaller, making it easier to adapt to the hardware resource limitations of FPGAs. FPGAs generally require less power to perform low-precision integer calculations because integer calculations are simpler and more energy-efficient than floating-point calculations. 8-bit quantization technology can further reduce power consumption, making FPGAs an energy-saving option for portable devices and embedded systems. FPGAs allow for the design of customized hardware accelerators to meet the specific requirements of deep learning tasks. 8-bit quantization technology makes it easier to design and implement these hardware accelerators because the hardware resources and computational requirements are relatively low. Although 8-bit quantization introduces a loss of precision, error correction techniques can be used to compensate for these losses to ensure that the performance of the model remains within an acceptable range. This requires implementing some compensation mechanisms on the FPGA to improve the accuracy of the model.

[0079] In FPGA:

[0080] Convolution module

[0081] Convolution is matrix multiplication. Matrix multiplication involves traversing the rows of matrix A and the columns of matrix B and calculating the product and sum. This method optimizes the storage reuse of matrix multiplication and can overwrite in place on the storage unit where the matrix is located, further reducing resource consumption. The data input is arranged in a streaming manner, and this data stream is presented in the form of a FIFO in the FPGA. Therefore, the block matrix needs to be cached in the FIFO. Each FIFO actually stores the vector data in the data block. Multiple copies of each FIFO data group are stored to provide parallel data for the circuit passing through. The data between the row and column FIFOs is controlled to flow in with a one-clock-cycle interval. The output result of the systolic array calculation is also cached in the FIFO and then the output data is passed to the next module through the FIFO.

[0082] Rectified Linear Unit (ReLU) module

[0083] The core of implementing ReLU in the FPGA is sign bit detection and multiplexer (MUX) control. If the sign bit is 0 (non-negative number), the original data is output; if the sign bit is 1 (negative number), all 0s are output. This module is embedded after the convolutional layer and directly applied as an activation function to the output of the convolutional operation to determine the sign bit of the output of the convolution.

[0084] Pooling layer

[0085] In HLS, hls::stream is used to process data streams. In a sliding window manner, the input stream is processed pixel by pixel, and the pooling operation is triggered at the appropriate output positions. For example, whenever enough row and column data are accumulated, a pooling calculation is triggered and the result is output. A row buffer array is used to store the data of the previous k - 1 rows of the window size. Especially when the stride is less than the kernel_size, the data reuse rate is high, and the row buffer can reduce external memory access. For each input pixel, it is read sequentially and filled into the row buffer. Whenever the row buffer is filled with enough rows, the pooling window starts to be processed. For example, for a 3x3 window, the data of the first two rows need to be saved. When processing the third row, it can be combined with the data of the first two rows to form a window.

[0086] 2. Model deployment:

[0087] Use the FPGA development tool to compile the hardware description into the hardware logic on the FPGA. In this step, the tool generates the logic circuit and then performs routing and timing analysis to ensure logical correctness and performance.

[0088] In the computing unit on the FPGA, multipliers and accumulators are used for convolution calculations. The image pixel data stream is read serially. In each clock cycle, the data read first flows in sequentially from left to right and from top to bottom. In each clock cycle, the convolution window moves one pixel from left to right and from top to bottom to perform the convolution operation under each window.

[0089] For the pooling layer, the output result is written back to the corresponding position of the output feature map. In each block, the maximum or average pooling operation is performed. The pooling result is written back to the corresponding position of the output feature map. The "block" of the pooling layer refers to the local area divided by the pooling window on the input feature map.

[0090] The traditional FPGA design process mainly uses VHDL, Verilog or System Verilog for project development, and then converges the design according to conditions such as delay, timing and resource usage, and synthesizes, places and routes the project written in the hardware description language. With the continuous improvement of the complexity of design tasks, HLS compilers have emerged, which can automatically convert high-level languages such as C / C++ into low-level hardware description languages (RTL) such as Verilog / VHDL. Use Vitis HLS to complete the project development based on C language or C++. The design process is as Figure 5 shown.

[0091] First, we use C / C++ to describe the algorithm, and we also need to prepare dependency functions. Vitis HLS has a dedicated graphical interface to set dependency functions, which constitute the input of the entire design.

[0092] Vitis HLS also integrates and provides C code libraries. These libraries cover arithmetic operations, video processing, signal / data processing, linear algebra, etc. Developers can directly call these library functions to speed up the description of their own C algorithms.

[0093] Subsequently, the above design is output as VHDL / Verilog code through the Vitis HLS platform, and a comprehensive report is generated, through which the performance evaluation, resource usage evaluation, interface information, etc. of the design can be viewed. Developers do not use these codes directly, but encapsulate them into IP cores, and then add the IP cores to the IPCatalog (IP core catalog) of Vivado for calling.

[0094] Add the IP module generated in the previous step to the pipeline for data stream processing. Taking video stream processing as an example, the data input by the camera is processed by the MIPI CSI-2Rx Subsytem (camera serial interface) module and the pixel data in raw10 format is output. Transmit 1920×1080 times to complete an image. The image data is processed by the designed IP module calculation unit, and then the demosaic module converts the raw10 data into RGB data to obtain an RGB image. The image can be further gamma corrected by the gamma_lut (gamma correction) module, enter the VDMA (Video Direct Memory Access, video register) module, and then enter the HP port display. The above steps can complete the pipeline array design in software and convert the neural network into a hardware circuit.

[0095] After hardware logic generation and optimization are completed, a bitstream file needs to be generated. Use FPGA development tools or JTAG (Joint Test Action Group) interface to load the generated bitstream file onto the target FPGA device, and the FPGA will be configured to perform the operations described by the hardware.

[0096] When designing a network, the network is divided into a frozen layer and an unfrozen layer. Among them, the frozen layer does not change during subsequent incremental processes and occupies a large amount of computing resources, making it suitable for the high-efficiency parallel computing capabilities of FPGAs. The unfrozen layer is the part that changes after class incremental learning and is deployed on the ARM side for modification. After retraining through incremental learning on the computer side, a new network model is generated, and data is transmitted through network communication to change the unfrozen layer on the ARM side, realizing the change from the original network on the heterogeneous embedded system to the new network after class incremental learning.

[0097] Specific application examples of the present invention are as follows:

[0098] (1) Resource-constrained detection devices such as drones and unmanned vehicles that need to carry target detection functions.

[0099] (2) Scenarios with complex and changeable roads, etc. in reality.

[0100] Based on the disclosure and teachings of the above specification, those skilled in the art to which the present invention pertains are also able to make changes and modifications to the above embodiments. Therefore, the present invention is not limited to the above specific embodiments, and any obvious improvements, substitutions, or variations made by those skilled in the art based on the present invention fall within the protection scope of the present invention. In addition, although some specific terms are used in this specification, these terms are only for convenience of description and do not constitute any limitation to the present invention.

Claims

1. A class-incremental object detection system based on heterogeneous embedding, characterized in that, It includes an embedded microprocessor ARM, a field-programmable gate array FPGA, a server, and a neural network module. The neural network module H(·) is decomposed into H(x) = G(F(x)), where F(·) is the fully-connected layer part of the neural network module, and G(·) is the part of the neural network module other than the fully-connected layer. F(·) is deployed on the ARM, and G(·) is deployed on the FPGA. First, use the basic initialization phase to initialize G(·), that is, perform offline training on half of the classes in the dataset. After the basic initialization phase, the layers in G(·) are frozen. Then, during incremental learning, only F(·) is plastic and is updated on new data: The original pixel-level samples are stored in the replay buffer, and a compressed representation of the stored feature map tensor is stored. Specifically, for the input image x, the output of G(x) is a feature map z of size p*q*d, where p*q is the spatial grid size and d is the feature dimension. After initializing G(x) on the basic initialization dataset, push all the basic initialization samples through G(x) to obtain these feature maps for training the quantization PQ (Product Quantization) model. The PQ model set in the server encodes each feature map tensor into a p*q*s integer array, where s is the number of indices required for storage, that is, the number of codebooks used by PQ. After training the PQ model, obtain the compressed representation of all the basic initialization samples, add the compressed samples to the memory replay buffer, and then input the new samples into the model H(·) one by one. Compress the new samples using the PQ model, reconstruct a random subset of the samples from the memory buffer, and update F(·) on this mixed sample set for one iteration. The replay buffer is set to the upper limit of the memory. If the memory buffer is full, add a new compressed example and select an existing example for deletion. Otherwise, just directly add the new compressed sample. Among them, each iteration forms a student network, that is, a new H(x). At the same time, the student network also serves as the teacher network for the student network formed in the next iteration, and knowledge distillation is performed between the student and teacher networks. After training is completed, use the trained neural network module for object detection.

2. The class incremental object detection system based on heterogeneous embedding as described in claim 1, wherein In the FPGA, 8-bit integers are used to replace floating-point numbers for the deployment of the partial neural network F(·). The specific process is as follows: (1) Neural network module training: Train the neural network module using standard floating-point parameters for training; (2) Parameter quantization: After training, quantize the weight parameters of the neural network module into 8-bit integers. First, calculate the parameter range. For each weight matrix, find the minimum value and the maximum value to determine the quantization range. Then, calculate the quantization factor. Divide the quantization range into 256 evenly spaced intervals and calculate the width of each interval, which will become the quantization factor. Finally, obtain the quantization parameters. Map the floating-point parameter values to the closest quantization factor to obtain 8-bit integer values; The specific formula is as follows: V q = Q * (V x - min(V x )) V′ x = V q / Q + min(V x ) Where Q = S / R, R = max(V x ) - min(V x ), S = 1 << $bits - 1, C x represents the original floating-point input, V q represents the quantized fixed-point value, V' x is the floating-point number restored according to the quantization parameter, and bits represents the quantization bit number; (3) Inference process: During the inference process, forward propagation is performed using 8-bit integers for the input data and the quantized model parameters; In the calculation, 8-bit integer operations, namely integer addition and multiplication, are used to perform convolution and fully connected operations; In the FPGA, a multiplier and an accumulator are used for convolution calculation. The image pixel data stream is read serially. In each clock cycle, the data read first flows in sequentially from left to right and from top to bottom. In each clock cycle, the convolution window moves one pixel from left to right and from top to bottom to implement the convolution operation under each window; A rectified linear unit (ReLU) module is designed to determine the sign bit of the convolution output; For the pooling layer, the output result is written back to the corresponding position of the output feature map. In each block, the maximum or average pooling operation is performed, and the pooling result is written back to the corresponding position of the output feature map.