Sensing and calculating integrated spectrum camera processor system for side end real-time analysis

Through the design of a three-level board architecture and a lightweight hyperspectral classification network, the real-time and power consumption issues of traditional hyperspectral imaging systems are solved, and efficient and real-time hyperspectral image classification is achieved, which is suitable for edge applications.

CN120743845AActive Publication Date: 2025-10-03HUNAN UNIV

Patent Information

Application Number
CN202511196046.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-10-03
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

The push-broom architecture of traditional hyperspectral imaging systems leads to large bandwidth and latency overheads, and relies on high-power general-purpose processors, making it difficult to meet the real-time, low-power, and compact requirements of the edge side. Existing deep learning models deployed on FPGAs have problems such as deep attention structure and irregular data flow, making it difficult to achieve high-precision spectral classification.

Method used

A sensing-computing integrated spectral camera processor system for edge real-time analysis is designed. The system adopts a three-level board architecture, integrates spectral preprocessing and accelerators, and a lightweight hyperspectral classification network PBViT. It combines the frequency domain encoding branch SSFE and the hardware-friendly ViT branch HFViT. Matrix multiplication and nonlinear calculations are performed in parallel through a time overlapping scheduling strategy. 8-bit quantized perception training is used to achieve hardware-friendly model deployment.

Benefits of technology

It achieves efficient and real-time hyperspectral image classification, reduces latency and power consumption, improves classification accuracy, is suitable for edge applications, and has good quantitative robustness and hardware adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743845A_ABST
    Figure CN120743845A_ABST
Patent Text Reader

Abstract

The invention discloses an edge real-time analysis-oriented sensing and calculation integrated spectrum camera processor system. Efficient edge calculation is realized through collaborative innovation of hardware, an algorithm and an accelerator. A hardware system adopts a three-level board card framework; a core board integrates a programmable logic unit and a processing system, and processes spectral data acquired by a CMOS board in real time; and the interface board accurately controls the power supply time sequence. In the lightweight classification network, the SSFE branch converts frequency domain filtering into cyclic matrix multiplication; carrying out normalization on the HFViT branch by using a batch normalization substitution layer; and an 8-bit full integer quantization strategy compression model is matched, and input spectrum identification characteristics are reserved. Time overlapping scheduling is designed for the accelerator, and SSFE matrix multiplication is executed in parallel in the HFViT calculation window period; the outer product pulsation architecture improves the matrix operation efficiency; and non-linear kernel multiplexing computing resources are unified. The system supports real-time classification of push-scan spectrum data, results are directly connected with sorting equipment, and the edge end embedded deployment requirement is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep learning and embedded image processing technology, and in particular relates to a sensing and computing integrated spectral camera processor system for edge real-time analysis. Background Art

[0002] Hyperspectral imaging systems, with their high spatial resolution and continuous spectral sampling capabilities, are widely used in fields such as the non-contact, non-destructive classification and detection of pharmaceutical and other chemical substances. However, traditional push-broom hyperspectral systems typically employ a separate "perception-computation" architecture, requiring the collected spectral data to be transmitted via a high-speed interface to an industrial computer for post-processing. This architecture not only introduces significant bandwidth and latency overhead but also relies on power-hungry general-purpose processors, making it difficult to meet the real-time, low-power, and compact requirements of edge applications.

[0003] Currently, a variety of processor platforms are being used for hyperspectral image processing, including CPUs, GPUs, ASICs, and FPGAs. CPUs offer architectural flexibility but limited computing power; GPUs offer high parallel computing capabilities but low energy efficiency; and ASICs offer the advantage of high customization but limited flexibility and high development costs. In contrast, FPGAs, with their reconfigurable logic structure, refined parallel capabilities, and low power consumption, are an ideal choice for building specialized hyperspectral computing platforms. Previous studies have implemented several traditional hyperspectral algorithms on FPGAs, such as hybrid decomposition algorithms based on image space reconstruction, object detection based on Gram–Schmidt orthogonalization, and anomaly detection based on multivariate Gaussian models, achieving some success in terms of real-time performance and resource efficiency. However, these methods generally rely on traditional analytical models, making it difficult to extract complex nonlinear relationships between spectra and spatial contextual information, resulting in low accuracy in fine-grained spectral classification. With the rapid development of deep learning in recent years, the Visual Transformer (ViT) architecture has become a new trend in hyperspectral image analysis due to its global receptive field and powerful spatial-spectral modeling capabilities. Furthermore, research has begun deploying networks such as lightweight CNNs on FPGA platforms, leveraging their parallel architecture for low-power inference. However, existing deployments still focus on CNN and RNN structures, and there is a lack of research on deploying ViT-based models on FPGAs. In addition, existing models based on ViT improvements often suffer from problems such as deep attention structures and irregular data flows, which pose great challenges to hardware implementation.

[0004] To address the above problems, the present invention proposes a sensing and computing integrated spectral camera processing solution for real-time edge analysis. Summary of the Invention

[0005] In response to the above technical problems, the present invention provides a spectral camera processor system with integrated sensing and computing for real-time edge analysis.

[0006] The technical solution adopted by the present invention to solve the technical problem is: A sensor-computing integrated spectral camera processor system for edge real-time analysis, including a hardware processing system, a lightweight hyperspectral classification network PBViT, and a PBViT hardware accelerator; The hardware processing system includes a core board, an interface board, and a CMOS board. The core board integrates a programmable logic unit (FPGA) and a processing system (PS), and connects to the CMOS board via a MIPI CSI-2 interface to collect hyperspectral data. The interface board is equipped with an STM32 microcontroller to manage multi-voltage domain power supply timing and external interfaces. The CMOS board is equipped with a CMOS sensor. The lightweight hyperspectral classification network PBViT takes as input a spatial patch generated from a hyperspectral image of the object being measured and outputs a category prediction result, which is displayed to an external display system via the UART interface on the PS side. The network includes a frequency domain encoding branch SSFE and a hardware-friendly ViT branch HFViT, which are executed in parallel. The frequency domain encoding branch converts frequency domain filtering into circulant matrix multiplication, and the hardware-friendly ViT branch replaces layer normalization with batch normalization. The network is trained with 8-bit quantization awareness to achieve full integer quantization. The PBViT hardware accelerator integrates linear computing cores, nonlinear computing cores and residual path cores, and executes SSFE branches in parallel during the Softmax calculation of HFViT through a time overlapping scheduling strategy. It adopts a streaming matrix multiplication architecture based on outer products and time-division multiplexing BN, quantization, Softmax and GELU operations through a unified nonlinear computing core.

[0007] Preferably, the core board in the hardware processing system has a built-in spectral pre-processing module and a PBViT accelerator, is configured with a DDR4 cache array and a 10G Ethernet communication interface, and uploads spectral data to a host computer via a UDP protocol; The interface board provides RJ45 Ethernet port, JTAG debugging interface, STLINK interface, USB to UART interface, 12V power supply interface and SD card slot; The CMOS board uses the PCA9306 level conversion chip to achieve compatible communication between the 1.8V I²C and the core board's 1.2V I / O Bank.

[0008] Preferably, the data collection of the core board includes: Use DPHY resources to decouple LP and HS signals, and use IDELAY primitives to achieve phase alignment of clock and data; Byte alignment is achieved through Bitslip technology, and multi-channel synchronization is achieved with the help of SoT packets; The spectral preprocessor module adopts a pipeline architecture to complete pixel merging, whiteboard correction and spectrum line smoothing operations; The pre-processed spectral lines are written into the DDR4 chip cache through the built-in DDR controller and assembled into a 9×9 spatial patch. The patch is represented as: ; in, Represents a spectral curve with a dimension of 63. Indicates the number of spectral curves contained in the spatial patch.

[0009] Preferably, the STM32 microcontroller of the interface board monitors the Power-Good signal of each voltage regulator and outputs the Enable signal according to the preset timing to control the power-on sequence of the core board and the CMOS board; the integrated buck regulator converts the 12V input into a multi-voltage domain power supply.

[0010] Preferably, in the frequency domain coding branch, it is assumed that the input image block is ,in, represents the spatial dimension, is the spectral number, and this branch performs the following frequency domain filtering operations in the spatial and spectral dimensions in turn: ; in, is a conjugate symmetric learnable filter, represents element-wise multiplication, is the Fourier transform, is the inverse Fourier transform, is the input image block, is the image block output after frequency domain filtering; According to the convolution theorem, the above formula can be equivalently converted into matrix multiplication form: ; in, Indicates based on This transformation is used during the network's inference process, and the standard matrix multiplication circuit can be reused.

[0011] Preferably, each ViTBlock of HFViT consists of the following two substructures: ; ; in, For input, is the intermediate output after multi-head attention, For the final output, represents the batch normalization operation, is multi-head self-attention, It is a two-layer feed-forward network with GELU activation.

[0012] Preferably, the quantization strategy of the quantization-aware training process includes token-level quantization of input activations, layer-level quantization of intermediate activations, and channel-level quantization of weights; Quantization perturbations are introduced during quantization-aware training to adapt the model to the errors caused by quantization. Quantization adopts a symmetric linear strategy: ; in is a floating point number, is the scaling factor obtained by the min-max algorithm, To quantize the bit width, the clip function limits the output to the representable range. Used to round input to an integer.

[0013] Preferably, the time overlap scheduling strategy is specifically: During the HFViT branch’s Softmax execution, the linear computation module is idle due to the delay in its exponential computation and normalization operations. The cyclic matrix multiplication of the SSFE branch is executed in parallel during this idle window. The idle GELU path DSP resources in the nonlinear computing core are used to process the BN operation of SSFE.

[0014] Preferably, the linear computation core uses a 32×32 outer product systolic array; each cycle, one column of the input matrix A and one row of B are directly calculated in the processing array; through a dual 8-bit operand packing strategy, two multiplications are performed in parallel within a single DSP48E2 unit; the array supports streaming data scheduling with results propagated to the right or downward.

[0015] Preferably, the unified nonlinear computing core performs BN, quantization, Softmax and GELU operations by time-division multiplexing hardware resources, wherein BN and quantization are fused into channel-level multiplication-addition operations, and quadratic polynomial integer approximation is adopted for Softmax and GELU.

[0016] The present invention proposes a sensor-computing integrated spectral camera processor system for real-time edge analysis, achieving triple breakthroughs through collaborative innovation of hardware, algorithm, and accelerator: the hardware system adopts a three-level board architecture, the core board integrates spectral preprocessing and accelerator, the CMOS board is equipped with a 400-700nm high quantum efficiency sensor to accurately capture the characteristic spectrum of the object under test, and the interface board controls the power timing through STM32 to ensure stable operation in complex environments; the PBViT network integrates dual branches: SSFE converts frequency domain filtering into circulant matrix multiplication to avoid FFT resource consumption, HFViT uses batch normalization to eliminate real-time computing overhead, and cooperates with 8-bit full integer quantization to compress the model while retaining the model's fitting ability; the accelerator design pioneers time overlapping scheduling, and executes SSFE matrix multiplication in parallel during the Softmax window period of HFViT to reduce latency. The outer product systolic architecture is combined with dual 8-bit operand packing, unified nonlinear kernel fusion BN and quantization, and integer approximation of Softmax and GELU to reduce resource overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a schematic diagram of the structure of a hyperspectral camera circuit hardware processing system in one embodiment of the present invention; Figure 2 Schematic diagram of the principle of a lightweight hyperspectral classification network PBViT in one embodiment of the present invention; Figure 3 FIG. 4 is a schematic diagram of the structure of a PBViT hardware accelerator in one embodiment of the present invention. DETAILED DESCRIPTION

[0018] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is further described in detail below with reference to the accompanying drawings.

[0019] In one embodiment, a sensor-computing integrated spectral camera processor system for edge real-time analysis includes a hardware processing system, a lightweight hyperspectral classification network PBViT, and a PBViT hardware accelerator; The hardware processing system includes a core board, an interface board, and a CMOS board. The core board integrates a programmable logic unit (FPGA) and a processing system (PS), and connects to the CMOS board via a MIPI CSI-2 interface to collect hyperspectral data. The interface board is equipped with an STM32 microcontroller to manage multi-voltage domain power supply timing and external interfaces. The CMOS board is equipped with a CMOS sensor. The lightweight hyperspectral classification network PBViT takes as input a spatial patch generated from the spectral image of the object being measured and outputs the category prediction result, which is output to an external display system via the UART interface on the PS side. The network includes a frequency domain encoding branch SSFE and a hardware-friendly ViT branch HFViT, which are executed in parallel. The frequency domain encoding branch converts frequency domain filtering into circulant matrix multiplication, and the hardware-friendly ViT branch replaces layer normalization with batch normalization. The network is trained with 8-bit quantization awareness to achieve full integer quantization. The PBViT hardware accelerator integrates linear computing cores, nonlinear computing cores and residual path cores, and executes SSFE branches in parallel during the Softmax calculation of HFViT through a time overlapping scheduling strategy. It adopts a streaming matrix multiplication architecture based on outer products and time-division multiplexing BN, quantization, Softmax and GELU operations through a unified nonlinear computing core.

[0020] Specifically, the first aspect of the present invention proposes a hyperspectral camera circuit hardware processing system, such as Figure 1 As shown, the hardware processor for a single camera consists of a core board, an interface board, and a CMOS board. The core board uses the Xilinx ZYNQ Ultrascale+ MPSoC ZU19EG as its core processor and integrates a power management module, memory (including PS / PL DDR, QSPI memory, and eMMC memory), an optical port (SFP+), and other related peripheral interfaces. The interface board, equipped with an STM32F103RBT6 microcontroller, is responsible for power sequencing and external interface management (ETH, UART, SD, etc.). The CMOS board, equipped with a SmartSens SC130GS CMOS sensor, provides high-precision image acquisition and communicates with the core board via a 4-lane MIPI interface.

[0021] The second aspect of this invention proposes a lightweight hyperspectral classification network, PBViT, which consists of two pathways: hardware-friendly ViT (HFViT) and spatial spectrum frequency-domain encoding (SSFE). After passing through a spectral preprocessing module, input spectral lines are cached as spatial patches in DDR4 on the PL side. PBViT then predicts the class of the object being measured. Quantization-aware training (QAT) is used to quantize the model to 8-bit integers, facilitating subsequent deployment on an FPGA accelerator.

[0022] The third aspect of the present invention proposes an FPGA-based PBViT accelerator. Using a time-overlapping scheduling strategy, the SSFE prediction time is hidden within the HFViT inference process, fully reusing HFViT's computational resources. An outer-product-based linear computation kernel is designed for all matrix multiplication operations in the network. A general nonlinear computation kernel is also designed for all nonlinear operations in the network, including quantization, dequantization, softmax, normalization, and GELU.

[0023] In one embodiment, the core board of the hardware processing system has a built-in spectral pre-processing module and a PBViT accelerator, is equipped with a DDR4 cache array and a 10G Ethernet communication interface, and uploads spectral data to a host computer via the UDP protocol; The interface board provides RJ45 Ethernet port, JTAG debugging interface, STLINK interface, USB to UART interface, 12V power supply interface and SD card slot; The CMOS board uses the PCA9306 level conversion chip to achieve compatible communication between the 1.8V I²C and the core board's 1.2V I / O Bank.

[0024] Specifically, the core board is the core unit of the optical circuit hardware system, responsible for spectral reception, on-chip computation, and communication management. Based on the ZU19EG chip platform, this board integrates a programmable logic (FPGA) and a processing system (PS). It boasts powerful computing power, memory bandwidth, and high-speed communication capabilities, making it suitable for the real-time processing requirements of push-broom hyperspectral imaging. Spectral data is acquired via a CMOS board, using the SmartSens SC130GS CMOS image sensor chip as its onboard photosensor. This chip features 1.3 megapixels, a global shutter, and high quantum efficiency in the 400–700 nm range. It outputs 10-bit spectral data at 120 frames per second via a four-lane MIPI CSI-2 interface. The sensor's required 27MHz input clock is provided by an onboard crystal oscillator. Furthermore, the CMOS uses a 1.8V I²C interface voltage level, while the SoC's I / O bank VCCO is 1.2V. Therefore, the design incorporates a PCA9306 level shifter to enable bidirectional voltage-compatible communication.

[0025] In one embodiment, data collection of the core board includes: Use DPHY resources to decouple LP and HS signals, and use IDELAY primitives to achieve phase alignment of clock and data; Byte alignment is achieved through Bitslip technology, and multi-channel synchronization is achieved with the help of SoT packets; The spectral preprocessor module adopts a pipeline architecture to complete pixel merging, whiteboard correction and spectrum line smoothing operations; The pre-processed spectral lines are written into the DDR4 chip cache through the built-in DDR controller and assembled into a 9×9 spatial patch. The patch is represented as: ; in, Represents a spectral curve with a dimension of 63. Indicates the number of spectral curves contained in the spatial patch.

[0026] Specifically, for spectral reception, the FPGA acquires high-speed data via the MIPI CSI-2 interface, using DPHY resources to decouple the low-frequency and high-frequency signals and using the IDELAY primitive to phase-align the clock and data. Byte alignment is then achieved using Bitslip technology, and multi-channel alignment is achieved using SoT packets. For on-chip computing, the system integrates a spectral preprocessing module and a PBViT accelerator. The spectral preprocessor module employs a pipelined architecture to perform pixel binning, whiteboard correction, and spectral line smoothing. Processed data is written to a DDR4 chip buffer via a built-in DDR controller. The FPGA is equipped with four MT40A1G16RC DDR4 chips, totaling 8GB and operating at 1600MHz, for data caching for hyperspectral processing. The PBViT accelerator performs real-time classification of spectral images. For data communication, control signals are transmitted via the JTAG interface for host debugging and the I²C interface for CMOS register configuration. Data transmission utilizes a 10G Ethernet network using GTH transceivers, combined with the UDP protocol, to stream spectral images to a host computer. PBViT's classification results are sent to external systems via the PS's UART interface for display and logging. Furthermore, the core board integrates QSPI and eMMC storage for program startup and configuration file storage, respectively. The PS also features 8GB of DDR4 memory, providing the ARM core with the memory resources required for software execution.

[0027] In one embodiment, the STM32 microcontroller of the interface board monitors the Power-Good signal of each voltage regulator and outputs the Enable signal according to the preset timing to control the power-on sequence of the core board and the CMOS board; the integrated buck regulator converts the 12V input into a multi-voltage domain power supply.

[0028] To accommodate the system's multiple voltage domains, the interface board incorporates multiple buck regulators to convert the 12V input voltage into the various voltages required by each module. Given the ZU19EG's stringent power-up sequencing requirements, an STM32F103 microcontroller is integrated into the interface board for power management and power-up control. This controller monitors the PG (Power-Good) signal of each regulator and outputs the EN (Enable) signal in a pre-set order, ensuring reliable system startup.

[0029] Furthermore, the present invention proposes a PBViT network suitable for hyperspectral image classification, such as Figure 2 As shown in the figure, the overall structure consists of two parallel branches: the frequency domain encoding branch (SSFE) and the hardware-friendly ViT branch (HFViT). This structure fuses the outputs of the two branches through soft voting to achieve high-precision, low-latency push-broom hyperspectral image recognition, suitable for FPGA deployment.

[0030] In one embodiment, the characteristic distribution law of the hyperspectral image in the frequency domain is used to design a frequency domain encoding branch SSFE. In the frequency domain encoding branch, it is assumed that the input image block is ,in, represents the spatial dimension, is the spectral number, and this branch performs the following frequency domain filtering operations in the spatial and spectral dimensions in turn: ; in, is a conjugate symmetric learnable filter, represents element-wise multiplication, is the Fourier transform, is the inverse Fourier transform, is the input image block, is the image block output after frequency domain filtering; According to the convolution theorem, the above formula can be equivalently converted into matrix multiplication form: ; in, Indicates based on This transformation is used in the network's inference process and can reuse the standard matrix multiplication circuit, thereby effectively avoiding the resource overhead brought by the frequency domain transformation.

[0031] In one embodiment, to reduce inference latency, the present invention designs HFViT, replacing layer normalization (LN) in standard ViT with batch normalization (BN) to avoid the overhead of real-time mean and variance calculations. Each ViTBlock in HFViT consists of the following two substructures: ; ; in, For input, is the intermediate output after multi-head attention, For the final output, represents the batch normalization operation, is multi-head self-attention, This is a two-layer feedforward network with GELU activations. This module has good hardware mapping and is easy to implement in pipelines.

[0032] In one embodiment, the quantization strategy of the quantization-aware training process includes token-level quantization of input activations, layer-level quantization of intermediate activations, and channel-level quantization of weights; Quantization perturbations are introduced during quantization-aware training to adapt the model to the errors caused by quantization. Quantization adopts a symmetric linear strategy: ; in is a floating point number, is the scaling factor obtained by the min-max algorithm, To quantize the bit width, the clip function limits the output to the representable range. Used to round input to an integer.

[0033] Specifically, to achieve full integer deployment, this invention utilizes quantization-aware training (QAT). The selection of quantization granularity is crucial. The quantization granularity employed in this invention includes: token-level quantization of input activations to preserve fine-grained spatial-spectral features; layer-level quantization of intermediate activations to reduce hardware deployment complexity; and channel-level quantization of weights to maintain model expressiveness. Furthermore, element-by-element operations such as batch normalization (BN), bias addition after linear layers, and quantization are integrated into a unified computational unit during inference to avoid redundant computation. Nonlinear functions (such as Softmax and GELU) are approximated using quadratic polynomials.

[0034] Furthermore, in order to achieve efficient deployment under the limited logic and storage resources of FPGA, the present invention proposes a hyperspectral network accelerator with overlapping execution architecture, such as Figure 3 As shown, the accelerator supports real-time inference of the PBViT model. This accelerator includes three specialized computing modules: linear cores, nonlinear cores, and residual path cores. Modules use on-chip SRAM for intermediate data caching. Image input is acquired through the MIG interface using external DDR4 memory, and model parameters are dynamically loaded via DMA. A finite state machine controls data flow and memory scheduling within each module, improving overall hardware utilization.

[0035] In one embodiment, the time overlap scheduling strategy is specifically: During the HFViT branch’s Softmax execution, the linear computation module is idle due to the delay in its exponential computation and normalization operations. The cyclic matrix multiplication of the SSFE branch is executed in parallel during this idle window. The idle GELU path DSP resources in the nonlinear computing core are used to process the BN operation of SSFE.

[0036] Specifically, parallel execution during the idle window period effectively hides inference latency. Furthermore, the SSFE output requires batch normalization (BN) processing, and the corresponding Softmax path in the nonlinear kernel already occupies the first DSP resource. To avoid conflicts, the present invention uses the idle DSP resources in the GELU path for BN calculations, thereby achieving inter-module resource reuse and timing coupling.

[0037] In one embodiment, the linear computation core employs a 32×32 outer product systolic array. Each cycle, one column of the input matrix A and one row of the input matrix B are directly computed in the processing array. Two multiplications are performed in parallel within a single DSP48E2 unit using a dual 8-bit operand packing strategy. The array supports streaming data scheduling with results propagated to the right or downward.

[0038] Specifically, the present invention employs an outer-product-based streaming matrix multiplication architecture, avoiding the input tensor reordering overhead of traditional architectures. Each cycle, one column of the input matrix A and one row of the input matrix B are directly computed within the processing array. The core architecture is a 32×32 quasi-systolic array, with each PE unit equipped with integer multiply-add capabilities, local registers, and result forwarding paths. DSP48E2 resources are used for multiplication mapping, and a dual 8-bit operand packing strategy is designed to execute two multiplications in parallel within each DSP, saving half the multiplication resources. The array supports rightward or downward propagation of results, adapting to different data flow directions.

[0039] In one embodiment, the unified nonlinear computing core performs batch normalization, quantization, softmax, and GELU operations by time-division multiplexing hardware resources, wherein the batch normalization and quantization are fused into a channel-level multiplication-addition operation, and quadratic polynomial integer approximation is used for softmax and GELU.

[0040] Specifically, the nonlinear computing core adopts a unified structural design and supports BN, quantization, Softmax, GELU and other operations through time division multiplexing. In the inference stage, BN and quantization are both channel-level element-by-element operations. The system integrates the two into an arithmetic path and uses pre-stored fixed-point parameters to implement one-time multiplication and addition. The bias addition and quantization in the linear layer are also merged and reuse the same hardware logic. In order to adapt to integer operations, the present invention performs polynomial approximation on the Softmax and GELU functions, and replaces the floating-point operations of the exponential and error functions with integer multiplication and addition to reduce resource overhead. These approximate operations are integrated into a unified nonlinear unit to reduce logic resource consumption.

[0041] In order to evaluate the accuracy and consistency of the hyperspectral image classification processing system proposed in this invention, the following three evaluation indicators are selected: overall accuracy (OA), average accuracy (AA) and Cohen's Kappa consistency coefficient (κ).

[0042] Comparative Experiments. The proposed PBViT model was compared with traditional machine learning models, CNN methods, and ViT-based network architectures. The experimental dataset used a traditional Chinese medicine dataset, and the results are summarized in Table 1. Traditional machine learning models lack the ability to model joint spatial-spectral features, resulting in poor overall classification performance. CNN methods outperform traditional models in terms of accuracy. In particular, three-dimensional convolutional networks (3DCNNs) can simultaneously model both spatial and spectral dimensions, achieving high accuracy. However, these networks have large parameter counts and rely on complex data rearrangement operations, resulting in inefficient deployment on hardware platforms. ViT-based models generally achieve optimal performance due to their global modeling capabilities. However, SSFTT and GCSViT introduce irregular network topologies, significantly increasing the complexity of hardware accelerator design. In contrast, the proposed PBViT model achieves optimal results across all metrics and has the following advantages: a regular network structure, minimal parameters, and hardware-friendly adaptability.

[0043] Table 1 Comparative experiment

[0044] Ablation and Quantification Experiments. To verify the independent contributions of each module proposed in this paper, ablation experiments were designed and conducted on the Chinese Herbal Medicine (CHM) dataset. The results are shown in Table 2. First, after replacing layer normalization (LayerNorm) in the ViT network with batch normalization (BatchNorm) to construct the HFViT structure, the overall accuracy (OA) of the model remained at 0.965, and the average accuracy (AA) increased from 0.914 to 0.921, indicating that BN, as a hardware-friendly alternative, can significantly reduce computational complexity without compromising performance. After introducing the frequency domain modeling branch SSFE, despite only adding 0.002M parameters, the model achieved an OA of 0.972, an AA of 0.916, and a Kappa coefficient of 0.937, verifying the effectiveness of frequency domain information in hyperspectral image classification. By further combining HFViT with SSFE, the performance of the model reached the optimal level. AA was significantly improved from 0.921 to 0.942, and Kappa was improved from 0.922 to 0.950, reflecting the synergistic gain effect between the two modules.

[0045] Furthermore, to verify the feasibility of the model's deployment on an FPGA, the HFViT and SSFE modules were each quantized to 8-bit integers, denoted as HFViT* and SSFE*, respectively. Experimental results show minimal performance degradation in the quantized submodules. Simultaneously quantizing both modules results in only 0.3% and 0.2% decreases in AA and Kappa, respectively. These results demonstrate that the proposed model exhibits good quantization robustness.

[0046] Table 2 Ablation and quantification experiments

[0047] Accelerator cross-platform experiment. To verify the performance of the processor of the present invention in terms of inference efficiency and power consumption control, the inference speed and power consumption were compared on the Intel i9-13900Hx CPU, RTX 4090 Laptop GPU and the FPGA accelerator of the present invention based on the ZU19EG platform. The experimental results are shown in Table 3. The CPU platform is limited by the serial computing structure and has the slowest inference speed. Although the GPU has a large number of floating-point units and can achieve the fastest inference speed, its power consumption is high and it is not suitable for edge deployment. In contrast, the FPGA accelerator proposed in the present invention achieves a good balance between inference latency and power consumption. Its average inference time is 75.95% of that of the GPU, while the power consumption is only 14.42W, which is significantly lower than that of the GPU and CPU. In terms of energy efficiency, the inference efficiency of the system of the present invention per unit power consumption is 6.24 times higher than that of the GPU and 63 times higher than that of the CPU, fully demonstrating its deployment advantages and energy efficiency in edge-side hyperspectral processing tasks.

[0048] Table 3 Comparison of cross-platform inference efficiency and power consumption

[0049] The present invention provides a sensor-computing integrated spectral camera processor system for edge real-time classification, which achieves a breakthrough in hyperspectral classification technology on the edge side through the three-in-one innovation of hardware architecture, lightweight network and dedicated accelerator. At the hardware level, the three-level board architecture of core board, interface board and CMOS board deeply integrates spectral acquisition, real-time processing and precise control capabilities: the core board integrates pre-processing modules and accelerators to directly process the 400-700nm band data collected by the CMOS board, and the interface board strictly controls the multi-voltage domain timing through the STM32 microcontroller to ensure the stable operation of the system in the complex electromagnetic environment of the production site. At the algorithm level, the lightweight PBViT network has an original dual-branch collaborative architecture - the frequency domain coding branch SSFE converts frequency domain filtering into cyclic matrix multiplication to avoid the resource bottleneck of traditional frequency domain transformation; the hardware-friendly ViT branch HFViT replaces layer normalization with batch normalization to eliminate the delay in real-time statistical calculation; combined with the 8-bit full integer quantization strategy, the token-level input quantization retains spatial-spectral details, and the channel-level weight quantization maintains the expression ability, while ensuring high-precision classification, significantly compressing the model scale. At the acceleration level, the PBViT hardware accelerator uses a time-overlapping scheduling strategy to parallelize SSFE matrix multiplication within the HFViT Softmax calculation window, effectively hiding branch computation latency. The outer product systolic array architecture, combined with dual 8-bit operand packing technology, achieves an order of magnitude increase in matrix multiplication computing power. The unified nonlinear computation core integrates batch normalization (BN), quantization, Softmax, and GELU into a time-division multiplexing unit, significantly reducing logic resource consumption. The system ultimately achieves a closed-loop "collection-computation-control" process for spectral classification: spectral data is processed on-chip in real time, and the classification results are directly connected to sorting equipment via UART, providing a domestically developed, high-performance edge computing solution for the intelligent manufacturing of pharmaceutical and other chemical substances.

[0050] The above is a detailed introduction to the sensor-computing integrated spectral camera processor system for edge real-time analysis provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the core idea of ​​the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.

Claims

1. A sensor-computing integrated spectral camera processor system for edge real-time analysis, characterized in that: Includes hardware processing system, lightweight hyperspectral classification network PBViT and PBViT hardware accelerator; The hardware processing system includes a core board, an interface board, and a CMOS board. The core board integrates a programmable logic unit (FPGA) and a processing system (PS), and connects to the CMOS board via a MIPI CSI-2 interface to collect hyperspectral data. The interface board is equipped with an STM32 microcontroller to manage multi-voltage domain power supply timing and external interfaces. The CMOS board is equipped with a CMOS sensor. The lightweight hyperspectral classification network PBViT takes as input a spatial patch generated from the spectral image of the object being measured and outputs the category prediction result, which is output to an external display system via the UART interface on the PS side. The network includes a frequency domain encoding branch SSFE and a hardware-friendly ViT branch HFViT, which are executed in parallel. The frequency domain encoding branch converts frequency domain filtering into circulant matrix multiplication, and the hardware-friendly ViT branch replaces layer normalization with batch normalization. The network is trained with 8-bit quantization awareness to achieve full integer quantization. The PBViT hardware accelerator integrates linear computing cores, nonlinear computing cores and residual path cores, and executes SSFE branches in parallel during the Softmax calculation of HFViT through a time overlapping scheduling strategy. It adopts a streaming matrix multiplication architecture based on outer products and time-division multiplexing BN, quantization, Softmax and GELU operations through a unified nonlinear computing core.

2. The system according to claim 1, wherein: The core board of the hardware processing system has a built-in spectral pre-processing module and PBViT accelerator, is equipped with a DDR4 cache array and a 10G Ethernet communication interface, and uploads spectral data to the host computer via the UDP protocol; The interface board provides RJ45 Ethernet port, JTAG debugging interface, STLINK interface, USB to UART interface, 12V power supply interface and SD card slot; The CMOS board uses the PCA9306 level conversion chip to achieve compatible communication between the 1.8V I²C and the core board's 1.2V I / O Bank.

3. The system according to claim 2, characterized in that The data collection of the core board includes: Use DPHY resources to decouple LP and HS signals, and use IDELAY primitives to achieve phase alignment of clock and data; Byte alignment is achieved through Bitslip technology, and multi-channel synchronization is achieved with the help of SoT packets; The spectral preprocessor module adopts a pipeline architecture to complete pixel merging, whiteboard correction and spectrum line smoothing operations; The pre-processed spectral lines are written into the DDR4 chip cache through the built-in DDR controller and assembled into a 9×9 spatial patch. The patch is represented as: ; in, Represents a spectral curve with a dimension of 63. Indicates the number of spectral curves contained in the spatial patch.

4. The system according to claim 3, characterized in that The STM32 microcontroller on the interface board monitors the Power-Good signals of each voltage regulator and outputs Enable signals according to the preset timing to control the power-on sequence of the core board and CMOS board; the integrated buck regulator converts the 12V input into multiple voltage domains for power supply.

5. The system according to claim 4, characterized in that In the frequency domain coding branch, assuming the input image block is ,in, represents the spatial dimension, is the spectral number, and this branch performs the following frequency domain filtering operations in the spatial and spectral dimensions in turn: ; in, is a conjugate symmetric learnable filter, represents element-wise multiplication, is the Fourier transform, is the inverse Fourier transform, is the input image block, is the image block output after frequency domain filtering; According to the convolution theorem, the above formula can be equivalently converted into matrix multiplication form: ; in, Indicates based on This transformation is used during the network's inference process, and the standard matrix multiplication circuit can be reused.

6. The system according to claim 5, characterized in that Each ViTBlock of HFViT consists of the following two substructures: ; ; in, For input, is the intermediate output after multi-head attention, For the final output, represents the batch normalization operation, is multi-head self-attention, It is a two-layer feed-forward network with GELU activation.

7. The system according to claim 6, characterized in that The quantization strategy of the quantization-aware training process includes token-level quantization of input activations, layer-level quantization of intermediate activations, and channel-level quantization of weights. Quantization perturbations are introduced during quantization-aware training to adapt the model to the errors caused by quantization. Quantization adopts a symmetric linear strategy: ; in is a floating point number, is the scaling factor obtained by the min-max algorithm, To quantize the bit width, the clip function limits the output to the representable range. Used to round input to an integer.

8. The system according to claim 7, characterized in that The specific time overlap scheduling strategy is: During the HFViT branch’s Softmax execution, the linear computation module is idle due to the delay in its exponential computation and normalization operations. The cyclic matrix multiplication of the SSFE branch is executed in parallel during this idle window. The idle GELU path DSP resources in the nonlinear computing core are used to process the BN operation of SSFE.

9. The system according to claim 8, characterized in that The linear computation core uses a 32×32 outer product systolic array; each cycle takes a column of the input matrix A and a row of the input matrix B, and directly computes the corresponding outer product in the processing array. Two multiplications are performed in parallel within a single DSP48E2 unit using a dual 8-bit operand packing strategy; Arrays support streaming data dispatch where results propagate right or down.

10. The system according to claim 9, characterized in that The unified nonlinear computing core performs batch normalization, quantization, softmax, and GELU operations by time-division multiplexing hardware resources. BN and quantization are integrated into channel-level multiplication-addition operations, and quadratic polynomial integer approximation is used for softmax and GELU.

Citation Information

Patent Citations

  • Stainless steel weld defect detection method based on multi-domain expression data enhancement and model self-optimization

    CN113129266A

  • Hyperspectral image recognition method based on deep sequence convolutional network

    CN115830461A

  • Signal processing techniques

    CN120151870A

  • Method for tracking and characterizing perishable goods in a store

    US20200111053A1

Cited By

  • Sensing and calculating integrated visual positioning method and system for robot target positioning

    CN120985682A

  • A sensor-computer integrated visual localization method and system for robot target localization

    CN120985682B

  • Visual language model quantification method, device, equipment, medium and program

    CN122088707A