Light and small sensing and computing integrated hyperspectral imaging and target positioning system

By integrating multispectral imaging, lidar, and aerospace-grade AI processor into a lightweight sensing and computing system, autonomous identification and rapid positioning of deep space exploration targets have been achieved, solving the problems of large response delay and low efficiency in existing technologies, and improving detection efficiency and adaptability.

CN121978031APending Publication Date: 2026-05-05NAT SPACE SCI CENT CAS +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NAT SPACE SCI CENT CAS
Filing Date
2026-03-12
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies rely on ground-based manual interpretation of remote sensing data in deep space exploration, which suffers from problems such as large response delays, low efficiency, and high costs. It is difficult to achieve autonomous identification and rapid positioning of exploration targets. Furthermore, existing intelligent identification algorithm models are highly complex and have poor adaptability, making them difficult to deploy on lightweight spaceborne platforms.

Method used

It adopts a lightweight, integrated hyperspectral imaging and target localization system that combines multispectral imaging, lidar, aerospace-grade FPGA+AI processor and 2D turntable. Combined with modules for spectral preprocessing, point cloud preprocessing, precise spatiotemporal registration, 3D spectral data fusion, target detection and path planning, it achieves autonomous data processing and target recognition.

Benefits of technology

It has achieved autonomous identification, screening and positioning of exploration targets in orbit, improved the intelligence level of the exploration process, and is suitable for large-scale mobile exploration and long-term on-orbit operation and control of the International Lunar Research Station, meeting the real-time requirements of deep space exploration missions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121978031A_ABST
    Figure CN121978031A_ABST
Patent Text Reader

Abstract

The invention provides a light and small sensing and computing integrated hyperspectral imaging and target positioning system, and the system comprises a multispectral imaging and hyperspectral point detection optical front end which comprises a multispectral camera and a point spectrometer; a laser radar; the aerospace-level FPGA + AI processor is used for processing data output by the multispectral imaging + hyperspectral point detection optical front end and the laser radar, outputting target morphology and mineral types, selecting potential sampling targets and performing path planning; and the two-dimensional turntable is used for bearing the multispectral imaging and hyperspectral point detection optical front end, the laser radar and the aerospace-level FPGA and AI processor, supporting rotation in the horizontal direction and feeding back a rotation angle. The method has the advantages that identification, screening and positioning of the detection target can be automatically completed in orbit, approaching detection can be carried out without depending on a ground control instruction, and the intelligent level of a detection process point is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of planetary surface science research, specifically involving a lightweight, integrated sensing and computing hyperspectral imaging and target positioning system, which can be used for target selection and positioning on the surfaces of exoplanets such as the Moon and Mars. Background Technology

[0002] Missions such as deep space exploration, lunar research station construction, and Mars rover exploration often require efficient and accurate selection of scientific exploration targets and autonomous positioning. Current technologies largely rely on manual interpretation of remote sensing data from the ground, resulting in problems such as large response delays, low efficiency, and high costs, making it difficult to meet the real-time and autonomous requirements of deep space exploration missions. Although some intelligent recognition algorithms are applied to planetary surface feature extraction, they generally suffer from high model complexity, poor adaptability, and difficulty in deploying on lightweight spaceborne platforms.

[0003] Existing exploration missions mainly rely on the full participation of ground-based scientists in their operation and control modes. This closed-loop operation and control mode between space and ground is inefficient and results in scattered targets, leading to slow movement of mobile probes and low detection efficiency. There is a lack of lightweight, compact, integrated hyperspectral imaging and target positioning systems that can quickly identify and locate targets in orbit and guide autonomous path planning.

[0004] Therefore, there is an urgent need for a highly integrated, low-power, autonomous, lightweight intelligent system to improve the accuracy and real-time performance of target identification and positioning on the surface of extraterrestrial objects, and to support subsequent scientific exploration and resource utilization missions. Summary of the Invention

[0005] The purpose of this application is to overcome the shortcomings of existing technologies that rely heavily on manual interpretation of remote sensing data on the ground, resulting in large response delays, low efficiency, and high costs.

[0006] To achieve the above objectives, this application proposes a lightweight, compact, integrated sensing and computing hyperspectral imaging and target localization system, the system comprising: Multispectral imaging + hyperspectral point detection optical front end, including multispectral camera and point spectrometer; LiDAR; An aerospace-grade FPGA+AI processor is used to process the data output from the multispectral imaging + hyperspectral point detection optical front-end and lidar, outputting target morphology and mineral types, selecting potential sampling targets, and performing path planning; and The two-dimensional turntable is used to support the multispectral imaging + hyperspectral point detection optical front end, lidar and aerospace-grade FPGA + AI processor. It supports horizontal rotation and provides feedback on the rotation angle.

[0007] As an improvement to the aforementioned system, the aerospace-grade FPGA+AI processor integrates a spectral preprocessing module, a point cloud preprocessing module, a spatiotemporal precise registration module, a three-dimensional spectral data fusion module, a target detection module, and a path planning module; among which... The spectral preprocessing module is used to calculate and compensate for the data generated by the multispectral imaging + hyperspectral point detection optical front end, including non-uniformity correction, bad pixel replacement, phase error compensation and fast Fourier transform. The point cloud preprocessing module is used to process the data generated by the lidar, including distance calculation, intensity calculation, point cloud generation, noise suppression, and motion compensation. The spatiotemporal precise registration module is used to perform spatiotemporal registration on the data processed by the spectral preprocessing module and the point cloud preprocessing module; The three-dimensional spectral data fusion module is used to process the data generated by the spatiotemporal precise registration module, and to perform three-dimensional spatial coordinate calculation, spectrum reconstruction and surface normal calculation for each point; The target detection module is used to perform target detection, material identification, anomaly marking and priority sorting on the data generated by the three-dimensional spectral data fusion module, and output the target morphology, mineral type and select potential sampling targets; The path planning module is used to plan the travel path based on the spatial location information of the sampled target output by the target detection module.

[0008] As an improvement to the aforementioned system, the object detection module includes a mapping layer, a Fourier transform layer, a feature extraction network layer, an inverse Fourier transform layer, a linear layer, a cross-attention fusion layer, and a recognition and output priority layer; among which, The mapping layer includes a parallel hyperspectral mapping layer and a lidar mapping layer; the hyperspectral mapping layer performs linear transformation on the data generated by the multispectral imaging + hyperspectral point detection optical front end; the lidar mapping layer performs linear transformation on the data output by the lidar. The Fourier transform layer includes a parallel first hyperspectral branch and a first lidar branch; the first hyperspectral branch performs a one-dimensional Fourier transform on the data output from the hyperspectral mapping layer; the first lidar branch performs a two-dimensional Fourier transform on the data output from the lidar mapping layer. The feature extraction network layer includes a parallel second hyperspectral branch and a second lidar branch; each branch contains N stacked residual units and parallel residual connections, and each residual unit includes a KANLinear subunit and a depthwise separable convolutional subunit; the second hyperspectral branch and the second lidar branch extract features from the outputs of the first hyperspectral branch and the first lidar branch, respectively. The inverse Fourier transform layer includes a parallel third hyperspectral branch and a third lidar branch; the third hyperspectral branch performs a one-dimensional inverse Fourier transform on the output of the second hyperspectral branch; the third lidar branch performs a two-dimensional inverse Fourier transform on the output of the second lidar. The linear layer includes a parallel fourth hyperspectral branch and a fourth lidar branch; the two branches perform 1×1 convolution processing on the outputs of the third hyperspectral branch and the third lidar branch respectively, and unify the number of channels; The cross-attention fusion layer is used to fuse the outputs of the fourth hyperspectral branch and the fourth lidar branch, and then perform feature flattening, attention calculation and feature reconstruction. The output priority layer is used to perform 1×1 convolution processing on the output of the cross-attention fusion layer, mapping the number of channels to the number of classification categories. After global average pooling, the classification probability is output using an activation function to obtain the recognition result. The recognition result and external expert knowledge are connected to a sub-linear layer, and the mapped features are fused element-wise to output a fused feature vector. The fused feature vector is connected to another sub-linear layer, mapped to a priority score vector, and the priority result is output.

[0009] As an improvement to the above system, the hyperspectral mapping layer takes a hyperspectral image with a dimension of 9×9×64 as input, performs a 1×1 convolution to map the number of channels from 64 to 100, and outputs a dimension of 9×9×100. The lidar mapping layer takes a lidar intensity image with a dimension of 9×9×1 as input, performs a 1×1 convolution to map the number of channels from 1 to 100, and outputs a dimension of 9×9×100.

[0010] As an improvement to the above system, the first hyperspectral branch performs a one-dimensional Fourier transform on the 9×9×100 input along the spectral channel dimension, downsampling the spectral dimension from 100 to 50, while keeping the spatial dimension unchanged, and the output dimension is 9×9×50. The second lidar branch performs a two-dimensional Fourier transform along the spatial dimension on a 9×9×100 input, downsampling the spatial dimension from 9×9 to 5×5 while keeping the number of channels unchanged, and the output dimension is 5×5×100.

[0011] As an improvement to the above system, the KANLinear subunit includes a SiLU activation function and a Linear layer connected in sequence, as well as a parallel Spline mapping layer; The depth-separable convolutional subunit includes a depthwise convolutional layer with a kernel size of 3×3 and a pointwise convolutional layer.

[0012] As an improvement to the above system, the third hyperspectral branch performs an inverse transformation along the channel dimension of the 9×9×50 input, restoring the channel dimension from 50 to 100, and the output dimension is 9×9×100. The third lidar branch performs an inverse transformation along the spatial dimension of the 5×5×100 input, restoring the spatial dimension from 5×5 to 9×9, while maintaining 100 channels, and the output dimension is 9×9×100.

[0013] As an improvement to the above system, the fourth hyperspectral branch performs a 1×1 convolution on a 9×9×100 input, mapping the number of channels from 100 to 96, maintaining the spatial dimension of 9×9, and outputting a dimension of 9×9×96. The fourth lidar branch performs a 1×1 convolution on a 9×9×100 input, mapping the number of channels from 100 to 96, maintaining the spatial dimension of 9×9, and outputting a dimension of 9×9×96.

[0014] As an improvement to the above system, the feature flattening is to transform the 9×9 spatial dimension into a one-dimensional sequence of length 81, forming a sequence feature with a dimension of 81×96, thus completing the conversion from spatial dimension to sequence dimension; The attention calculation involves linearly projecting the Query and Key of the sequence features into an 81×12 Q and K matrix, calculating the attention score, and then multiplying it by an 81×12 Value matrix to obtain an 81×96 attention-weighted feature. The feature reconstruction restores the 81×96 attention-weighted features to 9×9×96 spatial features, adds them element-wise to the original hyperspectral features and normalizes them using LayerNorm, outputting a 9×9×96 fused feature.

[0015] As an improvement to the above system, it also includes: The thermal control module, which includes multiple distributed heating elements and temperature measuring circuits, monitors and heats the system under the control of the aerospace-grade FPGA+AI processor.

[0016] Compared with existing technologies, the advantages of this application are: 1. Compared with existing technologies, this application can autonomously complete the identification, screening and positioning of detection targets in orbit, and can carry out close-range detection without relying on ground control commands, which significantly improves the intelligence level of the detection process. 2. Compared with existing technologies, this application is significantly more suitable for the application needs of large-scale mobile exploration intelligent agents, long-cycle on-orbit operation and control, and highly autonomous deep space exploration missions of the International Lunar Research Station; 3. This application represents the cutting-edge level of spectral imaging technology, lidar technology, and artificial intelligence technology in extraterrestrial exploration applications. Through key technologies such as solid-state modulation, dual-channel optimization, high-throughput design, and intelligent preprocessing, it not only meets basic detection requirements but also aims to create a high-performance, stable, and reliable "chemical vision eye" for the next generation of intelligent rover with autonomous scientific discovery capabilities. Attached Figure Description

[0017] Figure 1 The diagram shows the structure of a lightweight, compact, integrated sensing and computing hyperspectral imaging and target localization system. Figure 2 The diagram shown is a logic diagram of a lightweight, compact, integrated sensing and computing hyperspectral imaging and target localization system. Figure 3 The diagram shows the data processing flowchart of a lightweight, compact, integrated sensing and computing hyperspectral imaging and target localization system. Figure 4 The diagram shown is an architecture diagram of the target detection module. Detailed Implementation

[0018] The technical solution of this application will be described in detail below with reference to the accompanying drawings.

[0019] This application provides a lightweight, compact, integrated hyperspectral imaging and target localization system. Employing advanced engineering integration design technology, it integrates a multispectral camera, a point spectrometer, a lidar, an active scanning turntable, and a space computing chip into a single unit. This enables multispectral imaging of targets within the field of view, fusion of hyperspectral point detection data for intelligent spectral reconstruction, and, combined with target elevation information acquired by lidar and a spectral physics analysis model, target composition classification and localization. It also guides the mobile platform in path planning, conducts close-range detailed investigations of high-value targets, and replaces ground-based scientists in conducting fully autonomous, efficient on-orbit detection.

[0020] like Figure 1 As shown, the hardware of the lightweight and compact integrated hyperspectral imaging and target positioning system provided in this application includes: a multispectral imaging + hyperspectral point detection optical front end, a small lidar 1, a two-dimensional turntable 4, an aerospace-grade FPGA + AI processor, and a thermal control module configured for high-performance operation in short and medium waves.

[0021] Among them, the multispectral imaging + hyperspectral point detection optical front end, the small lidar 1, and the aerospace-grade FPGA + AI processor are all integrated into the two-dimensional turntable 4, forming a highly integrated sensing and computing device.

[0022] This system's spectral imaging channels cover 10 narrowband channels within the 0.4-1.5μm range, and employs grating beam splitting technology to achieve continuous output of the target point's spectrum. After acquiring the morphological information corresponding to the 10 narrowband spectra of the target, combined with the target point cloud data from the lidar, the system can completely establish the target's three-dimensional morphological information. Then, by combining the continuous spectral data of the target point, the system performs full-field spectral reconstruction to obtain mineral abundance distribution maps, three-dimensional geological models, and extract potential sampling and detection targets. By acquiring turntable pointing information and lidar positioning information, coordinate transformation is performed to form the path planning orientation for the autonomous navigation system, guiding the platform to approach the target for detailed investigation.

[0023] The two-dimensional turntable 4 is used to carry the optical front end of multispectral imaging + hyperspectral point detection, the small lidar 1, and the aerospace-grade FPGA + AI processor. It can rotate clockwise or counterclockwise in the horizontal direction and accurately feedback the rotation angle.

[0024] The multispectral imaging + hyperspectral point detection optical front end includes a multispectral camera 2 and a point spectrometer 3.

[0025] The multispectral camera 2 employs a filter wheel carrying 10 filters of different narrow bands, which are switched by rotation to acquire two-dimensional images of different spectral bands on the focal plane. Simultaneously, continuous spectral output is generated for the target point using grating beam splitting technology. The instrument includes a front-end optical module, a built-in calibration module, beam splitting and modulation components, and a focal plane detector.

[0026] The multispectral camera 2 uses existing mature products, and its technical parameters are as follows: 1) Spectral channels: The spectral range is not less than 480nm~1000nm, with 8 spectral bands and 8 channel spectral ranges, as shown in Table 1.

[0027] Table 1 8-channel spectral range

[0028] 2) Imaging distance: 0.5m to +∞. 3) Field of view: not less than 6°×6°. 4) Number of effective pixels: not less than 2048×2048. 5) Pixel angular resolution: ≤0.057mrad. 6) Quantization bits: ≥10 bits. 7) Dynamic range: ≥35dB (@550nm). 8) Image signal-to-noise ratio: ≥40dB (@550nm, target reflectivity 0.2, solar altitude angle 30°). 9) MTF: ≥0.2 (@550nm, target reflectivity 0.2, solar elevation angle 30°, Nyquist frequency, full field static transfer function). 10) Number of focus levels: ≥6. 11) Field distortion: ≤1% (@550nm, full field of view). 12) Stray light coefficient: ≤6% (@550nm).

[0029] The point spectrometer 3 also uses existing mature products, and the parameters are shown in Table 2.

[0030] Table 2. Parameters of the Point Spectrometer

[0031] The small lidar 1 generates pulsed laser light through a laser emission module and completes beam shaping. The laser light is diffused to a 20°×20° field of view through the emission optical module to illuminate the target area. The emitted light is received by the receiving optical module and forms response data in the detector array.

[0032] Small LiDAR is also a relatively mature technology product; you can refer to SLAMTEC's RPLIDAR S3 series products.

[0033] The aerospace-grade FPGA+AI processor integrates a spectral preprocessing module, a point cloud preprocessing module, a spatiotemporal precise registration module, a three-dimensional spectral data fusion module, a target detection module, and a path planning module.

[0034] The spectral preprocessing module performs calculations and compensation corrections on the raw data generated by the multispectral imaging + hyperspectral point detection optical front end. Its functions include non-uniformity correction, bad pixel replacement, phase error compensation, fast Fourier transform (according to the requirements of the spectroscopic scheme), radiometric calibration application, spectral channel switching, mode switching, control, data acquisition, autonomous fault and autonomous safety management, and outputs a calibration spectral data cube.

[0035] Non-uniformity correction: This involves data correction to address the inconsistencies in response between pixels of the photodetector. Using relative radiometric calibration results from a ground-based laboratory, real-time detection data is corrected in orbit. Relative radiometric calibration uses a uniformly radiating target as a reference to calibrate the flat-field matrix of the spectral camera, eliminating inconsistencies in the response output between detector pixels. Absolute radiometric calibration uses a standard radiation source as a reference to measure the quantitative relationship between the radiance of the instrument's input signal and the DN value of the output image. It also calibrates the relative changes in output between various states of the multispectral camera 2 (including exposure time, spectral channels, and focusing position), providing a fundamental basis for establishing a mathematical model for the system's target radiance inversion.

[0036] Bad Pixel Removal: By acquiring multiple frames of images under dark field (no light) and uniform bright field (calibration plate illumination) conditions, pixels whose signals are still close to the dark field level under bright field, have extremely low response, or whose signals are much higher than the average value under dark field and continue to "emit light" are identified. Spatial neighborhood interpolation (using the values ​​of surrounding normal pixels to estimate the value of bad pixels) and other methods are used to remove the directly read pixels and restore them to the calculated pixel values ​​to ensure the integrity of the image.

[0037] Phase error compensation: The core method of phase error compensation is the Mertz method, supplemented by the Forman method and the optimized sampling / calculation phase difference method. The main purpose is to eliminate spectral distortion caused by asymmetric sampling of the interferogram or optical path difference offset.

[0038] Fast Fourier Transform (FFT) converts the original interference fringes into standardized spectral curves that can be intelligently interpreted by AI in real time. Its efficient implementation is the key to whether the entire system can operate in real time.

[0039] Applications of radiometric calibration: Solving the problem of "how bright light does the value seen by the camera actually represent?" This is crucial for the accurate identification of material composition, transforming the camera from a "simple recorder" into a "physical quantity measuring instrument." It forms the data foundation for all subsequent intelligent interpretation and quantitative scientific analysis. Without it, the obtained data are merely "pictures" that cannot be scientifically compared.

[0040] Spectral channel switching: Multiple filters with different spectral bands are mounted on a wheel inside the multispectral camera 2. By rotating the wheel, the filters are switched and images are formed, similar to a revolver.

[0041] Mode switching: Switching between various working modes of the instrument, such as normal detection, instrument self-test, manual and automatic modes, etc. The specific switching method is injected by ground commands received by the mobile platform control computer.

[0042] Control: refers to setting instructions for the camera's operation, such as imaging and setting parameters.

[0043] Data acquisition: refers to the process by which the camera's internal main control chip reads data from the camera's focal plane detector array.

[0044] Autonomous fault diagnosis: refers to the software's ability to collect internal signal states of the camera to determine whether the camera has experienced hardware or software malfunctions, and autonomously complete fault reset, recovery, or shutdown protection according to a pre-defined fault handling plan.

[0045] The point cloud preprocessing module performs distance calculation, intensity calculation, point cloud generation, noise suppression, motion compensation, transmission and reception control, mode switching, data acquisition, autonomous fault management, and autonomous safety management on the lidar data channel, and outputs 3D point cloud, intensity map, and digital elevation model (DEM).

[0046] Distance Calculation: The lidar employs the direct time-of-flight method for ranging. A pulsed laser emits a specific spectral band of light, which is reflected by the target and received by a detector array. Each detector pixel independently records the precise round-trip time of the laser. The core component is the time-to-digital converter (TDC), which acts like an ultra-precise stopwatch, capable of resolving picosecond-level time differences (1 picosecond corresponds to a 0.15 mm distance change). The distance calculation formula is: Distance = (Speed ​​of light × Time of flight) / 2. Noise is suppressed through averaging multiple measurements and temperature compensation, ultimately generating a distance image with centimeter-level accuracy, which is strictly synchronized and aligned with the hyperspectral image.

[0047] Intensity Calculation: Intensity calculation for the lidar is performed simultaneously with ranging. The intensity of the echo light signal (photon count or electrical signal amplitude) received by each pixel of the detector is converted from analog to digital to obtain the raw intensity value. Then, background noise and dark current are subtracted, and normalization is performed using laboratory-calibrated parameters (such as laser emission power, optical transmittance, and detector responsivity) to finally generate an intensity image reflecting the target's reflectivity. This image is rigorously registered with the 3D point cloud and can assist in the classification of materials and terrain analysis using hyperspectral data.

[0048] Point Cloud Generation: The core of lidar point cloud generation is converting the polar coordinate measurements of each detected pixel into three-dimensional Cartesian coordinates. For each pixel, the calculated distance value, combined with the precise scan angle (determined by the solid-state beam deflection mechanism or the field of view of the array pixels) and the lidar's own pose, yields the three-dimensional coordinates (X, Y, Z) of the target point through geometric calculation. Simultaneously, the reflection intensity value (I) of the corresponding pixel is assigned as an attribute to that point. After denoising, registration, and time synchronization (strictly aligned with the hyperspectral camera), a dense point cloud dataset containing spatial location and reflection attributes is formed, providing a foundation for 3D terrain reconstruction and fusion with spectral data.

[0049] Noise Suppression: LiDAR noise suppression employs a multi-layered approach: First, at the hardware level, background light is suppressed through optical narrowband filtering, and detector cooling reduces thermal noise. During signal processing, multiple pulse accumulations are performed on each pixel to improve the signal-to-noise ratio; temporal filtering (such as moving average) is used to suppress random noise, and spatial filtering eliminates isolated outliers. For extreme deep-space environments, the algorithm dynamically adjusts the threshold and combines background estimation to subtract transient interferences such as cosmic rays in real time. Finally, reliable point cloud and intensity data are output through statistical consistency checks.

[0050] Motion Compensation: LiDAR motion compensation aims to correct point cloud distortion caused by the rovers' own movement. The core method utilizes high-frequency data from its own motion sensors (such as inertial measurement units and wheeled odometers) to accurately record the platform's pose changes during each frame of point cloud acquisition. Through coordinate transformation, the original points acquired at different times are uniformly compensated to the same reference coordinate system. This eliminates point cloud stitching misalignment and stretching caused by platform translation and rotation, ensuring the generation of a continuous and accurate 3D terrain model, and providing a geometrically consistent basis for pixel-level fusion with synchronously acquired hyperspectral data.

[0051] Transmission and Reception Control: The transmission and reception processes of the lidar are precisely synchronized and controlled by the main control unit. After receiving an external trigger command, the transmission module first drives the pulsed laser to generate nanosecond-level light pulses, and simultaneously generates a synchronization clock signal to start the timing circuit. The receiver detector turns on within a preset time window to collect the echo signal reflected from the target. The entire process achieves nanosecond-level synchronization accuracy through hardware circuitry and can dynamically adjust the laser energy and receiver gain according to the target distance to ensure detection efficiency and safety. This strict closed-loop control ensures high-precision spatiotemporal alignment between point cloud data and hyperspectral images.

[0052] Mode Switching: The lidar supports four main modes: Range Mode: switching between close-range high resolution (≤50m) and long-range reconnaissance (≥200m); Frame Rate Mode: switching between fast scanning while moving (5Hz) and high-precision scanning when stationary (1Hz); Trigger Mode: hard synchronization triggering with a hyperspectral camera or independent free acquisition; Energy Saving Mode: dynamically adjusting laser power and sampling density according to the mission stage. The software system can adaptively select the optimal mode based on terrain, lighting, and scientific objectives.

[0053] Data Acquisition: Data acquisition for the LiDAR begins with a synchronous trigger command from the FPGA+AI processor board. A pulsed laser emits a specific spectral beam to illuminate the target, simultaneously initiating precise timing. The detector array receives the reflected echo, and each pixel independently records the arrival time and intensity of photons. The raw time signal is converted to distance using a time-to-digital converter (TDC), and the intensity signal is amplified and converted from analog to digital to obtain reflectivity. Subsequently, the system calculates the polar coordinate measurements (distance, angle) of each pixel in real time into three-dimensional Cartesian coordinates and binds them with the intensity information to form the original point. This data stream immediately undergoes real-time processing such as spatiotemporal filtering, motion compensation, and noise suppression, ultimately generating a registered three-dimensional point cloud and intensity image, which is output through a high-speed interface, strictly synchronized with the hyperspectral data.

[0054] Autonomous fault and autonomous safety management: This refers to the software's ability to collect the status of some signals inside the lidar to determine whether the lidar has experienced hardware or software abnormalities and whether it may cause equipment safety issues. Based on the pre-defined fault handling plan, it can autonomously complete fault reset, recovery, or shutdown protection.

[0055] The spatiotemporal precise registration module is used to perform spatiotemporal registration on the data processed by the spectral preprocessing module and the point cloud preprocessing module.

[0056] The three-dimensional spectral data fusion module is used to calculate the three-dimensional spatial coordinates of each point (combined with the pointing angle information of the two-dimensional turntable 4), reconstruct the spectrum and calculate the surface normal of the data generated by the spatiotemporal precise registration module.

[0057] The target detection module is used to perform target detection, material identification, anomaly marking, and priority sorting on the data generated by the three-dimensional spectral data fusion module, and finally form scientific product output, including the morphology of the target (rocks, fresh impact craters, rubble piles, etc.), mineral types, and selection of potential sampling targets.

[0058] The target detection module adopts a lightweight network design: Since the radiation resistance of existing AI chips is generally not suitable for lunar and planetary environments, this application utilizes a combination of methods such as model architecture, model pruning, knowledge distillation, and model quantization to enable lightweight network models, thereby adapting them to high-reliability aerospace hardware systems, typically represented by the radiation-resistant AI chip Yulong810A. Specific designs include: 1) Lightweight model architecture design This application optimizes the model structure from the source, using core operators and topology connection methods with higher computational efficiency.

[0059] ① Depthwise separable convolution: As a technical foundation, it adopts a strategy of decomposing standard convolution into "depthwise convolution" and "pointwise convolution". This design can significantly reduce model parameters and computational cost, and is the cornerstone of the MobileNet series of models.

[0060] ② Inverted Residual Structure: To overcome the problem of feature information loss in depthwise convolution, an inverted residual structure is introduced, which involves "expansion followed by compression". This structure first increases the dimension of the feature channels and then compresses them again after depthwise convolution. This can retain richer feature information while maintaining lower computational cost, which is the core of advanced architectures such as MobileNetV2.

[0061] 2) Model pruning optimization ① Structured and Unstructured Pruning: Two approaches are studied in parallel: unstructured pruning (fine-grained removal of individual weights) to pursue the ultimate compression ratio; and structured pruning (removal of the entire neuron or channel) to obtain a better speedup on general-purpose hardware. This application selects the optimal pruning strategy based on the characteristics of the available high-reliability deployment platforms.

[0062] ② Iterative pruning and fine-tuning: To ensure the performance of the pruned model, a robust "iterative pruning-fine-tuning" process is adopted. That is, instead of completing all pruning at once, it is carried out in multiple rounds, with fine-tuning training after each round of pruning to maximize the recovery of model accuracy and ensure the reliability of the compression process.

[0063] 3) Knowledge distillation technology This application utilizes a "teacher-student" network paradigm to transfer knowledge from a large model to a compact model.

[0064] ① Soft label learning: This method guides student models to learn the probability distribution (soft labels) output by the teacher's model, rather than the original hard labels. These soft labels contain information about the relationships between categories, which can guide student models to build decision boundaries with greater generalization ability.

[0065] ② Temperature parameter control: In training, the Softmax temperature parameter (T) is introduced and optimized. By adjusting the smoothness of the probability distribution, the "concentration" of knowledge transfer is controlled, thereby more effectively mining and transferring the "hidden knowledge" in the teacher model.

[0066] 4) Model Quantization This application reduces storage overhead and improves inference speed by lowering the numerical precision of model parameters.

[0067] ① Quantization-aware training: As the preferred approach, this method simulates a low-precision computing environment during the forward propagation of model training, allowing the model weights to adapt to the quantization process during the training phase. Compared to post-training quantization, this method can more effectively maintain the model's performance after a decrease in precision.

[0068] ②Inference acceleration optimization: The volume reduction brought about by quantization is certain, but the improvement in inference speed is highly dependent on the target hardware's acceleration support for low-precision computing.

[0069] 5) Automation and Collaborative Optimization Strategies In pursuit of the optimal balance between accuracy and efficiency, this application adopts an advanced approach that integrates automation and multiple technologies.

[0070] ①Neural Architecture Search (NAS): Evaluate and utilize NAS technology to automatically search for the optimal lightweight model architecture under preset constraints (such as latency and number of parameters).

[0071] ② Integration of Technology Pipelines: Design and validate a pipeline for the collaborative operation of multiple technologies. First, prune the original model to remove redundant structures. Then, quantize the simplified model to further compress its size. Finally, use knowledge distillation to improve the final accuracy of the compressed student model with a stronger teacher model, thereby achieving performance preservation under multiple compressions.

[0072] The target detection module constructs an integrated end-to-end architecture of "dual-branch frequency domain feature extraction + KAN-enhanced lightweight residual module + cross-attention multimodal fusion + expert knowledge-driven priority decision-making." It performs differentiated frequency domain transformations on multi-source data, efficiently compressing redundant data information and significantly improving feature extraction efficiency and computational performance. The core residual module uses a KANLinear structure to replace the traditional fully connected layer, and collaborates with depthwise separable convolutions to achieve a lightweight structure design, balancing the model's nonlinear expressive power and edge inference efficiency. The target detection module as a whole consists of a hyperspectral input branch, a LiDAR input branch, a multimodal fusion unit, and an expert knowledge enhancement unit. Details of the topological connections, input / output feature dimensions, and operator combinations at each level are described below: 1) Input layer Hyperspectral input branch: The input is a hyperspectral image patch with dimensions of 9×9×64 (height H=9, width W=9, number of spectral channels C=64), representing a local region with a spatial size of 9×9 and containing 64 spectral bands. LiDAR Input Branch: The input is a LiDAR intensity image patch with dimensions of 9×9×1 (height H=9, width W=9, number of channels C=1), representing single-channel elevation intensity information with a spatial size of 9×9.

[0073] 2) Mapping layer The two branches execute the linear mapping independently, adapting the number of input channels to the subsequent modules: Hyperspectral mapping layer: Performs a linear transformation (1×1 convolution) on the 9×9×64 input, mapping the number of channels from 64 to 100, keeping the spatial dimension unchanged at 9×9, and the output dimension is 9×9×100.

[0074] LiDAR mapping layer: Performs a linear transformation (1×1 convolution) on the 9×9×1 input, mapping the number of channels from 1 to 50, keeping the spatial dimension unchanged at 9×9, and the output dimension is 9×9×100.

[0075] 3) Fourier transform layer Hyperspectral branch: 1D Fourier transform. For a 9×9×100 input, a 1D Fourier transform is performed along the spectral channel dimension, downsampling the spectral dimension from 100 to 50, while keeping the spatial dimension unchanged. The output dimension is 9×9×50.

[0076] LiDAR branch: 2D Fourier transform. For a 9×9×100 input, perform a 2D Fourier transform along the spatial dimension (H×W) to downsample the spatial dimension from 9×9 to 5×5, keep the number of channels unchanged at 100, and output the dimension as 5×5×100.

[0077] 4) Feature extraction network layer It is also divided into a parallel hyperspectral branch and a lidar branch. Each branch contains N stacked residual units (N is a hyperparameter that can be set according to task requirements) and parallel residual connections. Each residual unit includes a KANLinear subunit and a depthwise separable convolutional subunit. Taking the hyperspectral branch as an example (the lidar branch logic is completely the same): KANLinear subunit: Input: 9×9×50 feature map output from the previous stage. Operator Combinations: ① Activated by SiLU function (input and output dimensions remain unchanged: 9×9×50); ② Connect to a Linear layer (1×1 convolution), keep the number of channels at 50, and output 9×9×200; ③ The parallel branch is a spline mapping layer, which performs adaptive spline transformation on the input feature map, and the output dimension is the same as that of the linear layer; ④ Add the Linear layer output and the Spline output element by element to obtain the KANLinear sub-unit output: 9×9×50.

[0078] Depth-separable convolutional subunits: ①Depthwise convolution: Convolution kernel size 3×3 (stride 1, padding=1), one convolution kernel for each input channel, 50 output channels, and maintains a spatial dimension of 9×9; ② Pointwise convolution: 1×1 convolution, with the number of output channels remaining at 50 and the spatial dimension unchanged; Residual connection: The residual feature fusion is achieved by adding the input of the residual unit (the feature map before entering the KANLinear) to the output of the depthwise separable convolution element by element, with an output dimension of 9×9×50.

[0079] 5) Inverse Fourier Transform Layer The inverse operation of the Fourier transform layer restores the spatial dimension: Hyperspectral branch: The 1D inverse Fourier transform of the 9×9×50 input is inversely transformed along the channel dimension, the channel dimension is restored from 50 to 100, and the output dimension is 9×9×100.

[0080] LiDAR branch: 2D inverse Fourier transform transforms the 5×5×100 input inversely along the spatial dimension, restoring the spatial dimension from 5×5 to 9×9, keeping the number of channels at 100, and the output dimension is 9×9×100.

[0081] 6) Linear layer It is also divided into a parallel hyperspectral branch and a lidar branch. The two branches perform linear projection independently, and the number of channels is unified to prepare for subsequent fusion. Hyperspectral branched linear layer: Performs 1×1 convolution on 9×9×100 input, maps the number of channels from 100 to 96, maintains the spatial dimension of 9×9, and outputs a dimension of 9×9×96.

[0082] The linear branch layer of the LiDAR performs a 1×1 convolution on a 9×9×100 input, mapping the number of channels from 100 to 96, maintaining the spatial dimension of 9×9, and the output dimension is 9×9×96.

[0083] 7) Cross-attention fusion layer The 9×9×96 feature maps of the hyperspectral branch (Query) and the lidar branch (Key / Value) are input into the cross-attention fusion layer for feature flattening, attention calculation, and feature reconstruction. Feature flattening: Flatten the spatial dimension of 9×9 to 81 (sequence length), resulting in sequence features of 81×96.

[0084] Attention calculation: ① The Query and Key are linearly projected to generate Q and K matrices (dimension 81×12, divided into 8 attention heads, head dimension d_k=12). ② Calculate the attention score: Softmax(QKT / dk), with a dimension of 81×81; multiply it with the Value matrix (81×12) to obtain the attention-weighted feature, with a dimension of 81×96.

[0085] Feature reconstruction: The weighted features are reconstructed into 9×9×96 spatial features, which are then added element-wise to the original hyperspectral features and normalized using LayerNorm. The output fused features are 9×9×96.

[0086] 8) Identify and output priority layer Identify branches: The fused features are fed into a linear layer (1×1 convolution), and the number of channels is mapped from 96 to the number of classification categories M (e.g., 16 categories for scene classification tasks), with an output dimension of 9×9×M; A 1×1×M vector is obtained by global average pooling (GAP), and then the classification probability is output by softmax activation to obtain the recognition result (including class label and corresponding probability).

[0087] Expert knowledge integration: The recognition results and external expert knowledge are jointly integrated into the linear layer; element-wise weighted fusion is performed on the mapped features to output a fused feature vector.

[0088] Priority output: The fused feature vectors are fed into a linear layer, mapped to priority score vectors, and the priority results are output.

[0089] The path planning module, based on the target classification and spatial location information output by the target detection module, controls the mobile platform, controlled by the navigation unit, to plan a path to approach the target and perform preset scientific operations such as high-precision imaging, elemental composition analysis, or robotic arm sampling. The patrol mobile platform relies on its own GNC system for navigation, using star and solar optical sensors to perform high-precision identification of star charts and the sun, locate its own coordinates, and combine this with onboard navigation cameras, obstacle avoidance cameras, radar, and other equipment for path planning.

[0090] The thermal control module is used to monitor and heat the system temperature. It includes multiple distributed heating elements and temperature measuring circuits, and is controlled by an aerospace-grade FPGA+AI processor.

[0091] Current planetary missions rely on pre-set ground command sequences for fixed-pattern mobile exploration, lacking the ability to adapt to unpredictable changes during movement. Limited by communication latency and windows, ground intervention is also limited, leading to the loss of many high-value targets. The system described in this application allows the onboard computing module to intelligently analyze observed targets in real time during movement, outputting target composition information and location. This enables autonomous, detailed exploration of high-value targets, eliminating the regret of missed opportunities. Furthermore, after acquiring data on new categories of scientific targets, the system can adaptively update the classification model through an incremental learning mechanism, achieving compatible identification of both old and new categories. In future deep space exploration missions, this application empowers the mobile platform with autonomous discovery and continuous learning capabilities, significantly improving scientific exploration efficiency and mission autonomy.

[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit it. Although this application has been described in detail with reference to the embodiments, those skilled in the art should understand that modifications or equivalent substitutions to the technical solutions of this application do not depart from the spirit and scope of the technical solutions of this application, and should all be covered within the scope of the claims of this application.

Claims

1. A lightweight, compact, integrated sensing and computing hyperspectral imaging and target localization system, characterized in that, The system includes: Multispectral imaging + hyperspectral point detection optical front end, including multispectral camera and point spectrometer; LiDAR; An aerospace-grade FPGA+AI processor is used to process the data output from the multispectral imaging + hyperspectral point detection optical front-end and lidar, outputting target morphology and mineral types, selecting potential sampling targets, and performing path planning; and The two-dimensional turntable is used to support the multispectral imaging + hyperspectral point detection optical front end, lidar and aerospace-grade FPGA + AI processor. It supports horizontal rotation and provides feedback on the rotation angle.

2. The lightweight, compact, integrated sensing and computing hyperspectral imaging and target localization system according to claim 1, characterized in that, The aerospace-grade FPGA+AI processor integrates a spectral preprocessing module, a point cloud preprocessing module, a spatiotemporal precise registration module, a three-dimensional spectral data fusion module, a target detection module, and a path planning module; among them... The spectral preprocessing module is used to calculate and compensate for the data generated by the multispectral imaging + hyperspectral point detection optical front end, including non-uniformity correction, bad pixel replacement, phase error compensation and fast Fourier transform. The point cloud preprocessing module is used to process the data generated by the lidar, including distance calculation, intensity calculation, point cloud generation, noise suppression, and motion compensation. The spatiotemporal precise registration module is used to perform spatiotemporal registration on the data processed by the spectral preprocessing module and the point cloud preprocessing module; The three-dimensional spectral data fusion module is used to process the data generated by the spatiotemporal precise registration module, and to perform three-dimensional spatial coordinate calculation, spectrum reconstruction and surface normal calculation for each point; The target detection module is used to perform target detection, material identification, anomaly marking and priority sorting on the data generated by the three-dimensional spectral data fusion module, and output the target morphology, mineral type and select potential sampling targets; The path planning module is used to plan the travel path based on the spatial location information of the sampled target output by the target detection module.

3. The lightweight, compact, integrated sensing and computing hyperspectral imaging and target localization system according to claim 2, characterized in that, The object detection module includes a mapping layer, a Fourier transform layer, a feature extraction network layer, an inverse Fourier transform layer, a linear layer, a cross-attention fusion layer, and a recognition and output priority layer; among which, The mapping layer includes a parallel hyperspectral mapping layer and a lidar mapping layer; the hyperspectral mapping layer performs linear transformation on the data generated by the multispectral imaging + hyperspectral point detection optical front end; the lidar mapping layer performs linear transformation on the data output by the lidar. The Fourier transform layer includes a parallel first hyperspectral branch and a first lidar branch; the first hyperspectral branch performs a one-dimensional Fourier transform on the data output from the hyperspectral mapping layer; the first lidar branch performs a two-dimensional Fourier transform on the data output from the lidar mapping layer. The feature extraction network layer includes a parallel second hyperspectral branch and a second lidar branch; each branch contains N stacked residual units and parallel residual connections, and each residual unit includes a KANLinear subunit and a depthwise separable convolutional subunit; the second hyperspectral branch and the second lidar branch extract features from the outputs of the first hyperspectral branch and the first lidar branch, respectively. The inverse Fourier transform layer includes a parallel third hyperspectral branch and a third lidar branch; the third hyperspectral branch performs a one-dimensional inverse Fourier transform on the output of the second hyperspectral branch; the third lidar branch performs a two-dimensional inverse Fourier transform on the output of the second lidar. The linear layer includes a parallel fourth hyperspectral branch and a fourth lidar branch; the two branches perform 1×1 convolution processing on the outputs of the third hyperspectral branch and the third lidar branch respectively, and unify the number of channels; The cross-attention fusion layer is used to fuse the outputs of the fourth hyperspectral branch and the fourth lidar branch, and then perform feature flattening, attention calculation and feature reconstruction. The output priority layer is used to perform 1×1 convolution processing on the output of the cross-attention fusion layer, mapping the number of channels to the number of classification categories. After global average pooling, the classification probability is output using an activation function to obtain the recognition result. The recognition result and external expert knowledge are connected to a sub-linear layer, and the mapped features are fused element-wise to output a fused feature vector. The fused feature vector is connected to another sub-linear layer, mapped to a priority score vector, and the priority result is output.

4. The lightweight, compact, integrated sensing and computing hyperspectral imaging and target localization system according to claim 3, characterized in that, The hyperspectral mapping layer takes a hyperspectral image with a dimension of 9×9×64 as input, performs a 1×1 convolution to map the number of channels from 64 to 100, and outputs a dimension of 9×9×100. The lidar mapping layer takes a lidar intensity image with a dimension of 9×9×1 as input, performs a 1×1 convolution to map the number of channels from 1 to 100, and outputs a dimension of 9×9×100.

5. The lightweight, compact, integrated sensing and computing hyperspectral imaging and target localization system according to claim 4, characterized in that, The first hyperspectral branch performs a one-dimensional Fourier transform on the 9×9×100 input along the spectral channel dimension, downsampling the spectral dimension from 100 to 50, while keeping the spatial dimension unchanged, and the output dimension is 9×9×50. The second lidar branch performs a two-dimensional Fourier transform along the spatial dimension on a 9×9×100 input, downsampling the spatial dimension from 9×9 to 5×5 while keeping the number of channels unchanged, and the output dimension is 5×5×100.

6. The lightweight, compact, integrated sensing and computing hyperspectral imaging and target localization system according to claim 5, characterized in that, The KANLinear subunit includes a SiLU activation function and a Linear layer connected in sequence, as well as a parallel Spline mapping layer; The depth-separable convolutional subunit includes a depthwise convolutional layer with a kernel size of 3×3 and a pointwise convolutional layer.

7. The lightweight, compact, integrated sensing and computing hyperspectral imaging and target localization system according to claim 6, characterized in that, The third hyperspectral branch performs an inverse transformation along the channel dimension of the 9×9×50 input, restoring the channel dimension from 50 to 100, and the output dimension is 9×9×100. The third lidar branch performs an inverse transformation along the spatial dimension of the 5×5×100 input, restoring the spatial dimension from 5×5 to 9×9, while maintaining 100 channels, and the output dimension is 9×9×100.

8. The lightweight, compact, integrated sensing and computing hyperspectral imaging and target localization system according to claim 7, characterized in that, The fourth hyperspectral branch performs a 1×1 convolution on a 9×9×100 input, mapping the number of channels from 100 to 96, maintaining the spatial dimension of 9×9, and outputting a dimension of 9×9×96. The fourth lidar branch performs a 1×1 convolution on the 9×9×100 input, mapping the number of channels from 100 to 96, maintaining the spatial dimension of 9×9, and outputting a dimension of 9×9×96.

9. The lightweight, compact, integrated sensing and computing hyperspectral imaging and target localization system according to claim 8, characterized in that, The feature flattening is to transform the 9×9 spatial dimension into a one-dimensional sequence of length 81, forming a sequence feature with a dimension of 81×96, thus completing the conversion from spatial dimension to sequence dimension; The attention calculation involves linearly projecting the Query and Key of the sequence features into an 81×12 Q and K matrix, calculating the attention score, and then multiplying it by an 81×12 Value matrix to obtain an 81×96 attention-weighted feature. The feature reconstruction restores the 81×96 attention-weighted features to 9×9×96 spatial features, adds them element-wise to the original hyperspectral features and normalizes them using LayerNorm, outputting a 9×9×96 fused feature.

10. The lightweight, compact, integrated sensing and computing hyperspectral imaging and target localization system according to claim 1, characterized in that, Also includes: The thermal control module, which includes multiple distributed heating elements and temperature measuring circuits, monitors and heats the system under the control of the aerospace-grade FPGA+AI processor.