Unmanned surface vehicle environment coupling knowledge distillation multi-band fusion target detection method

By synchronously collecting multi-band data on an unmanned surface vessel and calculating inter-frame registration recursive compensation coefficients, feature enhancement weights are generated, and knowledge distillation processing of a multi-teacher network is realized. This solves the problem of insufficient dynamic linkage mechanism in the multi-band fusion target detection method for unmanned surface vessels, improves detection stability and performance, and is compatible with the real-time operation of edge computing platforms.

CN122368941APending Publication Date: 2026-07-10GUANGZHOU PANGAO LEADER TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU PANGAO LEADER TECH CO LTD
Filing Date
2026-03-13
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing multi-band fusion target detection methods for unmanned surface vessels lack a dynamic linkage mechanism for registration, feature enhancement, and knowledge distillation. This makes them unable to adapt to the dynamic changes in hull motion and environmental interference in complex sea scenarios, resulting in the accumulation of multimodal data registration deviations, low efficiency of effective feature extraction, and insufficient model detection stability. Consequently, they are ill-suited to the real-time operation requirements of edge computing platforms carried by unmanned surface vessels.

Method used

By simultaneously collecting multi-band sensing data, ship motion data, and marine environment data, calculating inter-frame registration recursive compensation coefficients, generating radar and visual feature enhancement weights, and performing knowledge distillation processing on the lightweight student network based on dynamic distillation weights, feature extraction and fusion of multi-teacher networks are realized for target detection.

Benefits of technology

It improves the stability and performance of target detection, adapts to the real-time operation requirements of unmanned surface vessel edge computing platforms, realizes dynamic linkage of parameters throughout the process, and adapts to the dynamic changes of complex sea surface scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122368941A_ABST
    Figure CN122368941A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of environmental perception and target detection technology for unmanned surface vessels. Specifically, it discloses a multi-band fusion target detection method for environmental coupling knowledge distillation on unmanned surface vessels. This method simultaneously collects multi-band perception data, hull motion data, and marine environment data, and sequentially calculates inter-frame registration recursive compensation coefficients, dual-modal feature enhancement weights, and dynamic distillation weights of multi-teacher networks to complete knowledge distillation, multi-modal feature fusion, and target detection using a lightweight student network. This invention achieves dynamic parameter linkage throughout the detection process, improves the stability of target detection in complex sea surface scenarios, and adapts to the deployment requirements of unmanned surface vessel edge computing platforms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of environmental perception and target detection technology for unmanned surface vessels, and specifically discloses a multi-band fusion target detection method for environmental coupling knowledge distillation of unmanned surface vessels. Background Technology

[0002] Existing multi-band fusion target detection methods for unmanned surface vessels (USVs) involve independent processing steps and lack a dynamic linkage mechanism between registration, feature enhancement, and knowledge distillation. This makes them unsuitable for adapting to the dynamic changes in hull motion and environmental interference in complex sea scenarios. Consequently, multimodal data registration bias accumulates continuously, effective feature extraction efficiency is low, and model detection stability is insufficient. Furthermore, they cannot balance model detection performance with lightweight requirements, making them unsuitable for the real-time operation requirements of edge computing platforms carried by USVs.

[0003] Based on the above problems, there is an urgent need for a multi-band fusion target detection technology solution that can achieve dynamic linkage of parameters throughout the entire process and adapt to complex sea surface scenarios. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and proposes a multi-band fusion target detection method for unmanned surface vessels, which includes the following steps: S1: Simultaneously collect multi-band sensing data, ship motion data, and marine environmental data; S2: Calculate the inter-frame registration recursive compensation coefficient based on the multi-band sensing data, the ship motion data, and the marine environment data; S3: Based on the inter-frame registration recursive compensation coefficient, generate radar feature enhancement weights and visual feature enhancement weights; S4: Calculate the dynamic distillation weights of the multi-teacher network based on the inter-frame registration recursive compensation coefficients, the radar feature enhancement weights, and the visual feature enhancement weights; S5: Based on the dynamic distillation weights, perform knowledge distillation on the lightweight student network; S6: Based on the lightweight student network, perform feature extraction and feature fusion processing on the multi-band sensing data to obtain fused features; S7: Based on the fused features, perform target detection processing.

[0005] Preferably, the step of simultaneously acquiring multi-band sensing data, ship motion data, and marine environmental data includes simultaneously acquiring millimeter-wave radar point cloud data, visible light image data, ship attitude disturbance time-series change rate data, ship speed data, effective wave height data, and light intensity data.

[0006] More preferably, the step of calculating the inter-frame registration recursive compensation coefficient based on multi-band sensing data, ship motion data, and marine environment data includes calculating the inter-frame registration recursive compensation coefficient based on ship speed data, ship attitude disturbance time-series change rate data, effective wave height data, sensor sampling period, ship reference length, and registration residual deviation of the previous frame.

[0007] More preferably, the step of generating radar feature enhancement weights and visual feature enhancement weights based on the inter-frame registration recursive compensation coefficients includes generating radar feature enhancement weights and visual feature enhancement weights based on the inter-frame registration recursive compensation coefficients, the proportion of radar echo specular reflection clutter, the signal-to-noise ratio of radar effective target signals, the proportion of overexposed pixels in visual images, and the signal-to-noise ratio of visual effective features.

[0008] Further preferably, the inter-frame registration recursive compensation coefficient, the radar feature enhancement weight, the visual feature enhancement weight, and the dynamic distillation weight of the multi-teacher network are calculated through a set of linked calculation formulas. This set of formulas includes an environment-coupled inter-frame registration recursive compensation coefficient formula, a coupled interference joint feature enhancement weight formula, and a time-reliability coupled multi-teacher dynamic distillation weight formula. The environment-coupled inter-frame registration recursive compensation coefficient formula is as follows: ; In the formula, The recursive compensation coefficients for inter-frame registration are used. This refers to the ship's real-time speed. The sensor sampling period, The effective wave height at sea surface. The reference length of the hull. Let be the temporal rate of change of the ship's attitude disturbance. To account for the registration residual error of the previous frame. , , The weighting coefficients are dimensionless and satisfy the following conditions: .

[0009] More preferably, the coupled interference joint feature enhancement weight formula includes a radar feature enhancement weight calculation formula and a visual feature enhancement weight calculation formula. The radar feature enhancement weight calculation formula uses the inter-frame registration recursive compensation coefficient as the reference input, and the visual feature enhancement weight calculation formula uses the inter-frame registration recursive compensation coefficient as the reference input.

[0010] More preferably, the time-series reliability coupled multi-teacher dynamic distillation weight formula includes the original weight calculation formula of the teacher network and the normalized weight calculation formula, wherein the original weight calculation formula of the teacher network uses the inter-frame registration recursive compensation coefficient, radar feature enhancement weight and visual feature enhancement weight as the reference input.

[0011] More preferably, the step of performing knowledge distillation processing on the lightweight student network based on dynamic distillation weights includes performing feature-level distillation processing and output-level distillation processing on the lightweight student network based on the dynamic distillation weights of the multi-teacher network. The multi-teacher network includes a sea clutter suppression teacher network, an illumination change compensation teacher network, and a hull attitude disturbance adaptation teacher network.

[0012] In a further preferred embodiment, the step of performing feature extraction and feature fusion processing on multi-band sensing data based on a lightweight student network to obtain fused features includes extracting radar depth features and visual depth features based on the lightweight student network, performing weighted fusion processing on the radar depth features and visual depth features based on the radar feature enhancement weights and the visual feature enhancement weights to obtain initial fused features, and performing multi-layer feature interaction processing on the initial fused features to obtain fused features.

[0013] Further preferably, the step of performing target detection processing based on fused features includes inputting the fused features into an anchorless target detection head, performing target category prediction processing and bounding box regression processing, outputting target detection results, and updating the calculation parameters of the inter-frame registration recursive compensation coefficient and the generation parameters of the feature enhancement weights for the next frame based on the target average confidence of the target detection results.

[0014] Technical effects: This invention solves the core problem of existing technologies, which operate independently and cannot adapt to complex dynamic changes on the sea surface, by establishing a full-process parameter linkage mechanism of inter-frame registration recursive compensation, dual-modal feature enhancement and multi-teacher dynamic knowledge distillation. It improves the stability of target detection, while taking into account the detection performance and lightweight requirements of the model, and adapts to the real-time operation requirements of the unmanned surface vessel edge computing platform. Attached Figure Description

[0015] Figure 1 This is a flowchart of the multi-band fusion target detection method for unmanned surface vessels coupled with environmental knowledge distillation, as described in this application. Figure 2 This is a flowchart illustrating the calculation of the inter-frame registration recursive compensation coefficients in this application. Figure 3 This is a flowchart for calculating the dynamic distillation weights in the multi-teacher network of this application. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0017] The existing technology has the following technical problems: the processing links of multi-band fusion target detection are independent of each other and no dynamic linkage mechanism has been established. It cannot adapt to the dynamic changes of ship movement and environmental interference in complex sea surface scenarios, resulting in the accumulation of registration deviation, low efficiency of effective feature extraction, insufficient detection stability, and inability to balance model performance and lightweight requirements. It is also difficult to adapt to the real-time operation requirements of edge computing platforms carried by unmanned surface vessels.

[0018] Based on this, please refer to Figures 1-3This embodiment provides a multi-band fusion target detection method for unmanned surface vessels, coupled with environmental knowledge distillation. The complete implementation process has a clear implementation path and repeatability. The core of the method is to establish a dynamic linkage mechanism for all process parameters, achieving collaborative optimization of each processing stage through the coupled calculation of multi-dimensional parameters. The method specifically includes the following steps: synchronously collecting multi-band sensing data, vessel motion data, and marine environment data; calculating inter-frame registration recursive compensation coefficients based on the multi-band sensing data, vessel motion data, and marine environment data; generating radar feature enhancement weights and visual feature enhancement weights based on the inter-frame registration recursive compensation coefficients; calculating the dynamic distillation weights of the multi-teacher network based on the inter-frame registration recursive compensation coefficients, radar feature enhancement weights, and visual feature enhancement weights; performing knowledge distillation processing on a lightweight student network based on the dynamic distillation weights; performing feature extraction and feature fusion processing on the multi-band sensing data based on the lightweight student network to obtain fused features; and performing target detection processing based on the fused features. It is worth mentioning that the steps of the method are sequentially related, with the output parameters of the previous step serving as the core inputs of the next. Furthermore, after completing the target detection processing, a closed-loop update of all parameters is achieved based on the detection results, forming a complete linkage system from data acquisition to target detection. The multi-band sensing data refers to the multi-band detection data of the unmanned surface vessel (USV) against sea surface targets; the hull motion data refers to the USV's own motion state data during navigation; and the marine environment data refers to the marine environment data of the USV's navigation area. The synchronous acquisition of these three types of data provides the foundation for all subsequent parameter calculations. The inter-frame registration recursive compensation coefficient is a core parameter for quantifying the impact of hull motion and marine environment coupling disturbances on multimodal data registration, serving as the benchmark for linking the registration stage with subsequent stages. The radar feature enhancement weight and visual feature enhancement weight are modal feature adjustment parameters adapted to sea surface interference, directly determining the enhancement of radar and visual features. The dynamic distillation weights of the multi-teacher network are parameters that allocate the knowledge distillation ratio of each teacher network according to changes in the sea surface scene, which is the key to realizing lightweight student network-specific knowledge learning; the lightweight student network is a core model for feature extraction and fusion adapted to edge computing platforms, which can balance lightweighting and detection performance after knowledge distillation; the fused features are comprehensive features of radar and visual features after weighted fusion and multi-layer interaction, which can give full play to the complementary advantages of dual modes; the target detection processing is the final target recognition and localization link based on the anchorless detection framework, and the output detection results can directly serve the autonomous navigation decision of unmanned surface vessels.All computations in this method are performed on the edge computing board mounted on the unmanned surface vessel (USV). The computation time for all parameters is controlled within milliseconds, meeting the real-time detection requirements of USVs. Furthermore, all technical features are integrated organically; changes in parameters at any stage trigger dynamic adjustments in subsequent stages, achieving adaptive detection in complex sea surface scenarios. The achieved technical results are as follows: By establishing a dynamic linkage mechanism for parameters throughout the entire process, the problem of independent processing stages is solved. This achieves deep coupling between vessel motion, the marine environment, and the detection algorithm, adapting to the dynamic changes in complex sea surface scenarios. It balances model detection performance with lightweight requirements, and can run stably on the edge computing platform of the USV, improving the stability and effectiveness of sea surface target detection.

[0019] The existing technology has the following technical problems: the multi-source data acquisition stage only collects a single type of sensing data, without simultaneously collecting multi-dimensional parameters of ship motion and marine environment, and the data collected by different devices lacks a time synchronization mechanism, resulting in a lack of comprehensive and unified data sources for subsequent calculations, causing various calculation deviations and affecting the final detection results.

[0020] Based on this, the step of simultaneously acquiring multi-band sensing data, ship motion data, and marine environmental data includes simultaneously acquiring millimeter-wave radar point cloud data, visible light image data, ship attitude disturbance temporal change rate data, ship speed data, effective wave height data, and light intensity data. It is worth noting that the acquisition of this multi-source data is accomplished through a multi-source sensing acquisition module mounted on the unmanned surface vessel. This module includes a millimeter-wave radar, a visible light camera, a ship IMU sensor, a sea clutter monitor, a light intensity sensor, and a BeiDou positioning module. All acquisition devices are designed with miniaturization and low power consumption to meet the installation requirements of the unmanned surface vessel. The millimeter-wave radar and visible light camera are mounted on the same detection plane at the bow of the unmanned surface vessel, ensuring that their perception fields are completely consistent. The millimeter-wave radar has a detection range of 5 to 500 meters and a point cloud acquisition frequency of 20Hz. The visible light camera has a resolution of 1920×1080 and an image acquisition frequency of 20Hz. Both simultaneously acquire radar point cloud and visual image data of sea surface targets. The hull IMU sensor is mounted at the center of gravity of the unmanned surface vessel, which can accurately acquire the hull's roll, pitch, and heave attitude data, and calculate the temporal changes of hull attitude disturbances in real time. The data acquisition system uses a 100Hz frequency. The BeiDou positioning module, installed on the top of the unmanned surface vessel, acquires real-time speed data with a positioning accuracy of 1 meter and a speed acquisition accuracy of 0.1 m / s. The sea clutter monitor, installed on the outer side of the vessel, detects effective wave height data in real-time, with a detection range of 0.1 to 10 meters and a acquisition frequency of 10Hz. The light intensity sensor, installed next to the visible light camera, acquires real-time light intensity data within the camera's field of view, with a detection range of 0.1 to 100,000 lux and a acquisition frequency of 20Hz. All acquisition devices are connected to a time synchronizer to achieve microsecond-level timestamp synchronization, ensuring complete time matching of data acquired by different devices. The synchronized raw data is transmitted to a high-speed data acquisition card. After outlier removal, format conversion, and numerical normalization by the data standardization unit, it is stored in a unified digital signal format in the edge computing board's cache. The cache capacity is 16GB, capable of storing at least 10 minutes of multi-source acquisition data to meet the real-time retrieval requirements of subsequent calculations. Technical effects achieved: Synchronous acquisition and standardized processing of multi-band sensing, ship motion, and marine environment data were realized, providing a comprehensive, unified, and accurate data source for subsequent full-process parameter calculations, and eliminating calculation deviations caused by data asynchrony and inconsistent formats.

[0021] The existing technology has the following technical problems: the calculation of the inter-frame registration recursive compensation coefficient only considers the single ship motion parameter and does not take into account the sea surface environment parameter and the registration residual deviation of the previous frame. As a result, the calculated compensation coefficient cannot accurately reflect the actual coupling disturbance, the registration compensation effect is poor, and the accumulation of inter-frame registration deviation cannot be effectively eliminated.

[0022] Based on this, the step of calculating the inter-frame registration recursive compensation coefficient based on multi-band sensing data, ship motion data, and marine environment data includes calculating the inter-frame registration recursive compensation coefficient based on ship speed data, ship attitude disturbance temporal change rate data, effective sea surface wave height data, sensor sampling period, ship reference length, and registration residual deviation of the previous frame. It is worth noting that the ship speed data is the real-time speed of the unmanned surface vessel output by the Beidou positioning module, a scalar value, directly used as the input parameter for coefficient calculation; the ship attitude disturbance temporal change rate data is the average rate of change of the roll, pitch, and heave attitude data collected by the ship's IMU sensors, calculated using first-order difference, which can comprehensively quantify the dynamic disturbance degree of the ship's attitude; the effective sea surface wave height data is the real-time wave height value output by the sea surface clutter monitor, a core parameter reflecting sea surface environmental disturbance; the sensor sampling period is the time interval for all samples collected. The unified sampling interval of the collection device is set by the time synchronizer, and in this embodiment, it is set to 0.05 seconds, which is a fixed parameter. The hull reference length is the basic length parameter of the unmanned surface vessel, which is preset according to the actual model of the unmanned surface vessel and is a fixed parameter, stored in the local database of the edge computing board. The registration residual deviation of the previous frame is the average deviation of the coordinate mapping between the radar point cloud and the visible light image after the registration processing of the multimodal data of the previous frame. It is output from the registration calculation process of the previous frame and stored in the cache in real time as a historical parameter for the calculation of the current frame. Before the coefficient calculation, the hull attitude disturbance temporal change rate data is subjected to moving average filtering processing with a filtering window of 5 frames, which can effectively remove high-frequency noise in the data and avoid the influence of instantaneous attitude disturbance on the coefficient calculation. At the same time, the registration residual deviation of the previous frame is thresholded. If the deviation value exceeds the preset maximum deviation threshold, it is reset to the threshold average value to avoid calculation distortion caused by extreme deviation values. All input parameters are converted to a uniform dimensionless form before subsequent calculations to ensure dimensional consistency throughout the calculation process. Parameter conversion and preprocessing are handled by the embedded processor of the edge computing board, with a processing time of less than 0.5 milliseconds. The achieved technical effect is as follows: By combining multi-dimensional coupling parameters of ship motion, sea surface environment, equipment parameters, and historical deviations, the inter-frame registration recursive compensation coefficients are calculated. This ensures that the compensation coefficients highly match the actual coupling disturbance situation, improving the accuracy of registration compensation, effectively eliminating the accumulation of inter-frame registration deviations, and providing accurate disturbance reference parameters for subsequent stages.

[0023] The existing technology has the following technical problems: the generation of dual-modal feature enhancement weights does not combine the inter-frame registration recursive compensation coefficients, nor does it consider the co-source interference characteristics of radar specular reflection clutter and visual highlight overexposure. The calculation of radar and visual weights is independent of each other, which makes the generated weights unable to adapt to the actual interference situation, resulting in poor feature enhancement and interference suppression effects.

[0024] Based on this, the step of generating radar feature enhancement weights and visual feature enhancement weights based on the inter-frame registration recursive compensation coefficients includes generating radar feature enhancement weights and visual feature enhancement weights based on the inter-frame registration recursive compensation coefficients, the proportion of specular reflection clutter in radar echoes, the signal-to-clutter ratio of radar effective target signals, the proportion of overexposed pixels in visual images, and the signal-to-noise ratio of visual effective features. It is worth mentioning that the proportion of specular reflection clutter in radar echoes is a parameter quantifying the degree of interference of sea surface specular reflections on radar echoes. It is obtained through frequency domain analysis of point cloud data collected by millimeter-wave radar. Specifically, after performing a fast Fourier transform on the radar echo signal, the ratio of specular reflection clutter energy to total echo energy is calculated; a larger ratio indicates more severe interference from specular reflection clutter. The signal-to-clutter ratio of radar effective target signals is the ratio of radar effective target signal energy to clutter energy, a core parameter reflecting radar signal quality; a larger signal-to-clutter ratio indicates clearer radar effective features. The percentage of overexposed pixels in the visual image is a parameter quantifying the interference of specular reflection from the sea surface on the visual image. It is obtained through grayscale analysis of image data acquired by a visible light camera. Specifically, a grayscale threshold for overexposed highlights is set, and the ratio of pixels with grayscale values ​​exceeding this threshold to the total number of pixels is counted. A larger ratio indicates more severe interference from overexposed highlights. The signal-to-noise ratio (SNR) of the effective visual features is the ratio of signal energy to noise energy of the effective target features in the visual image. It is a core parameter reflecting the quality of visual image features; a larger SNR indicates more obvious effective visual features. All four interference and signal quality parameters are calculated in real-time by the GPU of the edge computing board, with the calculation frequency consistent with the sensor sampling frequency at 20Hz. The inter-frame registration recursive compensation coefficient is the calculation output of the previous stage and is directly used as the benchmark parameter for weight generation, realizing parameter linkage between the registration stage and the feature enhancement stage. When generating weights, the inter-frame registration recursive compensation coefficient is coupled with the four parameters mentioned above for calculation. This ensures that the magnitude of the feature enhancement weights matches both the degree of coupling disturbance and the actual interference situation. A larger inter-frame registration recursive compensation coefficient indicates a more severe coupling disturbance, and the adjustment range of the feature enhancement weights increases accordingly. Similarly, more severe radar clutter or visual specular interference results in a correspondingly larger feature enhancement weight, achieving targeted interference suppression. The achieved technical effect is a high degree of matching between radar feature enhancement weights and visual feature enhancement weights and actual coupling disturbances and sea surface co-source interference. This enables deep linkage between the registration and feature enhancement stages, improves the enhancement effect of radar and visual features, effectively suppresses co-source coupling interference caused by specular reflection, and provides a precise modal weight benchmark for subsequent knowledge distillation.

[0025] The existing technology has the following technical problems: the inter-frame registration recursive compensation coefficient, radar feature enhancement weight, visual feature enhancement weight and dynamic distillation weight of multi-teacher network are calculated independently without establishing a progressive and linked calculation relationship. This makes it impossible to achieve deep coupling of parameters throughout the process, resulting in a lack of coordination in parameter adjustment at each stage and difficulty in adapting to the dynamic changes of complex sea surface scenarios.

[0026] Based on this, the inter-frame registration recursive compensation coefficient, the radar feature enhancement weight, the visual feature enhancement weight, and the dynamic distillation weight of the multi-teacher network are calculated through a set of linked calculation formulas. This set of formulas includes an environment-coupled inter-frame registration recursive compensation coefficient formula, a coupled interference joint feature enhancement weight formula, and a time-reliability coupled multi-teacher dynamic distillation weight formula. The environment-coupled inter-frame registration recursive compensation coefficient formula is as follows: ; It is worth mentioning that the theoretical design of this formula is based on the ship kinematics similarity criterion and the inter-frame registration error propagation law. When an unmanned surface vessel navigates on the sea surface, the inter-frame registration deviation of multimodal data is generated by the combined effect of three coupled disturbance sources: displacement deviation caused by the coupling of ship speed and sea wave height, coordinate mapping deviation caused by dynamic changes in ship attitude, and cumulative propagation of residual deviation from previous frames. Furthermore, the mutual influence of these three disturbance sources leads to a continuous amplification of the deviation. This formula achieves accurate quantification of the actual disturbance level by coupling calculations of these three disturbance sources. The logical derivation of the formula is based on the principle of dimensional homogeneity. First, the parameters of the three disturbance sources with different physical dimensions are respectively processed to be dimensionless, transformed into similarity numbers that can be directly coupled and calculated. Then, a comprehensive disturbance coefficient is obtained through linear weighting. Finally, a negative exponential function is used to map the comprehensive disturbance coefficient to the interval between 0 and 1, resulting in the final inter-frame registration recursive compensation coefficient. The definitions and acquisition methods of each parameter in the formula are as follows: The inter-frame registration recursive compensation coefficient is the output parameter of the formula, and its value ranges from 0 to 1. The larger the value, the higher the degree of coupling disturbance and the greater the strength of registration compensation. The real-time speed of the ship is collected by the BeiDou positioning module and is measured in meters per second. The sensor sampling period is set by the time synchronizer and is measured in seconds. The effective wave height is measured in meters, collected by a sea clutter monitor. The reference length of the hull is preset according to the model of the unmanned surface vessel, and the unit of measurement is meters; The metric is the temporal rate of change of the ship's attitude disturbance, calculated from the attitude data collected by the ship's IMU sensors, with dimensions of 1 / second. The residual error from the previous frame registration is output from the previous frame registration calculation process and is measured in meters. , , The weighting coefficients are dimensionless and satisfy the following conditions: It can be calibrated according to the actual navigation scenarios of unmanned surface vessels, in near-shore navigation scenarios. Set to 0.4 Set to 0.3. Set to 0.3 in the scenario of ocean navigation. Set to 0.3. Set to 0.4 Let it be 0.3. The dimensionless treatment of each term in the formula is as follows: Let be the displacement similarity number coupled between speed and wave height. The displacement of the ship's hull during the sampling period is expressed in meters. The ratio is made dimensionless, quantifying the displacement disturbance coupled between the ship's motion and the sea surface environment; Let be the angular similarity number of the attitude perturbation. and After multiplication, the dimension is 1, which quantifies the coordinate mapping disturbance of the dynamic change in the ship's attitude; The similarity number for deviation propagation, and The ratio is made dimensionless, quantifying the cumulative propagation perturbation of the residual bias of the preceding frame. Three dimensionless similarity numbers are obtained through... , , The linear weighted sum yields the composite perturbation coefficient, which remains dimensionless. This composite perturbation coefficient is the input to the negative exponential function. Since the input to the exponential function is dimensionless, its output is also dimensionless. The calculation result is also dimensionless, ensuring the homogeneity of the entire formula. The core innovation of this formula lies in achieving coupled quantization of multiple disturbance sources, organically integrating the influence of ship motion, sea surface environment, and historical deviations. Simultaneously, through dimensionless processing and negative exponential mapping, the output compensation coefficient is kept within a fixed range, avoiding over- or under-compensation. Furthermore, this coefficient serves as the benchmark parameter for all subsequent weight calculations, achieving deep parameter linkage between the registration stage and subsequent stages. The formula calculation is executed in real-time by the embedded processor of the edge computing board. All input parameters are retrieved from the cache in real-time, with a calculation time of less than 1 millisecond, synchronized with the sensor sampling frequency to ensure real-time coefficient updates. The achieved technical effects include: accurate quantization of multiple coupled disturbance sources; a high degree of matching between the generated inter-frame registration recursive compensation coefficients and the actual disturbance situation; providing a unified benchmark parameter for subsequent weight calculations; achieving parameter linkage between the registration stage and subsequent stages; and laying the foundation for deep parameter coupling throughout the entire process.

[0027] The existing technology has the following technical problems: the input parameters of the joint feature enhancement weight formula for coupled interference are not linked with the inter-frame registration recursive compensation coefficient, the calculation of radar and visual feature enhancement weights is independent of each other, and it is impossible to achieve joint suppression of co-source coupled interference caused by sea surface mirror reflection. The adaptability of feature enhancement is insufficient.

[0028] Based on this, the coupled interference joint feature enhancement weight formula includes a radar feature enhancement weight calculation formula and a visual feature enhancement weight calculation formula. The radar feature enhancement weight calculation formula uses the inter-frame registration recursive compensation coefficient as the reference input, and the visual feature enhancement weight calculation formula also uses the inter-frame registration recursive compensation coefficient as the reference input. It is worth noting that the two calculation formulas of the coupled interference joint feature enhancement weight formula are progressively linked, both using the output parameters of the environmental coupling inter-frame registration recursive compensation coefficient formula. Using the core baseline input, the deep parameter coupling between the registration and feature enhancement stages is achieved. The complete expressions for the two calculation formulas are as follows: ; ; The theoretical design of this formula is based on the joint suppression principle of co-source coupling interference from the sea surface. The specular reflection from the crest of the sea wave simultaneously triggers radar multipath clutter and visual overexposure of highlights. These two are co-source coupling interferences, requiring joint suppression through linked weight calculations. Furthermore, the degree of registration perturbation directly affects the quality of feature extraction; therefore, [the formula is then used]. As a baseline input, the feature enhancement weights are synchronized with the degree of coupling perturbation. The logical derivation of the formula is based on the interference suppression principle of signal processing. First, the registration compensation coefficients are... The interference-to-signal ratio of the modal is coupled to obtain the comprehensive interference coefficient. This coefficient is then mapped to the interval between 0 and 1 using a negative exponential function to obtain the feature enhancement weights. The magnitude of these weights is ensured to be positively correlated with the coupled perturbation and the actual interference level, thus achieving targeted feature enhancement and interference suppression. The definitions and acquisition methods of the parameters in the formula are as follows: The radar feature enhancement weight is an output parameter of the radar feature enhancement calculation formula. Its value ranges from 0 to 1. The larger the value, the stronger the radar feature enhancement and clutter suppression. The visual feature enhancement weight is the output parameter of the visual feature enhancement calculation formula. Its value ranges from 0 to 1. The larger the value, the stronger the visual feature enhancement and highlight suppression. The recursive compensation coefficients for inter-frame registration are calculated and output from the preceding formulas and are dimensionless. The proportion of clutter reflected from the radar echo mirror is obtained from the frequency domain analysis of radar point cloud data. It is dimensionless and ranges from 0 to 1. The signal-to-clutter ratio of the effective target signal of the radar is calculated from radar point cloud data. It is dimensionless and ranges from 0 to positive infinity. The percentage of overexposed pixels in the highlights of a visual image is obtained from the analysis of the grayscale values ​​of the visible light image. It is dimensionless and ranges from 0 to 1. The signal-to-noise ratio (SNR) of the visually effective features is calculated from visible light image data. It is dimensionless and ranges from 0 to positive infinity. In the calculation of the formula... This is the interference-signal integration ratio of the radar. The addition of 1 is to avoid the calculation error of the denominator being 0. The larger the ratio, the more severe the specular reflection clutter interference of the radar and the weaker the radar's effective characteristics. This is the visual interference-signal integration ratio. Similarly, adding 1 avoids calculation errors. A larger ratio indicates more severe visual highlight overexposure interference and weaker effective visual features. The two interference-signal integration ratios are then compared with... Multiplying these yields the combined radar and vision interference coefficient, which organically integrates the degree of coupled disturbance with the actual interference situation, achieving joint quantification of registration disturbance and modal interference. The input to the negative exponential function is this combined interference coefficient; since the input is dimensionless, the output is also dimensionless. The calculation result is also dimensionless, ensuring the homogeneity of the dimensions of the two calculation formulas. Simultaneously, the negative exponential mapping stabilizes the output weight values ​​within the range of 0 to 1, avoiding excessive or insufficient feature enhancement due to excessively large or small weights. The core innovation of this formula lies in achieving deep parameter linkage between the registration and feature enhancement stages, thus... As a common baseline input for the two weight calculation formulas, the feature enhancement weights of radar and vision are synchronized to match the degree of coupling perturbation, achieving joint suppression of co-source coupling interference. This solves the problem that traditional independent weight calculations cannot adapt to coupling interference. Simultaneously, the two weight values ​​serve as the core input to the subsequent knowledge distillation formula, realizing parameter linkage between the feature enhancement and knowledge distillation stages. The formula calculation is executed in parallel by the GPU of the edge computing board, and... The calculations are synchronized, with a computation time of less than 1 millisecond, and all input parameters are updated in real time to ensure that the weight values ​​can adapt to the dynamic changes of the sea surface scene. The achieved technical effects include: realizing deep parameter coupling between the registration and feature enhancement stages; the generated radar feature enhancement weights and visual feature enhancement weights can simultaneously match the degree of coupling disturbance and the situation of co-source interference on the sea surface; achieving joint suppression of radar clutter and visual highlights; improving the effect and adaptability of feature enhancement; and providing a precise modal weight benchmark for the subsequent knowledge distillation stage, realizing parameter linkage between feature enhancement and knowledge distillation stages.

[0029] The existing technology has the following technical problems: the multi-teacher dynamic distillation weight formula coupled with temporal reliability does not combine the inter-frame registration recursive compensation coefficient and the dual-modal feature enhancement weight of the preceding stage, nor does it introduce the constraint condition of modal temporal reliability. As a result, the allocation of distillation weight cannot adapt to the dynamic changes of the sea surface scene, the lightweight student network cannot learn the special detection knowledge of different scenarios in a targeted manner, and the detection performance and environmental adaptability are insufficient.

[0030] Based on this, the temporal reliability-coupled multi-teacher dynamic distillation weight formula includes the original weight calculation formula for the teacher network and the normalized weight calculation formula. The original weight calculation formula for the teacher network uses the inter-frame registration recursive compensation coefficient, radar feature enhancement weight, and visual feature enhancement weight as the baseline input. It is worth mentioning that the temporal reliability-coupled multi-teacher dynamic distillation weight formula is the core of achieving deep coupling of parameters throughout the entire process of registration, feature enhancement, and knowledge distillation. Its original weight calculation formula is based on the output parameters of the previous two formulas. , , Using the modal time-series reliability decay coefficient as the core input, a distillation weight constraint is introduced. The normalized weight calculation formula converts the original weights into effective distillation weights that sum to 1. The complete expression of the formula is: , , , ; The theoretical design of this formula is based on the principle of specialized adaptation of multi-teacher knowledge distillation and the temporal reliability constraint mechanism of modality. Each teacher network in the multi-teacher network performs specialized optimization for a core type of interference in the sea surface scenario. Different distillation weights need to be assigned according to the changes in the scenario. At the same time, the temporal reliability of the modality directly affects the effectiveness of knowledge distillation, and constraints need to be introduced to ensure that the student network only learns the effective knowledge of the high-reliability mode. The logical derivation of the formula is based on the principle of specialized adaptation of weights and reliability constraints. First, it uses... , , Based on this, the original weights of the three teacher networks are calculated using the modal temporal reliability decay coefficient. These original weights are then highly matched with the modal feature enhancement requirements, coupling perturbation degree, and modal temporal reliability. Normalization is then applied to convert the original weights into weight values ​​that meet the requirements of knowledge distillation, ensuring that the sum of the weights is 1, thus achieving a reasonable allocation of distillation weights. The definitions and acquisition methods of the parameters in the formula are as follows: , , The original weights of the three teacher networks are the output parameters of the original weight calculation formula. They are dimensionless and range from 0 to 1. , , The final dynamic distillation weights for the three teacher networks are the output parameters of the normalized weight calculation formula. They are dimensionless, range from 0 to 1, and satisfy the following conditions: ; The weights for enhancing radar features are calculated and output from the preceding formulas and are dimensionless. The weights for visual features are enhanced and the output is calculated from the preceding formula; they are dimensionless. The recursive compensation coefficients for inter-frame registration are calculated and output from the preceding formulas and are dimensionless. The radar mode timing reliability attenuation coefficient quantifies the timing reliability of the radar mode. It is obtained by the proportion of frames in a consecutive preset frame where the signal-to-clutter ratio of the effective target signal of the radar is lower than a threshold. It is dimensionless and ranges from 0 to 1. The larger the value, the lower the timing reliability of the radar mode. The visual modality temporal reliability attenuation coefficient quantifies the temporal reliability of the visual modality. It is obtained by calculating the proportion of frames with a visual effective feature signal-to-noise ratio below a threshold within consecutive preset frames. It is dimensionless and ranges from 0 to 1; a larger value indicates lower temporal reliability of the visual modality. The three original weights correspond to the three specialized teacher networks for optimization. Corresponding to the sea surface clutter suppression teacher network, and Positive correlation with Negative correlation occurs when radar clutter interference is severe and radar mode reliability is high. The network will be expanded to allow students to focus on learning specialized knowledge about sea surface clutter suppression. Corresponding to the teacher network that compensates for sudden changes in light intensity, and Positive correlation with Negative correlation, when visual illumination interference is severe and visual modality reliability is high. The network will be expanded to allow students to focus on learning specialized knowledge about compensation for sudden changes in light intensity. Corresponding to the ship attitude disturbance, the teacher network is adapted to the network. Positively correlated with the average reliability decay coefficient of the two modes, and negatively correlated with the average reliability decay coefficient of the two modes, when the hull attitude disturbance is severe and the overall modal reliability is high. The system synchronously increases the weights, allowing the student network to focus on learning specialized knowledge related to ship attitude perturbation adaptation. The normalized weight calculation formula normalizes the original weights by dividing each weight by the sum of the three original weights, ensuring that the final sum of the dynamic distillation weights is 1. This meets the weight allocation requirements of knowledge distillation, allowing the knowledge from the three teacher networks to be transferred to the lightweight student network in a reasonable proportion. The core innovation of this formula lies in achieving deep coupling of parameters throughout the entire process of registration, feature enhancement, and knowledge distillation. It uses the core parameters of the preceding steps as the basis for distillation weight calculation, while introducing a modal temporal reliability attenuation coefficient to achieve scene adaptation and reliability constraints for the distillation weights. This allows the lightweight student network to learn specialized optimization knowledge for different sea surface scenarios, solving the problem that traditional static knowledge distillation cannot adapt to dynamic scenarios. The calculation is performed in real-time by the GPU of the edge computing board during the knowledge distillation stage. All input parameters are updated in real-time, and the calculation time is less than 1.5 milliseconds, ensuring that the distillation weights can be dynamically adjusted according to changes in the sea surface scenario. Technical results achieved: Deep coupling of parameters throughout the entire process of the three core links was realized. The generated multi-teacher network dynamic distillation weights can accurately adapt to the dynamic changes and modal temporal reliability of the sea surface scene, allowing the lightweight student network to learn various specialized detection knowledge in a targeted manner. While maintaining the lightweight nature of the model, the detection performance and environmental adaptability of the model were greatly improved.

[0031] The existing technology has the following technical problems: the knowledge distillation of the lightweight student network only performs single-dimensional distillation processing and does not set up a special optimized teacher network for different core interferences on the sea surface. As a result, the student network cannot learn the special optimized knowledge for different sea surface scenarios, and its environmental adaptability and detection performance are insufficient.

[0032] Based on this, the step of performing knowledge distillation on the lightweight student network based on dynamic distillation weights includes performing feature-level distillation and output-level distillation on the lightweight student network based on the dynamic distillation weights of a multi-teacher network. The multi-teacher network includes a sea clutter suppression teacher network, an illumination change compensation teacher network, and a ship attitude disturbance adaptation teacher network. It is worth noting that all multi-teacher networks are constructed using high-capacity backbone networks to ensure that each teacher network has excellent detection performance in its corresponding specialized scenario. The sea clutter suppression teacher network uses ResNet101 as its backbone network and a radar point cloud and visual image fusion dataset containing sea clutter under different sea conditions as its training set. It is specifically trained until the loss converges. The trained network can effectively suppress various types of sea clutter interference and improve target detection capabilities in clutter backgrounds. The illumination change compensation teacher network also uses ResNet101 as its backbone network and a radar point cloud and visual image fusion dataset containing sea clutter under different sea conditions as its training set. It is specifically trained until the loss converges. The trained network can effectively suppress various types of sea clutter interference and improve target detection capabilities in clutter backgrounds. The training set includes sea surface target datasets with varying light intensities and abrupt changes in light intensity. The network is specifically trained until loss convergence. The trained network effectively recovers visual feature details under abrupt changes in light intensity, improving target detection capabilities in low-visibility scenarios. The teacher network for ship attitude perturbation adaptation uses ResNet101 as its backbone and a multimodal fusion dataset containing varying degrees of ship attitude perturbation as its training set. It is specifically trained until loss convergence. The trained network effectively adapts to roll, pitch, and heave perturbations, improving target detection capabilities under turbulent ship conditions. The lightweight student network uses MobileNetV3 as its backbone. This network has only 1 / 8 the number of parameters and 1 / 10 the computational cost of the teacher network, making it perfectly compatible with edge computing platforms for unmanned surface vessels. Before knowledge distillation, this network is pre-trained on a general sea surface target dataset, possessing basic sea surface target detection capabilities. The feature-level distillation process involves applying loss constraints to the intermediate layer features of the three teacher networks and the corresponding intermediate layer features of the lightweight student network. Specifically, it extracts the convolutional features of layers 4, 8, and 12 of both the teacher and student networks, calculates the mean squared error loss of each layer, and then uses the dynamic distillation weights of the multi-teacher networks as weighting coefficients to calculate the total feature-level distillation loss. This aims to make the intermediate layer features of the student network as close as possible to the intermediate layer features of the teacher network, thus learning the feature extraction capabilities of the teacher network. The output-level distillation process involves applying loss constraints to the output layer prediction results of the three teacher networks and the output layer prediction results of the lightweight student network. Specifically, it extracts the target category prediction probabilities and bounding box regression coordinates of both the teacher and student networks, calculates the cross-entropy loss of the category prediction and the smoothing L1 loss of the bounding box regression, and then uses the dynamic distillation weights of the multi-teacher networks as weighting coefficients to calculate the total output-level distillation loss. This aims to make the prediction results of the student network as close as possible to the prediction results of the teacher network, thus learning the target detection capabilities of the teacher network.The total loss of knowledge distillation is the weighted sum of the total distillation loss at the feature level and the total distillation loss at the output level, with weighting coefficients set to 0.4 and 0.6, respectively. During distillation training, a mini-batch stochastic gradient descent algorithm is used for optimization, with a learning rate of 0.001 and a batch size of 16, training continues until the total loss converges. In the actual navigation of the unmanned surface vessel, knowledge distillation is performed online and lightweight, dynamically adjusting distillation weights only according to changes in the sea surface scene, eliminating the need for retraining and ensuring real-time performance. The achieved technical results are as follows: Through joint distillation at the feature and output levels using three specialized optimized teacher networks, the lightweight student network simultaneously learns specialized optimized knowledge for sea surface clutter suppression, illumination abrupt change compensation, and ship attitude perturbation adaptation, as well as general sea surface target detection knowledge. While maintaining the model's lightweight nature, this significantly improves the student network's environmental adaptability and detection performance.

[0033] The existing technology has the following technical problems: the feature extraction and fusion of multi-band sensing data does not dynamically adjust the fusion ratio based on the radar feature enhancement weight and the visual feature enhancement weight, nor does it perform multi-layer feature interaction processing on the fused features. As a result, the fused features cannot give full play to the complementary advantages of radar and vision in the dual modes, and have weak feature representation capabilities for small targets and occluded targets on the sea surface.

[0034] Based on this, the step of performing feature extraction and feature fusion processing on multi-band sensing data using a lightweight student network to obtain fused features includes extracting radar depth features and visual depth features based on the lightweight student network; performing weighted fusion processing on the radar depth features and visual depth features based on the radar feature enhancement weights and the visual feature enhancement weights to obtain initial fused features; and performing multi-layer feature interaction processing on the initial fused features to obtain the fused features. It is worth mentioning that the lightweight student network contains two independent feature extraction branches: a radar feature extraction branch and a visual feature extraction branch. The two branches share the backbone network structure of MobileNetV3, with differences only in the input layer. The radar depth feature extraction process is as follows: The registered and compensated millimeter-wave radar point cloud data is voxelized, converting the 3D point cloud into a fixed-size voxel grid (32×32×32). The voxel grid is then input into the radar feature extraction branch, and processed through convolutional layers, depthwise separable convolutional layers, and pooling layers to extract radar depth features at four scales: 8×8, 16×16, 32×32, and 64×64. Each scale feature contains 256 feature channels, which can fully characterize the spatial structure and target features of the radar point cloud. The visual depth feature extraction process is as follows: The registered and compensated visible light image data is normalized and resized to a fixed size of 256×256. The resized image is then input into the visual feature extraction branch, processed through the same network layers as the radar feature extraction branch to extract four visual depth features corresponding to the radar depth feature scales. Each scale feature also contains 256 feature channels, which can fully characterize the texture details and target features of the visual image. The weighted fusion process involves performing a channel-by-channel weighted summation of radar depth features and visual depth features at the same scale. Specifically, this involves multiplying the radar depth features at each scale by the radar feature enhancement weight. Multiply the visual depth features at each scale by the visual feature enhancement weights. The calculation results from both are then added channel by channel to obtain initial fusion features at four scales. This fusion method can dynamically adjust the fusion ratio of radar and visual features according to changes in the sea surface scene, fully leveraging the environmental robustness of radar and the high-resolution advantages of vision. The multi-layer feature interaction processing adopts a feature fusion method combining upsampling and downsampling. Specifically, the initial fusion feature at the largest scale is upsampled to the second largest scale and concatenated with the initial fusion feature at the second largest scale. The concatenated feature is then upsampled to the next scale and concatenated with the initial fusion feature at the corresponding scale. This process is repeated until all scale features have been upsampled and concatenated. Then, the concatenated feature at the smallest scale is downsampled to each original scale to obtain the final multi-scale fusion features. The number of channels for each scale fusion feature is 512. This multi-layer feature interaction processing can fully integrate deep semantic features and shallow detail features. Deep semantic features can improve the target category recognition ability, while shallow detail features can improve the feature representation ability of small targets and occluded targets, effectively solving the problem of high detection difficulty for small targets and occluded targets on the sea surface. The extraction and fusion of the fused features are both executed in real time by the GPU of the edge computing board, with a processing time of less than 2 milliseconds, meeting the real-time detection requirements of unmanned surface vessels. The achieved technical effects include: weighted fusion of radar and visual features based on dynamic dual-modal feature enhancement weights, fully leveraging the complementary advantages of dual-modal approaches; multi-layer feature interaction processing strengthens the fusion of deep semantic features and shallow detail features, significantly improving the feature representation capabilities of small targets and occluded targets on the sea surface, and providing high-quality multi-scale fused features for subsequent target detection processing.

[0035] The existing technology has the following technical problems: The target detection processing adopts an open-loop detection mode and does not perform closed-loop updates of parameters throughout the process based on the average confidence of the target based on the detection results. This makes it unable to adapt to sudden changes in the sea surface scene, resulting in insufficient environmental adaptability of the detection system and poor detection stability.

[0036] Based on this, the step of performing target detection processing based on fused features includes inputting the fused features into an anchorless target detection head, performing target category prediction processing and bounding box regression processing, outputting target detection results, and updating the calculation parameters of the inter-frame registration recursive compensation coefficient and the generation parameters of the feature enhancement weights based on the target average confidence of the target detection results. It is worth mentioning that the anchorless target detection head is constructed based on a center regression detection framework, requiring no preset anchor boxes and adapting to the scale and shape changes of sea surface targets. This detection head includes a classification branch, a regression branch, and a center branch, all three branches being constructed using convolutional layers and activation functions, directly connected to the output of a lightweight student network. The classification branch is used to perform target category prediction processing, taking multi-scale fused features as input, and outputting the target category probability of each feature point through convolutional layers and a softmax activation function. The target categories covered include common sea surface targets such as unmanned vessels, fishing boats, cargo ships, navigation marks, floating bodies, and reefs. Each feature point corresponds to a category probability vector, and the dimension of the vector is consistent with the number of target categories. The regression branch performs bounding box regression processing, taking multi-scale fused features as input. Through convolutional layers and linear activation functions, it outputs the target bounding box coordinates corresponding to each feature point. These coordinates are represented as center coordinates and width / height, which can be directly converted to pixel coordinates for accurate target localization. The center branch determines whether a feature point is the center region of the target, outputting the center confidence score for each feature point. By setting a center confidence score threshold, false detection boxes can be effectively filtered out, improving the accuracy of the detection results. The target detection result generation process is as follows: First, the class probability output by the classification branch is multiplied by the center confidence score output by the center branch to obtain the overall target confidence score for each feature point. Then, a comprehensive confidence score threshold and a non-maximum suppression threshold are set, and a non-maximum suppression algorithm is used to filter out redundant detection boxes, obtaining the final target detection result. This result includes the target's class, bounding box coordinates, and comprehensive confidence score.

[0037] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A multi-band fusion target detection method for unmanned surface vessels, characterized in that: Includes the following steps: S1: Simultaneously collect multi-band sensing data, ship motion data, and marine environmental data; S2: Calculate the inter-frame registration recursive compensation coefficient based on the multi-band sensing data, the ship motion data, and the marine environment data; S3: Based on the inter-frame registration recursive compensation coefficients, generate radar feature enhancement weights and visual feature enhancement weights; S4: Calculate the dynamic distillation weights of the multi-teacher network based on the inter-frame registration recursive compensation coefficients, the radar feature enhancement weights, and the visual feature enhancement weights; S5: Based on the dynamic distillation weights, perform knowledge distillation on the lightweight student network; S6: Based on the lightweight student network, perform feature extraction and feature fusion processing on the multi-band sensing data to obtain fused features; S7: Based on the fused features, perform target detection processing.

2. The unmanned surface vessel environmental coupling knowledge distillation multi-band fusion target detection method according to claim 1, characterized in that, The steps of simultaneously acquiring multi-band sensing data, ship motion data, and marine environmental data include simultaneously acquiring millimeter-wave radar point cloud data, visible light image data, ship attitude disturbance time-series change rate data, ship speed data, effective wave height data, and light intensity data.

3. The unmanned surface vessel environmental coupling knowledge distillation multi-band fusion target detection method according to claim 2, characterized in that, The step of calculating the inter-frame registration recursive compensation coefficient based on multi-band sensing data, ship motion data, and marine environment data includes calculating the inter-frame registration recursive compensation coefficient based on ship speed data, ship attitude disturbance temporal change rate data, effective wave height data, sensor sampling period, ship reference length, and registration residual deviation of the previous frame.

4. The unmanned surface vessel environmental coupling knowledge distillation multi-band fusion target detection method according to claim 3, characterized in that, The step of generating radar feature enhancement weights and visual feature enhancement weights based on inter-frame registration recursive compensation coefficients includes generating radar feature enhancement weights and visual feature enhancement weights based on inter-frame registration recursive compensation coefficients, radar echo specular reflection clutter ratio, radar effective target signal-to-noise ratio, visual image highlight overexposure pixel ratio, and visual effective feature signal-to-noise ratio.

5. The unmanned surface vessel environmental coupling knowledge distillation multi-band fusion target detection method according to claim 4, characterized in that, The inter-frame registration recursive compensation coefficient, the radar feature enhancement weight, the visual feature enhancement weight, and the dynamic distillation weight of the multi-teacher network are calculated through a set of linked calculation formulas. This set of formulas includes a formula for the environment-coupled inter-frame registration recursive compensation coefficient, a formula for the coupled interference joint feature enhancement weight, and a formula for the time-series reliability-coupled multi-teacher dynamic distillation weight. The formula for the environment-coupled inter-frame registration recursive compensation coefficient is as follows: ; In the formula, The recursive compensation coefficients for inter-frame registration are used. This refers to the ship's real-time speed. The sensor sampling period, The effective wave height at sea surface. The reference length of the hull. Let be the temporal rate of change of the ship's attitude disturbance. To account for the registration residual error of the previous frame. , , The weighting coefficients are dimensionless and satisfy the following conditions: .

6. The unmanned surface vessel environmental coupling knowledge distillation multi-band fusion target detection method according to claim 5, characterized in that, The joint feature enhancement weight formula for coupled interference includes a radar feature enhancement weight calculation formula and a visual feature enhancement weight calculation formula. The radar feature enhancement weight calculation formula uses the inter-frame registration recursive compensation coefficient as the reference input, and the visual feature enhancement weight calculation formula uses the inter-frame registration recursive compensation coefficient as the reference input.

7. The unmanned surface vessel environmental coupling knowledge distillation multi-band fusion target detection method according to claim 5, characterized in that, The time-reliability coupled multi-teacher dynamic distillation weight formula includes the original weight calculation formula of the teacher network and the normalized weight calculation formula. The original weight calculation formula of the teacher network uses the inter-frame registration recursive compensation coefficient, radar feature enhancement weight and visual feature enhancement weight as the reference input.

8. The unmanned surface vessel environmental coupling knowledge distillation multi-band fusion target detection method according to claim 5, characterized in that, The step of performing knowledge distillation on the lightweight student network based on dynamic distillation weights includes performing feature-level distillation and output-level distillation on the lightweight student network based on the dynamic distillation weights of the multi-teacher network. The multi-teacher network includes a sea clutter suppression teacher network, an illumination change compensation teacher network, and a hull attitude disturbance adaptation teacher network.

9. The unmanned surface vessel environmental coupling knowledge distillation multi-band fusion target detection method according to claim 1, characterized in that, The steps of performing feature extraction and feature fusion processing on multi-band sensing data based on a lightweight student network to obtain fused features include: extracting radar depth features and visual depth features based on the lightweight student network; performing weighted fusion processing on the radar depth features and visual depth features based on the radar feature enhancement weights and the visual feature enhancement weights to obtain initial fused features; and performing multi-layer feature interaction processing on the initial fused features to obtain fused features.

10. The unmanned surface vessel environmental coupling knowledge distillation multi-band fusion target detection method according to claim 1, characterized in that, The steps of performing target detection processing based on fused features include inputting the fused features into an anchorless target detection head, performing target category prediction processing and bounding box regression processing, outputting target detection results, and updating the calculation parameters of the inter-frame registration recursive compensation coefficient and the generation parameters of the feature enhancement weights based on the target average confidence of the target detection results.