Unmanned aerial vehicle-oriented remote sensing cross-modal adversarial sample generation method, device and equipment

By acquiring real-time data and hardware status of UAV remote sensing platforms, dynamically optimizing strategies to generate adversarial examples, and performing closed-loop collaborative control, the problem of low attack efficiency and adaptability of UAV remote sensing platforms in resource-constrained environments is solved, achieving efficient and collaborative adversarial attacks.

CN122157011APending Publication Date: 2026-06-05HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS
Filing Date
2026-02-27
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

In dynamic, resource-constrained real-world mission environments, existing UAV remote sensing platforms cannot adaptively adjust algorithm parameters, resulting in low attack efficiency or resource waste. Cross-modal attacks are inconsistent in timing and easily filtered out. The generated adversarial examples lack adaptability and survivability and cannot be optimized in collaboration with hardware platforms.

Method used

By acquiring remote sensing data and real-time hardware status data, type recognition processing is performed to generate dynamic optimization strategies, image or video perturbation processing is performed to generate adversarial examples, and attack effectiveness is evaluated and hardware is controlled to achieve closed-loop collaboration between algorithms and hardware.

Benefits of technology

Generate highly mobile and highly covert adversarial samples in complex environments to improve attack effectiveness and platform resource utilization efficiency, adapt to changes in hardware status and electromagnetic interference, achieve coordinated control of software and hardware, and improve practical applicability and mission survivability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122157011A_ABST
    Figure CN122157011A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a method, device and equipment for generating remote sensing cross-modal adversarial samples for unmanned aerial vehicles. A specific embodiment of the method includes: acquiring remote sensing data and real-time hardware state data; performing type identification processing on the remote sensing data; generating a dynamic optimization strategy based on the data type determination result; in response to the type determination result representing an image, performing first perturbation processing on the remote sensing data; in response to the type determination result representing a video, performing second perturbation processing on the remote sensing data; performing superposition processing on the adversarial perturbation data or adversarial perturbation data sequence and the remote sensing data to obtain an adversarial sample; performing attack effectiveness evaluation on the adversarial sample to obtain an evaluation result; iteratively optimizing the dynamic optimization strategy to obtain a set of hardware control parameters; and adjusting an onboard computing platform according to the set of hardware control parameters. The embodiment provides a technical paradigm for the safety testing and adversarial defense research of an unmanned aerial vehicle remote sensing system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure relate to the field of computer technology, and more specifically to a method, apparatus, and device for generating cross-modal adversarial samples for remote sensing of unmanned aerial vehicles (UAVs). Background Technology

[0002] With the widespread application of UAV remote sensing platforms and the continuous development of adversarial attack techniques, the need for security assessment and robustness testing of airborne target recognition systems is becoming increasingly prominent. Currently, related technological explorations mainly focus on the algorithm level, such as improving gradient optimization methods (e.g., MIM, I-FGSM) and introducing input transformations (e.g., spatial cropping, frequency domain wavelet transform) to enhance the cross-model transferability of adversarial examples, and drawing on the idea of ​​image-to-video (I2V) cross-modal attacks to generate adversarial frames that can interfere with video recognition models using image models.

[0003] However, when applying the aforementioned existing methods to the dynamic, resource-constrained real-world UAV mission environment, the following technical problems often arise: First, the parameters (such as iteration step size and batch size) of existing adversarial example generation algorithms are mostly statically preset, unable to adaptively adjust according to the real-time fluctuations in hardware resources (such as video memory, power consumption, and communication bandwidth) of the UAV's onboard computing platform. This leads to a sharp drop in attack efficiency when computing power is a bottleneck, or an inability to fully utilize computing power when resources are abundant. Second, existing cross-modal attack methods fail to effectively model the spatiotemporal continuity between frames when processing remote sensing video. The generated adversarial perturbations lack temporal consistency, are easily filtered by the motion compensation mechanism of the video system, and are difficult to balance between attack strength and visual concealment. More importantly, existing research treats the generation of adversarial examples as a purely algorithmic problem. Its optimization process is completely disconnected from the specific flight mission stage of the UAV, the real-time electromagnetic environment encountered, and the platform's own maneuvering state, resulting in the generated adversarial strategies lacking adaptability and survivability in real complex battlefield environments. Meanwhile, how to feed back and control various processors (such as flight control processors and mission processors) and controlled hardware (such as data link transmitters and optoelectronic payload servo mechanisms) of the UAV's onboard computing platform in real time and with high precision, so as to achieve synergy between attack effectiveness and the platform's own safe and stable flight, constitutes a key technical obstacle.

[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0006] Some embodiments of this disclosure provide a method, apparatus, electronic device, and computer-readable medium for generating cross-modal adversarial samples for remote sensing of unmanned aerial vehicles (UAVs) to address one or more of the technical problems mentioned in the background section above.

[0007] In a first aspect, some embodiments of this disclosure provide a method for generating cross-modal adversarial examples for remote sensing applications on unmanned aerial vehicles (UAVs), comprising: acquiring remote sensing data and real-time hardware status data; performing type identification processing on the acquired remote sensing data to obtain a data type determination result; generating a dynamic optimization strategy based on the data type determination result and the acquired real-time hardware status data; in response to the type determination result being represented as an image, performing a first perturbation processing on the acquired remote sensing data based on the dynamic optimization strategy to obtain adversarial perturbation data; in response to the type determination result being represented as video, performing a second perturbation processing on the acquired remote sensing data based on the dynamic optimization strategy to obtain an adversarial perturbation data sequence; superimposing the obtained adversarial perturbation data or adversarial perturbation data sequence with the acquired remote sensing data to obtain an adversarial example; evaluating the attack effectiveness of the adversarial example to obtain an evaluation result; iteratively optimizing the dynamic optimization strategy based on the performance evaluation result and the acquired real-time hardware status data to obtain a hardware control parameter set; and adjusting each processor and each controlled hardware in the airborne computing platform according to the hardware control parameter set.

[0008] Secondly, some embodiments of this disclosure provide a remote sensing cross-modal adversarial example generation device for unmanned aerial vehicles (UAVs), comprising: an acquisition unit configured to acquire remote sensing data and real-time hardware status data; an identification unit configured to perform type identification processing on the acquired remote sensing data to obtain a data type determination result; a generation unit configured to generate a dynamic optimization strategy based on the data type determination result and the acquired real-time hardware status data; a first perturbation unit configured to, in response to the type determination result being represented as an image, perform a first perturbation processing on the acquired remote sensing data based on the dynamic optimization strategy to obtain adversarial perturbation data; and a second perturbation unit configured to, in response to the type determination result... The data is represented as video. Based on the aforementioned dynamic optimization strategy, the acquired remote sensing data undergoes a second perturbation process to obtain an adversarial perturbation data sequence. An overlay unit is configured to overlay the obtained adversarial perturbation data or the adversarial perturbation data sequence with the acquired remote sensing data to obtain adversarial examples. An evaluation unit is configured to evaluate the attack effectiveness of the adversarial examples to obtain an evaluation result. An optimization unit is configured to iteratively optimize the aforementioned dynamic optimization strategy based on the performance evaluation result and the acquired real-time hardware status data to obtain a set of hardware control parameters. An adjustment unit is configured to adjust each processor and each controlled hardware component in the airborne computing platform according to the aforementioned set of hardware control parameters.

[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0011] Fifthly, some embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0012] The above embodiments of the present invention have the following beneficial effects: The remote sensing cross-modal adversarial sample generation method for UAVs of the present invention achieves adaptive, closed-loop, and hardware-software collaborative capabilities for generating highly mobile and highly concealed adversarial samples in dynamic, resource-constrained UAV mission environments, improving the practical applicability, efficiency, and mission survivability of adversarial attacks in real and complex remote sensing scenarios. Specifically, traditional adversarial sample generation techniques (such as optimization algorithms relying on static parameters or cross-modal attack schemes detached from hardware) may experience a sharp drop in attack efficiency when encountering computing power bottlenecks on the UAV platform, the generated video adversarial perturbation may have temporal breaks that are easily filtered out, and the attack behavior may be "blindly" due to complete disconnection from the flight mission; if only a fixed strategy is used, it will be unable to adapt to fluctuations in hardware status during flight, interference from the electromagnetic environment, and changes in maneuvering attitude, ultimately leading to attack failure or self-exposure. Based on this, the remote sensing cross-modal adversarial sample generation method for UAVs of the present invention: First, acquire remote sensing data and real-time hardware status data. This implants environmental awareness "nerve endings" into the entire method, enabling not only the acquisition of target data but, more importantly, real-time monitoring of the UAV computing platform's "physical state" (such as GPU memory, CPU load, and bandwidth). This provides a dynamic physical basis for all subsequent decisions, ensuring tight coupling between the method and the hardware platform from the outset. Next, the acquired remote sensing data undergoes type identification processing to obtain data type determination results. This enables automated data splitting (images or videos), providing precise triggering conditions for launching two different, targeted optimized attack flows (SFCM or TSFCM-I2V), avoiding performance loss from a single processing mode for another type of data. Then, based on the data type determination results and the acquired real-time hardware status data, a dynamic optimization strategy is generated. This reveals the core innovation: dynamically binding algorithm parameters to hardware status. For example, increasing the batch size and model count when GPU memory is ample enhances attack power, while adjusting the optimization step size during high CPU load avoids system lag. This achieves real-time optimal matching between algorithm performance and hardware resources, solving the core problem of static parameters being unable to adapt to dynamic environments. Subsequently, in response to the aforementioned type determination result being represented as an image, the acquired remote sensing data undergoes a first perturbation process based on the aforementioned dynamic optimization strategy. This involves executing a frequency-space cooperative modulation attack (SFCM) specifically designed for remote sensing images. This attack decouples the target background through wavelet transform, employs block-based differential perturbation based on attention heatmaps, and combines gradient optimization adjusted by a dynamic strategy to generate highly transferable image adversarial perturbations. Simultaneously, in response to the aforementioned type determination result being represented as video, the acquired remote sensing data undergoes a second perturbation process based on the aforementioned dynamic optimization strategy.Therefore, a Spatiotemporal Frequency Domain Co-modulation Attack (TSFCM-I2V) specifically designed for remote sensing video was executed. Building upon image attacks, keyframe priority processing and inter-frame low-frequency consistency constraints were introduced to ensure that the generated video adversarial perturbations were both effective and temporally smooth and continuous, overcoming the challenge of temporal inconsistency in cross-modal attacks. Subsequently, the obtained adversarial perturbation data or sequences were superimposed on the original data to obtain adversarial samples. This completed the physical generation of adversarial samples, providing entity objects for subsequent performance evaluation and control execution. The attack performance of the adversarial samples was then evaluated, yielding the evaluation results. This established a real-time feedback loop of effectiveness, objectively measuring the effectiveness and stealth of the attack through quantitative indicators (such as attack success rate and structural similarity), providing a data-driven basis for strategy iteration. Furthermore, based on the above performance evaluation results and real-time hardware status data, the above dynamic optimization strategy was iteratively optimized to obtain a set of hardware control parameters. This enabled the online self-evolution of the strategy. Based on the effectiveness of the attack and the current hardware load, the optimization parameters for the next round (such as perturbation amplitude and number of iterations) are dynamically adjusted, and the optimization decisions are ultimately translated into executable hardware control instructions (such as adjusting computing resource allocation), forming a complete closed loop from "perception-decision-execution-evaluation" to "optimization". Finally, based on the above set of hardware control parameters, adjustments are made to each processor and controlled hardware in the airborne computing platform. This achieves the ultimate leap from digital strategy to physical control. The method not only generates "soft" adversarial examples but also directly drives "hard" components such as flight control processors, mission processors, data links, and optoelectronic payloads, enabling the UAV's computing resource allocation, communication waveforms, and even platform attitude to be collaboratively optimized for the current adversarial mission, fundamentally solving the problem of the disconnect between the algorithm and the physical system. Furthermore, because this method introduces dynamic adaptation and closed-loop feedback mechanisms throughout the entire chain of input perception, strategy generation, perturbation processing, effect evaluation, and hardware control, it can effectively address the challenges posed by UAV platform resource fluctuations, complex electromagnetic interference, and variable mission scenarios. By deeply integrating adversarial attack algorithms, real-time hardware scheduling, and flight mission context, this method transforms UAVs from passive "tools" executing pre-set attack programs into intelligent agents capable of proactively adjusting attack strategies and coordinating platform behavior based on their own state and environmental changes. Thus, by achieving the integration of algorithmic adaptation, cross-modal consistency, and software-hardware collaborative control, it significantly improves the reliability of UAV attack effectiveness, platform resource utilization efficiency, and overall mission survivability in realistic combat environments, providing a key technical paradigm for proactive security testing and adversarial defense research of UAV remote sensing systems. Attached Figure Description

[0013] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0014] Figure 1 This is a flowchart of some embodiments of the method for generating cross-modal adversarial examples for remote sensing of unmanned aerial vehicles according to the present disclosure; Figure 2 This is a schematic diagram of the structure of some embodiments of the remote sensing cross-modal adversarial sample generation device for UAVs according to the present disclosure; Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0016] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0017] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0018] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0019] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0020] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0021] Figure 1A flow 100 of some embodiments of a remote sensing cross-modal adversarial example generation method for unmanned aerial vehicles (UAVs) according to this disclosure is shown. This UAV-oriented remote sensing cross-modal adversarial example generation method includes the following steps: Step 101: Acquire remote sensing data and real-time hardware status data.

[0022] In some embodiments, the execution entity of the UAV-based remote sensing cross-modal adversarial example generation method (e.g., UAV-borne mission computer, ground control station server, or cloud-based collaborative computing node) can acquire remote sensing data and real-time hardware status data. The remote sensing data refers to digitized data containing ground target information collected by UAV-borne or spaceborne sensors, and its form can be static remote sensing images or dynamic remote sensing video sequences. Remote sensing images are typically single high-resolution optical, SAR (synthetic aperture radar), or multispectral images, with common file formats including TIFF and JPEG2000. Remote sensing video sequences are dynamic images composed of consecutive frames, recording the scene of the target changing over time; common encapsulation formats include MP4 and AVI. In practice, the remote sensing data can originate from real-time sensor data streams (such as camera video streams) or from historical mission data pre-stored in onboard memory or transmitted back via data link. The real-time hardware status data refers to a set of quantitative indicators characterizing the current operating status and resource availability of the UAV-borne computing platform and its related hardware modules. This data includes at least: computational resource metrics, such as GPU (Graphics Processing Unit) memory utilization and core occupancy, CPU (Central Processing Unit) load percentage and core frequencies, and system memory (RAM) occupancy and bandwidth; communication resource metrics, such as real-time uplink and downlink bandwidth of the wireless data link, signal interference intensity, and instantaneous power consumption of the communication sensing module; and platform status metrics, such as total system power consumption, key hardware temperatures (e.g., SoC temperature), and real-time maneuvering status feedback from the flight control system (e.g., acceleration and angular velocity). These data collectively characterize the upper limit of computational power available for generating adversarial examples, the power consumption and thermal constraints that must be followed, and the external environment that may affect the stability of the algorithm. In practice, the aforementioned execution entities can concurrently acquire both types of data through various hardware and software interfaces. For remote sensing data, the execution entity can read raw data or decoded data frames from optoelectronic payloads, SAR antennas, or onboard memory through camera driver interfaces (e.g., V4L2), high-speed serial buses (e.g., MIPI CSI, Camera Link), or file system interfaces. For real-time hardware status data, the executing entity can periodically collect and update various indicators through the performance monitoring interface provided by the operating system (such as the / proc file system of Linux), the dedicated tool interface provided by the hardware manufacturer (for example, by calling the nvidia-smi command line tool or its Python API to obtain the memory, utilization and power consumption information of NVIDIA GPU), and the status telemetry messages periodically issued by the flight control computer through the internal bus (such as CAN, Ethernet).For example, when an entity initiates an adversarial example generation task, it may simultaneously acquire the following data: a real-time electro-optical reconnaissance video stream with a resolution of 1920x1080 transmitted from the Camera Link interface (as remote sensing data); a query via nvidia-smi showing that the onboard Jetson AGX Orin module's GPU memory utilization is 65% and its temperature is 78°C; a reading via the system API showing the current average CPU load is 1.5 and available memory is 4GB; and a reading via the data link driver showing the current downlink communication bandwidth is 50Mbps with slight interference (as real-time hardware status data). This data collectively forms the input basis for all subsequent decisions and processing.

[0023] Step 102: Perform type identification processing on the acquired remote sensing data to obtain the data type determination result.

[0024] In some embodiments, the aforementioned execution entity may perform type identification processing on the acquired remote sensing data to obtain a data type determination result.

[0025] In some optional implementations of certain embodiments, the aforementioned execution entity may perform type identification processing on the acquired remote sensing data through the following steps to obtain a data type determination result: Step one involves parsing the file header information or data stream structure of the acquired remote sensing data to obtain the format identifier. The file header information refers to a specific sequence of bytes stored at the beginning of the data file, describing metadata such as file format, encoding parameters, and data structure. The data stream structure refers to the logical organization of continuous data packets read from a real-time sensor interface (such as a camera video stream) that follow a specific encapsulation protocol (such as RTP over UDP). The format identifier refers to the characteristic information extracted from the file header information or data stream structure that uniquely or highly likely indicates the specific encoding format and type of the data. Examples include file extensions (such as ".tiff"), magic numbers (such as "0x0000000C 0x6A502020" in JPEG2000 files), or the payload type field in streaming media protocols. In practice, the executing entity can call corresponding media processing libraries (such as libmagic or OpenCV's video capture function) or write a parser according to standard format specifications to read the header bytes of the data or analyze the streaming protocol header. By matching predefined feature patterns, key format identifiers are extracted from the raw data. For example, for data read from a storage file, the execution entity reads its first 64 bytes, parses out the magic number as "0x49492A00", which corresponds to the TIFF image format, and thus extracts "TIFF" as the format identifier. For a real-time stream captured from the Camera Link interface, the execution entity parses its packet header, finds that it conforms to the SMPTE 292M standard and contains a progressive scan frame identifier, and thus extracts "HD-SDI_Progressive" as the identifier.

[0026] Step two: In response to the extracted format identifier conforming to a pre-defined static image encoding format, a type determination result representing an image is generated. The aforementioned pre-defined static image encoding format refers to a set of industry-standard or universal encoding formats pre-defined to represent single, static remote sensing images. This set includes, but is not limited to: TIFF (Tagged Image File Format), JPEG2000, PNG (Portable Web Graphics), BMP (Bitmap), and specific remote sensing data formats (such as GeoTIFF, ENVI .hdr / .img formats). In practice, the executing entity internally maintains a whitelist of static image formats. When the format identifier extracted in step one (such as "TIFF" or "JPEG2000_CODESTREAM") successfully matches any item in this whitelist, the executing entity determines that the current remote sensing data is a single image and generates a specific Boolean or enumerated variable as the data type determination result, for example, assigning the value IMAGE to the data_type variable. For example, the extracted format identifier is "JPEG2000," which is located in the pre-defined whitelist. Based on this, the executing entity determines that the data is a static image and generates a judgment result {"type":"IMAGE","confidence":0.95}, where the confidence field represents the judgment confidence level.

[0027] Step 3: In response to the extracted format identifier conforming to a pre-defined dynamic video encoding format, a type determination result representing video is generated. The aforementioned pre-defined dynamic video encoding format refers to a set of container formats or encoding formats pre-defined to represent dynamic remote sensing video composed of a continuous sequence of image frames. This set includes, but is not limited to: MP4 (MPEG-4 Part 14), AVI (Audio Video Interleaved Format), MKV (Matroska), MOV (QuickTime File Format), and streaming formats for real-time transmission (such as MPEG-TS). In practice, the execution entity also maintains a dynamic video format whitelist. When the format identifier extracted in Step 1 (such as "MP4" or "AVI") successfully matches any item in this whitelist, the execution entity determines that the current remote sensing data is a video sequence and generates a corresponding data type determination result, for example, assigning the data_type variable to VIDEO. For example, the extracted format identifier is "MP4," which is located in the pre-defined video format whitelist. Based on this, the executing entity determines that the data is dynamic video and generates a determination result {"type":"VIDEO","confidence":0.98}.

[0028] Step four: In response to the extracted format identifier not conforming to either the preset static image encoding format or the preset dynamic video encoding format, content-assisted determination is performed on the acquired remote sensing data to obtain the data type determination result. Content-assisted determination refers to a backup determination logic that infers the data type by analyzing the actual pixel or signal content characteristics of the remote sensing data itself after the initial determination based on the format identifier fails. This is typically used to handle situations such as header information corruption, custom encapsulation formats, or raw sensor data streams. In practice, the execution entity can implement one or more of the following auxiliary determination strategies: 1) Frame count analysis: Attempt to decode the data as video or divide it into fixed-size blocks. If multiple frames are successfully parsed and there are timestamps or identifiable temporal changes between frames, it is determined to be video; if only one meaningful image frame can be parsed, it is determined to be an image. 2) Data volume inference: Estimate the number of frames based on the total data size and typical remote sensing image resolution. 3) Specific content feature detection: Use lightweight models or heuristic rules to detect whether there are obvious temporal change features (such as motion blur, foreground target displacement) in the data. For example, for a raw data file with missing header information, format identifier extraction failed. The executing entity attempted to load the data using an image decoder (such as OpenCV's imdecode) and a video decoder (such as OpenCV's VideoCapture). The image decoder successfully loaded a complete high-resolution image, while the video decoder reported an error. Based on the data size (approximately 20MB, consistent with a single high-resolution image), the executing entity used content-assisted determination to generate a data type determination result of {"type":"IMAGE","confidence":0.70}.

[0029] Step 103: Generate a dynamic optimization strategy based on the data type determination result and the acquired real-time hardware status data.

[0030] In some embodiments, the execution entity may generate a dynamic optimization strategy based on the data type determination results and the acquired real-time hardware status data.

[0031] In some optional implementations of certain embodiments, the aforementioned execution entity can generate a dynamic optimization strategy based on the data type determination result and the acquired real-time hardware status data through the following steps: Step 1: Based on the data type determination results mentioned above, generate the number of iterations and the upper limit of the perturbation amplitude corresponding to the gradient optimization. Here, gradient optimization refers to the core algorithm process in generating adversarial examples, which involves iteratively calculating the gradient of the loss function relative to the input data and updating the perturbation along the gradient direction to optimize the attack effect. Examples include MI-FGSM (Momentum Iterative Fast Gradient Signed Method) or its improved variants. The number of iterations refers to the total number of rounds of update operations performed in the gradient optimization process. The upper limit of the perturbation amplitude is a key parameter used to constrain the strength of the adversarial perturbation. It defines the maximum absolute value of the perturbation allowed to be added to each pixel of the original remote sensing data. In specific implementations, if the image pixel value range is normalized to [0, 255], the upper limit of the perturbation amplitude is usually taken within this range, for example, 16, meaning that the change in each pixel value of the adversarial example relative to the original value must not exceed 16. In practice, the execution entity has a pre-set baseline parameter mapping table associated with different data types. When the data type determination result is "image," representing a static image, the executing entity uses the baseline iteration count (e.g., 30 times) and baseline perturbation amplitude upper limit configured for image adversarial attacks (such as the SFCM algorithm) (e.g., after normalizing the pixel value variation range to 0-255, take 16). When the data type determination result is "video," representing a dynamic video, the baseline value configured for video cross-modal attacks (such as the TSFCM-I2V algorithm) is used. Since video attacks need to consider inter-frame temporal smoothness, their baseline iteration count is usually set higher (e.g., 60 times) for more refined optimization, while the baseline perturbation amplitude upper limit may be slightly lower (e.g., 12) to maintain better stealth. Based on the baseline value obtained from the query, the executing entity generates preliminary iteration count and perturbation amplitude upper limit parameters.

[0032] Step two: Based on the available GPU memory size in the acquired real-time hardware status data, generate the number of proxy models and batch size for processing. The available GPU memory size refers to the currently unused memory capacity on the graphics processor or dedicated AI accelerator that can be used to load deep learning models and compute intermediate data. The number of proxy models refers to the total number of pre-trained models with different structures that are simultaneously loaded and used for integrated gradient computation during adversarial example generation, denoted as M. The batch size refers to the number of sample images processed in parallel after input transformation during a single gradient computation, denoted as N. In practice, the execution entity determines M and N through a specific resource budget process. This process first uses preset or real-time measured parameters: the static memory usage of a single proxy model C_model (in megabytes) and the dynamic memory increment C_image (in megabytes / image) required to process a single transformed image. Next, it reads the currently available GPU memory M_avail from the real-time hardware status data and multiplies it by a preset safety factor k (e.g., 0.8) to calculate the available GPU memory M_budget for this task. Core decision-making follows the constraints: M C_model+N M C_image <= M_budget. The execution entity solves the problem under this constraint. For example, a heuristic approach can be used: first, try loading the maximum number of models (M_max) and set N=1. If memory exceeds the budget, gradually reduce M or N until the constraint is met. After determining the feasible M, try gradually increasing N to the maximum value allowed by the constraint. Finally, output the determined combination of (M, N) as the number of surrogate models generated and the batch size.

[0033] Step 3: Based on the processor utilization data obtained from the real-time hardware status data, generate the dynamic step size decay coefficient and initial value of the momentum factor corresponding to the gradient optimization process. The processor utilization refers to the average percentage of load on the computing cores of the central processing unit and / or graphics processing unit within the sampling time window, a key indicator reflecting the current computational load of the system. The dynamic step size decay coefficient, denoted as γ, is a hyperparameter controlling the exponential decay rate of the gradient optimization iteration step size over time; a larger value indicates faster step size decay. The initial value of the momentum factor, denoted as β_0, is the weight coefficient of the momentum term at the beginning of the iteration in the gradient optimization algorithm, used to accumulate historical gradient directions to stabilize the optimization path. In practice, the execution entity queries a pre-set strategy mapping table based on the real-time monitored processor utilization value to determine the corresponding γ and β_0. This mapping table, calibrated based on offline experiments, maps utilization intervals to different optimization strategies. For example, a simplified mapping rule could be: when utilization is below 40%, a "fine-tuning" strategy is adopted, setting γ=0.01 and β_0=0.95; when utilization is between 40% and 80%, a "balancing" strategy is adopted, setting γ=0.03 and β_0=0.9; when utilization is above 80%, a "fast convergence" strategy is adopted, setting γ=0.1 and β_0=0.8. The executing entity then directly obtains and outputs the aforementioned dynamic step size decay coefficient and initial values ​​of the momentum factor.

[0034] Step four involves integrating the aforementioned iteration count, perturbation amplitude upper limit, number of proxy models, batch size, dynamic step size decay coefficient, and initial momentum factor value into a dynamic optimization strategy. In practice, the executing entity organizes the specific parameter values ​​determined in the preceding steps into a structured, machine-readable strategy description object. This strategy description object encompasses the key operating parameters and resource configuration instructions required by the adversarial example generation algorithm during this execution, and it will be passed to the subsequent perturbation processing module as a complete "recipe." For example, this strategy can be a dictionary or JSON object containing specific fields, ensuring the accuracy of parameter passing and the clarity of interfaces between modules.

[0035] Step 104: In response to the type determination result being represented as an image, the acquired remote sensing data is subjected to a first perturbation process based on a dynamic optimization strategy to obtain adversarial perturbation data.

[0036] In some embodiments, the execution entity may, in response to the type determination result being represented as an image, perform a first perturbation process on the acquired remote sensing data based on the dynamic optimization strategy described above, to obtain adversarial perturbation data.

[0037] In some optional implementations of certain embodiments, the aforementioned execution entity may, in response to the aforementioned type determination result being represented as an image, perform a first perturbation process on the acquired remote sensing data based on the aforementioned dynamic optimization strategy to obtain adversarial perturbation data: Step 1: Based on the number of proxy models specified in the dynamic optimization strategy, load the corresponding number of pre-trained image recognition models to obtain the target proxy model set. The number of proxy models is an integer determined in step 103 based on the available GPU memory. The pre-trained image recognition models refer to deep learning models with different network architectures (e.g., VGG, ResNet, DenseNet) that have been trained on large image datasets (e.g., ImageNet). The target proxy model set refers to the set of model instances actually loaded and used for gradient calculation during this attack. In practice, the execution entity loads a specified number of model weight files and computation graph structures into memory (GPU memory) from a local model repository or a remote server's on-demand cache, according to a predetermined model selection strategy (e.g., selecting the top N models with the greatest architectural differences), and completes model initialization, making them ready for forward inference and gradient calculation. For example, the dynamic optimization strategy specifies a proxy model count of 3. The execution entity then loads three pre-trained models—VGG-16, ResNet-50, and DenseNet-121—to form a target proxy model set.

[0038] Step two involves performing a discrete wavelet transform on the remote sensing data to obtain low-frequency and high-frequency components. The discrete wavelet transform is a mathematical tool that transforms an image from the spatial domain to the frequency domain, decomposing it into low-frequency components representing overall contours and approximate information, and high-frequency components representing details, edges, and texture information. A single two-dimensional discrete wavelet transform typically produces one low-frequency subband and three high-frequency subbands in three directions (horizontal, vertical, and diagonal). In practice, the execution entity calls a specific wavelet transform library (such as PyWavelets) or uses a preset convolution kernel (such as Haar wavelet or Daubechies wavelet) to perform the transform operation on the input remote sensing image, decomposing the image into low-frequency components and multiple high-frequency components. For example, performing a single Haar wavelet transform on a remote sensing image yields a low-frequency component (LL) with its size halved, and three high-frequency components: horizontal (LH), vertical (HL), and diagonal (HH).

[0039] Step 3 involves performing random linear scaling on the low-frequency components while keeping the high-frequency components unchanged, resulting in the transformed frequency domain components. The random linear scaling refers to multiplying each pixel value in the low-frequency components by a scaling factor randomly generated within a specific range. This operation aims to preserve key structural information of the target while introducing controllable diversity to enhance the transferability of subsequent adversarial perturbations. Keeping the high-frequency components unchanged means not altering the values ​​of the high-frequency subbands. In practice, the execution entity generates a random matrix of the same size as the low-frequency components, where each element follows a uniform distribution U(1-δ, 1+δ), where δ is a predefined hyperparameter controlling the scaling magnitude (e.g., 0.1). This random matrix is ​​then multiplied element-wise with the low-frequency component matrix (Hadamard product) to complete the random linear scaling. The high-frequency components are retained as is. For example, setting δ=0.1, a random matrix is ​​generated with element values ​​between [0.9, 1.1]. The low-frequency component LL is multiplied by this random matrix to obtain the scaled low-frequency component LL. The high-frequency components LH, HL, and HH remain unchanged.

[0040] Step four involves performing an inverse wavelet transform on the transformed frequency domain components to obtain a frequency-domain transformed image. This inverse wavelet transform is the inverse process of the discrete wavelet transform; it recombines the processed frequency components and the unchanged high-frequency components to reconstruct the image back into the spatial domain. In practice, the execution entity uses the same discrete wavelet transform basis as in step two to perform the inverse transform on the scaled low-frequency components and the original high-frequency components, generating a reconstructed image of the same size as the original remote sensing image—the frequency-domain transformed image. For example, using LL, LH, HL, and HH as input and performing an inverse Haar wavelet transform yields a frequency-domain transformed image X_freq.

[0041] Step 5: Perform activation mapping processing on the frequency domain transformed image to obtain a semantic attention heatmap. This activation mapping processing is a visualization technique used to generate grayscale or color images that indicate the degree of attention a deep learning model pays to different regions of the input image, such as gradient-weighted class activation mapping. The semantic attention heatmap is a matrix corresponding to the spatial size of the input image. The intensity value of each pixel reflects the "attention" weight the model assigns to that location during classification decisions; highlighted areas typically correspond to the target object. In practice, the execution entity selects a model from the target proxy model set as a reference model. The frequency domain transformed image X_freq is input into this reference model for forward propagation. The gradient of the model with respect to the true class (or a specified class) is calculated, and combined with the final convolutional layer feature map, a semantic attention heatmap is generated using a specific weighted summation formula. For example, selecting VGG-16 as the reference model and calculating its Grad-CAM (gradient-weighted class activation mapping) with respect to X_freq yields a semantic attention heatmap M, where the aircraft target region is highlighted.

[0042] Step Six: Based on the aforementioned semantic attention heatmap, the remote sensing data is divided into target-related region blocks and background region blocks. This division refers to region segmentation of the original remote sensing image based on the intensity of the semantic attention heatmap. The target-related region blocks are image blocks corresponding to continuous regions with attention values ​​higher than a set threshold. The background region blocks are image blocks corresponding to regions with attention values ​​lower than the threshold. In practice, the execution entity first normalizes the semantic attention heatmap M. Then, an adaptive threshold determination method (e.g., the Otsu method) is used to automatically calculate the optimal global threshold τ for segmenting the foreground (target) and background. This threshold is used to binarize the normalized heatmap, marking pixels with values ​​greater than τ as 1 (target region) and others as 0 (background region). Then, through image morphological operations (such as erosion and dilation) and connected component analysis, several connected foreground regions are extracted from the binary image. For each foreground connected component, its minimum bounding rectangle (or a finer boundary) is calculated, and corresponding target-related region blocks are cropped from the original remote sensing image based on these rectangles. Background region blocks are divided based on the pixel positions with a value of 0; typically, areas in the entire image not covered by target blocks are considered as background region blocks.

[0043] Step 7: Perform geometric transformations on the aforementioned target-related region blocks and environmental transformations on the aforementioned background region blocks to obtain the transformed region blocks. The geometric transformations refer to image processing operations that simulate changes in the viewing angle, such as scaling, rotation, translation, and affine transformations. The environmental transformations refer to image processing operations that simulate changes in imaging conditions, such as adding noise, adjusting brightness, contrast, and color balance. In practice, the executing entity maintains two transformation operation pools: a geometric transformation pool and an environmental transformation pool. For each target-related region block, one or more transformations are randomly selected from the geometric transformation pool and applied sequentially. For each background region block, one or more transformations are randomly selected from the environmental transformation pool and applied sequentially. The parameters of all transformations (such as rotation angle and noise intensity) can be randomly sampled within a preset range. For example, a small-angle rotation and slight scaling are randomly applied to the target block B_target (target-related region block). Gaussian noise is randomly added to the background block B_background (background region block), and its contrast is fine-tuned. The transformed target block B_target and background block B_background are obtained respectively.

[0044] Step eight involves stitching together the transformed regions to obtain a spatially transformed image. This stitching refers to reassembling the processed regions into a complete image according to the layout of the original remote sensing image. In practice, the execution entity creates a blank canvas of the same size as the original remote sensing image. First, the transformed target region B_target is placed in its original position in the original image. Then, the transformed background region B_background is filled into the remaining area. If multiple discontinuous background regions exist, they are filled separately. This results in an image that incorporates multiple spatial transformations, i.e., a spatially transformed image X_space. For example, placing B_target back in the original position of the aircraft in the image, and then filling the sky and ground areas with B_background, yields X_space.

[0045] Step nine: Based on the aforementioned dynamic optimization strategy and the aforementioned target surrogate model set, the optimization objective is to minimize the preset loss function. Gradient optimization is then performed on the spatially transformed image by dynamically adjusting the step size and momentum factor to obtain adversarial perturbation data. The preset loss function is typically the classification loss function of the target surrogate model, such as cross-entropy loss. Minimizing this loss function means reducing the model's classification confidence in the input spatially transformed image or causing misclassification. Gradient optimization refers to using the target surrogate model set loaded in step one to calculate the gradient of the loss function with respect to the input image (i.e., X_space) and updating a perturbation tensor initialized to zero according to a specific optimization algorithm. Dynamically adjusting the step size and momentum factor means calculating the step size and momentum value of the current iteration in real time during the optimization iteration process based on the parameters provided in the aforementioned dynamic optimization strategy. In practice, the executing entity first extracts key parameters from the aforementioned dynamic optimization strategy: total number of iterations T, upper limit of perturbation amplitude ε, initial step size α_0, dynamic step size decay coefficient γ, and initial momentum factor value β_0. Subsequently, the adversarial perturbation tensor δ and momentum tensor g are initialized to zero. In the iterative loop from 1 to T, each step first follows the formula α_t=α_0. exp(-γ (t-1) Dynamically calculate the current step size α_t. Then, according to the batch size N specified by the strategy, perform multiple transformations on the spatial transformation image X_space to construct a small batch of input, and calculate the average gradient of this batch on all target surrogate models. J. Then, using the formula g=β_0 g+(1-β_0) J updates the momentum tensor g (where β_0 can remain constant or decay during iteration). Using the updated momentum and the current step size, according to δ=δ+α_t... The `sign(g)` function updates the adversarial perturbation `δ`. After each update, the operation `δ=clip(δ,-ε,ε)` is immediately executed to clip each element of the perturbation `δ` to the range [-ε,ε] to ensure that the perturbation amplitude constraint is satisfied. After the loop ends, the final perturbation tensor `δ` is the desired adversarial perturbation data.

[0046] Step 105: In response to the type determination result being represented as video, the acquired remote sensing data is subjected to a second perturbation process based on a dynamic optimization strategy to obtain an adversarial perturbation data sequence.

[0047] In some embodiments, the execution entity may, in response to the type determination result being characterized as video, perform a second perturbation process on the acquired remote sensing data based on the dynamic optimization strategy to obtain an anti-perturbation data sequence.

[0048] In some optional implementations of certain embodiments, the aforementioned execution entity may, in response to the aforementioned type determination result being characterized as video, perform a second perturbation process on the acquired remote sensing data based on the aforementioned dynamic optimization strategy to obtain an adversarial perturbation data sequence: Step one involves extracting keyframes from the aforementioned remote sensing data, resulting in a keyframe set and a non-keyframe index. Here, the remote sensing data specifically refers to a remote sensing video sequence composed of multiple frames arranged chronologically. Keyframes are frames in the video sequence where the content changes significantly or contains high information, such as scene transitions, the appearance of new targets, or sudden changes in target motion. The keyframe set is a subsequence composed of all extracted keyframes. The non-keyframe index is a list of the temporal positions (e.g., frame numbers) of frames in the video sequence that were not selected as keyframes. In practice, the execution entity can employ various strategies for keyframe extraction. A common method is an adaptive thresholding method based on inter-frame difference: calculating the difference between each pair of consecutive frames in the video (e.g., calculating the sum of the absolute differences of normalized pixel values ​​between two frames, or calculating the Bach distance of their grayscale histograms). Subsequently, the mean μ and standard deviation σ of all inter-frame differences in the entire video are calculated, and a threshold T = μ + k is set. σ, where k is a predefined coefficient (e.g., k=1.5). When the difference between two consecutive frames first exceeds this threshold T, the subsequent frame is marked as a keyframe. Another approach combines quantitative decision-making with object detection: using a lightweight object detection model (such as a lightweight version of YOLO) to process each frame, obtaining the object bounding box and category. Define a change metric: for example, calculate the change in the overlap (IoU) of the object boxes between adjacent frames, or statistically analyze the differences in the object category sets. Set a quantitative threshold (e.g., an IoU decrease of more than 50%, or the appearance of a new category); when the change metric exceeds this threshold, the current frame is marked as a keyframe. The first frame of the video is usually selected as a keyframe by default. After extraction, the video is divided into keyframes (forming the keyframe set) and transition frames located between keyframes (whose indices form the non-keyframe index).

[0049] Step two involves performing the first perturbation process on each keyframe in the aforementioned keyframe set to obtain keyframe adversarial perturbation data for each keyframe. Specifically, when performing the discrete wavelet transform and random linear scaling on any keyframe, a mean square error is generated between the low-frequency components of the current keyframe and the previous keyframe. This mean square error is then incorporated as a loss term into the optimization objective. The first perturbation process refers to the single-image-oriented frequency-space cooperative modulation (SFCM) adversarial example generation process defined in step 104. The keyframe adversarial perturbation data is a perturbation tensor generated independently for each keyframe, conforming to visual concealment constraints. The discrete wavelet transform and random linear scaling are specific steps within the first perturbation process. The mean square error is the average of the squared differences between corresponding elements of two matrices (here, the low-frequency component matrix). The optimization objective refers to the loss function that needs to be minimized during the gradient optimization process in step 104. In practice, the aforementioned execution entity sequentially executes step 104 for each frame in the keyframe set, but enhances it in two key aspects: 1) Temporal consistency enhancement: When processing the k-th keyframe (k>1), while executing the step of "randomly linearly scaling the aforementioned low-frequency components," it additionally calculates the mean square error (MSE(LL_k,LL_{k-1}) between the low-frequency component LL_k of the current frame and the low-frequency component LL_{k-1} of the previous keyframe. 2) Loss function expansion: Based on the original optimization objective (such as classification loss L_cls) in step 104, a feature consistency loss term composed of this mean square error is added. Specifically, the new optimization objective becomes: L_total=L_cls+α MSE(LL_k, LL_{k - 1}). Here, α is a weight coefficient used to balance the attack intensity and temporal smoothness. The setting of this coefficient can be based on experience: for example, for remote sensing video tasks, a grid search can be performed within the range of 0.1 to 1.0, or it can be dynamically fine-tuned according to the degree of motion in the video content (the more intense the motion, the appropriate increase in α to strengthen the smoothness constraint). This constraint forces the generated adversarial perturbations to maintain a smooth transition between adjacent key frames in the frequency domain (especially the low-frequency part representing the main structure), thereby significantly enhancing the temporal coherence of the adversarial video and avoiding inter-frame jitter.

[0050] Step 3: For each non-key frame included in the above non-key frame indices, according to the time stamp corresponding to the above non-key frame, perform linear interpolation on the two key frame adversarial perturbation data corresponding to two adjacent key frames to obtain the non-key frame adversarial perturbation data corresponding to the above non-key frame. Here, the time stamp corresponding to the above non-key frame refers to the absolute time position or relative frame number of the frame in the video sequence. The above linear interpolation is a mathematical method for estimating the value of an intermediate point based on the data of two known points. In this context, it means using the known adversarial perturbation data of the previous and the next adjacent key frames in time to estimate the perturbation data of a non-key frame located between them. In practice, assume that the time stamp of non-key frame F_t is t, and it is located between key frame A with time stamp t_a and key frame B with time stamp t_b (t_a < t < t_b), and the key frame adversarial perturbation data δ_a and δ_b are already available. The executor calculates an interpolation weight w = (t - t_a) / (t_b - t_a). Then, perform a linear interpolation calculation on the adversarial perturbation data δ_t of this non-key frame: δ_t = (1 - w) δ_a + w δ_b. This means that δ_t is a weighted average of δ_a and δ_b, and the weights are determined by the position of F_t on the time axis relative to A and B. This operation can efficiently generate perturbations for all non-key frames and ensure that these perturbations vary smoothly between the "anchor points" defined by adjacent key frames. For example, the key frames are at frame 10 (δ_10) and frame 20 (δ_20). For the non-key frame at frame 15, its interpolation weight w = (15 - 10) / (20 - 10) = 0.5, then its perturbation δ_15 = 0.5 δ_10 + 0.5 δ_20.

[0051] Step four: Based on the temporal sequence of the aforementioned remote sensing data, integrate the obtained adversarial perturbation data for each keyframe with the adversarial perturbation data for each non-keyframe to obtain an adversarial perturbation data sequence. The temporal sequence of the aforementioned remote sensing data refers to the original playback order of the video frames. The integration refers to arranging the adversarial perturbation data corresponding to all frames generated in steps two and three into an ordered data list or tensor according to the original order of the video frames. In practice, the execution entity creates a list or a four-dimensional tensor (dimensions: [total number of frames, image height, image width, number of channels]). Then, according to the frame numbers from 1 to N (N being the total number of frames), fill the corresponding perturbation data (whether keyframe or non-keyframe adversarial perturbation data) into the corresponding positions. Finally, this ordered set of perturbations corresponding frame-by-frame to the original video is the aforementioned adversarial perturbation data sequence. For example, in a 300-frame video, δ_1 (frame 1, keyframe), δ_2 (interpolation), δ_3 (interpolation) ... δ_15 (keyframe) ... δ_300 (keyframe or interpolation) are combined in sequence to obtain a sequence containing 300 perturbation tensors.

[0052] Step 106: Overlay the obtained adversarial perturbation data or adversarial perturbation data sequence with the acquired remote sensing data to obtain adversarial samples.

[0053] In some embodiments, the aforementioned executing entity may overlay the obtained adversarial perturbation data or adversarial perturbation data sequence with the acquired remote sensing data to obtain adversarial examples. The adversarial perturbation data is a perturbation tensor generated for a single remote sensing image, with the same size as the remote sensing image. The adversarial perturbation data sequence is a sequence of perturbation tensors generated for remote sensing video, arranged in chronological order, where the size of each perturbation tensor is the same as that of a video frame. The overlay process refers to applying the perturbation to the original data in a pixel-by-pixel addition manner. The adversarial example is the data obtained after overlay processing, intended to cause the target recognition model to produce incorrect outputs, and its form (single image or video sequence) corresponds to the original remote sensing data. In practice, the aforementioned executing entity performs specific overlay calculations. For the image case, the entity adds the adversarial perturbation data (denoted as δ) obtained in step 104 to the acquired remote sensing data (i.e., the original image, denoted as x) element-wise: x_adv = x + δ. For video scenarios, the entity performs a frame-by-frame, element-by-element addition of the adversarial perturbation data sequence (denoted as [δ_1,δ_2,...,δ_T]) obtained in step 105 with the acquired remote sensing data (i.e., the original video frame sequence, denoted as [x_1,x_2,...,x_T]): for each frame t, x_adv_t = x_t + δ_t is calculated, thus obtaining the adversarial video sample sequence [x_adv_1,x_adv_2,...,x_adv_T]. This superposition operation ensures that the perturbation, which has been constrained within the preset perturbation amplitude upper limit (ε) in steps 103 and 104, is precisely applied. The final generated adversarial samples (x_adv or [x_adv_t]) are typically organized into a multidimensional array (such as [H,W,C] or [T,H,W,C]) with the same dimensions as the original input data and stored in memory, or packaged into standard image / video file formats (such as PNG, MP4) for output for subsequent use or evaluation. For example, for a remote sensing image with a size of 1024x1024, its adversarial perturbation data δ is a tensor of 1024x1024x3. The executing agent reads the original image data x (also 1024x1024x3), performs the addition operation x_adv=x+δ, and generates an adversarial sample image x_adv, also of 1024x1024x3. For a 30-frame video, each frame being 1024x1024, the adversarial perturbation data sequence contains 30 tensors of 1024x1024x3. The executing agent reads the original video frames x_t frame by frame and performs the addition operation x_adv_t=x_t+δ_t 30 times, ultimately generating an adversarial sample video consisting of 30 adversarial frames.

[0054] Step 107: Evaluate the attack effectiveness of the adversarial sample and obtain the evaluation results.

[0055] In some embodiments, the aforementioned executing entity can evaluate the attack effectiveness of the aforementioned adversarial sample to obtain an evaluation result. The aforementioned attack effectiveness evaluation refers to the process of measuring the effectiveness and visual concealment of the generated adversarial sample in deceiving the target recognition model through quantitative indicators. The evaluation result is a structured data object containing one or more quantitative indicator values, used to objectively and comparatively reflect the overall effect of this adversarial attack. In practice, the aforementioned executing entity completes the evaluation through the following process: First, the adversarial sample generated in step 106 is fed as input to one or more predetermined evaluation models. These evaluation models include a proxy model used when generating the adversarial sample (used to evaluate the attack effect of white-box or known models), and one or more independent black-box models with different structures and training data that were not used in the generation process (used to evaluate the effect of cross-model transfer attacks). For each evaluated model, the executing entity calculates its prediction result for the adversarial sample and compares it with the true label of the original remote sensing data to calculate the core indicator—attack success rate, i.e., the proportion of samples where the model makes incorrect predictions about the adversarial sample. Secondly, to assess the visual stealth of the adversarial example (i.e., its imperceptibility to the human eye), the executing agent calculates a structural similarity index between the adversarial example and the original remote sensing data. This index compares the brightness, contrast, and structural information of the two images, yielding a value between -1 and 1. The closer the value is to 1, the smaller the visual difference and the better the stealth. Furthermore, depending on the task requirements, other auxiliary indicators can be calculated, such as the average amplitude of the perturbation and the spatial distribution entropy of the perturbation. Finally, the executing agent integrates all calculated indicator values ​​into an evaluation report or a structured data object (such as a dictionary), as the output of the above evaluation results. This result can be directly used for manual judgment or as input signals for strategy iterative optimization in step 108. For example, for a video adversarial example consisting of 300 adversarial frames, the executing agent first uses the proxy model ResNet-50 to classify each frame: if 270 of the 300 frames are misclassified, the attack success rate against this model is 90%. Next, the same test was performed using an untrained black-box model, VGG-16. If the attack success rate was 75%, then the cross-model transfer attack success rate was also 75%. Simultaneously, the SSIM index of each adversarial frame relative to the corresponding original frame was calculated, and the average of 300 frames was taken, resulting in an average SSIM of 0.92. Finally, the evaluation results were generated: {"white_box_asr":0.90,"black_box_asr":0.75,"avg_ssim":0.92}.

[0056] Step 108: Based on the performance evaluation results and the acquired real-time hardware status data, iteratively optimize the dynamic optimization strategy to obtain a set of hardware control parameters.

[0057] In some embodiments, the execution entity may iteratively optimize the dynamic optimization strategy based on the performance evaluation results and the acquired real-time hardware status data to obtain a set of hardware control parameters.

[0058] In some optional implementations of certain embodiments, the aforementioned execution entity may iteratively optimize the aforementioned dynamic optimization strategy based on the performance evaluation results and the acquired real-time hardware status data through the following steps to obtain a set of hardware control parameters: Step 1: In response to the attack success rate in the performance evaluation results being lower than the preset success rate threshold, the upper limit of the perturbation amplitude and the number of iterations in the dynamic optimization strategy are increased. The attack success rate (ASR) refers to the proportion of adversarial examples that successfully mislead the target model into producing erroneous outputs; it is a core indicator for measuring attack effectiveness. The preset success rate threshold is a pre-set, expected minimum attack success rate target value, such as 85%. The upper limit of the perturbation amplitude is a parameter in the dynamic optimization strategy that controls the maximum perturbation amplitude. The number of iterations is a parameter in the dynamic optimization strategy that controls the number of gradient optimization update rounds. In practice, the executing entity compares the attack success rate obtained in step 107 with the preset success rate threshold. If the attack success rate is lower than this threshold, it indicates that the current attack strength or optimization sufficiency is insufficient. To improve attack effectiveness, the executing entity will adjust the key parameters in the dynamic optimization strategy upwards according to certain incremental rules: for example, increasing the upper limit of the perturbation amplitude by a fixed step size (e.g., increasing by 2 / 255) and increasing the number of iterations by a fixed value (e.g., increasing by 10 times). This adjustment aims to allow for the generation of stronger perturbations and to give the optimization process more update rounds to find more effective attack directions, thereby hoping to improve the success rate in the next attack.

[0059] Step two: In response to the structural similarity index in the performance evaluation results being lower than the preset similarity index threshold, the upper limit of the perturbation amplitude in the dynamic optimization strategy is reduced. The Structural Similarity Index (SSIM) is an indicator used to measure the visual similarity between two images; a value closer to 1 indicates a smaller visual difference. The preset similarity index threshold is a pre-defined minimum visual concealment standard that the adversarial example must meet, for example, 0.90. In practice, the executing entity compares the structural similarity index obtained in step 107 with the preset similarity index threshold. If the index is lower than the threshold, it indicates that the currently generated adversarial example has excessive visual distortion and insufficient concealment. To generate more difficult-to-detect adversarial examples in the next attack, the executing entity will lower the upper limit of the perturbation amplitude in the dynamic optimization strategy: for example, by reducing it by a fixed step size (e.g., reducing it by 1 / 255). By limiting the maximum perturbation allowed, the subsequent optimization process is forced to find effective attack perturbations under stricter visual constraints, thereby improving concealment while maintaining a certain level of aggressiveness.

[0060] Step 3: Based on the current power consumption and temperature in the real-time hardware status data, adjust the dynamic step size decay coefficient and the initial value of the momentum factor in the dynamic optimization strategy. Here, the current power consumption refers to the real-time power consumption of the onboard computing platform (such as SoC, GPU). The temperature refers to the real-time operating temperature of the critical processing chip (such as CPU, GPU). The dynamic step size decay coefficient is a parameter that controls the exponential decay rate of the optimization step size. The initial value of the momentum factor is the initial weight of the momentum term in the gradient optimization algorithm. In practice, the execution entity monitors the power consumption and temperature in the real-time hardware status data. If the current power consumption or temperature exceeds a preset safety threshold, it indicates that the system may face the risk of overheating or power overload. To reduce computational intensity and alleviate thermal load, the execution entity will adjust the optimization parameters in the dynamic optimization strategy to guide the algorithm to converge faster. Specifically, the dynamic step size decay coefficient will be increased (e.g., from 0.03 to 0.1) to accelerate the decay of the optimization step size, thereby ending the high-intensity gradient update iterations more quickly. Simultaneously, the initial value of the momentum factor will be decreased (e.g., from 0.9 to 0.7) to reduce the optimization process's dependence on historical gradient directions, allowing for more flexible direction adjustments and avoiding unnecessary computational load during local oscillations. Conversely, if power consumption and temperature are at safe and low levels, the decay coefficient can be appropriately reduced and the initial momentum value increased for a more refined search.

[0061] Step four: Based on the memory and bandwidth utilization rates in the real-time hardware status data, dynamically adjust the number of proxy models and batch size in the dynamic optimization strategy. The memory utilization rate refers to the proportion of memory used by the GPU or AI accelerator card. The bandwidth utilization rate refers to the busy level of data transfer on the memory or data bus. The number of proxy models and batch size are two key parameters controlling the use of computing resources in the dynamic optimization strategy. In practice, the execution entity monitors the memory and bandwidth utilization rates in real time. If the utilization rate consistently exceeds a certain high threshold (e.g., 85%), it indicates resource scarcity, and continuing to use the current configuration may lead to memory overflow or performance bottlenecks. In this case, the execution entity will reduce the number of proxy models (e.g., from 4 to 2) and / or the batch size (e.g., from 8 to 4) in the dynamic optimization strategy to release resources and ensure stable system operation. Conversely, if the utilization rate remains below a certain low threshold (e.g., 50%), it indicates surplus resources, and the number of proxy models or the batch size can be appropriately increased to utilize more parallel computing power, potentially improving the mobility or generation efficiency of attacks. The adjustments follow the resource constraint model to ensure that the new parameter combination is within the range of available resources.

[0062] Step 5: Based on the adjusted dynamic optimization strategy, generate a set of hardware control instructions. This set of instructions is a group of low-level commands used to directly control the runtime parameters of various hardware units on the airborne computing platform. In practice, the executing entity takes the dynamic optimization strategy (containing all updated parameters) obtained from steps 1 to 4 as input and converts it into specific, executable hardware control instructions through an instruction mapping module. For example, based on the updated iteration count and the number of proxy models, instructions are generated to adjust the frequency and voltage of the task processor (e.g., GPU) computing core to match the expected computing load; and based on the adjusted batch size and data processing requirements, instructions are generated to reconfigure the memory access mode or DMA (Direct Memory Access) channel.

[0063] Step six involves converting the aforementioned hardware control instruction set into a hardware control parameter set. This hardware control parameter set is a structured, machine-readable data representation of the hardware control instruction set, typically directly writable into hardware configuration registers or driver layer interfaces. In practice, the executing entity encodes and serializes the generated hardware control instruction set, converting it into a specific data format recognizable by the target hardware platform driver or firmware. For example, this might be converted into a specific set of register address-value pairs, a structured configuration file (such as JSON or XML), or a series of data packets conforming to a specific bus protocol (such as I2C or SPI). This final hardware control parameter set is then sent to the corresponding hardware controller to physically adjust processor frequency, memory allocation, communication bandwidth, etc., thereby achieving coordinated configuration of computing resources and optimization strategies.

[0064] Step 109: Adjust each processor and each controlled hardware in the airborne computing platform according to the set of hardware control parameters.

[0065] In some embodiments, the aforementioned execution entity may adjust each processor and each controlled hardware in the airborne computing platform according to the aforementioned set of hardware control parameters.

[0066] In some optional implementations of certain embodiments, the aforementioned execution entity may adjust each processor and each controlled hardware in the airborne computing platform according to the aforementioned set of hardware control parameters through the following steps: The first step is to parse the current flight mission instructions of the UAV to obtain mission intent information, which includes the mission phase and the desired stealth level. The flight mission instructions refer to the set of instructions pre-planned or issued in real-time by the ground control station to the UAV flight control system, defining the mission objectives and procedures. The mission intent information is extracted from this information and contains semantic information related to the scheduling and strategy of adversarial sample generation. The mission phase refers to the specific stage the UAV is in during the mission, such as "cruise reconnaissance," "covert approach," "close observation," and "rapid disengagement." The desired stealth level is a quantitative indicator used to define the required level of electromagnetic and visual signal detectability of the UAV at the current phase, such as "high," "medium," or "low." In practice, the executing entity obtains a structured mission description by accessing the mission management module of the flight control system or parsing specific mission instruction data streams. From this, the "mission phase" field corresponding to the current moment and the "stealth level" parameter set in the mission planning or dynamically determined based on the phase are extracted, together constituting the mission intent information, which serves as the context for subsequent dynamic adjustment decisions.

[0067] The second step involves acquiring real-time flight status data and communication interference intensity data for the UAV. The real-time flight status data refers to parameters describing the UAV's motion and attitude, such as flight speed, acceleration, angular velocity, altitude, and heading angle, provided in real-time by sensors from the UAV's flight control system and inertial measurement unit. The communication interference intensity data refers to the strength and characteristics of interference signals targeting the UAV's communication frequency band in the current electromagnetic environment, monitored and assessed by onboard electronic support measurement equipment or data link receivers. In practice, the aforementioned entities subscribe to status telemetry messages issued by the flight control system via an internal data bus (such as CAN or Ethernet) to acquire flight status data in real time. Simultaneously, they obtain real-time monitored spectrum data and interference assessment reports from electronic warfare support units or communication systems through dedicated interfaces, extracting quantitative indicators characterizing the current communication interference intensity.

[0068] The third step involves generating correction instructions that, in response to the aforementioned task phase being "covert approach" and the desired covert level being "high," represent a reduction in the computational core frequency and an increase in the real-time priority parameter. Furthermore, the current hardware control parameter set is updated based on the generated correction instructions. Here, the "covert approach" phase refers to the flight phase where the UAV attempts to gradually approach the target area without being detected by enemy detection systems. The "computational core frequency" refers to the clock frequency of the onboard task processor (such as a GPU or NPU). The "real-time priority parameter" refers to the CPU time slice priority assigned to the adversarial example generation task relative to other system tasks when scheduling it in the operating system. In practice, when the parsed task intent information indicates a "covert approach" phase with a "high" covert level, it means that it is necessary to minimize electromagnetic radiation generated by computation (such as high-frequency processor operation) and ensure the real-time response capability of adversarial example generation to avoid exposure. Therefore, the executing entity generates correction instructions: on the one hand, the instruction task processor dynamically reduces the operating frequency of its computing cores to an energy-saving silent level to reduce electromagnetic leakage; on the other hand, in the operating system's task scheduler, the scheduling priority of the adversarial sample generation process is raised to the highest real-time level to ensure timely access to computing resources even under low computing power, avoiding task lag that could cause the attack to fail. Subsequently, the set of hardware control parameters to be executed is updated according to these instructions.

[0069] The fourth step involves generating a correction instruction, in response to the aforementioned communication interference intensity data indicating the presence of active electromagnetic interference, to increase anti-interference coding resources and decrease the visual disturbance intensity parameter. The current hardware control parameter set is then updated based on the generated correction instruction. Here, active electromagnetic interference refers to the deliberate emission of strong interference signals by an adversary in an attempt to block or disrupt the UAV's communication link. Anti-interference coding resources refer to the time-frequency resources, processing power, or coding gain used in the data link communication system to implement anti-interference functions (such as frequency hopping and direct sequence spread spectrum). The visual disturbance intensity parameter refers to the upper limit of the disturbance amplitude in the dynamic optimization strategy, which directly affects the visual difference between the adversarial example and the original image. In practice, when the detected communication interference intensity exceeds a predetermined threshold, indicating the presence of active interference, it means that the data transmission link may be unstable. To ensure the reliable transmission or reception of generated adversarial examples or attack commands, the executing entity generates correction instructions: First, the instruction data link system increases the redundancy of anti-jamming coding or switches to a more robust anti-jamming waveform, consuming more resources to ensure communication robustness. Second, considering that the enemy may rely more on photoelectric detection in a strong jamming environment, to reduce the probability of being detected by the optical system, the visual concealment of the adversarial examples needs to be further enhanced. Therefore, the instructions reduce the allowed visual perturbation intensity parameter (i.e., the upper limit of perturbation amplitude) when generating adversarial examples to generate more difficult-to-detect adversarial examples. Subsequently, the hardware control parameter set is updated to incorporate these adjustments.

[0070] The fifth step involves generating a correction command, in response to the flight status data indicating a high-speed maneuvering state, to increase the disturbance amplitude compensation parameter. The current hardware control parameter set is then updated based on this correction command. The high-speed maneuvering state refers to the UAV undergoing drastic attitude changes or high-speed movements, such as sharp turns, climbs, or dives. The disturbance amplitude compensation parameter is a dynamic incremental factor introduced on top of the original disturbance amplitude upper limit. It compensates for the decrease in attack effectiveness caused by the blurring, deformation, or displacement of the target in the image due to the platform's drastic movement. In practice, when flight status data (such as angular velocity and acceleration) exceeds a set threshold, the UAV is determined to be in a high-speed maneuvering state. At this time, the image captured by the airborne camera may exhibit motion blur, and the target characteristics may change. To ensure that the generated adversarial disturbance remains effective on such degraded images, the main generation correction command is executed: based on the original disturbance amplitude upper limit of the dynamic optimization strategy, a compensation coefficient greater than 1 (e.g., 1.2) is temporarily multiplied, effectively "increasing" the allowable disturbance amplitude and providing greater optimization space for disturbance generation to overcome the effects of motion. This compensation parameter has been integrated into the updated set of hardware control parameters.

[0071] Step 6: Based on the current set of hardware control parameters, execute the following control steps: The first sub-step involves adjusting the bandwidth allocation of the communication bus between the flight control processor and the task processor, and allocating high-priority time slices for the adversarial example generation task. In practice, the aforementioned execution entity, based on the configuration in the parameter set, dynamically adjusts the bandwidth allocation ratio of the internal high-speed bus (such as PCIe) connecting the flight control processor (responsible for flight control) and the task processor (responsible for adversarial example calculation) through the operating system kernel module or hardware resource management driver, reserving or increasing data transmission bandwidth for the task processor. Simultaneously, in the task processor's operating system scheduler, the scheduling policy for the adversarial example generation process is set to real-time (such as SCHED_FIFO), and a high-priority time slice is allocated to it, ensuring that its computational tasks can be responded to in a timely manner and reducing latency caused by resource contention.

[0072] The second sub-step involves dynamically allocating the computing resources of the graphics subprocessor and the neural network subprocessor within the task processor for executing adversarial example generation tasks. In practice, the aforementioned execution entity dynamically allocates the computing resources within heterogeneous task processors (such as GPUs with CUDA cores and NPUs with Tensor cores) based on the emphasis on different computing tasks in the parameter set. By calling the vendor-specific management library, the operating frequency, voltage, and computing task queue quotas are set for the general-purpose graphics computing unit (graphics subprocessor) and the AI-specific acceleration unit (neural network subprocessor), respectively. This precisely controls the resource allocation ratio between the two types of processors when performing image transformation, wavelet calculation (potentially biased towards the graphics subprocessor), and neural network gradient calculation (biased towards the neural network subprocessor), thereby optimizing energy efficiency and performance.

[0073] The third sub-step involves configuring the anti-jamming communication waveform and its transmission power for the data link transmitter. In practice, the aforementioned execution entity, based on the instructions in the parameter set, configures the transmitter's operating waveform through the data link device's control interface. For example, it switches from a conventional communication mode to an anti-jamming waveform mode with a frequency hopping pattern or spread spectrum sequence. Next, it precisely sets the transmission power of this anti-jamming waveform. This might involve appropriately increasing the power in an interference environment to ensure link budget, or reducing the power to decrease radiated signals under cover requirements.

[0074] The fourth sub-step involves driving the optoelectronic payload servo mechanism to adjust the gimbal's pitch and yaw angles, and simultaneously setting the camera integration time. In practice, the aforementioned execution entity, based on specific requirements for imaging quality or field of view that may be included in the parameter set (derived from mission intent or adversarial sample generation needs), sends angle control commands to the gimbal servo mechanism via the optoelectronic payload's control bus, driving it to adjust the pitch and yaw angles, thereby changing the camera's pointing direction. Simultaneously, to match the adjusted lighting conditions or motion state (such as high-speed maneuvering), the image sensor's integration time (exposure time) is simultaneously adjusted via the camera control interface to ensure that the acquired remote sensing image quality meets the basic requirements for subsequent adversarial attack algorithm processing.

[0075] Steps one through six of this disclosure address the technical problem that "in a dynamic, resource-constrained, and uncertain real battlefield environment, the dynamic optimization strategies generated by the algorithm of a UAV-based remote sensing cross-modal adversarial attack system cannot be automatically and accurately translated into physical control commands for complex heterogeneous hardware systems. This results in the algorithm's performance being unable to be fully utilized due to improper hardware resource configuration, and the overall platform behavior (computation, communication, detection) becoming disconnected from rapidly changing tactical intentions and environmental threats." Existing technologies have the following shortcomings in the above aspects: On the one hand, the algorithm optimization layer and the hardware control layer are often decoupled, with optimization parameters remaining at the software level, lacking a bridge to translate them into specific hardware control actions (such as bus bandwidth, core frequency, and waveform parameters); on the other hand, the system's response to higher-order mission intentions (such as "covert approach") and real-time physical states (such as "high-speed maneuvering" and "strong interference") is delayed or fragmented, failing to inject this contextual information into the hardware control loop in real time to achieve joint optimization of attack effectiveness, stealth, and platform survivability. Solving the above problems will enable a seamless transition from "intelligent algorithm decision-making" to "smart hardware execution," allowing the UAV's computing, communication, and sensor resources to respond instantly to the "brain's" (algorithm and mission planning) commands, much like "muscles," achieving a balance between maximizing attack effectiveness and minimizing platform risk in complex scenarios. To achieve this, this disclosure proposes the following: First, parsing the UAV's current flight mission commands to obtain mission intent information. By proactively understanding the UAV's stage in the macro-mission flow (such as cruise, approach, and disengagement) and its corresponding stealth level requirements, the system, for the first time, transforms high-level, semantic tactical intent into quantifiable inputs that drive hardware strategy adjustments. This solves the problem of hardware control lacking mission context awareness, making subsequent hardware resource configuration no longer static or blind, but endowed with clear tactical objectives (e.g., pre-setting a "silent priority" hardware configuration tone for the "stealthy approach" phase), achieving semantic alignment between mission planning and hardware behavior. Second, acquiring the UAV's real-time flight status data and communication interference intensity data. Thus, by continuously collecting data on the platform's own kinematic state and the degree of external electromagnetic environmental stress, the system establishes a real-time sensing channel for both the internal dynamic environment and the external signal environment. This solves the problem of hardware control decisions being detached from actual physical conditions, providing crucial realistic constraints and optimization basis for dynamic adjustments (e.g., knowing that it is undergoing violent maneuvers or is in a highly interfering environment), ensuring the physical feasibility and environmental adaptability of the control strategy. Thirdly, in response to the aforementioned task stage being a stealthy approach with a high desired stealth level, a correction instruction is generated to reduce the computational core frequency and increase the real-time priority parameter.Therefore, when the system identifies a tactical scenario requiring extreme concealment, it can automatically execute a pair of seemingly contradictory but actually synergistic fine-tuning adjustments: reducing processor frequency to decrease electromagnetic radiation signature, while simultaneously increasing task scheduling priority to ensure timely computational response under low computing power. This solves the problem that computational tasks may lag or fail due to performance limitations in scenarios requiring electromagnetic silence. Through "software-hardware synergy," the continuity of the attack mission is maximized while meeting the rigid constraints of concealment. The fourth step involves generating corrective instructions to increase anti-interference coding resources and reduce visual disturbance intensity parameters in response to active electromagnetic interference. Thus, when facing communication link threats, the system can simultaneously implement a dual strategy of "strengthening communication resilience" and "enhancing optical concealment." On the one hand, it resists interference by increasing coding redundancy or switching to stronger waveforms, ensuring reliable transmission of control commands and data; on the other hand, recognizing that the enemy may strengthen optical reconnaissance when communication is disrupted, it proactively reduces the visual disturbance intensity of adversarial examples, reducing the probability of being detected by optoelectronic systems. This solves the problem of being overwhelmed by complex threats and achieves synergistic concealment enhancement across the communication and visual domains. The fifth step involves generating a correction command to increase the perturbation amplitude compensation parameters in response to high-speed maneuvering. This allows the system to proactively compensate for the attack performance degradation caused by image blurring and deformation when the UAV's violent maneuvers lead to image quality degradation. By dynamically allowing the generation of larger perturbations, a "tolerance space" is provided for image degradation in adversarial example optimization algorithms. This solves the problem of the negative impact of platform physical motion on the attack algorithm's effectiveness, giving the attack system "immunity" to its own platform's motion state and maintaining the stability of attack performance. The sixth step involves executing the following control steps based on the current set of hardware control parameters. This transforms the final control parameters, which integrate task intent, environmental state, and compensation strategies, into tangible hardware state changes through the coordinated execution of the following four dimensions: The first sub-step (adjusting bus bandwidth and task priority) ensures unobstructed access to critical data paths required for the attack computation task and guarantees a definite computation time at the operating system level, resolving the latency uncertainty caused by internal competition for computing resources and providing a "green channel" for real-time attacks. The second sub-step (dynamically allocating heterogeneous computing resources) allocates resources precisely between the general-purpose computing cores of the GPU and the dedicated AI cores of the NPU, based on the different characteristics of image transformation and neural network computation. This solves the problem of low resource utilization or rigid configuration of heterogeneous computing platforms, achieving optimal matching of energy efficiency and performance. The third sub-step (configuring anti-interference waveforms and power) drives the robustness configuration of the communication link directly from software parameters to the RF hardware front end, enabling agile physical layer reconfiguration of communication strategies. This allows the UAV to quickly switch between modes such as conventional communication, covert communication, and anti-interference communication, dynamically adapting to the electromagnetic environment.The fourth sub-step (adjusting the gimbal angle and camera integration time) incorporates the control of the sensing payload into a unified closed loop. This allows for adjustments to the observation field of view based on mission intent and optimization of imaging parameters based on platform motion status, resolving the disconnect between hardware control in the "perception" and "attack" stages. This ensures that the quality of the acquired raw images meets the input requirements for high-precision adversarial attacks. In summary, step 109 of this embodiment constructs a complete, closed-loop hardware control chain from "tactical intent analysis" to "environmental state perception," then to "multimodal strategy generation," and finally to "cross-domain hardware collaborative execution." It acts as the "neuro-muscle" interface between intelligent algorithms and the physical world. Through this series of steps, the UAV is no longer a platform passively executing preset programs, but rather an intelligent organism capable of understanding the mission, perceiving the environment, and dynamically adjusting all its internal key hardware subsystems (computing, communication, and perception) to collaboratively achieve optimal attack effectiveness. This not only significantly improves the success rate and adaptability of adversarial attacks in complex dynamic environments but also fundamentally enhances the autonomy, survivability, and mission reliability of the UAV as an intelligent tool.

[0076] In addressing the aforementioned issues of adaptability, real-time performance, and stealth in deploying cross-modal adversarial samples on real UAV platforms using the remote sensing cross-modal adversarial sample generation method for UAVs, the following technical problem often arises when performing long-range adversarial attack missions in complex, highly adversarial battlefield environments with active electromagnetic countermeasures, network countermeasures, and physical threats: During the high-computational-intensity adversarial sample generation process, the system itself continuously generates specific electromagnetic radiation, communication traffic, and platform behavioral characteristics. On the one hand, these characteristics are easily captured by enemy electronic reconnaissance equipment, thus exposing the UAV's attack intent and real-time location; on the other hand, once the attack behavior is detected, it may trigger the enemy's active electronic jamming, network traffic filtering, or targeted physical countermeasures. This not only leads to the failure of the current adversarial attack mission but may also jeopardize the battlefield survivability of the carrier UAV platform. To address the following requirements for this application scenario: While ensuring the effectiveness of attack missions, it must possess real-time awareness of multi-dimensional battlefield threats, dynamic assessment of its own exposure risks, and rapid, autonomous, and tiered survival response capabilities in the event of countermeasures, we have decided to adopt the following solution: Optionally, the aforementioned implementing entity may also perform the following steps: The first step, in response to the start of generating counter-interference data or counter-interference data sequences, involves continuously scanning the operating frequency band using airborne electronic support measurement equipment, and conducting a risk assessment of the scanned operating frequency band to obtain a risk level. Here, the airborne electronic support measurement equipment refers to an electronic reconnaissance system used for passively receiving and analyzing electromagnetic environment signals. The operating frequency band refers to the range of radio frequencies used by UAV communication, navigation, and potentially involved countermeasure systems (such as the command link of a target recognition model). The risk assessment involves analyzing the scanned radio frequency signals to determine whether they contain enemy detection, tracking, or jamming intentions. The risk level is a quantitative classification of the current electromagnetic threat level, for example, divided into "no risk," "low risk" (Level 1), "medium risk" (Level 2), and "high risk" (Level 3). In practice, in response to the triggering of step 104 or 105, the aforementioned executing entity initiates a parallel monitoring thread. This thread instructs the electronic support measurement equipment to continuously scan and acquire signals from the preset operating frequency band. The acquired signals undergo feature analysis, including but not limited to: whether the signal strength is abnormally high, whether there are known enemy radar or communication signal characteristics, and whether dense pulse interference occurs. Based on a pre-set rule base or lightweight classification model, the above features are comprehensively scored, and a discrete risk level is finally output. For example, detecting an unknown but stable surveillance radar signal may be rated as "Level 1 Risk"; detecting targeting interference against the data link may be rated as "Level 2 Risk"; and detecting both a fire control radar lock-on signal and high-intensity communication interference simultaneously is rated as "Level 3 Risk".

[0077] The second step involves continuously identifying abnormal response characteristics in the feedback traffic during the attack effectiveness evaluation of the adversarial sample, thereby obtaining countermeasure feature identification data. The aforementioned feedback traffic refers to the response data stream generated by the evaluation model (especially a black-box model, when it exists as an online service) to the input adversarial sample during the attack effectiveness evaluation process in step 107, including returned prediction results, confidence levels, response delays, etc. The aforementioned abnormal response characteristics refer to patterns that may indicate the target system has initiated proactive defense or countermeasures. The aforementioned countermeasure feature identification data is the determination result of whether such characteristics exist. In practice, during the execution of step 107, the aforementioned executing entity analyzes the feedback from the evaluation model while receiving it. Abnormal characteristics include, but are not limited to: 1) Abrupt changes in response consistency: For the same or similar adversarial sample input, the returned label or confidence level shows drastic and unreasonable fluctuations, inconsistent with the normal behavior of the model. 2) Delayed attack detection: A significant and continuous abnormal increase in response time may indicate that the target system is running additional detection algorithms. 3) Defensive output: Returning special error codes or explicitly indicating that an adversarial attack has been detected. The executing entity identifies these features by comparing feedback with historical normal response baselines in real time, or by using a pre-trained anomaly detector, and generates structured countermeasure feature identification data, such as {"Feature Type": "Confidence Jitter", "Confidence": 0.85}.

[0078] The third step involves continuously acquiring platform anomaly data during the execution of either the first or second perturbation process. This platform anomaly data refers to unexpected internal state indicators that may be caused by malware or hardware failures when the onboard computing platform is running adversarial example generation algorithms. In practice, during the core computation cycle of steps 104 or 105, the aforementioned execution entity synchronously monitors multiple levels of the system: 1) Computational anomalies: such as uncorrectable ECC memory errors reported by the GPU or NPU, or large-scale computational errors in the computation cores (a surge in NaN or Inf values). 2) Behavioral anomalies: abnormal patterns in the system call sequence of the task processor, or unauthorized processes attempting to access the adversarial model weight memory region. 3) Performance anomalies: unexplained periodic peaks in processor utilization or power consumption without a significant increase in load. This data is acquired through the system performance monitoring interface, hardware error reporting mechanism, and security-enhanced operating system audit module, and is aggregated into a platform anomaly data stream.

[0079] The fourth step involves generating a comprehensive security posture identifier based on the obtained risk level, countermeasure feature identification data, or platform anomaly status data. This comprehensive security posture identifier is a concise identifier that comprehensively characterizes the overall security threat level faced by the UAV related to the counterattack mission, such as "Level 1 Risk," "Level 2 Risk," and "Level 3 Risk." In practice, the aforementioned implementing entity maintains a security information fusion engine. This engine receives asynchronous data streams from the first, second, and third steps. It performs a weighted evaluation of multi-source information according to predefined threat models and fusion rules. For example, using a rule engine: if the electronic support measurement equipment reports a Level 3 risk, a Level 3 risk identifier is directly generated regardless of other information. If the electronic support measurement reports a Level 2 risk and a clear countermeasure feature is identified from the feedback traffic, it is upgraded to a Level 3 risk identifier. If only internal platform reports a calculation anomaly, and there is no external risk, a Level 2 risk identifier may be generated, indicating a possible hardware failure or underlying software attack. Finally, a unified comprehensive security posture identifier is output to drive subsequent hardware adjustments.

[0080] Fifth, based on the generated security situation indicator, perform the following tiered hardware control operations: The first sub-step, in response to the aforementioned security situation identification indicating a Level 1 risk, involves reducing the generation intensity and frequency of the aforementioned adversarial disturbance data, and controlling the data link transmitter to switch to a backup frequency hopping pattern and reduce transmission power. In practice, when the security situation identification is "Level 1 risk" (low threat), the implementing entity performs preventative downgrade operations to reduce the probability of exposure. First, the generation intensity and frequency are reduced: by dynamically modifying the dynamic optimization strategy in step 103, the "disturbance amplitude upper limit" parameter is lowered by one level (e.g., from 16 / 255 to 12 / 255), and the period or event threshold for triggering a new round of adversarial sample generation is increased, thereby reducing the temporal density of attack behavior and the significance of each attack. Second, the communication strategy is adjusted: the data link system is instructed to switch from the currently used conventional frequency hopping pattern to a preset, less frequently used backup low-interception-probability frequency hopping pattern, while simultaneously reducing the transmission power to the minimum level required to maintain basic link quality, thereby reducing the detectable range of the communication signal.

[0081] The second sub-step, in response to the aforementioned security situation indicator being classified as Level 2 risk, suspends either the first or second disturbance processing and controls the task processor to clear the relevant computational cache and model weights; simultaneously, the flight control processor controls the UAV to execute a preset evasive maneuver. In practice, when the security situation indicator is "Level 2 risk" (medium threat, such as the detection of targeted interference or internal anomalies), the executing entity implements proactive defense and state cleanup operations. First, the ongoing adversarial disturbance generation task in step 104 or 105 is immediately suspended. Then, the task site is cleaned up: the task processor is instructed to release all video memory and RAM allocated for the current attack task and securely erase (e.g., zero out) the copies of the loaded proxy model weights in the cache to prevent potential memory scanning attacks from stealing model information. Simultaneously, a tactical maneuver is executed: an emergency command is sent to the flight control processor via the high-speed bus to trigger a preset evasive maneuver sequence (e.g., a sharp roll combined with an altitude change) to attempt to evade potential tracking or interference sources.

[0082] The third sub-step, in response to the aforementioned security situation assessment indicating a Level 3 risk, involves controlling the data link transmitter to emit a pre-set deceptive signal, and controlling the UAV to fly along a planned emergency escape route and increase its speed via the flight control processor. In practice, when the security situation assessment indicates a "Level 3 risk" (high threat, such as being locked on or confirmed to be under countermeasures), the implementing entity performs the highest level of survivability operations. First, electronic deception is implemented: the data link transmitter is instructed to immediately cease normal communication and instead emit a pre-loaded deceptive signal (e.g., simulating the communication characteristics of another UAV or emitting cluttered noise) to confuse the enemy's electronic warfare system and buy time to escape. Simultaneously, an emergency escape is executed: flight control is completely handed over to the flight control processor, which autonomously flies along a pre-set emergency escape route (usually a path moving away from the current threat area at maximum speed), and instructs the propulsion system to increase to the maximum usable flight speed to evacuate the danger zone as quickly as possible. During this process, the adversarial sample generation task is permanently terminated until the task is reset.

[0083] Steps one through five of the embodiments of this disclosure address the technical problem that "in complex, highly adversarial battlefield environments, the electromagnetic, communication, and behavioral characteristics generated by UAV-based remote sensing cross-modal adversarial attack methods during execution easily expose attack intentions and platform locations. Furthermore, when facing enemy electronic interference, network countermeasures, or physical threats, they lack real-time, autonomous threat perception and layered survival response capabilities, leading to attack mission interruption and a significant increase in platform survival risks." Existing technologies have the following shortcomings in the above aspects: Firstly, attack systems typically focus only on the generation and optimization of attack algorithms, lacking the ability to perceive and control the signal characteristics radiated during mission execution, becoming a "blind spot" exposure source. Secondly, the system lacks a fusion perception and unified assessment mechanism for external dynamic threats (such as enemy reconnaissance and interference) and internal anomalies (such as software attacks on the platform), resulting in delayed responses or one-sided decision-making. Thirdly, when encountering threats, there is a lack of automatically executable hardware and software coordinated control strategies that precisely match the threat level; often, it can only rely on manual operator intervention or execute a single "all or nothing" response, making it difficult to achieve a dynamic balance between protecting the mission and protecting the platform. Solving the above problems can significantly improve the stealth of the attack system in adversarial environments, its resilience to countermeasures, and the platform's autonomous survivability under threats, thereby ensuring the continuity and effectiveness of core attack missions. To achieve this effect, this disclosure proposes the following: First, in response to the generation of adversarial disturbance data or adversarial disturbance data sequences, the system continuously scans the operating frequency band using airborne electronic support measurement equipment and conducts a risk assessment of the scanned operating frequency band to obtain a risk level. Thus, by simultaneously initiating active and continuous reconnaissance of the operating electromagnetic frequency band when the attack mission is launched, the system can detect in real time whether there are abnormal radar scans, communication interference, or other active electronic threat signals from the enemy in the environment. This solves the "blind spot" problem of the attack system's perception of external electromagnetic threats, transforming the traditional "silent execution" into "simultaneous perception and execution," achieving early warning of exposed risks from the source, and providing crucial external environmental situational input for subsequent decision-making. Second, in response to the evaluation of the attack effectiveness of adversarial samples, the system continuously identifies abnormal response characteristics of the feedback traffic to obtain countermeasure feature identification data. Therefore, by deeply analyzing the feedback of the target recognition model (or system) to adversarial examples (such as prediction results, confidence levels, and response delays), it is possible to intelligently identify whether the opponent has activated an active defense algorithm based on anomaly detection. This solves the problem that attack systems cannot effectively detect "network-level countermeasures," upgrading the attack process from one-way sample delivery to two-way interactive detection. It can promptly detect whether the attack behavior has been perceived by the opponent and triggered countermeasures, providing direct evidence for judging the effectiveness and security of the attack strategy. The third step involves continuously acquiring platform anomaly status data during the execution of the first or second perturbation processing.Therefore, by monitoring hardware errors, abnormal process behavior, and performance indicators within the computing platform, it is possible to effectively detect whether the platform itself has been attacked by underlying software, experienced hardware failures, or malfunctions. This solves the problem of attack systems neglecting the security of the "internal fortress," establishes a continuous audit capability for the integrity and security of the computing platform itself, prevents attack algorithms from failing, models from being leaked, or erroneous outputs due to platform penetration or failure, and ensures the reliability of the attack execution carrier. The fourth step generates a comprehensive security posture indicator based on the obtained risk level, countermeasure feature identification data, or platform abnormal state data. Thus, by fusing and comprehensively evaluating heterogeneous and asynchronous threat information from three dimensions—external electromagnetic environment, network interaction feedback, and internal platform status—a unified, hierarchical comprehensive security posture indicator (such as Level 1, Level 2, and Level 3 risk) is output based on a pre-set threat model. This solves the problems of isolated multi-source threat information and complex and chaotic decision-making logic, constructing a centralized and intelligent "tactical brain," condensing complex battlefield situations into clear decision-making instructions, and providing a unique and authoritative input basis for executing precise and automated hierarchical responses. The fifth step, based on the generated security posture indicator, executes the following hierarchical hardware control operations. Thus, by establishing a pre-set set of hardware control instructions strictly bound to different risk levels, the system can automatically execute a progressive response from "flexible degradation" to "rigid defense" and then to "emergency disengagement" without relying on human intervention. This solves the problems of single, slow, or excessive response strategies when encountering threats. Specifically: In response to Level 1 risk, operations such as reducing attack intensity and switching to covert communication modes are executed, minimizing the platform's own signal characteristics while maintaining the continuity of the attack mission, achieving a fine balance between attack and covertness. In response to Level 2 risk, operations such as pausing the attack, clearing the scene, and performing evasive maneuvers are executed, achieving a rapid "meltdown" of the attack mission and proactive "tactical evasion" of the platform, sacrificing short-term mission progress in exchange for the opportunity to escape the threat and preserve combat power for future battles. In response to Level 3 risk, operations such as transmitting deceptive signals and disengaging at full speed along an emergency route are executed, activating the highest priority platform survival protocol, buying time through electronic deception, and using maximum maneuverability for physical disengagement, ensuring the survival of the platform itself in the worst-case scenario. In summary, the first to fifth steps of this embodiment cooperate with each other, from comprehensive perception including external electromagnetic reconnaissance, network feedback detection, and internal status monitoring, to multi-source information fusion assessment, and then to hierarchical hardware control precisely bound to threat levels, together forming a complete autonomous survival closed loop of "perception-assessment-response". Thus, it not only upgrades a simple attack algorithm execution unit into an organism with environmental awareness and survival intelligence, but also dynamically and elegantly resolves the core contradiction between "executing attacks" and "preserving oneself" through deep software and hardware synergy.This fundamentally enhances the resilience, reliability, and battlefield survivability of drones when carrying out high-risk combat missions, enabling advanced cross-modal combat attack technologies to be truly applied to severe real-world combat scenarios, achieving a key leap from "laboratory algorithms" to "practical tools."

[0084] The above embodiments of the present invention have the following beneficial effects: The remote sensing cross-modal adversarial sample generation method for UAVs of the present invention achieves adaptive, closed-loop, and hardware-software collaborative capabilities for generating highly mobile and highly concealed adversarial samples in dynamic, resource-constrained UAV mission environments, improving the practical applicability, efficiency, and mission survivability of adversarial attacks in real and complex remote sensing scenarios. Specifically, traditional adversarial sample generation techniques (such as optimization algorithms relying on static parameters or cross-modal attack schemes detached from hardware) may experience a sharp drop in attack efficiency when encountering computing power bottlenecks on the UAV platform, the generated video adversarial perturbation may have temporal breaks that are easily filtered out, and the attack behavior may be "blindly" due to complete disconnection from the flight mission; if only a fixed strategy is used, it will be unable to adapt to fluctuations in hardware status during flight, interference from the electromagnetic environment, and changes in maneuvering attitude, ultimately leading to attack failure or self-exposure. Based on this, the remote sensing cross-modal adversarial sample generation method for UAVs of the present invention: First, acquire remote sensing data and real-time hardware status data. This implants environmental awareness "nerve endings" into the entire method, enabling not only the acquisition of target data but, more importantly, real-time monitoring of the UAV computing platform's "physical state" (such as GPU memory, CPU load, and bandwidth). This provides a dynamic physical basis for all subsequent decisions, ensuring tight coupling between the method and the hardware platform from the outset. Next, the acquired remote sensing data undergoes type identification processing to obtain data type determination results. This enables automated data splitting (images or videos), providing precise triggering conditions for launching two different, targeted optimized attack flows (SFCM or TSFCM-I2V), avoiding performance loss from a single processing mode for another type of data. Then, based on the data type determination results and the acquired real-time hardware status data, a dynamic optimization strategy is generated. This reveals the core innovation: dynamically binding algorithm parameters to hardware status. For example, increasing the batch size and model count when GPU memory is ample enhances attack power, while adjusting the optimization step size during high CPU load avoids system lag. This achieves real-time optimal matching between algorithm performance and hardware resources, solving the core problem of static parameters being unable to adapt to dynamic environments. Subsequently, in response to the aforementioned type determination result being represented as an image, the acquired remote sensing data undergoes a first perturbation process based on the aforementioned dynamic optimization strategy. This involves executing a frequency-space cooperative modulation attack (SFCM) specifically designed for remote sensing images. This attack decouples the target background through wavelet transform, employs block-based differential perturbation based on attention heatmaps, and combines gradient optimization adjusted by a dynamic strategy to generate highly transferable image adversarial perturbations. Simultaneously, in response to the aforementioned type determination result being represented as video, the acquired remote sensing data undergoes a second perturbation process based on the aforementioned dynamic optimization strategy.Therefore, a Spatiotemporal Frequency Domain Co-modulation Attack (TSFCM-I2V) specifically designed for remote sensing video was executed. Building upon image attacks, keyframe priority processing and inter-frame low-frequency consistency constraints were introduced to ensure that the generated video adversarial perturbations were both effective and temporally smooth and continuous, overcoming the challenge of temporal inconsistency in cross-modal attacks. Subsequently, the obtained adversarial perturbation data or sequences were superimposed on the original data to obtain adversarial samples. This completed the physical generation of adversarial samples, providing entity objects for subsequent performance evaluation and control execution. The attack performance of the adversarial samples was then evaluated, yielding the evaluation results. This established a real-time feedback loop of effectiveness, objectively measuring the effectiveness and stealth of the attack through quantitative indicators (such as attack success rate and structural similarity), providing a data-driven basis for strategy iteration. Furthermore, based on the above performance evaluation results and real-time hardware status data, the above dynamic optimization strategy was iteratively optimized to obtain a set of hardware control parameters. This enabled the online self-evolution of the strategy. Based on the effectiveness of the attack and the current hardware load, the optimization parameters for the next round (such as perturbation amplitude and number of iterations) are dynamically adjusted, and the optimization decisions are ultimately translated into executable hardware control instructions (such as adjusting computing resource allocation), forming a complete closed loop from "perception-decision-execution-evaluation" to "optimization". Finally, based on the above set of hardware control parameters, adjustments are made to each processor and controlled hardware in the airborne computing platform. This achieves the ultimate leap from digital strategy to physical control. The method not only generates "soft" adversarial examples but also directly drives "hard" components such as flight control processors, mission processors, data links, and optoelectronic payloads, enabling the UAV's computing resource allocation, communication waveforms, and even platform attitude to be collaboratively optimized for the current adversarial mission, fundamentally solving the problem of the disconnect between the algorithm and the physical system. Furthermore, because this method introduces dynamic adaptation and closed-loop feedback mechanisms throughout the entire chain of input perception, strategy generation, perturbation processing, effect evaluation, and hardware control, it can effectively address the challenges posed by UAV platform resource fluctuations, complex electromagnetic interference, and variable mission scenarios. By deeply integrating adversarial attack algorithms, real-time hardware scheduling, and flight mission context, this method transforms UAVs from passive "tools" executing pre-set attack programs into intelligent agents capable of proactively adjusting attack strategies and coordinating platform behavior based on their own state and environmental changes. Thus, by achieving the integration of algorithmic adaptation, cross-modal consistency, and software-hardware collaborative control, it enhances the reliability of UAV attack effectiveness, platform resource utilization efficiency, and overall mission survivability in realistic combat environments, providing a key technical paradigm for proactive security testing and adversarial defense research of UAV remote sensing systems.

[0085] Further reference Figure 2 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of a remote sensing cross-modal adversarial sample generation device for UAVs. These device embodiments are similar to... Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0086] like Figure 2 As shown, a remote sensing cross-modal adversarial sample generation device 200 for UAVs according to some embodiments includes: an acquisition unit 201, an identification unit 202, a generation unit 203, a first perturbation unit 204, a second perturbation unit 205, an overlay unit 206, an evaluation unit 207, an optimization unit 208, and an adjustment unit 209. The acquisition unit 201 is configured to acquire remote sensing data and real-time hardware status data; the identification unit 202 is configured to perform type identification processing on the acquired remote sensing data to obtain a data type determination result; the generation unit 203 is configured to generate a dynamic optimization strategy based on the data type determination result and the acquired real-time hardware status data; the first perturbation unit 204 is configured to, in response to the type determination result being represented as an image, perform a first perturbation processing on the acquired remote sensing data based on the dynamic optimization strategy to obtain adversarial perturbation data; the second perturbation unit 205 is configured to, in response to the type determination result being represented as video, perform a first perturbation processing on the acquired remote sensing data based on the dynamic optimization strategy to obtain adversarial perturbation data. The acquired remote sensing data undergoes a second perturbation process to obtain an adversarial perturbation data sequence; the overlay unit 206 is configured to overlay the obtained adversarial perturbation data or adversarial perturbation data sequence with the acquired remote sensing data to obtain adversarial samples; the evaluation unit 207 is configured to evaluate the attack effectiveness of the adversarial samples to obtain an evaluation result; the optimization unit 208 is configured to iteratively optimize the dynamic optimization strategy based on the performance evaluation result and the acquired real-time hardware status data to obtain a set of hardware control parameters; and the adjustment unit 209 is configured to adjust each processor and each controlled hardware in the airborne computing platform according to the set of hardware control parameters.

[0087] It is understandable that the units described in the device 200 are related to the reference. Figure 1 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 200 and the units contained therein, and will not be repeated here.

[0088] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0089] like Figure 3As shown, the electronic device 300 may include a processing unit 301 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0090] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.

[0091] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.

[0092] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0093] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0094] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs. When the electronic device executes the aforementioned one or more programs, the electronic device causes the following actions: acquiring remote sensing data and real-time hardware status data; performing type identification processing on the acquired remote sensing data to obtain a data type determination result; generating a dynamic optimization strategy based on the aforementioned data type determination result and the acquired real-time hardware status data; in response to the aforementioned type determination result being represented as an image, performing a first perturbation processing on the acquired remote sensing data based on the aforementioned dynamic optimization strategy to obtain adversarial perturbation data; in response to the aforementioned type determination result being represented as video, performing a second perturbation processing on the acquired remote sensing data based on the aforementioned dynamic optimization strategy to obtain an adversarial perturbation data sequence; overlaying the obtained adversarial perturbation data or adversarial perturbation data sequence with the acquired remote sensing data to obtain an adversarial sample; evaluating the attack effectiveness of the aforementioned adversarial sample to obtain an evaluation result; iteratively optimizing the aforementioned dynamic optimization strategy based on the aforementioned performance evaluation result and the acquired real-time hardware status data to obtain a set of hardware control parameters; and adjusting each processor and each controlled hardware in the airborne computing platform according to the aforementioned set of hardware control parameters.

[0095] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0096] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0097] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including an acquisition unit, an identification unit, a generation unit, a first perturbation unit, a second perturbation unit, a superposition unit, an evaluation unit, an optimization unit, and an adjustment unit. The names of these units do not necessarily limit the specific unit; for example, the acquisition unit may also be described as "a unit that acquires remote sensing data and real-time hardware status data."

[0098] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0099] Some embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements any of the above-described methods for generating cross-modal adversarial examples for UAV remote sensing.

[0100] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A method for generating cross-modal adversarial examples for remote sensing using unmanned aerial vehicles (UAVs), comprising: Acquire remote sensing data and real-time hardware status data; The acquired remote sensing data is processed for type identification to obtain the data type determination result; Based on the data type determination result and the acquired real-time hardware status data, a dynamic optimization strategy is generated. In response to the type determination result being represented as an image, the acquired remote sensing data is subjected to a first perturbation process based on the dynamic optimization strategy to obtain adversarial perturbation data. In response to the type determination result being characterized as video, the acquired remote sensing data is subjected to a second perturbation process based on the dynamic optimization strategy to obtain an adversarial perturbation data sequence. The obtained adversarial perturbation data or adversarial perturbation data sequence is overlaid with the acquired remote sensing data to obtain adversarial samples; The attack effectiveness of the adversarial sample was evaluated, and the evaluation results were obtained. Based on the performance evaluation results and the acquired real-time hardware status data, the dynamic optimization strategy is iteratively optimized to obtain a set of hardware control parameters. Based on the set of hardware control parameters, adjustments are made to each processor and each controlled hardware component in the airborne computing platform.

2. The method according to claim 1, wherein, The process of performing type identification processing on the acquired remote sensing data to obtain data type determination results includes: The file header information or data stream structure of the acquired remote sensing data is parsed to obtain the format identifier; In response to the extracted format identifier conforming to the preset static image encoding format, a type determination result representing the image is generated; In response to the extracted format identifier conforming to the preset dynamic video encoding format, a type determination result representing a video is generated; In response to the fact that the extracted format identifier does not conform to the preset static image encoding format or the preset dynamic video encoding format, the acquired remote sensing data is subjected to content-assisted determination to obtain the data type determination result.

3. The method according to claim 1, wherein, The step of generating a dynamic optimization strategy based on the data type determination result and the acquired real-time hardware status data includes: Based on the data type determination result, generate the number of iterations and the upper limit of the perturbation amplitude corresponding to the gradient optimization; Based on the available video memory size in the acquired real-time hardware status data, generate the number of agent models to participate in the processing and the batch size; Based on the processor utilization in the acquired real-time hardware status data, the dynamic step size decay coefficient and the initial value of the momentum factor corresponding to the gradient optimization process are generated. The iteration number, the upper limit of the perturbation amplitude, the number of surrogate models, the batch size, the dynamic step size decay coefficient, and the initial value of the momentum factor are integrated into a dynamic optimization strategy.

4. The method according to claim 3, wherein, The response to the type determination result is represented as an image. Based on the dynamic optimization strategy, the acquired remote sensing data undergoes a first perturbation process to obtain adversarial perturbation data, including: Based on the number of proxy models in the dynamic optimization strategy, load the corresponding number of pre-trained image recognition models to obtain the target proxy model set. The remote sensing data is subjected to discrete wavelet transform to obtain low-frequency and high-frequency components; The low-frequency components are randomly linearly scaled while the high-frequency components remain unchanged to obtain the transformed frequency domain components. Perform inverse wavelet transform on the transformed frequency domain components to obtain the frequency domain transformed image; The frequency domain transform image is subjected to activation mapping processing to obtain a semantic attention heatmap; Based on the semantic attention heatmap, the remote sensing data is divided into target-related region blocks and background region blocks; Geometric transformations are performed on each target-related region block, and environmental transformations are performed on each background region block to obtain transformed region blocks. The transformed regions are stitched together to obtain a spatially transformed image; Based on the dynamic optimization strategy and the target proxy model set, the optimization objective is to minimize the preset loss function. By dynamically adjusting the step size and momentum factor, gradient optimization is performed on the spatially transformed image to obtain data that resists disturbances.

5. The method according to claim 4, wherein, The response to the type determination result is represented as video. Based on the dynamic optimization strategy, the acquired remote sensing data undergoes a second perturbation process to obtain an adversarial perturbation data sequence, including: Keyframes are extracted from the remote sensing data to obtain a set of keyframes and an index of non-keyframes. For each key frame in the set of key frames, the first perturbation processing is performed to obtain key frame adversarial perturbation data corresponding to each key frame. When performing the discrete wavelet transform and the random linear scaling on any key frame, the mean square error between the low-frequency components of the current key frame and the previous key frame is generated, and the mean square error is added as a loss term to the optimization objective. For each non-critical frame included in the non-critical frame index, based on the timestamp corresponding to the non-critical frame, linear interpolation is performed on the two key frame anti-disturbance data corresponding to two adjacent key frames to obtain the non-critical frame anti-disturbance data of the corresponding non-critical frame. Based on the time sequence of the remote sensing data, the obtained key frame adversarial perturbation data and non-key frame adversarial perturbation data are integrated to obtain an adversarial perturbation data sequence.

6. The method according to claim 3, wherein, Based on the performance evaluation results and the acquired real-time hardware status data, the dynamic optimization strategy is iteratively optimized to obtain a set of hardware control parameters, including: If the attack success rate in the performance evaluation result is lower than a preset success rate threshold, the upper limit of the perturbation amplitude and the number of iterations in the dynamic optimization strategy are increased. In response to the structural similarity index in the performance evaluation result being lower than a preset similarity index threshold, the upper limit of the perturbation amplitude in the dynamic optimization strategy is reduced; Based on the current power consumption and temperature in the real-time hardware status data, adjust the dynamic step size decay coefficient and the initial value of the momentum factor in the dynamic optimization strategy. Based on the video memory and bandwidth utilization rates in the real-time hardware status data, the number of proxy models and batch size in the dynamic optimization strategy are dynamically adjusted. Based on the adjusted dynamic optimization strategy, a set of hardware control instructions is generated. The set of hardware control instructions is converted into a set of hardware control parameters.

7. A remote sensing cross-modal adversarial example generation device for unmanned aerial vehicles (UAVs), comprising: The acquisition unit is configured to acquire remote sensing data and real-time hardware status data; The identification unit is configured to perform type identification processing on the acquired remote sensing data to obtain the data type determination result; The generation unit is configured to generate a dynamic optimization strategy based on the data type determination result and the acquired real-time hardware status data. The first perturbation unit is configured to respond to the type determination result represented as an image, and based on the dynamic optimization strategy, perform a first perturbation process on the acquired remote sensing data to obtain adversarial perturbation data. The second perturbation unit is configured to, in response to the type determination result being represented as video, perform a second perturbation process on the acquired remote sensing data based on the dynamic optimization strategy to obtain an adversarial perturbation data sequence. The overlay unit is configured to overlay the obtained adversarial perturbation data or adversarial perturbation data sequence with the acquired remote sensing data to obtain adversarial samples; An evaluation unit is configured to evaluate the attack effectiveness of the adversarial sample and obtain an evaluation result. The optimization unit is configured to iteratively optimize the dynamic optimization strategy based on the performance evaluation results and the acquired real-time hardware status data to obtain a set of hardware control parameters. The adjustment unit is configured to adjust each processor and each controlled hardware in the airborne computing platform according to the set of hardware control parameters.

8. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 6.

9. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 6.