Dynamic prediction method and system for spaceborne resource constraints and scene complexity
Patent Information
- Application Number
- CN202611116970.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-27
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-07-27
AI Technical Summary
[0005]有鉴于此,本申请实施例提供一种面向星载资源约束与场景复杂度的动态预测方法及系统,以解决星载平台无法适配在轨动态工况的问题
[0008]借由上述技术方案,本申请实施例提供一种面向星载资源约束与场景复杂度的动态预测方法及系统,所述方法通过采集星载平台的多维资源参数和遥感监测数据,基于多维资源参数预测超前状态参数,并根据从遥感监测图像中提取的场景关联特征计算场景复杂度评分。再基于超前状态参数和场景复杂度评分确定动态推理参数,然后根据动态推理参数重构视觉模型,从而使用重构后的视觉模型对遥感监测图像执行模型推理,得到模型预测结果。所述方法通过场景和资源双驱动的弹性适配机制,提升预测能效,可节约星载能源与算力资源。通过超前资源预判与分级安全控制机制,使得硬件运行安全可控,有效规避超温、算力过载、任务中断等硬件风险,在保证模型预测精度的同时,提高模型预测的稳定性。
Smart Images

Figure CN122637112B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of spaceborne artificial intelligence technology, and in particular to a dynamic prediction method and system for spaceborne resource constraints and scenario complexity. Background Technology
[0002] The spaceborne platform, based on on-orbit intelligent processing of large visual models such as Transformer, can be applied to remote sensing detection tasks such as target detection, scene interpretation, and disaster monitoring. However, the spaceborne platform exhibits significant resource constraints, specifically limited computing power, strict power consumption quotas, stringent chip thermal control conditions, and limited storage space. Furthermore, its on-orbit resource status fluctuates dynamically with satellite orbital position, operating time, and task load. In addition, remote sensing images encompass diverse scenes, including oceans, deserts, farmland, suburbs, urban areas, and disaster zones. The significant differences in texture features, target distribution, and information density across different scenes result in a markedly dynamic complexity in remote sensing scenes.
[0003] When the spaceborne platform makes predictions based on a large model, it can adopt a deployment method of offline compression, fixed quantization parameters, and fixed prediction structure on the ground. The model compression, quantization configuration, and prediction strategy are completed offline on the ground. Then, the spaceborne platform is launched into the space environment. In the space environment, the deployed large model and quantization configuration results are used to process the remote sensing monitoring data online according to the deployed prediction strategy to obtain prediction results. These results are then used to adjust the remote sensing monitoring method in real time.
[0004] However, because the prediction structure of the spaceborne platform is static and fixed, it cannot adapt to the dynamic fluctuations of spaceborne resources. Specifically, it cannot dynamically adjust the prediction mode based on changes in real-time computing load, remaining power consumption, chip temperature, and memory usage. This easily leads to problems such as resource overload, chip overheating, excessive prediction latency, and mission interruption, resulting in poor on-orbit stability. Furthermore, the spaceborne platform uses a uniform prediction scale for all remote sensing images. For simple scenarios, this results in computing redundancy and wasted power, while for complex scenarios, it leads to insufficient computing power and decreased prediction accuracy, resulting in unreasonable allocation of prediction resources. Summary of the Invention
[0005] In view of this, embodiments of this application provide a dynamic prediction method and system for spaceborne resource constraints and scenario complexity, in order to solve the problem that spaceborne platforms cannot adapt to dynamic on-orbit operating conditions.
[0006] According to a first aspect of this application, a dynamic prediction method for spaceborne resource constraints and scenario complexity is provided, the method comprising: Acquire multidimensional resource parameters and remote sensing monitoring data of the spaceborne platform; the remote sensing monitoring data includes any type of remote sensing monitoring image from optical remote sensing data, radar remote sensing data, and spectral remote sensing data. The state prediction model is used to predict leading state parameters based on the multidimensional resource parameters; the state prediction model includes at least one of a linear regression model and a time series prediction model. Scene association features are extracted from the remote sensing monitoring images, and a scene complexity score is calculated based on the scene association features. The scene association features include edge density, texture entropy, target density, and frequency domain complexity. The scene complexity score is a weighted calculation result of the scene association features based on adaptive weights. The adaptive weights are generated online by fine-tuning the spatial dispersion of various scene association features in a single remote sensing monitoring image, with a preset basic weight corresponding to the remote sensing monitoring data type as the benchmark. Dynamic inference parameters are determined based on the advanced state parameters and the scene complexity score. The dynamic inference parameters are dynamically adapted parameters obtained by quantifying and fine-tuning the baseline inference parameters under the target inference mode according to independent correction rules. The independent correction rules include at least one of remote sensing data type, task priority, and resource sufficiency. The target inference mode is obtained by interval matching based on the advanced state and the scene complexity score. The model parameters of the visual model are reconstructed based on the dynamic inference parameters, and the model parameters of the visual model include model depth, model width, and model accuracy. The reconstructed visual model is used to perform model inference on the remote sensing images to generate model prediction results.
[0007] According to a second aspect of this application, a dynamic prediction system for spaceborne resource constraints and scenario complexity is provided, the system comprising: The data acquisition module is used to acquire multi-dimensional resource parameters and remote sensing monitoring data of the spaceborne platform; the remote sensing monitoring data includes any kind of remote sensing monitoring image from optical remote sensing data, radar remote sensing data and spectral remote sensing data. A resource state prediction module is used to predict leading state parameters based on the multidimensional resource parameters using a state prediction model; the state prediction model includes at least one of a linear regression model and a time series prediction model. The scene detection module is used to extract scene association features from the remote sensing monitoring image and calculate a scene complexity score based on the scene association features. The scene association features include at least one of edge density, texture entropy, target density, and frequency domain complexity. The scene complexity score is a weighted calculation result of the scene association features based on adaptive weights. The adaptive weights are generated online by fine-tuning the spatial dispersion of various scene association features in a single remote sensing monitoring image, based on a preset basic weight corresponding to the remote sensing monitoring data type. The inference parameter determination module is used to determine dynamic inference parameters based on the advanced state parameters and the scene complexity score. The dynamic inference parameters are dynamically adapted parameters obtained by quantifying and fine-tuning the baseline inference parameters under the target inference mode according to independent correction rules. The independent correction rules include at least one of remote sensing data type, task priority, and resource abundance. The target inference mode is obtained by interval matching based on the advanced state and the scene complexity score. The model reconstruction module is used to reconstruct the model parameters of the visual model based on the dynamic inference parameters. The model parameters of the visual model include model depth, model width, and model accuracy. The model inference module is used to perform model inference on the remote sensing monitoring image using the reconstructed visual model to generate model prediction results.
[0008] By employing the above technical solutions, this application provides a dynamic prediction method and system for spaceborne resource constraints and scene complexity. The method collects multi-dimensional resource parameters and remote sensing monitoring data from the spaceborne platform, predicts advanced state parameters based on these parameters, and calculates a scene complexity score based on scene association features extracted from the remote sensing images. Then, dynamic inference parameters are determined based on the advanced state parameters and the scene complexity score. A visual model is then reconstructed based on these parameters, and the reconstructed visual model is used to perform model inference on the remote sensing images to obtain the model prediction results. This method improves prediction efficiency through a flexible adaptation mechanism driven by both scene and resources, saving spaceborne energy and computing resources. Through advanced resource prediction and a hierarchical safety control mechanism, hardware operation is safe and controllable, effectively avoiding hardware risks such as overheating, computing overload, and task interruption, thus improving the stability of model prediction while ensuring its accuracy.
[0009] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0010] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram of the dynamic prediction method provided in the embodiments of this application; Figure 2 This is a flowchart illustrating the dynamic prediction method provided in the embodiments of this application. Figure 3This is a schematic diagram of the spaceborne resource acquisition and advanced prediction process provided in the embodiments of this application; Figure 4 This is a schematic diagram of the multi-source remote sensing scene complexity assessment process provided in the embodiments of this application; Figure 5 This is a schematic diagram of the adaptive decision-making process for the flexible reasoning strategy provided in the embodiments of this application; Figure 6 This is a schematic diagram of the large model security dynamic reconstruction process provided in the embodiments of this application; Figure 7 This is a schematic diagram of the confidence closed-loop verification process provided in the embodiments of this application; Figure 8 This is a schematic diagram of the dynamic prediction system structure provided in an embodiment of this application. Detailed Implementation
[0011] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.
[0012] To address the issue of spaceborne platforms being unable to adapt to dynamic on-orbit operating conditions, some embodiments of this application provide a dynamic prediction method oriented towards spaceborne resource constraints and scenario complexity. This method improves prediction efficiency through a flexible adaptation mechanism driven by both scenario and resources, saving spaceborne energy and computing resources. By employing advanced resource prediction and a hierarchical safety control mechanism, hardware operation is made safe and controllable, effectively avoiding hardware risks such as overheating, computing overload, and task interruption, thus improving the stability of model predictions while ensuring their accuracy.
[0013] The method can be applied to a spaceborne platform or an electronic device that establishes a communication connection with the spaceborne platform and has data processing capabilities. The electronic device includes, but is not limited to, computers, servers, mobile terminals, smart wearable devices, and industrial control machines. For ease of description, the spaceborne platform is used as the execution subject in this embodiment. It should be understood that the method can also be applied to other types of execution subjects, which will not be shown in this embodiment.
[0014] To meet the needs of mainstream spaceborne intelligent processing hardware environments with limited hardware resources, strict power consumption constraints, and demanding thermal control conditions, this application is perfectly suited to its application scenarios. It is comprehensively compatible with intelligent interpretation tasks such as target detection, scene interpretation, disaster monitoring, ground feature survey, and on-orbit inspection of spaceborne optics, synthetic aperture radar (SAR), and hyperspectral multi-type remote sensing data. In some exemplary embodiments of this application, the dynamic prediction method for spaceborne resource constraints and scene complexity can be executed through the following hardware and software configurations and on-orbit constraints.
[0015] For the core hardware platform configuration, the spaceborne platform can adopt an aerospace-grade embedded neural network processor (NPU) chip with a main frequency of 800MHz to 1.2GHz, supporting multi-precision mixed quantization acceleration inference such as INT4, INT8, and FP16. The hardware INT8 peak AI computing power is ≥24TOPS, meeting the real-time inference requirements of lightweight large models on spaceborne platforms.
[0016] Equipped with 2GB of on-orbit solidified video memory and 8GB of onboard storage, it supports large-scale remote sensing image caching, persistent model parameter storage, and temporary storage of anomalous samples. The device's rated power consumption is ≤8W, and its short-term peak power consumption is ≤11W (single peak duration ≤30s), adapting to the strict power quota constraints of satellite platforms. The chip's standard operating temperature range is -40℃ to +80℃, supporting the complex thermal control environment of space, and possessing space irradiation hardening and single-event upset resistance capabilities, meeting the requirements for long-term stable on-orbit operation.
[0017] For basic operating parameter configuration, the onboard platform resource acquisition cycle is fixed at 100ms, the timing sliding window size is 10 consecutive samples, the single sampling interval is 100ms, the total sampling duration is 1s, and the strategy hysteresis stabilization control number is fixed at 3 times.
[0018] The base model adopts the ViT-Base standard visual Transformer backbone, which contains 12 cascaded Transformer coding layers (numbered from coding layer 1 to coding layer 12 in forward inference order). The input end is configured with patch embedding and positional encoding modules, and the output end is configured with the final normalization layer. No custom structural modifications were made. The model's native standard input resolution is 1024×1024, which, after 16×16 block encoding, generates a baseline token sequence of length 4096, providing a standard structural baseline for subsequent 3D elastic reconstruction.
[0019] For on-orbit operation constraints, the satellite platform can be set to require an end-to-end inference latency of ≤500ms for a single remote sensing image. It supports real-time processing of three mainstream types of satellite remote sensing data: optical, SAR, and hyperspectral. It also supports concurrent scheduling of multiple tasks such as disaster monitoring, ground feature surveys, and target inspection, meeting the requirements for long-term stable on-orbit operation of the satellite.
[0020] like Figure 1 As shown, the dynamic prediction method for spaceborne resource constraints and scenario complexity includes: S101. Acquire multi-dimensional resource parameters and remote sensing monitoring data of the spaceborne platform.
[0021] like Figure 2As shown, during operation, the spaceborne platform can perform remote sensing data acquisition according to specific remote sensing monitoring tasks. Furthermore, in order to execute dynamic remote sensing model predictions, it can also perform real-time acquisition of spaceborne resources, that is, obtain multi-dimensional resource parameters and remote sensing monitoring data from the spaceborne platform.
[0022] The remote sensing monitoring data includes images from any of the following sources: optical remote sensing data, radar remote sensing data, and spectral remote sensing data. For example, the remote sensing monitoring data may include images from multiple sources, such as optical remote sensing equipment, SAR remote sensing equipment, and hyperspectral remote sensing equipment.
[0023] The multi-dimensional resource parameters include at least one of the following collected within a preset time period: onboard computing power utilization, remaining power budget, chip real-time temperature, graphics memory occupancy, and satellite-to-ground communication bandwidth occupancy. As the first preliminary step in the dynamic prediction process, the onboard platform can collect multi-dimensional resource parameters over a short period and use them for time-series modeling, dual-model adaptive switching, and proactive safety prediction—a resource awareness and prediction mechanism. This allows for real-time capture of dynamic changes in on-orbit resources, prediction of short-term fluctuation trends, and output of resource status prediction values 100ms ahead of time. This provides accurate resource-dimensional constraints for subsequent adaptive adjustments to inference strategies, achieving proactive resource adaptation and risk prevention from the outset. For example, the multi-dimensional resource parameters can be set to a fixed 100ms collection period to collect multiple core operating parameters of the onboard computing unit in real time.
[0024] S102. Use the state prediction model to predict advanced state parameters based on multi-dimensional resource parameters.
[0025] After acquiring multidimensional resource parameters and remote sensing monitoring data from the spaceborne platform, the platform can perform advanced state prediction. This involves using a state prediction model to predict advanced state parameters based on the multidimensional resource parameters. The advanced prediction step size is one acquisition cycle, or 100ms, meaning that based on current and historical resource data within the last second, the resource state at the next acquisition time is predicted. The state prediction model includes at least one of linear regression models and time-series prediction models.
[0026] like Figure 3 As shown, in some embodiments, the leading state parameters include either a resource prediction value for the next acquisition step (100ms) or a resource prediction vector. Therefore, when using a state prediction model to predict leading state parameters based on multidimensional resource parameters, a resource fluctuation evaluation index can be calculated first based on the multidimensional resource parameters. This resource fluctuation evaluation index includes the resource fluctuation amplitude from multiple consecutive samples.
[0027] Then, a multi-dimensional resource state vector is constructed based on multi-dimensional resource parameters. For example, based on multi-dimensional resource parameters including onboard computing power utilization, remaining power budget, chip real-time temperature, graphics memory utilization, and satellite-to-ground communication bandwidth utilization, the following multi-dimensional resource state vector can be constructed:
[0028] in, R inst A multidimensional resource state vector; u To improve the utilization rate of onboard NPU computing power, p For the remaining power budget, T For real-time junction temperature of the chip, m This refers to the video memory usage rate. b This refers to the bandwidth utilization rate of satellite-to-ground communication.
[0029] After constructing a multidimensional resource state vector, the resource fluctuation amplitude can be compared with a preset fluctuation threshold. When the resource fluctuation amplitude is less than or equal to the preset fluctuation threshold, a linear regression model is used to fit an independent linear time series equation for each type of resource, and the future time prediction values of each type of resource are calculated. Finally, the vectors are spliced together to form an advanced state prediction vector that is consistent with the dimensions of the multidimensional resource parameters.
[0030] When resource fluctuations exceed a preset fluctuation threshold, a time-series prediction model can be used to memorize the near-term resource mutation and fluctuation characteristics of multidimensional resource parameters, and output future resource prediction vectors based on these characteristics. The resource mutation and fluctuation characteristics are used to fit the nonlinear variation patterns of the multidimensional resource parameters.
[0031] For example, a short-cycle sliding window time-series acquisition and prediction mechanism is adopted, which is based entirely on real-time short-term data in orbit and does not rely on long-term offline historical data, thus adapting to the real-time processing requirements of the satellite. To this end, the satellite platform samples multi-dimensional resource data from the satellite sequentially with a fixed acquisition cycle of 100ms to construct continuous time-series samples. The preset time-series sliding window size is 10 resource acquisition windows, corresponding to a total time span of 1 second (10 × 100ms). Only the most recent continuous real-time resource time-series data within the last second is retained, and expired historical sampling data is automatically discarded. This ensures that the prediction data source always closely matches the current real-time operating conditions of the satellite in orbit, mitigating prediction biases caused by outdated data.
[0032] Based on this 10-dimensional short-time series sample set, two lightweight state prediction models can be adaptively switched to predict leading state parameters to meet the requirements of low computing power overhead onboard systems. Both models ultimately output a unified 5-dimensional resource prediction vector, corresponding to five indicators: computing power utilization, remaining power budget, chip junction temperature, memory occupancy, and satellite-to-ground communication bandwidth, fully compatible with the input format of downstream decision-making modules. One of the state prediction models is a lightweight linear regression model. The input to the linear regression model is time-series data of the same type of resource collected in 10 consecutive sliding windows, such as 10 consecutive computing power utilization sequences or 10 consecutive chip temperature sequences, with a single resource input dimension of 1×10. The linear regression model can fit a one-dimensional linear time-series equation using the least squares method with time as the independent variable and resource value as the dependent variable. The optimal slope k and intercept b are obtained by solving the equation, and the short-term linear trend of the resource parameter is fitted.
[0033] Therefore, the linear regression model uses the least squares method to solve for the optimal slope. k With intercept b When the slope is found, the following formula can be used to solve for the slope. k :
[0034] In the formula, k The slope; n This represents the number of samples in the sliding window, such as a fixed value of n=10. The time sequence number of the i-th acquisition moment; These are the resource observation values at the corresponding time.
[0035] Similarly, the intercept can be solved using the following formula. b :
[0036] In the formula, b The intercept; k The slope; The mean of resource observations, i.e. ; The time average, i.e. .
[0037] After obtaining the fitted parameters slope k and intercept b, the future time... t f Substituting into the equation, we obtain the future time. t f The resource forecast value, i.e. y pred = k t f +b .
[0038] Taking chip temperature prediction as an example, 10 consecutive time points are selected as the independent variable t (t=1, 2, 3, 4, 5, 6, 7, 8, 9, 10). The corresponding dependent variable y for the 10 measured chip temperatures is: [72.1℃, 72.3℃, 72.5℃, 72.6℃, 72.8℃, 73.0℃, 73.2℃, 73.3℃, 73.5℃, 73.7℃]. Substituting the above samples into the least squares formula, given a sample size n=10, the slope k is calculated to be 0.0376, and the intercept b is 72.7932. The final fitted linear equation is then obtained as follows: y =0.0376 t +72.7932.
[0039] Based on the fitted linear equation, the predicted temperature for the next 100ms can be calculated, i.e., the predicted chip temperature at time 11 is 73.207℃. The calculation results show that the chip temperature exhibits a stable, linear, and gradual upward trend, with a predicted temperature of approximately 73.21℃ for the next 100ms. There is no risk of sudden overheating, and the onboard platform can maintain standard inference operation. This demonstrates that the linear regression model has extremely low computational complexity, very low latency, and high stability, making it suitable for the typical on-orbit conditions of stable fluctuations in satellite resources.
[0040] Another state prediction model is a single-layer Long Short-Term Memory (LSTM) time-series prediction model. The spaceborne platform adopts an extremely lightweight single-hidden-layer LSTM structure, without complex stacked networks, resulting in extremely low computational overhead and adaptability to spaceborne on-orbit computation constraints. The model structure parameters are fixed in orbit, without online training, and only perform forward prediction inference: Input dimension: 10×5, corresponding to 10 consecutive acquisition windows and 5 types of resource features; Hidden layer dimension: 32-dimensional, with Sigmoid as the gating activation function and Tanh as the cell state activation function, which is a standard LSTM three-gating structure; Output dimension: 1×5, corresponding to the predicted values of 5 types of resources in one future acquisition step (100ms); The model parameter size is controlled within 10KB, and the single inference latency is less than 1ms, fully meeting the requirements of low-overhead real-time prediction on spaceborne platforms.
[0041] The operational logic of the time-series prediction model relies on the LSTM's proprietary forget gate, input gate, and output gate three-gating mechanism to automatically eliminate invalid historical noise data, focusing on remembering recent resource mutations and fluctuation characteristics, and efficiently fitting the nonlinear variation patterns of resource parameters. Therefore, the output of the time-series prediction model is a multi-dimensional resource state prediction vector corresponding to future timeframes. It is evident that the time-series prediction model can accurately capture nonlinear resource jumps caused by satellite attitude switching and concurrent multi-task loading, compensating for the limitation of linear regression models that can only fit stable, gradually changing data, and improving the accuracy of resource prediction under complex disturbance conditions.
[0042] To adapt to both on-orbit stable and gradual changes and sudden disturbances, the onboard platform can automatically select and match the prediction model based on the real-time fluctuation characteristics of onboard resources, taking into account both low computing power and high prediction accuracy.
[0043] This involves employing a dual-judgment logic of trigger threshold verification based on operating conditions to avoid false triggering issues caused by a single judgment. Under default normal operating conditions, a lightweight linear regression model is used. If resource parameters such as onboard computing power, power consumption, chip temperature, video memory, and communication bandwidth change smoothly, and the fluctuation amplitude of resource samples is less than a preset fluctuation threshold after multiple consecutive samplings, it is determined to be a normal and stable on-orbit operating condition. The onboard platform can enable the linear regression model by default, relying on its advantages of extremely low computational load, zero latency, and high stability to complete resource timing prediction and ensure low-load operation of the entire system.
[0044] In complex, abrupt changes, the system switches to a single-layer LSTM time-series prediction model. For example, when the satellite experiences attitude maneuvers, simultaneous multi-task concurrency, or sudden changes in the external environment, and when the relative fluctuation of a single resource parameter exceeds a preset ±5% threshold within 10 consecutive sampling windows, and the time-series curve exhibits a non-linear jump trend, the system automatically switches to a lightweight LSTM model. Abrupt changes can include large jumps in chip junction temperature over multiple consecutive samplings, or a sudden surge in NPU computing power utilization without a smooth, linear gradual change pattern. Utilizing the time-series memory and non-linear fitting capabilities of LSTM, the system accurately captures resource abrupt change characteristics, improving prediction accuracy under complex disturbance conditions. If there are no significant resource fluctuations and only a change in operating conditions, the model switch is not triggered, and the system continues to operate in a low-power linear prediction mode. The format, dimensions, and physical meaning of the advanced state parameters output by both prediction models are completely unified. Downstream inference strategy decision modules do not need to distinguish between prediction model types and can directly perform subsequent calculations based on the unified format of advanced state parameters, ensuring seamless system logic during operating condition switching.
[0045] The state prediction model can ultimately output multi-dimensional resource prediction results for future moments, i.e., advanced state parameters. A multi-resource coupling safety judgment mechanism and a fixed safety margin threshold are configured simultaneously. Preset core safety thresholds include: predicted remaining power budget p < 3W, predicted chip junction temperature T > 75℃, and predicted computing power utilization u > 90%. Meeting any of these thresholds immediately triggers a low-power safety inference strategy, proactively mitigating risks such as hardware overload, over-temperature, and task interruption, thus achieving proactive safety control of on-orbit resources.
[0046] S103. Extract scene association features from remote sensing monitoring images, and calculate scene complexity score based on scene association features.
[0047] After acquiring remote sensing monitoring data, adaptive assessment of the multi-dimensional complexity of multi-source remote sensing scenes can be performed based on the remote sensing images within the data. This involves extracting scene association features from the remote sensing images and calculating scene complexity scores based on these features. In addition to providing advanced predictions of onboard resource status, the satellite-based augmentation system also performs precise quantitative assessments of scene complexity from multi-source remote sensing images. This enables accurate matching of scenes, resources, and computing power, mitigating the core shortcomings of rigid scene assessment logic, inability to adapt to multi-source remote sensing data, and weak adaptive capabilities.
[0048] When conducting scene evaluation, the input standard for feature calculation can be standardized first to eliminate size ambiguity. For example, the large-scale remote sensing image input in orbit should be preprocessed and cropped into standard 1024×1024 tile sub-images. All scene-related feature calculations and scene complexity scoring are performed on a per-1024×1024 tile sub-image basis, and each sub-image undergoes only one complete feature evaluation operation without repeated iterative calculations. Smaller sizes such as 512×512 are only used for dynamic adaptation in the subsequent inference stage and are not used for scene complexity evaluation.
[0049] The use of a unified fixed feature weight and static evaluation logic fails to differentiate the imaging mechanisms, feature validity, and noise characteristics of optical, SAR, and hyperspectral remote sensing data, leading to distorted complexity scores. This results in problems such as redundant computing power in simple scenarios, insufficient computing power in complex scenarios, and poor adaptability to multi-source data. Therefore, to address these pain points, the spaceborne platform can construct a fully adaptable evaluation mechanism that includes differentiated feature extraction from multi-source data, on-orbit adaptive dynamic weight matching, and standardized quantitative scoring. It can also develop specific feature selection and weight allocation rules for mainstream spaceborne remote sensing data, accurately adapting to the complexity characteristics of different data and providing reliable scene-dimensional quantitative basis for the adaptive generation of subsequent inference strategies. To balance evaluation accuracy with lightweight on-orbit computation requirements, lightweight image feature analysis is performed before model inference, resulting in an overall computational overhead of less than 5% and no additional computing power burden.
[0050] By combining the imaging differences among the three types of remote sensing data, scene association features are selected differentially. These scene association features include at least one of edge density, texture entropy, target density, and frequency domain complexity. Each of the four core evaluation dimensions—edge density, texture entropy, target density, and frequency domain complexity—can be assigned adaptive weights to calculate the scene complexity score. That is, the scene complexity score is a weighted calculation result based on the adaptive weights applied to the scene association features. The adaptive weights are generated online based on the preset base weights corresponding to the remote sensing data type, and are fine-tuned according to the spatial dispersion of various scene association features in a single remote sensing image.
[0051] Edge density D e Edge density is a core spatial feature index for quantifying the richness of ground feature outlines and the complexity of spatial structure in remote sensing images. The core physical meaning of edge density is the normalized proportion of the number of effective object edge pixels to the total number of pixels in the entire image. For example, flat and homogeneous scenes such as farmland and sea surfaces have few edges, resulting in extremely low De values; while scenes with intersecting and dense edges, such as urban buildings, road networks, and dense artificial targets, have extremely high De values, which can intuitively distinguish between simple empty scenes and complex structured scenes.
[0052] In some embodiments, a lightweight Sobel edge detection algorithm can be used for real-time on-orbit solution. The lightweight Sobel edge detection algorithm has low computational cost, is resistant to slight noise, and can be adapted for rapid traversal calculations of large 1024×1024 remote sensing images. Therefore, when extracting scene-related features from remote sensing images, pixel edge information can be extracted from the remote sensing images through bidirectional gradient convolution operations. This pixel edge information includes gradient values output by standard operators in multiple directions and abrupt changes in ground feature contours determined based on the gradient values. Then, pixel gradient magnitudes are synthesized according to the gradient values and the abrupt changes in ground feature contours.
[0053] Then, based on a preset edge filtering threshold, valid edge pixels are selected according to pixel gradient magnitude. Furthermore, global pixel statistics and normalization are performed on all pixels in the remote sensing image to obtain the edge density. Valid edge pixels are those whose pixel gradient magnitude is greater than the edge filtering threshold. Combining the valid edge pixels, the edge density is the ratio of the total number of valid edge pixels to the total number of all pixels.
[0054] For example, after acquiring remote sensing monitoring data, SAR images can be preprocessed. Because SAR synthetic aperture radar images inherently contain speckle coherence noise, generating numerous fragmented false edges, direct calculation can lead to inflated edge density and distorted complexity assessment. Therefore, for SAR input images, adaptive Lee filtering is first used to smooth and denoise, thoroughly filtering out random speckle noise and eliminating false edge interference while completely preserving the true contours of ground objects and the edge structures of strongly scattering targets, ensuring that subsequent edge statistics only target the true and effective object edges. Optical remote sensing images, on the other hand, have a high signal-to-noise ratio and no inherent false edge noise, requiring no filtering preprocessing and directly proceeding to the gradient calculation process.
[0055] Then, Sobel bidirectional gradient convolution is used to extract pixel edge information. An edge is essentially the location where the grayscale value of an image pixel changes abruptly or jumps; the more drastic the grayscale change, the more significant the edge feature. The Sobel operator uses two sets of 3×3 standard convolution kernels to traverse all pixels in the entire image, solving for the horizontal grayscale gradient of each pixel. With vertical grayscale gradient It accurately captures abrupt changes in the outlines of ground features from all directions.
[0056] The horizontal gray-level gradient represents the change in brightness in the horizontal direction, while the vertical gray-level gradient represents the change in brightness in the vertical direction. Therefore, the horizontal Sobel convolution kernel and the vertical Sobel convolution kernel can be set as fixed standard operators to solve the changes in gray-level intensity in the left and right directions and the changes in gray-level intensity in the up and down directions, respectively.
[0057] The horizontal gradient convolution kernel can be set to: {[-1, 0, 1], [-2, 0, 2], [-1, 0, 1]}; the vertical gradient convolution kernel can be set to {[-1, -2, -1], [0, 0, 0], [1, 2, 1]}. A weighted summation operation is performed between the convolution kernel and a local 3×3 pixel region of the image. This process is repeated for all pixels in the image, and the horizontal gradient value of each pixel is output point by point. Longitudinal gradient value .
[0058] Then, single-pixel gradient magnitude is synthesized to determine pixel edge strength. To fuse grayscale abrupt changes in both horizontal and vertical directions and comprehensively characterize the edge saliency of a single pixel, the overall gradient magnitude of the current pixel can be synthesized using the Euclidean distance formula. G The magnitude of the amplitude directly corresponds to the strength of the edge, that is .
[0059] Then, a fixed threshold binarization is used to filter effective edge pixels. The image has slight lighting noise and fine texture fluctuations, which will generate a large number of weak gradient pixels with extremely small gradient amplitudes. These pixels do not belong to effective object edges. If all of them are counted, it will lead to calculation errors.
[0060] By setting edge filtering thresholds Accurate edge pixel selection was completed. For 8-bit grayscale remote sensing images with a resolution of 1024×1024, the baseline value selection rule is as follows: optical remote sensing images are selected from... =30, SAR remote sensing image after Lee filter preprocessing is taken =25. This threshold can be adaptively adjusted according to imaging conditions: using 12% of the global grayscale mean of the image as a dynamic benchmark, it is moderately increased in brighter scenes and moderately decreased in darker scenes, with the adjustment range not exceeding ±30% of the benchmark value, taking into account the edge detection accuracy under different imaging qualities.
[0061] When pixel gradient magnitude If the pixel is determined to be a real and valid object edge pixel, it will be included in the statistics; while if the pixel gradient magnitude is... If a pixel is determined to be a flat background, a pixel with weak texture, or a pixel with noise interference, it will not be included in the edge statistics.
[0062] Next, we perform global pixel statistics and normalization calculations. After traversing all pixels of the entire image, we calculate two core statistics: the total number of all valid true edge pixels after the entire image has been filtered. N edge and the total number of pixels in the input remote sensing image N total By fixing the resolution to 1024×1024 and using a fixed value, the influence of image size differences is eliminated. Finally, the standardized edge density in the 0-1 interval is obtained by proportion normalization. D e =N edge / N total .
[0063] Edge density D e If the strict normalization is in the interval 0 to 1, then D e The closer a value is to 0, the higher the proportion of flat areas in the image, the fewer the outlines of ground features, and the simpler the scene structure. D e The closer a value is to 1, the denser the image edge contours, the more overlapping and complex the terrain features, and the more intricate the spatial structure, indicating a higher level of scene structural complexity. This feature can accurately distinguish between minimalist, homogeneous scenes such as oceans and deserts and fragmented, complex scenes such as cities, settlements, and disaster-stricken areas, and is used to assess spatial structural complexity.
[0064] Texture Entropy H tTexture entropy is a quantitative indicator of image texture complexity built upon the Gray-Level Co-occurrence Matrix (GLCM). Texture entropy characterizes the disorder and detail richness of the gray-level distribution of adjacent pixels in a remotely sensed image, and is a core feature for evaluating the regularity of surface texture and the complexity of scene details. The more disordered the pixel gray-level distribution and the richer the texture details, the higher the texture entropy value; conversely, the more uniform the pixel gray-level distribution and the more uniform and regular the texture, the lower the texture entropy value.
[0065] In some embodiments, a lightweight GLCM algorithm can be used to achieve efficient on-orbit calculation of texture entropy on spacecraft. H t By reducing computational overhead through grayscale hierarchical compression, and configuring differentiated processing logic for optical, SAR, and hyperspectral remote sensing data, the problem of feature failure caused by differences in imaging characteristics is avoided. When extracting scene-related features from remote sensing images, a standardized grayscale co-occurrence matrix of the remote sensing images can be constructed through directional neighborhood fusion statistics. Then, the original texture entropy is calculated based on the grayscale co-occurrence matrix, where the original texture entropy is used to quantify the degree of texture disorder in the remote sensing images through information entropy.
[0066] The original texture entropy is then normalized to a standard range to obtain normalized texture entropy. Texture entropy weights are then set based on the image type of the remote sensing image. Specifically, when the image type is optical remote sensing data, the texture entropy weight is set as the baseline weight; when the image type is radar remote sensing data, the texture entropy weight is set as a weakened weight; and when the image type is spectral remote sensing data, the texture entropy weight is set as an auxiliary weight. The auxiliary weight is less than the weakened weight, and the weakened weight is less than the baseline weight.
[0067] For example, for remote sensing images, lightweight grayscale compression preprocessing can be performed to uniformly compress the original remote sensing image's grayscale levels 0-255 to 16 levels, reducing the GLCM matrix dimension and on-orbit computation while fully preserving the image's true texture differences, thus adapting to the needs of real-time computing on satellite. This preprocessing rule is applicable to all remote sensing data types.
[0068] Then, a standardized gray-level co-occurrence matrix (GLCM) is constructed. A four-directional neighborhood fusion statistical method (0°, 45°, 90°, 135°) is used to traverse all adjacent pixel pairs in the entire image, and various gray-level combinations are statistically analyzed. i , j The frequency of occurrence of ) is used to obtain the gray-level co-occurrence probability matrix through frequency normalization. P ( i , j The sum of the probabilities of all elements in the matrix is 1, which accurately represents the correlation pattern of the global texture distribution in the image.
[0069] Then, texture entropy quantization is performed, that is, based on the normalized GLCM probability matrix, texture entropy is calculated using the information entropy formula to quantify the degree of texture disorder, i.e.:
[0070] in, i , j This is the pixel grayscale level index after grayscale compression, not pixel coordinates; i Represents the grayscale level of the center reference pixel. j Represents the grayscale level of the paired pixels in the neighborhood; P ( i , j ) is a grayscale combination ( i , j The normalized probability of adjacent co-occurrence represents the spatial correlation distribution law of pixel gray levels. Therefore, the more gray level combination types and the more uniform the probability distribution, the higher the texture disorder and texture entropy. H t The higher the value, the simpler the grayscale combination, the more concentrated the distribution, the stronger the texture regularity, and the higher the texture entropy. H t The smaller the value.
[0071] Then, through numerical normalization, the original texture entropy value is linearly normalized to the standard range of 0 to 1, which is consistent with the dimensions of edge density, target density, and frequency domain complexity, ensuring the dimensional equivalence and weight effectiveness of multi-feature weighted fusion scoring.
[0072] Then, multi-source remote sensing data differential adaptation processing is performed. Combining the differences in imaging mechanisms and texture feature effectiveness among the three types of remote sensing data, exclusive adaptation rules are configured to avoid interference from invalid features and ensure the accuracy of complexity scoring. For optical remote sensing data, the imaging signal-to-noise ratio is high, there are no inherent false textures, and the texture features can truly reflect the density and disorder of the surface texture. The texture entropy feature is fully effective and participates in the scene complexity fusion scoring with the benchmark weight. It is one of the core reference indicators for optical scene complexity.
[0073] SAR remote sensing data contains inherent point-like speckle coherent noise, which generates a large number of irregular false textures. This results in artificially high texture entropy values, failing to represent the true characteristics of ground features, and thus constituting ineffective and interfering features. Therefore, the weight of texture entropy is significantly weakened in subsequent fusion scoring, serving only as an auxiliary reference and not participating in the primary determination of scene complexity.
[0074] For hyperspectral remote sensing data, the core of scene complexity stems from subtle differences in the spectral dimension. Conventional spatial texture features have extremely low differentiation in terms of ground feature subdivision and scene identification, and texture entropy has negligible reference value. Therefore, their weight is weakened in fusion calculations, serving only as secondary auxiliary features, with the focus on relying on frequency domain complexity to complete scene quantization.
[0075] Normalized texture entropy H t The value range is 0 to 1. Texture entropy. H t A value close to 0 indicates uniform image texture, a regular surface, and a simple scene; texture entropy H t A value close to 1 indicates complex textures, disordered grayscale distribution, cluttered surface details, and high scene complexity.
[0076] The original remote sensing images have 256 gray levels, ranging from 0 to 255. Directly constructing a GLCM matrix results in extremely high dimensionality and computational overhead, making it unsuitable for real-time inference scenarios on spacecraft. To balance computational accuracy and on-orbit efficiency, the image gray levels can be uniformly compressed to 16 quantization levels. Dimensionality reduction is achieved through uniform mapping of gray levels, significantly reducing the matrix operation scale and on-orbit computing power consumption while preserving true texture differences. This preprocessing operation is performed uniformly on all types of remote sensing images.
[0077] Target density D o Target density is a core indicator for quantifying the density and complexity of target stacking in remote sensing images. Its core physical meaning is the normalized proportion of independent and effective targets per unit image scale. For open seas and deserts with no obvious targets, the target density approaches 0; for urban building clusters, ship formations, and disaster-damaged terrain, the target density is extremely high due to dense target stacking, accurately reflecting the difficulty of target interpretation in the scene.
[0078] In some embodiments, a lightweight visual saliency detection and connected component analysis algorithm is used to solve for the target density. This eliminates the need for deep learning inference, resulting in extremely low on-orbit computational overhead and adapting to the real-time solution requirements of spaceborne systems. When extracting scene association features from remote sensing images, a global contrast saliency detection algorithm can be used to generate a saliency response map based on the remote sensing images. Then, an adaptive segmentation threshold is set based on the illumination conditions and imaging brightness of the remote sensing images, and the remote sensing images are converted into binary images based on the adaptive segmentation threshold.
[0079] Then, false targets in the binarized image are filtered out by the connected component area threshold, and the total number of effective independent target connected components in the binarized image is counted. Then, based on the total number of effective independent target connected components, the total number of pixels in the remote sensing image is normalized to obtain the target density.
[0080] Density weights can also be set based on the image type of the remote sensing image. Specifically, when the image type is optical remote sensing data, the density weight is set as the standard weight; when the image type is radar remote sensing data, the density weight is set as the enhancement weight; and when the image type is spectral remote sensing data, the density weight is set as the reference weight. The reference weight is less than the standard weight, and the standard weight is less than the enhancement weight.
[0081] For example, in solving the target density D o In such cases, image preprocessing and background suppression can be performed first. Remote sensing images often contain large areas of uniform background (sea surface, desert, bare land) without effective target interpretation, necessitating prior suppression of background interference and highlighting of foreground targets. A lightweight global contrast saliency detection algorithm can quickly generate an image saliency response map, automatically highlighting small, weak, and edge targets while suppressing large areas of low-frequency uniform background, thus avoiding background interference with target statistical accuracy.
[0082] Next, adaptive threshold binarization segmentation is performed to extract the target mask. Adaptive Otsu threshold segmentation is applied to the saliency response map to automatically distinguish foreground targets from background regions and generate a binary foreground mask map. Highlighted areas are potential valid target regions, while dark areas are invalid background regions. The adaptive threshold can be adapted to remote sensing images with different lighting and imaging brightness, avoiding missed detections and false detections caused by fixed thresholds.
[0083] Then, connected component filtering and effective target statistics are performed to eliminate false targets. Since the binarized image contains a small number of fragmented noise points and isolated pixel false targets, they need to be eliminated through connected component area thresholding. All 8-neighbor connected components in the mask image are counted, and a minimum pixel area threshold is set to eliminate false targets. For 1024×1024 resolution remote sensing images, the baseline minimum area threshold is set to 16 pixels (corresponding to a 4×4 pixel area) to eliminate single-pixel noise, fragmented spots, and other false targets. For specific task scenarios, the threshold can be adjusted based on the prior target size: for example, for large target detection scenarios such as ships and large buildings, it can be increased to 32 pixels, and for small target detection scenarios such as vehicles, it can be decreased by 9 pixels to ensure the accuracy of target statistics. Finally, the total number of effective independent target connected components in the entire image is counted and denoted as [missing value]. N obj .
[0084] Then, cross-size normalization calculations are performed to output the standardized target density. To eliminate the influence of different image resolutions and different frame sizes on the target quantity statistics and to achieve unified benchmarking quantization across scenes and data types, a pixel total normalization method is used to solve for the target density, i.e.:
[0085] In the formula, The total number of valid independent target connected components; This is the normalized baseline target count for the corresponding resolution, used to characterize the statistical upper limit of the number of effective targets within this image size. The baseline value is set proportionally to the total area of the image pixels. For example, for a 1024×1024 standard resolution remote sensing image, the baseline target count... =1024; if the image resolution changes Scaled proportionally to the total pixel area to ensure comparability of target density across different image frames. (If the actual number of effective targets...) ≥ ,but Set the value to 1 to ensure that the value range is strictly limited to the range of 0 to 1.
[0086] Target density After normalization, the value range is 0 to 1. The smaller the value, the sparser the image target and the more open the background, making it easier to interpret; A larger value indicates a greater number of independent targets, denser stacking, more complex scenarios, higher risk of missed or false detections, and greater overall scenario complexity.
[0087] Furthermore, target density weights are specifically adapted for remote sensing data types. For optical remote sensing images, ground targets have complete and prominent shapes, and target density accurately reflects the density of man-made buildings, vehicles, and settlements, with high feature reliability, and is included in the scoring with normal weight. For SAR remote sensing images, strong scattering targets (ships, bridges, buildings) have stable connectivity features and strong noise resistance, and target density is a core effective feature, so its weight is emphasized and used as the core basis for judging complex SAR scenes. For hyperspectral remote sensing images, there are mostly subtle spectral differences and no obvious large concentrated targets, and the density of conventional spatial targets has poor distinguishability, so its weight is weakened and used only as an auxiliary reference.
[0088] Target density D o After normalization, the value range is 0 to 1. Target density D o A smaller value indicates sparse targets in the image, with a predominantly empty background, making interpretation easier; target density D o A larger value indicates a greater number of independent targets, denser stacking, more complex scenarios, higher risk of missed or false detections, and greater overall scenario complexity.
[0089] Frequency domain complexity F cIt is a core frequency domain indicator for quantifying the richness of high-frequency details, the intensity of gray-level abrupt changes, and the subtle differences in spectral density in remote sensing images, and is particularly suitable for the refined evaluation of hyperspectral remote sensing data. The core physical meaning of frequency domain complexity is the normalized proportion of high-frequency detail energy to the total energy of the entire image. In flat and uniform scenes, energy is concentrated in the low-frequency region, with a very low proportion of high-frequency energy and low frequency domain complexity; while complex scenes with rich details, fragmented textures, varied spectral gradients, and dense edges of ground features have a high proportion of high-frequency energy and high frequency domain complexity.
[0090] In some embodiments, a lightweight two-dimensional Fast Fourier Transform (FFT) can be used to achieve on-orbit frequency domain analysis with low frequency domain complexity. The two-dimensional Fast Fourier Transform is computationally efficient and adaptable to spaceborne hardware conditions. When extracting scene-related features from remote sensing images, grayscale normalization can be performed on the remote sensing images first, and the pixel grayscale value range can be unified to obtain a spatial domain image. Then, the spatial domain image is mapped to a frequency domain matrix using a two-dimensional Fast Fourier Transform, where the frequency domain matrix includes low-frequency and high-frequency components.
[0091] Then, the frequency domain matrix is partitioned based on a preset frequency domain boundary threshold to obtain low-frequency and high-frequency energy through integral statistics, and the frequency domain complexity is calculated based on the low-frequency and high-frequency energy. Low-frequency energy is used to represent the basic global information of the image, while high-frequency energy is used to represent the detailed and complex information of the image. The frequency domain complexity is the ratio of high-frequency energy to the total energy of the entire image, and the total energy of the entire image is the sum of the low-frequency and high-frequency energy.
[0092] For example, when performing on-orbit frequency domain resolution to reduce frequency domain complexity, image preprocessing and standardization can be performed first. Gray-scale standardization of the input remote sensing image can be performed to unify the range of pixel gray-scale values and eliminate frequency domain calculation deviations caused by differences in illumination intensity and imaging gain. Hyperspectral images require separate standardization preprocessing for each effective spectral band to ensure the accuracy of multi-band frequency domain resolution.
[0093] Then, a two-dimensional FFT frequency domain transformation is used to convert between the spatial and frequency domains. The two-dimensional Fast Fourier Transform maps the image from the spatial domain to the frequency domain, decomposing the image signal into two core signal types: low-frequency components and high-frequency components. The low-frequency components correspond to slowly changing information such as overall image brightness, global floor, and flat background; the high-frequency components correspond to rapidly changing information such as abrupt image edge changes, subtle textures, pixel jumps, spectral gradient differences, and minute ground details.
[0094] Then, high- and low-frequency energy separation and statistics are performed. The frequency domain matrix is divided into high- and low-frequency regions by setting a frequency domain boundary radius. For a frequency domain matrix of the same size generated from a 1024×1024 resolution image, the origin is taken as the center of the spectrum, and the reference boundary radius is set to 1 / 8 of the matrix side length, i.e., 128 pixels. The area within the radius is the low-frequency region, and the area outside the radius is the high-frequency region. The total low-frequency energy is obtained by integration and statistics. E low With high frequency total energy E high The boundary radius can be adjusted within the range of 1 / 10 to 1 / 6 of the matrix side length according to the scene detail requirements: the radius is reduced when global structure evaluation is emphasized, and expanded when detail complexity evaluation is emphasized. Low-frequency energy represents the basic global information of the image, and high-frequency energy represents the detailed and complex information of the image; the sum of the two is the total energy of the entire image.
[0095] Next, frequency domain complexity normalization calculation is performed. By using the ratio of high-frequency energy to total energy, the richness of image details is quantified, achieving a 0-1 normalized output, i.e.: F c = E high / ( E low + E high This ratio completely avoids interference from image brightness and overall grayscale base, focusing only on details and abrupt changes, exhibiting extremely high stability and adaptability to comparison and evaluation across multiple scenarios and data sources.
[0096] Similarly, regarding frequency domain complexity F c Alternatively, frequency domain complexity weights can be determined through data-specific adaptation logic. For optical remote sensing images, frequency domain features effectively reflect the richness of texture and edge details, serving as auxiliary features in complexity scoring. For SAR remote sensing images, speckle noise amplifies false high-frequency energy, so frequency domain complexity is only used for auxiliary verification and does not dominate the judgment. However, for hyperspectral remote sensing images, it is the core dominant feature. Hyperspectral ground feature subdivision, camouflage feature identification, and subtle surface changes are all reflected in the differences in high-frequency gradients across multiple spectral bands, exhibiting extremely low spatial feature distinguishability. Therefore, the frequency domain complexity weight is significantly increased, using the combined frequency domain energy difference across multiple bands as the core criterion for scene complexity determination.
[0097] Frequency domain complexity F c The value range is strictly 0 to 1. Frequency domain complexity F c A value close to 0 indicates that the image energy is concentrated in the low frequency range, resulting in a flat image, sparse details, gentle spectral changes, and a simple scene; frequency domain complexity F cA value close to 1 indicates a high proportion of high-frequency energy in the image, rich spatial details, varied spectral gradients, numerous subtle differences, extremely high scene refinement complexity, and significantly increased interpretation difficulty.
[0098] Based on the four standardized lightweight solution algorithms mentioned above, and combined with the characteristics of three types of remote sensing data to perform differentiated feature selection and weight adaptation, refined processing by type can be achieved. Specifically, for optical remote sensing data, full feature balance adaptation can be performed. Optical images have high imaging quality and no inherent noise; the four features—edge density, texture entropy, target density, and frequency domain complexity—are all realistic, effective, and highly reliable. By performing full-dimensional feature balance extraction and having the four features participate equally in complexity scoring, the true complexity of various optical scenes can be comprehensively and accurately reflected, adapting to the computing power and inference strategy matching requirements of conventional optical remote sensing interpretation scenarios.
[0099] For SAR remote sensing data, structural feature enhancement and texture noise filtering can be performed. SAR images contain inherent point-like coherent noise, and texture entropy features are easily rendered ineffective and cannot be used as a basis for evaluation; however, the edge, structural, and independent target features formed by backscattering of ground objects are highly stable. Adaptive Lee filtering can be used to suppress speckle noise, weaken the weight of ineffective texture entropy, and focus on enhancing the core features of edge density and target density. Frequency domain complexity is only used for auxiliary verification, effectively avoiding noise interference, accurately identifying complex structured SAR scenes, and matching adaptive elastic inference strategies.
[0100] For hyperspectral remote sensing data, frequency domain spectral enhancement and spatial feature weakening can be performed. The core advantage of hyperspectral data lies in its refined spectral differences. Scene complexity is concentrated in spectral frequency domain features, while the distinguishability of conventional spatial textures, edges, and target features is extremely low. By using the frequency domain complexity extracted through multi-band joint FFT transformation as the core evaluation index, the weight of various spatial features is significantly weakened, which aligns with the hyperspectral characteristic of "emphasizing spectral data and de-emphasizing spatial features." This allows for accurate identification of refined and complex scenes and meets the computational power requirements of high-precision, detailed hyperspectral interpretation.
[0101] After extracting differentiated features from three types of remote sensing images, this scheme adopts a two-level adaptive weight generation mechanism to unify the complexity evaluation standards of different data sources, achieve quantitative benchmarking of scene complexity, and take into account the feature distribution characteristics of individual images. First, preset basic weights are matched according to the type of remote sensing monitoring data to match the imaging mechanism characteristics of different data. Then, online fine-tuning is performed based on the spatial dispersion of various features in a single image. The fine-tuning range of a single feature is limited to ±15% of the corresponding basic weight, achieving precise adaptation of complexity at the single-image level. Therefore, weighted fusion calculations are performed on the four types of features to output a standardized scene complexity score in the range of 0 to 1. This mechanism is entirely based on lightweight statistical operations, with the overall scene evaluation overhead still less than 5% of the total inference computing power, without additional computing power burden, and is adapted to the constraints of real-time on-orbit processing.
[0102] For basic weight configuration, corresponding basic weight combinations are pre-configured based on the type of remote sensing monitoring data, matching the differences in imaging mechanisms and feature effectiveness of different data as the benchmark for weight calculation. The basic weight configuration rules for the three types of remote sensing data are as follows: Optical remote sensing data: A balanced weighting configuration is adopted, with the basic weight of each of the four feature classes being 0.25, comprehensively covering the multi-dimensional complexity of spatial structure, texture, target, and frequency domain; Radar (SAR) remote sensing data: Strengthen the weights of edge density and target density, weaken the weight of texture entropy, and configure it as edge density 0.35, target density 0.35, texture entropy 0.1, and frequency domain complexity 0.2 to avoid the interference of speckle noise on texture features; Hyperspectral remote sensing data: The weight of frequency domain complexity is strengthened and the weight of spatial features is weakened. The configuration is as follows: frequency domain complexity 0.45, edge density 0.2, target density 0.15, and texture entropy 0.2, which meets the needs of refined evaluation of the spectral dimension.
[0103] All basic weights satisfy the normalization constraint, that is, the sum of the basic weights of the four types of features is 1.
[0104] For online adaptive fine-tuning, the basic weights can be slightly adjusted online based on the spatial dispersion of features for a single remote sensing image: features with higher dispersion have stronger spatial heterogeneity and contain richer scene information, and their corresponding weights are slightly increased; features with lower dispersion have weaker scene discrimination, and their corresponding weights are slightly decreased, thus achieving single-image-level adaptive weight adaptation.
[0105] To address this, we can first divide the 1024×1024 remote sensing image into 64 non-overlapping 128×128 sub-blocks (8×8 blocks each). For each sub-block, we calculate the four types of feature values, resulting in a 64-dimensional sample sequence for each feature. Next, we calculate the dispersion coefficient (coefficient of variation) for each feature's sample sequence. The dispersion coefficient is the ratio of the sample standard deviation σ to the sample mean μ, i.e., CV = σ / μ. A larger dispersion coefficient indicates greater spatial distribution variation and higher information discrimination for that feature within the image. Then, we fine-tune the weights. Using the base weights for the corresponding data type as a benchmark, we adjust the weights according to the relative proportions of the dispersion coefficients of the four feature types. The adjustment range for a single feature should not exceed ±15% of its base weight to avoid significant weight fluctuations that could cause scoring distortion. Finally, we perform normalization correction, re-normalizing the fine-tuned weights to ensure the sum of the weights is 1, resulting in the final image-specific adaptive weights.
[0106] Therefore, the scene complexity score is obtained by weighted summation, i.e.:
[0107] In the formula, C score The final normalized scene complexity score takes a value in the range of [0,1]. D e Image edge density; H t For texture entropy; D o For target density; F c Frequency domain complexity; w De , w Ht , w Do , w Fc For the adaptive weighted weights corresponding to the four types of features, all weights satisfy the normalization constraint, i.e.; w De + w Ht + w Do + w Fc =1.
[0108] S104. Determine dynamic inference parameters based on advanced state parameters and scenario complexity scores.
[0109] After predicting the advanced state parameters and calculating the scene complexity score, the spaceborne platform can perform adaptive elastic inference strategy decision-making based on multi-constraint fusion, that is, determining dynamic inference parameters based on the advanced state parameters and the scene complexity score. The dynamic inference parameters are dynamically adapted parameters obtained by quantifying and fine-tuning the baseline inference parameters under the target inference mode according to independent correction rules; the independent correction rules include at least one of remote sensing data type, task priority, and resource surplus; the target inference mode is obtained by interval matching based on the advanced state and scene complexity score.
[0110] like Figure 4 As shown, the multi-constraint fusion adaptive elastic reasoning strategy decision-making serves as the core decision-making hub. It can take over the results of previous resource prediction and scenario quantification, complete the generation of multi-constraint fusion adaptive elastic reasoning strategy, and specifically solve the core defects of rigid reasoning strategy, inability to dynamically adapt to scenario and resource changes, and imbalance of computing power matching.
[0111] By constructing a complete decision-making system with three-level hierarchical progressive decision-making, three-dimensional fine-tuning, and dual hysteresis stability constraints, and integrating three constraints of hardware resource security, scenario complexity, and task priority, the system achieves precise adaptation of inference parameters from coarse to fine, realizing adaptive inference effects with on-demand allocation of onboard computing power, security controllability, and accuracy controllability.
[0112] To ensure the standardization and feasibility of policy decisions, the input to the multi-constraint fusion adaptive elastic inference policy decision-making process includes two types of core data: first, resource-dimensional data, which includes a multi-dimensional resource state vector predicted for the next 100ms, such as computing power utilization, remaining power, chip junction temperature, memory occupancy, communication bandwidth, etc., as well as the resource safety margin assessment results; second, scene-dimensional data, which includes a standardized scene complexity score for remote sensing images. C score (0-1) Current remote sensing data type and scene feature complexity. The module outputs four-dimensional elastic inference strategy parameters, namely the number of model activation layers, token retention rate, computational precision, and input resolution, which directly provide executable quantitative parameters for the subsequent large-scale model security dynamic reconstruction process.
[0113] In some embodiments, when determining inference parameters based on advanced state parameters and scenario complexity scores, a preset onboard resource security threshold can be obtained first, and the target inference mode can be determined based on the advanced state parameters and the onboard resource security threshold. Specifically, if any resource indicator in the advanced state parameters triggers the onboard resource security threshold, the target inference mode is determined to be an ultra-lightweight inference mode; if none of the resource indicators in the advanced state parameters trigger the onboard resource security threshold, the target inference mode is determined based on the scenario complexity score.
[0114] Next, the baseline inference parameters corresponding to the target inference mode are obtained, and the correction parameters are determined according to the independent correction rules. These correction parameters include at least one of type correction parameters, priority correction parameters, and dynamic margin correction parameters. Then, the baseline inference parameters are adjusted according to the correction parameters to obtain the dynamic inference parameters.
[0115] For example, by abandoning single-dimensional judgment logic, a three-tiered, step-by-step, and progressively optimized decision-making mechanism can be designed, balancing millisecond-level on-orbit response speed with accuracy adaptable to complex operating conditions. Each level of decision-making has independent constraints and is reinforced at each level. The first level in the three-tiered system is a millisecond-level rapid safety rule decision, used to prioritize ensuring on-orbit operational safety. This level is the highest priority fallback decision, requiring no complex calculations; it only makes rapid judgments based on preset onboard resource safety thresholds, prioritizing the avoidance of hardware risks such as equipment overload, overheating, and power outages.
[0116] A multi-resource coupling safety judgment mechanism and a fixed safety margin threshold are configured synchronously. The core safety thresholds are preset as follows: predicted remaining power budget p < 3W, predicted chip junction temperature T > 75℃, and predicted computing power utilization u > 90%. As long as any resource indicator triggers the safety threshold, the current situation is directly determined to be a high-risk situation with tight resources. The system will force downgrade and trigger the ultra-lightweight inference mode to prioritize satellite hardware safety and ensure that the mission is not interrupted. The system ignores the impact of scenario complexity and achieves on-orbit fallback protection with safety as the priority.
[0117] The second level of the three-tiered system is lightweight intelligent adaptation decision-making, used for accurate adaptation under normal operating conditions. If no safety alarm is triggered in the first level and the resource status is within a safe range, then the system proceeds to the second level for refined adaptation decision-making. This level uses scenario complexity scoring as the core criterion, combined with preset complexity range thresholds, to match the basic inference pattern: when... The scene is classified as a minimalist scenario, such as a vast ocean, a flat desert, or a continuous stretch of bare land—a homogeneous scene without a target—and is suitable for a very light reasoning mode; when It is judged as a simple scene, such as contiguous farmland, sparse villages, open suburbs and other low target density scenes, and is suitable for lightweight inference mode; when The scene is classified as moderately complex, typically including typical towns, dense farmland, and hilly terrain with medium target density, and is suitable for standard inference mode; when Typical examples include highly structured scenarios such as urban core areas, dense building clusters, disaster-damaged areas, and port anchorages, which are identified as highly complex scenarios and conditions for enhanced inference mode adaptation are reserved. This level achieves a preliminary match between scenario complexity and inference computing power through lightweight interval determination, adapting to most conventional on-orbit interpretation conditions, as shown in Table 1.
[0118] Table 1. Parameter Comparison Table for Inference Mode;
[0119] The number of activation layers specifically refers to the effective activation count of the Transformer encoding layer. The input embedding layer and output normalization mapping layer are fixed activation layers across all modes and are not included in this statistic. Based on the inference mode parameter comparison table shown in Table 1, the baseline inference parameters corresponding to the target inference mode can be obtained, namely the number of activation layers, token retention rate, computational accuracy, and input resolution.
[0120] The third level in the three-tiered system is for fine-tuning decisions, used for multi-condition adaptation and optimization. Because the second level relies solely on interval matching to output baseline parameters, it suffers from a lack of granularity, failing to distinguish between differences in scene, data type, and task value within the same interval. Therefore, the third level, based on the baseline parameters, uses three independent correction rules—remote sensing data type, task priority, and resource abundance—to perform small-scale quantitative fine-tuning of the four-dimensional inference parameters, eliminating coarse-matching errors and achieving fine-tuned optimal adaptation. All corrections are small and controllable adjustments, without overturning the basic model or causing parameter mutations, ensuring strategy smoothness while improving scene adaptation accuracy.
[0121] When correcting parameters specific to remote sensing data types, the inherent interpretation characteristics of the three types of remote sensing data can be combined to fine-tune the accuracy and token retention strategies. For optical remote sensing data, the scene details are uniform and the feature reliability is high, so no major corrections are made, and the second-level basic inference parameters remain unchanged to ensure efficient use of computing power. For SAR remote sensing data, SAR images rely on fine edges and weakly scattering targets to show complex features, and fixed-scale cropping can easily lose small targets. Therefore, the token retention rate is slightly increased (+10%) on the basic model, while the original accuracy level remains unchanged. By retaining more structural feature tokens, small and weak targets such as ships and small artificial buildings are not mistakenly cropped. For hyperspectral remote sensing data, hyperspectral interpretation depends on subtle spectral differences, is sensitive to computational accuracy, and has a higher tolerance for the number of tokens. Therefore, the computational accuracy level is increased (e.g., INT8~FP16) on the basic model, and the redundant token retention rate is moderately reduced (-5%). High-precision calculation is prioritized to retain spectral subdivision features, avoiding the loss of subtle ground feature differences due to low-precision quantization.
[0122] When prioritizing tasks based on their attributes, differentiated precision biases can be applied to different task types within the same complexity scenario, ensuring the interpretation effectiveness of high-value tasks. For high-priority disaster monitoring tasks: the token retention rate is increased by 10% and the computational precision level is increased, while appropriately sacrificing some computing power and latency to maximize the preservation of detailed features of disaster areas and ensure the accuracy of disaster boundary and damaged area identification; for routine survey tasks, the second-level basic parameters are used without additional adjustments to achieve balanced computing power consumption; for low-priority storage inspection tasks: the token retention rate is moderately reduced by 10% and the computational precision is lowered to further compress power consumption and reserve resources for core on-orbit tasks.
[0123] When dynamically adjusting margins based on resource abundance, the results of the first-level resource prediction can be combined to enhance accuracy when hardware safety margins are sufficient, and to implement lightweight backup when resources are scarce. Specifically, for resource-rich conditions (computing power utilization <60%, temperature <70℃, sufficient remaining power): the number of activation layers is increased (+1 to 2 layers), and the token retention rate is increased (+5% to 10%), making full use of idle computing power and slightly improving the interpretation accuracy in complex scenarios. For resource-critical conditions (all indicators approaching safety thresholds): accuracy and token quantity are no longer increased; instead, the token retention rate is slightly reduced (-5% to 10%), and the current accuracy level is locked without upgrading, ensuring stable inference processes that do not exceed limits or trigger hardware protection.
[0124] Through the three types of quantization correction rules mentioned above, the optimal four-dimensional inference parameters are finally output, taking into account safety, real-time performance, task priority, and scenario characteristics, achieving a two-layer precise adaptation of coarse matching and fine-tuning. To address the issues of frequent policy switching and system instability caused by small on-orbit disturbances and to ensure smooth and stable long-term operation, a dual constraint mechanism of dynamic task priority scheduling and three inference hysteresis stability control checks is also designed to provide a safety net for stability from both the task scheduling and policy output dimensions.
[0125] The first layer: Task priority scheduling mechanism. For scenarios involving multiple satellite tasks operating concurrently, a fixed task priority order is preset: disaster monitoring tasks > routine survey tasks > storage inspection tasks. When multiple tasks compete for computing resources, high-priority tasks can preferentially utilize NPU computing power and power resources, and preferentially match high-precision, high-computing-power inference strategies; low-priority tasks actively adapt to lightweight inference mode, achieving on-demand allocation and reasonable scheduling of computing resources, avoiding accuracy degradation of critical tasks.
[0126] The second layer: Hysteresis stabilization mechanism. To avoid frequent policy jumps and system operation fluctuations caused by instantaneous fluctuations in on-orbit resources and minor perturbations in scene complexity, this invention sets up a three-stage inference hysteresis judgment window to complete policy stability constraints. This invention uniformly defines a single inference as a single complete inference processing flow for a single remote sensing image, without using fixed duration or resource sliding windows as counting units. Each time the entire process of decision-making and inference for an image is completed, it is counted as one valid judgment result. The policy takes effect using a hysteresis verification logic: the inference mode of a single judgment does not take effect immediately; a mode switch can only be completed after three consecutive inferences with completely consistent inference strategies. Temporary policy fluctuations caused by one or two inferences are uniformly judged as invalid noise, and the previous stable mode is retained for continued operation. This mechanism eliminates frequent policy fluctuations and significantly improves the stability of long-term on-orbit operation.
[0127] By linking a three-level hierarchical decision-making system with a dual stability control constraint mechanism, it can output four-dimensional standardized, directly executable elastic inference parameters such as model activation layer number, token retention rate, computational accuracy, and input resolution. This achieves a comprehensive adaptive decision-making effect of "safety as a safety net without exceeding limits, scenario adaptation without waste, task priority to ensure accuracy, and stable strategy switching without fluctuations".
[0128] S105. Reconstruct the model parameters of the visual model based on the dynamic inference parameters.
[0129] After obtaining the dynamic inference parameters, a safe and controllable dynamic structural reconstruction can be performed on the visual model, that is, the model parameters of the visual model can be reconstructed based on the dynamic inference parameters. The model parameters of the visual model include model depth, model width, and model accuracy.
[0130] As an adaptive inference strategy, it can dynamically reconstruct visual models such as ViT-Base in orbit using four-dimensional elastic parameters, solving the core defects of dynamic reconstruction technology such as lack of security constraints, disordered layer skipping that easily damages feature links, crude pruning and quantization, and poor inference stability. For example, based on the native standard ViT-Base model, a dynamic reconstruction mechanism with three-dimensional coordination of depth, width, and accuracy is designed. All adjustments are lightweight adaptation operations within the original structure, without adding network parameters, destroying the topology, or causing abrupt changes. It accurately matches multi-source remote sensing scenarios and dynamic in-orbit conditions, achieving dual protection of computing power on demand and inference stability.
[0131] like Figure 5 As shown, in some embodiments, when reconstructing the model parameters of the visual model based on dynamic inference parameters, at least one reconstruction item can be executed. The first reconstruction item is: traversing the elastically adjustable layers of the visual model, where the elastically adjustable layers are intermediate encoding layers of the visual model; determining the number of consecutively activated layers of the elastically adjustable layers based on the dynamic inference parameters to reconstruct the model depth of the visual model.
[0132] The second reconstruction item is: obtaining the label scoring index, and scoring the value of the corresponding image serialization labels of the visual model according to the label scoring index. The label scoring index includes at least one of the following: fused attention weight, edge strength, texture variance, and position prior index; and progressively cropping the number of image serialization labels of the visual model according to the value scoring results to reconstruct the model width of the visual model. The third reconstruction item: Based on the inference mode and remote sensing data type corresponding to the dynamic inference parameters, set the quantization precision level of the activated elastically adjustable layer to reconstruct the model precision of the visual model.
[0133] The quantization precision level is matched to the four-level inference mode baseline: the ultra-light inference mode and lightweight inference mode use INT4 quantization precision by default; the standard inference mode uses INT8 quantization precision by default; and the enhanced inference mode uses FP16 floating-point precision by default. Based on this, fine-tuning is performed according to the remote sensing data type and task priority: SAR remote sensing data maintains the baseline precision unchanged; hyperspectral remote sensing data is locked at a minimum of INT8 precision, and the INT4 level is disabled; high-priority tasks can increase the precision by one level, and low-priority tasks can decrease the precision by one level.
[0134] For example, the reconstructed carrier is the native 12-layer ViT-Base standard backbone model. The 12-layer Transformer encoding layer is the inherent baseline structure of the model. The complete native topology of the model is: 1 input embedding layer, 12 Transformer encoding layers, and 1 output normalization mapping layer.
[0135] During visual model reconstruction, all elastic adjustments are performed within the original fixed topology. Adaptive reconstruction is achieved solely through controllable continuous layer activation, multi-stage feature selection, and hierarchical differentiated quantization, without structural alteration, random deletion, or topological destruction. Combining a four-level inference baseline model with three-level refined correction rules, a three-dimensional elastic constraint reconstruction mechanism can be used to strictly match the adaptation requirements of multiple scenarios and operating conditions.
[0136] For the depth elastic adjustment mechanism, the number of model inference layers can be dynamically and adaptively adjusted. The model depth corresponds to the effective number of inference layers in the Transformer encoding layer, which determines the model's global feature extraction and deep semantic representation capabilities. To address the issues of dynamic layer skipping, discreteness, disorder, and the potential for feature fragmentation, a safety control mechanism can be implemented, which locks the first and last layers of the visual model, continuously activates intermediate layers, and progressively adapts the depth, ensuring the integrity of feature transmission from the structural root.
[0137] Therefore, the input embedding layer and the output normalization mapping layer remain active throughout the process, without skipping or downgrading, locking the core link between shallow basic feature input and deep global feature output. The 12-layer Transformer coding layer adopts a continuous forward activation rule: starting from the first coding layer, K coding layers are activated consecutively in the inference order, and unactivated coding layers are skipped directly. Features are directly passed to the output normalization mapping layer from the last activated coding layer, ensuring that features are passed layer by layer in an orderly manner without the risk of skipping or breaking layers.
[0138] The depth layer matching is based on a four-level baseline inference mode, such as ultra-light mode with continuous activation of 2 layers, lightweight mode with continuous activation of 4 layers, standard mode with continuous activation of 8 layers, and enhanced mode with full activation of 12 layers. On top of the large depth matching, a fine-tuning mechanism of ±1 to 2 layers is added. Without changing the basic inference mode or causing abrupt changes in model depth, the number of layers is dynamically adjusted slightly based on on-orbit resource availability and task priority. When resources are plentiful, the number of network layers is appropriately increased to enhance detailed features; when resources are scarce, the number of layers is appropriately reduced to decrease computational load, achieving fine-tuned adaptation to different operating conditions and balancing inference stability with on-orbit adaptive optimization capabilities.
[0139] The width elastic adjustment mechanism enables dynamic cropping and control of the model's token feature dimensions. The model inference width corresponds to the number of image serialization token features, directly determining the amount of detailed feature information the model carries. The original 1024×1024 resolution input image, after block encoding, generates a fixed 4096-dimensional baseline token sequence. Through a refined width control mechanism involving multi-index value screening, progressive step-by-step cropping, and mandatory retention of key remote sensing features, the effective remote sensing features are preserved to the maximum extent while reducing computational overhead.
[0140] The algorithm scores and ranks all tokens based on a fusion of four metrics: attention weight, edge strength, texture variance, and location prior. High-value tokens from target regions, complex edge regions, and high-frequency detail regions are prioritized for retention, while redundant low-value tokens from flat, empty areas and featureless backgrounds are removed. A progressive pruning strategy is employed, with no more than 30% of redundant tokens removed in a single stage to prevent abrupt loss of detail due to large-scale pruning in a single step. The token retention rate benchmark strictly matches a four-level inference model, corresponding to fixed percentages of 20%, 30%, 60%, and 100%.
[0141] It also supports fine-grained correction for multiple scenarios. Specifically, for SAR remote sensing scenarios with small targets and weak edge scattering, it slightly improves the token retention rate to ensure that fine structural features such as ships and small man-made buildings are not lost. For hyperspectral remote sensing scenarios, it appropriately removes spatially redundant tokens and focuses on subdivided features in the spectral dimension. For high-priority disaster monitoring tasks, it prioritizes the preservation of high-value tokens in edge and damaged areas to ensure the accuracy of disaster interpretation.
[0142] The precision elastic adjustment mechanism can be used for dynamic switching and control of model quantization calculation precision. The model calculation precision determines the fineness of feature operations, directly affecting the representation capability of weak and subtle remote sensing features. Through a quantization adaptive mechanism that features high precision fixation at the beginning and end layers, multi-level elastic switching in the middle layers, precision locking within a single inference, and precision protection for multiple scenarios, computational waste is reduced.
[0143] By forcibly fixing high-precision computation in the core layers at the beginning and end of the model, and fixing FP16 high-precision floating-point computation throughout the input embedding layer and output mapping layer, without participating in low-precision quantization switching, the distortion of input original image features and drift of output inference results are avoided; the adjustable layers in the middle from the 2nd to the 11th layers support adaptive switching of three quantization precision levels: FP16, INT8, and INT4, which can accurately adapt to different complexity scenarios and resource conditions.
[0144] Furthermore, the precision level is strongly linked and matched with the scene complexity, task priority, and remote sensing data type. That is, FP16 high-precision computing is enabled for highly complex scenes, enhanced inference mode, and high-priority disaster tasks to preserve subtle details and subdivided features; INT8 balanced precision is enabled by default for medium-complexity standard scenes to achieve a two-way balance between computing power and precision; and INT4 low-precision quantization is adaptively switched for simple scenes and resource-constrained working conditions to maximize on-orbit power consumption and latency.
[0145] Simultaneously, dual dedicated protection constraints are implemented. Precision locking within a single inference iteration maintains a single precision level throughout the entire inference process for each image, prohibiting dynamic switching within a single inference iteration and avoiding computational instability caused by multi-precision mixed operations. Furthermore, dedicated precision protection for hyperspectral imaging, addressing the high sensitivity to subtle spectral differences, forcibly disables extremely low INT4 precision and retains at least INT8 precision, preventing weak spectral features from being obscured by quantization noise and ensuring high-resolution hyperspectral interpretation capabilities.
[0146] S106. Use the reconstructed visual model to perform model inference on the remote sensing monitoring images to generate model prediction results.
[0147] like Figure 6 As shown, after reconstructing the visual model based on the dynamic inference parameters, the reconstructed visual model can be used to perform adaptive inference, that is, to perform model inference on remote sensing monitoring images using the reconstructed visual model to generate model prediction results.
[0148] Dynamic inference features adjustable layers, customizable tokens, and switchable quantization precision. Although structural defects have been mitigated through security constraints, subtle differences in feature representation capabilities still exist between different inference modes. In lightweight and ultra-lightweight scenarios, minor accuracy deviations such as missed detection of small targets, misclassification of ground features, and positioning offsets can easily occur. To unify the output accuracy benchmark across multiple modes, compensate for controllable accuracy losses caused by flexible adaptation, and ensure consistency in inference across multiple source scenarios and operating conditions, a comprehensive guarantee mechanism can be constructed, including multi-dimensional fusion confidence assessment, three-level differentiated handling, inference-level stability verification, and on-orbit closed-loop feedback. This mechanism will establish a complete technical closed loop encompassing adaptive decision-making, secure reconstruction, accurate inference, and stable output.
[0149] To clearly illustrate the adaptive capabilities of the satellite-borne platform, the following section provides a complete example of typical high-value operational conditions under which the satellite is in orbit. The typical operational conditions for the satellite-borne platform involve the satellite performing SAR remote sensing disaster monitoring tasks in orbit. Real-time resource status includes a predicted chip junction temperature of 72°C for the next 100ms, a computing power utilization rate of 58%, and a remaining power of 12W, indicating ample resource margin. The input SAR disaster scene image is processed through complexity calculations to obtain… C score =0.72, which belongs to a medium-complexity scenario; the task priority is the highest level of disaster monitoring.
[0150] By predicting high-risk indicators, resources can be used to predict the absence of high-risk indicators, thus avoiding triggering the safety fallback strategy. The scenario score of 0.72 falls within the 0.5-0.8 range, matching the basic parameters of the standard mode, namely 8-layer activation, 60% token retention rate, INT8 accuracy, and 512 resolution. Further refined parameter corrections are made, increasing the token retention rate by 10% to 70% for the characteristics of small SAR targets, and increasing the number of activation layers by 1 to 9 layers based on sufficient resource margins, while maintaining the high accuracy level for high-priority disaster tasks. Then, a 3D safety reconstruction is performed to complete the model adaptive configuration according to the corrected parameters, with continuous activation, progressive pruning, and no quantization degradation throughout the process. On-orbit inference is completed to obtain high-confidence disaster interpretation results.
[0151] Therefore, the above example can effectively reduce inference power consumption and latency while ensuring the accuracy of disaster area and weak target identification. Compared with the fixed full inference mode, the computing power consumption is greatly reduced, while avoiding the problem of loss of disaster details caused by excessive compression in the lightweight mode, thus achieving the optimal balance between accuracy and energy efficiency.
[0152] By applying the technical solutions of the above embodiments, the dynamic prediction method for spaceborne resource constraints and scene complexity described in the above embodiments can construct a serial, closed-loop progressive, safe and controllable elastic prediction system for spaceborne models. The overall system follows the execution logic of resource state perception and advance prediction, multi-dimensional complex quantitative assessment of remote sensing scenes, multi-constraint fusion adaptive strategy decision-making, safe dynamic structural reconstruction of large models, and inference execution. Resource perception and prediction address the shortcomings of static, fixed inference structures that cannot adapt to dynamic fluctuations in spaceborne resources. Scene assessment and strategy decision-making address the shortcomings of rigid scene assessment, poor adaptive capability of inference strategies, and imbalance in computing power matching. Safe dynamic reconstruction addresses the shortcomings of unconstrained model dynamic reconstruction, easy feature fragmentation, and weak stability.
[0153] In some embodiments, as a refinement and extension of the specific implementation of the above embodiments, in order to fully illustrate the specific implementation process of this embodiment, some embodiments of this application also provide a dynamic prediction method for spaceborne resource constraints and scenario complexity. The difference between this method and the above embodiments is that a confidence level closed-loop verification can be performed after generating the model prediction results. Figure 7 As shown, the method includes: S201. Calculate the classification confidence and location confidence of the model prediction results; S202. Calculate the comprehensive confidence score for a single inference based on the classification confidence score and the location confidence score, wherein the comprehensive confidence score is the weighted sum of the classification confidence score and the location confidence score; S203. Obtain a preset confidence threshold, wherein the confidence threshold includes a first confidence threshold and a second confidence threshold; the first confidence threshold is greater than the second confidence threshold; S204. When the overall confidence level is greater than or equal to the first confidence threshold, output the model prediction result; S205. When the overall confidence level is less than the first confidence threshold but greater than or equal to the second confidence threshold, the adapter correction network is used to correct the feature loss of the model prediction results. S206. When the overall confidence level is less than the second confidence threshold, perform re-inference by temporarily optimizing the inference parameters.
[0154] After completing the model's safe dynamic reconstruction, the spaceborne platform can perform on-orbit inference operations according to the matched input resolution, number of activation layers, token retention rate, and quantization accuracy parameters. To balance the accuracy of remote sensing interpretation classification with the precision of target positioning, and to avoid the problems of biased and large errors in single confidence level evaluations, a classification and positioning weighted fusion confidence calculation model adapted to the characteristics of flexible inference can be constructed. This model standardizes the solution for the comprehensive confidence level of a single inference iteration, such as:
[0155] In the formula: Conf For overall confidence level; Conf cls The confidence score for remote sensing target classification characterizes the accuracy of ground feature category identification; Conf loc The location confidence score for the target can be the average intersection-over-union ratio (IoU), which characterizes the accuracy of the target bounding box and the outline of the disaster area. The weights are configured according to the characteristics of the remote sensing interpretation task, with classification weights being higher than positioning weights. This aligns with the interpretation requirements of prioritizing broad category identification and providing fine-grained positioning assistance in remote sensing scenarios, and is suitable for interpretation scenarios involving optical, SAR, and hyperspectral multi-source data.
[0156] The accuracy fluctuation characteristics of the four-level elastic inference mode can be addressed by setting a three-level hierarchical confidence judgment and differentiated handling mechanism. For inference results in different confidence intervals, a dedicated correction and output strategy can be matched to accurately adapt to the accuracy loss patterns of different reconstruction modes.
[0157] Level 1: High-confidence valid results ( Conf (≥0.7). The current inference result has accurate category identification, minimal target localization deviation, and no significant loss of accuracy during elastic reconstruction. It is determined to be a valid and compliant result, requiring no additional correction, and the final remote sensing interpretation result is directly output. Results in this range mostly appear in standard mode and enhanced mode inference scenarios. The model has sufficient activation layers, complete token features, and high quantization accuracy. The feature representation capability after reconstruction is stable, and it is suitable for most conventional on-orbit interpretation scenarios.
[0158] Level 2: Medium confidence bias results (0.5 ≤ Conf < 0.7). The current inference results exhibit slight feature loss, location offset, or category confusion, mostly normal accuracy fluctuations caused by lightweight or ultra-lightweight elastic inference modes, which are within a correctable range. In this case, due to reduced model activation layers, moderate token pruning, and low-precision quantization, weak features and small target features are not fully represented, but no complete inference failure occurs. By activating the onboard lightweight Adapter correction network, detail compensation, feature repair, and bias correction are performed on the elastically reconstructed feature map to compensate for feature loss caused by shallow inference, low-dimensional tokens, and low-precision quantization. The Adapter correction network adopts a bottleneck residual structure, embedded as a plug-in into the MLP module of each activated Transformer encoding layer, independent of the backbone weights. The single-layer Adapter structure consists of: a linear dimensionality reduction layer (768-dimensional input, 64-dimensional output) + a GELU activation layer + a linear dimensionality increase layer (64-dimensional input, 768-dimensional output), and a residual connection is introduced to add the input and output features. The structure has a very low number of parameters, with the number of parameters in a single Adapter layer being less than 2% of the backbone coding layer. It is only loaded and enabled during the correction phase, so it does not affect the computational overhead of regular inference. Corresponding Adapter weights are pre-trained for different inference modes, and the matching weights are directly called in orbit to complete feature correction without the need for online learning.
[0159] Level 3: Low-confidence failure results ( Conf <0.5). The current inference results show chaotic category identification and severe target positioning deviation, indicating inference failure. This often occurs under resource-critical conditions, in extremely complex scenarios, and in scenarios with multiple tasks and high loads. Low computing power configurations cannot support the interpretation requirements of complex scenarios. The system automatically triggers a high-precision forced re-inference process, temporarily optimizing inference strategy parameters: appropriately increasing the number of model activation layers, restoring high-value tokens, and switching to the INT8 / FP16 high-precision quantization level, and recompleting model reconstruction and the entire inference process. If the confidence level still fails to meet the standard after re-inference, it is determined that the current on-orbit resource conditions cannot adapt to the interpretation requirements of the complex scenario. Abnormal samples are automatically marked, the original remote sensing data and on-orbit operating parameters are retained, and the data is downloaded to the ground terminal for fine processing to avoid erroneous interpretation results.
[0160] Simultaneously, it can be deeply integrated with hysteresis stabilization and multi-scenario adaptive correction mechanisms to establish inference-level stability closed-loop verification logic. For example, when three consecutive inferences show the same type of low confidence deviation, it is determined that the current inference mode is not adaptable enough, and the parameter correction signal is actively fed back to the strategy generation process to slightly optimize the scenario scoring weight and strategy switching threshold, thereby achieving adaptive dynamic fine-tuning under operating conditions. For scenarios prone to distortion, such as weak features of small SAR targets and subtle spectral differences in hyperspectral data, the confidence verification sensitivity is improved separately, and the reliability of fine features is enhanced.
[0161] Through a multi-layered guarantee mechanism of graded confidence verification, on-board lightweight correction, high-precision forced re-inference, and on-orbit closed-loop feedback, the controllable accuracy fluctuation risk caused by the three-dimensional elastic reconstruction of depth, width, and precision is completely eliminated. On the basis of achieving on-demand adaptive adaptation of the computing power of the large on-board model, the system ensures controllable inference accuracy and high system stability, forming a complete, reliable, and long-term on-orbit elastic inference technology system.
[0162] By applying the technical solutions of the above embodiments and based on the complete implementation process described above, stable, efficient, and high-precision remote sensing intelligent interpretation can be achieved in a spaceborne confined environment. The overall implementation effect of the method is significantly better than that of traditional fixed inference schemes, and the various quantitative indicators and technical advantages are clear, verifiable, and feasible.
[0163] Through a flexible adaptation mechanism driven by both scenarios and resources, simple scenarios are extremely lightweight, while complex scenarios are precisely enhanced, resulting in a significant improvement in energy efficiency. Actual measurements show an average reduction of 54.7% in inference power consumption and a 56.9% reduction in average inference latency, greatly saving onboard limited energy and computing resources, and adapting to the long-term low-power operation requirements of satellites in orbit.
[0164] Through advanced resource prediction and a tiered security safety net mechanism, the hardware operation is made safe and controllable. The peak operating temperature of the chip is stably controlled below 78℃, effectively avoiding hardware risks such as overheating, computing overload, and task interruption, and meeting the stringent thermal control and load constraints of onboard equipment throughout the entire process.
[0165] Leveraging secure progressive reconstruction, priority preservation of high-value features, and on-board self-correction mechanisms, inference accuracy loss is controllable. Typical tests were conducted using the DIOR and NWPUVHR-10 optical remote sensing datasets, the SSDDSAR target detection dataset, and the Pavia University hyperspectral land cover classification dataset. Based on the inference results of the native model with all 12 coding layers, FP16 accuracy, and 1024×1024 resolution, the average accuracy loss for moderately complex scenes under standard inference mode is approximately 4.6%; the weighted average accuracy loss across all scenes is controlled within 8%. This ensures the core accuracy requirements of various remote sensing interpretation tasks while significantly reducing computational and power consumption.
[0166] With the integration of three-stage inference hysteresis stabilization, task priority scheduling, and closed-loop feedback optimization mechanisms, the system effectively suppresses dynamic inference strategy fluctuations, feature gaps, and result drift, thereby improving system stability. For medium-confidence bias results generated by elastic reconstruction, the accuracy rate after adapter correction network repair is approximately 92%. Combined with the low-confidence forced re-inference mechanism, the overall effective handling rate of abnormal results can reach over 85%, ensuring long-term stable operation in orbit.
[0167] The method can also adaptively adapt to three types of remote sensing data: optical, SAR, and hyperspectral, specifically addressing the problems of evaluation distortion and poor adaptability caused by different imaging mechanisms. It provides full coverage of various on-orbit mission scenarios, including routine surveys, detailed interpretation, and disaster monitoring, enhancing the generalization performance of multi-source scenarios and improving engineering feasibility.
[0168] In some embodiments, as a specific implementation of the dynamic prediction method for spaceborne resource constraints and scene complexity in the above embodiments, some embodiments of this application also provide a dynamic prediction system for spaceborne resource constraints and scene complexity, such as... Figure 8 As shown, the system includes: The data acquisition module is used to acquire multi-dimensional resource parameters and remote sensing monitoring data of the spaceborne platform; the remote sensing monitoring data includes any type of remote sensing monitoring image from optical remote sensing data, radar remote sensing data, and spectral remote sensing data. The resource state prediction module is used to predict leading state parameters based on multidimensional resource parameters using a state prediction model; the state prediction model includes at least one of a linear regression model and a time series prediction model. The scene detection module is used to extract scene association features from remote sensing monitoring images and calculate scene complexity scores based on these features. The scene association features include at least one of edge density, texture entropy, target density, and frequency domain complexity. The scene complexity score is a weighted calculation result of the scene association features based on adaptive weights. The adaptive weights are generated online by fine-tuning the spatial dispersion of various scene association features in a single remote sensing monitoring image, based on a preset basic weight corresponding to the remote sensing monitoring data type. The inference parameter determination module is used to determine dynamic inference parameters based on advanced state parameters and scene complexity scores. The dynamic inference parameters are dynamically adapted parameters obtained by quantifying and fine-tuning the baseline inference parameters under the target inference mode according to independent correction rules. The independent correction rules include at least one of remote sensing data type, task priority, and resource abundance. The target inference mode is obtained by interval matching based on advanced state and scene complexity scores. The model reconstruction module is used to reconstruct the model parameters of the visual model based on the dynamic inference parameters. The model parameters of the visual model include model depth, model width, and model accuracy. The model inference module is used to perform model inference on remote sensing images using the reconstructed visual model to generate model prediction results.
[0169] It should be noted that other corresponding descriptions of the functional units involved in the dynamic prediction system for spaceborne resource constraints and scene complexity provided in the embodiments of this application can be found in the corresponding descriptions in the dynamic prediction method for spaceborne resource constraints and scene complexity provided in the above embodiments, and will not be repeated here.
[0170] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0171] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application's patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application.
Claims
1. A dynamic prediction method for spaceborne resource constraints and scenario complexity, characterized in that, The method includes: Acquire multidimensional resource parameters and remote sensing monitoring data of the spaceborne platform; the remote sensing monitoring data includes any type of remote sensing monitoring image from optical remote sensing data, radar remote sensing data, and spectral remote sensing data. The state prediction model is used to predict leading state parameters based on the multidimensional resource parameters; the state prediction model includes a linear regression model and a time series prediction model. Scene association features are extracted from the remote sensing monitoring images, and a scene complexity score is calculated based on the scene association features. The scene association features include edge density, texture entropy, target density, and frequency domain complexity. The scene complexity score is a weighted calculation result of the scene association features based on adaptive weights. The adaptive weights are generated online by fine-tuning the spatial dispersion of various scene association features in a single remote sensing monitoring image, with a preset basic weight corresponding to the remote sensing monitoring data type as the benchmark. Determining dynamic inference parameters based on the advanced state parameters and the scene complexity score includes: obtaining a preset spaceborne resource security threshold; determining a target inference mode based on the advanced state parameters and the spaceborne resource security threshold; wherein the target inference mode is obtained by interval matching based on the advanced state parameters and the scene complexity score; when any resource indicator in the advanced state parameters triggers the spaceborne resource security threshold, the target inference mode is determined to be an extremely lightweight inference mode; when none of the resource indicators in the advanced state parameters trigger the spaceborne resource security threshold, the target inference mode is determined based on the scene complexity score; obtaining the baseline inference parameters corresponding to the target inference mode; determining correction parameters according to independent correction rules, wherein the independent correction rules include remote sensing data type, task priority, and resource surplus; the correction parameters include type correction parameters, priority correction parameters, and dynamic margin correction parameters; adjusting the baseline inference parameters according to the correction parameters to obtain the dynamic inference parameters, wherein the dynamic inference parameters are dynamically adapted parameters obtained by quantifying and fine-tuning the baseline inference parameters under the target inference mode according to the independent correction rules. The model parameters of the visual model are reconstructed based on the dynamic inference parameters, and the model parameters of the visual model include model depth, model width, and model accuracy. The reconstructed visual model is used to perform model inference on the remote sensing images to generate model prediction results.
2. The method according to claim 1, characterized in that, The advanced state parameters include one of the resource prediction values and resource prediction vectors at future times; predicting the advanced state parameters based on the multidimensional resource parameters using a state prediction model includes: The resource fluctuation evaluation index is calculated based on the multidimensional resource parameters, and the resource fluctuation evaluation index includes the resource fluctuation amplitude from multiple consecutive samples. Construct a multidimensional resource state vector based on the aforementioned multidimensional resource parameters; When the resource fluctuation amplitude is less than or equal to a preset fluctuation threshold, the linear regression model is used to fit the linear time series equation, and the resource prediction value at future time is calculated based on the linear time series equation and the multidimensional resource state vector. When the resource fluctuation amplitude is greater than a preset fluctuation threshold, The time-series prediction model is used to memorize the near-term resource mutation and fluctuation characteristics of the multidimensional resource parameters, and to output a resource prediction vector for future times based on the resource mutation and fluctuation characteristics. The resource mutation characteristics and fluctuation characteristics are used to fit the nonlinear variation law of the multidimensional resource parameters.
3. The method according to claim 1, characterized in that, Extracting scene association features from the remote sensing images includes: Pixel edge information is extracted from the remote sensing image by bidirectional gradient convolution operation. The pixel edge information includes gradient values output by multiple directional standard operators and ground feature contour abrupt change features determined based on the gradient values. Pixel gradient magnitudes are synthesized based on the gradient values and the abrupt change features of the ground cover contours; Based on a preset edge filtering threshold, valid edge pixels are filtered out according to the pixel gradient magnitude, wherein the valid edge pixels are pixels whose pixel gradient magnitude is greater than the edge filtering threshold; By combining the effective edge pixels, global pixel statistics and normalization calculations are performed on all pixels in the remote sensing image to obtain the edge density; the edge density is the ratio of the total number of effective edge pixels to the total number of all pixels.
4. The method according to claim 1, characterized in that, Extracting scene association features from the remote sensing images includes: A standardized gray-level co-occurrence matrix of the remote sensing image is constructed by directional neighborhood fusion statistics; The original texture entropy is calculated based on the gray-level co-occurrence matrix, and the original texture entropy is used to quantify the degree of texture disorder of the remote sensing image through information entropy; The original texture entropy is normalized to a standard range to obtain normalized texture entropy.
5. The method according to claim 1, characterized in that, Extracting scene association features from the remote sensing images includes: A global contrast saliency detection algorithm is used to generate a saliency response map based on the remote sensing monitoring image; An adaptive segmentation threshold is set based on the illumination conditions and imaging brightness of the remote sensing monitoring image, and the remote sensing monitoring image is converted into a binary image based on the adaptive segmentation threshold. False targets in the binarized image are removed by filtering out false targets using a connected component area threshold. Count the total number of valid independent target connected components in the binarized image; The target density is obtained by normalizing the solution based on the total number of effective independent target connected components and the total number of pixels in the remote sensing image.
6. The method according to claim 1, characterized in that, Extracting scene association features from the remote sensing images includes: The remote sensing image is subjected to grayscale normalization processing, and the pixel grayscale value range is unified to obtain a spatial domain image; The spatial domain image is mapped to a frequency domain matrix by a two-dimensional fast Fourier transform, and the frequency domain matrix includes low-frequency components and high-frequency components. The frequency domain matrix is partitioned in the frequency domain based on a preset frequency domain boundary threshold, so as to obtain low-frequency energy and high-frequency energy through integral statistics; the low-frequency energy is used to characterize the global basic information of the image, and the high-frequency energy is used to characterize the detailed and complex information of the image. The frequency domain complexity is calculated based on the low-frequency energy and the high-frequency energy. The frequency domain complexity is the ratio of the high-frequency energy to the total energy of the entire image, and the total energy of the entire image is the sum of the low-frequency energy and the high-frequency energy.
7. The method according to claim 1, characterized in that, The model parameters of the visual model are reconstructed based on the dynamic inference parameters, including at least one of the following reconstruction items: First reconstruction item: Traverse the adjustable layers of the visual model, where the adjustable layers are intermediate encoding layers of the visual model; determine the number of consecutively activated layers of the adjustable layers according to the dynamic inference parameters, so as to reconstruct the model depth of the visual model; The second reconstruction item is to obtain the label scoring index and to score the value of the corresponding image serialization label of the visual model according to the label scoring index. The label scoring index includes at least one of the following: fusion attention weight, edge strength, texture variance, and position prior index. The number of image serialization labels of the visual model is progressively cropped according to the value scoring results to reconstruct the model width of the visual model. The third reconstruction item: Based on the inference mode and remote sensing data type corresponding to the dynamic inference parameters, set the quantization precision level of the activated elastically adjustable layer to reconstruct the model precision of the visual model.
8. The method according to claim 1, characterized in that, After performing model inference on the remote sensing image using the reconstructed visual model to generate model prediction results, the method further includes: Calculate the classification confidence and location confidence of the model's prediction results; The comprehensive confidence score for a single inference is calculated based on the classification confidence score and the location confidence score, wherein the comprehensive confidence score is a weighted sum of the classification confidence score and the location confidence score. A preset confidence threshold is obtained, wherein the confidence threshold includes a first confidence threshold and a second confidence threshold; the first confidence threshold is greater than the second confidence threshold. When the overall confidence level is greater than or equal to the first confidence threshold, the model prediction result is output; When the overall confidence level is less than the first confidence threshold and greater than or equal to the second confidence threshold, an adapter correction network is used to correct the feature loss of the model prediction result; When the overall confidence level is less than the second confidence threshold, re-inference is performed by temporarily optimizing the dynamic inference parameters.
9. A dynamic prediction system for spaceborne resource constraints and scenario complexity, characterized in that, The system includes: The data acquisition module is used to acquire multi-dimensional resource parameters and remote sensing monitoring data of the spaceborne platform; the remote sensing monitoring data includes any type of remote sensing monitoring image from optical remote sensing data, radar remote sensing data, and spectral remote sensing data. The resource status prediction module is used to predict leading status parameters based on the multidimensional resource parameters using a status prediction model; the status prediction model includes a linear regression model and a time series prediction model. The scene detection module is used to extract scene association features from the remote sensing monitoring image and calculate a scene complexity score based on the scene association features. The scene association features include edge density, texture entropy, target density, and frequency domain complexity. The scene complexity score is a weighted calculation result of the scene association features based on adaptive weights. The adaptive weights are generated online by fine-tuning the spatial dispersion of various scene association features in a single remote sensing monitoring image, based on the preset basic weights corresponding to the remote sensing monitoring data type. The inference parameter determination module is used to determine dynamic inference parameters based on the advanced state parameters and the scene complexity score, including: obtaining a preset spaceborne resource security threshold; determining a target inference mode according to the advanced state parameters and the spaceborne resource security threshold; wherein, the target inference mode is obtained by interval matching based on the advanced state parameters and the scene complexity score; when any resource indicator in the advanced state parameters triggers the spaceborne resource security threshold, the target inference mode is determined to be an extremely lightweight inference mode; when none of the resource indicators in the advanced state parameters trigger the spaceborne resource security threshold, the target inference mode is determined based on the scene complexity score; obtaining the baseline inference parameters corresponding to the target inference mode; determining correction parameters according to independent correction rules, wherein the independent correction rules include remote sensing data type, task priority, and resource surplus; the correction parameters include type correction parameters, priority correction parameters, and dynamic margin correction parameters; adjusting the baseline inference parameters according to the correction parameters to obtain the dynamic inference parameters, wherein the dynamic inference parameters are dynamic adaptation parameters obtained by quantifying and fine-tuning the baseline inference parameters under the target inference mode according to the independent correction rules. The model reconstruction module is used to reconstruct the model parameters of the visual model based on the dynamic inference parameters. The model parameters of the visual model include model depth, model width, and model accuracy. The model inference module is used to perform model inference on the remote sensing monitoring image using the reconstructed visual model to generate model prediction results.
Citation Information
Patent Citations
Satellite-borne resource scheduling method under multi-task cooperation scene
CN120336037A
Scene and task dual-conditioned visual hybrid expert model construction and reasoning method
CN121811180A