Space-time remote sensing large model driven urban space intelligent monitoring method and system

CN122493318APending Publication Date: 2026-07-31BEIJING AITERAS INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING AITERAS INFORMATION TECH CO LTD
Filing Date
2026-06-26
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

常规方案因缺失地下管网拓扑边界条件与力学平衡偏微分方程的物理约束,无法将地表语义分割特征与地表形变梯度特征反向映射并解算为地下空间的应力分布状态,输出的监测结果仅停留在地表表象,无法识别地下空间隐患引发的早期地表异常,导致城市空间监测停留在地表二维层面

Benefits of technology

1、本发明通过在空天遥感大模型中引入嵌入城市地下空间拓扑约束的物理信息神经网络,将城市地下空间拓扑结构作为边界条件,以力学平衡偏微分方程作为物理约束,把多模态视觉编码器提取的地表语义分割特征与地表形变梯度特征映射为地下空间应力分布特征,打破了常规技术局限于地表二维像素层面的监测瓶颈,实现了地表语义分割特征与地表形变梯度特征向地下物理力学状态的跨维度耦合映射,使模型能够基于物理规律识别地下空间隐患引发的早期地表异常。同时,利用图注意力网络构建地表监测目标与地下空间节点的拓扑关联图,对不同节点的特征权重进行动态更新,明确了地表形变异常与地下隐患节点之间的拓扑传导逻辑。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493318A_ABST
    Figure CN122493318A_ABST
Patent Text Reader

Abstract

This invention relates to the field of image data technology and discloses a method and system for intelligent urban spatial monitoring driven by a large-scale aerospace remote sensing model. The method includes: receiving multi-source aerospace image data composed of optical remote sensing images, hyperspectral remote sensing images, and temporal differential InSAR deformation images; completing pixel-level feature alignment and spatiotemporal registration of the multi-source data using a multimodal visual encoder of a large-scale aerospace remote sensing model, and extracting surface semantic segmentation features and surface deformation gradient features; embedding the features into a physical information neural network of urban underground space topological constraints, using the urban underground space topology as boundary conditions and mechanical equilibrium partial differential equations as physical constraints, and mapping the fused features of surface semantic segmentation features and surface deformation gradient features to underground space stress distribution features. This invention achieves cross-dimensional coupling mapping of fused features to underground physical and mechanical states, improving the accuracy and depth of urban spatial monitoring, and enabling early identification of surface anomalies caused by underground hazards.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image data technology, specifically to a method and system for intelligent urban spatial monitoring driven by a large-scale aerospace remote sensing model. Background Technology

[0002] Current urban spatial monitoring schemes using aerospace remote sensing primarily rely on multi-source image fusion and deep learning networks for surface feature extraction. Conventional methods, after receiving optical remote sensing images, hyperspectral remote sensing images, and temporal differential interferometric synthetic aperture radar (TIA) deformation images, use a neural network encoder to perform pixel-level feature stitching and spatial coordinate alignment on these multi-source data. After feature stitching, conventional schemes use convolutional layers to extract two-dimensional spatial semantic features of surface buildings, road networks, vegetation, and water bodies, and combine this with radar imagery to extract the vertical deformation gradient of the surface. In subsequent processing, conventional schemes directly input the stitched surface semantic segmentation features and surface deformation gradient features into a decoder, outputting a semantic segmentation map of the urban surface space and the location results of deformation areas. The entire processing flow relies entirely on the statistical distribution patterns of image data in a two-dimensional pixel plane; the network model is driven by a large amount of labeled data to complete the classification mapping from pixels to surface semantics.

[0003] The aforementioned technical solutions, when handling urban surface anomaly monitoring, fail to establish a coupled mapping relationship between surface deformation features and underground space structures because they only perform visual feature stitching and statistical classification on a two-dimensional pixel plane of the surface, without incorporating the topological structure and physical and mechanical constraints of underground space. When slight surface deformation anomalies occur, these anomalies are often caused by stress changes propagating upwards due to underground pipe network leakage or underground space excavation. Conventional solutions, lacking the physical constraints of underground pipe network topological boundary conditions and mechanical equilibrium partial differential equations, cannot reverse-map surface semantic segmentation features and surface deformation gradient features to calculate the stress distribution state of underground space. The output monitoring results remain only at the surface level, failing to identify early surface anomalies caused by potential underground space hazards, thus limiting urban space monitoring to a two-dimensional surface level. Summary of the Invention

[0004] The purpose of this invention is to provide a method and system for intelligent urban spatial monitoring driven by a large-scale aerospace remote sensing model, which can effectively solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the technical solution adopted by this invention is: a large-scale urban spatial intelligent monitoring method driven by aerospace remote sensing models, comprising: Receive multi-source aerospace image data, which includes optical remote sensing images, hyperspectral remote sensing images, and temporal differential InSAR deformation images; The multi-modal visual encoder in the large multi-modal model of aerospace remote sensing performs pixel-level feature alignment and spatiotemporal registration on the multi-source aerospace image data, and extracts surface semantic segmentation features and surface deformation gradient features. The surface semantic segmentation features and surface deformation gradient features are input into a physical information neural network that is embedded with the topological constraints of urban underground space. The topological structure of urban underground space is used as the boundary condition and the mechanical equilibrium partial differential equation is used as the physical constraint to map the surface semantic segmentation features and surface deformation gradient features into the stress distribution features of underground space. A topological association graph between surface monitoring targets and underground space nodes is constructed using a graph attention network. The feature weights of different nodes in the topological association graph are dynamically updated to identify the association between deformation anomalies and underground hidden dangers. The decoding end of the aforementioned aerospace remote sensing multimodal large model outputs visualized monitoring results, including semantic segmentation results, abnormal area location, and hazard level classification.

[0006] Preferably, the pixel-level feature alignment and spatiotemporal registration include: A spatiotemporal registration reinforcement learning framework based on Markov decision process is constructed, and the minimization of multimodal image feature offset is used as the reward function; By using reinforcement learning agents to adaptively adjust the registration window size and spatial sampling step size of optical remote sensing images, hyperspectral remote sensing images, and temporal differential InSAR deformation images at different time scales, the spatiotemporal heterogeneity of multi-source aerospace image data can be eliminated.

[0007] Preferably, the features extracted from hyperspectral remote sensing images include: Hyperspectral remote sensing images are converted to a two-dimensional frequency domain space, and hyperspectral endmember features are extracted using a variational autoencoder. Physicochemical inversion of hyperspectral endmember features using the radiative transfer equation was performed to obtain the physicochemical parameters of the surface material. The physical and chemical parameters of the surface material are used as additional feature channels, and after being stitched together with the registered optical remote sensing image and temporal differential InSAR deformation image features, they are injected into the multimodal visual encoder.

[0008] Preferably, the physical constraints are constructed in the following ways: Breaking through the single solid mechanics framework, it introduces fluid dynamics and thermodynamic constraints for underground pipe networks; A multiphysics loss function is constructed by combining the partial differential equations of the seepage field and the partial differential equations of heat conduction in the underground pipe network. During the forward propagation process of mapping surface semantic segmentation features and surface deformation gradient features to underground space stress distribution features in the physical information neural network, the distribution features of fluid thermodynamic anomalies caused by underground pipe network leakage are simultaneously calculated.

[0009] Preferably, the physical information neural network that embeds feature inputs into the topological constraints of urban underground space includes: A horizontal federated learning framework was used to process underground space data across jurisdictions. In the forward propagation of the physical information neural network, homomorphic encryption technology is used to align the surface semantic segmentation features and surface deformation gradient features with the encrypted underground space topological boundary conditions of different jurisdictions. While maintaining the accuracy of solving mechanical equilibrium partial differential equations and multiphysics loss functions, it achieves safe fusion and joint mapping of cross-domain features.

[0010] Preferably, constructing a topological association graph using a graph attention network includes: The message passing mechanism of graph attention networks is improved by introducing the SIR infectious disease dynamics model from complex network theory. The stress evolution process of underground hidden dangers is simulated as an infection state, and the abnormal nodes of surface deformation are simulated as a susceptible state. The cross-layer propagation probability of potential risks in the topological association graph between surface monitoring targets and underground space nodes is calculated using the SIR infectious disease dynamics model, and the cross-layer propagation probability is used as a priori constraint on the edge weights in the graph attention network.

[0011] Preferably, identifying the correlation between deformation anomalies and underground hazards includes: A network for urban spatial risk propagation is constructed based on the identified hazard relationships and cross-layer propagation probabilities. The NSGA-II multi-objective genetic algorithm is used to solve the optimal dynamic sensor deployment path and edge computing power scheduling scheme in the risk propagation network. The optimal sensor dynamic deployment path and edge computing power scheduling scheme are transformed into vector representations, which are then used as context feature embeddings for the next round of dynamic feature weight updates in the graph attention network.

[0012] Preferably, dynamically updating the feature weights of different nodes in the topological association graph includes: After the graph attention network completes the node feature weight update, the Shapley value from game theory is introduced for feature attribution analysis and uncertainty quantification. The Shapley contribution of different surface semantic segmentation features, surface deformation gradient features, and underground physical constraint features to the final anomaly correlation identification results was calculated using Monte Carlo tree search. Based on Shapley contribution, all features are ranked according to their game disadvantages, and the disadvantaged features at the bottom of the ranking are removed to reduce the cognitive uncertainty of the output of the large multimodal aerospace remote sensing model.

[0013] Preferably, the visualized monitoring results output through the decoding end of the aforementioned space-air remote sensing multimodal large model include: Before the output at the decoding end, a counterfactual causal graph model is constructed. The spurious correlation between meteorological change confounding variables and surface deformation anomalies is removed by do-calculus, and purified causal features are extracted. The purified causal features and the node features after reducing cognitive uncertainty are mapped to the urban digital twin manifold space; In the urban digital twin manifold space, a three-dimensional visualization monitoring map is generated and output by combining deterministic constraints of engineering mechanics.

[0014] This invention provides an intelligent urban spatial monitoring system driven by a large-scale aerospace remote sensing model, comprising: A multi-source sensing data receiving component is used to receive multi-source aerospace image data, which includes optical remote sensing images, hyperspectral remote sensing images, and time-series differential InSAR deformation images. The feature alignment and extraction component is used to perform pixel-level feature alignment and spatiotemporal registration on the multi-source aerospace image data through the multimodal visual encoder in the aerospace remote sensing multimodal large model, and to extract surface semantic segmentation features and surface deformation gradient features. The physical and topology mapping component is used to input the surface semantic segmentation features and surface deformation gradient features into a physical information neural network that is embedded with the topological constraints of urban underground space. The surface semantic segmentation features and surface deformation gradient features are mapped to the stress distribution features of underground space using the topological structure of urban underground space as boundary conditions and the mechanical equilibrium partial differential equation as physical constraints. The graph association and anomaly identification component is used to construct a topological association graph between surface monitoring targets and underground space nodes through a graph attention network, dynamically update the feature weights of different nodes in the topological association graph, and identify the association between deformation anomalies and underground hidden dangers. The visualization decoding output component is used to output visualized monitoring results, including semantic segmentation results, abnormal area location, and hazard level classification, through the decoding end of the aforementioned space-air remote sensing multimodal large model.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention introduces a physical information neural network embedded with topological constraints of urban underground space into a large-scale aerospace remote sensing model. Using the topological structure of urban underground space as boundary conditions and mechanical equilibrium partial differential equations as physical constraints, it maps surface semantic segmentation features and surface deformation gradient features extracted by a multimodal visual encoder to stress distribution features in underground space. This breaks through the monitoring bottleneck of conventional technologies limited to the two-dimensional pixel level of the surface, achieving cross-dimensional coupling mapping of surface semantic segmentation features and surface deformation gradient features to the underground physical and mechanical state. This enables the model to identify early surface anomalies caused by underground space hazards based on physical laws. Simultaneously, a graph attention network is used to construct a topological association graph between surface monitoring targets and underground space nodes, dynamically updating the feature weights of different nodes, thus clarifying the topological transmission logic between surface deformation anomalies and underground hazard nodes.

[0016] 2. This invention eliminates the spatiotemporal heterogeneity of multi-source aerospace image data by constructing a spatiotemporal registration reinforcement learning framework based on Markov decision processes to adaptively adjust the registration window and sampling step size; it constructs a multi-physics loss function by introducing fluid dynamics and thermodynamic constraints of underground pipe networks, enabling the physical information neural network to synchronously solve the distribution characteristics of fluid thermodynamic anomalies caused by underground pipe network leakage; it aligns the topological boundary conditions of underground space across jurisdictional areas using a horizontal federated learning framework combined with homomorphic encryption technology, ensuring data security during the cross-domain feature joint mapping process; it improves the message passing mechanism of graph attention networks by introducing an infectious disease dynamics susceptible infection recovery model, using cross-layer propagation probability as the edge weight prior, thereby enhancing the physical rationality of cross-layer correlation identification of hidden dangers; it reduces the cognitive uncertainty of model output by performing feature attribution and game disadvantage ranking through Shapley value and Monte Carlo tree search; and it improves the certainty of visualized monitoring results under engineering mechanics constraints by combining operator calculations to remove the pseudo-correlation effects of meteorological confounding variables, extracting and purifying causal features, and mapping them to the urban digital twin manifold space to generate monitoring maps. Attached Figure Description

[0017] Figure 1 A flowchart of the overall process for intelligent urban spatial monitoring driven by a large-scale aerospace remote sensing model.

[0018] Figure 2 A flowchart for reinforcement learning for pixel-level feature alignment and spatiotemporal registration.

[0019] Figure 3 This is a flowchart of hyperspectral feature extraction and physicochemical parameter inversion.

[0020] Figure 4 This is a flowchart of multiphysics constraints and cross-domain federated learning mapping.

[0021] Figure 5 This is a flowchart of topology association and anomaly identification based on SIR-GAT.

[0022] Figure 6 A flowchart for causal inference and digital twin visualization. Detailed Implementation

[0023] The urban spatial intelligent monitoring system driven by a large-scale aerospace remote sensing model described in this invention includes a multi-source sensing data receiving component, a feature alignment and extraction component, a physical and topology mapping component, a graph association and anomaly identification component, and a visualization decoding output component. The multi-source sensing data receiving component is configured to receive multi-source aerospace image data, and its hardware carrier is the communication interface of a ground satellite receiving station, the Ethernet interface of an edge computing gateway, or the PCIe data interface of a cloud server. The feature alignment and extraction component, the physical and topology mapping component, the graph association and anomaly identification component, and the visualization decoding output component are all coupled to the multi-source sensing data receiving component, and their hardware carriers are the AI ​​acceleration chip of a cloud server, the FPGA of an edge computing node, or the CPU / GPU co-processing unit of a high-performance computing cluster. All components achieve data interaction and command transmission through a high-speed bus.

[0024] Please refer to Figure 1 In a preferred embodiment, the multi-source sensing data receiving component receives multi-source aerospace image data, which includes optical remote sensing images, hyperspectral remote sensing images, and temporal differential InSAR deformation images. Specifically, the optical remote sensing images originate from sub-meter-level high-resolution optical satellites, with a spatial resolution of 0.5m-2m, covering visible and near-infrared data of the entire urban area; the hyperspectral remote sensing images originate from hyperspectral satellites, with a spectral resolution of 5nm-10nm and 100-300 bands, covering spectral reflectance characteristics of the surface material; and the temporal differential InSAR deformation images originate from C-band temporal interferometric data from synthetic aperture radar satellites, which are processed through differential interferometry to obtain vertical deformation temporal data, with deformation measurement accuracy at the millimeter level and a time resolution of 6-12 days.

[0025] Please refer to Figure 2The feature alignment and extraction component performs pixel-level feature alignment and spatiotemporal registration on the multi-source aerospace image data through a multimodal visual encoder in the aerospace remote sensing multimodal large model, extracting surface semantic segmentation features and surface deformation gradient features. The aerospace remote sensing multimodal large model adopts a Transformer architecture. The multimodal visual encoder contains three parallel feature extraction branches, corresponding to optical remote sensing images, hyperspectral remote sensing images, and temporal differential InSAR deformation images, respectively. Each branch sequentially sets a convolutional embedding layer, a position coding layer, and a multi-head self-attention layer, with a cross-modal attention fusion layer at the end of each branch. Specifically, the pixel-level feature alignment and spatiotemporal registration involves uniformly mapping the multi-source aerospace image data to the CGCS2000 national geodetic coordinate system, unifying spatial resolution and temporal reference, and achieving rigid registration of multimodal images by maximizing mutual information, thus eliminating spatial coordinate deviations.

[0026] The surface semantic segmentation features are extracted through the feature extraction branch corresponding to the optical remote sensing image. These are pixel-by-pixel semantic classification features with dimensions H×W×C1, where H and W are the image height and width, and C1 is the number of semantic categories. The semantic categories include buildings, roads, pipeline ancillary facilities, vegetation, water bodies, and bare land. The surface deformation gradient features are extracted through the feature extraction branch corresponding to the temporal difference InSAR deformation image. These are the spatial gradient tensors of the deformation field, calculated using the central difference method. The expression for the deformation gradient tensor G is: , Where d represents the vertical deformation of the corresponding pixel in the temporal differential InSAR image. Let be the partial derivative of deformation in the x-direction. Let be the partial derivative of deformation in the y-direction. For the vertical temporal deformation partial derivative, , This represents the ground spatial resolution corresponding to each pixel. The extracted surface semantic segmentation features and surface deformation gradient features are concatenated along the channel dimension and then input into subsequent processing steps.

[0027] Please refer to Figure 4The physical and topological mapping component inputs the surface semantic segmentation features and surface deformation gradient features into a physical information neural network that embeds them into the topological constraints of urban underground space. Using the urban underground space topology as boundary conditions and mechanical equilibrium partial differential equations as physical constraints, it maps the surface semantic segmentation features and surface deformation gradient features into underground space stress distribution features. The physical information neural network is a fully connected neural network architecture, containing an input layer, eight hidden layers, and an output layer. Each hidden layer has 256 neurons, and the activation function is the Swish function. The input layer receives the concatenated surface feature vector with dimensions H×W×(C1+3). The output layer outputs the stress distribution features of the discrete grid nodes in the underground space.

[0028] The urban underground space topology is derived from urban underground pipeline survey data and underground space development and utilization planning data, including the pipe diameter, burial depth, material, and topological connection relationships of underground pipelines, as well as the spatial outline, burial depth, and support structure parameters of underground parking garages, civil defense projects, and subway tunnels. The urban underground space topology is discretized into three-dimensional mesh nodes, which serve as displacement boundary conditions for solving the physical information neural network. The boundary condition expression is as follows: ,in The boundary of the underground space topology. The displacement constraint values ​​for the boundary nodes are obtained by mapping the surface deformation characteristics.

[0029] The partial differential equations of mechanical equilibrium are linear elasticity equilibrium equations, expressed as follows: , in, For Cauchy stress tensor, This is a volume force vector, including gravity and groundwater pressure. Here, we have the divergence operator. For isotropic linear elastic materials, the stress-strain relationship satisfies the generalized Hooke's law: , in, , Lamé constant, derived from the elastic modulus of the subsurface medium. Compared to Poisson Calculations show that , ; For strain tensor, , It is a displacement vector; The trace of the strain tensor It is a unit tensor.

[0030] The total loss function of the physical information neural network includes data loss and physical loss, and its expression is: , in, For data loss, denoted as , and for mean square error between the model-predicted surface displacement and the InSAR-measured displacement. For physical losses, the residual mean square error of the partial differential equations of mechanical equilibrium is denoted as ; The weighting coefficients range from 0.1 to 10. The physical information neural network is trained by minimizing the total loss function. During forward propagation, it simultaneously maps surface semantic segmentation features and surface deformation gradient features to subsurface spatial stress distribution features. The output subsurface spatial stress distribution features are the stress tensor components of the discrete grid nodes in the subsurface space, including normal stress. , , With shear stress , , The dimension is N×6, where N is the total number of nodes in the discrete grid of the underground space.

[0031] Please refer to Figure 5 The graph association and anomaly identification component constructs a topological association graph between surface monitoring targets and underground space nodes using a graph attention network. It dynamically updates the feature weights of different nodes in the topological association graph to identify the correlation between deformation anomalies and underground hazards. The topological association graph is represented as follows: , where the node set Including surface monitoring target nodes With underground space nodes The surface monitoring target node The nodes are the center points of buildings, pipeline facilities, and deformation anomaly areas obtained from surface semantic segmentation. Each node's features are its corresponding surface semantic segmentation features and deformation gradient features; the underground space nodes... The nodes represent discrete grid nodes of the underground space topology, with each node characterized by its corresponding stress distribution. (Edge set) It includes adjacent edges of nodes on the same layer and spatial mapping edges of nodes across layers. The adjacent edges of nodes on the same layer include spatial adjacent edges between surface nodes and topological adjacent edges between underground nodes. The spatial mapping edges of nodes across layers are the connection edges between surface nodes and underground nodes within the corresponding vertical projection range.

[0032] The graph attention network adopts a multi-head attention architecture, dynamically updating the feature weights of nodes in the topological graph. The expression for calculating the attention coefficient is as follows: , in, , For nodes , The input feature vector, The weight matrix is ​​a learnable linear transformation. For learnable attention vectors, For feature splicing operations, The activation function is set to a negative slope of 0.2. The attention coefficients are normalized using the following expression: , in, For nodes The set of adjacent nodes. Node features are updated using a multi-head attention mechanism, and the output expression is: , in, For the number of attention heads, For activation function, For the first Normalized attention coefficients for each attention head. For the first Linear transformation weight matrix for each attention head.

[0033] The identification of the correlation between deformation anomalies and underground hazards specifically involves calculating the anomaly score of underground nodes based on the updated node features. The anomaly score is the product of the deviation between the node stress features and the preset normal stress threshold, and the attention weight of the corresponding surface deformation anomaly node. Underground nodes with anomaly scores exceeding the preset threshold are identified as underground hazard nodes, and the correlation between the hazard node and the corresponding surface deformation anomaly node is output simultaneously, including the correlation weight and stress transmission path.

[0034] The visualization decoding output component outputs visualized monitoring results, including semantic segmentation results, anomaly area location, and hazard level classification, through the decoding end of the aerospace remote sensing multimodal large model. The decoding end of the aerospace remote sensing multimodal large model adopts a Transformer decoder architecture, sequentially setting a cross-attention layer, a feedforward network layer, and an output layer. The input to the decoding end includes multimodal features extracted by the multimodal visual encoder, underground stress distribution features output by the physical information neural network, and node features updated by the graph attention network. The output of the decoding end includes pixel-by-pixel semantic segmentation masks, spatial location boxes of anomaly areas, and underground hazard level classification results, with the hazard levels divided into four levels according to risk severity. The visualization decoding output component couples the urban geographic information system platform with the display terminal, mapping the output monitoring results onto the urban 3D geographic model for visualization.

[0035] In a preferred embodiment, the feature alignment and extraction component constructs a spatiotemporal registration reinforcement learning framework based on Markov decision processes, using the minimization of multimodal image feature offsets as the reward function; and adaptively adjusts the registration window size and spatial sampling step size of optical remote sensing images, hyperspectral remote sensing images, and temporal differential InSAR deformation images at different time scales through a reinforcement learning agent, so as to eliminate the spatiotemporal heterogeneity of multi-source aerospace image data.

[0036] The Markov decision process is a quintuple. ,in: state space The current registration state of the multi-source images, and its state vector. ,in , This refers to the spatial offset between multimodal images. This is the time offset. To register the window size, For spatial sampling step size, This represents the mutual information value of the multimodal images under the current registration state; Action space The set of actions that a reinforcement learning agent can perform includes adjusting the size of the registration window, adjusting the spatial sampling step size, correcting the spatial offset, and correcting the temporal offset. The adjustment step size for each action is a discrete value. State transition probability : , indicating the state Next action Then transferred to state The probability of a transition is a deterministic one, determined by the result of the registration operation. reward function The optimization objective is to minimize the feature offset of the multimodal image, expressed as: , in, The L2 norm of the spatial offset It is the absolute value of the time offset. This represents the theoretical maximum mutual information value of the multimodal image. , , These are weighting coefficients, with values ​​ranging from 0.1 to 10; Discount factor The value ranges from 0.95 to 0.99, and is used to balance immediate rewards and long-term cumulative rewards.

[0037] The reinforcement learning agent employs a deep Q-network architecture, comprising an input layer, three hidden layers, and an output layer. The input is a state vector. The output is the Q-value corresponding to each action, and the training objective of the agent is to maximize the expected cumulative reward. The agent adaptively adjusts the registration parameters according to the land cover type and deformation characteristics of the registration area: for urban built-up areas with drastic deformation changes, the registration window size and spatial sampling step size are reduced to improve registration accuracy; for suburban areas with high vegetation coverage, the registration window size and spatial sampling step size are increased to improve registration efficiency, thereby eliminating the spatiotemporal heterogeneity of multi-source aerospace image data caused by differences in spatial resolution, temporal sampling frequency, and imaging modality.

[0038] As a preferred embodiment, please refer to Figure 3 When the feature alignment and extraction component extracts features from the hyperspectral remote sensing image, it first converts the hyperspectral remote sensing image to a two-dimensional frequency domain space and uses a variational autoencoder to extract hyperspectral endmember features. Then, it performs physicochemical inversion on the hyperspectral endmember features using the radiative transfer equation to obtain the physicochemical parameters of the land surface material. The physicochemical parameters of the land surface material are used as additional feature channels and stitched together with the registered optical remote sensing image and the temporal differential InSAR deformation image features before being injected into the multimodal visual encoder.

[0039] The three-dimensional data of the hyperspectral remote sensing image is represented as follows: ,in , These are spatial pixel coordinates. For each spatial pixel, a two-dimensional Fourier transform is performed on the spectral vector to convert it to the frequency domain. The transform expression is: , in, , Using frequency domain coordinates, Fourier transform is used to remove strip noise and atmospheric scattering noise from hyperspectral images, while preserving effective low-frequency spectral feature components.

[0040] The variational autoencoder includes an encoder and a decoder. The encoder maps the denoised hyperspectral pixel vectors to a Gaussian distribution in the latent space and outputs the mean value. With variance The decoder samples from the latent space. The hyperspectral pixel vector is reconstructed. The loss function expression of the variational autoencoder is: , in, To reconstruct the mean square error between the original vector and the original vector, Let KL divergence be the KL divergence. The weighting coefficients range from 0.1 to 1; the sampled vector output from the encoder's latent space. These are hyperspectral endmember features, corresponding to the spectral characteristics of different materials on the Earth's surface.

[0041] The radiative transfer equation is used to describe the radiative transfer process of the atmosphere-land surface system, and its simplified expression is: , in, The radiance at the sensor entrance pupil. Atmospheric path radiation, This represents the true reflectivity of the surface material. This refers to the downward solar irradiance. Atmospheric transmittance. The true reflectance of the surface material is obtained by inversion through the radiative transfer equation. Based on the spectral characteristics of the reflectance, the physicochemical parameters of the surface material are obtained by inversion, including soil moisture content, organic matter content, concrete weathering degree, porosity, and asphalt aging degree. The physicochemical parameters are calculated by a preset spectral characteristic inversion model, which is a linear regression model based on the material's physicochemical parameters and the absorption depth of the characteristic band.

[0042] The physical and chemical parameters of the surface material are encoded into a feature map with dimensions H×W×C2, where C2 is the number of physical and chemical parameters. This feature map is then stitched together with the registered optical remote sensing image features and temporal difference InSAR deformation image features in the channel dimension. The resulting feature vector has dimensions H×W×(C1+3+C2) and is input into the cross-modal attention fusion layer of the multimodal visual encoder to complete multimodal feature fusion.

[0043] As a preferred embodiment, when constructing physical constraints, the physical and topological mapping component breaks through the single solid mechanics framework and introduces fluid dynamics and thermodynamic constraints of underground pipe networks; it jointly constructs a multi-physics loss function by combining the partial differential equations of seepage field and heat conduction partial differential equations of underground pipe networks; and it simultaneously solves the abnormal distribution characteristics of fluid thermodynamics caused by leakage in underground pipe networks during the forward propagation process from physical information neural network mapping surface semantic segmentation features and surface deformation gradient features to underground space stress distribution features.

[0044] The partial differential equation of the seepage field of the underground pipe network is the saturated-unsaturated seepage control equation based on Darcy's law, and its expression is: , in, The water storage capacity of the underground medium. For pore water pressure head, For time, Let be the saturated permeability coefficient of the medium. The relative permeability coefficient is determined by the matrix suction. For source and sink items, the corresponding amount of water leaking from the pipeline network.

[0045] The partial differential equation for heat conduction is the governing equation for heat conduction considering convective heat transfer, and its expression is: , in, The density of the underground medium, The specific heat capacity of the underground medium. For the temperature of the medium, The thermal conductivity coefficient of the underground medium. The density of water, The specific heat capacity of water, Let be the seepage velocity, according to Darcy's law. Calculations show that This is the heat source, corresponding to the heat input caused by the temperature difference between the leaking water in the pipeline and the underground medium.

[0046] The expression for the multiphysics loss function is: , in, Data loss includes the mean square error between measured and model-predicted surface displacement values, the mean square error between monitored and model-predicted groundwater levels, and the mean square error between monitored and model-predicted groundwater temperatures. This represents the residual loss in the solid mechanics equilibrium equations; This represents the residual loss of the control equations for the seepage field. This represents the residual loss in the heat conduction control equation; The boundary condition loss includes the residual mean square error of the displacement boundary, head boundary, and temperature boundary of the underground pipe network topology. , , , These are the weighting coefficients for each loss term, with values ​​ranging from 0.01 to 100.

[0047] The output layer of the physical information neural network synchronously outputs the stress distribution characteristics, pore water pressure head distribution characteristics, and temperature distribution characteristics of the underground space. Through multi-physics field coupling calculation, it identifies seepage anomalies and temperature anomalies caused by pipeline leakage, as well as the resulting deterioration of soil mechanical parameters and stress redistribution, thereby achieving early identification of potential underground pipeline leakage hazards.

[0048] In a preferred embodiment, when the physical and topological mapping component embeds the feature input into the physical information neural network of urban underground space topological constraints, it uses a lateral federated learning framework to process underground space data across jurisdictional areas. During the forward propagation of the physical information neural network, homomorphic encryption technology is used to align surface semantic segmentation features and surface deformation gradient features with the encrypted underground space topological boundary conditions of different jurisdictional areas. While maintaining the accuracy of solving the mechanical equilibrium partial differential equations and multiphysics loss functions, it achieves secure fusion and joint mapping of cross-domain features.

[0049] The participants in the horizontal federated learning framework include a federated coordinator and multiple data holders. Each data holder corresponds to a city jurisdiction and possesses multi-source aerospace imagery, underground space topology data, and underground monitoring data for that region. The data holders do not directly share the raw data, but only share encrypted model parameters and features. The homomorphic encryption technology adopts the CKKS fully homomorphic encryption scheme, supporting homomorphic addition and multiplication operations of floating-point numbers, meeting the encrypted computation requirements during the forward propagation process of the neural network. The specific processing flow is as follows: Initialization phase: The federal coordinator initializes the global model parameters of the physical information neural network, and distributes the global model parameters to each data holder after homomorphic encryption; each data holder generates a local public key and private key, sends the public key to the federal coordinator, and keeps the private key locally. Local training phase: Each data holder receives the encrypted global model parameters, uses local data to complete the forward and backward propagation of the model, and calculates the local model gradient; during the forward propagation process, the data holder homomorphically encrypts the surface semantic segmentation features and surface deformation gradient features of the local area and sends them to other participants. Each participant completes the feature alignment operation between the encrypted surface features and the local encrypted underground space topological boundary conditions in the encrypted state, including feature matching, coordinate alignment and boundary condition mapping, without leaking the original data throughout the operation; Gradient aggregation phase: Each data holder sends the locally computed encrypted model gradient to the federated coordinator. The federated coordinator performs weighted average aggregation of all gradients in encrypted state to obtain the global encrypted gradient. Model update phase: The federation coordinator sends the aggregated global encrypted gradient to each data holder. Each data holder decrypts the gradient using its local private key, updates its local model parameters, and simultaneously sends the updated encrypted model parameters to the federation coordinator to complete the global model update. Iterative convergence: Repeat the local training, gradient aggregation, and model update steps until the change in the total loss function of the global model is less than a preset threshold, thus completing model convergence.

[0050] The federal coordinator is deployed in the trusted execution environment of the municipal government cloud, while the local processing modules of each data holder are deployed in the corresponding regional government intranet servers to ensure the security and compliance of cross-domain data processing.

[0051] In a preferred embodiment, when the graph association and anomaly identification component constructs a topological association graph through a graph attention network, it introduces the SIR infectious disease dynamics model from complex network theory to improve the message passing mechanism of the graph attention network; it simulates the stress evolution process of underground hazards as an infection state and simulates abnormal surface deformation nodes as susceptible states; it uses the SIR infectious disease dynamics model to calculate the cross-layer propagation probability of hazards in the topological association graph between the surface monitoring target and underground space nodes, and uses the cross-layer propagation probability as a priori constraint condition for the edge weights in the graph attention network.

[0052] The SIR infectious disease dynamics model categorizes nodes in the topological association graph into three states: susceptible state S, infected state I, and recovering state R. Infected state I corresponds to underground hazard nodes with stress anomalies, and the node state value is positively correlated with the degree of stress anomaly, ranging from 0 to 1. Susceptible state S corresponds to surface monitoring target nodes with deformation anomalies. Recovering state R corresponds to stable nodes without anomalies. The state transition differential equation of the SIR model is: , in, The infection rate is the probability that an infected node will infect a susceptible node within a unit of time. The recovery rate is the probability that an infected node will become a recovered node within a unit of time. , , These represent the percentage of the number of corresponding state nodes, satisfying... .

[0053] The expression for calculating the probability of the hidden danger propagating across layers is: , in, Surface node With underground nodes Topological connection weights between them , For nodes With nodes The Euclidean distance between them The mechanical conductivity coefficient of the underground medium; underground node The infection status value.

[0054] The cross-layer propagation probability As a prior constraint on the edge weights of the graph attention network, the original attention coefficients are modified, and the expression for the modified attention coefficients is as follows: , The corrected attention coefficients are normalized to obtain the final attention weights, which are used for message passing and node feature updates in the graph attention network, so that the attention weights conform to the physical law of the upward transmission of stress in underground hidden dangers.

[0055] In a preferred embodiment, when the graph association and anomaly identification component identifies the association between deformation anomalies and underground hazards, it constructs an urban spatial risk propagation network based on the identified hazard associations and cross-layer propagation probabilities; it uses the NSGA-II multi-objective genetic algorithm to solve for the optimal sensor dynamic deployment path and edge computing power scheduling scheme in the risk propagation network; and it transforms the optimal sensor dynamic deployment path and edge computing power scheduling scheme into a vector representation, which is used as the context feature embedding for the next round of dynamic feature weight update of the graph attention network.

[0056] The urban spatial risk propagation network uses surface monitoring target nodes and underground space nodes as network nodes, and the probability of cross-layer propagation of hazards and node association weights as edge weights to characterize the propagation path and intensity of hazard risks. The NSGA-II multi-objective genetic algorithm optimizes three objective functions, which are: Objective function for minimizing risk monitoring coverage gaps: ,in The number of high-risk nodes to be covered by sensor deployment points. This represents the total number of high-risk nodes. The objective function for minimizing the total lifecycle cost of the sensor is: ,in This represents the total number of sensors deployed. For the first The total lifecycle cost of a sensor; The objective function for minimizing edge computing processing latency is: ,in To calculate the number of nodes at the edge, For the first The monitoring data processing latency of each edge computing node.

[0057] The solution process of the NSGA-II multi-objective genetic algorithm is as follows: Population initialization: The initial population is randomly generated using real number encoding. The encoding content of each individual includes the coordinates of the sensor deployment points, the sensor type, and the computing power allocation ratio of the edge computing nodes. Non-dominated sorting: The individuals in the population are sorted in a non-dominated manner and divided into Pareto levels. Individuals in level 1 are the current Pareto optimal solutions. Crowding calculation: Calculate the crowding of individuals within the same Pareto level to maintain population diversity; Selection operation: A tournament selection strategy is adopted to select individuals with high Pareto rank and high crowding as parent individuals; Crossover mutation: Perform simulated binary crossover and polynomial mutation on the parent individuals to generate the offspring population; Population merging: Merge the parent population and the offspring population, re-perform the non-dominated sorting and crowding calculation, and select the top N individuals as the new population, where N is the population size; Iteration Termination: Repeat the above steps until the preset number of iterations is reached, and output the optimal sensor dynamic deployment path and edge computing power scheduling scheme in the Pareto optimal solution set.

[0058] The optimal output solution is encoded as a fixed-dimensional feature vector and concatenated into the input features of each node in the graph attention network. This serves as the context feature embedding for the next round of dynamic feature weight updates, thereby improving the global correlation of node features and the accuracy of anomaly detection.

[0059] In a preferred embodiment, when the graph association and anomaly identification component dynamically updates the feature weights of different nodes in the topological association graph, after the graph attention network completes the node feature weight update, the Shapley value from game theory is introduced for feature attribution analysis and uncertainty quantification. Monte Carlo tree search is used to calculate the Shapley contribution of different surface semantic segmentation features, surface deformation gradient features, and underground physical constraint features to the final anomaly association identification result. Based on the Shapley contribution, all features are ranked according to their game disadvantage, and the disadvantaged features at the bottom of the ranking are removed to reduce the cognitive uncertainty of the output of the large multimodal model of aerospace remote sensing.

[0060] The Shapley value is used to calculate the marginal contribution of each input feature to the abnormal association identification result, and the calculation formula is as follows: , in, The set of all input features. For the first One characteristic, For features not included Feature subset, To use only a subset of features The recognition accuracy of the time model, i.e., the profit function, Features The Shapley value.

[0061] The Monte Carlo tree search is used to find approximate solutions to the Shapley value, avoiding the computational explosion problem caused by an excessive number of features. The execution process includes four steps: selection, expansion, simulation, and backtracking. The root node is an empty feature set, and each tree node corresponds to a feature subset. Starting from the root node, nodes that have not been fully visited are selected for expansion. The recognition accuracy of the model under the feature subset is simulated and the profit value is calculated. The cumulative profit value of all nodes on the path is updated by backtracking. The approximate solution of the Shapley value for each feature is obtained through multiple iterations.

[0062] Based on the calculated Shapley values, all input features are sorted in ascending order. The smaller the Shapley value, the lower the contribution of the feature to the recognition result, and it is judged as a disadvantageous feature. The disadvantageous features at the bottom of the ranking are removed, and the remaining high-contribution features are input into the subsequent processing stage of the aerospace remote sensing multimodal large model to eliminate the model cognitive uncertainty caused by redundant features.

[0063] As a preferred embodiment, please refer to Figure 6 When the visualization decoding output component outputs the visualization monitoring results through the decoding end of the space-air remote sensing multimodal large model, it constructs a counterfactual causal graph model before outputting the results at the decoding end. It then uses do-calculus to remove the spurious correlation effects of meteorological change confounding variables on surface deformation anomalies and extracts purified causal features. The purified causal features and node features after reducing cognitive uncertainty are mapped to the urban digital twin manifold space. In the urban digital twin manifold space, a three-dimensional visualization monitoring map is generated and output by combining engineering mechanics deterministic constraints.

[0064] The counterfactual causal graph model is represented as follows: , where the node set Including processing variables Outcome variables With confounding variables The processing variable For surface deformation anomalies, the resulting variables For underground hidden dangers, the aforementioned confounding variables Meteorological change factors, including rainfall, temperature changes, and atmospheric deposition, are the confounding variables. Simultaneously, process variables With outcome variable It has a causal effect.

[0065] Processing variables through do-calculus Perform intervention operations Cut off confusing variables To process variables The causal path is determined by adjusting the formula via a backdoor to calculate the causal effect after intervention. The expression is as follows: , in, To obfuscate variables By integrating the values ​​of all confounding variables, the spurious correlation effects caused by the confounding variables are eliminated, and the true causal effect between surface deformation anomalies and underground hidden dangers is obtained, and the corresponding purified causal features are extracted.

[0066] The UMAP manifold learning algorithm maps purified causal features and node features with reduced cognitive uncertainty to the urban digital twin manifold space. The urban digital twin manifold space is constructed based on the urban three-dimensional geographic information model. During the high-dimensional feature mapping process, the topological structure and causal relationship between features are maintained. Each point in the manifold space corresponds to a spatial node on the surface or underground. The location of the point corresponds to three-dimensional geographic coordinates, and the attributes of the point correspond to feature values, anomaly scores, and hazard levels.

[0067] By incorporating deterministic constraints from engineering mechanics, including material strength limits, allowable deformation thresholds, and stress safety factors, a three-dimensional visualization monitoring map is generated in the urban digital twin manifold space. This map includes surface semantic segmentation results, spatial location of abnormal areas, three-dimensional distribution of underground hazards, hazard level classification, and risk propagation paths. Abnormal areas and hazard nodes exceeding safety thresholds are highlighted. This three-dimensional visualization monitoring map can be directly imported into the urban digital twin platform and geographic information system platform for visualization output.

Claims

1. A method for intelligent urban spatial monitoring driven by a large-scale aerospace remote sensing model, characterized in that, include: Receive multi-source aerospace image data, which includes optical remote sensing images, hyperspectral remote sensing images, and temporal differential InSAR deformation images; The multi-modal visual encoder in the large multi-modal model of aerospace remote sensing performs pixel-level feature alignment and spatiotemporal registration on the multi-source aerospace image data, and extracts surface semantic segmentation features and surface deformation gradient features. The surface semantic segmentation features and surface deformation gradient features are input into a physical information neural network that is embedded with the topological constraints of urban underground space. The topological structure of urban underground space is used as the boundary condition and the mechanical equilibrium partial differential equation is used as the physical constraint to map the surface semantic segmentation features and surface deformation gradient features into the stress distribution features of underground space. A topological association graph between surface monitoring targets and underground space nodes is constructed using a graph attention network. The feature weights of different nodes in the topological association graph are dynamically updated to identify the association between deformation anomalies and underground hidden dangers. The decoding end of the aforementioned aerospace remote sensing multimodal large model outputs visualized monitoring results, including semantic segmentation results, abnormal area location, and hazard level classification.

2. The urban spatial intelligent monitoring method driven by a large-scale aerospace remote sensing model according to claim 1, characterized in that, The pixel-level feature alignment and spatiotemporal registration include: A spatiotemporal registration reinforcement learning framework based on Markov decision process is constructed, and the minimization of multimodal image feature offset is used as the reward function; By using reinforcement learning agents to adaptively adjust the registration window size and spatial sampling step size of optical remote sensing images, hyperspectral remote sensing images, and temporal differential InSAR deformation images at different time scales, the spatiotemporal heterogeneity of multi-source aerospace image data can be eliminated.

3. The urban spatial intelligent monitoring method driven by a large-scale aerospace remote sensing model according to claim 2, characterized in that, Features extracted from hyperspectral remote sensing images include: Hyperspectral remote sensing images are converted to a two-dimensional frequency domain space, and hyperspectral endmember features are extracted using a variational autoencoder. Physicochemical inversion of hyperspectral endmember features using the radiative transfer equation was performed to obtain the physicochemical parameters of the surface material. The physical and chemical parameters of the surface material are used as additional feature channels and then stitched together with the registered optical remote sensing image and temporal differential InSAR deformation image features and injected into the multimodal visual encoder.

4. The urban spatial intelligent monitoring method driven by a large-scale aerospace remote sensing model according to claim 3, characterized in that, The physical constraints are constructed in the following ways: Breaking through the single solid mechanics framework, it introduces fluid dynamics and thermodynamic constraints for underground pipe networks; A multiphysics loss function is constructed by combining the partial differential equations of the seepage field and the partial differential equations of heat conduction in the underground pipe network. During the forward propagation process of mapping surface semantic segmentation features and surface deformation gradient features to underground space stress distribution features in the physical information neural network, the distribution features of fluid thermodynamic anomalies caused by underground pipe network leakage are simultaneously calculated.

5. The urban spatial intelligent monitoring method driven by a large-scale aerospace remote sensing model according to claim 4, characterized in that, Physical information neural networks that embed feature inputs into topological constraints of urban underground space include: A horizontal federated learning framework was used to process underground space data across jurisdictions. In the forward propagation of the physical information neural network, homomorphic encryption technology is used to align the surface semantic segmentation features and surface deformation gradient features with the encrypted underground space topological boundary conditions of different jurisdictions. While maintaining the accuracy of solving mechanical equilibrium partial differential equations and multiphysics loss functions, it achieves safe fusion and joint mapping of cross-domain features.

6. The urban spatial intelligent monitoring method driven by a large-scale aerospace remote sensing model according to claim 5, characterized in that, Constructing a topological association graph using graph attention networks includes: The message passing mechanism of graph attention networks is improved by introducing the SIR infectious disease dynamics model from complex network theory. The stress evolution process of underground hidden dangers is simulated as an infection state, and the abnormal nodes of surface deformation are simulated as a susceptible state. The cross-layer propagation probability of potential risks in the topological association graph between surface monitoring targets and underground space nodes is calculated using the SIR infectious disease dynamics model, and the cross-layer propagation probability is used as a priori constraint on the edge weights in the graph attention network.

7. The urban spatial intelligent monitoring method driven by a large-scale aerospace remote sensing model according to claim 6, characterized in that, Identifying the correlation between deformation anomalies and underground hazards includes: A network for urban spatial risk propagation is constructed based on the identified hazard relationships and cross-layer propagation probabilities. The NSGA-II multi-objective genetic algorithm is used to solve the optimal dynamic sensor deployment path and edge computing power scheduling scheme in the risk propagation network. The optimal sensor dynamic deployment path and edge computing power scheduling scheme are transformed into vector representations, which are then used as context feature embeddings for the next round of dynamic feature weight updates in the graph attention network.

8. The urban spatial intelligent monitoring method driven by a large-scale aerospace remote sensing model according to claim 7, characterized in that, Dynamically updating the feature weights of different nodes in the topological association graph includes: After the graph attention network completes the node feature weight update, the Shapley value from game theory is introduced for feature attribution analysis and uncertainty quantification. The Shapley contribution of different surface semantic segmentation features, surface deformation gradient features, and underground physical constraint features to the final anomaly correlation identification results was calculated using Monte Carlo tree search. Based on Shapley contribution, all features are ranked according to their game disadvantages, and the disadvantaged features at the bottom of the ranking are removed to reduce the cognitive uncertainty of the output of the large multimodal aerospace remote sensing model.

9. The urban spatial intelligent monitoring method driven by a large-scale aerospace remote sensing model according to claim 8, characterized in that, The visualized monitoring results output from the decoding end of the aforementioned space-air remote sensing multimodal large model include: Before the output at the decoding end, a counterfactual causal graph model is constructed. The spurious correlation between meteorological change confounding variables and surface deformation anomalies is removed by do-calculus, and purified causal features are extracted. The purified causal features and the node features after reducing cognitive uncertainty are mapped to the urban digital twin manifold space; In the urban digital twin manifold space, a three-dimensional visualization monitoring map is generated and output by combining deterministic constraints of engineering mechanics.

10. A large-scale urban spatial intelligent monitoring system driven by aerospace remote sensing models, characterized in that, include: A multi-source sensing data receiving component is used to receive multi-source aerospace image data, which includes optical remote sensing images, hyperspectral remote sensing images, and time-series differential InSAR deformation images. The feature alignment and extraction component is used to perform pixel-level feature alignment and spatiotemporal registration on the multi-source aerospace image data through the multimodal visual encoder in the aerospace remote sensing multimodal large model, and to extract surface semantic segmentation features and surface deformation gradient features. The physical and topology mapping component is used to input the surface semantic segmentation features and surface deformation gradient features into a physical information neural network that is embedded with the topological constraints of urban underground space. The surface semantic segmentation features and surface deformation gradient features are mapped to the stress distribution features of underground space using the topological structure of urban underground space as boundary conditions and the mechanical equilibrium partial differential equation as physical constraints. The graph association and anomaly identification component is used to construct a topological association graph between surface monitoring targets and underground space nodes through a graph attention network, dynamically update the feature weights of different nodes in the topological association graph, and identify the association between deformation anomalies and underground hidden dangers. The visualization decoding output component is used to output visualized monitoring results, including semantic segmentation results, abnormal area location, and hazard level classification, through the decoding end of the aforementioned space-air remote sensing multimodal large model.