Multimodal road change detection method based on reinforcement learning optimization
Through multimodal data fusion and parameter optimization based on reinforcement learning, the problems of insufficient multimodal data collaborative modeling and dynamic adaptability in road change detection in existing technologies are solved, and high-precision and robust road change detection is achieved.
Patent Information
- Application Number
- CN202511054647.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-09-16
AI Technical Summary
Existing road change detection technologies are weak in the collaborative modeling capabilities of multimodal data and lack the ability to collaboratively analyze three-dimensional spatial structural features, spectral response characteristics, and long-term observation data. In addition, traditional algorithms have poor feature coupling and weak dynamic adaptability in complex scenarios, resulting in low detection accuracy and high false alarm rates.
A multimodal road change detection method based on reinforcement learning is adopted. The natural language prompt is encoded into a semantic vector through the RoBERTa model. Multimodal data fusion and reinforcement learning optimization are combined, including three-dimensional spatial analysis, remote sensing data enhancement, time series data embedding, text-visual feature alignment and gated feature fusion. The improved PPO algorithm is used for parameter optimization to achieve adaptive adjustment of the detection model.
It achieves high-precision and robust road change detection in complex scenarios, improves detection accuracy and adaptability, and meets the requirements of 80% accuracy and 60% intersection-over-union ratio.
Smart Images

Figure CN120656031A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image recognition technology, and specifically relates to a multimodal road change detection method based on reinforcement learning optimization. Background Art
[0002] In recent years, with the rapid development of remote sensing technology, road change detection has played an increasingly important role in smart transportation, urban renewal, and disaster response. However, existing road change detection technologies still have significant flaws: First, the collaborative modeling capabilities of multimodal data are weak. Traditional algorithms often use isolated analysis modes and lack the ability to collaboratively analyze three-dimensional spatial structural features, spectral response characteristics, and long-term observation data. This leads to incomplete feature representation in refined tasks such as road crack identification and roadbed settlement detection. Second, traditional algorithms lack dynamic parameter adjustment mechanisms when faced with complex scenarios such as sudden changes in light intensity, seasonal changes in vegetation cover, and temporary obstructions. This often leads to increased false alarm rates and missed detections of minor changes. Finally, traditional algorithms rely on manual experience for parameter optimization, making it difficult to quickly adapt to different regional characteristics and detection task requirements, seriously restricting the large-scale application of the technology. Summary of the Invention
[0003] The purpose of the present invention is to provide a multimodal road change detection method based on reinforcement learning optimization, which can not only collaboratively process three-dimensional spatial structural characteristics, spectral response characteristics and long-term observation data, but also realize adaptive optimization of detection model parameters through reinforcement learning, effectively solving the problems of poor feature coupling and weak dynamic adaptability of traditional methods in complex scenarios, and providing a new road change detection paradigm with both innovation and practical application value for urban governance and intelligent transportation fields.
[0004] The technical solution adopted by the present invention is a multimodal road change detection method based on reinforcement learning optimization, comprising the following steps:
[0005] S1, natural language prompt input and semantic parsing of design specifications;
[0006] Convert user-entered natural language prompt instructions into structured, computable design constraints;
[0007] S2, text semantic vectorization;
[0008] The RoBERTa pre-trained model is used to build a natural language prompt text encoder, and deep semantic representation of natural language prompts is achieved through word embedding and hierarchical attention mechanism;
[0009] S3, road scene multimodal data input and preprocessing;
[0010] Construct a multimodal data fusion mechanism for road change detection scenarios to implement multimodal data input and preprocessing of road scenes. The specific steps include:
[0011] S31, 3D spatial data analysis;
[0012] S32, remote sensing data enhancement;
[0013] S33, time series data embedding;
[0014] S4, road image feature embedding and multimodal association, the specific steps include:
[0015] S41, three-dimensional space analysis branch;
[0016] S42, remote sensing enhancement coding branch;
[0017] S43, temporal spectral encoding branch;
[0018] S44, unified dimension;
[0019] S5, multimodal feature alignment and dynamic fusion, the specific steps include:
[0020] S51, text guides visual focus;
[0021] S52, visual feedback semantic revision;
[0022] S53, gated feature fusion;
[0023] S6, detection model parameter optimization and online simulation verification;
[0024] Users integrate a large multimodal model with a simulation engine, use processed feature data to build intelligent parameter configuration and verification, and convert the parameter vector output by the reinforcement learning strategy network into a specific parameter configuration scheme for the road change detection method. After semantic verification of the large model, the simulation engine evaluates the performance online.
[0025] S7, multi-objective reinforcement learning optimization and scenario verification;
[0026] Based on the improved PPO algorithm, an uncertainty-aware exploration strategy is used to dynamically adjust the action distribution variance σ σWhen the historical reward fluctuation of road change detection exceeds the set threshold, the exploration intensity is automatically increased to prevent the strategy from falling into the local optimal solution; a curriculum learning mechanism is introduced at the beginning of training to first relax the computing resource constraints to prioritize the exploration of the limits of detection accuracy and response performance; then the resource constraint penalty is gradually increased to push the model to converge to a more practical and efficient scenario; the training process adopts a priority experience replay mechanism, performs weighted sampling of the Pareto frontier solution according to the time difference error, and combines the advantage function and gradient clipping method to stably update the policy network parameters, ensuring the efficient convergence of the algorithm and its adaptability and robustness in real road change detection scenarios.
[0027] Furthermore, in S1, the natural language prompt command input by the user is converted into a structured and computable design constraint, specifically: the user inputs the road change image to be processed and completes the corresponding task requirement description T text "Based on inputting road change images of the same area at different times, the road change detection task is completed with an accuracy rate of no less than 80% and an intersection-over-union (IOU) index of no less than 60%."
[0028] Furthermore, in S2, the RoBERTa pre-trained model is used to construct a natural language prompt text encoder. The specific steps for achieving deep semantic representation of natural language prompts through word embedding and hierarchical attention mechanism are as follows:
[0029] The natural language prompt input by the user is segmented and mapped into a 768-dimensional semantic vector, which is then integrated with the embedding information in the field of position encoding and road change detection. Then, a local-global two-stage attention mechanism is used to capture the local semantic dependencies related to road changes in the natural language prompt in the first stage, and aggregate the overall context information of the natural language prompt in the second stage, ultimately generating a text feature vector T containing rich multimodal constraint information. text ∈R 512 , where R is the road domain relationship matrix.
[0030] Furthermore, in S41, the specific steps of the three-dimensional space analysis branch include:
[0031] S411, constructing geometric coding branch;
[0032] S412, constructing feature encoders for different types of geometric elements respectively, and using a hierarchical message passing mechanism to implement node feature updates;
[0033] S413, constructing point cloud coding branch;
[0034] S414, using the cross-attention mechanism to establish cross-modal associations between geometric features and point cloud features;
[0035] S415, joint representation of three-dimensional features.
[0036] Furthermore, in said S42, the specific steps of the remote sensing enhancement coding branch are:
[0037] The hyperspectral remote sensing image is input and the data enhancement operation is performed using the MixUp data enhancement and channel random dropout methods. Then the spectral attention module and spatial-spectral separation convolution are introduced. Finally, positive and negative samples are constructed to conduct comparative learning on the enhanced features, and the output spectral enhancement feature F is obtained. RS ∈R N×512 , where N is the number of positive and negative sample pairs constructed in contrastive learning.
[0038] Furthermore, in said S43, the specific steps of the time series spectrum encoding branch are:
[0039] Based on the 3D CNN network, the road hyperspectral remote sensing data and time series data fusion tensor are convolved to output the spatiotemporal perception feature F containing time series evolution and spectral differences. phy ∈R 512 .
[0040] Furthermore, in S44, the specific steps of unifying the dimensions are:
[0041] The multimodal features are uniformly expanded to 512 dimensions through the projection layer, and finally a visual feature vector V is formed for road change detection. fusion ∈R 1536 .
[0042] Furthermore, in S51, the specific steps of guiding visual focus by text are:
[0043] Take the text feature vector T text ∈R 512 is the Query vector, the visual feature vector V fusion ∈R 1536 The attention weights between multiple modalities are calculated as Key-Value vectors, the spatial region features closely related to the semantic requirements of natural language prompts are strengthened, and the text features and visual features are cross-fused to obtain the fused multimodal features.
[0044] Furthermore, in S52, the specific steps of visual feedback semantic correction are:
[0045] The text feature vector T is transformed into text ∈R 512 The dimension is expanded to 1536, making it consistent with the visual feature vector V fusion ∈R 1536The dimensions of are aligned, thereby achieving feedback enhancement of the semantic representation of natural language prompts using visual features.
[0046] Furthermore, in S53, the specific steps of gated feature fusion are:
[0047] Dynamically calculate the text feature vector T through the gating network text ∈R 512 and visual feature vector V fusion ∈R 1536 The contribution weights α1 and β1 are used to reflect the relative importance of the two in the current task. Then, the text features and visual features are weighted summed according to the contribution weights α1 and β1 to generate the joint decision vector Q:
[0048] Q=α1·T text +β1·V fusion
[0049] Among them, α1 is the text feature vector T text ∈R 512 The contribution weight of β1 is the visual feature vector V fusion ∈R 1536 contribution weight.
[0050] The beneficial effects of the present invention are:
[0051] The present invention can not only collaboratively process three-dimensional spatial structures, spectral characteristics and long-term observation data, but also realize adaptive optimization of detection model parameters through reinforcement learning, effectively solving the problems of poor feature coupling and weak dynamic adaptability of traditional methods in complex scenarios, and providing a high-precision and highly robust detection method for the field of road changes in multimodal data images. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a schematic diagram of the overall process of the present invention. DETAILED DESCRIPTION
[0053] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings.
[0054] A multimodal road change detection method based on reinforcement learning optimization, such as Figure 1 As shown, the specific steps include:
[0055] S1, natural language prompt input and semantic parsing of design specifications;
[0056] This invention converts natural language prompt instructions input by users into structured and computable design constraints, which can accurately capture user needs and provide clear and operational task definitions for subsequent multimodal large model reasoning and reinforcement learning optimization, thereby achieving efficient and accurate road change detection optimization.
[0057] Specifically, the user inputs the road change image to be processed and completes the corresponding task requirement description T text "Based on inputting road change images of the same area at different times, the road change detection task is completed with an accuracy rate of no less than 80% and an intersection-over-union (IOU) index of no less than 60%."
[0058] S2, text semantic vectorization;
[0059] The RoBERTa pre-trained model is used to build a natural language prompt text encoder. The deep semantic representation of natural language prompts is achieved through word embedding and hierarchical attention mechanism. The specific steps are as follows:
[0060] The natural language prompt input by the user is segmented and mapped into a 768-dimensional semantic vector, which is then integrated with the embedding information in the field of position encoding and road change detection. Then, a local-global two-stage attention mechanism is used to capture the local semantic dependencies related to road changes in the natural language prompt in the first stage, and aggregate the overall context information of the natural language prompt in the second stage, ultimately generating a text feature vector T containing rich multimodal constraint information. text ∈R 512 , where R is the road domain relationship matrix, which is used to encode the semantic association between natural language and road change detection in the attention mechanism.
[0061] Embedded information in the field of road change detection refers to prior knowledge and domain features specifically embedded in the natural language prompt encoding process, enhancing the professionalism and accuracy of semantic representation. This core information includes road entity topology, typical road change patterns, spatiotemporal dynamic characteristics, multimodal alignment signals, and regulatory constraints. This information is specifically encoded and integrated into text features, enabling the model to more accurately understand professional descriptions of road changes and improving multimodal detection performance.
[0062] This hierarchical semantic representation method can not only accurately parse complex instructions input by users, but also provide high-quality text feature input for subsequent multimodal data fusion and reinforcement learning optimization.
[0063] S3, road scene multimodal data input and preprocessing;
[0064] Construct a multimodal data fusion mechanism for road change detection scenarios to implement multimodal data input and preprocessing of road scenes. The specific steps include:
[0065] S31, 3D spatial data analysis;
[0066] The three-dimensional geometric structure and topological relationship of the road and its surrounding environment are extracted from road remote sensing images and represented as graph structure data suitable for large-scale model processing, providing a rich spatial information foundation for subsequent road change analysis.
[0067] S32, remote sensing data enhancement;
[0068] Denoising, registration and joint extraction of spatial-spectral feature points are performed on hyperspectral remote sensing images and lidar point cloud data to generate dense and robust multimodal feature maps P cloud ∈R 2048×3 .
[0069] S33, time series data embedding;
[0070] The remote sensing time series dataset of the AVHRR sensor is preprocessed to extract the key features of road and environmental area changes, and mapped into a three-dimensional tensor containing information such as deformation displacement and spectral feature changes, forming an input data channel that integrates geometric, spectral and time series information.
[0071] Key features of road and environmental changes primarily include geometric deformation, spectral variation, and temporal evolution. Geometric features are extracted through edge detection, morphological analysis, and displacement field calculation; spectral features utilize multi-band indices and spectral unmixing analysis; and temporal features employ time series decomposition and clustering to identify change patterns. Ultimately, these features are fused into a three-dimensional time-space-feature tensor, which serves as input for multimodal change detection.
[0072] By fusing three-dimensional spatial data, enhanced remote sensing data, and time series data, the present invention can comprehensively and multi-dimensionally characterize the changing characteristics of road scenes, providing a reliable multimodal data foundation for subsequent change detection and optimization, and significantly improving detection accuracy and robustness.
[0073] S4, road image feature embedding and multimodal association, the specific steps include:
[0074] S41, three-dimensional space analysis branch, the specific steps include:
[0075] S411, constructing geometric coding branch;
[0076] The graph structure data generated by CAD tools is used as input to generate node features of geometric elements.
[0077] S412: construct feature encoders for different types of geometric elements respectively, and use a hierarchical message passing mechanism to implement node feature updates.
[0078] S413, constructing point cloud coding branch;
[0079] The PointNet++ architecture is used to process laser point cloud data. The FPS sampling strategy is used to generate hierarchical point sets. Radius queries are applied at each sampling level to construct local regions, and adaptive weighted convolution kernels are used to extract local geometric features.
[0080] S414, uses the cross-attention mechanism to establish cross-modal associations between geometric graph features and point cloud features.
[0081] S415, joint representation of three-dimensional features;
[0082] The associated features are input into the graph pooling layer, node feature aggregation is achieved through attention weights, and the final output is a 512-dimensional joint feature vector F containing geometric topology and three-dimensional structure. 3D ∈R D×128 , where D is the intermediate dimension after the geometric features are associated with the point cloud features, and D is usually 256.
[0083] S42, remote sensing enhancement coding branch;
[0084] The hyperspectral remote sensing image is input, and the MixUp data enhancement and channel random dropout methods are used for data enhancement. Then the spectral attention module and spatial-spectral separation convolution are introduced. Finally, positive and negative samples are constructed to conduct comparative learning on the enhanced features, and the output spectral enhancement feature F can be obtained. RS ∈R N×512 , where N is the number of positive and negative sample pairs constructed in contrastive learning.
[0085] S43, temporal spectral encoding branch;
[0086] Based on the 3D CNN network, the road hyperspectral remote sensing data and time series data fusion tensor are convolved to output the spatiotemporal perception feature F containing time series evolution and spectral differences. phy ∈R 512 .
[0087] S44, unified dimension;
[0088] The above multimodal features are uniformly expanded to 512 dimensions through the projection layer, and finally a visual feature vector V is formed for road change detection. fusion :
[0089] V fusion =[F 3D ; F RS ; F phy ]∈R512×3 =R 1536 Formula 1
[0090] The visual feature vector V of the present invention fusion Including sub-task features such as change detection, temporal analysis, and scene understanding, through the collaborative training of a shared backbone network and task-specific heads, unified task optimization is achieved at three levels: input (multi-time series images of the same area), process (transmission of spatiotemporal features), and output (accuracy ≥ 80% and IOU ≥ 60%).
[0091] S5, multimodal feature alignment and dynamic fusion, the specific steps include:
[0092] S51, text guides visual focus;
[0093] Take the text feature vector T text ∈R 512 is the Query vector, the visual feature vector V fusion ∈R 1536 The attention weights between multiple modalities are calculated as Key-Value vectors, the spatial region features closely related to the semantic requirements of natural language prompts are strengthened, and the text features and visual features are cross-fused to obtain the fused multimodal features.
[0094] S52, visual feedback semantic revision;
[0095] The text feature vector T is transformed into text ∈R 512 The dimension is expanded to 1536, making it consistent with the visual feature vector V fusion ∈R 1536 , thereby achieving feedback enhancement of the semantic representation of natural language prompts by visual features and providing a consistent vector space for subsequent multimodal interactions.
[0096] S53, gated feature fusion;
[0097] Dynamically calculate the text feature vector T through the gating network text ∈R 512 and visual feature vector V fusion ∈R 1536 The contribution weights α1 and β1 are used to reflect the relative importance of the two in the current task. Then, the text features and visual features are weighted summed according to the contribution weights α1 and β1 to generate the joint decision vector Q:
[0098] Q=α1·T text +β1·V fusion Formula 2
[0099] Among them, α1 is the text feature vector T text ∈R512 The contribution weight of β1 is the visual feature vector V fusion ∈R 1536 Contribution weights; α1, β1∈[0,1]; α1+β1=1.
[0100] Based on the fused multimodal features and gated fusion technology, users can perform a comprehensive and in-depth holistic fusion of all the aforementioned feature variables. This process first pre-processes the feature data from different modalities (such as road geometry, spectral information, and temporal changes) to ensure data consistency and comparability. Subsequently, the gated fusion mechanism dynamically adjusts the weights of each feature variable, enabling the detection model to automatically learn and assign the importance of each modal feature based on the needs of the current road change detection task, thereby achieving more accurate feature expression.
[0101] S6, detection model parameter optimization and online simulation verification;
[0102] Users integrate multimodal large models (such as Deepseek) and simulation engines, use processed feature data to build intelligent parameter configuration and verification, and convert the parameter vector output by the reinforcement learning strategy network into a specific parameter configuration scheme for the road change detection method. After the large model semantics are verified, the simulation engine evaluates the performance online.
[0103] During the simulation process, the detection accuracy index P, real-time response delay T, and computing resource consumption C are dynamically obtained to construct a multi-objective reward function R for the road change detection task:
[0104]
[0105] in, η and γ are adjustable weight coefficients that guide the reinforcement learning policy network to dynamically generate the optimal detection parameter configuration that meets the requirements of multimodal input natural language prompts, realizing an intelligent closed loop of "data perception-strategy evolution-method tuning", thereby enhancing the accuracy and practicality of road change detection when users use the method of the present invention.
[0106] S7, multi-objective reinforcement learning optimization and scenario verification;
[0107] Based on the improved PPO algorithm, an uncertainty-aware exploration strategy is used to dynamically adjust the action distribution variance σ σWhen the historical reward fluctuations for road change detection exceed a set threshold, the exploration intensity is automatically increased to prevent the strategy from falling into a local optimal solution. A curriculum learning mechanism is introduced at the beginning of training, first relaxing computing resource constraints to prioritize exploring the limits of detection accuracy and response performance; then, resource constraint penalties are gradually increased to push the model to converge towards more practical and efficient scenarios. The training process uses a priority experience replay mechanism, performing weighted sampling of Pareto frontier solutions based on time-difference errors. The advantage function and gradient clipping method are combined to steadily update the policy network parameters, ensuring efficient convergence of the algorithm and its adaptability and robustness in real-world road change detection scenarios.
[0108] Furthermore, to further enhance the adaptability and accuracy of the detection solution, detection results and corresponding reward information are fed back to S5. In S5, this feedback data is further analyzed and processed using the improved PPO algorithm, fine-tuning the policy network to adapt to changing road conditions and user needs. This approach not only enables users to obtain an efficient detection solution but also enables them to continuously optimize and improve its performance in practice, meeting the diverse needs of road change detection applications.
[0109] This invention provides a multimodal road change detection method based on reinforcement learning optimization. The specific implementation method includes the following technical means: First, natural language prompts are encoded into 768-dimensional semantic vectors using the RoBERTa model and embedded with road domain knowledge. Second, feature extraction is performed on multi-source remote sensing data, using PointNet++ to process point cloud data and 3D CNN to analyze time-series spectral features, and a geometric-spectral-time series three-dimensional tensor is constructed. Then, a cross-modal attention mechanism is used to align text and visual features, and a gated network is used to dynamically fuse multimodal features. Finally, reinforcement learning optimization is performed based on an improved PPO algorithm, integrating curriculum learning and prioritized experience replay techniques to construct a multi-objective reward function that includes detection accuracy, response latency, and resource consumption, thereby automatically optimizing detection parameters. Through three key technical steps: hierarchical semantic parsing, multimodal feature fusion, and adaptive strategy optimization, this method enables those skilled in the art to build an end-to-end intelligent road change detection system that achieves balanced optimization of multiple objectives while ensuring detection accuracy ≥80% and IoU ≥60%.
[0110] The contents not described in detail in the specification of the present invention belong to the existing technologies disclosed in this field.
Claims
1. A multimodal road change detection method based on reinforcement learning optimization, characterized in that: The following steps are involved: S1, natural language prompt input and semantic parsing of design specifications; Convert user-entered natural language prompt instructions into structured, computable design constraints; S2, text semantic vectorization; The RoBERTa pre-trained model is used to build a natural language prompt text encoder, and deep semantic representation of natural language prompts is achieved through word embedding and hierarchical attention mechanism; S3, road scene multimodal data input and preprocessing; Construct a multimodal data fusion mechanism for road change detection scenarios to implement multimodal data input and preprocessing of road scenes. The specific steps include: S31, 3D spatial data analysis; S32, remote sensing data enhancement; S33, time series data embedding; S4, road image feature embedding and multimodal association, the specific steps include: S41, three-dimensional space analysis branch; S42, remote sensing enhancement coding branch; S43, temporal spectral encoding branch; S44, unified dimension; S5, multimodal feature alignment and dynamic fusion, the specific steps include: S51, text guides visual focus; S52, visual feedback semantic revision; S53, gated feature fusion; S6, detection model parameter optimization and online simulation verification; Users integrate a large multimodal model with a simulation engine, use processed feature data to build intelligent parameter configuration and verification, and convert the parameter vector output by the reinforcement learning strategy network into a specific parameter configuration scheme for the road change detection method. After semantic verification of the large model, the simulation engine evaluates the performance online. S7, multi-objective reinforcement learning optimization and scenario verification; Based on the improved PPO algorithm, an uncertainty-aware exploration strategy is used to dynamically adjust the action distribution variance σ σ When the historical reward fluctuation of road change detection exceeds the set threshold, the exploration intensity is automatically increased to prevent the strategy from falling into the local optimal solution; a curriculum learning mechanism is introduced at the beginning of training to first relax the computing resource constraints to prioritize the exploration of the limits of detection accuracy and response performance; then the resource constraint penalty is gradually increased to push the model to converge to a more practical and efficient scenario; the training process adopts a priority experience replay mechanism, performs weighted sampling of the Pareto frontier solution according to the time difference error, and combines the advantage function and gradient clipping method to stably update the policy network parameters, ensuring the efficient convergence of the algorithm and its adaptability and robustness in real road change detection scenarios.
2. The multimodal road change detection method based on reinforcement learning optimization according to claim 1, characterized in that: In S1, the natural language prompt command input by the user is converted into a structured and computable design constraint. Specifically, the user inputs the road change image to be processed and completes the corresponding task requirement description T text : "Based on the input of road change images of the same area at different times, the road change detection task is completed with an accuracy rate of no less than 80% and an intersection-over-union (IOU) index of no less than 60%." 3. The multimodal road change detection method based on reinforcement learning optimization according to claim 1, characterized in that: In S2, the RoBERTa pre-trained model is used to build a natural language prompt text encoder. The specific steps for achieving deep semantic representation of natural language prompts through word embedding and hierarchical attention mechanism are as follows: The natural language prompt input by the user is segmented and mapped into a 768-dimensional semantic vector, which is then integrated with the embedding information in the field of position encoding and road change detection. Then, a local-global two-stage attention mechanism is used to capture the local semantic dependencies related to road changes in the natural language prompt in the first stage, and aggregate the overall context information of the natural language prompt in the second stage, ultimately generating a text feature vector T containing rich multimodal constraint information. text ∈R 512 , where R is the road domain relationship matrix.
4. The multimodal road change detection method based on reinforcement learning optimization according to claim 1, characterized in that: In S41, the specific steps of the three-dimensional space analysis branch include: S411, constructing geometric coding branch; S412, constructing feature encoders for different types of geometric elements respectively, and using a hierarchical message passing mechanism to implement node feature updates; S413, constructing point cloud coding branch; S414, using the cross-attention mechanism to establish cross-modal associations between geometric features and point cloud features; S415, joint representation of three-dimensional features.
5. The multimodal road change detection method based on reinforcement learning optimization according to claim 1, characterized in that: In S42, the specific steps of the remote sensing enhancement coding branch are: The hyperspectral remote sensing image is input and the data enhancement operation is performed using the MixUp data enhancement and channel random dropout methods. Then the spectral attention module and spatial-spectral separation convolution are introduced. Finally, positive and negative samples are constructed to conduct comparative learning on the enhanced features, and the output spectral enhancement feature F is obtained. RS ∈R N×512 , where N is the number of positive and negative sample pairs constructed in contrastive learning.
6. The multimodal road change detection method based on reinforcement learning optimization according to claim 1, characterized in that: In S43, the specific steps of the time series spectrum encoding branch are: Based on the 3D CNN network, the road hyperspectral remote sensing data and time series data fusion tensor are convolved to output the spatiotemporal perception feature F containing time series evolution and spectral differences. phy ∈R 512 .
7. The multimodal road change detection method based on reinforcement learning optimization according to claim 1, characterized in that: In S44, the specific steps of unifying the dimensions are: The multimodal features are uniformly expanded to 512 dimensions through the projection layer, and finally a visual feature vector V is formed for road change detection. fusion ∈R 1536 .
8. The multimodal road change detection method based on reinforcement learning optimization according to claim 1, characterized in that: In S51, the specific steps of guiding visual focus with text are: Take the text feature vector T text ∈R 512 is the Query vector, the visual feature vector V fusion ∈R 1536 The attention weights between multiple modalities are calculated as Key-Value vectors, the spatial region features closely related to the semantic requirements of natural language prompts are strengthened, and the text features and visual features are cross-fused to obtain the fused multimodal features.
9. The multimodal road change detection method based on reinforcement learning optimization according to claim 1, characterized in that: In S52, the specific steps of visual feedback semantic correction are: The text feature vector T is transformed into text ∈R 512 The dimension is expanded to 1536, making it consistent with the visual feature vector V fusion ∈R 1536 The dimensions of are aligned, thereby achieving feedback enhancement of the semantic representation of natural language prompts using visual features.
10. The multimodal road change detection method based on reinforcement learning optimization according to claim 1, characterized in that: In S53, the specific steps of gated feature fusion are: Dynamically calculate the text feature vector T through the gating network text ∈R 512 and visual feature vector V fusion ∈R 1536 The contribution weights α1 and β1 are used to reflect the relative importance of the two in the current task. Then, the text features and visual features are weighted summed according to the contribution weights α1 and β1 to generate the joint decision vector Q: Q=α1·T text +β1·V fusion Among them, α1 is the text feature vector T text ∈R 512 The contribution weight of β1 is the visual feature vector V fusion ∈R 1536 contribution weight.
Citation Information
Cited By
Road structure design method based on large model and reinforcement learning
CN120910969A
A road structure design method based on large model and reinforcement learning
CN120910969B
Multi-modal fusion longitudinal section design optimization method
CN121859414A
A multi-modal fusion longitudinal profile design optimization method
CN121859414B