Method for positioning vegetation area in urban foggy weather based on streetscape and large model
By combining street scene data and multimodal large-modal model technology, including SAM, WRPM and SCP-Net, the problem of vegetation area identification in foggy weather is solved, accurate positioning and management of urban vegetation areas is achieved, and urban greening benefits are improved.
Patent Information
- Application Number
- CN202510506434.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-06-17
AI Technical Summary
In heavy foggy weather, it is difficult for intelligent identification technology to effectively identify urban vegetation areas, affecting urban management and greening benefits.
Street scene data, artificial intelligence multimodal large model SAM, image defogging module WRPM and image segmentation correction module SCP-Net are used to realize the localization of urban vegetation in heavy fog weather. Specific steps include data preprocessing, fog removal operation, vegetation area segmentation and result correction.
This method can accurately identify and position urban vegetation areas in heavy fog weather, and improve the efficiency and accuracy of urban greening management.
Smart Images

Figure CN120164115A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of remote sensing and machine learning, and particularly to a method for locating vegetation areas under urban foggy weather based on street views and large models. Background Art
[0002] Cities are important carriers of human social life. Urban street greening spreads throughout every corner of the city, providing services for humans living in the city, and is of great significance for achieving urban carbon neutrality, alleviating the heat island effect, and improving urban air quality. Urban managers need to locate street greening in order to master its distribution, maximize the benefits of street greening, and effectively carry out urban construction management planning. With the development of machine learning and artificial intelligence technologies, it has become a convenient and effective way to conduct intelligent analysis on the image data of urban greening and automatically identify greening areas from rich urban images. However, the weather in cities is changeable, and foggy weather will seriously affect the effect of intelligent identification, and there is an urgent need to improve the identification technology. Summary of the Invention
[0003] In view of the above-mentioned defects of the prior art, the present invention aims at the problem of locating vegetation areas under harsh weather scenarios such as urban fog. Using street view data, with the help of the artificial intelligence multi-modal large model SAM, the image dehazing module WRPM, and the image segmentation correction module SCP-Net, a method for locating vegetation areas in urban areas under foggy weather is proposed. This method first starts from data collection and preprocessing, crops and aligns the collected street view data, selects the research area, randomly selects 3% from the screened street view data for vegetation area segmentation, and after the segmentation is completed, manual annotation is carried out to mark the set of segmentation error points and the set of vegetation area points that have not been correctly segmented. Secondly, the processed image is dehazed using the WRPM module. Thirdly, the dehazed street view data is input into SAM for segmentation to extract the vegetation area in the image. Finally, the trained SCP-Net is used to correct the segmentation result to obtain the final vegetation area location result.
[0004] To achieve the above object, the present invention provides a method for locating vegetation areas under urban foggy weather based on street views and large models, which is characterized by including the following steps:
[0005] Step 1: Collect street view data within the research urban area and perform preprocessing operations of cropping and coordinate alignment;
[0006] Step 2: Use the WRPM module to perform dehazing operations on the processed street view data;
[0007] Step 3: Use SAM to segment the vegetation area from the dehazed street view data;
[0008] Step 4: Use SCP-Net to correct the results of SAM segmentation to obtain the final vegetation area.
[0009] Furthermore, Step 1 specifically includes:
[0010] A0. Organize the original street view data, crop the data into the 224×224 format, and set the region selection coefficient as bt. Through function screening operations, select the street view data within the research area;
[0011] A1. Randomly select 3% from the screened street view data for vegetation area segmentation. After the segmentation is completed, perform manual annotation to mark the set P of segmentation error points and the set T of vegetation area points that are not correctly segmented, where is the set of vegetation area points in the image, is the set of vegetation area points segmented by the model.
[0012] Furthermore, Step 2 specifically includes:
[0013] B0. Establish the thick fog dehazing module Landscape-FFA in WRPM, so that the original street view data S can obtain the thick fog recovery feature S after passing through Landscape-FFA high , where the expression of Landscape-FFA is Multiple residual blocks ResBlock are added to this module k and the channel attention module CA ffa and the pixel attention module PA ffa , represents N cascaded residual blocks ResBlock k , each residual block contains a convolutional layer and an activation function, and fuses shallow and deep features through skip connections to extract deep features in the image, effectively processing areas with thick haze and blurred features in the image;
[0014] B1. Establish the thin fog dehazing module Landscape-UNet in WRPM, so that the original street view data S can obtain the thin fog recovery feature S after passing through Landscape-UNet low , and the expression of Landscape-UNet is where, is the l-th downsampling module, and the output feature dimension is reduced to 1 / 2 of the input size l , is the Bottleneck layer, which integrates the dual attention mechanism, is the l-th upsampling module, which restores the size through transposed convolution and fuses the skip features, It is an upsampling cascade operation, and the final output S is a feature S with the same resolution as the input. low , specifically, the downsampling stage of Landscape-UNet can be expressed as Gradually extract low-dimensional features through L layers of convolutional blocks ConvBlock. Each layer contains convolution, activation function, and downsampling. In the middle bottleneck layer, through introduce channel attention CA ffa and pixel attention PA ffa The dual attention architecture of First apply channel attention CA ffa , and then perform spatial enhancement through pixel attention PA ffa . In the upsampling stage, the process can be expressed as Use deconvolution DeconvBlock to gradually restore the spatial dimension, and splice Concat the downsampled features with the upsampled features to retain detailed information;
[0015] B2. Establish the feature aggregation module Landscape-FIM in WRPM. Landscape-FIM aggregates the foggy restoration feature S high and the haze restoration feature S low , and obtain the final dehazed street view data Sw. The process can be expressed as Among them, is the channel attention weight, is the pixel attention weight, is the weighted feature channel splicing.
[0016] Furthermore, step three specifically includes:
[0017] C0. Input the dehazed street view data Sw into the Visual Encoder;
[0018] C1. Input the text prompt "vegetation" into the Prompt Encoder;
[0019] C2. Decode through the Mask Decoder to obtain an image marking the vegetation area;
[0020] C3. Restore the marked area to obtain the generated initial segmentation mask M SAM , M SAM represents that this is a vegetation area.
[0021] Furthermore, step four specifically includes:
[0022] D0. Use the labeled data in Step 1 to train SCP-Net. Set the network loss function to consist of two parts: cross-entropy and Dice loss. The calculation formula is where represents the total loss function, λ is the adjustment coefficient, and the specific forward propagation process of SCP-Net can be expressed as where f θ is the parameterized SCP-Net model, are the network parameters, is the predicted segmentation mask;
[0023] D1. Set automatic network parameter update during the training process. The update rule is where θ is the learning rate, is the gradient of the loss function with respect to the network parameters;
[0024] D2. After the training is completed, input the initial segmentation mask M SAM of the street view data processed by SAM into the encoder;
[0025] D3. The encoder part of SCP-Net extracts multi-scale features of M SAM . The decoder part fuses features of different scales through skip connections, gradually restores the spatial resolution, and finally outputs the corrected segmentation mask M SCP ;
[0026] D4. Restore the finally corrected segmentation mask M SCP to obtain the finally located vegetation area.
[0027] The beneficial effects of the present invention are as follows: Through street view data and by means of large model technologies such as SAM, WRPM, and SCP-Net, the present invention realizes the dehazing operation of street view data and completes the identification and location of urban vegetation areas in foggy weather. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 is the overall flow schematic diagram of the method proposed by the present invention;
[0029] Figure 2 is the flow schematic diagram of street view data preprocessing;
[0030] Figure 3 is the flow schematic diagram of WRPM for dehazing. DETAILED DESCRIPTION OF THE INVENTION
[0031] The present invention will be further described below with reference to the drawings and embodiments:
[0032] This embodiment locates the vegetation area in street view data under harsh weather such as fog.
[0033] As Figure 1 shown, this embodiment provides a method for locating vegetation areas under urban heavy fog weather based on street view and large models, including the following steps:
[0034] Step 1: Preprocessing operation of street view data, and the process is as Figure 2 shown. The specific steps include:
[0035] A0. Organize the original street view data, crop the data into the 224×224 format, and at the same time set the region selection coefficient as bt. Through the function screening operation, select the street view data within the research region;
[0036] A1. Randomly select 3% from the selected street view data for vegetation area segmentation. After the segmentation is completed, perform manual annotation to mark the set P of segmentation error points and the set T of vegetation area points that are not correctly segmented. The calculation method of P is The calculation method of T is is the set of vegetation area points in the image, is the set of vegetation area points segmented by the model.
[0037] Step 2: The operation of the street view data defogging module WRPM, and the process is as Figure 3 shown. The specific steps include:
[0038] B0. Establish the heavy fog defogging module Landscape-FFA in WRPM, so that the original street view data S can obtain the heavy fog recovery feature S high after passing through Landscape-FFA. Among them, the expression of Landscape-FFA is Multiple residual blocks ResBlock are added to this module k and the channel attention module CA ffa and the pixel attention module PA ffa , represents N cascaded residual blocks ResBlock k , and each residual block contains a convolutional layer and an activation function, and fuses shallow and deep features through skip connections to extract deep features in the image, effectively processing areas with thick haze and blurred features in the image;
[0039] B1. Establish the light fog defogging module Landscape-UNet in WRPM. The original street view data S obtains the light fog recovery feature S low after passing through Landscape-UNet. The expression of Landscape-UNet is Among them, is the l-th downsampling module, and the output feature dimension is reduced to 1 / 2 of the input sizel , is the Bottleneck layer, integrating the dual attention mechanism. is the upsampling module of the l-th layer, restoring the size through transposed convolution and fusing skip features. is the cascaded upsampling operation, and finally outputs the feature S with the same resolution as the input. low , specifically, the downsampling stage of Landscape-UNet can be expressed as gradually extract low-dimensional features through L layers of convolutional blocks ConvBlock. Each layer contains convolution, activation function, and downsampling. In the intermediate bottleneck layer, through introduce the channel attention CA ffa and the pixel attention PA ffa dual attention architecture, for the features of the deepest layer first apply the channel attention CA ffa , and then perform spatial enhancement through the pixel attention PA ffa . In the upsampling stage, the process can be expressed as use the transposed convolution DeconvBlock to gradually restore the spatial dimension, and splice the downsampled features with the upsampled features Concat to retain the detailed information;
[0040] B2. Establish the feature aggregation module Landscape-FIM in WRPM. Landscape-FIM aggregates the foggy restoration feature S high and the haze restoration feature S low , to obtain the final de-fogged street view data Sw. The process can be expressed as where, is the channel attention weight, is the pixel attention weight, is the weighted feature channel splicing.
[0041] Step 3. The segmentation work of the large model SAM, the specific steps include:
[0042] C0. Input the de-fogged street view data Sw into the Visual Encoder;
[0043] C1. Input the text prompt "vegetation" into the Prompt Encoder;
[0044] C2. Decode through the Mask Decoder to obtain the image marking the vegetation area;
[0045] C3. Restore the marked area to obtain the generated initial segmentation mask M SAM , MSAM It represents that this is a vegetation area here.
[0046] Step 4: The detailed process of SCP-Net calibration includes the following specific steps:
[0047] D0. Use the annotation data in Step 1 to train SCP-Net. Set the network loss function to consist of two parts: cross-entropy and Dice loss. The calculation formula is where represents the total loss function, λ is the adjustment coefficient, and the forward propagation process of SCP-Net can be specifically expressed as where f θ is the parameterized SCP-Net model, are the network parameters, is the predicted segmentation mask;
[0048] D1. Set automatic network parameter update during the training process. The update rule is where θ is the learning rate, is the gradient of the loss function with respect to the network parameters;
[0049] D2. After the training is completed, input the initial segmentation mask M SAM of the street view data processed by SAM into the encoder;
[0050] D3. The encoder part of SCP-Net extracts multi-scale features of M SAM . The decoder part fuses features of different scales through skip connections, gradually restores the spatial resolution, and finally outputs the calibrated segmentation mask M SCP ;
[0051] D4. Restore the finally calibrated segmentation mask M SCP to obtain the finally located vegetation area.
[0052] In summary, for the urban scene in foggy weather, this embodiment proposes a method for locating urban vegetation areas using the joint multi-feature defogging network WRPM and the vegetation area recognition and calibration network SCP-Net, and integrating the multi-modal large model SAM. Specifically, this method uses WRPM to defog the unclear street view images in foggy weather, uses SAM to segment the vegetation areas in the images to achieve the location of the vegetation areas, and finally uses SCP-Net to further calibrate the segmentation results to improve the accuracy of the location.
[0053] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field according to the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art shall fall within the protection scope determined by the claims.
Claims
1. A method for locating vegetation areas in urban foggy weather based on street scenes and large models, characterized in that: The steps include: Step 1: Collect street view data in the study city area and perform preprocessing operations such as cropping and coordinate alignment; Step 2: Use the WRPM module to perform defogging on the processed street view data; Step 3: Use the standard SAM (Segment Anything Model) to segment the vegetation area of the defogging street view data; Step 4: Use SCP-Net to correct the results of SAM segmentation to obtain the final vegetation area.
2. A method for locating vegetation areas in urban foggy weather based on street scenes and large models as claimed in claim 1, characterized in that: The step 1 specifically includes: A0. Arrange the original street view data, crop the data into 224×224 format, set the area selection coefficient to bt, and select the street view data in the study area through function screening operation; A1. Randomly select 3% of the selected street view data for vegetation area segmentation. After the segmentation is completed, manual labeling is performed to mark the set P of segmentation error points and the set T of vegetation area points that are not correctly segmented. is the set of vegetation area points in the image, It is the set of vegetation area points segmented by the model.
3. The method for locating vegetation areas in urban foggy weather based on street scenes and large models as claimed in claim 1, characterized in that: The specific steps of the WRPM module performing the defogging operation in step 2 include: B0. Establish the dense fog removal module Landscape-FFA in WRPM, so that the original street view data S can be obtained through Landscape-FFA to obtain dense fog recovery features S high , where the expression of Landscape-FFA is This module adds multiple residual blocks ResBlock k And the channel attention module CA ffa and pixel attention module PA ffa , Represents N series of residual blocks ResBlock k ,Each residual block contains a convolution layer and an activation function, and fuses shallow and deep features through jump connections to extract deep features in the image and effectively process areas with heavy haze and blurred features in the image; B1. Establish the mist removal module Landscape-UNet in WRPM, so that the original street view data S can be obtained through Landscape-UNet to obtain the mist recovery feature S low , the expression of Landscape-UNet is in, It is the l-th layer downsampling module, and the output feature dimension is reduced to 1 / 2 of the input size l , It is the Bottleneck layer, which integrates the dual attention mechanism. It is the l-th layer upsampling module, which restores the size and fuses the jump features through deconvolution. It is a cascade operation of upsampling layers, and the final output S has the same resolution as the input feature S low ,Specifically, the downsampling stage of Landscape-UNet can be expressed as Low-dimensional features are gradually extracted through L layers of convolutional blocks ConvBlock. Each layer contains convolution, activation function and downsampling. In the middle bottleneck layer, Introduced channel attention CA ffa and pixel attention PA ffa The dual attention architecture for the deepest features First apply channel attention CA ffa , and then through pixel attention PA ffa Perform spatial enhancement. In the upsampling stage, the process can be expressed as Use deconvolution DeconvBlock to gradually restore the spatial dimension and use skip connections to connect the downsampled features. Concatenate with the upsampled features to retain detail information; B2. Establish the feature aggregation module Landscape-FIM in WRPM. Landscape-FIM aggregates the dense fog recovery feature S high and mist recovery feature S low , the final dehazed street view data Sw is obtained, and the process can be expressed as in, is the channel attention weight, is the pixel attention weight, It is the weighted feature channel concatenation.
4. The method for locating vegetation areas in urban foggy weather based on street scenes and large models as claimed in claim 1, characterized in that: The specific steps of using SAM to perform image segmentation in step 3 include: C0, input the dehazed street view data Sw into Visual Encoder; C1. Input the text prompt word "vegetation" into Prompt Encoder; C2, decode the image of the marked vegetation area through Mask Decoder; C3, restore the marked area to generate the initial segmentation mask M SAM , M SAM This indicates that this is a vegetation area.
5. The method for locating vegetation areas in urban foggy weather based on street scenes and large models as claimed in claim 1, characterized in that: The specific steps of using SCP-Net to correct the results of SAM segmentation in step 4 include: D0. Use the labeled data in step 1 to train SCP-Net. Set the network loss function to consist of cross entropy and Dice loss. The calculation formula is in represents the total loss function, λ is the adjustment coefficient, and the SCP-Net forward propagation process can be specifically expressed as where f θ is the parameterized SCP-Net model, θ is the network parameter, is the predicted segmentation mask; D1. Set automatic network parameter update during training. The update rule is: where θ is the learning rate, is the gradient of the loss function with respect to the network parameters; D2. After training, the initial segmentation mask M of the street view data processed by SAM is SAM Pass in the encoder; D3, SCP-Net encoder part extracts M SAM The decoder fuses features of different scales through skip connections, gradually restores the spatial resolution, and finally outputs the corrected segmentation mask M SCP ; D4, the final corrected segmentation mask M SCP Restore and obtain the final located vegetation area.