A method for built-up area environmental assessment based on remote sensing image scene understanding

Through the method based on remote sensing image scene understanding, the problem of insufficient coverage of street scene image data is solved, and efficient and comprehensive assessment of the urban built-up environment is achieved. It can extract landform relations and attribute information and provide accurate environmental evaluation.

CN116307857BActive Publication Date: 2025-08-12CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310161265.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-24
Publication Date
2025-08-12
Estimated Expiration
2043-02-24

AI Technical Summary

Technical Problem

In the prior art, the street view image data is updated at a low frequency, making it difficult to cover all areas of the city, and the update frequency of quantitative indicators and indicator weights is also low, resulting in the inability to assess the urban built environment in a timely and comprehensive enough.

Method used

Using a method based on the understanding of remote sensing image scenes, remote sensing images are acquired by region-by-regional regions, environmental evaluation indicators are determined, image description model is trained, environmental evaluation map is drawn using ArcGIS software, and continuous distribution results are achieved through interpolation processing.

Benefits of technology

It has achieved comprehensive coverage and high frequency assessment of the environment of urban built-up areas, and can extract spatial relationships and attribute information between land objects and provide objective and accurate environmental evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116307857B_ABST
    Figure CN116307857B_ABST
Patent Text Reader

Abstract

The present invention provides a method for evaluating the environment of built-up areas based on scene understanding of remote sensing images, comprising step S1: determining an environmental evaluation index; step S2: using the environmental evaluation index to calculate an environmental evaluation dataset for the built-up area; dividing the environmental evaluation dataset into a training set and a test set in proportion; step S3: using the training set to train an image description model; step S4: obtaining a remote sensing image of the built-up area to be evaluated and cropping it into a plurality of scene images; using the image description model to predict each of the scene images; calculating an environmental score corresponding to each of the scene images using the environmental evaluation index according to step S2; and using ArcGIS software to plot the environmental scores of each of the scene images into a map for environmental evaluation of the built-up area to be evaluated. The present invention achieves objective evaluation of the environment of the urban built-up area to be evaluated and has the advantage of good usability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of urban planning, and in particular to a method for evaluating the environment of built-up areas based on remote sensing image scene understanding. Background Art

[0002] A high-quality built environment is crucial for improving the health and well-being of residents. Urban planners and managers need to evaluate the built environment and minimize adverse impacts on citizens by renovating areas with poor built environments. In other words, assessing the quality of the urban built environment can help planners and managers make urban renewal decisions, increase the livability of the environment, and improve resident satisfaction. Existing research on assessing the quality of the urban built environment has mostly been conducted from a streetscape perspective. However, streetscape image data is acquired from panoramic vehicle-mounted street images taken along roads. Environmental assessments are limited to areas on both sides of the road, making it difficult to cover the built environment in every area of the city. Furthermore, streetscape image data is updated infrequently (specifically, updates are typically every two or three years, with the fastest taking about one year), making it difficult to meet the requirements for building environment assessments in rapidly developing cities.

[0003] In addition, the use of quantitative indicators and the calculation of indicator weights are also the main contents of urban built environment assessment research. These indicators are often derived from social statistics, surveys, street views and GIS data. For example, Nguyen et al. used Google Street View images to extract street greenness, crosswalks and building type derived indicators to describe the built environment at the block level in three American cities. Chan and Liu analyzed the impact of neighborhood environment on builder health by collecting questionnaires and establishing indicators. Although the above indicators can be obtained, there are problems with low data update frequency (specifically, social statistics and GIS data are often counted on an annual basis; survey data require personnel to conduct field visits and surveys and are only updated when necessary) and time-consuming.

[0004] At present, although the use of high-resolution remote sensing imagery can increase the frequency of data updates, its data information has not been fully utilized in assessing the environment of urban built-up areas. For example, existing environmental assessment methods can only use the classification information of land objects in high-resolution remote sensing images, while the spatial relationship between land objects (such as whether buildings are surrounded by vegetation) and the attribute information of land objects (such as high or low building density) are not used. Summary of the Invention

[0005] The present invention aims to provide a method for building area environmental assessment based on remote sensing image scene understanding, which has the advantages of wide coverage of earth observation data, high update frequency and good usability. The specific technical solution is as follows:

[0006] A method for building area environmental assessment based on remote sensing image scene understanding includes the following steps:

[0007] Step S1: Divide the built-up area into a plurality of continuous regions, obtain remote sensing images corresponding to each region of the built-up area, and determine environmental evaluation indicators for each remote sensing image; the environmental evaluation indicators include building density, road connectivity, vegetation cover, industrial area distribution, and leisure activity area distribution;

[0008] Step S2: using the environmental evaluation index to calculate the environmental score corresponding to each of the remote sensing images, and combining the environmental scores to obtain an environmental evaluation dataset of the built-up area; dividing the environmental evaluation dataset into a training set and a test set in proportion;

[0009] Step S3: using the training set to train and obtain an image description model;

[0010] Step S4, obtaining a remote sensing image of the built-up area to be evaluated, and cropping it into a plurality of scene images; using the image description model to predict each of the scene images, and then outputting a sentence corresponding to the environmental evaluation index; according to step S2, using the environmental evaluation index to calculate the environmental score corresponding to each of the scene images; using ArcGIS software to draw the environmental score of each of the scene images into a map for environmental evaluation of the built-up area to be evaluated.

[0011] Optionally, the method for built-up area environment assessment based on remote sensing image scene understanding further includes step S5, specifically using interpolation processing to process the map into an environmental assessment map of the built-up area to be assessed with a continuous distribution result.

[0012] Optionally, step S5 includes:

[0013] Step S5.1: Mark the areas corresponding to the scene images as corresponding spatial grids, and move the spatial grid located in the center of all scene images outward along each perimeter of the spatial grid until the moved area covers all scene images, wherein each movement distance is one-half of the spatial grid range;

[0014] Step S5.2: Use the image description model to evaluate the scene image corresponding to the moved spatial grid to obtain the environment score corresponding to each spatial grid; assign the environment score of each spatial grid to the center point x of each spatial grid. i ;

[0015] Step S5.3: Use the inverse distance weighted method in ArcGIS to calculate the center point x obtained in step S5.2. i Perform interpolation processing, where the center point x i The inverse distance weight is calculated as follows:

[0016]

[0017] In formula (1), x0 is the position of the estimated center point; x i is the position of the known center point; n is the total number of known center points; i is the i-th known center point; the estimated value X(x0) is the n-th measured value X(x i ) weighted average; ω i For each known center point x i The weight of

[0018] The ω i Calculate using the following formula (2):

[0019] ω i =d 0,i -p Formula (2)

[0020] In formula (2), d 0,i is the estimated center point x0 and the known center point x i The Euclidean distance between them; p is the exponential power parameter;

[0021] Step S5.4: Use ArcGIS mapping software to produce the interpolation results into an environmental assessment map of the built-up area to be evaluated.

[0022] Optionally, in step S3, the image description model includes an image feature extraction network ResNet-152, a two-layer fully connected network, a single-layer feedforward network, a sentence LSTM network, and a word LSTM network that sequentially process the training set data. The specific processing process is as follows:

[0023] Step S3.1, inputting the corresponding remote sensing image in the training set into the image feature extraction network ResNet-152 to obtain the visual feature v;

[0024] Step S3.2: Input the visual feature v into a two-layer fully connected network to obtain a predicted identification phrase corresponding to the environmental evaluation index. The predicted identification phrase Converted into semantic features m ;

[0025] Step S3.3: The semantic feature a m And the visual feature v is input into a single-layer feedforward network to obtain the context feature ctx (s) ;

[0026] Step S3.4: the context feature ctx (s) Input into the sentence LSTM network to obtain feature t(s) ;

[0027] Step S3.5: The feature t (s) Input into the word LSTM network, and finally generate the predicted image description sentence

[0028] Step S3.6: Calculate the total error L, specifically by adding the predicted identification phrase obtained in step S3.2 The ground truth phrase annotated with the corresponding remote sensing image Use cross entropy loss to calculate the error L1; the predicted image description sentence obtained in step S3.5 Real image description sentences annotated with corresponding remote sensing images The error L2 is calculated using cross entropy loss; the total error L is the sum of the error L1 and the error L2, specifically expressed by the following formula (3):

[0029]

[0030] In formula (3), N represents the total number of remote sensing images in the training set; i represents the i-th remote sensing image in the training set; c represents the c-th sentence in each remote sensing image in the training set; represents the true value of the identification phrase corresponding to the cth sentence of the i-th image in the training set; represents the predicted value of the identification phrase corresponding to the cth sentence of the i-th image in the training set; represents the true value of the cth sentence in the i-th image in the training set, represents the predicted value of the cth sentence in the i-th image in the training set;

[0031] Step S3.7: Backpropagate the total error L in step S3.6 and use the Adam algorithm to update the network parameters; the network parameters include the parameters W1 of the image feature extraction network ResNet-152 in step S3.1, the parameters W2 of the two-layer fully connected network in step S3.2, the parameters W3 of the single-layer feedforward network in step S3.3, the parameters W4 of the sentence LSTM network in step S3.4, and the parameters W5 of the word LSTM network in step S3.5;

[0032] Step S3.8, repeat steps S3.1 to S3.7 until the total error L no longer decreases, and save the updated network weight parameters W1, W2, W3, W4 and W5 to the weight file; load the weight file into the feature extraction network ResNet-152, the two-layer fully connected network, the single-layer feedforward network, the sentence LSTM network and the word LSTM network to obtain the image description model.

[0033] Optionally, step S3 also includes step S3.9, which specifically uses the test set data to test the prediction accuracy of the image description model obtained in step S3.8; if the prediction accuracy of the test is not less than 80%, the prediction accuracy of the image description model meets the standard; if the prediction accuracy of the test is less than 80%, it is necessary to repeat steps S3.1 to S3.8 to reconstruct the image description model.

[0034] Optionally, in step S1, the building density and industrial area distribution are negatively correlated with the environmental evaluation, and the corresponding environmental score is recorded as 0 points; the road connectivity, vegetation coverage and leisure activity area distribution are positively correlated with the environmental evaluation, and the corresponding environmental score is recorded as 1 point.

[0035] Optionally, in step S1 , the resolution of each remote sensing image is 0.5 m, and the size of each remote sensing image is 250 pixels×250 pixels.

[0036] Optionally, in step S2, the process of obtaining the environmental score in the environmental assessment dataset of the built-up area is as follows:

[0037] Step S2.1, marking the geographical entities corresponding to each of the remote sensing images in step S1;

[0038] Step S2.2: Label the geographical entity with ground features, and calculate the environmental scores of the ground features using the environmental evaluation index, thereby obtaining the environmental scores of the remote sensing images.

[0039] Optionally, in step S2, the data ratio of the training set and the test set in the environmental assessment dataset is 4:1 or 7:3.

[0040] Optionally, in step S4, the resolution of the remote sensing image of the built-up area to be evaluated is 0.5 m, and the size of the scene image is 250 pixels × 250 pixels.

[0041] The application of the technical solution of the present invention has at least the following beneficial effects:

[0042] (1) Compared with the method in which street view image data can only cover streets, the built-up area environment evaluation method based on remote sensing image scene understanding described in the present invention can achieve coverage of the entire built-up area environment of the city, and with the short update cycle of remote sensing images, it can obtain urban environment evaluation results with good timeliness; the environmental evaluation indicators determined in step S1 of the present invention select 5 representative evaluation indicators, which can realize quantitative analysis of remote sensing images; at the same time, the environmental evaluation indicators can be directly and quickly obtained from remote sensing images, solving the problems of low frequency and long time consumption of obtaining quantitative indicators and indicator weight updates in the existing technology; the image description model obtained by steps S1-S3 can not only utilize the object category information in the remote sensing image, but also utilize the spatial relationship between objects (such as whether the building is surrounded by vegetation) and the attribute information of the objects (such as whether the building density is high or low), that is, the image description model can extract objective scene information; in step S4, the objective evaluation of the built-up area environment to be evaluated in the city is realized by applying the image description model, which has the advantage of good usability.

[0043] (2) In step S5, the present invention uses interpolation processing to process the map obtained in step S4 into an environmental evaluation map of the built-up area to be evaluated with a continuous distribution result, and the result after interpolation processing is more consistent with the actual geographical scene, and the environmental score obtained by the evaluation is more accurate.

[0044] In addition to the above-described objects, features and advantages, the present invention has other objects, features and advantages. The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:

[0046] Figure 1 This is an example of a scene image with a score of 5 in Example 1;

[0047] Figure 2 This is an example of a scene image with a score of 4 in Example 1;

[0048] Figure 3 This is an example of a scene image with a score of 3 in Example 1;

[0049] Figure 4 This is an example of a scene image with a score of 2 in Example 1;

[0050] Figure 5 This is an example of a scene image with a score of 1 in Example 1;

[0051] Figure 6 is an example of a scene image with a score of 0 in Example 1;

[0052] Figure 7 This is an example of the map drawn by steps S1-S4 in Example 1;

[0053] Figure 8 The spatial grid is moved before and after step S5 in embodiment 1 and the spatial grid center point x is assigned. i Schematic diagram of the process;

[0054] Figure 9 The environmental evaluation map of the built-up area to be evaluated obtained by processing steps S1-S5 in Example 1;

[0055] Figure 10 This is an example of the visibility analysis results of Example 1 on the built-up area within the main urban area of Changsha City. DETAILED DESCRIPTION

[0056] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of the present invention.

[0057] Example 1:

[0058] A method for building area environmental assessment based on remote sensing image scene understanding includes the following steps:

[0059] Step S1: Divide the built-up area into a plurality of continuous regions, obtain remote sensing images corresponding to each region of the built-up area, and determine environmental evaluation indicators for each remote sensing image;

[0060] In step S1, the environmental evaluation indicators include building density, road connectivity, vegetation coverage, industrial area distribution, and leisure activity area distribution; therefore, the environmental evaluation indicators involve five evaluation indicators, that is, there are five evaluation description statements when evaluating the remote sensing image;

[0061] Among them, the building density and industrial area distribution are negatively correlated with the environmental evaluation, and the corresponding environmental score is recorded as 0 points; the road connectivity, vegetation coverage and leisure activity area distribution are positively correlated with the environmental evaluation, and the corresponding environmental score is recorded as 1 point;

[0062] Step S2: using the environmental evaluation index to calculate the environmental score corresponding to each of the remote sensing images, and combining the environmental scores to obtain an environmental evaluation dataset of the built-up area; dividing the environmental evaluation dataset into a training set and a test set in proportion;

[0063] Step S3: using the training set to train and obtain an image description model;

[0064] Step S4: obtain a remote sensing image of the built-up area to be evaluated (selected from the built-up area within the main urban area of Changsha City) and cut it into several scene images; use the image description model to predict each of the scene images, and then output a sentence corresponding to the environmental evaluation index; according to step S2, use the environmental evaluation index to calculate the environmental score corresponding to each of the scene images, see Figures 1-6 (For scene images with a score of 5, the quality of the five evaluation indicators is relatively high, as shown by open building layout, high vegetation coverage, good road connectivity, no industrial interference, and space for leisure activities; for scene images with a score of 4, most of them perform well in the four evaluation indicators of building density, vegetation coverage, road connectivity, and industrial area distribution, but fail to meet the standard in the distribution indicator of leisure activity areas; for scene images with a score of 3, their performance is relatively more diverse, with some geographical scenes having insufficient vegetation coverage but relatively open buildings, while some geographical scenes have good vegetation coverage but relatively dense buildings; for scene units with scores of 0, 1, and 2, the building density index performs poorly, and most scenes lack vegetation coverage). The environmental scores of each scene image are plotted into a map for environmental evaluation of the built-up area to be evaluated using ArcGIS software, see Figure 7 .

[0065] The method for built-up area environment assessment based on remote sensing image scene understanding further includes step S5, specifically using interpolation processing to process the map into an environmental assessment map of the built-up area to be assessed with a continuous distribution result.

[0066] See also Figure 8 , the step S5 comprises:

[0067] Step S5.1: Mark the areas corresponding to the scene images as corresponding spatial grids. Move the spatial grid located in the center of all scene images outward along each perimeter of the spatial grid until the moved area covers all scene images. Each movement distance is one-half of the spatial grid range (i.e., 125 pixels).

[0068] Step S5.2: Use the image description model to evaluate the scene image corresponding to the moved spatial grid to obtain the environment score corresponding to each spatial grid; assign the environment score of each spatial grid to the center point x of each spatial grid. i (center point x i The attribute is the environmental score of the spatial environment centered at the point);

[0069] Step S5.3: Use the inverse distance weighted method in ArcGIS to calculate the center point x obtained in step S5.2. i Perform interpolation processing, where the center point x i The inverse distance weight is calculated as follows:

[0070]

[0071] In formula (1), x0 is the position of the estimated center point; x i is the position of the known center point; n is the total number of known center points; i is the i-th known center point; the estimated value X(x0) is the n-th measured value X(x i ) weighted average; ω i For each known center point x i The weight of

[0072] The ω i Calculate using the following formula (2):

[0073] ω i =d 0,i -p Formula (2)

[0074] In formula (2), d 0,i is the estimated center point x0 and the known center point x i The Euclidean distance between them; p is the exponential power parameter (specifically, p = 1);

[0075] Step S5.4: Use ArcGIS mapping software to produce the interpolation results into an environmental assessment map of the built-up area to be evaluated.

[0076] In step S3, the image description model includes an image feature extraction network ResNet-152, a two-layer fully connected network, a single-layer feedforward network, a sentence LSTM network, and a word LSTM network that sequentially process the training set data. The specific processing process is as follows:

[0077] Step S3.1, inputting the corresponding remote sensing image in the training set into the image feature extraction network ResNet-152 to obtain a 14×14×512-dimensional visual feature v;

[0078] Step S3.2: Input the visual feature v into a two-layer fully connected network with a dimension of 512×5 to obtain a predicted identification phrase corresponding to the environmental evaluation index. The predicted identification phrase Converted into 512-dimensional semantic features a m ;

[0079] Step S3.3: The semantic feature a mAnd the visual feature v is input into a 512-dimensional single-layer feedforward network to obtain a 512-dimensional context feature ctx (s) ;

[0080] Step S3.4: transform the 512-dimensional context feature ctx (s) Input into the sentence LSTM network with a hidden layer size of 512 dimensions, and obtain 512-dimensional features t that describe the topic and structure of the sentence (s) ;

[0081] Step S3.5: The feature t (s) Input into the word LSTM network with a hidden layer size of 512 dimensions, and finally generate 5 predicted image description sentences

[0082] Step S3.6: Calculate the total error L, specifically by adding the predicted identification phrase obtained in step S3.2 The ground truth phrase annotated with the corresponding remote sensing image Use cross entropy loss to calculate the error L1; the predicted image description sentence obtained in step S3.5 Real image description sentences annotated with corresponding remote sensing images The error L2 is calculated using cross entropy loss; the total error L is the sum of the error L1 and the error L2, specifically expressed by the following formula (3):

[0083]

[0084] In formula (3), N represents the total number of remote sensing images in the training set; i represents the i-th remote sensing image in the training set; c represents the c-th sentence in each remote sensing image in the training set; represents the true value of the identification phrase corresponding to the cth sentence of the i-th image in the training set; represents the predicted value of the identification phrase corresponding to the cth sentence of the i-th image in the training set; represents the true value of the cth sentence in the i-th image in the training set, represents the predicted value of the cth sentence in the i-th image in the training set;

[0085] Step S3.7: Backpropagate the total error L in step S3.6 and use the Adam algorithm to update the network parameters; the network parameters include the parameters W1 of the image feature extraction network ResNet-152 in step S3.1, the parameters W2 of the two-layer fully connected network in step S3.2, the parameters W3 of the single-layer feedforward network in step S3.3, the parameters W4 of the sentence LSTM network in step S3.4, and the parameters W5 of the word LSTM network in step S3.5;

[0086] Step S3.8, repeat steps S3.1 to S3.7 until the total error L no longer decreases, and save the updated network weight parameters W1, W2, W3, W4 and W5 to a weight file with the suffix .pth; load the weight file into the feature extraction network ResNet-152, two-layer fully connected network, single-layer feedforward network, sentence LSTM network and word LSTM network to obtain the image description model.

[0087] In the step S3, step S3.9 is also included, which specifically uses the test set data to test the prediction accuracy of the image description model obtained in step S3.8; if the prediction accuracy of the test is not less than 80%, the prediction accuracy of the image description model meets the standard; if the prediction accuracy of the test is less than 80%, it is necessary to repeat steps S3.1 to S3.8 to reconstruct the image description model.

[0088] The test process in step S3.9 is as follows:

[0089] First, use the corresponding remote sensing images in the test set to execute steps S3.1 to S3.5 to obtain the five predicted description sentences for each remote sensing image in the test set;

[0090] Secondly, the five description sentences predicted for each remote sensing image in the test set are compared with the five real description sentences of the corresponding remote sensing image. If the sentences are completely consistent after comparison, it means that the image description model has made a correct prediction. If the sentences are not completely consistent after comparison, it means that the image description model has made an incorrect prediction.

[0091] Finally, the number of correct sentences in each remote sensing image in the test set is counted, and the proportion of remote sensing images with different numbers of correct sentences in the entire test set to the total number of remote sensing images in the test set is counted; if the proportion of remote sensing images predicted correctly by all five description sentences is not less than 80% of the total number of remote sensing images in the test set, it indicates that the image description model is well trained and can provide the correct language description of the environmental evaluation indicators for the input remote sensing images.

[0092] In this case, the image description model obtained in step S3.8 can be applied to the remote sensing image of the built-up area to be evaluated to obtain an objective description of the scene image.

[0093] Table 1 shows the test results of the image description model obtained in step S3.8 on the test set. In Table 1, i represents the number of correct sentences in the five description sentences of each remote sensing image in the test set; i It represents the proportion of remote sensing images with i correct sentences in the test set to the total number of remote sensing images in the test set.

[0094] Table 1 Test results of the image description model obtained in step S3.8 on the test set

[0095] i <![CDATA[Proportion i ]]> 5 85.69% 4 12.52% 3 1.46% 2 0.33% 1 0 0 0

[0096] The statistical results in Table 1 show that the image description model accurately predicted 85.69% of remote sensing images (i.e., all five description statements were accurately predicted); 12.52% of remote sensing images had one statement incorrectly predicted (i.e., all four statements were accurately predicted); and 1.46% of remote sensing images had two statement incorrectly predicted (i.e., all three statements were accurately predicted). None of the remote sensing images had a prediction error exceeding two statements. This analysis shows that the image description model achieved a prediction accuracy of 85.69%.

[0097] In step S1 , the resolution of each remote sensing image is 0.5 m, and the size of each remote sensing image is 250 pixels×250 pixels.

[0098] In step S2, the process of obtaining the environmental score in the environmental assessment dataset of the built-up area is as follows:

[0099] Step S2.1, marking the geographical entities corresponding to each of the remote sensing images in step S1;

[0100] Step S2.2: Label the geographical entity with ground features, and calculate the environmental scores of the ground features using the environmental evaluation index, thereby obtaining the environmental scores of the remote sensing images.

[0101] In step S2, the data ratio of the training set and the test set in the environmental assessment dataset is 4:1.

[0102] In step S4, the resolution of the remote sensing image of the built-up area to be evaluated is 0.5m, and the size of the scene image is 250 pixels×250 pixels.

[0103] See also Figure 7 In Example 1, the map drawn by steps S1-S4 shows that the environmental scores of the built-up areas to be evaluated are separated by scene images. However, the fact is that the scene environment in reality is continuous, so the environmental scores of the scene environment should also be continuous. In order to make the environmental scores of the scene environment geographically continuous, step S5 is also set in Example 1, see Figure 9 , that is, interpolation processing is used to process the map into an environmental assessment map of the built-up area to be evaluated with continuous distribution results.

[0104] Depend on Figure 7 and Figure 9By comparison, after interpolation in step S5, the environmental scores are no longer discretely distributed across the map by spatial grids, but are now continuously distributed. Furthermore, the environmental scores are not discrete values of 0, 1, 2, 3, 4, and 5, but real numbers within the range [0, 5]. The interpolated results are more consistent with the actual geographic scene and are more accurate.

[0105] Figure 10 This is an example of the visibility analysis results of Example 1 on the built-up area within the main urban area of Changsha City. Figure 10 The darker areas in the circle in (2) indicate higher environmental scores. Figure 10 The light-colored areas in the circle in (4) indicate lower environmental scores. Figure 10 (2) Corresponding Figure 10 (1) In the middle, the buildings are neatly arranged, the vegetation is adequate, and there are lakes and playgrounds, which gives a higher environmental score; Figure 10 (4) Corresponding Figure 10 (3) The buildings are densely packed, there is a lack of vegetation cover, and the environmental score is low.

[0106] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A method for building area environmental assessment based on remote sensing image scene understanding, characterized in that: The following steps are involved: Step S1: Divide the built-up area into a plurality of continuous regions, obtain remote sensing images corresponding to each region of the built-up area, and determine environmental evaluation indicators for each remote sensing image; the environmental evaluation indicators include building density, road connectivity, vegetation cover, industrial area distribution, and leisure activity area distribution; Step S2: using the environmental evaluation index to calculate the environmental score corresponding to each of the remote sensing images, and combining the environmental scores to obtain an environmental evaluation dataset of the built-up area; dividing the environmental evaluation dataset into a training set and a test set in proportion; Step S3: using the training set to train and obtain an image description model; Step S4: obtaining a remote sensing image of the built-up area to be evaluated and cropping it into a plurality of scene images; using the image description model to predict each of the scene images, and then outputting a sentence corresponding to the environmental evaluation index; according to step S2, using the environmental evaluation index to calculate the environmental score corresponding to each of the scene images; using ArcGIS software to plot the environmental score of each of the scene images into a map for environmental evaluation of the built-up area to be evaluated; The method further includes step S5, specifically, using interpolation processing to process the map into an environmental assessment map of the built-up area to be assessed with a continuous distribution result; The step S5 comprises: Step S5.1: Mark the areas corresponding to the scene images as corresponding spatial grids, and move the spatial grid located in the center of all scene images outward along each perimeter of the spatial grid until the moved area covers all scene images, wherein each movement distance is one-half of the spatial grid range; Step S5.2: Use the image description model to evaluate the scene image corresponding to the moved spatial grid to obtain the environment score corresponding to each spatial grid; assign the environment score of each spatial grid to the center point x of each spatial grid. i ; Step S5.3: Use the inverse distance weighted method in ArcGIS to calculate the center point x obtained in step S5.

2. i Perform interpolation processing, where the center point x i The inverse distance weight is calculated as follows: In formula (1), x0 is the position of the estimated center point; x i is the position of the known center point; n is the total number of known center points; i is the i-th known center point; the estimated value X(x0) is the n-th measured value X(x i ) weighted average; ω i For each known center point x i The weight of The ω i Calculate using the following formula (2): ω i =d 0,i -p Formula (2) In formula (2), d 0,i is the estimated center point x0 and the known center point x i The Euclidean distance between them; p is the exponential power parameter; Step S5.4: Use ArcGIS mapping software to produce the interpolation results into an environmental assessment map of the built-up area to be evaluated.

2. The method for building area environment assessment based on remote sensing image scene understanding according to claim 1 is characterized in that: In step S3, the image description model includes an image feature extraction network ResNet-152, a two-layer fully connected network, a single-layer feedforward network, a sentence LSTM network, and a word LSTM network that sequentially process the training set data. The specific processing process is as follows: Step S3.1, inputting the corresponding remote sensing image in the training set into the image feature extraction network ResNet-152 to obtain the visual feature v; Step S3.2: Input the visual feature v into a two-layer fully connected network to obtain a predicted identification phrase corresponding to the environmental evaluation index. The predicted identification phrase Converted into semantic features m ; Step S3.3: The semantic feature a m And the visual feature v is input into a single-layer feedforward network to obtain the context feature ctx (s) ; Step S3.4: The context feature ctx (s) Input into the sentence LSTM network to obtain feature t (s) ; Step S3.5: The feature t (s) Input into the word LSTM network, and finally generate the predicted image description sentence Step S3.6: Calculate the total error L, specifically by adding the predicted identification phrase obtained in step S3.2 The ground truth phrase annotated with the corresponding remote sensing image Use cross entropy loss to calculate the error L1; the predicted image description sentence obtained in step S3.5 Real image description sentences annotated with corresponding remote sensing images The error L2 is calculated using cross entropy loss; the total error L is the sum of the error L1 and the error L2, specifically expressed by the following formula (3): In formula (3), N represents the total number of remote sensing images in the training set; i represents the i-th remote sensing image in the training set; c represents the c-th sentence in each remote sensing image in the training set; represents the true value of the identification phrase corresponding to the cth sentence of the i-th image in the training set; represents the predicted value of the identification phrase corresponding to the cth sentence of the i-th image in the training set; represents the true value of the cth sentence in the i-th image in the training set, represents the predicted value of the cth sentence in the i-th image in the training set; Step S3.7: Backpropagate the total error L in step S3.6 and use the Adam algorithm to update the network parameters; the network parameters include the parameters W1 of the image feature extraction network ResNet-152 in step S3.1, the parameters W2 of the two-layer fully connected network in step S3.2, the parameters W3 of the single-layer feedforward network in step S3.3, the parameters W4 of the sentence LSTM network in step S3.4, and the parameters W5 of the word LSTM network in step S3.5; Step S3.8, repeat steps S3.1 to S3.7 until the total error L no longer decreases, and save the updated network weight parameters W1, W2, W3, W4 and W5 to the weight file; load the weight file into the feature extraction network ResNet-152, the two-layer fully connected network, the single-layer feedforward network, the sentence LSTM network and the word LSTM network to obtain the image description model.

3. The method for building area environment assessment based on remote sensing image scene understanding according to claim 2, characterized in that: In the step S3, step S3.9 is also included, which specifically uses the test set data to test the prediction accuracy of the image description model obtained in step S3.8; if the prediction accuracy of the test is not less than 80%, the prediction accuracy of the image description model meets the standard; if the prediction accuracy of the test is less than 80%, it is necessary to repeat steps S3.1 to S3.8 to reconstruct the image description model.

4. The method for building area environment assessment based on remote sensing image scene understanding according to claim 1, characterized in that: In step S1, the building density and industrial area distribution are negatively correlated with the environmental evaluation, and the corresponding environmental score is recorded as 0 points; the road connectivity, vegetation coverage and leisure activity area distribution are positively correlated with the environmental evaluation, and the corresponding environmental score is recorded as 1 point.

5. The method for building area environment assessment based on remote sensing image scene understanding according to claim 1, characterized in that: In step S1 , the resolution of each remote sensing image is 0.5 m, and the size of each remote sensing image is 250 pixels×250 pixels.

6. The method for building area environment assessment based on remote sensing image scene understanding according to claim 1, characterized in that: In step S2, the process of obtaining the environmental score in the environmental assessment dataset of the built-up area is as follows: Step S2.1, marking the geographical entities corresponding to each of the remote sensing images in step S1; Step S2.2: Label the geographical entity with ground features, and calculate the environmental scores of the ground features using the environmental evaluation index, thereby obtaining the environmental scores of the remote sensing images.

7. The method for building area environment assessment based on remote sensing image scene understanding according to claim 1, characterized in that: In step S2, the data ratio of the training set and the test set in the environmental assessment dataset is 4:1 or 7:

3.

8. The method for building area environment assessment based on remote sensing image scene understanding according to claim 1, characterized in that: In step S4, the resolution of the remote sensing image of the built-up area to be evaluated is 0.5m, and the size of the scene image is 250 pixels×250 pixels.