Urban block change detection method, system, device and storage medium

By analyzing time-series satellite images of urban blocks, constructing time-series segments and calculating semantic coherence scores, the problem of low detection accuracy caused by light and cloud interference in existing technologies is solved, and more accurate block change monitoring is achieved.

CN120451797BActive Publication Date: 2025-09-19CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510905181.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-09-19
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

Existing urban block change detection methods rely on spectral and texture features and are easily affected by light and cloud cover, resulting in low detection accuracy and inability to accurately monitor urban block changes.

Method used

By analyzing the land cover type characteristics of multiple time-series satellite images, constructing time-series segments, calculating semantic coherence scores, and using the ResNet50 backbone network and computational model, we can reduce shallow feature errors and determine the boundaries and types of changes.

Benefits of technology

It improves the accuracy of change detection in urban blocks, can ignore errors caused by illumination and cloud cover, provides more comprehensive change analysis, and supports timely updates of urban planning and traffic navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451797B_ABST
    Figure CN120451797B_ABST
Patent Text Reader

Abstract

The present application provides a method, system, device and storage medium for detecting changes in urban blocks, relating to the field of image processing technology. The method comprises: obtaining N time-series satellite images of a target block, where N≥3 and one satellite image is used as a slice; constructing a time-series segment with the i-th slice as the center and a length of k, where 3≤k≤N; calculating the semantic coherence score of the i-th slice based on the temporal correlation between multiple slices in the time-series segment, where the semantic coherence score is used to reflect the degree of change of the target block at the time point corresponding to the i-th slice; and obtaining a detection result based on the semantic coherence score. The present application determines the change boundary and change type of the target block when it changes by analyzing the semantic coherence of the corresponding satellite images of the target block within the time period covered by the time-series segment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method, system, device and storage medium for detecting changes in urban blocks. Background Art

[0002] With the advancement of urban infrastructure, urban renewal has significantly accelerated, significantly shortening the construction cycle of urban blocks. However, if these dynamic changes in urban blocks are not detected and updated in a timely and accurate manner, they will impact a variety of areas, including traffic navigation, urban planning, and environmental monitoring. For example, outdated block information can cause navigation systems to provide incorrect routes, increasing traffic congestion and safety hazards. In this context, remote sensing technology, with its advantages of large-scale and multi-temporal observation, provides a new technical approach for monitoring urban block changes.

[0003] Currently, the acquisition of massive multi-temporal remote sensing data provides a solid data foundation for block change detection research. Significant progress has been made in using remote sensing data to monitor urban development, particularly in time series change detection and analysis methods. For example, the following two detection methods extend the change detection method of dual-temporal imagery to time series analysis, and the other employs a single-temporal change detection method to determine changes in urban blocks by continuously monitoring the shallow information contained in the image. However, these methods still have significant limitations. They primarily rely on shallow information such as spectral and texture features. This shallow information is susceptible to changes in lighting conditions and cloud cover. Both varying light intensities and cloud cover can distort the true spectral curves and texture characteristics of urban blocks, thereby reducing the accuracy of urban block change detection.

[0004] In view of this, the present invention is proposed. Summary of the Invention

[0005] To address the above technical issues, this application provides a method, system, device, and storage medium for detecting changes in urban blocks. These methods are primarily used to analyze the semantic coherence of remote sensing images corresponding to urban blocks over a long time period to determine whether the urban blocks have changed, and what the change boundaries and change types are. The technical solution is as follows:

[0006] In a first aspect, a method for detecting changes in urban blocks is provided, comprising:

[0007] Obtain N time-series satellite images of the target block, N ≥ 3, with one satellite image as a slice;

[0008] Construct a time series segment centered at the i-th slice and of length k, 3≤k≤N;

[0009] Calculate the semantic coherence score of the i-th slice based on the temporal correlation between the multiple slices in the temporal segment, where the semantic coherence score is used to reflect the degree of change of the target block at the time point corresponding to the i-th slice;

[0010] A detection result is obtained according to the semantic coherence score.

[0011] In a possible implementation, calculating the semantic coherence score of the i-th slice includes:

[0012] Extracting land cover type features of each slice in the time series segment, wherein the land cover type features at least include stable type features, structural change features, and functional transformation features;

[0013] Comparing the stable type features, the structural change features, and the functional transition features of the plurality of slices to obtain temporal correlations between the plurality of slices;

[0014] The semantic coherence score is obtained according to the temporal correlation.

[0015] In one possible implementation, the structural change characteristics include changes in impervious surface coverage caused by expansion and / or demolition of the target block.

[0016] In a possible implementation, the semantic coherence score is calculated using the following formula:

[0017]

[0018] in, represents the semantic coherence score of the i-th slice, is the activation function used to constrain the output semantic coherence score, Indicates the temporal correlation of multiple slices in the temporal segment where the i-th slice is located. and They respectively represent the weight matrix and bias term of a preset calculation model, where the calculation model is used to calculate the semantic coherence score of the i-th slice.

[0019] In a possible implementation, before obtaining the semantic coherence score, the method further includes: performing convolution, maximum pooling, and flattening operations on the land cover type features of the plurality of slices.

[0020] In a possible implementation, the semantic coherence score is input into a preset classification model to obtain a detection result, where the detection result includes a change boundary and a change type, where the change type is any one of expansion, demolition, and unchanged of the target block.

[0021] In a possible implementation, when the change type is that the target block remains unchanged, the method further includes:

[0022] Along the forward direction of the time axis of N slices, one or more slices are moved with a window of length k to obtain a new time sequence segment.

[0023] In a second aspect, a city block change detection system is provided, comprising:

[0024] The data acquisition module is used to obtain N time-series satellite images of the target block, where N is greater than or equal to 3, and one satellite image is used as a slice.

[0025] The data processing module is suitable for constructing a time series segment centered on the i-th slice and of length k, 3≤k≤N;

[0026] a data calculation module adapted to calculate a semantic coherence score of the i-th slice based on the temporal correlation between the multiple slices in the temporal segment, wherein the semantic coherence score is used to reflect the degree of change of the target block at the time point corresponding to the i-th slice;

[0027] The data generation module is adapted to obtain a detection result according to the semantic coherence score.

[0028] In a third aspect, an electronic device is provided, comprising a processor and a memory, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute any of the above-mentioned urban block change detection methods.

[0029] In a fourth aspect, a storage medium is provided, wherein the storage medium stores a computer program, wherein the computer program is configured to execute any of the above-mentioned urban block change detection methods when running.

[0030] The technical solutions provided in the embodiments of the present application can achieve the following technical effects:

[0031] (1) This application selects some satellite images from N time-series satellite images as time-series segments, and then extracts the land cover type features of each slice in the time-series segment. By comparing and analyzing the land cover type features of multiple slices, the semantic coherence of the multiple slices is obtained, that is, the degree of change of the target block at the time point corresponding to the slice in the time-series segment is obtained, and then the degree of change is used as the basis for judging whether the target block has changed. It can be seen that in the analysis process, the application uses the change trend and dependency relationship of the land cover type features of multiple slices over time, that is, after the time point of the change, the land cover type features also reflect the change. Therefore, the application can ignore the errors caused by shallow features (such as changes in light intensity and cloud cover), and have a more comprehensive understanding of the status of the target block at different time points, ensuring the accuracy of the semantic coherence of the multiple slices obtained by analysis.

[0032] (2) In addition to using multiple slices with a temporal relationship as time series segments to facilitate the analysis of the semantic coherence of multiple slices in a time series segment, this application also provides multiple models to provide technical support for obtaining accurate monitoring results through the collaborative work between models. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application.

[0034] Figure 1 This is a flow chart of a method for detecting changes in urban blocks according to an embodiment of the present application.

[0035] Figure 2 This is an example diagram of generating timing segments in an embodiment of the method of the present application.

[0036] Figure 3 It is a structural diagram of the backbone network of the feature extraction model of the embodiment of the method of the present application.

[0037] Figure 4 It is a structural diagram of the feature extraction model and calculation model of the embodiment of the method of the present application.

[0038] Figure 5 This is a block diagram of a city block change detection system according to an embodiment of the present application.

[0039] Figure 6 This is a structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0040] The following describes exemplary embodiments of the present application in more detail with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0041] It should be noted that the terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that such usage is interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the term "including" and its variations are to be interpreted as open-ended terms meaning "including but not limited to."

[0042] Currently, since both the single-phase change detection method and the dual-phase change detection method analyze shallow features such as texture and spectrum of satellite images, neither of them can ignore the detection errors caused by changes in light intensity and cloud occlusion. Of course, some scholars have also proposed an arbitrary boundary detection framework at the shot level, event level, and scene level. This framework is based on a deep learning model and uses ResNet50 in the residual network as the backbone network to densely sample video frames, while calculating the semantic coherence scores of the video frames. When using this method, in order to reduce the training time and inference time of the deep learning model, latent queries are used to compress the video, reducing the computational time complexity from high-dimensional to low-dimensional. Reduced to low dimension However, the disadvantage is that the deep learning model compression method is only applicable to long videos, because long videos have many basic units (video frames). For example, a 45-minute video often has no less than 60,000 video frames. It is not applicable to the small amount of satellite images obtained using satellite remote sensing technology. Moreover, this method can only detect the boundary of changes, but cannot detect the type of change boundary, that is, it cannot determine the type of change.

[0043] In order to solve the above problems, this application provides a method for detecting changes in urban blocks, such as Figure 1 As shown, the method may include the following steps S101 to S104:

[0044] Step S101: Obtain N time-series satellite images of the target block, where N≥3, and one satellite image is used as a slice.

[0045] The target block is the urban block to be analyzed for changes. In real-world scenarios, urban block changes can occur due to expansion, demolition, and functional conversion. Functional conversion refers to a shift in land use, such as converting vacant land into residential areas, or converting residential areas back into vacant land. Since converting vacant land into residential areas is considered urban block expansion, while converting residential areas into vacant land is considered urban block demolition, this application primarily analyzes changes caused by urban block expansion and urban block demolition.

[0046] For the target neighborhood described above, remote sensing technology is used to capture multiple satellite images of the target neighborhood at different time points, such as at intervals of ten days, fifteen days, or one month. The acquisition time points can be determined based on the timeliness of changes and updates to the target neighborhood. For example, the higher the timeliness of changes and updates to the target neighborhood, the shorter the intervals between acquisition time points. In this embodiment, to conserve satellite imagery resources, the target neighborhood is captured once at any time on a randomly selected day within each month, resulting in 12 satellite images in a single year. In this embodiment, the dimensions of the captured satellite images include height, width, and number of channels. The height and width corresponding to different neighborhoods may vary. However, for ease of processing, the captured satellite images are resized to have the same height and width. For example, both the height and width of the satellite images are reshaped to a standard 64×64 image with 10 channels.

[0047] Based on the multiple satellite images obtained, they are arranged in the order of their acquisition time points to obtain multiple satellite images with a time-series relationship. These multiple satellite images obtained from N acquisitions with a time-series relationship are used as N time-series satellite images of the target neighborhood. Furthermore, to facilitate subsequent analysis of any one of the N time-series satellite images, each satellite image is treated as a slice in the time-series. Therefore, the N time-series satellite images contain N slices.

[0048] Step S102: construct a time sequence segment with the i-th slice as the center and a length of k, where 3≤k≤N.

[0049] Based on the N slices obtained, one of the N slices is selected, denoted as the i-th slice. With the i-th slice as the center, multiple slices of length k are selected to form a time sequence segment. Since the time sequence segment is obtained by selecting multiple slices with the i-th slice as the center, the time sequence segment includes at least three slices: the i-1th slice, the i-th slice, and the i+1th slice, i.e., k ≥ 3.

[0050] It should be noted that the time series segment contains at least three slices in order to combine three consecutive slices and analyze the consistency of the i-th slice with the i-1th slice and the i+1th slice respectively. That is, through the consistency analysis of the three time series slices, the dependence on the shallow features appearing on a single slice is reduced. For example, if the target block undergoes urban expansion changes, the slices collected after the time point of the change will reflect the change and have consistency of the change. However, if there is cloud occlusion, multiple time series slices will not all have the same cloud occlusion, and there will always be a certain deviation. Therefore, the present application analyzes the consistency of at least three slices with a time series relationship to reduce the detection error. In this embodiment, the consistency of multiple slices can also be expressed by semantic coherence.

[0051] In order to ensure the accuracy of detection, the length of the time sequence segment of this application is set to 5, that is, k=5, so as to analyze the semantic coherence between the i-th slice and the two previous slices and the two following slices on the time axis where N slices are located. Figure 2 As shown, assuming that there are 12 slices, a segment containing 5 slices is taken as a time sequence segment, and the time sequence segment is represented by the portion selected by the solid line frame.

[0052] Step S103 , based on the temporal correlation between multiple slices in the time sequence segment, the semantic coherence score of the i-th slice is calculated. The semantic coherence score is used to reflect the degree of change of the target block at the time point corresponding to the i-th slice.

[0053] First, call the pre-trained feature extraction model and calculation model. Among them, the feature extraction model mainly uses ResNet50 as the backbone network, such as Figure 3 As shown, the ResNet50 adopted in this embodiment mainly includes three convolutional layers, and the three convolutional layers are connected in series in sequence. Among them, after the first convolutional layer receives the time series segment, it performs a dimensionality reduction operation on each slice in the time series segment to reduce the number of channels of the model and reduce the amount of calculation. The second convolutional layer mainly performs feature extraction. After the second convolutional layer extracts the features, the third convolutional layer performs a channel number recovery operation to adjust the number of channels of the data output by the feature extraction model to the number of channels before the slice enters the model to ensure the integrity of the output data. In addition, the feature extraction model is also configured with the function of splicing feature vectors, so that after extracting the features of each slice, the extracted features are spliced ​​to ensure the integrity of the data output by the feature extraction model. As Figure 4 As shown in the figure, the time series segments entering the feature extraction model are first extracted by ResNet50 to extract features, which are represented by long bars. Based on the obtained features, the features are spliced ​​to obtain semantic splicing vectors, and the semantic splicing vectors are output to the calculation model.

[0054] like Figure 4 As shown in the figure, the computational model consists of a series of convolutional layers, pooling layers, flattening layers, and fully connected layers. The convolutional layers are temporal convolutional layers, which are used to extract local features of the semantic splicing vector, such as splicing edges. In practical applications, a long short-term memory (LSTM) network can be used to replace this convolutional layer to perform the corresponding tasks. The pooling layer is mainly used to implement maximum pooling, which is to extract the maximum value in the local features so that the change boundary can be determined based on this maximum value. For example, when the maximum value of a local feature exceeds a preset threshold, the location of the largest local feature is used as the change boundary. The flattening layer is used to convert the largest local feature output by the pooling layer into a one-dimensional vector, which serves as the input to the fully connected layer. The fully connected layer then calculates the semantic coherence score of the temporal segment based on the input of the flattening layer.

[0055] In order to make the feature extraction model and the computing model work together better, a transposition layer is set between the feature extraction model and the computing model. The transposition layer is used to adjust the dimension of the data output by the feature extraction model to match the input dimension supported by the convolution layer in the computing model.

[0056] It should be noted that the models and various functional layers involved above can be adaptively increased or decreased according to actual needs, and this application does not impose any restrictions. At the same time, the parameters of each model and functional layer can be determined during the training phase, or dynamically adjusted during actual use based on the parameters determined during the training phase, and this application does not impose any restrictions.

[0057] The aforementioned feature extraction model and calculation model are stored in the current execution entity, which represents the entity executing the detection method of this application, such as a processor, server, terminal, component, etc. The feature extraction model and calculation model can also be stored in an external memory. When the execution entity needs to execute the detection method of this application, the feature extraction model and calculation model are retrieved and applied.

[0058] Based on the retrieved feature extraction model, the feature extraction model extracts the land cover type features of each slice in the time series segment. The land cover type features include at least stable type features, structural change features and functional transformation features. Among them, stable type features refer to the features of land cover types that remain unchanged for a long time, such as features of road surfaces, water bodies and other types. Structural change features refer to changes in the coverage rate of impervious surfaces caused by the construction or demolition of urban blocks. Impervious surfaces include hardened surfaces such as building roofs, parking lots, squares, etc. If the hardened surfaces change, the coverage rate of impervious surfaces will also change. Functional transformation features refer to the features of land use type changes. For example, when the vacant land and residential area in step S101 are converted to each other, the changed vacant land or residential area is used as the functional transformation feature. The above-mentioned step S101 also illustrates that the functional transformation is essentially a structural change. Therefore, the feature extraction model of the present application extracts the stable type features and structural change features of each slice in the time series segment.

[0059] Specifically, if you use represents the i-th slice, then the time series segment X is obtained:

[0060] (1)

[0061] After the time series segment X is input into the feature extraction model, the feature extraction model uses ResNet50 to extract features from each slice in the time series segment X and obtains the feature vector corresponding to each slice, specifically:

[0062] (2)

[0063] represents the feature vector of the i-th slice, Represents the land cover type features extracted by ResNet50, Represents the parameter matrix of ResNet50, Represents the i-th slice.

[0064] From the above formula, we can see that in the calculation When the corresponding eigenvector is added itself, that is, " ", on the one hand, is to avoid the problem of information loss after extracting land cover type features, so that the feature extraction model can at least retain On the other hand, it can also alleviate the gradient vanishing problem, promote the propagation of information along the gradient direction, and ensure the accuracy of the extracted land cover type features.

[0065] It should be noted that when calculating the feature vector of any slice in the time series segment, it is necessary to calculate it in the same way as the above formula (2), so this embodiment will not describe them one by one.

[0066] Based on the feature vector corresponding to each slice in the time series segment, the feature vectors obtained by splicing are obtained to obtain the semantic splicing vector :

[0067] (3)

[0068] It can be seen that after the time series segment X is input into the feature extraction model, the feature extraction model extracts the land cover type features of each slice, and performs vector mapping and vector splicing on the land cover type features, and outputs the semantic splicing vector .

[0069] The computational model receives the semantic concatenation vector , and based on the semantic splicing vector By analyzing the temporal correlation of multiple slices, it can be considered that the computational model obtains the temporal correlation between multiple slices by comparing the stable type features and structural change features of multiple slices, and then obtains the semantic coherence score based on the temporal correlation.

[0070] Specifically, the convolution layer in the computational model first performs a convolution operation on the semantic splicing vector. Assuming that the temporal convolution function of the convolution layer is Conv1D, we get:

[0071] (4)

[0072] Then, the convolutional semantic concatenation vector is subjected to maximum pooling through the pooling layer to obtain:

[0073] (5)

[0074] Maxpool represents the maximum pooling function of the pooling layer. is the semantic concatenation vector after convolution.

[0075] Furthermore, the semantic splicing vector after maximum pooling is flattened through the flattening layer to obtain a one-dimensional semantic splicing vector:

[0076] (6)

[0077] Flatten represents the flattening function of the flattening layer, F pool Represents the semantic concatenation vector after maximum pooling. The flattened semantic concatenation vector reflects the semantic information of each slice in one dimension, such as the area of ​​residential areas, water areas, etc., so the data obtained after flattening It covers: the relationship between multiple slices in the time series segment X and the time period covered by the time series segment X. For example, if the residential area of ​​one slice increases, the residential areas of subsequent slices will also increase by the same area. Therefore, the data obtained after flattening is Used to represent the temporal correlation of multiple slices.

[0078] Finally, the fully connected layer is based on the flattened data , calculate the semantic coherence score of the i-th slice :

[0079] (7)

[0080] FC represents the calculation function of the fully connected layer, which mainly involves the weight matrix and bias term. The weight matrix and bias term have been determined during training, so formula (7) can be converted into:

[0081] (8)

[0082] in, It is an activation function. The activation function is used to introduce nonlinearity so that the computational model can learn complex computational processes. At the same time, the activation function is also used to constrain the output semantic coherence score so that the output semantic coherence score falls within a preset range for subsequent use. Indicates the temporal correlation of multiple slices in the temporal segment where the i-th slice is located, that is, , and Represent the weight matrix and bias term of the fully connected layer respectively.

[0083] In summary, the convolutional layer in the computational model obtains local features of the semantic splicing vector through convolution operations, such as splicing edges and splicing textures, and then undergoes processing through the pooling layer and flattening layer to obtain the temporal correlation of multiple slices in the time series segment. Finally, the fully connected layer performs further processing based on the temporal correlation of multiple slices. Each neuron in the fully connected layer is connected to all neurons in the previous layer. Through matrix multiplication and addition of bias terms, the weights of different features are adjusted, so that various features can be comprehensively considered to calculate the semantic coherence score. This semantic coherence score reflects the similarity or coherence of the target block at different time points. Since the fully connected layer can capture the changing trend and dependency of the feature vector over time during the entire calculation process, that is, the feature vector also reflects the change after the time point of the change, the present application can ignore the errors caused by shallow features (such as changes in light intensity and cloud occlusion), have a more comprehensive understanding of the state of the target block at different time points, and ensure the accuracy of the calculated semantic coherence score.

[0084] Therefore, a preset score can be set in advance. The preset score can be calculated through a limited number of experiments. When the semantic coherence score is greater than the preset score, it means that the target block has changed, such as the above , If it is a preset score, it means that the target block has changed at the time point corresponding to the i-th slice, and the i-th slice serves as the change boundary.

[0085] In one possible implementation, when the semantic coherence score is greater than a preset score, the process proceeds to step S104, i.e., after determining that the target block has changed, the process proceeds to step S104 to analyze the specific cause of the change and obtain the type of change. When the semantic coherence score is less than or equal to the preset score, it indicates that the target block has not changed. In this case, a new time segment can be selected to analyze whether the target block has changed within the new time segment, thereby reducing repeated calculations and lowering monitoring costs.

[0086] Specifically, the process of selecting a new time series segment includes: first, determining the forward direction of the time axis of N slices, where the forward direction refers to the direction from the past time point to the current time point. Then, determining the position of the current time series segment, based on the position of the current time series segment, moving one or more slices along the forward direction of the time axis of N slices with a window length of k to obtain a new time series segment, such as Figure 2 In the example, the new time series segment is a window of length 5, which is obtained by shifting one slice along the positive direction of the time axis of 12 slices.

[0087] In another possible implementation, after obtaining the semantic coherence score, the process directly proceeds to step S104, in which the change type of the target block is analyzed. Specifically, when the semantic coherence score is less than or equal to the preset score, if the semantic coherence score continues to be input into step S104, the change type obtained is unchanged, that is, the target block has not changed.

[0088] In practical applications, after obtaining the semantic coherence score, a specific processing flow can be selected as needed, and this application does not impose any restrictions.

[0089] Step S104: obtaining a detection result according to the semantic coherence score.

[0090] First, retrieve a pre-trained classification model. The classification model primarily uses the Softmax function, which is used to convert semantic coherence scores into probability distribution vectors representing the probabilities of each category. In this embodiment, the categories in the classification model include target block expansion, demolition, and unchanged. Therefore, during the training phase, the training samples are used to train the classification model to distinguish between the three categories of target block expansion, demolition, and unchanged. After training, the classification model is saved in the execution entity or stored in an external memory, which is not limited by this application.

[0091] Based on the retrieved classification model, the semantic coherence score is processed by the classification model. The processing process includes: First, the semantic coherence score is obtained. The semantic coherence score is an output vector containing multiple elements. Each element in the output vector corresponds to a type of change in the target block, such as the addition of buildings, the reduction of vegetation, or the expansion of a road. Then, the output vector is converted into a probability distribution vector using the Softmax function. Each element in the vector represents the probability of the corresponding change type occurring in the target block.

[0092] Based on the probability of occurrence of the change type of the target block output by the classification model, the change type with the greatest probability is selected as the change type of the target block.

[0093] In one possible implementation, when the target block's change type is expansion or demolition, the change type and change boundary of the target block expansion or demolition are used as detection results, and the detection results are output to facilitate relevant management personnel to understand the change boundary and change type of the target block, thereby timely updating the target block's changes. When the target block's change type is unchanged, since the target block has not changed, only the unchanged target block is used as the detection result.

[0094] In another possible implementation, when the target block change type is any of the following: expansion, demolition, or unchanged, the target block change type and change boundary are output as the detection result, thereby unifying the output format of the detection results. Specifically, when the target block change type is unchanged, the change boundary in the detection result is a space or other value to indicate that the change boundary in the detection result does not exist.

[0095] It should be noted that the order of execution of the steps in the above embodiments does not necessarily imply a specific order of execution. The order of execution of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. In practical applications, all possible implementation methods described above can be combined in any manner to form possible embodiments of the present application, and will not be described in detail here.

[0096] Based on the urban block change detection methods provided in the above embodiments, and based on the same inventive concept, the embodiments of the present application also provide an urban block change detection system.

[0097] Figure 5 This is a structural diagram of the urban block change detection system provided by the embodiment of the present application. Figure 5 As shown, the system 500 may specifically include a data acquisition module 501 , a data processing module 502 , a data calculation module 503 and a data generation module 504 .

[0098] The data acquisition module 501 is adapted to acquire N time-series satellite images of a target block, where N≥3, and one satellite image is regarded as a slice.

[0099] The data processing module 502 is adapted to construct a time sequence segment centered at the i-th slice and having a length of k, where 3≤k≤N.

[0100] The data calculation module 503 is adapted to calculate the semantic coherence score of the i-th slice based on the temporal correlation between multiple slices in the temporal segment. The semantic coherence score is used to reflect the degree of change of the target block at the time point corresponding to the i-th slice.

[0101] The data generation module 504 is adapted to obtain a detection result according to the semantic coherence score.

[0102] In a possible implementation, the data calculation module 503 includes a feature extraction model and a calculation model. The data calculation module obtains the semantic coherence score of the i-th slice in the time sequence segment through the feature extraction model and the calculation model.

[0103] In one possible implementation, the data generation module 504 includes a classification model, based on which the change type of the target block is obtained, and the detection results are output based on the change type and / or semantic coherence score, so that relevant management personnel can know the actual situation of the target block.

[0104] Based on the same inventive concept, an embodiment of the present application also provides an electronic device, including a processor and a memory, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the urban block change detection method of any one of the above embodiments.

[0105] In an exemplary embodiment, an electronic device is provided, such as Figure 6 As shown, Figure 6The electronic device 600 shown includes a processor 601 and a memory 603. The processor 601 and the memory 603 are connected, for example, via a bus 602. Optionally, the electronic device 600 may further include a transceiver 604. It should be noted that in actual applications, the number of transceivers 604 is not limited to one, and the structure of the electronic device 600 does not constitute a limitation on the embodiments of the present application.

[0106] Processor 601 may be a CPU (Central Processing Unit), a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 601 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, or a combination of a DSP and a microprocessor.

[0107] Bus 602 may include a path for transmitting information between the above components. Bus 602 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. Bus 602 may be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0108] The memory 603 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0109] The memory 603 is used to store computer program codes for executing the solution of the present application, and the execution is controlled by the processor 601. The processor 601 is used to execute the computer program codes stored in the memory 603 to implement the contents shown in the above method embodiment.

[0110] Among them, electronic devices include but are not limited to: mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0111] Based on the same inventive concept, an embodiment of the present application further provides a storage medium storing a computer program, wherein the computer program is configured to execute the urban block change detection method of any one of the above embodiments when running.

[0112] Those skilled in the art will clearly understand that the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the aforementioned method embodiments, and for the sake of brevity, they will not be further described here.

[0113] Those skilled in the art will understand that the technical solution of the present application, in essence, or in whole or in part, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of program instructions that cause an electronic device (such as a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application when the program instructions are executed. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0114] Alternatively, all or part of the steps of implementing the aforementioned method embodiments may be accomplished by hardware related to program instructions (such as electronic devices such as personal computers, servers, or network devices), and the program instructions may be stored in a computer-readable storage medium. When the program instructions are executed by a processor of an electronic device, the electronic device executes all or part of the steps of the methods described in the various embodiments of the present application.

[0115] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that, within the spirit and principles of the present application, they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate from the protection scope of the present application.

Claims

1. A method for detecting changes in urban blocks, characterized in that: include: Obtain N time-series satellite images of the target block, N ≥ 3, with one satellite image as a slice; Construct a time series segment centered at the i-th slice and of length k, 3≤k≤N; Calculating a semantic coherence score of the i-th slice based on the temporal correlation between the multiple slices in the temporal segment, including: extracting land cover type features of each slice in the temporal segment, the land cover type features including at least stable type features, structural change features, and functional transformation features; comparing the stable type features, the structural change features, and the functional transformation features of the multiple slices to obtain temporal correlation between the multiple slices; obtaining the semantic coherence score based on the temporal correlation, the semantic coherence score being used to reflect the degree of change of the target block at the time point corresponding to the i-th slice, the semantic coherence score being an output vector containing multiple elements, each element in the output vector corresponding to a change type of the target block; Obtaining a detection result according to the semantic coherence score includes: converting the output vector into a probability distribution vector, wherein each element in the vector represents the probability of occurrence of a change type of the corresponding target block; and selecting the change type with the greatest probability as the change type of the target block based on the probability of occurrence of the change type of the target block.

2. The method according to claim 1, characterized in that The structural change characteristics include changes in impervious surface coverage caused by expansion and / or demolition of the target block.

3. The method according to claim 1, characterized in that The semantic coherence score is calculated using the following formula: , in, represents the semantic coherence score of the i-th slice, is the activation function used to constrain the output semantic coherence score, Indicates the temporal correlation of multiple slices in the temporal segment where the i-th slice is located. and They respectively represent the weight matrix and bias term of a preset calculation model, where the calculation model is used to calculate the semantic coherence score of the i-th slice.

4. The method according to claim 3, characterized in that Before obtaining the semantic coherence score, the method further includes: performing convolution, maximum pooling and flattening operations on the land cover type features of the plurality of slices.

5. The method according to claim 1, wherein The semantic coherence score is input into a preset classification model to obtain a detection result, wherein the detection result includes a change boundary and a change type, and the change type is any one of expansion, demolition, and unchanged of the target block.

6. The method according to claim 5, characterized in that When the change type is that the target block remains unchanged, the method further includes: Along the forward direction of the time axis of N slices, one or more slices are moved with a window of length k to obtain a new time sequence segment.

7. A city block change detection system, characterized in that: include: The data acquisition module is used to obtain N time-series satellite images of the target block, where N is greater than or equal to 3, and one satellite image is used as a slice. The data processing module is suitable for constructing a time series segment centered on the i-th slice and of length k, 3≤k≤N; The data calculation module is adapted to calculate a semantic coherence score of the i-th slice based on the temporal correlation between the multiple slices in the temporal segment, comprising: extracting land cover type characteristics of each slice in the temporal segment, the land cover type characteristics including at least stable type characteristics, structural change characteristics, and functional transformation characteristics; comparing the stable type characteristics, the structural change characteristics, and the functional transformation characteristics of the multiple slices to obtain the temporal correlation between the multiple slices; and obtaining the semantic coherence score based on the temporal correlation, the semantic coherence score being used to reflect the degree of change of the target block at the time point corresponding to the i-th slice, the semantic coherence score being an output vector containing multiple elements, each element in the output vector corresponding to a type of change of the target block; A data generation module is adapted to obtain a detection result based on the semantic coherence score, comprising: converting the output vector into a probability distribution vector, wherein each element in the vector represents the probability of occurrence of a change type corresponding to a target block; and based on the probability of occurrence of the change type of the target block, selecting the change type with the highest probability as the change type of the target block.

8. An electronic device, characterized in that: The system comprises a processor and a memory, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the urban block change detection method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the urban block change detection method according to any one of claims 1 to 6 when running.