Urban green land multi-label classification method and system and computer storage medium
Through the dynamic cross-grouping deep convolution and sparse attention mechanism of the dual-branch neural network, combined with the lightweight visual Transformer, the problems of weak feature generalization ability and insufficient intelligence in traditional methods are solved, and efficient and accurate multi-label classification of urban green spaces is achieved.
Patent Information
- Application Number
- CN202510631417.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional urban green space classification methods have weak feature generalization capabilities, strong subjectivity in feature screening, and insufficient intelligence, which leads to limited classification accuracy and efficiency.
A dual-branch neural network is adopted, combined with dynamic cross-grouping deep convolution and sparse attention mechanism. The number of parameters is reduced by dynamically adjusting the grouping structure and sharing the convolution kernel. Combined with the lightweight visual Transformer, global semantic information is captured to achieve efficient feature extraction and classification.
It significantly improves the accuracy and efficiency of urban green space classification, and can accurately identify green space types in complex scenarios, especially mixed scenarios of buildings and green spaces, reduce computing resource consumption, and improve the flexibility and adaptability of the model.
Smart Images

Figure CN120807990A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of urban green space change monitoring, and particularly relates to a method and system for multi-label classification of urban green space and a computer storage medium. BACKGROUND
[0002] Fine classification of urban green space refers to high-precision and multi-level classification of urban green space according to multi-dimensional information such as vegetation type, functional attribute, spatial distribution characteristics and ecological service value of urban green space, which is of great significance for urban planning, ecological research and environmental management. Traditional urban green space identification methods include manual experience identification and machine learning methods. For manual experience identification, the work is complex, time-consuming, and requires high professional skills of the operator, and there is a difficulty of high labor cost. The machine learning-based method mainly includes vegetation index extraction, supervised classification, unsupervised classification, mixed pixel decomposition, object-oriented classification method, information fusion method, support vector machine classification method, decision tree method, etc. However, the feature generalization ability of the experimental sample is poor, and it is difficult to accurately classify the urban green space type of the test sample, mainly because the feature information of different types of urban green space in the experimental sample is similar, and the feature information of non-urban green space is more. The manual feature selection method usually cannot accurately eliminate invalid feature information. In addition, the traditional machine learning method is limited by the manual feature selection step and cannot realize completely intelligent application.
[0003] Therefore, in order to solve the problems of weak feature generalization ability, strong subjectivity of feature selection, and insufficient intelligence of the traditional method, it is necessary to explore an intelligent classification method that fuses multi-source heterogeneous data, automatically extracts deep features, and has strong generalization ability, so as to realize efficient, accurate and automatic classification of urban green space. SUMMARY
[0004] In order to solve the problems of weak feature generalization ability, strong subjectivity of feature selection, and insufficient intelligence of the traditional method, the purpose of the present application is to provide a method and system for multi-label classification of urban green space and a computer storage medium, and the technical solutions adopted are as follows:
[0005] In the first aspect, the present application discloses a method for multi-label classification of urban green space, which comprises:
[0006] S1, obtaining a pre-processed target remote sensing image data set for green space multi-label classification;
[0007] S2. Inputting the target remote sensing image dataset into a two-branch neural network for model training, and improving the model's adaptability to complex remote sensing scenes through a dynamic optimization strategy, wherein the first branch in the network adopts dynamic cross-grouping deep convolution, dynamically adjusts the grouping structure based on the similarity between input features, and achieves inter-group feature complementarity through a sparse attention mechanism and reduces the number of parameters by sharing convolution kernels;
[0008] S3. Perform end-to-end reasoning based on the trained dual-branch neural network to obtain the classification results of green space types in remote sensing images.
[0009] Furthermore, in step S2, the dynamic cross-group deep convolution is used to dynamically adjust the grouping structure based on the similarity between the input features, and the sparse attention mechanism is used to achieve inter-group feature complementarity and reduce the number of parameters by sharing the convolution kernel, including:
[0010] (1) Determine the normalized inner product between channels in the input feature map by calculating the cosine similarity, and generate a correlation matrix whose size is proportional to the square of the number of channels;
[0011] (2) Based on the correlation matrix, channel groups are gradually constructed in an iterative manner through a greedy algorithm, wherein in each iteration, a channel is selected from the channels that have not yet been grouped and added to the existing group to maximize the total similarity of the channels in the group, thereby ensuring that the channels in the group are highly correlated in feature expression;
[0012] (3) A sparse attention mechanism is introduced between each channel grouping. By calculating the attention weights of the features between different channel groups, the features of each channel are weightedly fused based on the attention weights, thereby achieving feature complementarity;
[0013] (4) Share the same depth convolution kernel among multiple channel groups to reduce the number of model parameters and computational complexity.
[0014] Furthermore, in step S2, the first branch in the network also adopts a dynamic channel pruning mechanism based on a gating mechanism to dynamically adjust the model structure according to the feature complexity of the input data.
[0015] Furthermore, in step S2, dynamically adjusting the model structure according to the feature complexity of the input data includes:
[0016] (1) Calculate the contribution score of each neuron channel under the current input based on the dynamic correlation between the spatial structure characteristics of the input sample and the channel weight;
[0017] (2) Based on the contribution score of each neuron channel, the dynamic threshold for channel pruning decision is calculated through dynamic statistical distribution analysis;
[0018] (3) based on the dynamic threshold, enabling or disabling neurons in the neural network through a comparison operation.
[0019] Further, in step S2, the contribution score of each neuron channel under the current input is calculated according to the dynamic correlation between the spatial structure features of the input sample and the channel weights, including:
[0020] (1) For the current input sample, extract its spatial structure features and generate a gating coefficient for dynamic channel selection;
[0021] (2) For each neuron channel, calculate the Frobenius norm of its weight matrix;
[0022] (3) Multiply the gating coefficient by the Frobenius norm of the channel weight to obtain the contribution score of each neuron channel under the current input.
[0023] Further, in step S2, during the channel pruning process, the method further comprises: constraining the pruning rate by the following correlation constraint term:
[0024]
[0025] Where λ·KL(p(m|x)||q(m)) represents the KL divergence term, p(m|x) represents the posterior distribution of the gating mask m given the input x, q(m) represents the prior distribution, λ represents the coefficient controlling the weight of the term, and β·Σ m | represents the absolute value summation term of the gating mask m, and β represents the coefficient controlling the weight of the term.
[0026] Further, in step S2, the second branch in the network adopts a lightweight visual Transformer as the backbone network, and captures global semantic information through a multi-scale fusion self-attention mechanism. After fusing feature maps of different scales and performing self-attention calculation, the model can understand image information from multiple levels.
[0027] Further, during the model training process, the method further comprises:
[0028] (1) Perform channel transformation processing on the ViT branch and the CNN branch respectively through a 1x1 convolution block to align the channel dimensions between the two branches;
[0029] (2) Perform dynamic weight allocation processing on the aligned ViT branch features and CNN branch features respectively to obtain weighted features F c-weighted , F s-weighted of each branch;
[0030] (3) Perform feature fusion on the weighted features F c-weighted, F s-weighted The addition is performed, and the final fusion feature F is obtained through a layer normalization operation fusion .
[0031] In a second aspect, the application discloses a city green land multi-label classification system, the system comprises a data acquisition module, a double-branch neural network training module, and a green land multi-label classification module, wherein:
[0032] The data acquisition module is configured to acquire a target remote sensing image data set for green land multi-label classification after preprocessing.
[0033] The double-branch neural network training module is configured to input the target remote sensing image data set into a double-branch neural network for model training, and improve the adaptability of the model to complex remote sensing scenes through a dynamic optimization strategy, wherein a first branch in the network adopts a dynamic cross-group deep convolution, dynamically adjusts a group structure based on the similarity between input features, and realizes feature complementation between groups through a sparse attention mechanism and reduces the parameter quantity through a shared convolution kernel.
[0034] The green land multi-label classification module is configured to perform end-to-end inference based on the trained double-branch neural network, and obtain a classification result of green land types in a remote sensing image.
[0035] In a third aspect, the application discloses a computer storage medium, the computer storage medium is used for storing computer execution instructions, and the computer execution instructions are used for executing the city green land multi-label classification method.
[0036] The application has the following beneficial effects:
[0037] (1) The application proposes a design based on a double-branch neural network, which combines the local feature extraction capability of CNN and the global information capturing capability of Transformer. Among them, the first branch adopts a MobileNetV2 based on dynamic grouping pruning architecture, efficiently extracts local texture features through grouped deep convolution and dynamic channel pruning mechanism, and significantly reduces the parameter quantity and computational complexity of the model. Through an asymmetric path, the difference learning of spectral-spatial features is realized, and the respective advantages of CNN and Transformer are fully utilized. The second branch adopts a lightweight visual Transformer as the backbone network, captures global semantic information by using the self-attention mechanism, and solves the potential correlation problem between labels in complex scenes.
[0038] (2) In view of the fact that there is a lack of information interaction between traditional grouped convolutions, which leads to limited feature fusion, and the pruning process relies on a fixed threshold, which cannot dynamically adjust the model structure according to the complexity of input data, the present application proposes a dynamic grouping pruning architecture, including dynamic channel pruning and grouped depth convolution optimization. Among them, the dynamic channel pruning dynamically adjusts the model structure according to the complexity of the input data through the gating mechanism, reduces the amount of calculation while maintaining the model performance, can adaptively adjust the computing resources according to the complexity of the input data, and significantly improves the efficiency and flexibility. The grouped depth convolution optimization dynamically adjusts the grouping structure through the cosine similarity, introduces the sparse attention mechanism to realize the complementary features between groups, and reduces the parameter amount through shared convolution kernels, which makes up for the performance loss of traditional grouped convolution, and realizes the balance between efficiency and performance.
[0039] (3) The present application proposes a space-channel dual attention mechanism, which strengthens the multi-modal feature fusion efficiency through the decoupling design of independent channel and spatial attention path; by modeling the relationship between different channels, the importance of each feature channel is automatically evaluated, and higher weight is allocated to the key channel. In complex scenes, the recognition ability of the model for overlapping targets (such as mixed scenes of buildings and green land) can be improved by suppressing the noise interference of low correlation areas. BRIEF DESCRIPTION OF DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, and the advantages thereof, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.
[0041] Figure 1 The method flow chart of the urban green land multi-label classification method provided by an embodiment of the present application;
[0042] Figure 2 The overall implementation flow chart of the urban green land multi-label classification method provided by an embodiment of the present application;
[0043] Figure 3 The system structure diagram of the urban green land multi-label classification system provided by an embodiment of the present application;
[0044] Figure 4 The structure schematic diagram of the computer storage medium provided by an embodiment of the present application. DETAILED DESCRIPTION
[0045] In order to further illustrate the technical means and effects taken by the present application to achieve the predetermined inventive objectives, the following describes in detail the specific implementation, structure, features and effects of a kind of urban green space multi-label classification method, system and computer storage medium according to the present application, combined with the preferred embodiments and the drawings. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.
[0047] The specific scheme of the urban green space multi-label classification method, system and computer storage medium provided by the present application is described in detail below in combination with the drawings.
[0048] Please refer to Figure 1 and Figure 2 which shows the method flowchart of a kind of urban green space multi-label classification method provided by an embodiment of the present application, the method comprises:
[0049] Step S1, obtaining a pre-processed target remote sensing image data set for green space multi-label classification.
[0050] It should be noted that the remote sensing image data set needs to cover multiple time phases, multiple spectral bands, and support the generation of true color, false color and NDVI image, while ensuring that the data is strictly consistent in terms of spatial range, resolution, projection coordinate system, etc. to meet the standardization requirements of multi-source data fusion and model training.
[0051] Specifically, the pre-processing method involved in step S1 includes at least one of geometric transformation, color transformation, MixUp data enhancement, GAN enhancement strategy for generating realistic new data through adversarial training, and searching for the optimal data enhancement strategy based on AutoML technology. The purpose of data preprocessing is to enhance the diversity and robustness of data and reduce the dependence on a large amount of labeled data.
[0052] Step S2, inputting the target remote sensing image data set into a double-branch neural network for model training, and improving the adaptability of the model to complex remote sensing scenes through a dynamic optimization strategy, wherein the first branch in the network adopts dynamic cross-group deep convolution, dynamically adjusts the grouping structure based on the similarity between input features, and realizes feature complementation between groups through sparse attention mechanism and reduces parameter quantity through shared convolution kernel.
[0053] It should be noted that the dynamic cross-grouping depth convolution drives the model to dynamically divide the channel subsets according to the similarity threshold and optimize the connection mode between the subsets to flexibly reorganize the channels according to the current input characteristics, so as to improve the synergistic effect of the features in the group, based on the similarity calculation results between the input features. In the application process of the sparse attention mechanism, the features between different dynamic groups are weighted and fused to realize the complementation of the features between the groups, and the convolution kernel parameters are cross-grouped by sharing the convolution kernel to reduce the model parameter quantity.
[0054] Step S3: performing end-to-end inference based on the trained double-branch neural network to obtain the classification result of the green land type in the remote sensing image.
[0055] As can be seen from the above, the urban green land multi-label classification method disclosed in the present application proposes a design based on a double-branch neural network, which combines the local feature extraction capability of CNN and the global information capturing capability of Transformer. Among them, the first branch adopts a dynamic grouping pruning architecture MobileNetV2, which efficiently extracts local texture features and significantly reduces the model parameter quantity and computational complexity through grouping depth convolution and dynamic channel pruning mechanism. The asymmetric path realizes the differentiated learning of spectral-spatial features, and fully gives play to the respective advantages of CNN and Transformer. The second branch adopts a lightweight visual Transformer as the backbone network, which captures global semantic information by using its self-attention mechanism and solves the potential correlation problem between labels in complex scenes. In view of the fact that the traditional grouping convolution lacks information interaction between groups, which leads to limited feature fusion, and the pruning process relies on a fixed threshold, which cannot dynamically adjust the model structure according to the complexity of the input data, the present application proposes a dynamic grouping pruning architecture, including dynamic channel pruning and grouping depth convolution optimization. Among them, the dynamic channel pruning dynamically adjusts the model structure according to the complexity of the input data through the gating mechanism, reduces the computational complexity while maintaining the model performance, can adaptively adjust the computing resources according to the complexity of the input data, and significantly improves the efficiency and flexibility. The grouping depth convolution optimization dynamically adjusts the grouping structure through cosine similarity, introduces a sparse attention mechanism to realize the complementation of the features between groups, and reduces the parameter quantity by sharing the convolution kernel, which makes up for the performance loss of the traditional grouping convolution and realizes the balance between efficiency and performance. Finally, the present application also proposes a spatial-channel dual attention mechanism, which strengthens the multi-modal feature fusion efficiency through the decoupling design of independent channel and spatial attention path, models the relationship between different channels, automatically evaluates the importance of each feature channel, and allocates higher weights to the key channels. In complex scenes, the model can suppress the noise interference in the low correlation region to improve the resolution capability of the model for overlapping targets (such as building and green land mixed scenes).
[0056] In one of the embodiments, in step S2, the dynamic cross-grouping deep convolution is used to dynamically adjust the grouping structure based on the similarity between input features, and to realize the complementarity of features between groups through sparse attention mechanism and to reduce the parameter quantity through shared convolution kernel, including:
[0057] (1) By calculating the cosine similarity, the normalized inner product between each channel in the input feature map is determined to generate a correlation matrix whose size is proportional to the square of the number of channels;
[0058] (2) Based on the correlation matrix, the channel grouping is gradually constructed in an iterative manner through a greedy algorithm, wherein in each iteration, a channel is selected from the channels that have not been grouped and is added to the existing group to maximize the total similarity of the channels in the group, thereby ensuring that the channels in the group are highly correlated in feature expression.
[0059] Specifically, for each possible grouping operation, the application calculates the change in the total similarity of the channels in the group before and after the channel is added, and selects the operation with the largest change to execute. This process is repeated until all channels are grouped and meet the constraint condition that the number of groups |G k | = M (where, C is the total number of channels, and G is the preset number of groups) to dynamically adjust the grouping structure, maximize the total similarity of the channels in the group, and ensure that the channels in each group are highly correlated in feature expression.
[0060] (3) A sparse attention mechanism is introduced between the channel groups to calculate the attention weight of the features between different channel groups, and to weight and fuse the channel features based on the attention weight, thereby realizing the complementarity of the features.
[0061] Specifically, the application will interact with the Query vector Q p of the pth group and the Key vector K q and the Value vector V q of the qth group to calculate the attention weight A p,q between the two channel groups through the following formula:
[0062]
[0063] where d represents the feature dimension between the Query vector Q p and the Key vector K q , which must be consistent.
[0064] Therefore, by introducing the sparse attention mechanism between the channel groupings, the model can selectively aggregate useful information from other channel groupings, effectively improving the model's ability to capture complex features, thereby enhancing the model's performance on related tasks and solving the problem of information isolation between groups.
[0065] (4) Sharing the same deep convolution kernel among multiple channel groupings to reduce the parameter quantity of the model and reduce the computational complexity.
[0066] Specifically, considering that traditional grouped convolution needs to assign independent convolution kernels to each channel grouping, which will cause the parameter quantity of the model to increase significantly with the increase in the number of channel groupings, greatly occupying storage resources and increasing the difficulty of training and inference of the model. Therefore, in order to reduce the parameter quantity of the model and reduce the computational complexity, the application adopts the strategy of sharing the same deep convolution kernel among multiple channel groupings, so that multiple channel groupings can use the same deep convolution kernel for feature extraction, thereby optimizing the model structure.
[0067] In one embodiment, in step S2, the first branch in the network also adopts a dynamic channel pruning mechanism based on a gating mechanism to dynamically adjust the model structure according to the feature complexity of the input data.
[0068] It should be noted that in the process of dynamic channel pruning, the contribution score of each neuron channel under the current input is calculated according to the dynamic correlation between the spatial structure features of the current input sample and the channel weights. Through threshold comparison operation to determine whether each neuron channel should be retained or pruned, to enable or disable the neurons in the neural network, and to reduce unnecessary neuron calculation to reduce the amount of calculation.
[0069] In one embodiment, in step S2, the dynamic adjustment of the model structure according to the feature complexity of the input data includes:
[0070] (1) According to the dynamic correlation between the spatial structure features of the input sample and the channel weights, the contribution score of each neuron channel under the current input is calculated.
[0071] Specifically, for the current input sample, the application first extracts its spatial structure features and generates gating coefficients for dynamic channel selection, then calculates the Frobenius norm of each neuron channel weight matrix, and finally multiplies the gating coefficients and the Frobenius norm of the channel weights to obtain the contribution score of each neuron channel under the current input.
[0072] (2) Based on the contribution score of each neuron channel, the dynamic threshold for channel pruning decision is calculated through dynamic statistical distribution analysis.
[0073] Specifically, the application performs dynamic statistical distribution analysis by the following formula:
[0074] τ(x) = μ · Quantile(S (c) , k) + b;
[0075] wherein τ(x) represents a dynamic threshold, μ represents a learnable scaling factor, k and b represent quantile parameters and bias terms, S (c) represents the contribution score of each neuron channel, and Quantile(*) represents a quantile function used to find the corresponding quantile value from the set of contribution scores of neuron channels according to the given quantile point parameter k.
[0076] (3) Based on the dynamic threshold, enabling or disabling neurons in the neural network through a comparison operation.
[0077] It should be noted that after obtaining the dynamic threshold, the application compares the contribution score of each neuron channel with the threshold. If the score is greater than or equal to the threshold, the corresponding neuron channel is enabled; otherwise, if the score is less than the threshold, the corresponding neuron channel is disabled. In this way, the enabling state of the neuron channel in the neural network is dynamically adjusted according to the current input, which effectively reduces unnecessary computational resource consumption while ensuring model performance, and improves the inference efficiency of the model.
[0078] In the above embodiment, the model structure is dynamically adjusted according to the complexity of the input data through the gating mechanism, which not only reduces the amount of calculation while maintaining the model performance, but also adaptively adjusts the computing resources according to the complexity of the input data, significantly improving the calculation efficiency and flexibility.
[0079] In one of the embodiments, in step S2, the contribution score of each neuron channel under the current input is calculated according to the dynamic correlation between the spatial structure feature of the input sample and the channel weight, including:
[0080] (1) For the current input sample, its spatial structure feature is extracted and a gating coefficient for dynamic channel selection is generated.
[0081] (2) For each neuron channel, the Frobenius norm of its weight matrix is calculated.
[0082] (3) Multiply the gating coefficient by the Frobenius norm of the channel weight to obtain the contribution score of each neuron channel under the current input.
[0083] Specifically, the contribution score of each neuron channel under the current input can be understood by the following formula:
[0084] S (c) = fθ (x)·||W (c) || F ;
[0085] wherein, f θ (x) represents a gating coefficient for dynamic channel selection, ||W (c) || F represents the Frobenius norm corresponding to the weight matrix of the neuron channel c.
[0086] In one of the embodiments, in the process of channel pruning in step S2, the method further comprises: constraining the pruning rate by the following associated constraint term:
[0087]
[0088] wherein, λ·KL(p(m|x)‖q(m)) represents the KL divergence term, p(m|x) represents the posterior distribution of the gating mask m given the input x, q(m) represents the prior distribution, λ represents the coefficient controlling the weight of the term, β·∑ m |m| represents the absolute value summation term of the gating mask m, and β represents the coefficient controlling the weight of the term.
[0089] It should be noted that the present application forces the distribution of the gating mask m to be close to the prior distribution q(m) through the KL divergence term, so as to avoid the destruction of the stability of the model due to extreme sparsification, and controls the weight of different constraints through the λ coefficient to coordinate the performance and lightweight.
[0090] In one of the embodiments, in step S2, the second branch in the network adopts a lightweight visual Transformer as the backbone network, and captures global semantic information through multi-scale fusion self-attention mechanism, so that the model can understand image information from multiple levels after fusing the feature maps of different scales and then performing self-attention calculation.
[0091] In one of the embodiments, in the process of model training, the method further comprises:
[0092] (1) The ViT branch and the CNN branch are respectively subjected to channel transformation processing through a 1×1 convolution block, so as to realize the alignment of the channel dimensions between the two branches.
[0093] Specifically, the 1×1 convolution block adjusts the number of convolution kernels to remap the channel number of the input feature without changing the spatial size of the feature map, so as to realize the alignment of the channel dimensions between the two branches and lay a foundation for subsequent feature fusion.
[0094] (2) The aligned ViT branch feature and CNN branch feature are respectively subjected to dynamic weight allocation processing to obtain the weighted features Fc-weighted s-weighted .
[0095] It should be noted that the information interaction efficiency across branches is currently optimized through dynamic weight allocation, similar to the multi semantic space alignment strategy in the SCSA module. Such design can avoid the information bottleneck problem caused by global pooling (such as SE-Net) or simple superposition (such as CBAM), and is more suitable for capturing fine-grained features (such as vegetation texture, shape difference).
[0096] (3) adding the weighted features F c-weighted s-weighted and obtaining the final fusion feature F fusion through layer normalization operation.
[0097] Specifically, the weighted features F c-weighted s-weighted are added element by element to realize the fusion of the feature information of the two branches. Subsequently, the fused features are normalized through layer normalization operation, and the mean and variance of the features are adjusted to make them have better numerical stability and generalization ability, and finally the fusion feature F fusion is obtained.
[0098] Please refer to Figure 3 , the urban green space multi-label classification system disclosed in the present application comprises a data acquisition module, a double-branch neural network training module, and a green space multi-label classification module, wherein:
[0099] The data acquisition module is configured to acquire a preprocessed target remote sensing image data set for green space multi-label classification.
[0100] The double-branch neural network training module is configured to input the target remote sensing image data set into a double-branch neural network for model training, and improve the adaptability of the model to complex remote sensing scenes through a dynamic optimization strategy. In the network, the first branch adopts dynamic cross-group deep convolution, dynamically adjusts the grouping structure based on the similarity between input features, and realizes the complementation of inter-group features through a sparse attention mechanism and reduces the parameter amount through shared convolution kernels.
[0101] The green space multi-label classification module is configured to perform end-to-end inference based on the trained double-branch neural network to obtain the classification result of the green space type in the remote sensing image.
[0102] In one embodiment, the above modules are also configured to implement the urban green space multi-label classification method according to any one of the preceding method embodiments, which is not limited in the present application.
[0103] As can be known from the above, the urban green land multi-label classification system disclosed by the application proposes a design based on a double-branch neural network, which combines the local feature extraction capability of CNN and the global information capturing capability of Transformer. Among them, the first branch adopts a dynamic grouping pruning architecture MobileNetV2, which efficiently extracts local texture features through a grouping depth convolution and a dynamic channel pruning mechanism, while significantly reducing the parameter quantity and computational complexity of the model. Through asymmetric paths, the differences in spectral-spatial features are learned, and the respective advantages of CNN and Transformer are fully utilized. The second branch adopts a lightweight visual Transformer as the backbone network, which captures global semantic information using its self-attention mechanism and solves the potential correlation problem between labels in complex scenes. In view of the fact that traditional grouped convolutions lack information interaction, which leads to limited feature fusion, and the pruning process relies on a fixed threshold, which cannot dynamically adjust the model structure according to the complexity of the input data, the application proposes a dynamic grouping pruning architecture, including dynamic channel pruning and grouping depth convolution optimization. Among them, the dynamic channel pruning dynamically adjusts the model structure according to the complexity of the input data through a gating mechanism, reduces the amount of calculation while maintaining the performance of the model, can adaptively adjust the computing resources according to the complexity of the input data, and significantly improves the efficiency and flexibility. The grouping depth convolution optimization dynamically adjusts the grouping structure through cosine similarity, introduces a sparse attention mechanism to realize complementary features between groups, and reduces the parameter quantity through shared convolution kernels, which makes up for the performance loss of traditional grouped convolutions, realizes the balance between efficiency and performance. Finally, the application also proposes a spatial-channel dual attention mechanism, which strengthens the multi-modal feature fusion efficiency through the decoupling design of independent channel and spatial attention paths; by modeling the relationship between different channels, the importance of each feature channel is automatically evaluated, and higher weights are allocated to key channels. In complex scenes, the recognition ability of the model for overlapping targets (such as building and green land mixed scenes) can be improved by suppressing the noise interference of low correlation regions.
[0104] Please refer to Figure 4 The computer storage medium disclosed by the application is used to store computer execution instructions, and the computer execution instructions are used to execute the urban green land multi-label classification method described in any of the preceding embodiments.
[0105] As can be known from the above, the computer storage medium disclosed in the application proposes a design based on a double-branch neural network, which combines the local feature extraction capability of CNN and the global information capturing capability of Transformer. Among them, the first branch adopts a dynamic grouping pruning architecture MobileNetV2, which efficiently extracts local texture features through a grouping depth convolution and a dynamic channel pruning mechanism, while significantly reducing the parameter quantity and computational complexity of the model. Through asymmetric paths, the differences in spectral-spatial features are learned, and the respective advantages of CNN and Transformer are fully utilized. The second branch adopts a lightweight visual Transformer as the backbone network, which captures global semantic information using its self-attention mechanism and solves the potential correlation problem between labels in complex scenes. In view of the fact that traditional grouped convolutions lack information interaction, resulting in limited feature fusion, and the pruning process relies on a fixed threshold, which cannot dynamically adjust the model structure according to the complexity of the input data, the application proposes a dynamic grouping pruning architecture, including dynamic channel pruning and grouping depth convolution optimization. Among them, the dynamic channel pruning dynamically adjusts the model structure according to the complexity of the input data through a gating mechanism, reduces the amount of calculation while maintaining the performance of the model, can adaptively adjust the computing resources according to the complexity of the input data, and significantly improves the efficiency and flexibility. The grouping depth convolution optimization dynamically adjusts the grouping structure through cosine similarity, introduces a sparse attention mechanism to realize complementary features between groups, and reduces the parameter quantity through shared convolution kernels, making up for the performance loss of traditional grouped convolutions, achieving a balance between efficiency and performance. Finally, the application also proposes a spatial-channel dual attention mechanism, which strengthens the multi-modal feature fusion efficiency through the decoupling design of independent channel and spatial attention paths; by modeling the relationship between different channels, the importance of each feature channel is automatically evaluated, and higher weights are allocated to key channels. In complex scenes, the recognition ability of the model for overlapping targets (such as building and green mixed scenes) can be improved by suppressing the noise interference of low correlation regions.
[0106] It should be noted that the above-mentioned order of the embodiments of the application is only for description, and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or may be advantageous.
[0107] Each embodiment in the specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment mainly describes the differences from other embodiments.
Claims
1. A multi-label classification method for urban green space, characterized by: The method comprises: S1. Obtain the preprocessed target remote sensing image dataset for green space multi-label classification; S2. Inputting the target remote sensing image dataset into a two-branch neural network for model training, and improving the model's adaptability to complex remote sensing scenes through a dynamic optimization strategy, wherein the first branch in the network adopts dynamic cross-grouping deep convolution, dynamically adjusts the grouping structure based on the similarity between input features, and achieves inter-group feature complementarity through a sparse attention mechanism and reduces the number of parameters by sharing convolution kernels; S3. Perform end-to-end reasoning based on the trained dual-branch neural network to obtain the classification results of green space types in remote sensing images.
2. The method according to claim 1, characterized in that In step S2, the dynamic cross-group deep convolution is used to dynamically adjust the grouping structure based on the similarity between the input features, and the sparse attention mechanism is used to achieve inter-group feature complementarity and reduce the number of parameters by sharing the convolution kernel, including: (1) Determine the normalized inner product between the channels in the input feature map by calculating the cosine similarity, and generate a correlation matrix whose size is proportional to the square of the number of channels; (2) Based on the correlation matrix, channel groups are gradually constructed in an iterative manner through a greedy algorithm, wherein in each iteration, a channel is selected from the channels that have not yet been grouped and added to the existing group to maximize the total similarity of the channels in the group, thereby ensuring that the channels in the group are highly correlated in feature expression; (3) A sparse attention mechanism is introduced between each channel grouping. By calculating the attention weights of the features between different channel groups, the features of each channel are weightedly fused based on the attention weights, thereby achieving feature complementarity; (4) Share the same depth convolution kernel among multiple channel groups to reduce the number of model parameters and computational complexity.
3. The method according to claim 1, characterized in that In step S2, the first branch in the network also adopts a dynamic channel pruning mechanism based on a gating mechanism to dynamically adjust the model structure according to the feature complexity of the input data.
4. The method according to claim 3, characterized in that In step S2, dynamically adjusting the model structure according to the feature complexity of the input data includes: (1) Calculate the contribution score of each neuron channel under the current input based on the dynamic correlation between the spatial structure characteristics of the input sample and the channel weight; (2) Based on the contribution score of each neuron channel, the dynamic threshold for channel pruning decision is calculated through dynamic statistical distribution analysis; (3) Based on the dynamic threshold, neurons in the neural network are enabled or disabled through comparison operations.
5. The method according to claim 4, characterized in that In step S2, the contribution score of each neuron channel under the current input is calculated based on the dynamic correlation between the spatial structure characteristics of the input sample and the channel weight, including: (1) For the current input sample, extract its spatial structure features and generate gating coefficients for dynamic channel selection; (2) For each neuron channel, calculate the Frobenius norm of its weight matrix; (3) Multiply the gating coefficient by the Frobenius norm of the channel weight to obtain the contribution score of each neuron channel under the current input.
6. The method according to claim 3, characterized in that In step S2, during the channel pruning process, the method further includes: constraining the pruning rate by the following associated constraint items: Among them, λ·KL(p(m|x)||q(m)) represents the KL divergence term, p(m|x) represents the posterior distribution of the gated mask m given the input x, q(m) represents the prior distribution, λ represents the coefficient controlling the weight of this term, β·∑ m |m| represents the absolute value summation term of the gated mask m, and β represents the coefficient controlling the weight of this term.
7. The method according to claim 1, characterized in that In step S2, the second branch of the network uses a lightweight visual Transformer as the backbone network, captures global semantic information through a multi-scale fusion self-attention mechanism, fuses feature maps of different scales and then performs self-attention calculation, enabling the model to understand image information from multiple levels.
8. The method according to claim 1, characterized in that During the model training process, the method further includes: (1) Perform channel transformation on the ViT branch and the CNN branch through a 1×1 convolution block to achieve alignment of the channel dimensions between the two branches; (2) Dynamically weight the aligned ViT branch features and CNN branch features to obtain the weighted features F of each branch. c-weighted 、F s-weighted ; (3) The weighted feature F c-weighted 、F s-weighted Add them together and perform layer normalization to obtain the final fusion feature F fusion .
9. A multi-label classification system for urban green space, characterized by: The system includes a data acquisition module, a dual-branch neural network training module, and a green space multi-label classification module, wherein: The data acquisition module is used to obtain the pre-processed target remote sensing image dataset for green space multi-label classification; The dual-branch neural network training module is used to input the target remote sensing image dataset into the dual-branch neural network for model training, and improve the model's adaptability to complex remote sensing scenes through a dynamic optimization strategy, wherein the first branch in the network adopts dynamic cross-group deep convolution to dynamically adjust the grouping structure based on the similarity between input features, and achieve inter-group feature complementarity through a sparse attention mechanism and reduce the number of parameters by sharing convolution kernels; The green space multi-label classification module is used to perform end-to-end reasoning based on the trained two-branch neural network to obtain the classification results of green space types in remote sensing images.
10. A computer storage medium, characterized in that The computer storage medium is used to store computer-executable instructions, and the computer-executable instructions are used to execute the urban green space multi-label classification method according to any one of claims 1 to 8.