A method and system for remote sensing image scene classification that guides deep network learning using a knowledge graph

By constructing the land cover concept knowledge graph and cross-modal alignment constraints, the problem of deep network dependence on training samples in remote sensing image scene classification is solved, and remote sensing image scene classification with higher accuracy and stronger generalization capabilities is achieved.

CN116543232BActive Publication Date: 2025-08-01WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310630415.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-30
Publication Date
2025-08-01
Estimated Expiration
2043-05-30

AI Technical Summary

Technical Problem

The existing remote sensing image scene classification method based on deep networks highly relies on a large number of manually labeled training samples and cannot effectively integrate and utilize the rich prior knowledge in the remote sensing field, resulting in insufficient generalization capabilities of the model and limited accuracy performance.

Method used

By constructing a land cover concept knowledge graph, using knowledge graph representation learning to generate remote sensing scene semantic benchmarks, and guiding deep network learning through cross-modal alignment constraints, reducing dependence on manual annotation training samples and improving classification performance.

Benefits of technology

It effectively improves the accuracy and generalization ability of remote sensing image scene classification, reduces dependence on training samples, and improves the flexibility and accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116543232B_ABST
    Figure CN116543232B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for remote sensing image scene classification that guides deep network learning using a knowledge graph. First, a land cover concept knowledge graph is constructed, and then, through knowledge graph representation learning methods, the remote sensing scene semantic categories in the land cover concept knowledge graph are expressed as semantic vectors, forming a semantic benchmark for remote sensing scene categories. In the knowledge graph-guided deep network training stage, by imposing cross-modal alignment constraints between the semantic vectors of remote sensing scene categories and the shallow visual feature vectors of the deep network, the shallow part of the deep network is guided to more effectively learn the shared features of different categories of remote sensing image scenes. In the deep part of the deep network, the scene category labels are still used as constraints to guide the deep network to learn discriminative features for distinguishing different remote sensing scenes. In the testing stage, the optimized deep network model can achieve high-precision remote sensing image scene classification without relying on any prior knowledge.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - field of artificial intelligence, knowledge graph and remote sensing image interpretation, and relates to a method for remote sensing image scene classification guided by knowledge graph - based deep network learning. Specifically, it includes a cross - modal alignment constraint between the remote sensing scene semantic vector obtained by applying the knowledge graph and the shallow - layer visual feature vector of the deep network to guide the shallow part of the deep network to more effectively learn the shared features of different - category remote sensing image scenes, and overall improve the performance of the deep network in remote sensing image scene classification. Background Art

[0002] With the rapid development of satellite remote sensing technology, the means and capabilities of observing the Earth have achieved revolutionary breakthroughs. Humans can efficiently obtain a large amount of high - resolution remote sensing image observation data at regional and global scales with the help of satellite remote sensing technology, and the era of remote sensing big data has arrived. For a large amount of high - resolution remote sensing image data, using computer technology to interpret it from simple observation data into remote sensing geoscience knowledge that can serve humans is an important research content in the current remote sensing field.

[0003] However, due to the typical phenomenon of "same object with different spectra, same spectrum with different objects" in high - resolution remote sensing images, traditional remote sensing image land - cover classification methods based on pixels and objects have great limitations and cannot effectively meet the interpretation capabilities and application requirements for high - resolution remote sensing images. Therefore, the task of remote sensing image scene classification has received more attention from remote sensing researchers at home and abroad. Specifically, the goal of remote sensing image scene classification is to mine the relationship between basic visual elements and basic visual elements in remote sensing images, so as to predict the semantic categories of ground - object scenes in remote sensing image blocks, thereby effectively reducing the confusion that may occur during pixel - level or object - level interpretation of ground objects, making the interpretation results of high - resolution remote sensing images more stable and accurate. Therefore, remote sensing image scene classification is one of the important means to understand high - resolution remote sensing images and plays a crucial role in tasks such as traffic control, disaster monitoring, and land - cover classification.

[0004] Most traditional remote sensing image scene classification methods classify remote sensing image scenes based on artificially constructed features. For example, early remote sensing image scene classification methods would use information such as color, texture, and edges to construct low - level visual features, or further composite low - level visual features into middle - level visual features and then apply them to remote sensing image scene classification. However, the middle - and low - level visual features constructed using these simple visual information are often difficult to express the global information of high - resolution remote sensing image blocks, and the artificial construction of visual features completely depends on prior knowledge and is difficult to cope with the current large number of high - resolution remote sensing images with complex ground - object distributions.

[0005] In recent years, with the development of deep learning technology, the data-driven method of deep neural networks has demonstrated excellent performance in related tasks of natural image processing. Therefore, it has been gradually applied to the interpretation tasks related to remote sensing images and achieved excellent results. However, there is another problem with the data-driven deep learning method: Although the deep network can capture more and higher-level visual features beneficial to the remote sensing image scene classification task through training on training samples, it can hardly make full and effective use of the prior knowledge in the remote sensing field. Therefore, this defect makes the deep network extremely prone to overfitting to the training samples when facing insufficient training samples, which greatly restricts the generalization ability of the remote sensing image scene classification model and limits the accuracy performance of the model. Therefore, it is necessary to guide the deep network through the prior knowledge in the remote sensing field, improve the existing remote sensing image scene classification method based on the deep network, and obtain better remote sensing image scene classification results. Summary of the Invention

[0006] The present invention mainly solves the problem that the existing remote sensing image scene classification method based on the deep network highly depends on a large number of manually labeled training samples and cannot effectively integrate and utilize the rich prior knowledge in the remote sensing field, and proposes a new remote sensing image scene classification method guided by a knowledge graph to deep network learning. By constructing a land cover concept knowledge graph and a remote sensing scene semantic benchmark obtained from the semantic expression of land cover, the remote sensing image method based on the deep network is guided to improve the performance of remote sensing image scene classification while reducing the dependence of the model on manually labeled training samples.

[0007] The remote sensing image scene classification method guided by the knowledge graph to deep network learning first constructs a land cover concept knowledge graph to manage and utilize the prior knowledge of remote sensing scenes more comprehensively, flexibly, and conveniently; then obtains a remote sensing scene semantic benchmark through the semantic expression of land cover concepts for guiding deep network learning; finally, in the knowledge graph-guided deep network training stage, by imposing a cross-modal alignment constraint between the remote sensing scene category semantic vector and the shallow-layer visual feature vector of the deep network, the shallow part of the deep network is guided to more effectively learn the shared features of different categories of remote sensing image scenes, and in the deep part of the deep network, the scene category label constraint is still used to guide the deep network to learn the discriminant features for distinguishing different remote sensing scenes. In the test stage, the optimized deep network model can complete high-precision remote sensing image scene classification without relying on any prior knowledge.

[0008] The technical solution adopted by the present invention is: The remote sensing image scene classification method guided by the knowledge graph to deep network learning reduces the dependence of the remote sensing image scene classification method based on the deep network on manually labeled training samples through the guidance of the prior knowledge in the knowledge graph, and includes the following steps:

[0009] Step 1, construct a land cover concept knowledge graph;

[0010] Step 2, perform semantic expression of land cover concepts based on the land cover classification knowledge graph to obtain a remote sensing scene semantic benchmark, that is, a set A of remote sensing scene semantic category feature vectors;

[0011] Step 3, use the remote sensing scene semantic category vectors to guide and optimize the deep network for remote sensing image scene classification;

[0012] The deep network is an open-source model, which is divided into a shallow part and a deep part. The input remote sensing image scene obtains visual features through the shallow part, and then obtains classification results through the deep part; during the guidance and optimization process, first transform the visual features obtained by the shallow part into a set V of visual feature vectors, and then use the knowledge-guided learning module to perform cross-modal alignment constraints on the visual feature vector v and the scene semantic category vector a in the set, guiding the shallow part of the deep network to more effectively learn the shared features between remote sensing scene categories;

[0013] The knowledge-guided learning module includes a visual feature encoder E v 、a visual feature decoder D v 、a semantic feature encoder E a and a semantic feature decoder D a , and sequentially implement the cross-modal alignment operation of the visual feature vector v and the scene category semantic category vector a, where v ∈ V and a ∈ A;

[0014] Step 4, use the optimized deep network for remote sensing image scene classification.

[0015] Furthermore, the land cover concept knowledge graph constructed in Step 1 includes: domain standards, expert knowledge, and general knowledge;

[0016] Among them, the domain standards include the third national land survey classification system, which hierarchically classifies the remote sensing scene categories in the land cover classification task according to the secondary classification structure, and strictly defines each category in words;

[0017] Expert knowledge includes the domain knowledge sorted out by domain experts in the practice tasks of the remote sensing field, including the common internal structures of remote sensing scenes and the spatial relationship knowledge between remote sensing scenes.

[0018] Furthermore, in step 2, the knowledge graph representation learning model DistMult is adopted to perform the semantic expression of land cover concepts, and the semantic category representation vector of the remote sensing scene is obtained. The constructed remote sensing scene knowledge graph G is represented as a set of fact triples G = {<h, r, t>|h, t ∈ E, r ∈ R}, where E and R represent the entity set and the relationship set of the remote sensing scene knowledge graph respectively, and h, r, and t represent the head entity, relationship, and tail entity of the triple respectively. The DistMult model uses a triple scoring function based on a bilinear model to measure the semantic similarity between the head and tail entity pairs connected by the same relationship:

[0019]

[0020] Among them, M r represents the transformation matrix corresponding to the relationship r, and y h , y t represent the semantic representation vectors of the head and tail entities respectively;

[0021] Based on the constraint relationship of the above-mentioned land cover concept semantic expression method of knowledge graph representation learning, the overall land cover concept knowledge graph is subjected to representation vector mapping. After learning, all triples G = {<h, r, t>|h, t ∈ E, r ∈ R} in the land cover concept knowledge graph can obtain the semantic expression vector set of all entities in the knowledge graph Among them is the total number of entities in the remote sensing scene knowledge graph. In Y e select entity semantic vectors corresponding to the remote sensing scene semantic categories, and obtain the remote sensing scene semantic category feature vector set That is, the remote sensing scene semantic benchmark.

[0022] Furthermore, in step 3, the open-source model efficientnet_b3 network is adopted as the deep network. All convolutional layers before the activation function located at Swish54 in this open-source model are regarded as the shallow part, and the extracted shallow visual features are led out and input into the knowledge guidance module for cross-modal alignment with the semantic category vector. The model structure after the activation function is correspondingly regarded as the deep part.

[0023] Furthermore, the overall loss function adopted in the step 3 guidance optimization includes the knowledge embedding loss of the shallow part of the deep network and the data learning loss

[0024]

[0025] Among them, ω is the loss balance weight, which is used to balance the learning effects of the above two different parts guiding the deep network model.

[0026] Furthermore, the knowledge embedding loss has the following expression;

[0027]

[0028] Among them, α and β are weight balance parameters, is the VAE loss, is the cross-modal loss, is the hidden layer space distribution loss.

[0029] Furthermore, the specific calculation method of the loss is as follows;

[0030] First, the visual encoder E v and the semantic encoder E a are used to encode the visual feature v and the scene category semantic vector a to obtain the means and variances μ (v) , Σ (v) and μ (a) , Σ (a) of the corresponding Gaussian hidden layer distributions; then, through the reparameterization operation, the above two sets of means and variances are transformed into the hidden layer space features Z (v) and Z (a) ; in order to enable the visual and semantic VAE models to correctly reconstruct the original features respectively, the hidden layer space feature Z (v) is passed through the visual decoder D v , Z (a) is passed through the semantic decoder D a to obtain the visual feature V v-v generated by the visual distribution and the semantic feature A a-a generated by the semantic distribution; in order to reconstruct the original features, the visual feature V v-v generated by the visual distribution should be consistent with the original visual feature v, while the semantic feature A a-a generated by the semantic distribution should be consistent with the original semantic feature a; in order to enable the knowledge guidance module to achieve this feature reconstruction goal, following the feature reconstruction loss construction method of the VAE model, the basic VAE loss is adopted to constrain the training:

[0031]

[0032] Among them, x and y respectively represent the input visual feature vector v and the scene category semantic vector a; i and j respectively represent the counts of the input visual feature vector v and the scene category semantic vector a in the visual feature vector set V and the semantic category vector set A; q v (Z(v) |x) and q a (Z (a) |y (j) ) represent the posterior distribution probabilities of the features obtained by the corresponding visual and semantic VAE encoders respectively; p v (x (i) |Z (v) ) and p a (y (j) |Z (a) ) represent the prior distribution probabilities of the features obtained by the corresponding visual and semantic VAE decoders respectively; p v (Z (v) ) and p a (Z (a) ) represent the prior distribution of the hidden layer features obtained by the corresponding visual and semantic VAE encoders, which is assumed to be a Gaussian distribution satisfying in the VAE model; represents maximizing the feature prior distribution probability p(·) given the feature posterior distribution probability q(·); D KL represents the KL divergence calculation function.

[0033] Further, the specific calculation method of the loss is as follows;

[0034] Based on the visual encoder E v and the semantic encoder E a to obtain the visual hidden layer space feature Z (v) and the semantic hidden layer space feature Z (a) , on this basis, pass Z (v) through the semantic decoder D a to obtain the semantic feature A v-a generated by the visual distribution, and pass Z (a) through the visual decoder D v to obtain the visual feature V a-v generated by the semantic distribution. Based on the premise that there is a cross-modal alignment relationship between the visual feature and the semantic feature, the semantic feature A v-a generated by the visual distribution is consistent with the original semantic feature a, and the visual feature V a-v generated by the semantic distribution is consistent with the original visual feature v; in order to enable the knowledge guidance module to achieve this reconstruction goal, a cross-modal loss is adopted to constrain the training:

[0035]

[0036] where, N represents the total number of training samples, x (i) and y (i) represent the visual feature vector v and the scene category semantic vector a corresponding to the same training sample i respectively, E vWith E a respectively represent the visual and semantic VAE encoders corresponding to this feature, D v and D a respectively represent the corresponding visual and semantic VAE decoders.

[0037] Furthermore, The specific calculation method of the loss is as follows;

[0038] The knowledge-guided learning module also needs to align the mean and variance μ (v) , Σ (v) of the original visual feature vector in the hidden layer space with the mean and variance μ (a) , Σ (a) of the original semantic feature vector in the hidden layer space. To enable the knowledge-guided module to achieve this reconstruction goal, a hidden layer space distribution loss is adopted to constrain the training:

[0039]

[0040] where N represents the total number of training samples, respectively represent the mean and variance of the visual feature vector v of sample i in the corresponding hidden layer space, respectively represent the mean and variance of the semantic vector a of the scene category to which sample i belongs in the hidden layer space, ‖·‖2 and ‖·‖ F respectively represent the 2-norm and the F-norm.

[0041] The present invention also provides a remote sensing image scene classification system for knowledge graph-guided deep network learning, including the following modules:

[0042] A land cover concept knowledge graph construction module for constructing a land cover concept knowledge graph;

[0043] A remote sensing scene semantic benchmark acquisition module for performing land cover concept semantic expression based on the land cover classification knowledge graph to obtain a remote sensing scene semantic benchmark, that is, a set A of remote sensing scene semantic category feature vectors;

[0044] A deep network optimization module that uses the remote sensing scene semantic category vector to guide the optimization of the deep network for remote sensing image scene classification;

[0045] The deep network is an open-source model, which is divided into a shallow part and a deep part. The input remote sensing image scene obtains visual features through the shallow part, and then obtains a classification result through the deep part; during the guidance and optimization process, first transform the visual features obtained by the shallow part into a set V of visual feature vectors, and then use the knowledge-guided learning module to perform cross-modal alignment constraints on the visual feature vector v and the scene semantic category vector a in the set, guiding the shallow part of the deep network to more effectively learn the shared features between remote sensing scene categories;

[0046] The knowledge-guided learning module includes a visual feature encoder E v , a visual feature decoder D v , a semantic feature encoder E a and a semantic feature decoder D a , sequentially realize the cross-modal alignment operation of the visual feature vector v and the scene category semantic category vector a, where v∈V, a∈A;

[0047] The classification module uses the optimized deep network to classify remote sensing image scenes.

[0048] Compared with the existing remote sensing image scene classification methods, the present invention has the following advantages and positive effects:

[0049] (1) This paper constructs a land cover concept knowledge graph to manage and utilize remote sensing scene prior knowledge in a more comprehensive, flexible, and convenient manner. Based on the land cover concept knowledge graph, a more accurate and effective remote sensing scene semantic benchmark is formed through land cover semantic representation learning based on the knowledge graph.

[0050] (2) The knowledge-guided learning module proposed in this paper effectively improves the remote sensing image scene classification capability of the deep network by imposing cross-modal alignment constraints between the remote sensing scene category semantic vectors and the shallow visual feature vectors of the deep network, and significantly reduces the deep network's dependence on training samples. Compared with existing remote sensing image scene classification methods, the knowledge-guided learning module demonstrates advanced remote sensing image scene classification performance and application flexibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 : is an overall flow chart of an embodiment of the present invention.

[0052] Figure 2 : A schematic diagram of the land cover concept knowledge graph according to an embodiment of the present invention.

[0053] Figure 3 : Schematic diagram of the semantic expression method of land cover concept in an embodiment of the present invention. DETAILED DESCRIPTION

[0054] In order to facilitate ordinary technicians in this field to understand and implement the present invention, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the implementation examples described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.

[0055] Please see Figure 1 The remote sensing image scene classification method provided by the embodiment of the present invention using a knowledge graph to guide deep network learning includes the following steps:

[0056] Step 1: Construct a land cover concept knowledge graph. To manage and utilize the prior knowledge of remote sensing scenes, the present invention first constructs a land cover concept knowledge graph. When constructing the land cover concept knowledge graph, various knowledge types such as domain standards, expert knowledge, and general knowledge are fully considered. Domain standards mainly include the national third land survey classification system. This system hierarchically classifies the remote sensing scene categories in the land cover classification task according to a secondary classification structure and strictly defines each category in words. According to this domain standard, the land cover concept knowledge graph can organize the semantic hierarchical information of remote sensing scene categories and obtain relatively accurate remote sensing scene category attribute information according to the category definition, ensuring the domain professionalism in the ontology structure. Expert knowledge mainly includes the domain knowledge sorted out by domain experts in the practice tasks of the remote sensing field, including knowledge such as the common internal structure of remote sensing scenes and the spatial relationships between remote sensing scenes. By constructing the land cover concept knowledge graph with expert knowledge, the knowledge graph can describe the land cover concept from a professional perspective and organize the prior knowledge of remote sensing scenes in a task-oriented form. The present invention constructs the land cover concept knowledge graph in the form of multi-person online collaborative editing with the self-developed remote sensing knowledge graph service engine software RKE, avoiding the limitations of collecting expert knowledge from a single domain task perspective and ensuring the generalization ability of the land cover classification knowledge graph. After constructing the land cover concept knowledge graph with domain standards and expert knowledge, the knowledge graph already has strong knowledge professionalism. To further enhance the completeness of the land cover concept knowledge graph in expressing the prior knowledge of remote sensing scenes, general knowledge is introduced to fill in the human common sense knowledge related to remote sensing scenes. Specifically, common sense triples related to remote sensing scene categories are retrieved according to the open-source concept knowledge graph ConceptNet and added to the constructed land cover concept knowledge graph to supplement the factual descriptions of remote sensing scene concepts and related relationships in the human common sense scenario. Based on the above construction principles, the land cover concept knowledge graph constructed by the present invention contains 8,521 entities and 13,892 triple facts, as Figure 2 shown.

[0057] Step 2: Perform semantic expression of land cover concepts based on the land cover classification knowledge graph. To generate a semantic benchmark for remote sensing scenes through the constructed land cover concept knowledge graph and use it to guide the remote sensing image scene classification model, the present invention adopts the knowledge graph representation learning model DistMult to perform semantic expression of land cover concepts and obtain the semantic category representation vector of remote sensing scenes, as Figure 3As shown in the figure. The constructed remote sensing scene knowledge graph G can be represented as a set of fact triples G = {<h, r, t>|h, t ∈ E, r ∈ R}, where E and R represent the entity set and relationship set of the remote sensing scene knowledge graph respectively, and h, r, and t represent the head entity, relationship, and tail entity of the triple. The DistMult model uses a triple scoring function based on a bilinear model to measure the semantic similarity between the head and tail entity pairs connected by the same relationship:

[0058]

[0059] Among them, M r represents the transformation matrix corresponding to the relationship r, and y h , y t represent the semantic representation vectors of the head and tail entities respectively. For a triple existing in a remote sensing scene knowledge graph, the score of S(h, r, t) should be close to 1, while for a non-existing triple, S(h, r, t) should be close to 0. Considering that the land cover concept knowledge graph has a large number of entities and relationships, without restrictions, the number of parameters of the relationship transformation matrix M r will be quite large, which is not conducive to the training of the knowledge graph representation model. Therefore, the relationship transformation matrix M r is restricted to a diagonal matrix to reduce the number of model parameters and facilitate the semantic expression of the land cover concept.

[0060] Based on the constraint relationship of the above land cover concept semantic expression method for knowledge graph representation learning, the overall land cover concept knowledge graph is subjected to a representation vector mapping. After learning, all triples G = {<h, r, t>|h, t ∈ E, r ∈ R} in the land cover concept knowledge graph can obtain the semantic expression vector set of all entities in the knowledge graph where is the total number of entities in the remote sensing scene knowledge graph. From Y e , entity semantic vectors corresponding to the remote sensing scene semantic categories are selected to obtain the remote sensing scene semantic category vector set

[0061] Step 3: The land cover concept knowledge graph guides the deep network learning. After obtaining the set A of remote sensing scene semantic category vectors, the remote sensing scene semantic category vectors need to be used to guide the deep network for remote sensing image scene classification. The remote sensing scene category semantic benchmark may make it difficult for the deep network model to learn the discriminative features between remote sensing scene categories with high semantic similarity, resulting in a decline in the final accuracy performance of the remote sensing scene classification model. To solve this problem, during the training process, a cross-modal alignment constraint between the visual feature vector and the scene semantic category vector is imposed on the shallow part of the deep network through the knowledge-guided learning module, guiding the shallow part of the deep network to more effectively learn the shared features between remote sensing scene categories. This part is guided by the knowledge embedding loss to guide the learning. In the deep part of the deep network, it is still guided by the remote sensing scene category label Y to maintain the ability of the deep network to learn the discriminative features between remote sensing scene categories. This part is guided by the data learning loss to guide the learning, and the data learning loss uses the conventional cross-entropy loss function. Specifically in the present invention, the open-source model efficientnet_b3 network is used as the backbone network. All convolutional layers before the activation function at position Swish54 in this model are regarded as the shallow part, and the shallow visual features extracted from it are led out and input into the knowledge-guided module proposed by the present invention for cross-modal alignment with the semantic vector. The model structure after this activation function is correspondingly regarded as the deep part. The overall loss function of the knowledge graph-guided deep network learning method proposed by the present invention is shown in the following formula:

[0062]

[0063] where ω is the loss balance weight, in order to balance the learning effects of the above two different parts guiding the deep network model.

[0064] Considering that the variational autoencoder (VAE) has characteristics and advantages such as being able to effectively reconstruct the feature hidden layer space distribution and the reconstructed hidden layer space distribution being continuous, the present invention uses the VAE model to complete the feature reconstruction and cross-modal alignment of the shallow visual features and semantic features of the deep network, so as to realize guiding the shallow part of the deep network to learn shared features through the remote sensing scene category semantic vectors. When specifically implemented, the present invention leads out the visual feature blocks obtained from the shallow part of the backbone network, transforms the visual feature blocks into visual feature vectors v through a global average pooling layer GAP, and then inputs them into the knowledge-guided learning module for cross-modal alignment learning between the visual feature vectors and the semantic category vectors.

[0065] The knowledge-guided learning module realizes the cross-modal alignment operation between the visual feature vector v and the scene category semantic vectors a, a ∈ A through two VAE models, and it mainly includes a visual feature encoder E v, a visual feature decoder D v , a semantic feature encoder E a and a semantic feature decoder D a . When performing cross-modal alignment, the knowledge-guided learning module mainly imposes three constraints:

[0066] (1) Reconstruction of the hidden layer space of the original features; First, the visual encoder E v and the semantic encoder E a are respectively used to encode the visual feature vector v and the scene category semantic vector a to obtain the means and variances μ (v) , Σ (v) and μ (a) , Σ (a) of the corresponding Gaussian hidden layer distributions. Then, through the reparameterization operation, the above two sets of means and variances are transformed into the hidden layer space features Z (v) and Z (a) . In order to enable the visual and semantic VAE models to correctly reconstruct the original features respectively, the hidden layer space features Z (v) are passed through the visual decoder D v , Z (a) are passed through the semantic decoder D a to obtain the visual features V v-v generated by the visual distribution and the semantic features A a-a generated by the semantic distribution. In order to reconstruct the original features, the visual features V v-v generated by the visual distribution should be consistent with the original visual feature v, while the semantic features A a-a generated by the semantic distribution should be consistent with the original semantic feature a. In order for the knowledge-guided module to achieve this feature reconstruction goal, following the feature reconstruction loss construction method of the VAE model, the basic VAE loss is adopted to constrain the training:

[0067]

[0068]

[0069] where x and y respectively represent the input visual feature vector v and the scene category semantic vector a; i and j respectively represent the counts of the input visual feature vector v and the scene category semantic vector a in the visual feature vector set V and the semantic category vector set A; q v (Z (v) |x) and q a (Z (a) |y (j) ) respectively represent the posterior distribution probabilities of the features obtained by the corresponding visual and semantic VAE encoders; p v (x (i) |Z (v) ) and p a(y (j) |Z (a) ) represent the prior distribution probabilities of the features obtained by the corresponding visual and semantic VAE decoders respectively; p v (Z (v) ) and p a (Z (a) ) represent the prior distributions of the hidden layer features obtained by the corresponding visual and semantic VAE encoders, which are assumed to follow a Gaussian distribution satisfying in the VAE model; represents maximizing the prior distribution probability p(·) of the features given the posterior distribution probability q(·) of the features; D KL represents the KL divergence calculation function.

[0070] (2) Reconstruction of the hidden layer space of cross-modal features; Based on the reconstruction of the hidden layer space of the original features, the knowledge guidance module needs to perform cross-modal reconstruction on the visual feature vectors and semantic category vectors that can correspond to each other. Specifically, based on the visual encoder E v and the semantic encoder E a to obtain the visual hidden layer space feature Z (v) and the semantic hidden layer space feature Z (a) , Z (v) is passed through the semantic decoder D a to obtain the semantic feature A v-a generated by the visual distribution, and Z (a) is passed through the visual decoder D v to obtain the visual feature V a-v generated by the semantic distribution. Based on the premise that there is a cross-modal alignment relationship between the visual features and the semantic features, the semantic feature A v-a generated by the visual distribution should be consistent with the original semantic feature a, and the visual feature V a-v generated by the semantic distribution should be consistent with the original visual feature v. To enable the knowledge guidance module to achieve this reconstruction goal, a cross-modal loss is adopted to constrain the training:

[0071]

[0072] where N represents the total number of training samples, x (i) and y (i) represent the visual feature vector v and the scene category semantic vector a corresponding to the same training sample i respectively, E v and E a represent the visual and semantic VAE encoders corresponding to the features respectively, and D v and D a represent the corresponding visual and semantic VAE decoders respectively.

[0073] (3) Alignment of the hidden layer space distribution; further, the knowledge-guided learning module also needs to align the mean and variance μ (v) and Σ (v) of the original visual feature vector in the hidden layer space with the mean and variance μ (a) and Σ (a) of the original semantic feature vector in the hidden layer space. To enable the knowledge-guided module to achieve this reconstruction goal, a hidden layer space distribution loss is adopted to constrain the training:

[0074]

[0075] where N represents the total number of training samples, represent the mean and variance of the visual feature vector v of training sample i in the corresponding hidden layer space respectively, represent the mean and variance of the semantic vector a of the scene category to which training sample i belongs in the corresponding hidden layer space respectively, and ‖·‖2 and ‖·‖ F represent the 2-norm and F-norm respectively.

[0076] Through the above three loss constraints, the knowledge-guided learning module learns and reconstructs the visual feature vector and the semantic category vector in the hidden layer space, and aligns them cross-modally, so as to guide the shallow part of the backbone network to learn shared visual features through the semantic category vector. The overall loss function of the knowledge-guided learning module is as follows:

[0077]

[0078] where α and β are weight balance parameters.

[0079] Step 4: Obtain the remote sensing image scene classification result. Use the remote sensing image scene classification method of the knowledge graph-guided deep network that has been optimized to perform remote sensing image scene classification. The optimized backbone network efficientnet_b3 can directly perform remote sensing image scene classification without relying on any prior knowledge. Assume that the original data of the remote sensing image is d, and the remote sensing image scene category obtained by classifying the given observation sample d is

[0080] The method of the present invention addresses the problem that existing remote sensing image scene classification models are unable to fully utilize the prior knowledge of remote sensing scenes. By constructing a land cover concept knowledge graph that includes domain standards, expert knowledge, and general knowledge, it manages and utilizes the prior knowledge of remote sensing scenes in a more comprehensive, flexible, and convenient manner. Furthermore, based on the land cover concept knowledge graph, a more accurate and effective remote sensing scene semantic benchmark is formed through the land cover semantic expression based on knowledge graph representation learning. Finally, the remote sensing scene semantic benchmark is used to guide the shallow part of the deep network to learn more effective shared features, reducing the deep network's dependence on training samples. The knowledge-guided learning module proposed in the present invention can effectively improve the remote sensing image scene classification capabilities of various deep networks and significantly reduce the deep network's dependence on training samples. Compared with existing remote sensing image scene classification methods, the knowledge-guided learning method proposed in the present invention demonstrates advanced remote sensing image scene classification performance and flexibility in application.

[0081] The present invention also provides a remote sensing image scene classification system guided by knowledge graph deep network learning, including the following modules:

[0082] Land cover concept knowledge graph construction module, used to construct land cover concept knowledge graph;

[0083] The remote sensing scene semantic benchmark acquisition module is used to express the semantics of land cover concepts based on the land cover classification knowledge graph and obtain the remote sensing scene semantic benchmark, that is, the remote sensing scene semantic category feature vector set A;

[0084] A deep network optimization module uses the semantic category vectors of remote sensing scenes to guide the optimization of deep networks for remote sensing image scene classification;

[0085] The deep network is an open-source model that is divided into a shallow layer and a deep layer. The input remote sensing image scene is used to obtain visual features through the shallow layer, and then the classification results are obtained through the deep layer. During the guided optimization process, the visual features obtained in the shallow layer are first transformed into a set of visual feature vectors V. Then, a knowledge-guided learning module is used to perform cross-modal alignment constraints on the visual feature vectors v in the set and the scene semantic category vectors a, guiding the shallow layer of the deep network to more effectively learn shared features between remote sensing scene categories.

[0086] The knowledge-guided learning module includes a visual feature encoder E v , a visual feature decoder D v , a semantic feature encoder E a and a semantic feature decoder D a , sequentially realize the cross-modal alignment operation of the visual feature vector v and the scene category semantic category vector a, where v∈V, a∈A;

[0087] A classification module that uses an optimized deep network for remote sensing image scene classification.

[0088] The specific implementation methods of each module correspond to the respective steps, and are not described in this invention.

[0089] It should be understood that the parts not elaborated in this specification belong to the prior art.

[0090] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art to which the present invention pertains may make various modifications or supplements to the described specific embodiments or use similar methods for substitution, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

Claims

1. A method for remote sensing image scene classification that guides deep network learning with a knowledge graph, characterized in that: Reducing the dependence of the remote sensing image scene classification method based on a deep network on manually labeled training samples through the guidance of prior knowledge in the knowledge graph, including the following steps: Step 1, construct a land cover concept knowledge graph; Step 2, perform land cover concept semantic expression based on the land cover classification knowledge graph to obtain a remote sensing scene semantic benchmark, that is, a set A of remote sensing scene semantic category feature vectors; Step 3, use the remote sensing scene semantic category vector to guide and optimize the deep network for remote sensing image scene classification; The deep network is an open-source model, which is divided into a shallow part and a deep part. The input remote sensing image scene obtains visual features through the shallow part and then obtains a classification result through the deep part; during the guidance and optimization process, first transform the visual features obtained by the shallow part into a set V of visual feature vectors, and then use the knowledge-guided learning module to perform cross-modal alignment constraints on the visual feature vector v and the scene semantic category vector a in the set, guiding the shallow part of the deep network to more effectively learn the shared features between remote sensing scene categories; In step 3, the open-source model efficientnet_b3 network is used as the deep network. All convolutional layers before the activation function at Swish54 in the open-source model are regarded as the shallow part, and the shallow visual features extracted by it are led out and input into the knowledge guidance module for cross-modal alignment with the semantic category vector. The model structure after the activation function is correspondingly regarded as the deep part; The knowledge-guided learning module includes a visual feature encoder E v , a visual feature decoder D v , a semantic feature encoder E a and a semantic feature decoder D a , and successively implements the cross-modal alignment operation of the visual feature vector v and the scene category semantic category vector a, where v ∈ V and a ∈ A; Step 4, use the optimized deep network for remote sensing image scene classification.

2. A method for remote sensing image scene classification that guides deep network learning based on a knowledge graph, characterized in that: The land cover concept knowledge graph constructed in step 1 includes: domain standards, expert knowledge, and general knowledge; Among them, the domain standard includes the third national land survey classification system, which hierarchically classifies the remote sensing scene categories in the land cover classification task according to the secondary classification structure and strictly defines each category in words; Expert knowledge includes the domain knowledge sorted out by domain experts in the practice tasks of the remote sensing field, including the common internal structures of remote sensing scenes and the spatial relationship knowledge between remote sensing scenes.

3. A method for remote sensing image scene classification that guides deep network learning based on a knowledge graph, characterized in that: In step 2, the knowledge graph representation learning model DistMult is adopted to perform land cover concept semantic expression and obtain the remote sensing scene semantic category representation vector. The constructed remote sensing scene knowledge graph G is represented as a set of fact triples G = {<h, r, t>|h, t ∈ E, r ∈ R}, where E and R represent the entity set and relationship set of the remote sensing scene knowledge graph respectively, and h, r, t represent the head entity, relationship, and tail entity of the triple respectively. The DistMult model uses a triple scoring function based on a bilinear model to measure the semantic similarity between the head and tail entity pairs connected by the same relationship: Among them, M r represents the transformation matrix corresponding to the relationship r, and y h , y t respectively represent the semantic representation vectors of the head and tail entities; Based on the constraint relationships of the semantic expression method of land cover concepts in the above knowledge graph representation learning, perform a representation vector mapping on the overall land cover concept knowledge graph. After learning, all triples G = {<h, r, t>|h, t ∈ E, r ∈ R} in the land cover concept knowledge graph can obtain the semantic expression vector set of all entities in the knowledge graph Among them is the total number of entities in the remote sensing scene knowledge graph. Select e from Y entity semantic vectors corresponding to remote sensing scene semantic categories to obtain the remote sensing scene semantic category feature vector set That is, the remote sensing scene semantic benchmark.

4. A method for classifying remote sensing image scenes by guiding deep network learning with a knowledge graph according to claim 1, characterized in that: The overall loss function adopted in Step 3 for guiding optimization includes the knowledge embedding loss of the shallow part of the deep network and the data learning loss of the deep part of the deep network Among them, the data learning loss uses the conventional cross-entropy loss function, and the overall loss function is shown as follows: Among them, ω is the loss balance weight, in order to balance the learning effects of the above two different parts on the deep network model.

5. A method for classifying remote sensing image scenes by guiding deep network learning with a knowledge graph according to claim 1, characterized in that: Knowledge embedding loss The expression is as follows; where α and β are weight balance parameters, is the VAE loss, is the cross-modal loss, is the latent space distribution loss.

6. A method for remote sensing image scene classification that guides deep network learning based on a knowledge graph, characterized in that: The specific calculation method of the loss is as follows; First, the visual encoder E v and the semantic encoder E a are used to encode the visual feature v and the scene category semantic vector a to obtain the means and variances μ (v) , Σ (v) and μ (a) , Σ (a) of the corresponding Gaussian hidden layer distributions; then, through the reparameterization operation, the above two sets of means and variances are transformed into the hidden layer space features Z (v) and Z (a) ; in order to enable the visual and semantic VAE models to correctly reconstruct the original features respectively, the hidden layer space feature Z (v) is passed through the visual decoder D v , Z (a) is passed through the semantic decoder D a to obtain the visual feature V v-v generated by the visual distribution and the semantic feature A a-a generated by the semantic distribution; in order to reconstruct the original features, the visual feature V v-v generated by the visual distribution should be consistent with the original visual feature v, while the semantic feature A a-a generated by the semantic distribution should be consistent with the original semantic feature a; in order to enable the knowledge guidance module to achieve this feature reconstruction goal, following the feature reconstruction loss construction method of the VAE model, the basic VAE loss is adopted to constrain the training: Among them, x and y respectively represent the input visual feature vector v and the scene category semantic vector a; i and j respectively represent the counts of the input visual feature vector v and the scene category semantic vector a in the visual feature vector set V and the semantic category vector set A; q v (Z (v) |x) and q a (Z (a) |y (j) ) respectively represent the posterior distribution probabilities of the features obtained by the corresponding visual and semantic VAE encoders; p v (x (i) |Z (v) ) and p a (y (j) |Z (a) ) respectively represent the prior distribution probabilities of the features obtained by the corresponding visual and semantic VAE decoders; p v (Z (v) ) and p a (Z (a) ) represent the prior distributions of the hidden layer features obtained by the corresponding visual and semantic VAE encoders, which are assumed to be Gaussian distributions satisfying in the VAE model; represents maximizing the prior distribution probability p(·) of the features given the posterior distribution probability q(·) of the features; D KL represents the KL divergence calculation function.

7. A method for remote sensing image scene classification that guides deep network learning based on a knowledge graph, characterized in that: The specific calculation method of the loss is as follows; Based on the visual encoder E v and the semantic encoder E a to obtain the visual hidden layer spatial feature Z (v) and the semantic hidden layer spatial feature Z (a) On this basis, Z (v) is passed through the semantic decoder D a to obtain the semantic feature A generated by the visual distribution v-a , and Z (a) is passed through the visual decoder D v to obtain the visual feature V generated by the semantic distribution a-v , based on the premise that there is a cross-modal alignment relationship between the visual feature and the semantic feature, the semantic feature A generated by the visual distribution v-a is consistent with the original semantic feature a, while the visual feature V generated by the semantic distribution a-v is consistent with the original visual feature v; in order to enable the knowledge guidance module to achieve this reconstruction goal, a cross-modal loss is adopted to constrain the training: Among them, N represents the total number of training samples, x (i) and y (i) respectively represent the visual feature vector v and the scene category semantic vector a corresponding to the same training sample i, E v and E a respectively represent the visual and semantic VAE encoders corresponding to this feature, D v and D a respectively represent the corresponding visual and semantic VAE decoders.

8. A method for classifying remote sensing image scenes by guiding deep network learning with a knowledge graph according to claim 5, characterized in that: The specific calculation method of the loss is as follows; The knowledge-guided learning module also needs to align the mean and variance μ (v) , Σ (v) of the original visual feature vector in the hidden layer space with the mean and variance μ (a) , Σ (a) of the original semantic feature vector in the hidden layer space. To enable the knowledge-guided module to achieve this reconstruction goal, a hidden layer space distribution loss is adopted to constrain the training: Among them, N represents the total number of training samples, respectively represent the mean and variance of the hidden layer space corresponding to the visual feature vector v of sample i, respectively represent the mean and variance of the hidden layer space corresponding to the semantic vector a of the scene category to which sample i belongs, ||·||2 and ||·|| F respectively represent the 2-norm and the F-norm.

9. A remote sensing image scene classification system for guiding deep network learning by a knowledge graph, characterized in that: Reducing the dependence of the remote sensing image scene classification method based on a deep network on manually labeled training samples through the guidance of prior knowledge in the knowledge graph, including the following modules: The land cover concept knowledge graph construction module is used to construct a land cover concept knowledge graph; The remote sensing scene semantic benchmark acquisition module is used to perform semantic expression of land cover concepts based on the land cover classification knowledge graph, and obtain the remote sensing scene semantic benchmark, that is, the remote sensing scene semantic category feature vector set A; The deep network optimization module uses the remote sensing scene semantic category vector to guide and optimize the deep network for remote sensing image scene classification; The deep network is an open-source model, which is divided into a shallow part and a deep part. The input remote sensing image scene obtains visual features through the shallow part, and then obtains the classification result through the deep part. During the guidance and optimization process, first transform the visual features obtained by the shallow part into the visual feature vector set V, and then use the knowledge-guided learning module to perform cross-modal alignment constraints on the visual feature vector v and the scene semantic category vector a in the set, guiding the shallow part of the deep network to more effectively learn the shared features between remote sensing scene categories; The open-source model efficientnet_b3 network is used as the deep network. All convolutional layers before the activation function at Swish54 in the open-source model are regarded as the shallow part, and the shallow visual features extracted by it are led out and input into the knowledge guidance module for cross-modal alignment with the semantic category vector. The model structure after the activation function is correspondingly regarded as the deep part; The knowledge-guided learning module includes a visual feature encoder E v , a visual feature decoder D v , a semantic feature encoder E a and a semantic feature decoder D a , and successively implements the cross-modal alignment operation between the visual feature vector v and the scene category semantic category vector a, where v ∈ V and a ∈ A; The classification module uses the optimized deep network to perform remote sensing image scene classification.