House floor space-time image information processing method based on multi-scale weak supervised learning
By combining deep convolutional networks and self-attention transformation networks, low-rank-sparse matrix decomposition and multi-layer graph model are used to solve the problem of spatiotemporal image data processing in traditional methods, efficient understanding and precise positioning of complex spatiotemporal data are achieved, and the identification and clustering capabilities of the model are improved.
Patent Information
- Application Number
- CN202510297862.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-01
AI Technical Summary
Traditional image recognition methods are difficult to effectively process large-scale, high-resolution, and complex and changeable spatio-temporal image data, lack precise positioning and quantitative modeling methods, and information interoperability is difficult when combining multi-scale weak-supervised learning and self-attention transformation networks.
A deep convolutional network and a self-attention transformation network are combined, and redundant information is removed through low-rank-sparse matrix decomposition, weighted category activation maps are generated, topological analysis and clustering are used for multi-layer graph models, and the 2-Wasserstein distance is used to measure the shape similarity of node neighborhood shapes.
The model's understanding of complex spatiotemporal data is improved, the identification accuracy and discrimination ability of house areas are improved, and the clustering and pattern discovery of spatiotemporal data of floor houses is realized, and the flexibility and adaptability of the model is enhanced.
Smart Images

Figure CN120236193A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a method for processing spatio-temporal image information of building floors by multi-scale weakly supervised learning. Background Art
[0002] In recent years, with the increasing demand for refined management of building distribution and floor structures in urban and regional planning, spatio-temporal data analysis and recognition technologies for buildings under different floor and structural conditions have gradually become the focus of research. Traditional image recognition and calculation methods are difficult to effectively process large-scale, high-resolution, and complex spatio-temporal image data, and there are also a lack of accurate positioning and quantitative modeling means for the key features existing in various floor building areas in the city.
[0003] To address the above challenges, researchers have proposed a multi-scale weakly supervised learning method, which combines a deep convolutional network and a self-attention transformation network to extract features of key regions of buildings and their floors in spatio-temporal images; and based on the idea of low-rank-sparse matrix decomposition, irrelevant redundant and noise information is removed, while maintaining good recognition accuracy when extracting the key structures of buildings. In addition, through topological data analysis (TDA) methods such as multi-layer network topology analysis and persistent diagram clustering (CPD), the spatio-temporal distribution pattern and spatial dependence relationship of floor buildings within an administrative region can be further quantified. The comprehensive application of the above series of new methods can not only finely depict the characteristics of floor building areas, but also accurately distinguish the spatio-temporal data of buildings with different floors or structures during subsequent retrieval and analysis.
[0004] However, when combining multi-scale weakly supervised learning with a self-attention transformation network, there are still great difficulties in achieving efficient connection and information exchange among the deep convolutional network, key discriminant matrix fusion analysis, and multi-layer network topology clustering. Summary of the Invention
[0005] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a method for processing spatio-temporal image information of building floors by multi-scale weakly supervised learning, which can adapt to spatio-temporal image features of different magnifications, comprehensively consider the specificity and commonality of data samples, and improve the flexibility and adaptability of the model in practical applications; through further analysis of spatio-temporal data of buildings with specific structures or floors, the spatio-temporal distribution characteristics of buildings can be deeply understood, providing support for subsequent retrieval and analysis.
[0006] To achieve the above purpose, the present invention provides the following solution:
[0007] A method for processing spatio-temporal image information of building floors by multi-scale weakly supervised learning, comprising:
[0008] Obtain the original image data containing the spatio-temporal information of houses within the administrative region, and use the class labels of houses and non-houses to screen and clean the original image data to obtain an image dataset for multi-scale weakly supervised learning;
[0009] Abstract the image dataset into a mixed form of a target low-rank matrix and a sparse matrix, use a deep convolutional network to perform multi-scale extraction of the spatio-temporal features of houses, and remove the irrelevant redundant information and noise in the spatio-temporal features of houses through a regularizer to obtain target features;
[0010] According to the identified target features, generate class activation maps at different scales respectively, and perform weighted fusion on the class activation maps at multiple scales to obtain the target region of interest for houses and a multi-scale fusion feature map containing the target region of interest for houses;
[0011] Convert the multi-scale fusion feature map into visual tokens, model the dependency relationship between the tokens through a self-attention mechanism and a learnable projection matrix, obtain the encoded result output by the self-attention layer, and decode the encoded result in combination with a key discriminant matrix to obtain the floor house structure information;
[0012] Determine the house region or the floor region as a graph node, and determine the floor house structure information as the feature of the graph node. Use a multi-layer graph method to represent the graph nodes and the connections between the graph nodes under different floor properties or resolution conditions to obtain a multi-layer graph model;
[0013] Perform multi-resolution filtering and topological analysis on the multi-layer graph model, use the 2-Wasserstein distance to measure the shape similarity of the node neighborhoods, and obtain persistent clusters under different thresholds to achieve clustering and pattern discovery of the spatio-temporal data of the floor houses.
[0014] Preferably, the target features include house and floor categories.
[0015] Preferably, the method for constructing the key discriminant matrix includes:
[0016] Define the corresponding class labels according to the types of houses and the features of non-house regions; the class labels will be used for discriminant analysis;
[0017] Perform statistical analysis on the extracted target features to identify the features that are significantly different between different classes;
[0018] Select the features that are significantly different between classes as key features;
[0019] Organize the selected key features into a matrix form to form a key discrimination matrix; the rows of the key discrimination matrix represent different housing categories, and the columns represent the extracted key features; each element of the key discrimination matrix represents the statistical value of the key feature under a specific category or the correlation measure between the key feature and the category.
[0020] Preferably, the expression of the optimization function of the regularizer is:
[0021]
[0022] Where is the loss between the true value and the predicted value, R(f θ (M)) is a regularization loss, f θ (M) is the predicted output obtained by processing the input spatio-temporal image M through the model parameters θ, that is, the softmax segmentation result generated by the network, f θ (M) ∈ [0, 1] |Ω|×K , R(1 - f θ (M)) is another regularization loss for processing the supplementary information of the model output; K represents the number of categories; Ω represents all pixel positions of the image; λ and μ are both weight parameters of the regularization loss, and C is the label corresponding to the input spatio-temporal image M.
[0023] Preferably, the expression of the visual marker is:
[0024] T = UOFTMAX HW (XW A ) T X
[0025] Where T is the visual marker, X is the matrix representation of the feature map extracted by the convolutional network, W A is a learnable weight and H, W, and C’ respectively represent the respective dimensions of the feature, L represents the number of visual markers T, and L << HW.
[0026] Preferably, the expression of the 2-Wasserstein distance is:
[0027]
[0028] Where W2(x i , x j ) represents the 2-Wasserstein distance between node x i and node x j , PD(i) and PD(j) respectively represent node x i and node x jPersistence diagram, inf γ represents the infimum over all possible bijective mappings γ, ||x - γ(x)|| ∞ represents the distance between node x and its corresponding point under its mapping γ in the infinity norm, where γ is controlled by all bijective mappings from PD(i) ∪ Δ to PD(j) ∪ Δ, and Δ represents a set containing all point pairs (x, x).
[0029] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0030] The present invention provides a method for processing spatio-temporal image information of house floors based on multi-scale weakly supervised learning. The method includes: obtaining original image data containing spatio-temporal information of houses within an administrative region, and screening and cleaning the original image data using class labels of houses and non-houses to obtain an image data set for multi-scale weakly supervised learning; abstracting the image data set into a mixed form of a target low-rank matrix and a sparse matrix, extracting multi-scale spatio-temporal features of houses using a deep convolutional network, and removing irrelevant redundant information and noise in the spatio-temporal features of houses through a regularizer to obtain target features; generating class activation mapping diagrams at different scales respectively according to the identified target features, and performing weighted fusion on the class activation mapping diagrams at multiple scales to obtain a target region of interest for houses and a multi-scale fusion feature map containing the target region of interest for houses; converting the multi-scale fusion feature map into visual tokens, modeling the dependency relationship between the tokens through a self-attention mechanism and a learnable projection matrix to obtain an encoded result output by the self-attention layer, and decoding the encoded result in combination with a key discriminant matrix to obtain floor house structure information; determining the house region or the floor region as a graph node, and determining the floor house structure information as the feature of the graph node, and characterizing the graph nodes and the connections between the graph nodes under different floor properties or resolution conditions using a multi-layer graph method to obtain a multi-layer graph model; performing multi-resolution filtering and topological analysis on the multi-layer graph model, measuring the shape similarity of the node neighborhoods using the 2-Wasserstein distance, and obtaining persistent clusters under different thresholds to achieve clustering and pattern discovery of the spatio-temporal data of the floor houses.The present invention extracts multi-scale spatio-temporal features of houses through a deep convolutional network, which can effectively capture feature information at different scales, thereby improving the model's ability to understand complex spatio-temporal data; a regularizer is used to remove irrelevant redundant information and noise to ensure that the extracted features are more accurate and reliable, thereby enhancing the model's prediction performance; by generating a weighted class activation mapping diagram, the model can better identify and locate the house areas of interest, improving the discrimination ability between house and non-house areas; the self-attention mechanism is used to model the dependence relationship between visual tokens, which can better capture the interaction between features and enhance the model's expressive ability; by determining the house area or floor area as graph nodes and using a multi-layer graph to represent the graph nodes and their connections under different floor properties or resolution conditions, the structural features of house spatio-temporal data can be effectively analyzed and understood; the 2-Wasserstein distance is used to measure the shape similarity of node neighborhoods, and persistent clusters can be obtained under different thresholds to achieve clustering and pattern discovery of floor house spatio-temporal data, helping to identify potential structural patterns; this method can adapt to spatio-temporal image features of different magnifications, comprehensively consider the specificity and commonality of data samples, and improve the flexibility and adaptability of the model in practical applications; through further analysis of specific structural or floor house spatio-temporal data, the spatio-temporal distribution characteristics of houses can be deeply understood, providing support for subsequent retrieval and analysis. Brief Description of the Drawings
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0032] Figure 1 It is a flowchart of the method provided by the embodiment of the present invention;
[0033] Figure 2 It is a flowchart for accurately positioning and identifying floor house spatio-temporal data provided by the embodiment of the present invention. Detailed Embodiments
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0035] The object of the present invention is to provide a method for processing spatio-temporal image information of building floors based on multi-scale weakly supervised learning, which can adapt to spatio-temporal image features at different magnifications, comprehensively consider the specificity and commonality of data samples, and improve the flexibility and adaptability of the model in practical applications; through further analysis of the spatio-temporal data of houses with specific structures or floors, the spatio-temporal distribution characteristics of houses can be deeply understood, providing support for subsequent retrieval and analysis.
[0036] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0037] Figure 1 The following is a flowchart of the method provided by the embodiment of the present invention. As Figure 1 shown, the present invention provides a method for processing spatio-temporal image information of building floors based on multi-scale weakly supervised learning, including:
[0038] Step 100: Obtain the original image data containing the spatio-temporal information of houses within the administrative region, and use the class labels of houses and non-houses to screen and clean the original image data to obtain an image data set for multi-scale weakly supervised learning;
[0039] Step 200: Abstract the image data set into a mixed form of a target low-rank matrix and a sparse matrix, use a deep convolutional network to perform multi-scale extraction of the spatio-temporal features of houses, and remove irrelevant redundant information and noise in the spatio-temporal features of houses through a regularizer to obtain target features;
[0040] Step 300: According to the identified target features, generate class activation maps at different scales respectively, and perform weighted fusion on the class activation maps at multiple scales to obtain the target interested house regions and the multi-scale fusion feature maps containing the target interested house regions;
[0041] Step 400: Convert the multi-scale fusion feature maps into visual tokens, model the dependency relationships between the tokens through a self-attention mechanism and a learnable projection matrix, obtain the encoded results output by the self-attention layer, and decode the encoded results in combination with a key discriminant matrix to obtain the floor house structure information;
[0042] Step 500: Determine the house regions or floor regions as graph nodes, and determine the floor house structure information as the features of the graph nodes. Use a multi-layer graph to represent the graph nodes and the connections between the graph nodes under different floor properties or resolution conditions to obtain a multi-layer graph model;
[0043] Step 600: Perform multi-resolution filtering and topological analysis on the multi-layer graph model, use the 2-Wasserstein distance to measure the shape similarity of node neighborhoods, and obtain persistent clusters at different thresholds to achieve clustering and pattern discovery of floor housing spatio-temporal data.
[0044] Optionally, in step 100 of this embodiment, researchers or data annotators will mark each image according to the features in the image (such as the shape, color, texture, etc. of the building), and determine which areas belong to houses and which areas belong to non-houses (such as roads, green spaces, etc.). This annotation can be carried out through manual annotation or training using existing annotation datasets, so as to generate a cleaned dataset containing house and non-house class labels, providing a basis for subsequent multi-scale weakly supervised learning.
[0045] House spatio-temporal information refers to the comprehensive data of time and space characteristics related to houses. This includes the geographical location of the house (such as longitude and latitude), building height, number of floors, house type (such as residential, commercial, etc.), and dynamic information that changes over time (such as the usage status of the house, changes in the surrounding environment, etc.). By analyzing this spatio-temporal information, the characteristics and behaviors of houses under different time and space conditions can be better understood, providing an important basis for subsequent image processing and feature extraction.
[0046] Specifically, in step 200 of this embodiment, the features of key regions in the spatio-temporal data are constructed according to the spatio-temporal image-level labels, irrelevant redundant information and noise are removed, and the original spatio-temporal data information is abstracted into two major categories of data matrices. The target structure information is a low-rank matrix, and the redundant and noise information is a sparse matrix. The two matrices are solved respectively, and finally the feature information of the target structure data is obtained, and a specific regularizer is proposed for this loss. Let the spatio-temporal image be M, and its label be C. Let f θ (M) be the output of the segmentation network parameterized by θ. Generally, the optimization problem trained by CNN using the joint regularization loss is:
[0047] Among them, is the loss between the true value and the predicted value, R(U) is the regularization loss, and the parameter U = f θ (M) ∈ [0, 1] |Ω|×K, namely the softmax segmentation result generated by the network. To fuse the feature information of spatio-temporal images at different magnifications into the learning process of the algorithm, the model input integrates the spatio-temporal image information of multiple magnifications to comprehensively pay different attentions to the spatio-temporal data of floor houses at high magnification and the spatio-temporal distribution of block houses at medium and low magnifications, fully considering the specificity and commonality of data samples and learning the key features of the data. At the same time, the business analysis process of data retrieval is simulated, and different attention weights are given to the spatio-temporal image features at different magnifications, fully considering the data features under the spatio-temporal images of each magnification. Its corresponding optimization function for multi-magnification regularization loss is as follows:
[0048]
[0049] Among them, M d represents the spatio-temporal image input at d magnification, and f θ,η represents the feature calculation under the attention weight of η with parameter θ; in addition, η is mainly calculated by Softmax(f θ (M)). By learning and optimizing the above parameters through a deep convolutional network, it has the ability to efficiently extract features, enabling the model to learn data priors better and faster.
[0050] Furthermore, after the parametric model completes learning, according to the recognized target category, the features at multiple scales are fused and calculated to obtain a weighted class activation mapping diagram, and finally the target interested house area is processed.
[0051] Due to the large-scale nature of the spatio-temporal images of administrative regions, based on the above research foundation, the spatio-temporal distribution area of block houses of interest to houses is obtained, as well as the key features that convey the discriminant characteristics between house and non-house areas. In the subsequent retrieval and analysis process, this spatio-temporal distribution feature of houses needs to be further analyzed, including but not limited to the analysis of the spatio-temporal data of houses with specific structures or floors. Since the spatio-temporal data of houses with specific structures and floors are not only numerous in types but also small in proportion, and at the same time, the spatio-temporal data of some floor houses, such as the spatio-temporal data between different types of floor houses, are very similar in morphological performance, the above problems pose a huge challenge to whether the algorithm can accurately locate and identify.
[0052] Optionally, according to the differences and correlations between the spatio-temporal data of floor houses and the structure itself, a transformation network based on self-attention is used to implement the encoding and decoding calculations of the spatio-temporal data and structures of each floor of the spatio-temporal image, and at the same time, the key discriminant matrix obtained from the previous research is fused and analyzed. The specific implementation is as follows:
[0053] First, a series of feature maps extracted by the convolutional network are converted into visual tokens T, T = UOFTMAX HW (XW A )T Among them, W A is a learnable weight and H, W, and C respectively represent the respective dimensions of the features, L represents the number of visual markers T, and L << HW.
[0054] After obtaining T, self-attention transformation will be used to model the dependency relationship between Ts and project it onto the dimension of the normal feature map. At the same time, combined with the previous key discrimination matrix,
[0055] X out = X in + UOFTMAX L ((X in W Q )(TW K )) T )T + G; where W Q and W K are respectively learnable weight parameters. After constructing the feature relationships of the spatio-temporal image, a large amount of data learning is carried out to realize the recognition and positioning of the spatio-temporal data, structure of different types of floor houses, and the processing method is as Figure 2 shown:
[0056] Optionally, in order to analyze the spatial distribution expression of the coexistence of multi-floor house spatio-temporal data and structure types to quantitatively characterize the interaction between the spatio-temporal data of floor houses and non-floor houses, structures, it is proposed to adopt the method of multi-layer graph modeling of multi-layer networks. The multi-layer graph is a set of adjacency matrices of single-layer graphs with weights, including the interaction of intra-layer relationships and inter-layer relationships. The specific implementation includes the construction of a multi-layer graph topological space model and the clustering calculation of a multi-layer network, and finally realizes the construction of a spatial distribution expression model of the spatio-temporal data of floor houses.
[0057] Specifically, in this embodiment, the topological space heterogeneous high-order characteristic modeling of the multi-layer network is first carried out. The single-layer graph network is defined as: G = (V, E, ω), where V is the set of nodes, is the set of edges. The total number of points in graph G is n = |V|. is an edge weight function, and the weight of edge e uv ∈ E is ω uv , and the adjacency matrix A is a symmetric matrix, A ij = A ji . According to the definition of the single-layer graph, a multi-layer network G can be constructed. G consists of non-overlapping m layers, and each layer is modeled by a weighted graph G i with an adjacency matrix of A i , i = 1,..., m. The set A = {A1, A2,..., A mThe elements in} are called intra-layer matrices, representing the connections within a single layer, i.e., intra-layer connections. For modeling the relationship between two graphs, G k and G l and their adjacency matrices can be respectively represented as, A k and A l (k, l = 1, 2, …, m; k ≠ l), which represents the one-to-one symmetric internal connections between the nodes of two related graphs. Thus, a set D p ={A l,k , k ≠ l} is obtained, representing the edges between the nodes of different layers. p represents the number of related graphs. In summary, a multi-layer network g has an inter-layer connection set E M (G) that connects the cross-layer nodes. For an edge (u, v) ∈ E M (G), there are u ∈ V(G k ) and v ∈ V(G l ), and k ≠ l. The defined super-adjacency matrix of the multi-layer network g has a block matrix structure:
[0058]
[0059] The diagonal elements in the set A are intra-layer matrices, and the non-diagonal elements A kl (k, l = 1, 2, …, m; k ≠ l) represent the inter-layer connections that connect the nodes in the G k layer to the nodes in the G l layer. In the definition, the same-layer nodes represent the spatio-temporal data of houses of the same type on the same floor, and the connections between different layers represent the spatial connection relationships between the spatio-temporal data or structures of houses of different types on different floors. Taking the geographical distribution structure and the spatio-temporal data of houses on the floors as an example, the establishment of the inter-layer relationship between the spatio-temporal data and the structure of the houses on the floors can obtain the value of the non-diagonal element A kl based on the size of the spatial distance. The houses or the spatio-temporal data of the immune floors that are close to the geographical distribution have a strong connection with the structure layer, and vice versa, a weak connection. The diagonal elements are also intra-layer matrices and are obtained through the Euclidean distance between the spatio-temporal data of the houses on the floors.
[0060] Let X and Y represent two spatio-temporal image feature information, representing the spatio-temporal image information of the observations of the same variable at different nodes (when A is defined) or the records of different variables at the same node (when D is defined). For example, in, X and Y represent a certain type of spatio-temporal data information of two retrieval responses or prognoses or the spatial distribution of different types of spatio-temporal data of the same retrieval outcome situation. The goal is to find a non-linear transformation between X and Y that maximizes the correlation between the transformed variables. This step is achieved by using alternating conditional expectation, which is an algorithm for finding the optimal additive model, making the maximum linear effect generated between the transformed response and the predictive variable, defined as follows: ω = r* (X, C) = max φ,θ r[φ(X), θ(C)]; where r* is the maximum correlation between the optimal transformations φ(X) and θ(C) of X and Y. To find such transformations, it can also be transformed into the error e 2 (θ, φ) = E[θ(C) - φ(X)] 2 alternating minimization, first find the minimization with respect to θ(C) (while keeping E(θ 2 ) = 1), and then obtain the minimization with respect to φ(X) according to the given θ(C). The result can be written as: The minimization process starts from an initial value of the function (θ(C) = C / ||C||), and each iteration performs the minimization of a pair of single functions until a complete iteration process can no longer make e 2 decrease. The algorithm converges to the optimal transformations θ and φ.
[0061] Furthermore, considering that extracting meaningful information from complex networks requires a large amount of computation and memory, to solve these two problems, this embodiment transforms the network into a low-dimensional space through node embedding and retains its structural information. The project plans to use the multilayered network embedding method (MANE) based on the idea of matrix factorization to achieve dimensionality reduction. To describe MANE, define to represent all nodes in the i-th layer (i = 1,..., m), n i is the number of nodes in the i-th layer, and d i is the embedding dimension. The goal of this algorithm is to find a low-dimensional vector representation that can retain the proximity of nodes in the network topology. The objective function is as follows: where D ij represents the network connection between the i-th layer and the j-th layer; L i is the normalized Laplacian matrix, and the embedding representation F i is a matrix that satisfies Embedding is essentially obtained by concatenating the top d i eigenvectors of L i +. The first term corresponds to embedding a single-layer network into a low-dimensional representation to maintain the proximity of nodes in the original single-layer structure. The second term corresponds to embedding cross-layer connections (i.e., connections across single-layer networks). Here, the interaction between the posterior features of nodes in different layers is used to approximately simulate the real dependence connections. One advantage of this matrix factorization technique is that it reduces the number of debugging parameters.
[0062] Then, based on the multi-layer network topology analysis method of persistent graph clustering, deduce the spatio-temporal data types of multi-storey houses and the prior knowledge of the administrative regions of the spatio-temporal data of the house floors. Forming clusters based on shape dynamics helps to discover clusters of persistent nodes with similar patterns. The concept of topological data analysis (TDA) is introduced into complex multi-layer network analysis. Assume a weighted graph G. If a threshold ∈ j > 0 is selected and only the edges with weights satisfying ω uv ≤ ∈ j are retained, a graph G with an adjacency matrix of can be obtained. j o If the threshold is changed to ∈1 < ∈2 < … < ∈ n a hierarchical nested sequence of graphs is obtained which is called "network filtration". Among them, the Vietoris-Rips (VR) complex is the most widely used simplicial complex. The VR complex at threshold v j is defined as By means of network filtration, the changes in network topology induction are evaluated to detect persistent features on a large range of thresholds ∈ j . The goal is to detect persistent features exceeding different thresholds ∈, and such persistent features are the features of the internal spatial distribution.
[0063] Furthermore, most current clustering methods for multi-layer networks are based on spectral decomposition to embed the graph into Euclidean space, without explicitly considering local graph geometry and topology. This embodiment innovatively proposes a multi-layer network clustering method, which starts from the perspective of the shape similarity of data recorded at multiple resolutions and performs clustering calculations on multi-layer networks in an unsupervised manner. In order to quantify the shape dynamics of multi-layer networks at the evolutionary similarity scale, a geographical distribution tool of TDA is introduced into the method CPD. The basic principle is as follows: If the local neighborhoods of two points are similar in shape at all resolution scales, the distance between them is close enough to be clustered into a cluster.
[0064] The specific implementation steps of shape comparison are as follows:
[0065] 1) Consider X n =(x1,…,x n ) in some metric space (X, D);
[0066] 2) Set resolution thresholds v1 < v2 < … < v k , and establish a VR filtration
[0067] 3) Calculate χ in the form of persistent diagrams PD(i), i = 1,..., ni Local topological summary;
[0068] 4) For all neighborhoods N(i) of x i and all neighborhoods N(j) of x j where i, j = 1, …, n, calculate the pairwise topological or data shape dissimilarity as the 2-Wasserstein distance between their respective persistence diagrams PD(i) and PD(j);
[0069] where γ is controlled by all bijective mappings from PD(i) ∪ Δ to PD(j) ∪ Δ, and the multiplicity is calculated. The 2-Wasserstein distance allows quantifying the similar shapes of the neighborhoods of two nodes, and calculates and compares all cycles, holes, and other topological features in the neighborhoods of each node;
[0070] 5) Construct a distance graph G based on W2(N(i), N(j)), where i, j = 1, 2, ···, n, and the adjacency matrix is A:
[0071]
[0072] The truncation point κ is defined by an elbow graph or cross-validation;
[0073] 6) The connected components of G are the clusters of the clustering results. Therefore, CPD utilizes the distance function and the local spatial information around the points.
[0074] The beneficial effects of the present invention are as follows:
[0075] (1) The present invention performs multi-scale extraction of the spatio-temporal features of houses through a deep convolutional network, which can effectively capture the feature information at different scales, thereby improving the model's understanding ability of complex spatio-temporal data.
[0076] (2) The present invention uses a regularizer to remove irrelevant redundant information and noise, ensuring that the extracted features are more accurate and reliable, and thus improving the prediction performance of the model.
[0077] (3) By generating a weighted class activation mapping graph, the model of the present invention can better identify and locate the house areas of interest, improving the discrimination ability between house and non-house areas.
[0078] (4) The present invention uses a self-attention mechanism to model the dependence relationships between visual tokens, which can better capture the interactions between features and enhance the expressive ability of the model.
[0079] (5) By determining the housing area or floor area as graph nodes and using a multi-layer graph method to represent the graph nodes and their connections under different floor properties or resolution conditions, the present invention can effectively analyze and understand the structural characteristics of housing spatio-temporal data.
[0080] (5) The present invention uses the 2-Wasserstein distance to measure the shape similarity of node neighborhoods, can obtain persistent clusters under different thresholds, realizes the clustering and pattern discovery of floor housing spatio-temporal data, and helps to identify potential structural patterns.
[0081] (6) The method of the present invention can adapt to spatio-temporal image features of different magnifications, comprehensively consider the specificity and commonality of data samples, and improves the flexibility and adaptability of the model in practical applications.
[0082] (7) By further analyzing the spatio-temporal data of specific structures or floor housing, the present invention can deeply understand the spatio-temporal distribution characteristics of housing and provide support for subsequent retrieval and analysis.
[0083] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same and similar parts between the various embodiments, reference can be made to each other.
[0084] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A multi-scale weakly supervised learning method for processing spatiotemporal image information of house floors, characterized in that: include: Obtaining original image data containing spatiotemporal information of houses in the administrative area, and filtering and cleaning the original image data using category labels of houses and non-houses to obtain an image dataset for multi-scale weakly supervised learning; Abstracting the image dataset into a hybrid form of a target low-rank matrix and a sparse matrix, using a deep convolutional network to perform multi-scale extraction of house spatiotemporal features, and removing irrelevant redundant information and noise in the house spatiotemporal features through a regularizer to obtain target features; According to the identified target features, class activation maps are generated at different scales respectively, and the class activation maps at multiple scales are weightedly fused to obtain a target house area of interest and a multi-scale fused feature map containing the target house area of interest; The multi-scale fusion feature map is converted into visual tags, the dependency relationship between the tags is modeled through the self-attention mechanism and the learnable projection matrix, the encoding result output by the self-attention layer is obtained, and the encoding result is decoded in combination with the key discriminant matrix to obtain the floor structure information of the building; Determine the house area or floor area as a graph node, and determine the floor house structure information as the feature of the graph node, use a multi-layer graph method to characterize the graph nodes and the connections between the graph nodes under different floor properties or resolution conditions, and obtain a multi-layer graph model; Multi-resolution filtering and topological analysis are performed on the multi-layer graph model, the shape similarity of the node neighborhood is measured using the 2-Wasserstein distance, and persistent clusters are obtained under different thresholds to achieve clustering and pattern discovery of floor house spatiotemporal data.
2. The multi-scale weakly supervised learning house floor spatiotemporal image information processing method according to claim 1 is characterized in that: The target features include house and floor categories.
3. The multi-scale weakly supervised learning house floor spatiotemporal image information processing method according to claim 1 is characterized in that: The method for constructing the key discriminant matrix includes: Define the corresponding category labels according to the type of housing and the characteristics of non-housing areas; the category labels will be used for discriminant analysis; Performing statistical analysis on the extracted target features to identify features with significant differences between different categories; Select features that differ significantly between categories as key features; The selected key features are organized into a matrix form to form a key discriminant matrix; the rows of the key discriminant matrix represent different housing categories, and the columns represent the extracted key features; each element of the key discriminant matrix represents the statistical value of the key feature under a specific category or the correlation measure between the key feature and the category.
4. The multi-scale weakly supervised learning house floor spatiotemporal image information processing method according to claim 1 is characterized in that: The expression of the optimization function of the regularizer is: minutes θ l(f θ (M),C)+λ·R(f θ (M))+μ·R(1-f θ (M)) Among them, l(U,C) is the loss between the true value and the predicted value, R(f θ (M)) is a regularization loss, f θ (M) is the predicted output obtained by processing the input spatiotemporal image M through the model parameters θ, that is, the softmax segmentation result generated by the network, f θ (M)∈[0,1] |Ω|×K , R(1-f θ (M)) is another regularization loss used to process the supplementary information of the model output; K represents the number of categories; Ω represents all pixel positions of the image; λ and μ are the weight parameters of the regularization loss, and C is the label corresponding to the input spatiotemporal image M.
5. The multi-scale weakly supervised learning house floor spatiotemporal image information processing method according to claim 1, characterized in that: The expression of the visual mark is: T=UOFTMAX HW (XW A ) T X Among them, T is the visual mark, X is the matrix representation of the feature map extracted by the convolutional network, W A are learnable weights and H, W, and C' represent the dimensions of the features, respectively. L represents the number of visual markers T, and L< <HW。 6. The multi-scale weakly supervised learning house floor spatiotemporal image information processing method according to claim 1 is characterized in that: The expression of the 2-Wasserstein distance is: Among them, W2(x i ,x j ) represents node x i and node x j , PD(i) and PD(j) represent the 2-Wasserstein distance between nodes x and i and node x j Persistence graph of γ represents the infimum of all possible bijective mappings γ, ||x-γ(x)|| ∞ represents the distance between a node x and its corresponding point under its mapping γ in the infinite norm, where γ is controlled by all bijective mappings from PD(i)∪Δ to PD(j)∪Δ, and Δ represents a set of all point pairs (x,x).