Building function classification method and equipment

By using technical means such as RoBERTa model, graph convolution network and SAM large model in the functional classification of buildings, multimodal data feature expression and feature fusion are carried out, and the problem of low efficiency in the existing technology is solved, and the functional classification of urban buildings with high accuracy is achieved.

CN120217293APending Publication Date: 2025-06-27CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510286706.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, multimodal data feature expression and feature fusion methods are inefficient, making it difficult to achieve high-accuracy functional classification of urban buildings.

Method used

The RoBERTa model is used to extract the sequence features of the building area, combine the graph convolution network and SAM large model to perform cross-modal feature fusion, and efficient fusion of building-level and regional-level features is achieved through multi-level cross-modal fusion network and SE-Net network.

Benefits of technology

A high-accuracy functional classification of urban buildings is achieved, and the buildings' own characteristics in multi-source and multi-modal data and their interactive characteristics with the regional environment are fully explored and captured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217293A_ABST
    Figure CN120217293A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cartographic synthesis, and discloses a building function classification method and equipment, and the method comprises the steps: extracting the sequence features of a region through employing a RoBERTa model, inputting the sequence features into a graph convolution network based on an attention mechanism, obtaining the feature representation of region POI information, screening a region street scene image through employing an SAM large model, and reserving a building mask. The method comprises the following steps: extracting visual features of a building mask, carrying out cross-modal feature fusion based on Cross-Attention model fusion, carrying out representation learning on POI text features by using a BERT model, representing building geometric features by using Transform, processing sign-in data by using an Inform model to extract human activity features, and fusing multi-modal data through a multi-level cross-modal fusion network. And region-level and building-level cross-level feature fusion is carried out by using SE-Net to obtain building category representation, and the accuracy is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cartographic generalization, and particularly to a method and device for classifying building functions. Background Art

[0002] Maps are recognized as one of the three major universal languages in the world (painting, music, maps), and are masterpieces for interpreting the world. Cartographic generalization refers to the engineering, technology, and science of abstracting and generalizing spatial data when large-scale spatial data is reduced to small-scale spatial data. It is one of the basic means for spatial data scale transformation, integration and fusion, analysis and mining, etc. Due to its complexity and the difficulty of solving, cartographic generalization has always been the most challenging and innovative research field in modern cartography, whether in the past, present, or future.

[0003] Buildings, as an important part of the geographic spatial vector database, are one of the bases for expressing spatial phenomena, supporting spatial analysis, and providing spatial services, and are the core elements of large-scale urban maps, having an important impact on the effect of map expression. With the vigorous development of computer cartography and geographic information systems, building merging has evolved from traditional manual cartography to automatic cartographic generalization and has made great progress.

[0004] Currently, users have put forward higher requirements for the timeliness and consistency of digital spatial information. With the development of geographic information systems (GIS) and remote sensing technology, the representation of maps at different scales has become increasingly important. Summarizing the related research on building function classification, the current research lacks efficient multi-modal data feature expression and feature fusion methods.

[0005] The above content is only used to assist in understanding the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention

[0006] The main purpose of the present invention is to provide a method and device for classifying building functions, aiming to solve the technical problem of low efficiency of multi-modal data feature expression and feature fusion methods in the prior art.

[0007] To achieve the above object, the present invention provides a method for classifying building functions, and the method for classifying building functions includes:

[0008] Using the RoBERTa model to extract the sequence features of the entire area where the target building is located;

[0009] Inputting the sequence features into a graph convolutional network based on the attention mechanism to obtain the feature representation of regional POI information;

[0010] According to the above-mentioned regional POI information features representation, use the SAM large model to screen the regional street view images and retain the building mask;

[0011] According to the building mask, use an image encoder to extract the visual features of the building mask in the regional street view image;

[0012] According to the regional POI information features and the visual features of the building mask, perform cross-modal feature fusion based on the Cross-Attention model fusion to obtain the feature expression of the area where the building is located;

[0013] Use the BERT model to perform representation learning on the POI text features in the regional POI information features;

[0014] Use Transformer to represent the extracted building geometric features;

[0015] Use the Informer model to process the platform check-in data and extract human activity features;

[0016] Based on the generated POI text features, the building geometric features, and the human activity features, fuse the multi-modal data through a multi-level cross-modal fusion network to obtain building-level cross-level features;

[0017] Based on the feature expression of the area where the building is located and the building-level cross-level features, use SE-Net to perform regional-level and building-level cross-level feature fusion to obtain the functional classification result of the target building.

[0018] Preferably, input the sequence features into a graph convolutional network based on the attention mechanism to obtain the regional POI information features representation, including:

[0019] Input the sequence features into a graph convolutional network based on the attention mechanism and convert them into word vectors through a word embedding layer;

[0020] The word vectors are input into the Transformer model, and the semantic information in the text sequence is captured through the multi-head self-attention mechanism and the feed-forward neural network layer. The feature representation of the text in the highest layer of the Transformer is layer-normalized and linearly projected into the multi-modal embedding space, and the masked self-attention mechanism is used in the text encoder to process and obtain the sequence features;

[0021] Convert the sequence features into a graph structure, and use a graph convolutional neural network to perform feature aggregation based on the graph structure. The graph convolutional neural network updates the feature representation of the nodes by transmitting and aggregating information between the nodes to obtain the regional POI information features representation.

[0022] Preferably, according to the regional POI information feature representation, use the SAM large model to screen the regional street view images and retain the building mask, including:

[0023] According to the regional POI information feature representation, collect multiple original street view images in the region that are closest to the target building to form an overall regional street view image;

[0024] Use the SAM large model to segment the regional street view image to obtain the target mask;

[0025] Multiply the target mask back to the original street view image to obtain the building mask.

[0026] Preferably, according to the building mask, use an image encoder to extract the visual features of the building mask in the regional street view image, including:

[0027] According to the building mask, perform characterization through the Image Encoder to obtain the visual features of the building mask as a one-dimensional vector.

[0028] Preferably, according to the regional POI information feature and the visual features of the building mask, perform cross-modal feature fusion based on the Cross-Attention model fusion to obtain the feature expression of the area where the building is located, including:

[0029] According to the regional POI information feature and the visual features of the building mask, based on Cross-Attention, learn the directional pairwise attention between cross-modal elements to perform cross-modal interaction between elements, and obtain the feature expression of the area where the building is located.

[0030] Preferably, use the BERT model to perform representation learning on the POI text features in the regional POI information feature, including:

[0031] Use the BERT model to capture the relationships and semantic information between the words in the regional POI information feature. The output of the middle layer is used as the feature representation of the input text. The preprocessed text is input into the BERT model, and features are extracted from the last layer to replace the features of the POI name.

[0032] Preferably, use the Informer model to process the check-in data and extract human activity features, including:

[0033] Take the number of people in the buffer area of each building as a time series, divide it into several time periods at a preset time interval, and count the number of platform check-ins in each time period as an indicator of the number of people.

[0034] Input the index of the pedestrian flow into the Informer model, extract features from the last layer of the Informer model, and obtain the expression of the human activity characteristics of the building.

[0035] Preferably, based on the generated POI text features, the building geometric features, and the human activity characteristics, fuse the multi-modal data through a multi-level cross-modal fusion network to obtain building-level cross-level features, including:

[0036] Based on the generated POI text features, the building geometric features, and the human activity characteristics, through a multi-level cross-modal fusion network, process the dependencies between two different sequences based on the cross-attention mechanism, learn the directional pairwise attention between the source modality and the target modality, use the source modality information to strengthen the target modality, aggregate features from single-modal and multi-modal, and obtain building-level cross-level features.

[0037] Preferably, based on the feature expression of the area where the building is located and the building-level cross-level features, use SE-Net to perform regional-level and building-level cross-level feature fusion to obtain the functional classification result of the target building, including:

[0038] Based on the feature expression of the area where the building is located and the building-level cross-level features, use SE-Net to convert the feature map of each channel into a scalar through global average pooling to obtain global description information;

[0039] In the Squeeze stage, introduce the full connection operations of the compression layer and the excitation layer, map the scalar representation of each channel to an intermediate representation, and introduce non-linear transformation through the activation function;

[0040] In the Excitation stage, learn the weight vector of each channel, normalize it through the softmax activation function, and capture key features through feature weighting operations;

[0041] Fuse the building-level multi-modal features and regional-level multi-modal features after feature selection through splicing to obtain the functional classification result of the target building.

[0042] In addition, to achieve the above object, the present invention also proposes a building function classification device, on which a building function classification program is stored, and when the building function classification program is executed by a processor, the steps of the building function classification method described above are implemented.

[0043] Compared with the prior art, the beneficial effects that the present invention can achieve are at least as follows:

[0044] The present invention proposes a deep learning network structure, aiming to achieve high-accuracy classification of urban building functions through efficient feature representation learning and feature fusion. This network structure can fully explore and capture the building's own features and the interaction features with the surrounding regional environment contained in multi-source and multi-modal data. In the present invention, a regional feature representation learning method is proposed, including obtaining OpenStreetMap (OSM) data and preprocessing the vector data of the building's location area using a scale factor, extracting the sequence information of regional points of interest (POIs) through a BERT model, and using a graph convolutional neural network (GCN) to perform efficient representation learning on these sequence information. At the same time, the visual features in the street view map of the building's location area are extracted using the SAM large model, and different modal features are fused and strengthened through Cross-Transformer to obtain a high-level expression of regional features. In the present invention, a multi-modal processing unit (MPU) and a multi-level cross-modal fusion network (MCFC) are designed for the fused and strengthened expression of building-level single-modal features. Aiming at the problems of high time complexity and parameter complexity in the MCFC fusion process, a global-local interaction learning mode is adopted for improvement. The building-level features are derived from Weibo check-in data, building vector contour data, and Baidu POI data. In the present invention, the SE-Net network is used to achieve the fusion of regional-level and building-level cross-level features. The SE-Net network can dynamically adjust the weights of different channels, learn the effective information of different modalities, and finally the fused features pass through a fully connected layer to obtain the final building category representation. In the present invention, the problem of cross-scale intelligent building merging tasks under scale factor constraints is considered, which can more effectively convey geographical information and improve the usability and readability of the map. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a schematic structural diagram of a building function classification method for the hardware operating environment involved in the embodiment of the present invention;

[0046] Figure 2 is a schematic flowchart of the first embodiment of the building function classification method of the present invention;

[0047] Figure 3 are schematic diagrams of five building geometric features in the embodiment of the building function classification method of the present invention;

[0048] Figure 4 is a structural diagram of the MPU in the embodiment of the building function classification method of the present invention.

[0049] The realization, functional features, and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0051] Reference Figure 1 , Figure 1 It is a schematic diagram of the structure of a building function classification device in the hardware operating environment involved in the embodiment of the present invention.

[0052] like Figure 1 As shown, the building function classification device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), and the user interface 1003 may also include a standard wired interface and a wireless interface. The wired interface of the user interface 1003 may be a USB interface in the present invention. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a wireless fidelity (WIreless-FIdelity, WI-FI) interface). The memory 1005 may be a high-speed random access memory (Random Access Memory, RAM) memory, or a stable memory (Non-volatile Memory, NVM), such as a disk memory. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0053] Those skilled in the art will understand that Figure 1 The structure shown in the figure does not constitute a limitation on the functional classification equipment of the building, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.

[0054] like Figure 1 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a building function classification program.

[0055] exist Figure 1 In the building function classification device shown, the network interface 1004 is mainly used to connect to the background server and communicate data with the background server; the user interface 1003 is mainly used to connect to the user device; the building function classification device calls the building function classification program stored in the memory 1005 through the processor 1001, and executes the building function classification method provided by the embodiment of the present invention.

[0056] Based on the above hardware structure, an embodiment of the building function classification method of the present invention is proposed.

[0057] Refer to Figure 2 , Figure 2 which is a schematic flowchart of the first embodiment of the building function classification method of the present invention, and the first embodiment of the building function classification method of the present invention is proposed.

[0058] In the first embodiment, the building function classification method includes the following steps:

[0059] Step S1: Use the RoBERTa model to extract the sequence features of the entire area where the target building is located.

[0060] In a specific implementation, the execution subject of this embodiment is the building function classification device. Among them, the building function classification device can be an electronic device such as a personal computer or a server, and this embodiment does not limit this. The sequence features of the entire area can be extracted using the area POI and street view images. The street view image can be a Baidu street view image, or a street view image in other application software, or a captured street view image, and this is not limited.

[0061] Step S2: Input the sequence features into a graph convolutional network based on the attention mechanism to obtain the feature representation of the area POI information.

[0062] It can be understood that the feature representations corresponding to the texts are obtained for all the POI information in the area through the RoBERTa model. For the text information of the POI, the text input is first converted into word vectors through the word embedding layer and then input into the Transformer model.

[0063] Furthermore, in this embodiment, step S2 includes: inputting the sequence features into a graph convolutional network based on the attention mechanism, and converting them into word vectors through the word embedding layer; the word vectors are input into the Transformer model, and the semantic information in the text sequence is captured through the multi-head self-attention mechanism and the feed-forward neural network layer. The feature representation of the text in the highest layer of the Transformer is layer-normalized and linearly projected into the multi-modal embedding space. The masked self-attention mechanism is used in the text encoder to process and obtain the sequence features; the sequence features are converted into a graph structure, and the graph convolutional neural network is used for feature aggregation based on the graph structure. The graph convolutional neural network updates the feature representation of the nodes by transmitting and aggregating information between the nodes to obtain the feature representation of the area POI information.

[0064] Note that the Transformer captures semantic information in text sequences through the multi-head self-attention mechanism and the feed-forward neural network layer. The text sequence is enclosed by [SOS] and [EOS] tokens, and at the top layer of the Transformer, the activation at the [EOS] token is regarded as the feature representation of the text. This representation is layer-normalized and then linearly projected into the multi-modal embedding space. The masked self-attention mechanism is used in the text encoder to retain the ability to initialize with a pre-trained language model or use language modeling as an auxiliary objective. The text representation learned by the model goes through multiple levels of abstraction and contains rich semantic information. The formula is as follows:

[0065] T i = t1, t2,..., t n

[0066] H t = RoBERTa(T i )

[0067] Where T i represents the original text, t i represents the words or phrases that make up the original text, and n represents the sequence length of the original text. represents the context representation of the output text, represents the context representation of the i-th word t i in the text. Among them, taking represents the feature of the current POI information. After that, all the POI features in the current area are concatenated to obtain the sequence information at the whole area level. We define the sequence feature after the multi-head self-attention module as After that, the sequence feature is converted into a graph structure. The area-level sequence information within a batch size is regarded as the nodes in the graph network structure. After the nodes are constructed, the edges are created through the self-attention mechanism on the nodes:

[0068] Where, represents the adjacency matrix, W q , Let the parameter for learning be \(d\) representing the dimension of the entire sequence, and \(T\) representing the transpose operation. Using ReLU as our activation function can effectively filter out the negative connections between nodes, and a negative connection means that the direct connection between these two nodes is unnecessary. After completing the construction of nodes and edges, a graph containing rich and reliable connections between relevant nodes is formed. Finally, a graph convolutional neural network (GCN) is used for feature aggregation. The graph convolutional neural network updates the feature representation of nodes by transmitting and aggregating information between nodes, thereby obtaining a rich information representation. By introducing GCN, both the graph structure and sequence information can be considered simultaneously, thus better understanding the local and global dependencies in graph-structured data.

[0069] Define the nodes input to the graph convolution as Define its adjacency matrix as Define the identity matrix of the graph structure as For performing the self-loop operation, the entire process is as follows:

[0070]

[0071] f g = ReLU(W f n)

[0072] where \(W\) k , represents the parameter for graph convolution learning, represents the output after learning through the graph convolutional network.

[0073] Step S3: According to the regional POI information feature representation, use the SAM large model to screen the regional street view images and retain the building mask.

[0074] It should be understood that when using the SAM large model to screen the regional street view images, only the building mask is retained.

[0075] Furthermore, in this embodiment, the step S3 includes: according to the regional POI information feature representation, collect multiple original street view images closest to the target building in the region to form an overall regional street view image; use the SAM large model to segment the regional street view image to obtain the target mask; multiply the target mask back to the original street view image to obtain the building mask.

[0076] It should be noted that multiple street view images, such as the four closest images, can also be other numbers of the closest street view images, and this embodiment does not limit this.

[0077] Specifically: First, collect the four street view images closest to the target building within the collection area to form an overall area image. Second, use the SAM image to segment the street view map to obtain the target mask. Finally, multiply the target mask back to the original street view map to obtain an image with only the building mask retained. For topologically adjacent buildings, perform the operation of deleting the shared arc segments of the buildings to achieve the merger of topologically adjacent buildings. Generally speaking, SAM can effectively screen data by emphasizing the parts of interest within the area and suppressing irrelevant or redundant information.

[0078] Step S4: According to the building mask, use an image encoder to extract the visual features of the building mask in the regional street view image.

[0079] It should be noted that for the regional street view image preprocessed in step S3, a building mask is obtained. The image obtained after SAM data screening is characterized by an Image Encoder, and the characterized features are one-dimensional vectors.

[0080] Furthermore, in this embodiment, step S3 includes:

[0081] According to the building mask, perform characterization through an Image Encoder to obtain the visual features of the one-dimensional vector building mask.

[0082] It should be understood that the purpose of this step is to convert the image into a high-dimensional feature vector to capture the key information and abstract features in the image.

[0083] f s = SAM(E1, E2,..., E i ,...E n )

[0084] Where each E i represents the input of the street view map with only the building retained. Represents the visual features of the building extracted by the SAM model. It should be noted that in this model, the SAM model does not update parameters.

[0085] Step S5: According to the regional POI information features and the visual features of the building mask, perform cross-modal feature fusion based on the model fusion of Cross-Attention to obtain the feature expression of the area where the building is located.

[0086] It should be understood that based on the regional POI information features generated in step S2 and the visual features of the building mask obtained in step S4, the above two will use the model fusion method of Cross-Attention for cross-modal feature fusion to obtain the feature expression of the area where the building is located.

[0087] Further, in this embodiment, step S5 includes:

[0088] Based on the regional POI information features and the visual features of the building mask, through Cross-Attention, learn the directional pairwise attention between cross-modal elements to perform cross-modal interaction between elements, and obtain the feature expression of the area where the building is located.

[0089] Specifically: Cross-Attention can achieve cross-modal interaction between elements by learning the directional pairwise attention between cross-modal elements, and strengthen the target modality by introducing a modality reinforcement unit using information from the source modality. The Transformer encoder is connected by applying residuals to multi-head attention (MSA), layer normalization (LN), and multi-layer perceptron (MLP) blocks, and is specifically defined as follows:

[0090]

[0091] h l+1 = MLP(LN(y l )) + y l

[0092] Among them, taking the street view map features as an example, define the input of the sequence encoded as h m ,l m represents the length of the modality input, and d m represents the dimension of the modality input. The above formula represents the encoding result of one layer of the Transformer, and y l is the intermediate result after multi-head attention. The MSA operation calculates dot-product attention, where query, key, and value are all linear projections of the same tensor h m MSA(h m ) = Attention(W Q h m ,W K h m ,W V h m ). This operation enables the high-order features of the POI modality to perform feature selection, making it more focused on the features that have a greater impact on the result.

[0093] The cross-modal attention operation learns the directional pairwise attention between the source modality and the target modality, and strengthens the target modality using the source modality information. For this purpose, define a Cross-Transformer layer:

[0094]

[0095] The Cross-Transformer follows the original Transformer operation, except that the formula becomes:

[0096]

[0097] Based on the above description, we fuse the regional-level SAM feature f s and the regional-level POI feature f g , and use the two modalities to reinforce each other to obtain the enhanced information f s and f g . After that, the regional-level feature

[0098] Step S6: Use the BERT model to perform representation learning on the POI text features in the regional POI information features.

[0099] It is understandable that the RoBERTa model is used to extract features of the POI names. To match the POIs with the buildings, a buffer analysis with a radius of 5m is performed on each building vector contour to generate a one-to-many correspondence between the POIs and the buildings, and one building corresponds to all the POIs within the buffer.

[0100] Furthermore, in this embodiment, the step S6 includes:

[0101] Use the BERT model to capture the relationships and semantic information between the words in the regional POI information features. The output of the middle layer is used as the feature representation of the input text. The preprocessed text is input into the BERT model, and features are extracted from the last layer to replace the features of the POI names.

[0102] It should be noted that the RoBERTa model is used to extract features of all the POI names within the buffer of each building, and the following steps are adopted: First, convert the POI names into the input format acceptable to the model. We divide the POI names into words and add special tokens, such as [CLS] (used to represent the start of the sequence) and [SEP] (used to represent the end of the sequence), then input the preprocessed text into the RoBERTa model, and finally extract features from the last layer of the RoBERTa model to replace the features of the POI names.

[0103] The formal representation is as follows:

[0104] If we have a POI name represented as a sequence of N words: (W1, W2,..., W N ), where each W i is a word vector, then the representation of the entire POI name can be expressed as:

[0105] h t = RoBERTa(w1, w2,..., w N )

[0106] where represents the obtained POI feature, where l t is the sequence length, and d t is the vector dimension of the POI representation. In this way, we obtain a feature vector for representing the POI name, which captures the understanding of the input text by the BERT model. This feature vector can be used for subsequent building classification tasks.

[0107] Step S7: Use Transformer to represent the extracted building geometric features.

[0108] In a specific implementation, in order to fully capture the feature similarity of the building, representative building geometric features are selected, including at least one of the base perimeter, base area, number of corner points, height, volume, minimum bounding rectangle area, and orientation angle. As Figure 3 shown, Figure 3 is a schematic diagram of five building geometric features. Figure 3 Among them, (a) is the initial contour, (b) is the perimeter, and the value of the perimeter can be represented by P. (c) is the area, and the value of the area can be represented by A. (d) is the corner point, and the corner points can be represented by n1, n2, n3......n i represents the corner points. There are a total of 8 corner points n1 to n8 in the figure. (e) is the height, and the value of the height can be represented by H. (f) is the volume, and the value of the volume can be represented by V.

[0109] In this embodiment, seven representative building geometric features are selected as examples for illustration, including the base perimeter, base area, number of corner points, height, volume, minimum bounding rectangle area, and orientation angle. The detailed introduction of the seven building vector contour features is as follows:

[0110] (1) Base perimeter, base area, number of corner points, height, volume

[0111] The base perimeter of the building vector contour refers to the sum of the lengths of the boundary shape of the building. The base area of the building refers to the geometric area of the building vector contour. The number of corner points of the building vector contour refers to the number of vertices of the boundary shape of the building. The height and volume describe the three-dimensional features of the building, and buildings with different functions show different heights and volumes. The volume of the building refers to the space size of the building, which is obtained by multiplying the building area by the height of the building.

[0112] (2) Minimum bounding rectangle area

[0113] The area of the minimum bounding rectangle provides the size information of the building. The minimum bounding rectangle and area of a polygon have been widely used in Geographic Information Science (GIS) and are helpful for identifying the size and proportion of buildings.

[0114] (3) Direction angle

[0115] This feature is used to describe the consistency of the building's direction with other buildings. In this study, the direction angle of a building is defined as the direction of the minimum area bounding rectangle corresponding to the building's footprint. A direction vector is used to describe the building's direction and can be calculated from the vertex coordinates of the minimum area bounding rectangle. The formula is as follows:

[0116]

[0117] where (x1, y1), (x2, y2), (x3, y3) are the coordinates of the vertices, and l1 and l2 are the lengths of the sides of the minimum area bounding rectangle.

[0118] Finally, these numerical features are concatenated into a one-dimensional vector, and the entire sequence is encoded by a Transformer. The resulting structured feature is represented as l t is the sequence length, and d t is the vector dimension of the structured representation.

[0119] Step S8: Use the Informer model to process the platform check-in data and extract human activity features.

[0120] It should be understood that a buffer with a radius of 10m is generated for each building vector contour, the number of people flowing in within a certain time range within the buffer is counted, and a one-to-one correspondence between the number of people flowing in the building within a certain time and the building is generated as an expression of the human activity characteristics of the building. The number of people flowing in different time periods can be used as the human activity characteristics of the building to identify its function. Based on this, we use the Informer model to extract features from the building pedestrian flow data.

[0121] Furthermore, in this embodiment, the step S8 includes:

[0122] Regarding the number of people flowing in each building buffer as a time series, dividing it into several time periods according to a preset time interval, and counting the number of platform check-ins within each time period as an indicator of the number of people flowing in;

[0123] Inputting the indicator of the number of people flowing in into the Informer model, and extracting features from the last layer of the Informer model to obtain an expression of the human activity characteristics of the building.

[0124] It should be noted that when using the Informer model to extract the characteristics of the pedestrian flow in each building buffer area, the following steps are adopted: First, the pedestrian flow in each building buffer area is regarded as a time series, which is divided into several time periods at a certain time interval, and the number of check-ins in each time period is counted as an indicator of the pedestrian flow. In this article, the time interval is set to seven days. Then, the preprocessed data is input into the Informer model. Finally, features are extracted from the last layer of the Informer model to serve as the feature expression of the pedestrian flow. The formal representation is as follows:

[0125] The pedestrian flow series in different time periods within a building buffer area is represented as (X t1 , X t2 ,..., X tn ), where each X ti is the pedestrian flow in a time period. Then, the pedestrian flow in the building can be expressed as:

[0126] h s = Informer(X t1 , X t2 ,..., X tn )

[0127] where represents obtaining the characteristics of human activities in the building, where l t is the sequence length, and d t is the vector dimension of the representation of human activity characteristics. In this way, the expression of the characteristics of human activities in the building by the Informer model is obtained.

[0128] Step S9: Based on the generated POI text features, the building geometric features, and the human activity features, the multi-modal data is fused through a multi-level cross-modal fusion network to obtain building-level cross-level features.

[0129] It should be noted that based on the POI text features generated in step S6, the building geometric features in step S7, and the human activity features extracted in step S8, the building-level multi-modal fusion is realized based on the multi-modal processing unit (MPU), as Figure 4 shown, Figure 4 is the structure diagram of the MPU. The MPU forms target modality reinforcement by learning the directional pairwise attention between the target modality and the source modality. Figure 4 The multi-modal interaction network framework shown mainly includes multi-head cross-modal attention, multi-head self-attention, and a feed-forward network. The following is forFigure 4 Detailed explanations of each part in

[0130] (1) Input part H m and H n represent different modal information, and the source modal information will be used to strengthen the target modal in the subsequent part of this module. H m and H n After the input of H and H, it first undergoes normalization (Norm) processing and then enters the multi - head cross - modal attention module. In this module, the query comes from the target modal H m , and the key and value come from the source modal H n . In this way, information is extracted from the source modal H n to enhance the target modal H m , realizing the directional pairwise attention between different modalities.

[0131] (2) The output after multi - head cross - modal attention and normalization will enter the multi - head self - attention module. In the multi - head self - attention module, the attention weights between the internal elements of the sequence are calculated here to capture the dependencies between the elements within the same modality. Using multiple heads, with the same calculation method but different parameters, feature representations are carried out from multiple sub - spaces, so as to capture richer feature information.

[0132] (3) The output after multi - head self - attention and normalization will enter the feed - forward network. After the data is further processed by the feed - forward network, the final outputs H m→n and H n→m represent the transformation results of different modalities after a series of interactions and processes.

[0133] In the entire multi - modal interaction network architecture, the normalization operation appears multiple times, before and after the multi - head cross - modal attention module and the multi - head self - attention module respectively. Its purpose is to ensure the stability of the data distribution entering these modules, which helps the model to better learn and train.

[0134] Furthermore, in this embodiment, the step S9 includes:

[0135] Based on the generated POI text features, the building geometric features, and the human activity features, through a multi - level cross - modal fusion network, the dependencies between two different sequences are processed based on the cross - attention mechanism. By learning the directional pairwise attention between the source modal and the target modal, the source modal information is used to strengthen the target modal, and the features from single - modal and multi - modal are aggregated to obtain the building - level cross - level features.

[0136] It is understandable that based on this, a Multi-level Cross-Codal Fusion Cetworks (MCFC) is further designed. The architecture of MCFC is based on MPU to achieve the exchange of beneficial information between sequences, enabling different modalities to promote each other. This network has four layers. To allow for further information integration, self-attention is used to simulate the temporal dependencies in each feature sequence.

[0137] The MPU is described in detail as follows. The self-attention mechanism is mainly used to calculate the attention weights between elements in a sequence to capture the dependencies between elements. The multi-head self-attention mechanism uses multiple heads with the same calculation method and different parameters to perform feature representation from multiple subspaces, capable of capturing richer feature information. Taking the POI text modality as an example, the input of the multi-head self-attention mechanism is defined as where d t represents the encoding dimension of the text modality. The whole process can be expressed as:

[0138]

[0139] MSA(H q ) = concat(head1, head2..., head n )W o

[0140] where Q i , K i , V i respectively represent the results of the linear transformation of the input vector by the i-th head, W q , W k , W v are the weight parameters for Query, Key, and Value mappings respectively, mapping the input to the output of d dimensions. concat represents the concatenation operation, and W o is the weight matrix for the final linear transformation. The cross-attention mechanism is used to handle the dependencies between two different sequences. By learning the directional pairwise attention between the source modality and the target modality, it uses the source modality information to strengthen the target modality, that is, the query comes from the target modality H m , while the key and value come from the source modality H n , Q j = H m W q , K j = H i W k , V j = H n W v . In this way, it provides information from modality H mto H n Interaction:

[0141]

[0142] MSA(H m , H n ) = concat(head1, head2..., head n )W o

[0143] Where: H m , H n represent different modalities, and MSA(H m , H n ) is the calculation result of the multi-head cross-attention mechanism.

[0144] Use the efficient multi-level cross-modal fusion network MCFC to achieve better interaction between and within modalities. It mainly uses symmetric cross-modal attention to explore the internal correlation between two input feature sequences, realizes the exchange of beneficial information between the two sequences, so that they can promote each other. To allow further information integration, self-attention is used to simulate the temporal dependencies in each feature sequence. MCFC takes two sequences H t and H i as inputs and outputs the information that promotes each other. Specifically, is calculated as follows:

[0145] H' t→i = MCA(LN(H t ), LN(H i )) + H t

[0146] H'' t→i = MSA(LN(H' t→i )) + H' t→i

[0147] H t→i = FFN(LN(H'' t→i )) + H'' t→i

[0148] Where LN represents layer normalization, and FFN is the feed-forward neural network in Transformer. Similarly, we can obtain Considering that the computational time complexity of MCA(H t , H i ) is O(T t T i ), while the complexity of MSA(H t ) is O(T t 2), the total time complexity of an MCFC is O(T t T i +T t 2 +T t T i +T i 2 ) = O((T t +T i ) 2 ), where T t respectively represents the length of one modality, and T i represents the sequence length of another modality.

[0149] This study notes that based on the Cross-Attention interaction method, one modality needs to be updated twice during the interaction to achieve modality enhancement. For example, the interaction between text and image and the interaction between image and text need to be executed twice. This way of processing is inefficient and will bring redundant features to the sequence. The fact that a certain Token in the large-scale pre-model can represent the historical experience of the entire sequence can further enhance the efficiency of modality interaction. Based on this, this paper designs a global-local learning mode based on linear computational cost to further improve parameter efficiency. We believe that the utterance-level representation of each modality can replace the common information and interact with the local unimodal features as the global multimodal context G. Specifically, set the global multimodal context information where i represents the number of layers of global-local interaction. By interacting the global information and local modality information, modality consistency and modality specificity are learned. The entire interaction process is as follows:

[0150]

[0151] In this way, this strategy can capture one-to-many global-local cross-modal interactions in two CAPs. By stacking multiple layers, the global multimodal context and local unimodal features can promote each other and gradually refine themselves. Since the length of the global multimodal context is very small, the overall time complexity is reduced to (Actually, M << T m ), and it degenerates to O(MT 2 ) in the case of modality alignment. Therefore, the default global-local fusion strategy in MCFC not only has linear space complexity but also enjoys linear computation on the involved modalities.

[0152] At the same time, after each global-local fusion, it will pass through a pooling layer to aggregate the enhanced information of different modalities to promote subsequent fusion. This paper uses a fully connected layer based on the Tanh non-linear activation to implement this operation. Define 3 enhanced global multimodal contexts, namely A new global multi-modal context can be obtained as follows:

[0153]

[0154] The entire learning process is processed in layers, and the features at different main stages of each layer model are obtained. Then, we aggregate the features from unimodal and multi-modal and pass them through a Transformer layer to obtain the corresponding building multi-modal fusion features:

[0155]

[0156] where h q is the feature output after multi-modal fusion, and G [L] represents the aggregated features output after four layers of MCFC. W and b are trainable parameters.

[0157] Step S10: Based on the feature expression of the area where the building is located and the cross-level features at the building level, use SE-Net to perform cross-level feature fusion at the regional and building levels to obtain the functional classification result of the target building.

[0158] It can be understood that based on the feature expression of the area where the building is located obtained in step S5 and the cross-level features at the building level generated in step S9, use SE-Net to perform cross-level feature fusion at the regional and building levels.

[0159] Further, in this embodiment, step S10 includes:

[0160] Based on the feature expression of the area where the building is located and the cross-level features at the building level, use SE-Net to convert the feature map of each channel into a scalar through global average pooling to obtain global description information;

[0161] In the Squeeze stage, introduce the full connection operations of the compression layer and the excitation layer, map the scalar representation of each channel to an intermediate representation, and introduce non-linear transformation through the activation function;

[0162] In the Excitation stage, learn the weight vector of each channel, normalize it through the softmax activation function, and capture key features through feature weighting operations;

[0163] Fuse the multi-modal features at the building level and the multi-modal features at the regional level after feature selection through splicing to obtain the functional classification result of the target building.

[0164] In a specific implementation, SE-Net mainly models the relationship between channels to improve feature expression and selection. Specifically, this module is implemented through two key stages. First, the feature map of each channel is converted into a scalar through global average pooling to obtain global descriptive information. For the input region-level features and building-level features, global average pooling operations are performed on each channel to obtain the global descriptive information of each channel. Subsequently, in the Squeeze stage, fully connected operations of the compression layer and the excitation layer are introduced to map the scalar representation of each channel to an intermediate representation, and a non-linear transformation is introduced through an activation function. Among them, Squeeze reduces the dimension through a fully connected layer and maps the scalar representation of each channel to a smaller intermediate representation. The excitation layer Excitation then uses ReLU to introduce a non-linear transformation. In the Excitation stage, the weight vector of each channel is learned and normalized through the softmax activation function to dynamically adjust the importance of the channels. Finally, through the feature weighting operation, the network realizes the adaptive weighting of the original feature map to capture key features and improve the overall performance. The specific formula is as follows:

[0165]

[0166] Among them, is the enhanced feature expression, W a ,W v are the linear transformation weights, represents the broadcast multiplication of tensors and vectors, and ⊕ represents the addition operation of tensors and vectors. This operation mainly fuses the building-level multi-modal features and region-level multi-modal features after feature selection through concatenation to achieve feature enhancement.

[0167] In this embodiment, a deep learning network structure is proposed, aiming to achieve high-accuracy classification of urban building functions through efficient feature representation learning and feature fusion. This network structure can fully explore and capture the building's own features and the interaction features with the surrounding regional environment contained in multi-source and multi-modal data. In this embodiment, a regional feature representation learning method is proposed, which includes obtaining OpenStreetMap (OSM) data and preprocessing the vector data of the building's location area using a scale factor, extracting the sequence information of regional points of interest (POIs) through a BERT model, and using a graph convolutional neural network (GCN) to perform efficient representation learning on these sequence information. At the same time, the SAM large model is used to extract the visual features in the street view map of the building's location area, and different modal features are fused and enhanced through Cross-Transformer to obtain a high-level expression of regional features. In this embodiment, a multi-modal processing unit (MPU) and a multi-level cross-modal fusion network (MCFC) are designed for the fused and enhanced expression of building-level single-modal features. Aiming at the problems of high time complexity and parameter complexity in the MCFC fusion process, a global-local interaction learning mode is adopted for improvement. The building-level features come from Weibo check-in data, building vector contour data, and Baidu POI data. In this embodiment, the SE-Net network is used to achieve the fusion of cross-level features at the regional and building levels. The SE-Net network can dynamically adjust the weights of different channels, learn the effective information of different modalities, and finally the fused features pass through a fully connected layer to obtain the final building category representation. In this embodiment, the problem of cross-scale intelligent building merging tasks under scale factor constraints is considered, which can more effectively convey geographical information and improve the practicality and readability of the map.

[0168] In addition, an embodiment of the present invention also proposes a storage medium, on which a building function classification program is stored. When the building function classification program is executed by a processor, the steps of the building function classification method described above are implemented.

[0169] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or system including that element.

[0170] The serial numbers of the embodiments of the present invention above are only for description and do not represent the superiority or inferiority of the embodiments. Among the several device unit claims listing a number of devices, several of these devices may be embodied by the same hardware item. The use of the terms first, second, third, etc. does not denote any order and these terms may be construed as identifiers.

[0171] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as a magnetic disk, an optical disc) and includes several instructions for causing a terminal device (which may be a mobile phone, a computer, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0172] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A building function classification method, characterized in that: The building function classification method includes: Use the RoBERTa model to extract sequence features of the entire area where the target building is located; Inputting the sequence features into a graph convolutional network based on an attention mechanism to obtain a feature representation of regional POI information; According to the characteristic representation of the regional POI information, the regional street view image is screened using the SAM large model, and the building mask is retained; Extracting visual features of the building mask in the regional street view image using an image encoder according to the building mask; According to the regional POI information features and the visual features of the building mask, cross-modal feature fusion is performed based on the Cross-Attention model fusion to obtain the feature expression of the area where the building is located; Using the BERT model to learn the representation of POI text features in the regional POI information features; Use Transformer to represent the extracted building geometric features; Use the Informer model to process the platform check-in data and extract human activity features; Based on the generated POI text features, the building geometric features and the human activity features, multimodal data is fused through a multi-level cross-modal fusion network to obtain building-level cross-level features; Based on the feature expression of the area where the building is located and the building-level cross-level features, SE-Net is used to fuse the area-level and building-level cross-level features to obtain the functional classification result of the target building.

2. The building function classification method according to claim 1, characterized in that: The sequence features are input into a graph convolutional network based on an attention mechanism to obtain a feature representation of regional POI information, including: Input the sequence features into a graph convolutional network based on an attention mechanism and convert them into word vectors through a word embedding layer; The word vectors are input into the Transformer model, which captures the semantic information in the text sequence through a multi-head self-attention mechanism and a feed-forward neural network layer. In the highest layer of the Transformer, the feature representation of the text is layer-normalized and linearly projected into a multimodal embedding space. The masked self-attention mechanism is used in the text encoder to obtain sequence features. The sequence features are converted into a graph structure, and a graph convolutional neural network is used to perform feature aggregation based on the graph structure. The graph convolutional neural network updates the feature representation of the nodes by transferring and aggregating information between nodes to obtain the feature representation of the regional POI information.

3. The building function classification method according to claim 1, characterized in that: According to the regional POI information feature representation, the regional street view image is screened using the SAM large model, and the building mask is retained, including: According to the regional POI information feature representation, a plurality of original street view images closest to the target building in the region are collected to form an overall regional street view image; Use the SAM large model to segment the regional street view image and obtain the target mask; The target mask is multiplied back to the original street view image to obtain a building mask.

4. The building function classification method according to claim 1, characterized in that: Extracting visual features of the building mask in the regional street view image using an image encoder according to the building mask, including: According to the building mask, the image encoder is used to characterize the building mask to obtain a visual feature of the building mask in the form of a one-dimensional vector.

5. The building function classification method according to claim 1, characterized in that: According to the regional POI information features and the visual features of the building mask, cross-modal feature fusion is performed based on the Cross-Attention model fusion to obtain the feature expression of the area where the building is located, including: According to the regional POI information features and the visual features of the building mask, based on Cross-Attention, the feature expression of the area where the building is located is obtained by learning the directional pairwise attention between cross-modal elements to perform cross-modal interaction between elements.

6. The building function classification method according to claim 1, characterized in that: The BERT model is used to perform representation learning on the POI text features in the regional POI information features, including: The BERT model is used to capture the relationship and semantic information between words in the POI information features of the area. The output of the middle layer is used as the feature representation of the input text. The preprocessed text is input into the BERT model, and features are extracted from the last layer to replace the features of the POI name.

7. The building function classification method according to claim 1, characterized in that: Use the Informer model to process the check-in data and extract human activity features, including: The flow of people in each building buffer zone is taken as a time series, divided into several time periods according to the preset time intervals, and the number of platform check-ins in each time period is counted as the indicator of the flow of people; The index of the human flow is input into the Informer model, and features are extracted from the last layer of the Informer model to obtain an expression of the characteristics of human activities in the building.

8. The building function classification method according to claim 1, characterized in that: Based on the generated POI text features, the building geometric features and the human activity features, the multimodal data is fused through a multi-level cross-modal fusion network to obtain building-level cross-level features, including: Based on the generated POI text features, the building geometric features and the human activity features, a multi-level cross-modal fusion network is used to process the dependency between two different sequences based on the cross-attention mechanism. By learning the directional pairwise attention between the source modality and the target modality, the source modality information is used to strengthen the target modality, and the features from single modality and multi-modality are aggregated to obtain building-level cross-level features.

9. The building function classification method according to any one of claims 1 to 8, characterized in that: Based on the feature expression of the area where the building is located and the building-level cross-level features, SE-Net is used to fuse the area-level and building-level cross-level features to obtain the functional classification result of the target building, including: Based on the feature expression of the area where the building is located and the building-level cross-level features, the feature map of each channel is converted into a scalar by global average pooling using SE-Net to obtain global description information; In the Squeeze stage, the fully connected operations of the compression layer and the excitation layer are introduced to map the scalar representation of each channel to an intermediate representation, and nonlinear transformation is introduced through the activation function; In the Excitation stage, the weight vector of each channel is learned, normalized by the softmax activation function, and the key features are captured through feature weighting operations; The building-level multimodal features and the region-level multimodal features after feature selection are spliced ​​and fused to obtain the functional classification result of the target building.

10. A building function classification device, characterized in that: The building function classification device stores a building function classification program, and when the building function classification program is executed by the processor, the steps of the building function classification method according to any one of claims 1 to 9 are implemented.