Model framework and creation method for creating building space-time object

Through the address analysis and matching model of the shared coding network, combined with geographic coordinate coding and multi-task learning, the identification and association modeling problems in the integration of multi-dimensional data of urban buildings are solved, high-precision semantic analysis and full-time object construction are realized, and the semantic alignment capabilities of cross-departmental data are improved.

CN120386828AActive Publication Date: 2025-07-29CHINA UNIV OF PETROLEUM (EAST CHINA)
12 Cites 0 Cited by

Patent Information

Application Number
CN202510876109.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

It is difficult for the existing technology to effectively integrate multi-dimensional complex data of urban buildings. Especially when cross-departmental data standards are not unified, the accuracy of building entity object recognition and relationship extraction is low. The traditional model has limited dynamic correlation modeling capabilities for all space-time objects, which affects the construction effect of spatiotemporal knowledge graphs and the semantic alignment ability across data sources.

Method used

The address resolution model and address matching model are used to share the coding network. Through multi-task joint training, the coding network, address resolution network and address matching network are used, and combined with geographic coordinate coding and multi-task learning, to generate token-level representations with both semantic and geographical features, realizing accurate semantic analysis and full-time and spatial correlation modeling of multi-source data of buildings.

Benefits of technology

It significantly improves the accuracy, consistency and efficiency of semantic analysis and spatiotemporal correlation modeling of cross-departmental multi-source heterogeneous urban building data, provides a unified semantic understanding standard, and supports the dynamic representation of all-time and space-time objects of urban buildings and the precise construction of space-time knowledge graphs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386828A_ABST
    Figure CN120386828A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computers, and discloses a model framework for creating a building space-time object and a creating method. The model framework comprises an address resolution model and an address matching model, a coding network is shared between the address resolution model and the address matching model, the input of the coding network is a text address, and the output of the coding network is a text geographic fusion vector; the address resolution model and the address matching model are obtained by training in a multi-task joint training mode; the address analysis model is used for converting an address text in the building multi-source data into an address tag sequence; the address matching model is used for identifying address texts indicating the same building in the building multi-source data; the building space-time object of the target building comprises building information extracted from building multi-source data indicating the target building and an address label sequence. Therefore, accurate semantic analysis and full space-time correlation modeling of cross-department multi-source heterogeneous data are realized through multi-task learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a model framework and a creation method for creating spatio-temporal objects of buildings. Background Art

[0002] Data of buildings, especially urban buildings, often has multi-dimensional complexity, such as including physical attribute information (such as building structure, area, floor height), multi-format address expressions (both standardized and non-standardized coexist), and multi-source business data (such as house maintenance records, elevator installation projects, historical renovation and expansion information).

[0003] When processing multi-source heterogeneous data of buildings, it is difficult to effectively integrate geographical knowledge, resulting in insufficient accuracy and consistency of semantic parsing. Especially in the case where cross-departmental data standards are not unified, the accuracy of building entity object recognition and relationship extraction is relatively low. In addition, in the case where object attribute information has multi-dimensional characteristics, the traditional model has limited ability to dynamically associate and model full spatio-temporal objects, and it is difficult to capture the complex semantic relationships of buildings in the time and space dimensions, which affects the construction effect of spatio-temporal knowledge graphs and the semantic alignment ability across data sources. Summary of the Invention

[0004] This disclosure section is provided to introduce concepts in a brief form, which will be described in detail in the following Detailed Implementation section. This disclosure section is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] In a first aspect, an embodiment of this application provides a model framework for creating spatio-temporal objects of buildings. The model framework includes: an address parsing model and an address matching model. Among them, a coding network is shared between the address parsing model and the address matching model. The input of the coding network is a text address, and the output of the coding network is a text geographical fusion vector. The address parsing model and the address matching model are trained by a multi-task joint training method. The address parsing model includes a coding network and an address parsing network. The input of the address parsing model is an address text, and the output of the address parsing model is an address label sequence. The address parsing model is used to convert the address text in the multi-source data of buildings into an address label sequence. The address matching model includes a coding network and an address matching network. The input of the address matching model is a pair of address texts, and the output of the address matching model is the relationship category between the pair of address texts. The address matching model is used to identify the address texts indicating the same building in the multi-source data of buildings. The spatio-temporal object of the target building includes the building information and the address label sequence extracted from the multi-source data of buildings indicating the target building.

[0006] Optionally, the encoding network includes a tokenizer, an address label embedding module, a geographic coordinate encoding module, and a feature extraction layer; the tokenizer is used to convert the input address text into an address text vector; the address label embedding module is used to convert the address elements in the address text into address label vectors; the geographic coordinate encoding module is used to convert the address elements in the address text into geographic coordinate vectors; the input of the feature extraction layer is a text-geography splicing vector generated based on the address text vector, the address label vector, and the geographic coordinate vector, and the output of the feature extraction layer is a text-geography fusion vector.

[0007] Optionally, the geographic coordinate encoding adopts a spatial continuity preservation technique to convert longitude and latitude coordinates into multi-dimensional continuous vectors; the feature extraction layer adopts a 12-layer Transformer encoder structure, with each layer having a 768-dimensional hidden state and 12 attention heads.

[0008] Optionally, the building spatio-temporal object includes: a building identifier, which is the unique identifiable identifier of the building object; a building spatio-temporal reference, which is the spatio-temporal reference of the building object; a building spatial location; a building spatial form; building attribute features; the behavior ability of the building, the response of the building to changes in the natural environment and to social planning measures; a set of changes of the building.

[0009] Optionally, the address parsing network includes a dilated convolutional neural network, a bidirectional long short-term memory network, and a conditional random field layer; the dilated convolutional neural network is used to extract local multi-scale features from the text-geography fusion vector output by the encoding network; the bidirectional long short-term memory network is used to extract the dependencies in the text-geography fusion vector; the input of the conditional random field layer is a first fusion vector, which is obtained by splicing the outputs of the dilated convolutional neural network and the bidirectional long short-term memory network; the output of the conditional random field layer is an address label sequence.

[0010] Optionally, the input of the address matching network is the paired text-geography fusion vectors output by the encoding network; the address matching network includes a pooling layer, a feature interaction layer, and a fully connected network; the pooling layer is used to pool the input paired text-geography fusion vectors to obtain a first pooled result vector and a second pooled result vector; the feature interaction layer is used to fuse the first pooled result vector and the second pooled result vector to obtain a second fusion vector; the fully connected network is used to map the second fusion vector to the address category matching space.

[0011] Optionally, the pooling layer performs average pooling and max pooling, and fuses the results of average pooling and max pooling; the feature interaction layer performs: calculating the difference vector between the first pooling result vector and the second pooling result vector, and calculating the dot product vector between the first pooling result vector and the second pooling result vector; concatenating the first pooling result vector, the second pooling result vector, the difference vector, and the dot product vector to obtain a second fusion vector.

[0012] Optionally, the loss function of the address parsing model is a conditional random field loss function , and the loss function of the address matching model uses a multi-class cross-entropy loss function ; the joint training process of the address parsing model and the address matching model includes: constructing a joint loss function; where the joint loss function is: ; where, and are task weights, initially set to 0.5, is the regularization coefficient, set to , are model parameters.

[0013] Optionally, the joint training process of the address parsing model and the address matching model includes: determining the performance change rate during the joint training process; adjusting the first task weight and the second task weight based on the performance change rate; where the performance change rate is defined as: ; where, is the performance metric of the th round, where the parsing task uses accuracy, and the matching task uses score; where the weight update formula is: ; ; where, is the adjustment rate, set to 0.1.

[0014] Second aspect, an embodiment of the present application provides a creation method for creating a building spatio-temporal object, including: obtaining heterogeneous data from different departments and text data including addresses; for the text data, extracting structured address information through an address parsing model to generate standard address data; inputting the standard address data and heterogeneous data from different departments into an address matching model together; the address matching model performing matching processing on the received standard address data and heterogeneous data from different departments; if the matching is successful, extracting the matching address pairs, and based on the matching address pairs and the heterogeneous data corresponding to the address pairs, constructing a building spatio-temporal object, further integrating spatio-temporal attributes, and storing the generated spatio-temporal object; if the matching fails, recording the unmatched data and feeding it back for manual verification; wherein, the address matching model and / or the address parsing model are determined based on the model framework in any one of the first aspect.

[0015] The model framework and creation method for creating a building spatio-temporal object provided by the embodiments of the present application effectively enhance the generalization ability and robustness of the model framework through a joint training mechanism, and build a unified semantic understanding standard for cross-department building data. Furthermore, integrating the encoding network of geographical knowledge and multi-task learning significantly improves the accuracy, consistency and efficiency of semantic parsing and spatio-temporal association modeling of cross-department multi-source heterogeneous urban building data. Brief Description of the Drawings

[0016] Combined with the drawings and referring to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present application will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale.

[0017] Figure 1 is a flowchart of an embodiment of a model framework for creating a building spatio-temporal object according to the present application; Figure 2 is a schematic structural diagram of an encoding network integrating geographical knowledge; Figure 3 is a schematic structural diagram of an address parsing model for multi-scale feature fusion and sequence annotation; Figure 4 is a schematic structural diagram of a building address matching model for deep feature interaction; Figure 5 is a flowchart of an embodiment of a creation method for a building spatio-temporal object according to the present application; Figure 6 is a schematic diagram of the basic structure of an electronic device provided by an embodiment of the present application. Detailed Description of the Embodiments

[0018] Embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are only for illustrative purposes and are not used to limit the protection scope of the present application.

[0019] It should be understood that the various steps recited in the method embodiments of the present application can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this regard.

[0020] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0021] It should be noted that the concepts such as "first", "second", etc. mentioned in the present application are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions executed by these devices, modules or units or their interdependent relationships.

[0022] It should be noted that the modifications of "one" and "multiple" mentioned in the present application are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0023] The names of the messages or information exchanged between multiple devices in the embodiments of the present application are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0024] In one or more embodiments of the present application, a semantic association method for urban buildings based on a large language model is proposed, aiming to achieve precise semantic parsing and full-time and space association modeling of cross-departmental multi-source heterogeneous data by integrating geographical knowledge and multi-task learning. The core technical solution uses a shared encoder (such as GeoRoBERTa), and by injecting address label embeddings and geographical coordinate encodings, generates token-level representations with both semantic and geographical features. In the address parsing branch, IDCNN and BiLSTM are used to extract multi-scale features, and combined with CRF to predict the geographical name label sequence, realizing high-precision building entity recognition and the construction of full-time and space building objects. In the address matching branch, pooling processing is performed on the address pair features output by GeoRoBERTa to predict the matching category, ensuring semantic alignment across data sources. Through multi-task joint training, the address parsing and matching tasks are optimized synchronously to improve the model efficiency and generalization ability. This solution constructs a unified semantic standard by deeply integrating geographical knowledge and the semantic analysis ability of the large language model, supports the dynamic representation of full-time and space building objects in the city and the precise construction of a spatio-temporal knowledge graph, and solves the consistency problem in cross-departmental data integration.

[0025] In one or more embodiments of the present application, an implementation manner of a semantic association method for urban buildings based on a large language model is proposed. Through a shared encoder and a multi-task learning framework, semantic parsing and full-time and space object association modeling of cross-departmental multi-source heterogeneous data are realized. The overall implementation solution can use GeoRoBERTa as the basic encoder, inject geographical knowledge (address label embeddings and geographical coordinate encodings), and generate token-level representations with both semantic and geographical features. The system includes two branches: the address parsing branch uses IDCNN and BiLSTM to extract multi-scale features, and combined with CRF to predict the geographical name label sequence; the address matching branch performs pooling processing on the address pair features to predict the matching category. The two branches are optimized through multi-task joint training to ensure the high efficiency and generalization ability of the model.

[0026] Please refer to Figure 1 , which shows the process of an embodiment of the model framework for creating building spatio-temporal objects according to the present application. As Figure 1 shown, the model framework for creating building spatio-temporal objects.

[0027] The model framework includes an address parsing model and an address matching model.

[0028] In this embodiment, an encoding network is shared between the address parsing model and the address matching model. The input of the encoding network is a text address, and the output of the encoding network is a text geographical fusion vector; the address parsing model and the address matching model are trained by means of multi-task joint training.

[0029] In this embodiment, the address parsing model includes an encoding network and an address parsing network; the input of the address parsing model is address text, and the output of the address parsing model is an address tag sequence; the address parsing model is used to convert the address text in the multi-source data of buildings into an address tag sequence.

[0030] In this embodiment, the address matching model includes an encoding network and an address matching network; the input of the address matching model is paired address text, and the output of the address matching model is the relationship category between the paired address texts; the address matching model is used to identify the address texts indicating the same building in the multi-source data of buildings.

[0031] In this embodiment, the building spatio-temporal object can integrate the multi-dimensional information of the building physical store, and support the expression, association and dynamic change analysis of the building.

[0032] For example, represents a building object.

[0033] ; wherein, Id refers to the Id value of the building object, which is the unique identifiable identifier of the building object, Refer is the spatio-temporal reference of the building object, Locations represents the spatial position of the building object, Shapes represents the spatial form of the building object, Features represents the attribute features of the building object, Actions represents the behavioral ability of the building object, and Variations represents the set of changes of the building object.

[0034] It should be noted that for the multi-dimensional complexity of building data, including physical attribute information (such as building structure, area, floor height), multi-format address expressions (both standardized and non-standardized), and multi-source business data (such as house maintenance records, elevator installation projects, historical renovation and expansion information), the multi-task learning framework effectively enhances the generalization ability and robustness of the model through a joint training mechanism, and constructs a unified semantic understanding standard for cross-department building data. Furthermore, the encoding network integrating geographical knowledge and multi-task learning significantly improves the accuracy, consistency and efficiency of semantic parsing and spatio-temporal association modeling of cross-department multi-source heterogeneous urban building data.

[0035] In some embodiments, the building spatio-temporal object, which can also be referred to as the full spatio-temporal object model, has a higher level of logical and cognitive abstraction and can build a unified data framework for spatio-temporal entities. For multi-dimensional building entities, this model can integrate their attributes to support the expression, association, and dynamic change analysis of buildings. The core of the objectified expression of buildings lies in treating them as independent entities and describing and managing information in the form of spatio-temporal objects. This method transforms the attributes such as the history, current situation, and development potential of buildings into computable data streams, not only reconstructing the information system but also providing a theoretical basis for full spatio-temporal data modeling.

[0036] In some embodiments, the building spatio-temporal object may include: a building identifier, which is the unique identifiable identifier of the building object; a building spatio-temporal reference, which is the spatio-temporal reference of the building object; the building's spatial location; the building's spatial form; the building's attribute characteristics; the building's behavioral ability, which is the building's response to changes in the natural environment and social planning measures; and the set of changes of the building.

[0037] When establishing the data model of the building object, it is necessary to take it as the core entity category and build a systematic framework based on the key elements of objectified description. An example of the specific modeling method is as follows: ; Among them, Id refers to the Id value of the building object, which is the unique identifiable identifier of the building object, Refer is the spatio-temporal reference of the building object, Locations represents the spatial location of the building object, Shapes represents the spatial form of the building object, Features represents the attribute characteristics of the building object, Actions represents the behavioral ability of the building object, and Variations represents the set of changes of the building object.

[0038] Optionally, the spatio-temporal reference can determine the time and space position of the full spatio-temporal object in the real world. Based on the spatio-temporal reference as the benchmark, this reference can be divided into time and space and can be expressed as a formula: ; Among them, TemporalRefer represents the time reference of the building object, and SpatialRefer represents the spatial reference of the building object. Currently, commonly used time reference systems include International Atomic Time (TAI), Coordinated Universal Time (UTC), Greenwich Mean Solar Time (GMST), BeiDou Time (BDT), etc. The spatial reference of buildings can adopt commonly used domestic spatial reference systems: China Geodetic Coordinate System 2000 (CGCS2000), WGS84, etc.

[0039] Optionally, the attribute of the spatial location determines the spatial range and centroid coordinates of the building object, and the representation method of the spatial location is as follows: ; Among them, SpatialExtent represents the two-dimensional spatial range of the building object, and Centroid represents the two-dimensional centroid coordinates of the building. The SpatialExtent in the formula is the framing of the spatial range, and its formal expression is the formula: ; In the formula, are the upper left and lower right coordinates of the two-dimensional square spatial range of the building object.

[0040] Optionally, the spatial form details the two-dimensional spatial boundary range, form, and expression form of the building object. At different scales, angles, and states, the expressions of the full-time and full-space building objects are not the same. The representation method of the spatial form is as follows: ; Among them, Form represents the different forms of the building object. For the display of different forms, it is necessary to explain the Type type, MinScale represents the minimum scale for displaying this form, MaxScale represents the maximum scale for displaying this form, and Geometry represents the display method of this form.

[0041] Optionally, the attribute features are used to describe the specific attributes of the object, such as the building area, the actual use of the building, the construction year of the building, the building structure, the building height, the number of above-ground and underground floors of the building, and other specific attributes. The attribute features can be expressed as a formula: ; Among them, Features represents the set of attribute features of the building object, Feature1, Featurer2, Feature3 represent an attribute feature in the set of attribute features, Name represents the name of the specific attribute feature, and Value represents the attribute value corresponding to the specific attribute feature.

[0042] Optionally, the behavioral ability of the building describes how the building responds to time changes and social planning and development, including the building's response to changes in the natural environment and social planning measures. For example, the building structure will change due to artificial economic or natural factors such as renovation, expansion, decoration, and natural disasters. The behavioral ability can be expressed as a formula: ; Among them, Type and Name represent the type and name of the behavior, and Parameters, Conditions, Triggers, Receptors, Models represent the behavioral ability parameters, environmental impact factors, behavior trigger conditions, behavior action objects, and behavior calculation models respectively.

[0043] Optionally, the variation set describes the dynamic changes that occur to a building object over time. For example, urban planning can change the building area and form of a building structure. The variation set can be expressed as a formula: ; where Variations represents all variation sets of the building object, Variation represents each specific variation, Time is the time when a certain variation occurs, and Refer, Locations, Shapes, Features, and Action respectively represent the characteristic values corresponding to the spatio-temporal reference, spatial location, spatial form, attribute features, and behavioral capabilities after the variation.

[0044] Thus, the full spatio-temporal object model provides a unified framework for the accurate and multi-dimensional description and management of building entities in a computer. It can express the spatial form of a building object, describe the attribute features of a building object, and record the dynamic changes of a building object, which provides a basic framework and subsequent data management guidance for cross-departmental multi-source heterogeneous data processing, address matching, and semantic association of building entity objects.

[0045] In some embodiments, the encoding network includes a tokenizer, an address tag embedding module, a geographic coordinate encoding module, and a feature extraction layer. The tokenizer is used to convert the input address text into an address text vector. The address tag embedding module is used to convert the address elements in the address text into address tag vectors. The geographic coordinate encoding module is used to convert the address elements in the address text into geographic coordinate vectors.

[0046] The input of the feature extraction layer is a text-geographic splicing vector generated based on the address text vector, address tag vector, and geographic coordinate vector, and the output of the feature extraction layer is a text-geographic fusion vector.

[0047] It should be noted that the encoding network realizes the accurate semantic modeling of multi-source heterogeneous data by injecting geographic knowledge and a multi-layer Transformer structure into the encoder. This module establishes a deep connection between language and space by fusing semantic and geographic features, provides a high-quality contextualized feature representation for subsequent address parsing and matching tasks, and thus significantly improves the accuracy and generalization ability of building semantic association.

[0048] As an example, the encoding network can be based on a large language model as a pre-training basis.

[0049] As an example, the encoding network can include a 12-layer Transformer encoder structure, with each layer having a 768-dimensional hidden state and 12 attention heads, and is optimized through the Dynamic Masked Language Modeling (MLM) technique.

[0050] Thus, the dynamic mask applies different mask patterns in each training batch, significantly enhancing the model's semantic capture ability and generalization for complex texts. The core of the Transformer architecture lies in its self-attention mechanism, which allows the model to consider information at all positions when processing a sequence, thereby capturing long-range semantic dependencies in the building description text.

[0051] Optionally, a deep transformation of geographical knowledge injection can be performed based on the Transformer architecture to enhance the model's ability to model address and spatial information to meet the specific task requirements of building semantic associations. Geographical knowledge injection is a multi-modal information fusion technology that significantly improves the model's performance in geographical-related tasks by combining semantic information with spatial information.

[0052] Geographical knowledge injection mainly includes two key parts: address label embedding and geographical coordinate encoding that adopts a spatial continuity preservation technique.

[0053] In one embodiment, the address embedding can be achieved in the following way. Design 64-dimensional label embedding vectors for address components (such as province, city, district, street, house number). Let the set of address components be C = {c1, c2, ..., c n}, and each component c i is mapped to an embedding vector , which is generated through a learnable embedding matrix . The introduction of address label embedding not only enhances the semantic representation of tokens but also injects prior knowledge of the address structure into the model, improving the sensitivity and parsing accuracy for address information.

[0054] In one embodiment, the geographical coordinate encoding adopts a spatial continuity preservation technique to convert longitude and latitude coordinates into multi-dimensional continuous vectors.

[0055] The geographical coordinate encoding adopts a spatial continuity preservation technique to convert longitude and latitude coordinates into 128-dimensional continuous vectors, using sine-cosine positional encoding: ; In this expression, is the position after normalizing the coordinate values, = 128 is the encoding latitude, is the dimension index. The core advantage of sine-cosine encoding is that it can preserve the continuity characteristics of space, making geographically close locations have similar representations in the embedding space, which is crucial for the geographical clustering and spatial relationship reasoning of buildings. In addition, this encoding method also has scale invariance and can adapt to the location representation requirements of different geographical granularities.

[0056] In one embodiment, the feature extraction layer adopts a 12-layer Transformer encoder structure, with each layer having a 768-dimensional hidden state and 12 attention heads.

[0057] As Figure 2 shown, the encoding network adopts the GeoRoBERTa model, which provides high-quality token-level representations for subsequent address parsing and matching tasks by deeply integrating semantic and geographical features.

[0058] In terms of building the basic model, this solution can select RoBERTa-base as the pre-training basis. This model contains a 12-layer Transformer encoder structure, with each layer having a 768-dimensional hidden state and 12 attention heads. RoBERTa is optimized through the Dynamic Masked Language Modeling (MLM) technique. Compared with the static masking strategy, dynamic masking applies different masking patterns in each training batch, significantly enhancing the model's semantic capture ability and generalization for complex texts. The core of the Transformer architecture lies in its self-attention mechanism, which allows the model to consider information at all positions when processing sequences, thereby capturing long-range semantic dependencies in the building description text. The mathematical basis expression of self-attention calculation is: ; Among them, represent the Query, Key, and Value matrices respectively, which are generated from the same input data through linear transformation. is the dimension of the key vector (64), which is used to scale the dot product operation to prevent the gradient vanishing problem. The uniqueness of this attention mechanism is that it can adaptively assign weights to each token, enabling the model to dynamically adjust the focus according to the context, which is crucial for complex address parsing and matching tasks.

[0059] To meet the specific task requirements of semantic association of urban buildings, this solution has deeply transformed RoBERTa by injecting geographical knowledge, thereby enhancing the model's ability to model address and spatial information. Geographical knowledge injection is a multi-modal information fusion technology that significantly improves the model's performance in geographical-related tasks by combining semantic information with spatial information. Specifically, geographical knowledge injection mainly consists of two key parts: Address tag embedding: Design 64-dimensional tag embedding vectors for address components (such as province, city, district, street, house number). Let the set of address components be C = {c1, c2, ..., c n}, and each component c i is mapped to the embedding vector , which is generated through a learnable embedding matrix . The introduction of address tag embedding not only enhances the semantic representation of tokens but also injects prior knowledge of address structure into the model, improving the sensitivity and parsing accuracy of address information.

[0060] Geographical coordinate encoding adopts a technology to maintain spatial continuity, converting longitude and latitude coordinates into 128-dimensional continuous vectors, using sine-cosine positional encoding: ; In this expression, is the position after normalizing the coordinate value, is the encoded latitude, and is the dimension index. The core advantage of sine-cosine encoding is that it can retain the continuous characteristics of space, making geographically close locations have similar representations in the embedding space, which is crucial for geographical clustering and spatial relationship reasoning of buildings. In addition, this encoding method also has scale invariance and can adapt to the position representation requirements of different geographical granularities.

[0061] In the model pre-training stage, a large-scale cross-department building-related text corpus is used, including building attribute information (such as structure type, function description), address information (such as province, city, district, street, house number), house maintenance records, elevator installation, and renovation and expansion data, etc. The pre-training tasks include: Masked Language Model (MLM): Randomly mask 15% of the tokens and predict the original vocabulary. The loss function is cross-entropy: ; Address component prediction: Predict the address component label corresponding to the token to enhance geographical knowledge modeling. The loss function is the same as above. Use the Adam optimizer (initial learning rate , with a batch size of 32), and the training cycle is 10 epochs. The model parameters are optimized to adapt to the semantic scenario of urban buildings.

[0062] The input text is converted into a token sequence by a tokenizer (based on Byte-Pair Encoding). Each token is mapped to a 768-dimensional word embedding vector. Combining the address label embedding and geographical coordinate encoding forms a comprehensive input representation. The formula for the embedding vector is: ; where are the word embedding, label embedding, and coordinate encoding respectively. This multi-modal fusion method integrates semantic and geographical information in the vector space, providing a rich feature basis for subsequent attention calculation.

[0063] The attention calculation process is implemented by cascading 12 layers of Transformer encoders. Each layer of the encoder calculates the context-related token representation based on the self-attention mechanism. This multi-level encoding architecture can gradually extract and integrate the semantic and geographical dependencies in the building description, from surface lexical features to deep semantic associations, constructing a hierarchical representation learning system. In each layer of the Transformer, the multi-head attention mechanism decomposes the attention operation into multiple parallel attention heads, and each head focuses on different feature subspaces, further enhancing the model's expressive ability. The output of each layer is a 768-dimensional vector, retaining rich context information and serving as the input for the next layer, forming a deep cascaded feature extraction network.

[0064] GeoRoBERTa finally outputs the contextualized representation (dimension 768) of each token as the input feature for the address parsing and matching tasks. The output vector sequence is: ; In summary, through the geographical knowledge injection of GeoRoBERTa and the multi-layer Transformer structure, the encoding network realizes the accurate semantic modeling of multi-source heterogeneous data. By fusing semantic and geographical features, this module establishes a deep connection between language and space, provides high-quality contextualized feature representations for subsequent address parsing and matching tasks, and thus significantly improves the accuracy and generalization ability of urban building semantic associations.

[0065] In some embodiments, the address parsing network includes a dilated convolutional neural network, a long short-term memory network, and a conditional random field layer; the dilated convolutional neural network is used to extract local multi-scale features from the text-geographical fusion vector output by the encoding network; the long short-term memory network is used to extract dependencies in the text-geographical fusion vector; the input of the conditional random field layer is a first fusion vector, which is obtained by concatenating the outputs of the dilated convolutional neural network and the long short-term memory network; the output of the conditional random field layer is an address label sequence.

[0066] As Figure 3 shown, the address parsing branch is built on the semantic representation generated by an encoding network (such as GeoRoBERTa). Through a multi-level feature extraction and sequence annotation architecture, namely, integrating an Iterated Dilated Convolutional Neural Network (IDCNN), a Bidirectional Long Short-Term Memory (BiLSTM), and a Conditional Random Field (CRF), it realizes high-precision semantic deconstruction and entity recognition of building address texts. As Figure 3 shown, through a hierarchical feature extraction and inference process, this branch converts unstructured address texts into structured semantic representations, providing basic support for cross-department building data integration.

[0067] As an example, the Iterated Dilated Convolutional Neural Network (IDCNN) is a neural network designed specifically for sequence modeling. By dilating the convolutional kernel, it expands the receptive field and efficiently captures local context information without stacking too many layers. Therefore, by introducing gaps (dilation rates) in the convolutional kernel, the model can cover longer sequence ranges, which is particularly suitable for parsing the hierarchical structure of address texts (such as streets and house numbers).

[0068] As an example, the IDCNN is configured with 3 layers of dilated convolution, with dilation rates of 1, 2, and 4 respectively, a convolutional kernel size of 3×3, and an output channel number of 256. The convolution operation is defined as: ; where is the input feature window, is the convolutional kernel weight, is the bias, which is determined by the dilation rate. The increasing dilation rate makes the receptive field of each layer grow exponentially, covering short, medium, and long-range context information respectively, and generating a 256-dimensional local feature vector.

[0069] As an example, the core advantages of the Bidirectional Long Short-Term Memory Network (BiLSTM) include its ability to effectively model long-range dependencies in sequential data. BiLSTM processes sequential information bidirectionally, from front to back and from back to front, by combining a forward and a backward LSTM unit. It is particularly suitable for capturing long-distance semantic associations that may exist in address text (such as the subordinate relationship between building names and administrative divisions, and the distinction between main address components and auxiliary information). The uniqueness of the LSTM unit lies in its carefully designed gating mechanism - including the forget gate, input gate, and output gate. These gating structures can intelligently control the flow of information, determining which historical information needs to be retained, which new information needs to be integrated, and which internal states need to be output, effectively alleviating the vanishing gradient problem in traditional recurrent neural networks.

[0070] As an example, BiLSTM can be configured with 2 layers and a hidden layer dimension of 256. Its output is calculated as follows: ; ; where, represents the hidden state of the forward LSTM at position , which depends on the current input and the hidden state at the previous position, ; represents the hidden state of the backward LSTM at position , which depends on the current input and the hidden state at the next position ; represents vector concatenation. By concatenating the hidden states in both forward and backward directions, BiLSTM generates a 512-dimensional feature representation for each token (a combination of 256-dimensional forward LSTM features and 256-dimensional backward LSTM features). This bidirectional processing mechanism of BiLSTM significantly enhances the model's ability to model complex context dependencies in address sequences, enabling the system to understand long-distance association patterns and semantic integrity in address text.

[0071] The two network structures, IDCNN and BiLSTM, are respectively good at capturing features of different scales and natures: IDCNN efficiently extracts local multi-scale features, while BiLSTM is good at modeling long-term dependencies in sequences. To fully utilize the complementary advantages of the two structures, the system integrates their outputs through a feature fusion mechanism. Specifically, for each position in the sequence, the output features of IDCNN and BiLSTM are fused into a comprehensive feature vector through a concatenation operation: ; The feature dimension after splicing is 768 (256 - dimensional IDCNN + 512 - dimensional BiLSTM). To reduce the computational complexity, it is reduced to 512 dimensions through a linear layer: ; Among them, 、 are learnable parameters. The features after dimensionality reduction retain the key semantic information and are suitable for subsequent sequence labeling.

[0072] Here, the Conditional Random Field (CRF) is a probabilistic graphical model for sequence labeling tasks. By modeling the transition probabilities between labels, it ensures the global consistency of sequence prediction. CRF optimizes the label sequence by maximizing the log - likelihood probability of the sequence, which is particularly suitable for tasks that require considering label dependencies in address parsing (e.g., in the Chinese address format, after "Province", "City" should usually follow rather than "District", and "Street" usually follows "District", etc.). The probability distribution of CRF is defined as: ; Among them, is the label score of the th token, is the label transition score. The BIOES (Begin - Inside - Outside - End - Single) annotation scheme is adopted, and the label set includes Province (provincial administrative region), City (municipal administrative region), District (district - level administrative region), Street (street / township), Community (community / residential area), Building (building name), Number (house number), POI (point of interest), Other (other information). CRF decodes through a dynamic programming algorithm (such as Viterbi) to output the optimal label sequence.

[0073] The negative log - likelihood loss function is used in the training process: ; The Adam optimizer is used. Among them, is the true label sequence, is the input feature sequence, is the conditional probability calculated by CRF. The initial learning rate is , the weight decay is , and the batch size is 16. The accuracy on the validation set is monitored during training. When the accuracy no longer improves, the model convergence is optimized through a learning rate decay strategy (reduced to 1 / 10 of the original learning rate). Joint training ensures the collaborative optimization of the IDCNN, BiLSTM, and CRF modules, improving the recognition accuracy of address entities (such as building names and house numbers) in cross - departmental data.

[0074] During the entire training process, the three core modules of IDCNN, BiLSTM, and CRF are jointly optimized in an end-to-end manner, ensuring the coordinated development of feature extraction and sequence annotation. This overall optimization strategy enables the system to automatically adjust the parameters of each module, improving the recognition accuracy of address entities in cross-departmental heterogeneous data, especially the ability to accurately extract key entities such as building names and house numbers.

[0075] Through this multi-level architecture, the address parsing network integrates the multi-scale local feature extraction ability of IDCNN, the long-distance dependence modeling ability of BiLSTM, and the sequence consistency guarantee mechanism of CRF to achieve precise semantic parsing of building address texts. The system can not only handle standardized standard address formats but also demonstrates good robustness to unstructured, partially missing, or diversely expressed address texts, providing key technical support for semantic alignment and integration of cross-departmental multi-source heterogeneous data.

[0076] In some embodiments, the input of the address matching network is the paired text-geographical fusion vector output by the encoding network; the address matching network includes a pooling layer, a feature interaction layer, and a fully connected network; the pooling layer is used to pool the input paired text-geographical fusion vector to obtain a first pooled result vector and a second pooled result vector; the feature interaction layer is used to fuse the first pooled result vector and the second pooled result vector to obtain a second fusion vector; the fully connected network is used to map the second fusion vector to the address category matching space.

[0077] In some embodiments, the pooling layer performs average pooling and max pooling and fuses the results of average pooling and max pooling; the feature interaction layer performs: calculating the difference vector between the first pooled result vector and the second pooled result vector, calculating the dot product vector between the first pooled result vector and the second pooled result vector; concatenating the first pooled result vector, the second pooled result vector, the difference vector, and the dot product vector to obtain the second fusion vector.

[0078] As Figure 4 shown, the address matching network can be constructed on the semantic representation generated by an encoding network (such as GeoRoBERTa) to process the similarity assessment of address pairs and determine whether two address records point to the same building entity in the real world. This function is crucial for building identity authentication and semantic alignment in cross-departmental data integration. As Figure 4 shown, this matching mechanism realizes high-precision address entity matching through a series of steps such as deep feature extraction, multi-strategy feature pooling, multi-dimensional feature interaction, and hierarchical classification decision-making.

[0079] The address matching network is based on the output features of the encoding network (i.e., the text-geographical fusion vector). As an example, for the two input address texts and , context representations at the token level are generated through the encoding network respectively.

[0080] Optionally, the encoding network adopts an improved Transformer architecture, which can not only capture the linguistic semantic information in the address text, but also extract the geographical spatial correlation through a special geographical knowledge enhancement mechanism, so as to generate a sequence of rich 768-dimensional feature vectors for each token: ; wherein, and are the number of tokens of addresses and respectively. This step utilizes the geographical knowledge injection of the encoding network (such as address label embedding and coordinate encoding) to ensure that the features contain the semantic information of building attributes, addresses and business information.

[0081] Optionally, a complementary feature pooling strategy can be adopted, combining two operations of average pooling and max pooling, to effectively aggregate the local features at the token level into a global representation that can represent the entire address entity.

[0082] Here, average pooling effectively captures the overall semantic distribution and general characteristics of the address by calculating the mean of all token features.

[0083] Here, max pooling highlights the key identification elements in the address by extracting the most significant features in each dimension. The mathematical definitions of these two pooling operations are: ; wherein, is the number of tokens, is the token feature. Average pooling and max pooling are calculated for addresses and respectively to obtain two 768-dimensional vectors. To retain the complementary information captured by different pooling strategies, the results of average pooling and max pooling are further concatenated to generate a richer 1536-dimensional address representation vector: ; This concatenation operation not only retains the global semantics and local significant features of the address, but also enhances the robustness of the representation, can adapt to address descriptions in different formats and expressions, and has strong adaptability to non-standard address inputs.

[0084] To capture the semantic relationship between two addresses, calculate the address representation vectors and The interaction features. These interaction features mainly include the difference vector and the dot product vector.

[0085] The difference vector reflects the degree of difference between two addresses in each dimension through element-wise subtraction operation.

[0086] The dot product vector quantifies the similarity strength between two addresses in each dimension through element-wise multiplication operation: ; where represents element-wise multiplication. These two interaction features respectively characterize the semantic relationship between address pairs from different perspectives.

[0087] Furthermore, the original address representation vector (i.e., the pooling result vector processed by the pooling layer), the difference vector, and the dot product vector can be comprehensively concatenated to form the final high-dimensional feature representation: ; This 6144-dimensional fusion feature vector (formed by concatenating 4 1536-dimensional vectors) comprehensively integrates the original semantics, relative differences, and interaction similarities of the address pair, providing a rich and multi-dimensional input representation for subsequent refined classification.

[0088] Optionally, in the matching classification stage, through a carefully designed fully connected network, the high-dimensional feature space is mapped to the address matching category space. The fully connected network gradually extracts the complex relationships and abstract patterns between features through multiple non-linear transformations.

[0089] Optionally, the fully connected network can adopt a network architecture with three hidden layers, and the dimensions of each layer are set to 1024, 512, and 256 respectively to achieve a progressive mapping from the high-dimensional feature space to the low-dimensional decision space. The ReLU activation function is used between each layer to introduce non-linear transformation and enhance the expression ability of the network: ; where , are the weights and biases of the th layer. The last layer of the network outputs the matching probability through the Softmax function: ; The matching categories include full match (the addresses point to the same building), partial match (the addresses partially overlap, such as the same street but different house numbers), and non-match (the addresses are irrelevant). Softmax ensures that the sum of probabilities is 1, enhancing the reliability of classification.

[0090] The training uses the cross-entropy loss function to measure the difference between the predicted probability and the true label: ; Among them, is the true label, is the predicted probability. Using the gradient descent algorithm with an adaptive learning rate (such as the Adam optimizer) can automatically adjust the parameter update step size. The initial learning rate is , and the batch size is 24. To promote the model to converge more stably to the optimal solution, after every 5 complete epochs of training, the learning rate will decay to 0.9 times the original value to promote convergence: ; This learning rate decay strategy maintains a larger parameter update step size in the initial stage of training to quickly approach the optimal solution region, while in the later stage of training, a smaller step size is used for fine-tuning, effectively avoiding the overfitting problem and improving the generalization ability of the model. The training process optimizes the parameters of the fully connected network to ensure the matching accuracy of building addresses, attributes, and business information in cross-departmental data.

[0091] It should be noted that through this hierarchical feature extraction, multi-strategy pooling, multi-dimensional interaction, and deep classification architecture design, the address matching network realizes the accurate semantic matching of address pairs, providing reliable technical support for the semantic alignment of cross-departmental data and the full-time and space association modeling of buildings. The address matching network can not only process standardized address formats, but also has strong adaptability to unstructured, partially missing, or diverse address descriptions, effectively solving the challenges of address data heterogeneity in practical applications.

[0092] In some embodiments, the loss function of the address parsing model is the conditional random field loss function . The loss function of the address matching model adopts the multi-class cross-entropy loss function . The joint training process of the address parsing model and the address matching model includes: constructing a joint loss function; among them, the joint loss function is: ; Among them, and are task weights, initially set to 0.5, is the regularization coefficient, set to , are model parameters.

[0093] By jointly optimizing the two closely related but distinct tasks of address parsing and address matching through joint training, the semantic parsing and association capabilities of the model for cross-departmental building data can be significantly enhanced. The feature representation sharing mechanism and task-specific scaffolding architecture, combined with the adaptive dynamic weight adjustment technique, enable efficient knowledge transfer and fusion learning in heterogeneous data environments. Aiming at the multi-dimensional complexity of building data - including physical attribute information (such as building structure, area, floor height), multi-format address expressions (both standardized and non-standardized coexist), and multi-source business data (such as house maintenance records, elevator installation projects, historical renovation and expansion information) - the multi-task learning framework effectively enhances the generalization ability and robustness of the model through the joint training mechanism, and constructs a unified semantic understanding standard for cross-departmental building data.

[0094] Defining specific losses for each task and combining them with weights can coordinate the optimization processes of the address parsing and matching tasks. The address parsing task adopts the conditional random field (CRF) loss, which measures the log-likelihood of the prediction probability of sequence labeling and the true label: ; where, is the true label sequence, is the input sequence, is calculated from the CRF probability distribution, representing the probability of the label sequence appearing given the input sequence . The adoption of the CRF loss function fully considers the context dependencies and transition constraints between address components, can capture the implicit grammar rules in the urban hierarchical structure and address format specifications, and improve the global consistency of sequence labeling.

[0095] The address matching task is constructed as a three-class classification problem, and the multi-class cross-entropy loss function is used to evaluate the information difference between the classification prediction and the true class label: ; where, is the one-hot encoding of the true class, is the prediction probability, and the matching categories are finely divided into three semantic relationships: exact match (two addresses point to exactly the same building entity), partial match (there is an inclusion or overlap relationship between addresses, such as the same block but different building units), and no match (addresses point to completely different building entities).

[0096] The joint loss function balances the task contributions and prevents overfitting by weighted summation and introducing L2 regularization: ; where, and is the task weight, initially set to 0.5, is the regularization coefficient, set to , are model parameters. The joint loss ensures the collaborative optimization of the two tasks and improves the processing ability of building attributes, addresses, and business information.

[0097] In some embodiments, the dynamic weight adjustment is optimized through task performance monitoring and , to solve the problem of unbalanced task convergence.

[0098] In some embodiments, during the joint training process, the performance change rate can be determined; based on the performance change rate, the first task weight and the second task weight are adjusted; wherein, the weight update formula is: ; ; wherein, is the adjustment rate, set to 0.1.

[0099] Optionally, based on the relative change in task performance, the weights can be dynamically adjusted to preferentially optimize the task with lagging performance, enhancing the generalization ability of the overall model. The weight adjustment is based on the performance change rate, defined as: ; wherein, is the round performance metric (accuracy for the parsing task and F1 score for the matching task).

[0100] When the performance of the parsing task stagnates , increases to strengthen the optimization of the parsing task, and vice versa to enhance the matching task. This adaptive adjustment mechanism based on performance differences not only balances the learning progress between tasks but also dynamically adjusts the optimization strategy according to data characteristics and model status, significantly improving the adaptability of the model in complex heterogeneous data environments.

[0101] Optionally, for the multi-task learning framework, a parameter sharing mechanism can be set up. By strategically reusing some parameters in the neural network hierarchy, it reduces the computational resource overhead while enhancing the implicit knowledge transfer between tasks.

[0102] The parameter sharing in the embodiments of this application may include a basic feature extraction layer, i.e., an encoding network. For example, the basic structure can adopt the GeoRoBERTa encoder. GeoRoBERTa is a geographically enhanced pre-trained language model based on the Transformer architecture. By sharing all the parameters of its 12-layer encoder, it provides a unified feature representation that fuses semantic and geographical information for address parsing and matching tasks: ; Among them, is the input sequence, is the token-level feature. Task-specific branches (IDCNN-BiLSTM-CRF for parsing and fully connected network for matching) use independent parameters to ensure the extraction of task-specific features. This strategy promotes cross-task learning by sharing the encoder while retaining branch flexibility.

[0103] The training adopts an alternating training strategy to enhance the co-optimization between tasks. In each round of mini-batch, a task is randomly selected for forward propagation and backward propagation, and the corresponding loss is calculated and the parameters are updated: ; Among them, is the learning rate, is the loss of the selected task. In the initial stage, a warm-up period of 2 epochs is set, and only the parameters of the GeoRoBERTa encoder are optimized to stabilize the shared feature learning. The total number of training epochs is 20. The best model is selected based on the comprehensive performance (parsing accuracy and matching F1 score) on the validation set. The training uses the Adam optimizer with an initial learning rate of , and the batch size is 16 to ensure the adaptability to multi-source heterogeneous data.

[0104] It should be noted that the multi-task learning framework has successfully achieved the efficient co-optimization of address parsing and address matching tasks through one or more of the following: joint loss function, strategic parameter sharing, alternating training mechanism, and innovative dynamic weight adjustment technology, providing strong technical support for the full-time and full-space semantic association of cross-department building data. This framework not only improves the model's processing ability for multi-source heterogeneous data but also enhances the system's robustness to different address expression forms and changes in building attributes, laying a solid foundation for data fusion in smart city construction.

[0105] Please refer to Figure 5 , the Figure 5 provides a method for creating building spatio-temporal objects based on the model framework provided in any embodiment of this application. As shown in Figure 5 , the exemplary process of creating building spatio-temporal objects based on the model is as follows: Obtain heterogeneous data from different departments and text data containing addresses.

[0106] For the text data, extract structured address information through an address parsing model to generate standard address data.

[0107] As an example, the text data can be long text data, such as text data with the number of characters greater than a preset quantity threshold. The quantity threshold can be set according to the actual application scenario, for example, 200.

[0108] Input the standard address data and heterogeneous data from different departments into the address matching model together.

[0109] The address matching model performs matching processing on the received standard address data and heterogeneous data from different departments.

[0110] If the matching is successful, extract the matching address pairs, and based on the matching address pairs and the heterogeneous data corresponding to the address pairs, construct building spatio-temporal objects, further integrate spatio-temporal attributes, and store the generated spatio-temporal objects.

[0111] If the matching fails, record the unmatched data and feedback for manual verification.

[0112] In this embodiment, the address matching model and / or the address parsing model are determined based on the embodiments of any model framework in the present application.

[0113] Through this method, effective integration and spatio-temporal objectification processing of building-related information from different sources are achieved.

[0114] Next, refer to Figure 6 , which shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present application. The terminal device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0115] As Figure 6As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0116] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0117] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above functions defined in the method of the embodiment of the present application are executed.

[0118] It should be noted that the computer-readable medium described above in the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. And in the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the program code readable by a computer is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0119] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0120] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.

[0121] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0122] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0123] The units involved in the embodiments described in this application can be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation on the unit itself in some cases. For example, the selection unit can also be described as "the unit for selecting the first type of pixel".

[0124] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, the exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.

[0125] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0126] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present application.

[0127] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present application. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0128] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A model framework for creating spatio-temporal objects of buildings, characterized in that, The model framework includes: an address parsing model and an address matching model. Among them, an encoding network is shared between the address parsing model and the address matching model. The input of the encoding network is a text address, and the output of the encoding network is a text-geographical fusion vector. The address parsing model and the address matching model are trained by a multi-task joint training method; The address parsing model includes an encoding network and an address parsing network; the input of the address parsing model is an address text, and the output of the address parsing model is an address label sequence; the address parsing model is used to convert the address text in the building multi-source data into an address label sequence; The address matching model includes an encoding network and an address matching network; the input of the address matching model is a pair of address texts, and the output of the address matching model is the relationship category between the pair of address texts; the address matching model is used to identify the address texts indicating the same building in the building multi-source data; The building spatio-temporal object of the target building includes building information and an address label sequence extracted from the building multi-source data indicating the target building.

2. The model framework according to claim 1, wherein The encoding network includes a tokenizer, an address label embedding module, a geographical coordinate encoding module, and a feature extraction layer; The tokenizer is used to convert the input address text into an address text vector; The address label embedding module is used to convert the address elements in the address text into address label vectors; The geographical coordinate encoding module is used to convert the address elements in the address text into geographical coordinate vectors; The input of the feature extraction layer is a text-geographical concatenated vector generated based on the address text vector, the address label vector, and the geographical coordinate vector, and the output of the feature extraction layer is a text-geographical fusion vector.

3. The model framework according to claim 2, wherein The geographical coordinate encoding adopts a spatial continuity preservation technique to convert longitude and latitude coordinates into multi-dimensional continuous vectors; The feature extraction layer adopts a 12-layer Transformer encoder structure, with each layer having a 768-dimensional hidden state and 12 attention heads.

4. The model framework according to claim 1, characterized in that, The building spatio-temporal object includes: A building identifier, which is the unique identifiable identifier of the building object; A building spatio-temporal reference, which is the spatio-temporal reference of the building object; The building spatial location; The building spatial form; The building attribute features; The behavior ability of the building, the response of the building to changes in the natural environment and to social planning measures; The change set of the building.

5. The model framework according to claim 1, wherein The address parsing network includes a dilated convolutional neural network, a bidirectional long short-term memory network, and a conditional random field layer; The dilated convolutional neural network is used to extract local multi-scale features from the text-geographical fusion vector output by the encoding network; The bidirectional long short-term memory network is used to extract the dependency relationship in the text-geographical fusion vector; The input of the conditional random field layer is a first fusion vector, which is obtained by splicing the outputs of the dilated convolutional neural network and the bidirectional long short-term memory network; The output of the conditional random field layer is an address label sequence.

6. The model framework according to claim 1, wherein The input of the address matching network is the paired text-geographical fusion vectors output by the encoding network; The address matching network includes a pooling layer, a feature interaction layer, and a fully connected network; The pooling layer is used to pool the input pairwise text geographical fusion vectors to obtain a first pooled result vector and a second pooled result vector; The feature interaction layer is used to fuse the first pooled result vector and the second pooled result vector to obtain a second fusion vector; The fully connected network is used to map the second fusion vector to the address category matching space.

7. The model framework according to claim 6, wherein The pooling layer performs average pooling and max pooling, and fuses the results of average pooling and max pooling; The feature interaction layer performs: calculating the difference vector between the first pooled result vector and the second pooled result vector, and calculating the dot product vector between the first pooled result vector and the second pooled result vector; concatenating the first pooled result vector, the second pooled result vector, the difference vector, and the dot product vector to obtain a second fusion vector.

8. The model framework according to claim 1, characterized in that The loss function of the address resolution model is a conditional random field loss function , and the loss function of the address matching model adopts a multi-class cross-entropy loss function ; The joint training process of the address parsing model and the address matching model includes: constructing a joint loss function; wherein, the joint loss function is: ; Among them, and is the task weight, initially set to 0.5, is the regularization coefficient, set to , is the model parameter.

9. The model framework according to claim 8, wherein The joint training process of the address parsing model and the address matching model includes: Determining the performance change rate during the joint training process; Adjusting the first task weight and the second task weight based on the performance change rate; wherein, the performance change rate is defined as: ; Among them, is the performance metric for the round. For the parsing task, accuracy is used, and for the matching task, score is used. Among them, the weight update formula is as follows: ; ; Among them, is the adjustment rate, set to 0.

1.

10. A creation method for creating a spatio-temporal object of a building, characterized in that, including: Obtaining heterogeneous data from different departments and text data containing addresses; For the text data, extracting structured address information through the address parsing model to generate standard address data; Inputting the standard address data and the heterogeneous data from different departments into the address matching model together; The address matching model performs matching processing on the received standard address data and the heterogeneous data from different departments; If the matching is successful, extracting the matching address pairs, and based on the matching address pairs and the heterogeneous data corresponding to the address pairs, constructing building spatio-temporal objects, further integrating spatio-temporal attributes, and storing the generated spatio-temporal objects; If the matching fails, recording the unmatched data and feedbacking for manual verification; wherein, the address matching model and / or the address parsing model is determined based on the model framework in any one of claims 1-9.

Citation Information

Patent Citations

  • Machine reading comprehension method based on multi-task joint training, and computer storage medium

    CN110309305A

  • Multi-task cascaded human face frame selection and comparison method

    CN111539351A

  • Entity interaction detection method and method and device for establishing entity interaction detection model

    CN115457529A

  • Chinese address sequence labeling method, system and device based on word segmentation and labeling task sharing and storage medium

    CN116542248A

  • Chinese address resolution method and system based on multi-modal geographic text pre-training

    CN117892718A