A method for embedding urban spatiotemporal data in multimodal large models
By constructing a spatiotemporal heterogeneous graph and hierarchical index of urban spatiotemporal data, and performing multimodal prompt word embedding and retrieval, the problem of insufficient multimodal data processing capabilities is solved, and the accuracy and efficiency of urban management are improved, resources are optimized, emergency response is strengthened, and support is provided for smart city construction and policy evaluation.
Patent Information
- Application Number
- CN202510670387.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-05-23
AI Technical Summary
Existing technologies cannot effectively process multimodal data, especially spatiotemporal data, cannot achieve urban calculations and predictions, and lack numerical computing capabilities.
Construct a spatiotemporal heterogeneous graph of urban spatiotemporal data, design a hierarchical index, perform multimodal prompt word embedding and spatiotemporal hierarchical retrieval, reconstruct prompt words into blocks in the form of HTML tags, and realize multi-scale mining and analysis of urban spatiotemporal data.
Improve the accuracy and efficiency of urban management, optimize resource allocation and planning, strengthen emergency response and risk management, promote smart city construction, and support policy formulation and evaluation.
Smart Images

Figure CN120179883B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of smart city technology, and in particular relates to an urban spatiotemporal data embedding method for a multimodal large model. Background Art
[0002] In recent years, the continuous maturity of technologies such as artificial intelligence, big data analysis, and multimodal large models has made large-scale model-based urban operations and management computation possible. The compatibility of multimodal large models with spatiotemporal data has become a common concern in both AI and spatiotemporal data applications.
[0003] Patent application number 202411443611.3, entitled "Method and Device for Evaluating a Multimodal Large Model of Urban Governance Based on Multi-layered Text," discloses a method and device for evaluating a multimodal large model of urban governance based on multi-layered text. The method comprises the following steps: obtaining the true labeled text of a sample image and the output text of a multimodal large model of urban governance; using an urban management vocabulary to segment the text into independent words and calculate the similarity scores of the word units; using the urban management vocabulary to create a bag-of-words model to convert the text data into a word frequency matrix and calculate the low-level word frequency similarity scores; using a trained language model to extract semantic features with contextual context from the text and convert them into vector representations and calculate the semantic similarity scores; using a general large language model for formatting requests with business guide words and evaluation requirements to derive the corresponding high-level semantic similarity scores for the text; and combining the similarity scores at each layer to calculate the comprehensive evaluation score of the multimodal large model. However, this method uses a traditional word segmentation language model, first segmenting long texts, then analyzing the word units therein, and then analyzing the text semantics to comprehensively evaluate whether two sentences are consistent from the perspective of urban management business. This method cannot achieve urban calculations and predictions based on urban data, but only provides an assessment of the accuracy of textual meaning. It only analyzes and calculates textual semantics and lacks the ability to process multimodal data, spatiotemporal data, or perform numerical calculations. Summary of the Invention
[0004] The problem to be solved by the present invention is to achieve in-depth mining and analysis of multi-scale spatiotemporal data involved in multi-type urban tasks, and propose an urban spatiotemporal data embedding method for multimodal large models.
[0005] To achieve the above object, the present invention is implemented through the following technical solutions:
[0006] A method for embedding urban spatiotemporal data in a multimodal large model includes the following steps:
[0007] S1. Based on the collected multi-scale spatiotemporal data of the city, construct a spatiotemporal heterogeneous graph of the urban spatiotemporal data and a spatiotemporal heterogeneous graph feature set of the urban spatiotemporal data;
[0008] S2. Design a hierarchical index for the spatiotemporal heterogeneous graph of the urban spatiotemporal data obtained in step S1, using spatial and temporal dimensions as the first-level index and assigning an index to each node entity as the second-level index.
[0009] S3. Construct a spatiotemporal hierarchical retrieval method for large-scale multimodal prompt words. First, embed the large-scale multimodal prompt words. Then, based on the hierarchical index of the spatiotemporal heterogeneous graph obtained in step S2, perform spatiotemporal hierarchical retrieval based on similarity metrics to obtain spatiotemporal hierarchical retrieval results.
[0010] S4. Based on the spatiotemporal hierarchical retrieval results obtained in step S3, the initial prompt words and retrieval results of the large model are reconstructed. The reconstructed prompt words are divided into blocks in the form of HTML tags to obtain new prompt words with retrieval content of the large model.
[0011] Furthermore, the specific implementation method of step S1 includes the following steps:
[0012] S1.1. Constructing a spatiotemporal heterogeneous graph of urban spatiotemporal data:
[0013] Based on the collected multi-scale spatiotemporal data of the city, a spatiotemporal heterogeneous map of the urban spatiotemporal data is constructed. , the expression is:
[0014] ;
[0015] in, Represents a A node set consisting of spatial entities, , For the N nodes, Represents the edges in the node set, , For the M edges, each node and each edge corresponds to a different type, is the feature set, Including the features corresponding to all nodes, is a collection of node types, is a set of edge types;
[0016] S1.2. Associate the nodes constructed in step S1.1:
[0017] for A spatial entity, defined from the spatial dimension Different spatial divisions are defined from the time dimension Different time partitions are used. In a spatiotemporal heterogeneous graph, the expression for the set of node types is:
[0018] ;
[0019] The size of the scales in the definition of space division and time division is arranged from large to small, and the definition Indicates The node types in the spatial and temporal regions represented by , , ,by express The number of nodes under the node type;
[0020] Based on establishing only one association between different nodes, the expression for the set of edge types is:
[0021] ;
[0022] S1.3. Feature Set for Constructing a Spatiotemporal Heterogeneous Graph of Urban Spatiotemporal Data:
[0023] Set the node type The feature space includes the natural language description space , a vector space describing statistical characteristics , and the vector space describing the image , represents the set of all possible strings, Represents node type The characteristic dimension of Represent the pixel size and number of channels of the image respectively;
[0024] The node type The feature set The expression is:
[0025] ;
[0026] in, Indicates The natural language description of the spatial and temporal regions represented by Indicates The set of features included in the nodes in the spatial and temporal regions represented by Indicates The image of the nodes in the spatial and temporal regions represented by ,
[0027] ;
[0028] ;
[0029] in, Indicates composition String, represents the set of all possible characters, Indicates The statistical characteristics of the nodes in the spatial and temporal regions represented by .
[0030] Furthermore, the specific implementation method of step S2 includes the following steps:
[0031] S2.1. Design the second-level index layer as the layer where the node is located. Build the second-level index for any node. The expression is:
[0032] ;
[0033] in, For the i nodes, For the i The second-level index of nodes, 、 、 They correspond to the trained word embedding model, vector encoding model, and image encoding model respectively. 、 、 Respectively represent i The natural language description, statistical features, and images corresponding to each node, Represents the concatenation of the two vectors before and after;
[0034] S2.2. Based on the second-level index designed in step S2.1, construct the first-level index layer, which is the layer where the time-space partition is located. The expression is:
[0035] ;
[0036] ;
[0037] in, represents a differentiable and permutation-invariant function, and represent the first differentiable function and the second differentiable function respectively, , , Represents different nodes, Indicates that in layer 2 The adjacent nodes of is a value close to 0, represents the attention coefficient of two nodes, represents the trainable parameter matrix, represents the natural exponential function, Indicates the node type in the first index layer The corresponding index.
[0038] Furthermore, the specific implementation method of step S3 includes the following steps:
[0039] S3.1. Perform large-scale multimodal cue word embedding. The embedding process is expressed as:
[0040] ;
[0041] in, Respectively represent u i The natural language description, statistical features, and images corresponding to the segment prompt words, Indicates that for u i The query index constructed by the segment prompt words;
[0042] If the u i If the natural language description, statistical features, or image information is missing from the segment prompt, the missing information is filled with the data obtained from the following corresponding calculations. The expression is:
[0043] ;
[0044] ;
[0045] ;
[0046] in, is the mean of the natural language description, is the mean of the statistical characteristics, is the mean of the image;
[0047] S3.2. Constructing a spatiotemporal hierarchical retrieval method based on similarity measurement:
[0048] S3.2.1. First, search the spatiotemporal scale. u i Query index constructed by segment prompt words , calculate its relationship with all space-time partitions Node similarity , the expression is:
[0049] ;
[0050] in, is the 2-norm;
[0051] Then for all Sort, select The node corresponding to the largest time-space partition calculates its relationship with all nodes under the partition Similarity , the expression is:
[0052] ;
[0053] Then for all Sort and select the largest The spatiotemporal features corresponding to the indexes are expressed as the retrieval results:
[0054] .
[0055] Furthermore, the specific implementation method of step S4 includes the following steps:
[0056] S4.1. Based on the search results obtained in step S3, reconstruct the initial prompt words and search results of the large model. The reconstructed expression is:
[0057] ;
[0058] in, is the new prompt word after reconstruction, is the assignment symbol, Indicates the overall description, adding a reference to the search results based on the initial prompt word. It is the description of the current task, and its content is the initial prompt word. for Search results, Output description, including requirements for output format;
[0059] S4.2. Divide the prompt words reconstructed in step S4.1 into blocks in the form of HTML tags, where the overall description section adds a reference to the indexed content, the current task description is the original prompt word, and the search content is from the first to the The items are listed in sequence, and each search content contains the area description corresponding to the retrieved spatial entity , area vector , and regional images , where the region description and region vector are given in character form, and the region image is converted into Base64 characters and input through the image interface of the multimodal large model.
[0060] Beneficial effects of the present invention:
[0061] The proposed method for embedding urban spatiotemporal data within a multimodal large-scale model can improve the accuracy and efficiency of urban management. By integrating and analyzing multi-source data on urban transportation, the environment, population, and economy, the model can monitor urban operations in real time, identify potential problems promptly, and provide accurate decision-making support. For example, in traffic management, the model can predict traffic flow, optimize signal timing, reduce congestion, and improve traffic efficiency.
[0062] The proposed method for embedding urban spatiotemporal data within a multimodal large-scale model optimizes resource allocation and planning. By analyzing urban regions, the model identifies imbalances in resource distribution and assists in developing rational resource allocation plans. In urban planning, the model can simulate the impact of different planning options on urban development, helping decision-makers select the optimal solution and enhancing the sustainability of urban development.
[0063] The proposed method for embedding urban spatiotemporal data within a multimodal large-scale model can enhance emergency response and risk management. When emergencies or natural disasters occur, the model can quickly assess the impact, predict potential subsequent developments, and assist in developing emergency plans, thereby improving the speed and effectiveness of emergency responses. For example, in flood warnings, the model can predict flood spread paths, allowing for early evacuation and minimizing losses.
[0064] The proposed method for embedding urban spatiotemporal data within a multimodal large-scale model can facilitate the development of smart cities. This model provides data support and decision-making for smart city development. By deeply mining various urban data, the model can provide technical support for smart transportation, smart healthcare, smart security, and other fields, enhancing the intelligence of urban services.
[0065] The proposed method for embedding urban spatiotemporal data in a multimodal large-scale model can support policy formulation and evaluation. The model can simulate the impact of different policies on urban development, assisting policymakers in evaluating their effectiveness and optimizing their design. For example, when formulating environmental protection policies, the model can predict changes in air quality after policy implementation and assess their effectiveness. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 This is a flow chart of a method for embedding urban spatiotemporal data in a multimodal large model according to the present invention;
[0067] Figure 2 This is an example diagram of the search results of the present invention;
[0068] Figure 3 This is an example diagram of the large model prompt words after reconstruction of the present invention. DETAILED DESCRIPTION
[0069] In order to make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present invention and are not intended to limit the present invention. That is, the specific embodiments described herein are only some embodiments of the present invention, not all embodiments. Generally, the components of the specific embodiments of the present invention described and illustrated in the drawings herein can be arranged and designed in various different configurations, and the present invention can also have other embodiments.
[0070] Therefore, the following detailed description of the specific embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but is merely representative of selected specific embodiments of the present invention. All other specific embodiments obtained by those skilled in the art based on the specific embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0071] In order to further understand the content, features and effects of the present invention, the following specific embodiments are given as examples, and the attached Figure 1 -Attached Figure 3 The detailed instructions are as follows:
[0072] Example 1:
[0073] A method for embedding urban spatiotemporal data in a multimodal large model includes the following steps:
[0074] S1. Based on the collected multi-scale spatiotemporal data of the city, construct a spatiotemporal heterogeneous graph of the urban spatiotemporal data and a spatiotemporal heterogeneous graph feature set of the urban spatiotemporal data;
[0075] Furthermore, the specific implementation method of step S1 includes the following steps:
[0076] S1.1. Constructing a spatiotemporal heterogeneous graph of urban spatiotemporal data:
[0077] Based on the collected multi-scale spatiotemporal data of the city, a spatiotemporal heterogeneous map of the urban spatiotemporal data is constructed. , the expression is:
[0078] ;
[0079] in, Represents a A node set consisting of spatial entities, , For the N nodes, Represents the edges in the node set, , For theM edges, each node and each edge corresponds to a different type, is the feature set, Including the features corresponding to all nodes, is a collection of node types, is a set of edge types;
[0080] S1.2. Associate the nodes constructed in step S1.1:
[0081] for A spatial entity, defined from the spatial dimension Different spatial divisions are defined from the time dimension Different time partitions are used. In a spatiotemporal heterogeneous graph, the expression for the set of node types is:
[0082] ;
[0083] The size of the scales in the definition of space division and time division is arranged from large to small, and the definition Indicates The node types in the spatial and temporal regions represented by , , ,by express The number of nodes under the node type;
[0084] For example Indicates the node type corresponding to the first spatial division and the third temporal division, and The spatial scale represented is To be big, such as Represents the entire city area, Indicates a district or county level area; for example The time scale represented is To be big, such as Represents monthly statistical data, Indicates daily statistical data.
[0085] Based on establishing only one association between different nodes, the expression for the set of edge types is:
[0086] ;
[0087] S1.3. Feature Set for Constructing a Spatiotemporal Heterogeneous Graph of Urban Spatiotemporal Data:
[0088] Set the node type The feature space includes the natural language description space , a vector space describing statistical characteristics , and the vector space describing the image , represents the set of all possible strings, Represents node type The characteristic dimension of Represent the pixel size and number of channels of the image respectively;
[0089] The node type The feature set The expression is:
[0090] ;
[0091] in, Indicates The natural language description of the spatial and temporal regions represented by Indicates The set of features included in the nodes in the spatial and temporal regions represented by Indicates The image of the nodes in the spatial and temporal regions represented by ;
[0092] ;
[0093] ;
[0094] in, Indicates composition String, represents the set of all possible characters, Indicates The statistical characteristics of the nodes in the spatial and temporal regions represented by .
[0095] For example Contains the node in a Various features in the spatial and temporal regions represented by Represents the statistical data of a certain day in a certain district or county (such as weather conditions, number of visitors, average road congestion index, cumulative travel volume, and number of takeout orders). Represents some unstructured natural language descriptions of the area on that day (such as the area location description, the description of the events that occurred in the area), Indicates the layout image of the area on that day (such as the layout of charging stations open for use, POI layout). The number of nodes is different for different types of nodes. Generally speaking, and The smaller the value of is, the fewer the spatial and temporal regions of this type are, and the fewer the number of nodes is. For all nodes, the spatial division and time division The finer the The more node types are included, the more and The value of is discussed in detail.
[0096] S2. Design a hierarchical index for the spatiotemporal heterogeneous graph of the urban spatiotemporal data obtained in step S1, using spatial and temporal dimensions as the first-level index and assigning an index to each node entity as the second-level index.
[0097] Furthermore, a hierarchical index construction technology based on heterogeneous graph embedding network is proposed, which divides the index into two layers, thereby significantly reducing the size of each search range. Specifically, for urban spatiotemporal heterogeneous graph data , using space and time dimension division as the first layer index, the number of indexes is ,in Indicates the size of a set. The second-level index is all node entities, each node entity corresponds to an index, and the number of indexes is Using this method, for a node-level data retrieval task, if a conventional index is used for a retrieval, the original retrieval range is ; After adopting the hierarchical index, the size of the two-layer retrieval range is and , the sizes of these two search ranges are much smaller than , thus significantly improving the retrieval efficiency.
[0098] Furthermore, the specific implementation method of step S2 includes the following steps:
[0099] S2.1. Design the second-level index layer as the layer where the node is located. Build the second-level index for any node. The expression is:
[0100] ;
[0101] in, For the i nodes, For the i The second-level index of nodes, 、 、 They correspond to the trained word embedding model, vector encoding model, and image encoding model respectively. 、 、 Respectively representi The natural language description, statistical features, and images corresponding to each node, Represents the concatenation of the two vectors before and after;
[0102] S2.2. Based on the second-level index designed in step S2.1, construct the first-level index layer, which is the layer where the time-space partition is located. The expression is:
[0103] ;
[0104] ;
[0105] in, represents a differentiable and permutation-invariant function, and represent the first differentiable function and the second differentiable function respectively, , , Represents different nodes, Indicates that in layer 2 The adjacent nodes of is a value close to 0, represents the attention coefficient of two nodes, represents the trainable parameter matrix, represents the natural exponential function, Indicates the node type in the first index layer The corresponding index;
[0106] Further, Indicates one of the following: sum, mean, minimum, and maximum;
[0107] S3. Construct a spatiotemporal hierarchical retrieval method for large-scale multimodal prompt words. First, embed the large-scale multimodal prompt words. Then, based on the hierarchical index of the spatiotemporal heterogeneous graph obtained in step S2, perform spatiotemporal hierarchical retrieval based on similarity metrics to obtain spatiotemporal hierarchical retrieval results.
[0108] Furthermore, the specific implementation method of step S3 includes the following steps:
[0109] S3.1. Perform large-scale multimodal cue word embedding. The embedding process is expressed as:
[0110] ;
[0111] in, 、 、 Respectively represent u i The natural language description, statistical features, and images corresponding to the segment prompt words, Indicates that for u i The query index constructed by the segment prompt words;
[0112] If the u i If the natural language description, statistical features, or image information is missing from the segment prompt, the missing information is filled with the data obtained from the following corresponding calculations. The expression is:
[0113] ;
[0114] ;
[0115] ;
[0116] in, is the mean of the natural language description, is the mean of the statistical characteristics, is the mean of the image;
[0117] S3.2. Constructing a spatiotemporal hierarchical retrieval method based on similarity measurement:
[0118] S3.2.1. First, search the spatiotemporal scale. u i Query index constructed by segment prompt words , calculate its relationship with all space-time partitions Node similarity , the expression is:
[0119] ;
[0120] in, is the 2-norm;
[0121] Then for all Sort, select The node corresponding to the largest time-space partition calculates its relationship with all nodes under the partition Similarity , the expression is:
[0122] ;
[0123] Then for all Sort and select the largest The spatiotemporal features corresponding to the indexes are expressed as the retrieval results:
[0124] .
[0125] Based on the index of the large model prompt word (i.e., the embedding vector) and the hierarchical index in the spatiotemporal heterogeneous graph, the similarity between the two is first calculated, and then the index content with higher similarity is filtered from the spatiotemporal heterogeneous graph to complete the retrieval.
[0126] S4. Based on the spatiotemporal hierarchical retrieval results obtained in step S3, the initial prompt words and retrieval results of the large model are reconstructed. The reconstructed prompt words are divided into blocks in the form of HTML tags to obtain new prompt words with retrieval content of the large model.
[0127] Furthermore, the specific implementation method of step S4 includes the following steps:
[0128] S4.1. Based on the search results obtained in step S3, reconstruct the initial prompt words and search results of the large model. The reconstructed expression is:
[0129] ;
[0130] in, is the new prompt word after reconstruction, is the assignment symbol, Indicates the overall description, adding a reference to the search results based on the initial prompt word. It is the description of the current task, and its content is the initial prompt word. for Search results, Output description, including requirements for output format;
[0131] For a search result , and its corresponding content is as follows Figure 2 As shown, respectively A spatial region at a certain scale, and Regional description at a certain time scale , area vector and regional images .
[0132] S4.2. Divide the prompt words reconstructed in step S4.1 into blocks in the form of HTML tags, where the overall description section adds a reference to the indexed content, the current task description is the original prompt word, and the search content is from the first to the The items are listed in sequence, and each search content contains the area description corresponding to the retrieved spatial entity , area vector , and regional images , where the region description and region vector are given in character form, and the region image is converted into Base64 characters and input through the image interface of the multimodal large model.
[0133] Furthermore, the order of all inputs is as follows Figure 3 shown.
[0134] The key technical points and points to be protected of the present invention are:
[0135] The multi-scale modeling technology of urban spatiotemporal data in the present invention is the multi-scale modeling technology of urban spatiotemporal data based on heterogeneous graphs proposed in step 1.
[0136] The urban spatiotemporal data index construction technology in the present invention is the spatiotemporal heterogeneous graph hierarchical index construction technology based on graph embedding proposed in step 2.
[0137] The multimodal prompt word related content retrieval technology in the present invention is the spatiotemporal hierarchical retrieval technology for the multimodal large model proposed in step 3.
[0138] The multimodal prompt word construction technology of spatiotemporal data in the present invention is the multimodal prompt word reconstruction technology based on the spatiotemporal heterogeneous graph retrieval results proposed in step 4.
[0139] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
[0140] Although the present application has been described above with reference to specific embodiments, various modifications may be made thereto and components may be substituted with equivalents without departing from the scope of the present application. In particular, as long as there are no structural conflicts, the various features of the embodiments disclosed herein may be combined with each other in any manner, and the omission of an exhaustive description of these combinations in this specification is solely for the sake of space and resource conservation. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions within the scope of the claims.
Claims
1. A method for embedding urban spatiotemporal data in a multimodal large model, characterized by: The steps include: S1. Based on the collected multi-scale spatiotemporal data of the city, construct a spatiotemporal heterogeneous graph of the urban spatiotemporal data and a spatiotemporal heterogeneous graph feature set of the urban spatiotemporal data; S2. Design a hierarchical index for the spatiotemporal heterogeneous graph of the urban spatiotemporal data obtained in step S1, using spatial and temporal dimensions as the first-level index and assigning an index to each node entity in all node entities as the second-level index; S3. Construct a spatiotemporal hierarchical retrieval method for large-scale multimodal prompt words. First, embed the large-scale multimodal prompt words. Then, based on the hierarchical index of the spatiotemporal heterogeneous graph obtained in step S2, perform spatiotemporal hierarchical retrieval based on similarity metrics to obtain spatiotemporal hierarchical retrieval results. The specific implementation method of step S3 includes the following steps: S3.
1. Perform large-scale multimodal prompt word embedding. The embedding process is expressed as: in, Respectively represent the uth i The natural language description, statistical features, and images corresponding to the segment prompt words, Indicates that for the uth i The query index constructed by the segment prompt word; Emb L 、Emb Q 、Emb M They correspond to the trained word embedding model, vector encoding model, and image encoding model respectively; ‖ represents the concatenation of the two vectors; If u i If the natural language description, statistical features, or image information is missing from the segment prompt, the missing information is filled with the data obtained from the following corresponding calculations. The expression is: Among them, Emb L ′ is the mean of natural language description, Emb Q ' is the mean of the statistical characteristics, Emb M ' is the mean of the image; Respectively represent the natural language description, statistical features, and image corresponding to the i-th node; v i is the i-th node; V represents a node set consisting of N spatial entities; S3.
2. Constructing a spatiotemporal hierarchical retrieval method based on similarity measurement: S3.2.
1. First, perform a retrieval of the spatiotemporal scale. i Query index constructed by segment prompt words Calculate its relationship with all space-time partitions The node similarity d s,t , the expression is: Among them, ||·||2 is the 2-norm; Then for all d s,t To sort, select d s,t The node corresponding to the largest time-space partition calculates its relationship with all nodes under the partition The similarity d n , the expression is: Then for all d n Sort and select the spatiotemporal features corresponding to the largest K indexes as the retrieval results. The expression is: in, Represents a natural language description in the spatial region and time region represented by (s, t); Represents the statistical characteristics of the nodes in the spatial and temporal regions represented by (s, t); represents the images included in the nodes in the spatial and temporal regions represented by (s, t); n is any one of K, and K is the total number of search results; S4. Based on the spatiotemporal hierarchical retrieval results obtained in step S3, the initial prompt words and retrieval results of the large model are reconstructed, and the reconstructed prompt words are divided into blocks in the form of HTML tags to obtain new prompt words with retrieval content of the large model.
2. The urban spatiotemporal data embedding method for a multimodal large model according to claim 1 is characterized in that: The specific implementation method of step S1 includes the following steps: S1.
1. Constructing a spatiotemporal heterogeneous graph of urban spatiotemporal data: Based on the collected multi-scale spatiotemporal data of the city, a spatiotemporal heterogeneous map of the urban spatiotemporal data is constructed. The expression is: in, Represents a node set consisting of N spatial entities, v N is the Nth node, ε represents the edge in the node set, ε={e1,…,e M }, e M For the Mth edge, each node and each edge corresponds to a different type, is the feature set, Including the features corresponding to all nodes, Γ v is the set of node types, Γ ε is a set of edge types; S1.
2. Associate the nodes constructed in step S1.1: For N spatial entities, define N from the spatial dimension s Different spatial divisions are defined from the time dimension. t Different time partitions are used. In a spatiotemporal heterogeneous graph, the expression for the set of node types is: C v ={1,…,N s }×{1,…,N t }; The size of the scales in the definition of space division and time division is arranged from large to small, and the definition Indicates the node type in the spatial region and time region represented by (s, t), s∈{1,…,N s }, t∈{1,…,N t }, with N (s,t) express The number of nodes under the node type; Based on establishing only one association between different nodes, the expression for the set of edge types is: C ε =C v ×C V ; S1.
3. Constructing a feature set of spatiotemporal heterogeneous graphs for urban spatiotemporal data: Set the node type The feature space includes the natural language description space Σ * , a vector space describing statistical characteristics And the vector space describing the image Σ * represents the set of all possible strings, F (s,t) Represents node type The feature dimension of , w, c represent the pixel size and number of channels of the image respectively; The node type The feature set The expression is: in, represents the natural language description in the spatial region and time region represented by (s, t), Represents the set of features included in the nodes in the spatial region and time region represented by (s, t), Represents the image of the nodes in the spatial region and time region represented by (s, t), in, Indicates composition String, Σ * represents the set of all possible characters, Represents the statistical characteristics of the nodes in the spatial and temporal regions represented by (s, t), 3. The urban spatiotemporal data embedding method for a multimodal large model according to claim 2 is characterized in that: The specific implementation method of step S2 includes the following steps: S2.
1. Design the second-level index layer as the layer where the node is located. Build the second-level index for any node. The expression is: Among them, v i is the i-th node, is the second-level index of the i-th node, Emb L 、Emb Q 、Emb M They correspond to the trained word embedding model, vector encoding model, and image encoding model respectively. Respectively represent the natural language description, statistical features, and image corresponding to the i-th node, and ‖ represents the concatenation of the two vectors before and after; S2.
2. Based on the second-level index designed in step S2.1, construct the first-level index layer, which is the layer where the time and space division is located. The expression is: in, represents a differentiable and permutation-invariant function, φ Θ and ψ Θ Represent the first differentiable function and the second differentiable function, v i , v j , v k Represents different nodes, Indicates that in the second layer with v i The adjacent nodes, ∈ is a value close to 0, α ij Represents the attention coefficient of the two nodes, W represents the trainable parameter matrix, exp represents the natural exponential function, Indicates the node type in the first index layer The corresponding index.
4. The urban spatiotemporal data embedding method for a multimodal large model according to claim 3 is characterized in that: The specific implementation method of step S4 includes the following steps: S4.
1. Based on the search results obtained in step S3, the initial prompt words and search results of the large model are reconstructed. The reconstructed expression is: Among them, {prompt} is the new prompt word after reconstruction, := is the assignment symbol, prompt foreword Indicates the overall description, adding a reference to the search results based on the initial prompt word, prompt task It is the description of the current task, and its content is the initial prompt word. For K search results, prompt result Output description, including requirements for output format; S4.
2. The prompt words reconstructed in step S4.1 are divided into blocks in the form of HTML tags. The overall description part adds a reference to the index content. The current task description is the original prompt word. The search content is listed from the 1st to the Kth item. Each search content contains the area description corresponding to the retrieved spatial entity. Area Vector and regional images The region description and region vector are given in character form, and the region image is converted into Base64 characters and input through the image interface of the multimodal large model.
Citation Information
Patent Citations
Multimodal Large Model Evaluation Method and Device for Urban Governance Based on Multiple Layers of Text
CN119399006B
Distributed intelligent retrieval method and system for spatio-temporal data
CN118838937A
Device, Method and Program for Managing Area Information
US20080133484A1