Multi-modal large model-oriented city spatio-temporal data embedding method
By constructing the spatiotemporal heterogeneous diagram and feature set of urban spatiotemporal data, designing a spatiotemporal hierarchical search method for hierarchical indexes and multimodal prompt words, the problem of in-depth mining and analysis of multi-scale spatiotemporal data in the existing technology is solved, and the precise analysis and prediction of urban multi-type tasks is realized, and the accuracy and efficiency of urban management are improved.
Patent Information
- Application Number
- CN202510670387.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-23
AI Technical Summary
The prior art is difficult to achieve in-depth mining and analysis of multi-scale spatiotemporal data involved in urban multi-type tasks, and lacks the processing capabilities of multi-modal data and spatiotemporal data.
A method of embedding urban spatiotemporal data for multimodal large models is proposed. By constructing spatiotemporal heterogeneous diagrams and feature sets of urban spatiotemporal data, designing hierarchical indexes, and using a spatiotemporal hierarchical search method for multimodal prompt words, the initial prompt words of the large model are reconstructed to realize data embedding.
It realizes accurate analysis and prediction of urban multi-type tasks, improves the accuracy and efficiency of urban management, and supports smart city construction and policy formulation.
Smart Images

Figure CN120179883A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of smart cities, and specifically relates to a method for embedding urban spatio-temporal data for multi-modal large models. Background Art
[0002] In recent years, with the continuous maturity of technologies such as artificial intelligence, big data analysis, and multi-modal large models, it has become possible to perform urban operation and management calculations based on large models. The compatibility of multi-modal large models with spatio-temporal data has become a common concern in the fields of AI technology and spatio-temporal data applications.
[0003] The invention patent with the application number 202411443611.3 and the invention name "Method and device for evaluating a multi-modal large model for urban governance based on text multi-levels" discloses a method and device for evaluating a multi-modal large model for urban governance based on text multi-levels, including the following steps: obtaining the true label text of a sample image and the output text of a multi-modal large model for urban governance; using an urban management thesaurus to segment the text into independent words and calculating the similarity score of lexical units; using an urban management thesaurus to create a bag-of-words model to convert the text data into a term frequency matrix and calculating the low-level term frequency similarity score; using a trained language model to extract the semantic features with context in the text and converting them into vector representations and calculating the semantic similarity score; using a formatted request general large language model with business guiding words and evaluation requirements to obtain the high-level semantic similarity score corresponding to the text; combining the similarity scores at each level to calculate the comprehensive evaluation score of the multi-modal large model. However, this method uses a traditional word segmentation language model, first segments the long text, then analyzes the lexical units therein, and then analyzes the text semantics, so as to comprehensively evaluate whether two sentences are consistent from the perspective of urban management operations. This method cannot achieve urban calculations and predictions based on urban data, and only realizes the evaluation of the accuracy of text meaning. This method only analyzes and calculates the text semantics, does not have the processing ability for multi-modal data and spatio-temporal data, and does not have the ability of numerical calculation. Summary of the Invention
[0004] The problem to be solved by the present invention is to realize the in-depth mining and analysis of multi-scale spatio-temporal data involved in various urban tasks, and propose a method for embedding urban spatio-temporal data for multi-modal large models.
[0005] To achieve the above object, the present invention is realized through the following technical solutions: A method for embedding urban spatio-temporal data for multi-modal large models, including the following steps: S1. Based on the collected multi-scale spatio-temporal data of the city, construct a spatio-temporal heterogeneous graph of the urban spatio-temporal data and a spatio-temporal heterogeneous graph feature set of the urban spatio-temporal data; S2. Design a hierarchical index for the spatio-temporal heterogeneous graph of the urban spatio-temporal data obtained in step S1. Use the spatial and temporal dimension divisions as the first-level index, and assign an index to each node entity in all node entities as the second-level index; S3. Construct a spatio-temporal hierarchical retrieval method for large model multi-modal prompts. First, perform large model multi-modal prompt embedding, and then perform spatio-temporal hierarchical retrieval based on similarity measurement using the hierarchical index of the spatio-temporal heterogeneous graph obtained in step S2 to obtain the spatio-temporal hierarchical retrieval result; S4. Based on the spatio-temporal hierarchical retrieval result obtained in step S3, reconstruct the initial prompt and retrieval result of the large model. The reconstructed prompt is chunked in the form of HTML tags to obtain a new prompt of the large model with retrieved content.
[0006] Further, the specific implementation method of step S1 includes the following steps: S1.1. Construct a spatio-temporal heterogeneous graph of urban spatio-temporal data: Based on the collected multi-scale spatio-temporal data of the city, construct a spatio-temporal heterogeneous graph of urban spatio-temporal data , the expression is: ; Among them, represents a node set composed of spatial entities, , is the N th node, represents the edges formed in the node set, , is the M th edge, and each node and each edge correspond to different types, is the feature set, including the features corresponding to all nodes, is the set of node types, is the set of edge types; S1.2. Associate the nodes constructed in step S1.1: For spatial entities, define types of different spatial divisions from the spatial dimension and define types of different temporal divisions from the temporal dimension. In the spatio-temporal heterogeneous graph, the expression for obtaining the set of node types is: ; Define that the scales of the spatial division and the temporal division are arranged from large to small, and define to represent The node types within the represented spatial and temporal regions , , , to represent the number of nodes under the node type; Based on different nodes, establish associations only once, and the expression for the set of edge types is: ; S1.3. Construct the feature set of the spatio-temporal heterogeneous graph of urban spatio-temporal data: Set the feature space of the node type to include the natural language description space , the vector space describing statistical features , and the vector space describing images , representing all possible string sets, representing the feature dimension of the node type , respectively representing the pixel size and the number of channels of the image; Then the feature set of the node type has the expression: ; Among them, represents the natural language description within the spatial and temporal regions represented by , represents the set of features included in the nodes within the spatial and temporal regions represented by , represents the images included in the nodes within the spatial and temporal regions represented by , , ; ; Among them, represents the string that makes up , represents all possible character sets, represents the statistical features included in the nodes within the spatial and temporal regions represented by , .
[0007] Furthermore, the specific implementation method of step S2 includes the following steps: S2.1. Design the second-layer index layer as the layer where the nodes are located, and construct the second-layer index for any node. The expression is: ; Among them, is the i th node, is the second - layer index of the i th node, , , correspond to a trained existing word - embedding model, vector - encoding model, and image - encoding model respectively, , , respectively represent the natural - language description, statistical features, and image corresponding to the i th node; represents the concatenation of two vectors; S2.2. Based on the second - layer index designed in step S2.1, construct the first - layer index layer, which is the layer where spatio - temporal partitioning is located, and obtain the expression: ; ; Among them, represents a differentiable and permutation - invariant function, and respectively represent the first differentiable function and the second differentiable function, , , represent different nodes, represents the adjacent node to in the second layer, is a value close to 0, represents the attention coefficient between two nodes, represents a trainable parameter matrix, represents the natural - exponential function, represents the index corresponding to the node type in the first - layer index layer.
[0008] Furthermore, the specific implementation method of step S3 includes the following steps: S3.1. Perform large - model multi - modal prompt - word embedding, and the expression during the embedding process is: ; Among them, respectively represent the natural - language description, statistical features, and image corresponding to the u i th segment of prompt words, represents the query index constructed for the u i th segment of prompt words; If the ui If the natural language description, statistical features, and information in the image are missing in the segment prompt words, the missing information is filled with null values using the data calculated by the following corresponding calculations. The expression is: ; ; ; where, is the mean of the natural language description, is the mean of the statistical features, is the mean of the image; S3.2. Construct a spatio-temporal hierarchical retrieval method based on similarity measurement: S3.2.1. First, perform spatio-temporal scale retrieval. For the query index constructed by the u i segment prompt words , calculate its node similarity with all spatio-temporal partitions . The expression is: ; where, is the 2-norm; Then sort all , and select the node corresponding to the spatio-temporal partition with the largest to calculate its similarity with all nodes under this partition. The expression is: ; Then sort all , and select the largest indices to obtain the spatio-temporal features corresponding to them as the retrieval result. The expression is: .
[0009] Furthermore, the specific implementation method of step S4 includes the following steps: S4.1. Based on the retrieval result obtained in step S3, reconstruct the initial prompt word and the retrieval result of the large model. The reconstruction expression is: ; where, is the new prompt word after reconstruction, is the assignment symbol, represents the overall description, which adds a reference method to the retrieval result on the basis of the initial prompt word, is the current task description, and its content is the initial prompt word, is retrieval results is the output description, including the requirements for the output format; S4.2. Chunk the reconstructed prompt words in step S4.1 in the form of HTML tags. In the overall description part, add a reference prompt for the indexed content. The current task description is the original prompt words. The retrieved content is listed sequentially from the 1st to the number of items. Each retrieved content includes the regional description corresponding to the retrieved spatial entity , regional vector , and regional image . Among them, the regional description and regional vector are given in character form, and the regional image is converted into a Base64 character and input through the image interface of the multimodal large model.
[0010] Advantages of the present invention: The urban spatio-temporal data embedding method for a multimodal large model according to the present invention can improve the accuracy and efficiency of urban management; through the integration and analysis of multi-source data such as urban traffic, environment, population, and economy, this model can monitor the urban operation status in real time, timely discover potential problems, and provide accurate decision-making support. For example, in traffic management, the model can predict traffic flow, optimize signal timing, reduce traffic congestion, and improve traffic efficiency.
[0011] The urban spatio-temporal data embedding method for a multimodal large model according to the present invention can optimize resource allocation and planning; the model analyzes each region of the city to identify the imbalance of resource distribution and assist in formulating a reasonable resource allocation plan. In urban planning, the model can simulate the impact of different planning schemes on urban development, help decision-makers select the optimal scheme, and enhance the sustainability of urban development.
[0012] The urban spatio-temporal data embedding method for a multimodal large model according to the present invention can strengthen emergency response and risk management; in the event of an emergency or natural disaster, the model can quickly evaluate the impact scope of the event, predict possible subsequent developments, assist in formulating emergency plans, and improve the speed and effectiveness of emergency response. For example, in flood warning, the model can predict the flood spread path, evacuate people in advance, and reduce losses.
[0013] The urban spatio-temporal data embedding method for a multimodal large model according to the present invention can promote the construction of smart cities; this model provides data support and decision-making basis for the construction of smart cities. Through in-depth mining of various urban data, the model can provide technical support for fields such as smart transportation, smart healthcare, and smart security, and improve the intelligent level of urban services.
[0014] A method for embedding urban spatio-temporal data for multi-modal large models according to the present invention can support policy formulation and evaluation; the model can simulate the impacts of different policies on urban development, assist policymakers in evaluating policy effects, and optimizing policy design. For example, when formulating environmental protection policies, the model can predict the changes in air quality after the implementation of the policies and evaluate the effectiveness of the policies. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a flowchart of a method for embedding urban spatio-temporal data for multi-modal large models according to the present invention; Figure 2 is an example diagram of the retrieval results of the present invention; Figure 3 is an example diagram of the prompt words of the large model reconstructed according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the specific embodiments described are only a part of the embodiments of the present invention, rather than all of the specific embodiments. Usually, the components of the specific embodiments of the present invention described and shown in the accompanying drawings herein can be arranged and designed in various different configurations, and the present invention can also have other embodiments.
[0017] Therefore, the detailed description of the specific embodiments of the present invention provided in the accompanying drawings below is not intended to limit the scope of the claimed invention, but merely represents the selected specific embodiments of the present invention. All other specific embodiments obtained by those skilled in the art based on the specific embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0018] To further understand the content, features and effects of the present invention, the following specific embodiments are exemplified and combined with the attached Figure 1 - attached Figure 3 are described in detail as follows:
[0019] Embodiment 1: A method for embedding urban spatio-temporal data for multi-modal large models includes the following steps: S1. Based on the collected multi-scale spatio-temporal data of the city, construct a spatio-temporal heterogeneous graph of the urban spatio-temporal data and a spatio-temporal heterogeneous graph feature set of the urban spatio-temporal data; Further, the specific implementation method of step S1 includes the following steps: S1.1. Construct a spatio-temporal heterogeneous graph of the urban spatio-temporal data: Construct a spatio-temporal heterogeneous graph of urban spatio-temporal data based on the collected multi-scale spatio-temporal data of the city , the expression is: ; Among them, represents a set of nodes composed of spatial entities, , is the N th node, represents the edges formed in the set of nodes, , is the M th edge, and each node and each edge correspond to different types, is the set of features, including the features corresponding to all nodes, is the set of node types, is the set of edge types; S1.2. Associate the nodes constructed in step S1.1: For spatial entities, define types of different spatial partitions from the spatial dimension, and define types of different temporal partitions from the temporal dimension. In the spatio-temporal heterogeneous graph, the expression for obtaining the set of node types is: ; Define that the scales of the spatial partition and the temporal partition are arranged from large to small. Define to represent the node types within the spatial region and temporal region represented by , , , . Take to represent the number of nodes under the node type; For example represents the node type corresponding to the 1st spatial partition and the 3rd temporal partition, while represents a spatial scale that is larger than . For example represents the entire urban area, represents the district or county-level area; Another example is represents a temporal scale that is larger than . For example represents monthly statistical data, represents daily statistical data.
[0020] Only establish associations once for different nodes, and the expression for obtaining the set of edge types is: ; S1.3. Construct the feature set of the spatio-temporal heterogeneous graph of urban spatio-temporal data: Set the feature space of the node type to include the natural language description space , the vector space of descriptive statistical features , and the vector space of descriptive images , represents the set of all possible strings, represents the feature dimension of the node type , respectively represent the pixel size and the number of channels of the image; Then the feature set of the node type has the expression: ; Among them, represents the natural language description within the spatial and temporal regions represented by , represents the set of features included in the nodes within the spatial and temporal regions represented by , represents the images included in the nodes within the spatial and temporal regions represented by , ; ; ; Among them, represents the string that makes up , represents the set of all possible characters, represents the statistical features included in the nodes within the spatial and temporal regions represented by , .
[0021] For example contains various features of the node within a spatial and temporal region represented by , represents the statistical data (such as weather conditions, number of staying people, average road congestion index, cumulative travel volume, number of takeaway orders) on a certain day in a certain district or county-level region, represents some unstructured natural language descriptions (such as regional location descriptions, descriptions of events occurring within the region) on that day in that region, Indicates the layout image of the area on that day (such as the layout diagram of charging stations in open use, the POI layout diagram). For different types of nodes, the number of nodes is different. Generally speaking, and The smaller the value, the fewer the spatial and temporal regions under this type, and thus the fewer the number of nodes. For all nodes, the finer the spatial division and the temporal division , the more types of nodes it contains. The present invention does not specifically discuss the and values.
[0022] S2. Design a hierarchical index for the spatio-temporal heterogeneous graph of the urban spatio-temporal data obtained in step S1. Use the spatial and temporal dimension divisions as the first-level index, and correspond each node entity in all node entities to an index as the second-level index; Furthermore, a hierarchical index construction technique based on the heterogeneous graph embedding network is proposed, and the index is designed with upper and lower two layers, thereby significantly reducing the size of the retrieval range each time. Specifically, for the urban spatio-temporal heterogeneous graph data , use the spatial and temporal dimension divisions as the first-level index, and the number of indexes is , where represents the size of a set. The second-level index is all node entities, and each node entity corresponds to an index, and the number of indexes is . Using this method, for a node-level data retrieval task, for one retrieval, if a conventional index is used, the size of its original retrieval range is ; after using the hierarchical index, the sizes of its two-layer retrieval ranges are and , and the sizes of both of these two retrieval ranges are much smaller than , so the retrieval efficiency is significantly improved.
[0023] Furthermore, the specific implementation method of step S2 includes the following steps: S2.1. Design the second-level index layer as the layer where the node is located, and construct the second-level index for any node. The expression is: ; where is the i th node, is the second-level index of the i th node, , , correspond to the trained existing word embedding model, vector coding model, and image coding model respectively, , , respectively represent the natural language description, statistical features, and image corresponding to the i th node; represents the concatenation of the previous and next vectors; S2.2. Based on the second-layer index designed in step S2.1, construct the first-layer index layer, which is the layer where spatio-temporal partitioning is located, and obtain the expression: ; ; where represents a differentiable and permutation-invariant function, and respectively represent the first differentiable function and the second differentiable function, , , represent different nodes, represents the adjacent node to in the second layer, is a value close to 0, represents the attention coefficient between two nodes, represents a trainable parameter matrix, represents the natural exponential function, represents the index corresponding to the node type in the first-layer index layer; Furthermore, represents one of summation, mean, minimum, and maximum; S3. Construct a spatio-temporal hierarchical retrieval method for large model multi-modal prompt words. First, perform large model multi-modal prompt word embedding, and then perform spatio-temporal hierarchical retrieval based on similarity measurement using the hierarchical index of the spatio-temporal heterogeneous graph obtained in step S2 to obtain the spatio-temporal hierarchical retrieval result; Furthermore, the specific implementation method of step S3 includes the following steps: S3.1. Perform large model multi-modal prompt word embedding, and the expression for the embedding process is: ; where , , respectively represent the natural language description, statistical features, and image corresponding to the u i th segment of the prompt word, represents the query index constructed for the u i th segment of the prompt word; If the u iIf the natural language description, statistical features, and information in the image are missing in the segment prompt, the missing information is filled with null values using the data calculated by the following corresponding calculations. The expression is: ; ; ; Among them, is the mean of the natural language description, is the mean of the statistical features, is the mean of the image; S3.2. Construct a spatio-temporal hierarchical retrieval method based on similarity measurement: S3.2.1. First, perform spatio-temporal scale retrieval. For the query index u i constructed by the segment prompt , calculate its node similarity with all spatio-temporal partitions . The expression is: ; Among them, is the 2-norm; Then, sort all , and select the node corresponding to the spatio-temporal partition with the largest to calculate its similarity with all nodes under this partition. The expression is: ; Then, sort all , and select the largest indices. The corresponding spatio-temporal features are used as the retrieval result. The expression is: .
[0024] Based on the index (i.e., the embedding vector) of the large model prompt and the hierarchical index in the spatio-temporal heterogeneous graph, first calculate the similarity between the two, and then screen the index content with higher similarity from the spatio-temporal heterogeneous graph to complete the retrieval.
[0025] S4. Based on the spatio-temporal hierarchical retrieval result obtained in step S3, reconstruct the initial prompt and retrieval result of the large model. The reconstructed prompt is divided into blocks in the form of HTML tags to obtain a new prompt of the large model with retrieval content.
[0026] Furthermore, the specific implementation method of step S4 includes the following steps: S4.1. Based on the retrieval results obtained in step S3, reconstruct the initial prompt word and retrieval results of the large model. The reconstruction expression is: ; Among them, is the new prompt word after reconstruction, is the assignment symbol, represents the overall description, which adds a reference method for the retrieval results on the basis of the initial prompt word, is the current task description, and its content is the initial prompt word, is retrieval results, is the output description, which includes the requirements for the output format; For one retrieval result , the corresponding content is as Figure 2 shown, which respectively represent a certain spatial region at a scale, and the regional description of a certain time period at a scale , regional vector and regional image .
[0027] S4.2. Block the prompt word reconstructed in step S4.1 in the form of HTML tags. Among them, a reference prompt for the indexed content is newly added to the overall description part, the current task description is the original prompt word, and the retrieval content is listed in sequence from the 1st to the th. Each retrieval content includes the regional description , regional vector , and regional image corresponding to the retrieved spatial entity. Among them, the regional description and regional vector are given in character form, and the regional image is converted into a Base64 character and input through the image interface of the multimodal large model.
[0028] Furthermore, the order of all inputs is as Figure 3 shown.
[0029] The technical key points and points to be protected in the present invention are: The multi-scale modeling technology of urban spatio-temporal data in the present invention, that is, the multi-scale modeling technology of urban spatio-temporal data based on heterogeneous graphs proposed in step 1.
[0030] The urban spatio-temporal data index construction technology in the present invention, that is, the spatio-temporal heterogeneous graph hierarchical index construction technology based on graph embedding proposed in step 2.
[0031] The related content retrieval technology of multi-modal prompt words in the present invention, that is, the spatio-temporal hierarchical retrieval technology for multimodal large models proposed in step 3.
[0032] The multi-modal prompt word construction technology for spatio-temporal data in the present invention, that is, the multi-modal prompt word reconstruction technology proposed in step 4 based on the spatio-temporal heterogeneous graph retrieval results.
[0033] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0034] Although the present application has been described above with reference to specific embodiments, various improvements can be made to it and components therein can be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the features in the specific embodiments disclosed in the present application can be combined with each other in any way, and the fact that the combinations are not exhaustively described in this specification is only for the sake of saving space and resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A method for embedding urban spatio-temporal data for multi-modal large models, characterized in that, The steps include: S1. Based on the collected multi-scale spatiotemporal data of the city, construct the spatiotemporal heterogeneous graph of the urban spatiotemporal data and the spatiotemporal heterogeneous graph feature set of the urban spatiotemporal data; S2. Design a hierarchical index for the spatiotemporal heterogeneous graph of the urban spatiotemporal data obtained in step S1, using the spatial and temporal dimension division as the first-level index, and assigning an index to each node entity in all node entities as the second-level index; S3. Construct a spatiotemporal hierarchical retrieval method for large-model multimodal prompt words. First, embed the large-model multimodal prompt words, and then perform spatiotemporal hierarchical retrieval based on similarity measurement based on the hierarchical index of the spatiotemporal heterogeneous graph obtained in step S2 to obtain spatiotemporal hierarchical retrieval results; S4. Based on the spatiotemporal hierarchical retrieval results obtained in step S3, the initial prompt words and retrieval results of the large model are reconstructed, and the reconstructed prompt words are divided into blocks in the form of HTML tags to obtain new prompt words with retrieval content of the large model.
2. The method for embedding urban spatio-temporal data for multi-modal large models according to claim 1, characterized in that, The specific implementation method of step S1 includes the following steps: S1.
1. Constructing a spatiotemporal heterogeneous graph of urban spatiotemporal data: Construct a spatio-temporal heterogeneous graph of urban spatio-temporal data based on the collected multi-scale spatio-temporal data of the city , the expression is: ; Among them, represents a set of nodes composed of spatial entities, , is the N th node, represents the edges formed in the set of nodes, , is the M th edge, and each node and each edge correspond to different types, is the set of features, including the features corresponding to all nodes, is the set of node types, is the set of edge types; S1.
2. Associate the nodes constructed in step S1.1: For spatial entities, different types of spatial partitions are defined in terms of spatial dimensions and different types of temporal partitions are defined in terms of temporal dimensions. In a spatio-temporal heterogeneous graph, the expression for obtaining the set of node types is as follows: ; Define that the sizes of the scales in spatial partitioning and temporal partitioning are arranged from large to small. Define to represent the node types within the spatial and temporal regions represented by , , , and to represent the number of nodes under the node type represented by Based on different nodes, only one association is established, and the expression for the set of edge types is: ; S1.
3. Feature set for constructing spatiotemporal heterogeneous graphs of urban spatiotemporal data: Set node type The feature space of includes the natural language description space , the vector space for describing statistical features , and the vector space for describing images represents the set of all possible strings and represents the feature dimension of node type which respectively represent the pixel size and the number of channels of the image; Then the node type feature set The expression is as follows: ; Among them, represents the natural language description within the spatial and temporal regions denoted by ; represents the set of features included in the nodes within the spatial and temporal regions denoted by ; represents the images included in the nodes within the spatial and temporal regions denoted by ; ; ; ; Among them, represents the composition of the string, represents all possible character sets, represents starting from the statistical characteristics included in the nodes within the spatial and temporal regions represented by .
3. The method for embedding urban spatio-temporal data for multi-modal large models according to claim 2, characterized in that, The specific implementation method of step S2 includes the following steps: S2.
1. Design the second-level index layer as the layer where the node is located. Build the second-level index for any node. The expression is: ; Among them, is the i th node, is the second-layer index of the i th node, , , correspond to the pre-trained existing word embedding model, vector encoding model, and image encoding model respectively, , , represent the natural language description, statistical features, and image corresponding to the i th node respectively, represents the concatenation of the front and back vectors; S2.
2. Based on the second-level index designed in step S2.1, the first-level index layer is constructed, which is the layer where the time-space division is located, and the expression is obtained: ; ; Among them, represents a differentiable and permutation-invariant function, and represent the first differentiable function and the second differentiable function respectively, , , represent different nodes, represents the adjacent nodes to in the second layer, is a value close to 0, represents the attention coefficient between two nodes, represents a trainable parameter matrix, represents the natural exponential function, represents the index corresponding to the node type in the first-layer index layer.
4. The method for embedding urban spatio-temporal data for multi-modal large models according to claim 3, characterized in that, The specific implementation method of step S3 includes the following steps: S3.
1. Perform large-model multimodal prompt word embedding. The expression of the embedding process is: ; Among them, , , respectively represent the natural language description, statistical features, and image corresponding to the u i th segment of prompt words, represents the query index constructed for the u i th segment of prompt words; If the natural language description, statistical features, and information in the image are missing in the u i prompt words of the segment, the missing information is filled with null values using the data calculated by the following corresponding calculations. The expression is: ; ; ; Among them, is the mean value described by natural language, is the mean value of statistical features, is the mean value of the image; S3.
2. Constructing a spatiotemporal hierarchical retrieval method based on similarity measurement: S3.2.
1. First, perform a spatio-temporal scale retrieval. For the query index constructed by the prompt words in the u i th segment , calculate its node similarity with all spatio-temporal partitions . The expression is: ; Among them, is the 2-norm; Then, for all perform sorting and select the node corresponding to the largest spatio-temporal division and calculate its similarity with all nodes under this division . The expression is: ; Then, for all of the perform sorting and select the spatiotemporal features corresponding to the largest indices as the retrieval result. The expression is: 。 5. The method for embedding urban spatio-temporal data for a multimodal large model according to claim 4, wherein, The specific implementation method of step S4 includes the following steps: S4.
1. Based on the search results obtained in step S3, the initial prompt words and search results of the large model are reconstructed. The reconstructed expression is: ; Among them, is the newly reconstructed prompt, is the assignment symbol, represents the overall description, which adds a reference method for the retrieval results based on the initial prompt, is the current task description, and its content is the initial prompt, is the number of retrieval results, is the output description, which includes the requirements for the output format; S4.
2. Chunk the prompt words reconstructed in step S4.1 in the form of HTML tags. A reference hint for the indexed content is added to the overall description part. The current task description is the original prompt words, and the retrieved content is listed sequentially from the 1st to the th item. Each retrieved content contains the regional description corresponding to the retrieved spatial entity , regional vector , and regional image . The regional description and regional vector are given in character form, and the regional image is converted into a Base64 character and input through the image interface of the multimodal large model.
Citation Information
Patent Citations
Multimodal Large Model Evaluation Method and Device for Urban Governance Based on Multiple Layers of Text
CN119399006B
Distributed intelligent retrieval method and system for spatio-temporal data
CN118838937A
Device, Method and Program for Managing Area Information
US20080133484A1
Cited By
Retrieval enhancement generation method based on multi-level space-time grid feature alignment
CN122153021A