Urban space general representation learning method and device based on multi-modal spatio-temporal data fusion, terminal and storage medium
By generating a multi-view fusion representation matrix of urban spatial units and performing global aggregation, the applicability of multimodal spatiotemporal data in urban analysis tasks is solved, achieving multi-task adaptability and knowledge sharing, and supporting refined urban governance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN UNIV
- Filing Date
- 2026-04-08
- Publication Date
- 2026-05-05
AI Technical Summary
In existing tasks for predicting urban population distribution, urban travel flow, and urban environmental quality, the problem of data distribution differences among various modal spatiotemporal data has not been effectively resolved, resulting in insufficient applicability of general urban representations.
By acquiring multimodal spatiotemporal data of each spatial unit in the target city, setting corresponding views, generating single-view representations, and forming a general urban representation through multi-view fusion and global aggregation, it is applicable to a variety of prediction tasks.
It achieves a unified representation of spatiotemporal data of multiple modalities, improves the applicability of general urban representation, and enables it to be applied to diverse urban analysis tasks, with good generalization ability and task adaptability.
Smart Images

Figure CN121980249A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of urban data representation technology. More specifically, this application relates to a general urban spatial representation learning method, device, terminal, and storage medium based on multimodal spatiotemporal data fusion. Background Technology
[0002] Existing urban population distribution prediction, urban traffic flow prediction, and urban environmental quality prediction tasks all require multimodal spatiotemporal data of the city. This multimodal spatiotemporal data typically includes remote sensing imagery, street view imagery, points of interest (i.e., the geographic locations of the city on a map), and vehicle movement trajectories. The multimodal spatiotemporal data used in different prediction tasks often exhibit differences in data distribution. Currently, there is no method to transform multimodal spatiotemporal data into a universal urban representation applicable to multiple prediction tasks. Therefore, existing technologies require further improvement and enhancement. Summary of the Invention
[0003] The purpose of this application is to provide a method, apparatus, terminal, and storage medium for learning a general urban spatial representation based on multimodal spatiotemporal data fusion. This method improves the applicability of the general urban representation, making it suitable for diverse urban analysis tasks. This application is mainly achieved through the following technical solutions: A first aspect of this application provides a general urban spatial representation learning method based on multimodal spatiotemporal data fusion, comprising: Acquire multimodal spatiotemporal data for each spatial unit in the target city, and set a corresponding view for each modal spatiotemporal data. Based on the spatiotemporal data of each spatial unit for each modality, generate a single-view representation of each spatial unit under the view corresponding to each modality spatiotemporal data; Based on the single-view representation of each spatial unit under the views corresponding to all modal spatiotemporal data, a multi-view fusion representation corresponding to each spatial unit is generated, and the multi-view fusion representations corresponding to all spatial units are combined to form a multi-view fusion representation matrix. The multi-view fusion representation matrix is subjected to global aggregation processing to obtain a general urban representation of the target city.
[0004] According to one embodiment of this application, the multimodal spatiotemporal data of each spatial unit includes the target visual features, target semantic representation, and target trajectory representation corresponding to each spatial unit in the target city.
[0005] According to one embodiment of this application, the step of generating a single-view representation of each spatial unit under the view corresponding to each modal spatiotemporal data based on each modal spatiotemporal data of each spatial unit includes: Construct a set of relationships based on all spatial units in the target city; Within the set of relationships, determine the set of associated units for each spatial unit under the view corresponding to each modal spatiotemporal data; The feature aggregation function is used to fuse each spatial unit and the set of associated units of each spatial unit under the view corresponding to each modal spatiotemporal data to obtain the single-view representation of each spatial unit under the view corresponding to each modal spatiotemporal data.
[0006] According to one embodiment of this application, the step of generating a multi-view fusion representation for each spatial unit based on the single-view representation of each spatial unit under the views corresponding to all modal spatiotemporal data includes: An attention mechanism is used to calculate the fusion weight of the single-view representation of each spatial unit under the view corresponding to each modality of spatiotemporal data, so as to obtain the target fusion weight of each spatial unit under the view corresponding to each modality of spatiotemporal data. The first preset algorithm is used to perform weighted fusion calculation on the target fusion weight of each spatial unit under the views corresponding to all modal spatiotemporal data and the single-view representation of each spatial unit under the views corresponding to all modal spatiotemporal data, so as to obtain the multi-view fusion representation of each spatial unit.
[0007] According to one embodiment of this application, after the step of performing global aggregation processing on the multi-view fusion representation matrix to obtain the general urban representation of the target city, the urban spatial general representation learning method based on multimodal spatiotemporal data fusion further includes: Transform the general urban representation into a task-specific representation for the first objective task; The second preset algorithm is used to calculate and process the task-specific representation and at least one second target task to obtain the representation matrix of the first target task fused with all the second target tasks; The first objective task and all the second objective tasks are selected from one of the following tasks: estimating the economic level of the target city, predicting population distribution, predicting travel flow, assessing urban health, or predicting environmental quality. The first objective task and all the second objective tasks are different tasks.
[0008] According to one embodiment of this application, the step of transforming the general urban representation into a task-specific representation of the first target task includes: Obtain the task prompt embedding of the first target task; The task prompt embedding is used to guide the general representation of the city to focus on the task features of the first target task, thereby obtaining the task-specific representation of the first target task.
[0009] According to one embodiment of this application, the step of obtaining the task prompt embedding of the first target task includes: Obtain the task requirement information of the first target task; The task requirement information is encoded into a machine-readable task description; The task description is encoded into a task prompt embedding for the first target task using a learnable task prompt encoder.
[0010] A second aspect of this application provides a general urban spatial representation learning device based on multimodal spatiotemporal data fusion, comprising: The multimodal spatiotemporal data acquisition module is used to acquire multimodal spatiotemporal data of each spatial unit in the target city and set a corresponding view for each type of spatiotemporal data. The single-view in-spatial representation generation module is used to generate a single-view in-spatial representation of each spatial unit under the view corresponding to each modal spatiotemporal data, based on each modal spatiotemporal data of each spatial unit. The multi-view fusion representation matrix construction module is used to generate a multi-view fusion representation for each spatial unit based on the single-view representation of each spatial unit under the views corresponding to all modal spatiotemporal data, and to construct a multi-view fusion representation matrix by combining the multi-view fusion representations of all spatial units. The city general representation acquisition module is used to perform global aggregation processing on the multi-view fusion representation matrix to obtain the city general representation of the target city.
[0011] A third aspect of this application provides a terminal device, including a processor and a memory, the memory being used to store a computer program, and the processor being used to call and run the computer program stored in the memory to execute the steps of the urban spatial general representation learning method based on multimodal spatiotemporal data fusion provided in the first aspect of this application.
[0012] A fourth aspect of this application provides a computer-readable storage medium for storing a computer program that causes a computer to execute the steps of the urban spatial general representation learning method based on multimodal spatiotemporal data fusion provided in the first aspect of this application.
[0013] The beneficial effects of the embodiments of this application include: This application embodiment acquires multimodal spatiotemporal data for each spatial unit in a target city and sets a corresponding view for each modal spatiotemporal data. Based on each modal spatiotemporal data of each spatial unit, a single-view in-place representation of each spatial unit under the view corresponding to each modal spatiotemporal data is generated. Based on the single-view in-place representation of each spatial unit under the views corresponding to all modal spatiotemporal data, a multi-view fusion representation corresponding to each spatial unit is generated, and the multi-view fusion representations corresponding to all spatial units are combined to form a multi-view fusion representation matrix. The multi-view fusion representation matrix is then globally aggregated to obtain a general urban representation of the target city. Compared with the prior art, this application embodiment can transform multimodal spatiotemporal data into a general urban representation applicable to various prediction tasks. Therefore, this application embodiment can improve the applicability of the general urban representation, making it suitable for diverse urban analysis tasks. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of this application or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 The flowcharts for some embodiments of the general urban spatial representation learning method based on multimodal spatiotemporal data fusion of this application are shown below. Figure 2 The flowcharts are shown in some other embodiments of the general urban spatial representation learning method based on multimodal spatiotemporal data fusion of this application; Figure 3 This is a block diagram illustrating the principle of the urban spatial general representation learning device based on multimodal spatiotemporal data fusion in some embodiments of this application. Figure 4 This is a schematic block diagram of the terminal device of this application in some embodiments. Detailed Implementation
[0016] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the specific embodiments of this application are described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.
[0017] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0018] The terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0019] The terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units that are expressly listed, but may include other steps or units that are not expressly listed or that are inherent to such process, method, product, or apparatus.
[0020] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The term "and / or" as used in this application includes any and all combinations of one or more of the associated listed items.
[0021] The specific embodiments of this application will be further described below with reference to the accompanying drawings.
[0022] refer to Figure 1 The diagram shown is a flowchart of a general urban spatial representation learning method based on multimodal spatiotemporal data fusion, provided in the first aspect of an embodiment of this application. Figure 1 The urban spatial general representation learning method based on multimodal spatiotemporal data fusion includes the following steps S1, S2, S3 and S4.
[0023] S1. Obtain multimodal spatiotemporal data for each spatial unit in the target city, and set a corresponding view for each modal spatiotemporal data.
[0024] This application's embodiments divide the target city into multiple spatial regions, and each spatial region is considered a spatial unit. A spatial unit refers to a spatial extent, such as a building plot, a grid size of 500 meters or 1 kilometer, an administrative street unit, or an administrative district. The spatial scale of a spatial unit can vary, and this document does not impose any limitations on it. The spatial extent of the spatial unit can be extended to buildings, plots, blocks, and administrative districts, etc.
[0025] The aforementioned multimodal spatiotemporal data is multimodal spatiotemporal data.
[0026] The multimodal spatiotemporal data of each spatial unit includes the target visual features, target semantic representation, and target trajectory representation corresponding to each spatial unit in the target city.
[0027] In other implementations, the multimodal spatiotemporal data of each spatial unit may also include road networks, buildings, or social media text, and corresponding encoders may be designed for the new modalities.
[0028] The use of multimodal spatiotemporal data enables a unified representation of multi-source heterogeneous urban data.
[0029] Furthermore, the steps for acquiring the target visual features corresponding to each spatial unit (see reference) Figure 2 The "visual modality encoding" step includes: acquiring the target image corresponding to each spatial unit; applying a pre-trained visual encoder with an autoencoder to perform feature extraction processing on the target image corresponding to each spatial unit to obtain the initial features corresponding to each spatial unit; and overlaying geographic coordinates onto the initial features corresponding to each spatial unit to obtain the original visual features corresponding to each spatial unit (see reference). Figure 2 (The "geographic coordinate embedding" step in the process) compresses the original visual features corresponding to each spatial unit into a latent representation corresponding to each spatial unit; and applies the decoder of the autoencoder to reconstruct the latent representation corresponding to each spatial unit into the target visual features corresponding to each spatial unit.
[0030] The target image is a remote sensing image block or a street view image sampled along a road network. In other embodiments, the target image may be other images, which can be set by those skilled in the art according to actual needs.
[0031] The target visual features corresponding to each spatial unit are visual features containing spatial semantics.
[0032] Furthermore, the calculation formula for the step "applying a pre-trained visual encoder with an autoencoder to perform feature extraction processing on the target image corresponding to each spatial unit to obtain the initial features corresponding to each spatial unit; superimposing geographic coordinates on the initial features corresponding to each spatial unit to obtain the original visual features corresponding to each spatial unit" is as follows: ; in, It is the first The original visual features corresponding to each spatial unit; It is pooling and linear projection; It is the pre-trained visual encoder; It is the first The target image corresponding to each spatial unit; It is a location encoder based on geographic coordinates, which can also be understood as geographic coordinate embedding; It is the first The geographic coordinates corresponding to each spatial unit; yes A set of real numbers of dimension 1. For example, The value can be 768.
[0033] Furthermore, the pre-trained visual encoder is trained using a training dataset. The data format of each element in the training dataset is identical to that of the target image. During training, the gradient is backpropagated through the decoder to the initial visual encoder of the autoencoder, updating the parameters of the initial visual encoder to form the pre-trained visual encoder. This allows the visual features extracted in this embodiment to retain visual information while being more similar to the visual features of adjacent units.
[0034] Furthermore, during training, the parameters of the initial visual encoder are updated based on the loss function.
[0035] The loss function The calculation formula is: ; in, It is the total number of all spatial units; It is the first The visual features of the target corresponding to each spatial unit; It is a geographically smoothed weight; It is the first The original visual features corresponding to the first spatial unit, the first The spatial unit is the... Adjacent spatial units of a spatial unit; It is a spatial adjacency edge set; It is an L2 norm.
[0036] Furthermore, the steps for obtaining the target semantic representation corresponding to each spatial unit (see reference) Figure 2 The "POI (Point of Interest) modality encoding" step includes: obtaining text descriptions and category labels for multiple points of interest in each spatial unit; converting the text descriptions and category labels for each point of interest in each spatial unit into individual feature vectors corresponding to each point of interest; estimating the spatial coordinates of all points of interest in each spatial unit using a Gaussian kernel-based spatial density estimation function to obtain the spatial distribution pattern of all points of interest in each spatial unit; and using a multilayer perceptron to perform nonlinear mapping processing on the individual feature vectors corresponding to all points of interest in each spatial unit and the spatial distribution pattern of all points of interest in each spatial unit to obtain the target semantic representation corresponding to each spatial unit.
[0037] The use of the Gaussian kernel-based spatial density estimation function enables the joint embedding of text, category, and space.
[0038] The target semantic representation is a semantic representation that includes urban functions.
[0039] Furthermore, the formula for calculating the step of converting the text description and category label of each point of interest in each spatial unit into the individual feature vector corresponding to each point of interest is as follows: ; in, It is the first In the _ spatial unit, the _ ... Individual feature vectors corresponding to points of interest; It is an embedding layer; It is the first In the _ spatial unit, the _ ... Textual descriptions of points of interest; It is the first In the _ spatial unit, the _ ... Category labels for points of interest; It is a location encoder for points of interest, which can also be understood as point of interest location embedding; It is the first In the _ spatial unit, the _ ... Spatial coordinates of a point of interest; , It is the first The total number of points of interest within each spatial unit.
[0040] Furthermore, the calculation formula for estimating the spatial coordinates of all points of interest in each spatial cell using a Gaussian kernel-based spatial density estimation function to obtain the spatial distribution pattern of all points of interest in each spatial cell is as follows: ; in, It is the first Spatial distribution pattern of all points of interest in a spatial unit; It is the spatial density estimation function based on Gaussian kernel; yes The set of real numbers of dimension ; The value can be set by those skilled in the art according to actual needs.
[0041] Furthermore, the calculation formula for obtaining the target semantic representation corresponding to each spatial unit by using a multilayer perceptron to perform nonlinear mapping processing on the individual feature vectors corresponding to all points of interest in each spatial unit and the spatial distribution pattern of all points of interest in each spatial unit is as follows: ; in, It is the first The target semantic representation corresponding to each spatial unit; It is the multilayer sensor; It is a sequence encoder; yes The set of real numbers of dimension ; The value can be set by those skilled in the art according to actual needs; It is a splicing symbol.
[0042] Furthermore, the steps for obtaining the target trajectory representation for each spatial unit (refer to...) Figure 2 The "trajectory modal encoding" step includes: acquiring the inflow and outflow features of each spatial unit; performing multilayer perceptron encoding on the inflow and outflow features of each spatial unit to obtain the flow embedding of each spatial unit; acquiring the spatiotemporal map of each spatial unit, and using a spatiotemporal map encoder to extract the spatiotemporal distribution pattern representation of each spatial unit to obtain the target spatiotemporal distribution pattern representation corresponding to each spatial unit; acquiring the periodic sequence of each spatial unit, and using a sequence encoder to extract the periodic fluctuation representation of each spatial unit to obtain the target periodic fluctuation representation corresponding to each spatial unit; and merging the flow embedding, the target spatiotemporal distribution pattern representation, and the target periodic fluctuation representation of each spatial unit to obtain the target trajectory representation corresponding to each spatial unit.
[0043] The inflow characteristics of each space unit refer to the mobile phone location information, vehicle GNSS (Global Navigation Satellite System) trajectory information, and public transportation card swipe information entering each space unit.
[0044] The outflow characteristics of each spatial unit refer to the mobile phone location information, vehicle GNSS trajectory information, and public transportation card swipe information leaving each spatial unit.
[0045] The periodic sequence of each spatial unit can refer to the inflow or outflow characteristics of each spatial unit within a predetermined period. The specific time of the predetermined period can be set by those skilled in the art according to actual needs.
[0046] The target trajectory representation can be understood as a unified representation of trajectory modes.
[0047] Furthermore, the inflow and outflow features of each spatial unit are processed using multilayer perceptron encoding to obtain the flow embedding for each spatial unit. The calculation formula for this step is as follows: ; in, It is the first Flow embedding of each spatial unit; It is the first Inflow characteristics of each spatial unit; It is the first Outflow characteristics of each spatial unit; yes The set of real numbers of dimension ; The value can be set by those skilled in the art according to actual needs.
[0048] Furthermore, the spatiotemporal distribution pattern representation extraction process for the spatiotemporal graph of each spatial unit is performed using a spatiotemporal graph encoder. The calculation formula for obtaining the target spatiotemporal distribution pattern representation corresponding to each spatial unit is as follows: ; in, It is the first Characterization of the spatiotemporal distribution pattern of the target corresponding to each spatial unit; It is the spatiotemporal graph encoder; It is the first Spatiotemporal diagram of a spatial unit.
[0049] Furthermore, the calculation formula for the step of extracting the periodic fluctuation characterization of the periodic sequence of each spatial unit using a sequence encoder to obtain the target periodic fluctuation characterization corresponding to each spatial unit is as follows: ; in, It is the first The target periodic fluctuation characterization corresponding to each spatial unit; It is the sequence encoder; It is the first A periodic sequence of spatial units.
[0050] Furthermore, the calculation formula for merging the flow embedding, the target spatiotemporal distribution pattern representation, and the target periodic fluctuation representation for each spatial unit to obtain the target trajectory representation for each spatial unit is as follows: ; in, It is the first The target trajectory representation corresponding to each spatial unit; It is a fusion function; It is a merge function; yes The set of real numbers of dimension ; The value can be set by those skilled in the art according to actual needs.
[0051] S2. Based on the spatiotemporal data of each spatial unit for each modality, generate a single-view in-place representation of each spatial unit under the view corresponding to each modality spatiotemporal data (refer to...). Figure 2 The "Single View Learning" step in the process.
[0052] Further, step S2 includes: constructing a relationship set based on all spatial units in the target city; determining the set of associated units for each spatial unit under the view corresponding to each modal spatiotemporal data in the relationship set; and using a feature aggregation function to fuse each spatial unit and the set of associated units for each spatial unit under the view corresponding to each modal spatiotemporal data to obtain the single-view representation of each spatial unit under the view corresponding to each modal spatiotemporal data.
[0053] Furthermore, the step of constructing a relationship set based on all spatial units in the target city includes: connecting spatial units according to spatial adjacency, road connection strength, and / or mobility interaction strength to obtain the relationship set.
[0054] In other embodiments, spatial units can be connected based on temporal co-occurrence to obtain the set of relationships.
[0055] The set of relations can be expressed as .
[0056] The relationships in the set of relationships represent the connections between urban spaces in dimensions such as geometry, road network, or human activities.
[0057] Furthermore, the first The spatial unit in the first The set of associated units under the view corresponding to the spatiotemporal data of a certain modality can be expressed as: .
[0058] No. The view corresponding to the spatiotemporal data of a modality can also be expressed as .
[0059] Furthermore, the calculation formula for the step of fusing each spatial unit and the set of associated units of each spatial unit under the view corresponding to each modal spatiotemporal data using a feature aggregation function to obtain the single-view representation of each spatial unit under the view corresponding to each modal spatiotemporal data is as follows: ; in, It is the first The spatial unit in the first In-view representation of spatiotemporal data of various modalities under a single view; It is the first Aggregation functions under the view corresponding to the spatiotemporal data of various modalities; yes The set of real numbers of dimension ; The value can be set by those skilled in the art according to actual needs; It is the first The spatial unit in the first Input features under the view corresponding to the spatiotemporal data of each modality; It is the first The spatial unit in the first Input features under the view corresponding to the spatiotemporal data of each modality; It is the set difference symbol.
[0060] The meaning of the input features mentioned in the embodiments of this application is the same as the meaning of the multimodal spatiotemporal data of each spatial unit. That is, the input features include the target visual features, target semantic representation and target trajectory representation corresponding to each spatial unit in the target city.
[0061] S3. Based on the single-view representation of each spatial unit under the views corresponding to all modal spatiotemporal data, generate a multi-view fusion representation for each spatial unit, and construct a multi-view fusion representation matrix by combining the multi-view fusion representations of all spatial units (refer to [reference]). Figure 2 The "Cross-view merging" step in the document.
[0062] Furthermore, the step of generating a multi-view fusion representation for each spatial unit based on its single-view in-view representation under views corresponding to all modal spatiotemporal data includes: using an attention mechanism to calculate the fusion weight of the single-view in-view representation of each spatial unit under views corresponding to each modal spatiotemporal data, to obtain the target fusion weight of each spatial unit under views corresponding to each modal spatiotemporal data; and using a first preset algorithm to perform weighted fusion calculation on the target fusion weight of each spatial unit under views corresponding to all modal spatiotemporal data and the single-view in-view representation of each spatial unit under views corresponding to all modal spatiotemporal data, to obtain the multi-view fusion representation for each spatial unit.
[0063] Furthermore, the fusion weight calculation process for the single-view in-view representation of each spatial unit under each modal spatiotemporal data view using an attention mechanism is as follows: ; in, It is the first The spatial unit in the first Target fusion weights under the view corresponding to the various modal spatiotemporal data; It is an exponential function; It is a learnable attention vector. , yes The set of real numbers of dimension , The value can be set by those skilled in the art according to actual needs; Represents transposition; It is the hyperbolic tangent activation function, used to introduce nonlinear transformations; It is the first Learnable transformation matrices under the views corresponding to various modal spatiotemporal data. , yes The set of real numbers of dimension ; It refers to the total number of views, or the total number of modal spatiotemporal data. It is the first Learnable transformation matrices under the views corresponding to various modal spatiotemporal data; It is the first The spatial unit in the first The intra-view representation of a modal spatiotemporal data under the corresponding view.
[0064] Furthermore, .
[0065] Furthermore, the first preset algorithm is used to perform weighted fusion calculation on the target fusion weight of each spatial unit under the views corresponding to all modal spatiotemporal data and the single-view in-view representation of each spatial unit under the views corresponding to all modal spatiotemporal data. The calculation formula for obtaining the multi-view fusion representation corresponding to each spatial unit is as follows: ; in, It is the first Multi-view fusion representation corresponding to each spatial unit; It is the first The fusion projection matrix of the view corresponding to the spatiotemporal data of the three modalities. , yes The set of real numbers of dimension , The value can be set by those skilled in the art according to actual needs; It is the bias vector; yes The set of real numbers of dimension , The value can be set by those skilled in the art according to actual needs. It can be a unified multi-view fusion representation dimension.
[0066] The weighted fusion calculation process adaptively allocates weights based on the correlation between the cell and the view, so that the final representation reflects the dominant role of different views in different cells.
[0067] Furthermore, the multi-view fusion representation matrix can be expressed as... , It is the multi-view fusion representation corresponding to the first spatial unit. It is the first Multi-view fusion representation corresponding to each spatial unit yes The set of real numbers of dimension .
[0068] S4. Perform global aggregation processing on the multi-view fusion representation matrix to obtain the general urban representation of the target city (refer to...). Figure 2 The "global-scale aggregation" step in the process.
[0069] The aforementioned general urban characterization can be applied to tasks such as estimating urban economic levels, predicting population distribution, predicting travel flow, assessing urban health, or predicting environmental quality. Estimating urban economic levels can involve GDP (Gross Domestic Product), tax revenue, or employment rate. Predicting population distribution can involve population density or age distribution. Predicting travel flow can involve real-time traffic flow or public transportation usage. Assessing urban health can involve morbidity assessment or medical resource distribution assessment. Predicting environmental quality can involve air quality index or pollutant concentration prediction.
[0070] Furthermore, the calculation formula for step S4 is as follows: ; ; in, It is a general urban representation of the target city; It is the universal representation vector of the first spatial unit; It is the first A universal representation vector for each spatial unit; yes The set of real numbers of dimension , It is a universal representation dimension; yes The Line, i.e., the first A universal representation vector for each spatial unit. ; It is a global aggregation function used to integrate all view information at a global scale; and The calculation formula can be referred to The calculation formula.
[0071] Furthermore, the global aggregation function can be an identity mapping, a linear transformation with normalization, or an attention mechanism with a global receptive field. In other embodiments, the global aggregation function can also be other functions, which can be set by those skilled in the art according to actual needs.
[0072] Through the above implementation methods, the embodiments of this application can transform multimodal spatiotemporal data into a general urban representation that can be applied to a variety of prediction tasks. Thus, the embodiments of this application can improve the applicability of the general urban representation and make it suitable for diverse urban analysis tasks.
[0073] In some implementations, after the step of globally aggregating the multi-view fusion representation matrix to obtain a general urban representation of the target city, the urban spatial general representation learning method based on multimodal spatiotemporal data fusion further includes: transforming the general urban representation into a task-specific representation of a first target task; and using a second preset algorithm to calculate and process the task-specific representation and at least one second target task to obtain a representation matrix of the first target task fused with all second target tasks (see reference). Figure 2 (The "task fusion representation" step in the text); the first target task and all second target tasks are selected from one of the following tasks: estimating the economic level of the target city, predicting population distribution, predicting travel flow, assessing urban health, or predicting environmental quality. The first target task and all second target tasks are different tasks.
[0074] The number of the at least one second target task can be set by those skilled in the art according to actual needs.
[0075] Furthermore, the step of transforming the general urban representation into a task-specific representation of the first target task includes: obtaining the task cue embedding of the first target task (see reference). Figure 2 The "cue embedding" step in the text); using the task cue embedding to guide the city's general representation to focus on the task features of the first target task, to obtain the task-specific representation of the first target task (refer to...). Figure 2 The "task-specific representation" step in the process.
[0076] Further, the step of obtaining the task prompt embedding of the first target task includes: obtaining task requirement information of the first target task; encoding the task requirement information into a machine-parseable task description; and using a learnable task prompt encoder to encode the task description into the task prompt embedding of the first target task.
[0077] The task requirement information includes estimates of the target city's economic level, population distribution predictions, travel flow predictions, urban health assessments, or environmental quality predictions.
[0078] Furthermore, the calculation formula for the step of encoding the task description into a task prompt embedding of the first target task using a learnable task prompt encoder is as follows: ; in, It is the first target task Task prompt embedding; It is the learnable task prompt encoder; This is the task description; yes The set of real numbers of dimension , This is a prompt dimension.
[0079] Furthermore, the calculation formula for the step of using the task prompt embedding to guide the city's general representation to focus on the task features of the first target task, and obtaining the task-specific representation of the first target task, is as follows: ; in, It is the first target task Task-specific representation; It is a representation-focusing operation based on prefix tuning or task-specific attention.
[0080] Furthermore, the calculation formula for the step of using a second preset algorithm to calculate and process the task-specific representation and at least one second target task to obtain the representation matrix of the first target task fused with all second target tasks is as follows: ; in, It is the representation matrix of the first objective task fused with all the second objective tasks; It is the first target task The independence parameter is a learnable parameter or hyperparameter. It is the first target task With the Normalized similarity weights for each secondary objective task; It is the first A task-specific representation of a second objective task. This can also be understood as traversing all tasks except the first target task. The weights used to control the representation of this task (i.e., the task-specific representation of the first target task) and the fusion of other task representations (i.e., the task-specific representations of all second target tasks).
[0081] Through the above process, the representation matrix integrates multi-view information, global structural information, and inter-task knowledge, possessing good generalization ability and task adaptability, and can be directly applied to various urban spatial analysis tasks.
[0082] This application's embodiments unify multimodal spatiotemporal data fusion, urban representation learning, and multi-task collaborative optimization, learning a universal urban representation with expressiveness, robustness, and task adaptability. These embodiments enable knowledge sharing across related tasks while suppressing negative transfer, providing technical support for refined and intelligent urban governance.
[0083] refer to Figure 3The diagram shown is a principle block diagram of a general urban spatial representation learning device based on multimodal spatiotemporal data fusion, provided in the second aspect of an embodiment of this application. Figure 3 The urban spatial general representation learning device 100 based on multimodal spatiotemporal data fusion includes: The multimodal spatiotemporal data acquisition module 101 is used to acquire multimodal spatiotemporal data of each spatial unit in the target city and set a corresponding view for each type of spatiotemporal data. The single-view in-spatial representation generation module 102 is used to generate a single-view in-spatial representation of each spatial unit under the view corresponding to each modal spatiotemporal data based on each modal spatiotemporal data of each spatial unit. The multi-view fusion representation matrix construction module 103 is used to generate a multi-view fusion representation for each spatial unit based on the single-view representation of each spatial unit under the views corresponding to all modal spatiotemporal data, and to construct a multi-view fusion representation matrix by combining the multi-view fusion representations of all spatial units. The city general representation acquisition module 104 is used to perform global aggregation processing on the multi-view fusion representation matrix to obtain the city general representation of the target city.
[0084] A third aspect of this application provides a terminal device, the schematic diagram of which is as follows: Figure 4 As shown, the terminal device includes a processor, memory, network interface, display screen, and temperature sensor connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface of the terminal device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a general urban spatial representation learning method based on multimodal spatiotemporal data fusion. The display screen can be a liquid crystal display (LCD) or an e-ink display. The temperature sensor is pre-installed inside the terminal device to detect the operating temperature of the internal components.
[0085] Those skilled in the art will understand that Figure 4 The schematic diagram shown is only a partial structural diagram related to the present invention and does not constitute a limitation on the terminal device to which the present invention is applied. The specific terminal device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0086] In some embodiments, this application provides a terminal device, which includes a processor and a memory for storing computer programs. The processor is used to call and run the computer programs stored in the memory to execute the steps of the urban spatial general representation learning method based on multimodal spatiotemporal data fusion provided in the first aspect of this application.
[0087] A fourth aspect of this application provides a computer-readable storage medium for storing a computer program that causes a computer to execute the steps of the urban spatial general representation learning method based on multimodal spatiotemporal data fusion provided in the first aspect of this application.
[0088] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0089] The technical features of the above embodiments can be combined without changing the basic principles of this application. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0090] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the patent protection scope of this application should be determined by the appended claims.
Claims
1. A general urban spatial representation learning method based on multimodal spatiotemporal data fusion, characterized in that, include: Acquire multimodal spatiotemporal data for each spatial unit in the target city, and set a corresponding view for each modal spatiotemporal data. Based on the spatiotemporal data of each spatial unit for each modality, generate a single-view representation of each spatial unit under the view corresponding to each modality spatiotemporal data; Based on the single-view representation of each spatial unit under the views corresponding to all modal spatiotemporal data, a multi-view fusion representation corresponding to each spatial unit is generated, and the multi-view fusion representations corresponding to all spatial units are combined to form a multi-view fusion representation matrix. The multi-view fusion representation matrix is subjected to global aggregation processing to obtain a general urban representation of the target city.
2. The urban spatial general representation learning method based on multimodal spatiotemporal data fusion according to claim 1, characterized in that, The multimodal spatiotemporal data of each spatial unit includes the target visual features, target semantic representation, and target trajectory representation corresponding to each spatial unit in the target city.
3. The urban spatial general representation learning method based on multimodal spatiotemporal data fusion according to claim 1, characterized in that, The steps for generating a single-view representation of each spatial unit based on each modal spatiotemporal data of each spatial unit include: Construct a set of relationships based on all spatial units in the target city; Within the set of relationships, determine the set of associated units for each spatial unit under the view corresponding to each modal spatiotemporal data; The feature aggregation function is used to fuse each spatial unit and the set of associated units of each spatial unit under the view corresponding to each modal spatiotemporal data to obtain the single-view representation of each spatial unit under the view corresponding to each modal spatiotemporal data.
4. The urban spatial general representation learning method based on multimodal spatiotemporal data fusion according to claim 1, characterized in that, The steps for generating a multi-view fusion representation for each spatial unit, based on the single-view representation of each spatial unit across all modal spatiotemporal data, include: An attention mechanism is used to calculate the fusion weight of the single-view representation of each spatial unit under the view corresponding to each modality of spatiotemporal data, so as to obtain the target fusion weight of each spatial unit under the view corresponding to each modality of spatiotemporal data. The first preset algorithm is used to perform weighted fusion calculation on the target fusion weight of each spatial unit under the views corresponding to all modal spatiotemporal data and the single-view representation of each spatial unit under the views corresponding to all modal spatiotemporal data, so as to obtain the multi-view fusion representation of each spatial unit.
5. The urban spatial general representation learning method based on multimodal spatiotemporal data fusion according to claim 1, characterized in that, After performing global aggregation processing on the multi-view fusion representation matrix to obtain the general urban representation of the target city, the urban spatial general representation learning method based on multimodal spatiotemporal data fusion further includes: Transform the general urban representation into a task-specific representation for the first objective task; The second preset algorithm is used to calculate and process the task-specific representation and at least one second target task to obtain the representation matrix of the first target task fused with all the second target tasks; The first objective task and all the second objective tasks are selected from one of the following tasks: estimating the economic level of the target city, predicting population distribution, predicting travel flow, assessing urban health, or predicting environmental quality. The first objective task and all the second objective tasks are different tasks.
6. The urban spatial general representation learning method based on multimodal spatiotemporal data fusion according to claim 5, characterized in that, The steps for transforming the general urban representation into a task-specific representation for the first objective task include: Obtain the task prompt embedding of the first target task; The task prompt embedding is used to guide the general representation of the city to focus on the task features of the first target task, thereby obtaining the task-specific representation of the first target task.
7. The urban spatial general representation learning method based on multimodal spatiotemporal data fusion according to claim 6, characterized in that, The steps for obtaining the task hint embedding of the first target task include: Obtain the task requirement information of the first target task; The task requirement information is encoded into a machine-readable task description; The task description is encoded into a task prompt embedding for the first target task using a learnable task prompt encoder.
8. A general urban spatial representation learning device based on multimodal spatiotemporal data fusion, characterized in that, include: The multimodal spatiotemporal data acquisition module is used to acquire multimodal spatiotemporal data of each spatial unit in the target city and set a corresponding view for each type of spatiotemporal data. The single-view in-spatial representation generation module is used to generate a single-view in-spatial representation of each spatial unit under the view corresponding to each modal spatiotemporal data, based on each modal spatiotemporal data of each spatial unit. The multi-view fusion representation matrix construction module is used to generate a multi-view fusion representation for each spatial unit based on the single-view representation of each spatial unit under the views corresponding to all modal spatiotemporal data, and to construct a multi-view fusion representation matrix by combining the multi-view fusion representations of all spatial units. The city general representation acquisition module is used to perform global aggregation processing on the multi-view fusion representation matrix to obtain the city general representation of the target city.
9. A terminal device, characterized in that, include: A processor and a memory, the memory for storing a computer program, the processor for calling and running the computer program stored in the memory, and performing the steps of the urban spatial general representation learning method based on multimodal spatiotemporal data fusion as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store a computer program that causes a computer to perform the steps of the urban spatial general representation learning method based on multimodal spatiotemporal data fusion as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-modal information enhancement recommendation method based on knowledge graph
CN118296226A
Parallel fusion method and system for urban stock space multi-modal data, terminal and storage medium
CN120012026A
Flood monitoring system and method for mountain hydrological environment
CN121093263A
Cited By
A multi-view urban area embedding method based on spatial function consistency
CN122336073A
A multi-view urban area embedding method based on spatial function consistency
CN122336073B