Parallel fusion method and system for urban stock space multi-modal data, terminal and storage medium

By adopting multimodal spatiotemporal index structure and neural network model in the parallel fusion method of multimodal data in urban stock space, combined with deep learning and federated learning technology, the problem of low efficiency of multi-source data fusion processing in the existing technology is solved, efficient data fusion and complex urban analysis are achieved, and strong support for urban planning is provided.

CN120012026AActive Publication Date: 2025-05-16SHENZHEN UNIV

Patent Information

Application Number
CN202510473812.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-16
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The existing technology is inefficient when integrating and processing large-scale and diversified urban stock space multi-source data, which limits the ability of urban planners and managers to fully understand the current situation and development trends of the city, thereby affecting the quality and efficiency of decision-making.

Method used

A parallel fusion method of multimodal data in urban stock space is adopted. By acquiring and classifying multimodal data, designing multimodal spatiotemporal index structures and building neural network models, the distributed processing of data and the application of deep learning algorithms is realized, and data is initially fused with multi-head attention mechanisms, and global aggregation is carried out through federated learning methods to generate a target aggregation model.

Benefits of technology

It significantly improves the efficiency of the integration and processing of multi-source data in large-scale and diversified urban stock space, supports complex and detailed urban analysis and simulation, and provides strong technical support for urban planning and management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012026A_ABST
    Figure CN120012026A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data fusion, and discloses a parallel fusion method and system for urban stock space multi-modal data, a terminal and a storage medium, and the method comprises the steps: sequentially carrying out the classification, feature vector crude extraction and normalization of the urban stock space multi-modal data, and obtaining a full-vector stock space data set; the master node divides a full-vector stock space data set and a neural network model based on a multi-mode spatio-temporal index structure, and correspondingly distributes the data to each sub-node; feature vector fine extraction is carried out on each child node through a deep learning algorithm, and multi-modal data preliminary fusion is carried out through a multi-head attention mechanism; and the main node globally converges the local fusion feature vector output by each sub-node by adopting a federated learning method, and generates a target convergence model through iterative training. According to the method, the fusion processing efficiency of large-scale and diversified urban stock space multi-source data is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data fusion technology, and in particular to a parallel fusion method, system, terminal and computer-readable storage medium for urban stock space multimodal data. Background Art

[0002] With the continuous advancement of urbanization, the urban development model has entered a new stage of "stock improvement" from "incremental expansion". In order to improve the utilization efficiency of urban stock space, it is necessary to aggregate and integrate the stock space data and conduct in-depth analysis to identify inefficient land use and implement reasonable transformation and upgrading. This process involves the integration and analysis of multi-source, multi-type and multi-modal spatial data, including but not limited to vector data, remote sensing images, three-dimensional models, and spatiotemporal flow data reflecting dynamic changes such as traffic flow and population flow. At present, most urban data fusion methods have two main limitations: one is that they focus on a single data format, or although they can process multiple formats, they are all isomorphic data; the other is that most existing data fusion technologies are suitable for limited-scale data sets, but their computational efficiency and processing capabilities are insufficient when faced with multi-modal and massive urban spatial data. This limits the ability of urban planners and managers to fully understand the current situation and development trends of the city, thereby affecting the quality and efficiency of decision-making.

[0003] Therefore, the prior art still needs to be improved and developed. Summary of the invention

[0004] The main purpose of the present invention is to provide a parallel fusion method, system, terminal and computer-readable storage medium for multimodal data of urban stock space, aiming to solve the problem of low processing efficiency of traditional data fusion methods when fusing and processing large-scale and diversified multi-source data of urban stock space.

[0005] To achieve the above-mentioned object of the invention, the present invention provides a parallel fusion method for multimodal data of urban stock space, and the parallel fusion method for multimodal data of urban stock space comprises: Acquire multimodal data of urban stock space, and classify, roughly extract feature vectors and normalize the multimodal data of urban stock space in sequence to obtain a full-vector stock space data set; Designing a multimodal spatiotemporal index structure including a main node and a plurality of sub-nodes and constructing a neural network model, wherein the main node divides the stock spatial data set of the full vector and the neural network model based on the multimodal spatiotemporal index structure to obtain a plurality of sub-feature vectors and a plurality of sub-neural network models, and assigning the plurality of sub-feature vectors and the plurality of sub-neural network models to each of the sub-nodes; Based on the sub-feature vectors and the sub-neural network model, each of the sub-nodes performs fine feature vector extraction through a deep learning algorithm to obtain a fine feature vector, and performs preliminary multi-modal data fusion on the fine feature vector through a multi-head attention mechanism to obtain a local fused feature vector; The master node obtains the local fusion feature vector output by each of the child nodes, performs global aggregation based on the multiple local fusion feature vectors by using a federated learning method to obtain a global fusion feature vector, and generates a target aggregation model based on the global fusion feature vector through iterative training; Acquire the multimodal data of urban stock space to be fused, input the multimodal data of urban stock space to be fused into the sub-neural network model to obtain the target local fusion feature vector, and input the target local fusion feature vector into the target convergence model to obtain the data fusion result.

[0006] Optionally, the acquiring of multimodal data of urban stock space, and sequentially classifying, roughly extracting feature vectors and normalizing the multimodal data of urban stock space to obtain a full vector stock space data set specifically includes: Acquire urban stock space multimodal data, and divide the urban stock space multimodal data into raster data, graph data and time series data; Preprocessing the raster data, the graph data and the time series data respectively to obtain preprocessed raster data, preprocessed graph data and preprocessed time series data, wherein the preprocessing includes denoising and missing value filling; Respectively performing unified abstraction and modeling on the preprocessed raster data, the preprocessed graph data, and the preprocessed time series data to obtain a raster data model, a graph data model, and a time series data model; Using the raster data model, the graph data model and the time series data model, respectively, the preprocessed raster data, the preprocessed graph data and the preprocessed time series data are subjected to coarse feature vector extraction to obtain coarse feature vectors corresponding to the raster data, coarse feature vectors corresponding to the graph data and coarse feature vectors corresponding to the time series data; Normalizing the coarse feature vectors corresponding to the raster data, the graph data, and the time series data, respectively, to obtain a target coarse feature vector corresponding to the raster data, a target coarse feature vector corresponding to the graph data, and a target coarse feature vector corresponding to the time series data; The target coarse feature vector corresponding to the raster data, the target coarse feature vector corresponding to the graph data and the target coarse feature vector corresponding to the time series data are combined to obtain a stock spatial data set of the full vector.

[0007] Optionally, the design includes a multimodal spatiotemporal index structure of a main node and multiple sub-nodes and constructs a neural network model, wherein the main node divides the stock spatial data set of the full vector and the neural network model based on the multimodal spatiotemporal index structure to obtain multiple sub-feature vectors and multiple sub-neural network models, and assigns the multiple sub-feature vectors and the multiple sub-neural network models to each of the sub-nodes, specifically including: A multimodal spatiotemporal index structure is designed. A quadtree is used for recursive partitioning in the spatial dimension to divide the two-dimensional space into multiple sub-regions. A linear index is used in the time dimension to sort and retrieve data by timestamp. The quadtree includes a main node and multiple sub-nodes. The main node represents the two-dimensional space, and each sub-node represents each sub-region. Constructing a neural network model, the neural network model includes a raster data feature extraction module, a graph data feature extraction module, a time series data feature extraction module and a multimodal data preliminary fusion module; Based on the multimodal spatiotemporal index structure, the master node divides the stock spatial data set of the full vector into sub-regions and time ranges to obtain a plurality of sub-feature vectors, where the sub-feature vectors correspond to the sub-regions one by one; Based on the multimodal spatiotemporal index structure, the master node divides the neural network model into sub-regions to obtain a plurality of sub-neural network models, and the sub-neural network models correspond to the sub-regions one by one; Assigning the sub-feature vectors and the sub-neural network models corresponding to the respective sub-regions to the corresponding respective sub-nodes; Establish a task allocation strategy, dynamically adjust the task allocation ratio according to the computing power and memory capacity of each child node, and give priority to allocating tasks to child nodes whose distance to the data storage location is less than the set value. In addition, introduce a backup mechanism to allocate some tasks to multiple child nodes at the same time.

[0008] Optionally, the sub-feature vectors and the sub-neural network models corresponding to each sub-region are assigned to corresponding sub-nodes, and the sub-feature vectors corresponding to each sub-region include sub-feature vectors corresponding to raster data, sub-feature vectors corresponding to graph data and sub-feature vectors corresponding to time series data; the sub-neural network models corresponding to each sub-region include sub-raster data feature extraction modules, sub-graph data feature extraction modules, sub-time series data feature extraction modules and sub-multimodal data preliminary fusion modules.

[0009] Optionally, the fine feature vector includes a raster data feature vector, a graph data feature vector and a time series data feature vector; Based on the sub-feature vectors and the sub-neural network model, each of the sub-nodes extracts feature vectors through a deep learning algorithm to obtain a fine feature vector, and performs preliminary multi-modal data fusion on the fine feature vectors through a multi-head attention mechanism to obtain a local fusion feature vector, specifically including: At each child node, the sub-feature vector corresponding to the raster data is sequentially extracted through the 3D convolution layer, GeLU (Gaussian Error Linear Unit) activation function layer, maximum pooling layer, residual block, global pooling layer and the first fully connected layer of the sub-raster data feature extraction module to obtain a raster data feature vector. ; The sub-feature vector corresponding to the graph data is sequentially extracted through the graph convolution network, the first ReLU (Rectified Linear Unit) activation function layer, the multi-head attention network and the second fully connected layer of the sub-graph data feature extraction module to obtain the graph data feature vector. ; The sub-feature vector corresponding to the time series data is sequentially extracted through the multi-layer long short-term memory network, the third fully connected layer, the second ReLU activation function layer, the discard layer and the fourth fully connected layer of the sub-time series data feature extraction module to obtain the time series data feature vector ; Through the multi-head attention mechanism, the raster data feature vector , the graph data feature vector and the time series data feature vector Perform preliminary fusion of multimodal data to obtain local fusion feature vectors.

[0010] Optionally, the grid data feature vector is , the graph data feature vector and the time series data feature vector Perform preliminary fusion of multimodal data to obtain local fusion feature vectors, including: The raster data feature vector , the graph data feature vector and the time series data feature vector Through different linear transformation layers, the raster data feature vector is obtained. The corresponding first query vector Q1, the first key vector K1 and the first value vector V1, the graph data feature vector The corresponding second query vector Q2, the second key vector K2, the second value vector V2 and the time series data feature vector corresponding third query vector Q3, third key vector K3 and third value vector V3; The first similarity between the first query vector Q1 and the third key vector K3 is calculated in parallel by multiple attention heads , the second similarity between the second query vector Q2 and the second key vector K2 , the third similarity between the third query vector Q3 and the first key vector K1 : ; ; ; Where T represents transpose, represents the dimension of the first key vector, the second key vector, or the third key vector; According to the first similarity The second similarity and the third similarity , use the Softmax function to calculate the first attention weight corresponding to the raster data , the second attention weight corresponding to the graph data The third attention weight corresponding to the time series data : ; ; ; According to the first attention weight , the second attention weight and the third attention weight , respectively perform weighted summation on the third value vector V3, the second value vector V2 and the first value vector V1 to obtain a local fusion feature vector O: ; ; ; ; in, Represents the attention mechanism.

[0011] Optionally, the master node obtains a local fusion feature vector output by each of the child nodes, performs global aggregation based on the local fusion feature vector using a federated learning method to obtain a global fusion feature vector, and generates a target aggregation model based on the global fusion feature vector through iterative training, specifically including: The master node obtains the local fusion feature vector output by each of the child nodes, and based on the weight corresponding to each of the child nodes, uses the federated learning algorithm to perform weighted fusion on the local fusion feature vector output by each of the child nodes to obtain a global fusion feature vector : ; Among them, i represents the i-th child node, represents the local fusion feature vector output by the i-th child node, Represents the weight corresponding to the i-th child node; Among them, the weight corresponding to the i-th child node According to the number of samples of the i-th child node The calculation results are: ; in, Indicates the number of samples of the first child node, Indicates the number of samples of the second child node, represents the number of samples of the third child node, Indicates the number of samples of the fourth child node; The global fusion feature vector generated by the master node is used to update the aggregation weight parameter of the master node, and the updated aggregation weight parameter is fed back to each child node to obtain the target aggregation model; The aggregation weight parameter of the master node includes the weight corresponding to each of the child nodes.

[0012] To achieve the above-mentioned object of the invention, the present invention further provides a parallel fusion system for multimodal data of urban stock space, and the parallel fusion system for multimodal data of urban stock space comprises: Data preliminary processing module: used to obtain urban stock space multimodal data, and classify, roughly extract feature vectors and normalize the urban stock space multimodal data in turn to obtain a full vector stock space data set; Distributed task allocation module: used to design a multimodal spatiotemporal index structure including a main node and multiple sub-nodes and to construct a neural network model, wherein the main node divides the stock spatial data set of the full vector and the neural network model based on the multimodal spatiotemporal index structure to obtain multiple sub-feature vectors and multiple sub-neural network models, and allocates the multiple sub-feature vectors and the multiple sub-neural network models to each of the sub-nodes; Preliminary fusion module: used for extracting fine feature vectors of each sub-node through a deep learning algorithm based on the sub-feature vectors and the sub-neural network model to obtain fine feature vectors, and performing preliminary multi-modal data fusion on the fine feature vectors through a multi-head attention mechanism to obtain local fused feature vectors; Global aggregation module: used for the master node to obtain the local fusion feature vector output by each of the child nodes, and based on the multiple local fusion feature vectors, a federated learning method is used to perform global aggregation to obtain a global fusion feature vector, and based on the global fusion feature vector, a target aggregation model is generated through iterative training; Multimodal data fusion module: used to obtain the multimodal data of the urban stock space to be fused, input the multimodal data of the urban stock space to be fused into the sub-neural network model to obtain the target local fusion feature vector, and input the target local fusion feature vector into the target convergence model to obtain the data fusion result.

[0013] In order to achieve the above-mentioned purpose of the invention, the present invention also provides a terminal, which includes: a memory, a processor, and a parallel fusion program of urban stock space multimodal data stored in the memory and run on the processor, and when the parallel fusion program of urban stock space multimodal data is executed by the processor, the steps of the parallel fusion method of urban stock space multimodal data are implemented as described above.

[0014] In order to achieve the above-mentioned purpose of the invention, the present invention also provides a computer-readable storage medium, which stores a parallel fusion program of urban stock space multimodal data. When the parallel fusion program of urban stock space multimodal data is executed by a processor, the steps of the parallel fusion method of urban stock space multimodal data as described above are implemented.

[0015] In the present invention, multimodal data of urban stock space is obtained, and the multimodal data of urban stock space is classified, feature vectors are roughly extracted and normalized in turn to obtain a stock space data set of the full vector; a multimodal spatiotemporal index structure including a main node and multiple sub-nodes is designed and a neural network model is constructed, the main node divides the stock space data set of the full vector and the neural network model based on the multimodal spatiotemporal index structure to obtain multiple sub-feature vectors and multiple sub-neural network models, and the multiple sub-feature vectors and the multiple sub-neural network models are correspondingly assigned to each of the sub-nodes; based on the sub-feature vectors and the sub-neural network models, each of the sub-nodes performs feature vector rough extraction through a deep learning algorithm Take, obtain fine feature vectors, and perform preliminary multimodal data fusion on the fine feature vectors through a multi-head attention mechanism to obtain local fusion feature vectors; the master node obtains the local fusion feature vectors output by each of the sub-nodes, and based on multiple local fusion feature vectors, uses a federated learning method to perform global aggregation to obtain a global fusion feature vector, and based on the global fusion feature vector, generates a target aggregation model through iterative training; obtain the multimodal data of the urban stock space to be fused, input the multimodal data of the urban stock space to be fused into the sub-neural network model, obtain the target local fusion feature vector, and input the target local fusion feature vector into the target aggregation model to obtain the data fusion result. The present invention significantly improves the fusion processing efficiency of large-scale and diversified multi-source data of urban stock space, supports complex and sophisticated urban analysis and simulation, and provides strong technical support for urban planning and management. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flow chart of a preferred embodiment of the method for parallel fusion of urban stock space multimodal data of the present invention; Figure 2 is a schematic diagram of a multimodal spatiotemporal index structure of the present invention; Figure 3 It is a structural schematic diagram of the sub-neural network model of the present invention; Figure 4 It is a structural schematic diagram of a sub-grid data feature vector extraction module of the present invention; Figure 5 It is a structural schematic diagram of a sub-graph data feature vector extraction module of the present invention; Figure 6 It is a structural schematic diagram of a sub-time series data feature vector extraction module of the present invention; Figure 7 It is a schematic diagram of the structure of the present invention for preliminary fusion of multimodal data through a multi-head attention mechanism; Figure 8It is a structural diagram of a preferred embodiment of the parallel fusion system of urban stock space multimodal data of the present invention; Fig. 9 It is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical solution and advantages of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0018] With the continuous advancement of urbanization, the urban development model has entered a new stage of "stock improvement" from "incremental expansion". In order to improve the utilization efficiency of urban stock space, it is necessary to aggregate and fuse the stock space data, and conduct in-depth analysis to identify inefficient land use and implement reasonable transformation and upgrading. This process involves the integration and analysis of multi-source, multi-type, and multi-modal spatial data, including but not limited to vector data, remote sensing images, three-dimensional models, and spatiotemporal flow data that reflect dynamic changes such as traffic flow and population flow. Due to the huge amount of these data, the wide range of sources and the complexity, they each carry different geographic information and time dimension characteristics. Therefore, how to effectively aggregate and fuse these massive and diverse data has become a key technical problem that needs to be urgently solved in current urban planning and management. At present, most urban data fusion methods have two main limitations: first, they focus on a single data format, or although they can process multiple formats, they are all isomorphic data; second, most of the existing data fusion technologies are suitable for limited-scale data sets, but their computational efficiency and processing capabilities are insufficient when faced with multi-modal and massive urban spatial data. This limits the ability of urban planners and managers to fully understand the current status and development trends of the city, which in turn affects the quality and efficiency of decision-making.

[0019] In order to solve the above technical problems, the present invention provides a parallel fusion method for urban stock space multimodal data, which obtains urban stock space multimodal data, and classifies, roughly extracts feature vectors and normalizes the urban stock space multimodal data in turn to obtain a stock space data set of the full vector; designs a multimodal spatiotemporal index structure including a main node and multiple sub-nodes and constructs a neural network model, the main node divides the stock space data set of the full vector and the neural network model based on the multimodal spatiotemporal index structure to obtain multiple sub-feature vectors and multiple sub-neural network models, and assigns the multiple sub-feature vectors and the multiple sub-neural network models to each of the sub-nodes; based on the sub-feature vectors and the sub-neural network models, each of the sub-nodes The feature vector is finely extracted through a deep learning algorithm to obtain a fine feature vector, and the fine feature vector is preliminarily fused with multimodal data through a multi-head attention mechanism to obtain a local fusion feature vector; the master node obtains the local fusion feature vector output by each of the sub-nodes, and based on multiple local fusion feature vectors, a federated learning method is used for global aggregation to obtain a global fusion feature vector, and based on the global fusion feature vector, a target aggregation model is generated through iterative training; the multimodal data of the urban stock space to be fused is obtained, and the multimodal data of the urban stock space to be fused is input into the sub-neural network model to obtain a target local fusion feature vector, and the target local fusion feature vector is input into the target aggregation model to obtain a data fusion result. The present invention significantly improves the fusion processing efficiency of large-scale and diversified multi-source data of urban stock space, supports complex and sophisticated urban analysis and simulation, and provides strong technical support for urban planning and management.

[0020] The application content is further explained below through the description of embodiments in conjunction with the accompanying drawings.

[0021] A preferred embodiment of the parallel fusion method of urban stock space multimodal data of the present invention is as follows: Figure 1 As shown, specifically including: S1. Acquire multimodal data of urban stock space, and classify, roughly extract feature vectors and normalize the multimodal data of urban stock space in sequence to obtain a full-vector stock space data set.

[0022] In an implementation of this embodiment, the step of acquiring multimodal data of urban stock space, and sequentially classifying, roughly extracting feature vectors, and normalizing the multimodal data of urban stock space to obtain a full-vector stock space dataset specifically includes: Acquire urban stock space multimodal data, and divide the urban stock space multimodal data into raster data, graph data and time series data; Preprocessing the raster data, the graph data and the time series data respectively to obtain preprocessed raster data, preprocessed graph data and preprocessed time series data, wherein the preprocessing includes denoising and missing value filling; Respectively performing unified abstraction and modeling on the preprocessed raster data, the preprocessed graph data, and the preprocessed time series data to obtain a raster data model, a graph data model, and a time series data model; Using the raster data model, the graph data model and the time series data model, respectively, the preprocessed raster data, the preprocessed graph data and the preprocessed time series data are subjected to coarse feature vector extraction to obtain coarse feature vectors corresponding to the raster data, coarse feature vectors corresponding to the graph data and coarse feature vectors corresponding to the time series data; Normalizing the coarse feature vectors corresponding to the raster data, the graph data, and the time series data, respectively, to obtain a target coarse feature vector corresponding to the raster data, a target coarse feature vector corresponding to the graph data, and a target coarse feature vector corresponding to the time series data; The target coarse feature vector corresponding to the raster data, the target coarse feature vector corresponding to the graph data and the target coarse feature vector corresponding to the time series data are combined to obtain a stock spatial data set of the full vector.

[0023] Specifically, first, the multimodal data of urban stock space (multimodal data is data of different modes, such as images, texts, streaming data, etc.) are divided into raster data, graph data, time series data, etc., and then these multi-source heterogeneous data are preprocessed, including denoising and missing value filling. Denoising removes noise through filtering algorithms to ensure data quality. For remote sensing image data (i.e. raster data), adaptive filtering technology is used to reduce the impact of random noise; for time series data, sliding window smoothing method or Kalman filter is used to eliminate short-term fluctuations. For numerical data, interpolation method is used to fill missing values; for time series data, time series prediction model is used to fill missing values. Then, these multi-source heterogeneous data are uniformly abstracted and modeled, and uniformly abstracted into raster data model, graph data model and time series data model for expression. Raster data model is used to describe continuous spatial data and support efficient pixel-level analysis; graph data model is to capture the connectivity and adjacency relationship between nodes by converting vector data into graph structure, and then construct entity-relationship graph, calculate the adjacency matrix and degree matrix of the graph; time series data model is used to capture the change law of time dimension. Next, the feature vectors of the multimodal data of the urban stock space are roughly extracted using the raster data model, the graph data model and the time series data model. Rough feature extraction of raster data: the main features of the raster data are extracted using principal component analysis, and the texture features of the local area are extracted using the sliding window method to enhance the modeling ability of spatial locality; rough feature extraction of graph data: the points, lines and surfaces in the vector data are converted into graph structures, and the connectivity and adjacency relationships between the nodes are extracted using the topological analysis method, and the adjacency matrix and degree matrix of the graph are generated, and then the topological features of the nodes and edges are extracted using the graph neural network; rough feature extraction of time series data: the long short-term memory network is used to extract the long-term dependency relationship in the time series, and the statistical features and frequency domain features of the time series data are extracted. Finally, the extracted rough feature vectors are normalized, and the training set and the test set are divided to form a unified and complete feature vector space (i.e., the stock space data set of the full vector), ensuring that the data can be effectively input into the deep neural network (i.e., the neural network model). The present invention extracts the global feature vector (i.e., the target rough feature vector) through the raster data model, the graph data model and the time series data model, and finally forms the stock space data set D of the full vector: The i-th data in the stock space dataset D of the full vector The format is as follows: ; Among them, F1, F2, and F3 represent the relevant features extracted from the raster data model, the graph data model, and the time series data model, respectively (i.e., the target coarse feature vector. It should be noted that the target coarse feature vector is a global feature vector). Represents spatial coordinates, lat is latitude, lon is longitude, and t is timestamp.

[0024] S2. Design a multimodal spatiotemporal index structure including a main node and multiple sub-nodes and construct a neural network model. The main node divides the existing spatial data set of the full vector and the neural network model based on the multimodal spatiotemporal index structure to obtain multiple sub-feature vectors and multiple sub-neural network models, and assigns the multiple sub-feature vectors and the multiple sub-neural network models to each of the sub-nodes.

[0025] In an implementation of this embodiment, the design includes a multimodal spatiotemporal index structure of a main node and multiple sub-nodes and constructs a neural network model. The main node divides the stock spatial data set of the full vector and the neural network model based on the multimodal spatiotemporal index structure to obtain multiple sub-feature vectors and multiple sub-neural network models, and assigns the multiple sub-feature vectors and the multiple sub-neural network models to each of the sub-nodes, specifically including: A multimodal spatiotemporal index structure is designed. A quadtree is used for recursive partitioning in the spatial dimension to divide the two-dimensional space into multiple sub-regions. A linear index is used in the time dimension to sort and retrieve data by timestamp. The quadtree includes a main node and multiple sub-nodes. The main node represents the two-dimensional space, and each sub-node represents each sub-region. Constructing a neural network model, the neural network model includes a raster data feature extraction module, a graph data feature extraction module, a time series data feature extraction module and a multimodal data preliminary fusion module; Based on the multimodal spatiotemporal index structure, the master node divides the stock spatial data set of the full vector into sub-regions and time ranges to obtain a plurality of sub-feature vectors, where the sub-feature vectors correspond to the sub-regions one by one; Based on the multimodal spatiotemporal index structure, the master node divides the neural network model into sub-regions to obtain a plurality of sub-neural network models, and the sub-neural network models correspond to the sub-regions one by one; Assigning the sub-feature vectors and the sub-neural network models corresponding to the respective sub-regions to the corresponding respective sub-nodes; Establish a task allocation strategy, dynamically adjust the task allocation ratio according to the computing power and memory capacity of each child node, and give priority to allocating tasks to child nodes whose distance to the data storage location is less than the set value. In addition, introduce a backup mechanism to allocate some tasks to multiple child nodes at the same time.

[0026] In an implementation of the present embodiment, the sub-feature vectors and the sub-neural network models corresponding to each sub-region are assigned to corresponding sub-nodes, and the sub-feature vectors corresponding to each sub-region include sub-feature vectors corresponding to raster data, sub-feature vectors corresponding to graph data, and sub-feature vectors corresponding to time series data; the sub-neural network models corresponding to each sub-region include a sub-raster data feature extraction module, a sub-graph data feature extraction module, a sub-time series data feature extraction module, and a sub-multimodal data preliminary fusion module.

[0027] Specifically, first, a multimodal spatiotemporal index structure is designed. The multimodal spatiotemporal index structure is a unified spatial and temporal benchmark built based on the above-mentioned full-vector stock spatial dataset, which is used to improve the query efficiency of multimodal datasets in space, time, attributes and semantics, such as Figure 2 As shown. In the spatial dimension, quadtree is used for recursive partitioning. Quadtree is a hierarchical spatial index structure that can divide the two-dimensional space into multiple sub-areas. Each quadtree node (main node) represents a spatial range A(x y), recursively divide it into four sub-areas, A1, A2, A3, and A4 represent the four sub-areas corresponding to the main node respectively. By setting the conditions for stopping the recursive division, the depth of the division is controlled: the amount of data in the spatial range is less than the threshold N, and the spatial division depth reaches the maximum limit L. After recursive division, each sub-area corresponds to a unique spatial ID and contains all the data falling into the sub-area: ; in, , Respectively represent the horizontal and vertical coordinates of the spatial range A, Indicates sub-area and subregions The horizontal axis of Indicates sub-area and subregions The vertical coordinate of Indicates sub-area and subregions The horizontal axis of Indicates sub-area and subregions The vertical coordinate of The time dimension is managed using linear indexing, and data is sorted and retrieved by timestamp: ; Where T represents the entire time set, Indicates the nth timestamp; By specifying a time range To retrieve data, pass the data that meets the time range to the spatial index. After spatial division, the output is the data set R that falls into a certain spatial sub-area within the specified time range: ; Among them, R represents the space set after time and space division, Indicates the starting point of the time range, Indicates the end point of the time range. Figure 2 middle, Represents the data with subscript index 1 in the stock spatial dataset D of the full vector (which can be understood as the first data in the stock spatial dataset D of the full vector), Represents the data with subscript index 2 in the stock spatial dataset D of the full vector (which can be understood as the second data in the stock spatial dataset D of the full vector), Represents the data with subscript index 3 in the stock spatial dataset D of the full vector (which can be understood as the third data in the stock spatial dataset D of the full vector).

[0028] Then, the whole space of feature vectors is divided based on the multimodal spatiotemporal index structure. The master node divides the global feature vector (referring to the target coarse feature vector in the stock spatial data set of the full vector) and the neural network model into multiple sub-feature vectors and multiple sub-neural network models for subsequent distributed computing. First, the sub-feature vectors are divided: according to the results of the multimodal spatiotemporal index, the global feature vector is divided according to the spatial sub-region and time range; each sub-feature vector corresponds to a specific spatial sub-region and time range, and contains the feature vectors of all data points in the region. Specifically, the output result R of the multimodal spatiotemporal index is traversed to extract the data set of each spatial sub-region; the feature vectors in each data set are stored independently to form sub-feature vectors. Then, the sub-neural network model is divided: the neural network model (i.e., the global neural network model) is divided into multiple sub-neural network models according to the number of spatial sub-regions; each sub-neural network model is responsible for processing the calculation tasks of the corresponding sub-feature vector, extracting local features, performing preliminary data fusion, and outputting intermediate results for subsequent global fusion.

[0029] Finally, distributed task allocation is performed. First, the computing nodes (i.e., sub-nodes) are selected and initialized. The master node allocates the divided sub-feature vectors and sub-neural network models to the processes of each computing node. Each computing node is responsible for processing its corresponding sub-tasks and loading sub-feature vectors and sub-neural network models. Then, a task allocation strategy is established. According to the computing power and memory capacity of each computing node, the task allocation ratio is dynamically adjusted to ensure the workload balance of each computing node; tasks are preferentially allocated to computing nodes close to data storage locations to reduce data transmission overhead; in order to prevent computing node failures from causing task failures, a backup mechanism is introduced to allocate some tasks to multiple computing nodes to improve the reliability of the system.

[0030] S3. Based on the sub-feature vectors and the sub-neural network model, each of the sub-nodes performs fine feature vector extraction through a deep learning algorithm to obtain a fine feature vector, and performs preliminary multi-modal data fusion on the fine feature vector through a multi-head attention mechanism to obtain a local fused feature vector.

[0031] In an implementation of this embodiment, the fine feature vector includes a raster data feature vector, a graph data feature vector, and a time series data feature vector; Based on the sub-feature vectors and the sub-neural network model, each of the sub-nodes extracts feature vectors through a deep learning algorithm to obtain a fine feature vector, and performs preliminary multi-modal data fusion on the fine feature vectors through a multi-head attention mechanism to obtain a local fusion feature vector, specifically including: At each child node, the sub-feature vector corresponding to the raster data is sequentially extracted through the 3D convolution layer, GeLU activation function layer, maximum pooling layer, residual block, global pooling layer and the first fully connected layer of the sub-raster data feature extraction module to obtain a raster data feature vector. ; The sub-feature vector corresponding to the graph data is sequentially extracted through the graph convolution network, the first ReLU activation function layer, the multi-head attention network and the second fully connected layer of the sub-graph data feature extraction module to obtain the graph data feature vector ; The sub-feature vector corresponding to the time series data is sequentially extracted through the multi-layer long short-term memory network, the third fully connected layer, the second ReLU activation function layer, the discard layer and the fourth fully connected layer of the sub-time series data feature extraction module to obtain the time series data feature vector ; Through the multi-head attention mechanism, the raster data feature vector , the graph data feature vector and the time series data feature vector Perform preliminary fusion of multimodal data to obtain local fusion feature vectors.

[0032] In one implementation of this embodiment, the grid data feature vector is , the graph data feature vector and the time series data feature vector Perform preliminary fusion of multimodal data to obtain local fusion feature vectors, including: The raster data feature vector , the graph data feature vector and the time series data feature vector Through different linear transformation layers, the raster data feature vector is obtained. The corresponding first query vector Q1, the first key vector K1 and the first value vector V1, the graph data feature vector The corresponding second query vector Q2, the second key vector K2, the second value vector V2 and the time series data feature vector corresponding third query vector Q3, third key vector K3 and third value vector V3; The first similarity between the first query vector Q1 and the third key vector K3 is calculated in parallel by multiple attention heads , the second similarity between the second query vector Q2 and the second key vector K2 , the third similarity between the third query vector Q3 and the first key vector K1 : ; ; ; Where T represents transpose, represents the dimension of the first key vector, the second key vector, or the third key vector; According to the first similarity The second similarity and the third similarity , use the Softmax function to calculate the first attention weight corresponding to the raster data , the second attention weight corresponding to the graph data The third attention weight corresponding to the time series data : ; ; ; According to the first attention weight , the second attention weight and the third attention weight , respectively perform weighted summation on the third value vector V3, the second value vector V2 and the first value vector V1 to obtain a local fusion feature vector O: ; ; ; ; in, Represents the attention mechanism.

[0033] It should be noted that the structure of the neural network model of the present invention and the sub-neural network model are the same. Figure 3 The structure of the grid data feature extraction module and the sub-grid data feature extraction module are the same, as shown in Figure 4 The structure of the graph data feature extraction module is the same as that of the subgraph data feature extraction module. Figure 5 The structure of the time series data feature extraction module and the sub-time series data feature extraction module are the same, both as Figure 6 shown.

[0034] Specifically, Figure 3 As shown in the figure, first, the sub-neural network model uses different network layers to access the feature vectors of each source data. For example, the raster data uses convolutional layers and pooling layers for feature extraction; the graph data is processed using graph convolutional layers, graph attention networks, and autoencoders; and the time series data is processed using recurrent neural networks. Then, the outputs of each sub-module are fused using a multi-head attention mechanism to achieve the initial fusion of multi-source data in the urban stock space (i.e., local fusion). Figure 3 and Figure 4 As shown in the figure, for raster data such as remote sensing images, the spatial and spectral features of the image are first extracted through a 3D convolution layer, and the size of each convolution kernel is 7*7*3; and the GeLU activation function is used to enhance the nonlinear expression ability of the model; then the feature dimension is reduced through a maximum pooling layer (i.e., maximum pooling) while retaining important information; then, the residual block contains two 3D convolution layers with a convolution kernel of 3*3*3 and a GeLU activation function, and a multi-layer residual network is used to extract deeper features to alleviate the gradient disappearance problem; finally, a global pooling layer (i.e., global average pooling) is used to generate a fixed-length feature vector, and the feature dimension is adjusted through a fully connected layer to obtain , for subsequent network processing. Figure 3 and Figure 5As shown in the figure, for graph data, the structural information of nodes and their neighborhoods is extracted through the graph convolutional network, the ReLU activation function is combined to enhance the nonlinear expression ability, and the multi-head attention network is used to calculate the relationship weights between nodes to highlight the important node relationships. Finally, the feature vector is further extracted through the fully connected layer. .like Figure 3 and Figure 6 As shown in the figure, for time series data, the dynamic dependencies in time are captured through a multi-layer LSTM (Long Short-Term Memory) network, and then the fully connected layer is used to reduce the feature dimension. The generalization ability of the model is enhanced through the ReLU activation function and the discard layer, and finally the time series data feature vector is generated. .

[0035] In each LSTM network, the calculation process of each timestamp t is realized through four key components: input gate, forget gate, cell state and output gate, to achieve dynamic modeling of time series data. Determines how much of the input information of the current timestamp will be written into the cell state. Through a sigmoid activation function ( Figure 6 Medium indicates) calculation, combined with the current input and the hidden state at the previous timestamp , and get a weight value between 0 and 1. Forget gate Using the sigmoid function, based on the current input and the hidden state of the previous timestamp, a weight value is output to determine the degree of retention of each element in the cell state. Updated by the output of the forget gate and the input gate. The forget gate controls which information in the cell state needs to be discarded, while the input gate determines which new information needs to be written. The update of the cell state combines the forget operation with the writing of new information, and the new information is scaled by a tanh function to ensure that the value of the cell state remains in the appropriate range. Output gate Determines the hidden state of the current timestamp The hidden state not only contains the information of the current timestamp, but also retains the historical information through the cell state, so as to capture the long-term dependencies in the time series.

[0036] In multimodal data fusion, Figure 7As shown, the feature vectors extracted from raster data, graph data, and time series data are effectively fused through a multi-head attention mechanism. First, the three feature vectors are passed through three different linear transformation layers to obtain query, key, and value matrices, respectively, and the correlation between different modalities is learned by adjusting the weights. The raster data feature vector is linearly transformed to obtain K1, Q1, and V1. The graph data feature vector is linearly transformed to obtain K2, Q2, and V2. The time series data feature vector is linearly transformed to obtain K3, Q3, and V3. Cross-modal information fusion is performed on raster data and time series data. The similarity between Q1 and K3, and Q3 and K1 is calculated in parallel through multiple attention heads. The dot product can reflect the correlation between different modalities and calculate the similarity Score. The attention weight is obtained using the SoftMax function. , , , so that the model can make weighted decisions between multiple modalities, thereby focusing on the most relevant information. The values ​​V1, V2, and V3 are weighted according to the calculated attention weights, and through weighted summation, the model can dynamically select the most relevant information. The multi-head attention mechanism captures different attention patterns through parallel calculations, allowing the model to learn and fuse information from different modalities from multiple perspectives, and ultimately obtain a local fusion feature vector O that integrates all modal features, which can more comprehensively represent the input data.

[0037] S4. The master node obtains the local fusion feature vector output by each of the child nodes, and based on the multiple local fusion feature vectors, uses a federated learning method to perform global aggregation to obtain a global fusion feature vector, and based on the global fusion feature vector, generates a target aggregation model through iterative training.

[0038] In an implementation of this embodiment, the master node obtains a local fusion feature vector output by each of the child nodes, performs global aggregation based on multiple local fusion feature vectors using a federated learning method to obtain a global fusion feature vector, and generates a target aggregation model based on the global fusion feature vector through iterative training, specifically including: The master node obtains the local fusion feature vector output by each of the child nodes, and based on the weight corresponding to each of the child nodes, uses the federated learning algorithm to perform weighted fusion on the local fusion feature vector output by each of the child nodes to obtain a global fusion feature vector : ; Among them, i represents the i-th child node, represents the local fusion feature vector output by the i-th child node, Represents the weight corresponding to the i-th child node; Among them, the weight corresponding to the i-th child node According to the number of samples of the i-th child node The calculation results are: ; in, Indicates the number of samples of the first child node, Indicates the number of samples of the second child node, represents the number of samples of the third child node, Indicates the number of samples of the fourth child node; The global fusion feature vector generated by the master node is used to update the aggregation weight parameter of the master node, and the updated aggregation weight parameter is fed back to each child node to obtain the target aggregation model; The aggregation weight parameter of the master node includes the weight corresponding to each of the child nodes.

[0039] Specifically, after processing the data, each child process (child node) will calculate the result vector etc. are returned to the master node. After receiving the feature vectors of all child processes, the master node uses the federated average algorithm to perform weighted fusion on the parameters uploaded by each child node to obtain the final convergence result; then the final convergence result (i.e., the global fusion feature vector) is used to update the master node's convergence weight parameters (i.e., the weights corresponding to each child node); finally, the updated convergence weight parameters are fed back to each child node to achieve a global collaborative training iteration, and finally generate a target convergence model with stronger generalization ability. For the master node, this convergence process is a convergence model.

[0040] S5. Obtain the multimodal data of the urban stock space to be fused, input the multimodal data of the urban stock space to be fused into the sub-neural network model to obtain the target local fusion feature vector, and input the target local fusion feature vector into the target convergence model to obtain the data fusion result.

[0041] Specifically, the present invention solves the problem of poor fusion performance of multimodal urban stock spatial data through a neural network for fusion of multimodal data of urban stock space; and solves the problem of poor fusion performance of massive urban stock spatial data through a distributed parallel fusion method of spatial data. The present invention implements distributed computing of each sub-node through the sub-process of each sub-node, that is, distributed preliminary fusion of multimodal data, and then realizes parallel fusion of multimodal data of urban stock space, which significantly improves the fusion processing efficiency of large-scale and diversified multi-source data of urban stock space, supports complex and sophisticated urban analysis and simulation, and provides strong technical support for urban planning and management.

[0042] In addition, based on the above-mentioned parallel fusion method of urban stock space multimodal data, the present invention also provides a parallel fusion system of urban stock space multimodal data, wherein a preferred embodiment of the parallel fusion system of urban stock space multimodal data is as follows: Figure 8 As shown, specifically including: Data preliminary processing module 01: used to obtain urban stock space multimodal data, and classify, roughly extract feature vectors and normalize the urban stock space multimodal data in sequence to obtain a full vector stock space data set; Distributed task assignment module 02: used to design a multimodal spatiotemporal index structure including a master node and multiple sub-nodes and to construct a neural network model, wherein the master node divides the stock spatial data set of the full vector and the neural network model based on the multimodal spatiotemporal index structure to obtain multiple sub-feature vectors and multiple sub-neural network models, and assigns the multiple sub-feature vectors and the multiple sub-neural network models to each of the sub-nodes; Preliminary fusion module 03: for extracting fine feature vectors of each sub-node through a deep learning algorithm based on the sub-feature vectors and the sub-neural network model to obtain fine feature vectors, and performing preliminary multi-modal data fusion on the fine feature vectors through a multi-head attention mechanism to obtain local fused feature vectors; Global aggregation module 04: used for the master node to obtain the local fusion feature vector output by each of the child nodes, and based on the multiple local fusion feature vectors, a federated learning method is used to perform global aggregation to obtain a global fusion feature vector, and based on the global fusion feature vector, a target aggregation model is generated through iterative training; Multimodal data fusion module 05: used to obtain the multimodal data of the urban stock space to be fused, input the multimodal data of the urban stock space to be fused into the sub-neural network model to obtain the target local fusion feature vector, and input the target local fusion feature vector into the target convergence model to obtain the data fusion result.

[0043] In addition, based on the above-mentioned method and system for parallel fusion of multimodal data of urban stock space, the present invention also provides a terminal accordingly, wherein a preferred embodiment of the terminal is as follows: Fig. 9 As shown, it specifically includes a processor 10, a memory 20 and a display 30. Fig. 9 Only some components of the terminal are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0044] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, and a flash card (Flash Card) equipped on the terminal. Further, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal, such as program codes of the storage terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, the parallel fusion program 40 of urban stock space multimodal data is stored on the memory 20, and the parallel fusion program 40 of urban stock space multimodal data can be executed by the processor 10, thereby realizing the steps of the parallel fusion method of urban stock space multimodal data in the present application.

[0045] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program codes or process data stored in the memory 20, such as executing a parallel fusion program 40 of multimodal data of urban stock space.

[0046] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 30 is used to display information on the terminal and to display a visual user interface.

[0047] In one embodiment, when the processor 10 executes the parallel fusion program 40 of urban stock space multimodal data in the memory 20, the steps of the parallel fusion method of urban stock space multimodal data as described above are implemented.

[0048] The present invention also provides a computer-readable storage medium accordingly, wherein the computer-readable storage medium stores a parallel fusion program for urban stock space multimodal data, and when the parallel fusion program for urban stock space multimodal data is executed by a processor, the steps of the parallel fusion method for urban stock space multimodal data as described above are implemented.

[0049] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the existence of other identical elements in the process, method, article or terminal including the element.

[0050] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the program can be stored in a computer-readable storage medium that can be read by a computer, and the program can include the processes of the above-mentioned method embodiments when executed. The computer-readable storage medium can be a memory, a disk, an optical disk, etc.

[0051] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A parallel fusion method for multimodal data of urban stock space, characterized in that: The parallel fusion method of urban stock space multimodal data includes: Acquire multimodal data of urban stock space, and classify, roughly extract feature vectors and normalize the multimodal data of urban stock space in sequence to obtain a full-vector stock space data set; Designing a multimodal spatiotemporal index structure including a main node and a plurality of sub-nodes and constructing a neural network model, wherein the main node divides the stock spatial data set of the full vector and the neural network model based on the multimodal spatiotemporal index structure to obtain a plurality of sub-feature vectors and a plurality of sub-neural network models, and assigning the plurality of sub-feature vectors and the plurality of sub-neural network models to each of the sub-nodes; Based on the sub-feature vectors and the sub-neural network model, each of the sub-nodes performs fine feature vector extraction through a deep learning algorithm to obtain a fine feature vector, and performs preliminary multi-modal data fusion on the fine feature vector through a multi-head attention mechanism to obtain a local fused feature vector; The master node obtains the local fusion feature vector output by each of the child nodes, performs global aggregation based on the multiple local fusion feature vectors by using a federated learning method to obtain a global fusion feature vector, and generates a target aggregation model based on the global fusion feature vector through iterative training; Acquire the multimodal data of urban stock space to be fused, input the multimodal data of urban stock space to be fused into the sub-neural network model to obtain the target local fusion feature vector, and input the target local fusion feature vector into the target convergence model to obtain the data fusion result.

2. The parallel fusion method of urban stock space multimodal data according to claim 1 is characterized in that: The method of obtaining the multimodal data of urban stock space, and sequentially classifying, roughly extracting and normalizing the multimodal data of urban stock space to obtain a stock space data set of full vectors specifically includes: Acquire urban stock space multimodal data, and divide the urban stock space multimodal data into raster data, graph data and time series data; Preprocessing the raster data, the graph data and the time series data respectively to obtain preprocessed raster data, preprocessed graph data and preprocessed time series data, wherein the preprocessing includes denoising and missing value filling; Respectively performing unified abstraction and modeling on the preprocessed raster data, the preprocessed graph data, and the preprocessed time series data to obtain a raster data model, a graph data model, and a time series data model; Using the raster data model, the graph data model and the time series data model, respectively, the preprocessed raster data, the preprocessed graph data and the preprocessed time series data are subjected to coarse feature vector extraction to obtain coarse feature vectors corresponding to the raster data, coarse feature vectors corresponding to the graph data and coarse feature vectors corresponding to the time series data; Normalizing the coarse feature vectors corresponding to the raster data, the graph data, and the time series data, respectively, to obtain a target coarse feature vector corresponding to the raster data, a target coarse feature vector corresponding to the graph data, and a target coarse feature vector corresponding to the time series data; The target coarse feature vector corresponding to the raster data, the target coarse feature vector corresponding to the graph data and the target coarse feature vector corresponding to the time series data are combined to obtain a stock spatial data set of the full vector.

3. The parallel fusion method of urban stock space multimodal data according to claim 2 is characterized in that: The design includes a multimodal spatiotemporal index structure of a main node and multiple sub-nodes and a constructed neural network model. The main node divides the stock spatial data set of the full vector and the neural network model based on the multimodal spatiotemporal index structure to obtain multiple sub-feature vectors and multiple sub-neural network models, and assigns the multiple sub-feature vectors and the multiple sub-neural network models to each of the sub-nodes, specifically including: A multimodal spatiotemporal index structure is designed. A quadtree is used for recursive partitioning in the spatial dimension to divide the two-dimensional space into multiple sub-regions. A linear index is used in the time dimension to sort and retrieve data by timestamp. The quadtree includes a main node and multiple sub-nodes. The main node represents the two-dimensional space, and each sub-node represents each sub-region. Constructing a neural network model, wherein the neural network model includes a raster data feature extraction module, a graph data feature extraction module, a time series data feature extraction module and a multimodal data preliminary fusion module; Based on the multimodal spatiotemporal index structure, the master node divides the stock spatial data set of the full vector into sub-regions and time ranges to obtain a plurality of sub-feature vectors, where the sub-feature vectors correspond to the sub-regions one by one; Based on the multimodal spatiotemporal index structure, the master node divides the neural network model into sub-regions to obtain a plurality of sub-neural network models, and the sub-neural network models correspond to the sub-regions one by one; Assigning the sub-feature vectors and the sub-neural network models corresponding to the respective sub-regions to the corresponding respective sub-nodes; Establish a task allocation strategy, dynamically adjust the task allocation ratio according to the computing power and memory capacity of each child node, and give priority to allocating tasks to child nodes whose distance to the data storage location is less than the set value. In addition, introduce a backup mechanism to allocate some tasks to multiple child nodes at the same time.

4. The parallel fusion method of urban stock space multimodal data according to claim 3 is characterized in that: The sub-feature vectors and sub-neural network models corresponding to each sub-region are assigned to corresponding sub-nodes, wherein the sub-feature vectors corresponding to each sub-region include sub-feature vectors corresponding to raster data, sub-feature vectors corresponding to graph data, and sub-feature vectors corresponding to time series data; the sub-neural network models corresponding to each sub-region include sub-raster data feature extraction modules, sub-graph data feature extraction modules, sub-time series data feature extraction modules, and sub-multimodal data preliminary fusion modules.

5. The parallel fusion method of urban stock space multimodal data according to claim 4 is characterized in that: The detailed feature vectors include raster data feature vectors, graph data feature vectors and time series data feature vectors; Based on the sub-feature vectors and the sub-neural network model, each of the sub-nodes extracts feature vectors through a deep learning algorithm to obtain a fine feature vector, and performs preliminary multi-modal data fusion on the fine feature vectors through a multi-head attention mechanism to obtain a local fusion feature vector, specifically including: At each child node, the sub-feature vector corresponding to the raster data is sequentially extracted through the 3D convolution layer, GeLU activation function layer, maximum pooling layer, residual block, global pooling layer and the first fully connected layer of the sub-raster data feature extraction module to obtain a raster data feature vector. ; The sub-feature vector corresponding to the graph data is sequentially extracted through the graph convolution network, the first ReLU activation function layer, the multi-head attention network and the second fully connected layer of the sub-graph data feature extraction module to obtain the graph data feature vector ; The sub-feature vector corresponding to the time series data is sequentially extracted through the multi-layer long short-term memory network, the third fully connected layer, the second ReLU activation function layer, the discard layer and the fourth fully connected layer of the sub-time series data feature extraction module to obtain the time series data feature vector ; Through the multi-head attention mechanism, the raster data feature vector , the graph data feature vector and the time series data feature vector Perform preliminary fusion of multimodal data to obtain local fusion feature vectors.

6. The parallel fusion method of urban stock space multimodal data according to claim 5 is characterized in that: The multi-head attention mechanism is used to transform the grid data feature vector , the graph data feature vector and the time series data feature vector Perform preliminary fusion of multimodal data to obtain local fusion feature vectors, including: The raster data feature vector , the graph data feature vector and the time series data feature vector Through different linear transformation layers, the raster data feature vector is obtained. The corresponding first query vector Q1, the first key vector K1 and the first value vector V1, the graph data feature vector The corresponding second query vector Q2, the second key vector K2, the second value vector V2 and the time series data feature vector corresponding third query vector Q3, third key vector K3 and third value vector V3; The first similarity between the first query vector Q1 and the third key vector K3 is calculated in parallel by multiple attention heads , the second similarity between the second query vector Q2 and the second key vector K2 , the third similarity between the third query vector Q3 and the first key vector K1 : ; ; ; Where T represents transpose, represents the dimension of the first key vector, the second key vector, or the third key vector; According to the first similarity The second similarity and the third similarity , use the Softmax function to calculate the first attention weight corresponding to the raster data , the second attention weight corresponding to the graph data The third attention weight corresponding to the time series data : ; ; ; According to the first attention weight , the second attention weight and the third attention weight , respectively perform weighted summation on the third value vector V3, the second value vector V2 and the first value vector V1 to obtain a local fusion feature vector O: ; ; ; ; in, Represents the attention mechanism.

7. The parallel fusion method of urban stock space multimodal data according to claim 6 is characterized in that: The master node obtains the local fusion feature vector output by each of the child nodes, performs global aggregation based on the multiple local fusion feature vectors using a federated learning method to obtain a global fusion feature vector, and generates a target aggregation model based on the global fusion feature vector through iterative training, specifically including: The master node obtains the local fusion feature vector output by each of the child nodes, and based on the weight corresponding to each of the child nodes, uses the federated learning algorithm to perform weighted fusion on the local fusion feature vector output by each of the child nodes to obtain a global fusion feature vector : ; Among them, i represents the i-th child node, represents the local fusion feature vector output by the i-th child node, Represents the weight corresponding to the i-th child node; Among them, the weight corresponding to the i-th child node According to the number of samples of the i-th child node The calculation results are: ; in, Indicates the number of samples of the first child node, Indicates the number of samples of the second child node, represents the number of samples of the third child node, Indicates the number of samples of the fourth child node; The global fusion feature vector generated by the master node is used to update the aggregation weight parameter of the master node, and the updated aggregation weight parameter is fed back to each child node to obtain the target aggregation model; The aggregation weight parameter of the master node includes the weight corresponding to each of the child nodes.

8. A parallel fusion system for multimodal data of urban stock space, characterized in that: The parallel fusion system of urban stock space multimodal data includes: Data preliminary processing module: used to obtain urban stock space multimodal data, and classify, roughly extract feature vectors and normalize the urban stock space multimodal data in turn to obtain a full vector stock space data set; Distributed task allocation module: used to design a multimodal spatiotemporal index structure including a main node and multiple sub-nodes and to construct a neural network model, wherein the main node divides the stock spatial data set of the full vector and the neural network model based on the multimodal spatiotemporal index structure to obtain multiple sub-feature vectors and multiple sub-neural network models, and allocates the multiple sub-feature vectors and the multiple sub-neural network models to each of the sub-nodes; Preliminary fusion module: used for extracting fine feature vectors of each sub-node through a deep learning algorithm based on the sub-feature vectors and the sub-neural network model to obtain fine feature vectors, and performing preliminary multi-modal data fusion on the fine feature vectors through a multi-head attention mechanism to obtain local fused feature vectors; Global aggregation module: used for the master node to obtain the local fusion feature vector output by each of the child nodes, and based on the multiple local fusion feature vectors, a federated learning method is used to perform global aggregation to obtain a global fusion feature vector, and based on the global fusion feature vector, a target aggregation model is generated through iterative training; Multimodal data fusion module: used to obtain the multimodal data of the urban stock space to be fused, input the multimodal data of the urban stock space to be fused into the sub-neural network model to obtain the target local fusion feature vector, and input the target local fusion feature vector into the target convergence model to obtain the data fusion result.

9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a parallel fusion program for urban stock space multimodal data stored in the memory and executable on the processor. When the parallel fusion program for urban stock space multimodal data is executed by the processor, the steps of the parallel fusion method for urban stock space multimodal data as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a parallel fusion program for urban stock space multimodal data. When the parallel fusion program for urban stock space multimodal data is executed by a processor, the steps of the parallel fusion method for urban stock space multimodal data as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Multi-modal data fusion method and system, readable storage medium and computer equipment

    CN119128790A

  • Live broadcast room content identification and intelligent distribution method and system based on multi-modal fusion

    CN119377895A

Cited By

  • Data interaction sharing method and system suitable for digital asset management

    CN120596282A

  • Urban space general representation learning method and device based on multi-modal spatio-temporal data fusion, terminal and storage medium

    CN121980249A

  • A general urban spatial representation learning method, device, terminal, and storage medium based on multimodal spatiotemporal data fusion.

    CN121980249B