A multi-scale cross-modal data enhancement method and system under risk disturbance

By constructing a spatiotemporal coupling relationship graph and a cross-modal association network, the challenge of multi-scale cross-modal data enhancement under risk disturbances is solved, and comprehensive data enhancement and stability improvement are achieved, which is suitable for data analysis and decision support in urban agglomerations.

CN118643344BActive Publication Date: 2025-09-26BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410952983.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-16
Publication Date
2025-09-26
Estimated Expiration
2044-07-16

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively cope with the enhancement of multi-scale and cross-modal data in the face of risk disturbances, resulting in problems such as degraded data quality, abnormal samples, missing information, and time sequence disorder, which affect the operation and management of urban agglomerations.

Method used

A data enhancement method based on multi-spatiotemporal scale coupling and cross-modal association network is adopted to achieve comprehensive enhancement and compensation of cross-modal data by constructing spatiotemporal coupling relationship graph, feature fusion network, feature alignment and migration technology.

Benefits of technology

The robustness and stability of data enhancement have been improved, which can better cope with risk disturbances and provide a richer and more reliable data analysis foundation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118643344B_ABST
    Figure CN118643344B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-scale cross-modal data enhancement method and system under risk disturbance, thereby providing a more comprehensive and accurate data foundation for digital intelligence-driven urban agglomeration applications; multi-domain cross-modal perception data of urban agglomerations at multiple spatiotemporal scales have problems such as local missing, sparse sampling, noise disturbance, and unclear patterns. By exploring data enhancement methods under risk disturbance, the originally discrete human mobility behavior data can be serialized and the missing values ​​in the samples can be filled to ensure the continuity of the samples in time, space and multiple scales, thereby improving the overall quantity and quality of the samples. With the help of data enhancement technology, various noises and disturbances can be introduced to simulate the real situation of the real world, thereby improving the robustness and generalization of the model. The process of samples from sparse to continuous, quality from missing to complete, model from local to global, and patterns from fuzzy to clear is realized, so as to better understand the inherent properties and associations of the data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi-scale cross-modal data enhancement method and system under risk disturbance, belonging to the cross-application field of big data and data enhancement processing. Background Art

[0002] Risk disturbances are primarily categorized into three types: 1) sudden disturbances, such as weather emergencies, major social events, and traffic accidents; 2) cyclical disturbances, such as major holidays, the start of the school year, and rush hour periods; and 3) long-term trend disturbances, such as social evolution, climate change, and industrial restructuring. Multi-scale, cross-modal data exposed to risk disturbances requires data augmentation for optimal use.

[0003] Data augmentation is a technique that transforms and processes raw data to generate new data samples. Its purpose is to expand datasets and improve model performance and generalization capabilities. By analyzing the distribution of data features, raw data can be augmented in a targeted manner. Therefore, the consistency and stability of data features are fundamental to effective data augmentation. However, urban data is modally diverse, widely distributed, and subject to a large amount of uncontrollable noise interference, resulting in degraded data quality, anomalous samples, missing information, and time series distortion. Currently, spatiotemporal prediction compensation and cross-modal fusion are two common methods for anti-disturbance data augmentation.

[0004] Data enhancement methods based on spatiotemporal prediction compensation: This type of method aims to utilize the characteristic distribution and sequence trends of spatiotemporal data to achieve data enhancement. To address the problem of abnormal samples, there is currently a hypernetwork Kalman filter used to track cross-modal data changes, which improves the generalization and robustness of fused features; there is also a spatiotemporal denoising graph autoencoding model that achieves the recovery and enhancement of lost data under environmental disturbances. To address the problem of missing information, a square root Kalman filter with self-learning ability has been proposed, which can provide continuous navigation observations when GPS data fails; a long-term multidimensional spatiotemporal graph convolutional network LMSTGCN has been proposed, which uses a multidimensional graph convolution module to simultaneously model spatial and short-term temporal information, and on this basis, a spatiotemporal adjacency matrix construction method is designed to process short-term correlations in sequence data based on a gated time module.

[0005] Data augmentation methods based on cross-modal fusion: These methods aim to leverage the complementarity of multimodal data to achieve cross-validation and feature fusion. To address data quality degradation, current technologies can achieve a cascaded combination of low-level and high-level features by concatenating vectorized features from multiple sources. Some researchers have combined neural architecture search with progressive exploration to propose a method for constructing fusion functions that adaptively selects fusion layers and concatenation weights. Attention models are also a common approach in cross-modal fusion. To overcome feature perturbations caused by outliers, stacked attention networks can be used. This multi-layer attention model allows for multiple queries and dynamically fuses cross-modal features based on the query results. Current dual attention networks simultaneously consider the feature distributions of each modality, constructing an attention distribution and memory vector model for the feature vector. Subsequently, current technologies utilize high-dimensional convolution operators to capture local features, further enriching data diversity. Current technologies employ Bayesian deep learning methods to mitigate unimodal bias, constructing a cross-modal likelihood formula and particle filter to improve the interference resistance of the fused features. A multi-scale fusion generalization module has been proposed to aggregate global reference information. This module uses global perceptrons and location-aware mapping to predict urban traffic flow in noisy environments.

[0006] While these methods achieve data enhancement from various perspectives, the traditional data enhancement methods they represent fail to consider the impact of risk perturbations on multi-scale, cross-modal data enhancement. Consequently, they often fail to effectively address the challenges posed by risk perturbations and suffer from various shortcomings and deficiencies. In the context of real urban agglomerations, the combined effects of various types of risk perturbations pose significant challenges to the operation and management of modern urban agglomerations. Summary of the Invention

[0007] The technology of the present invention solves the problem: Overcoming the shortcomings of the existing technology, providing a multi-scale cross-modal data enhancement method and system under risk disturbance, comprehensively utilizing risk disturbance factors, improving the robustness and stability of data enhancement, making data enhancement more comprehensive and accurate, and better able to cope with the challenges of risk disturbance.

[0008] Technical solution of the present invention:

[0009] In a first aspect, the present invention provides a multi-scale cross-modal data enhancement method under risk disturbance, comprising a data enhancement step based on multi-spatiotemporal scale coupling and a data compensation step based on a cross-modal association network, wherein:

[0010] The data enhancement step based on multi-spatiotemporal scale coupling includes: collecting multi-scale unimodal data on the temporal and spatial movement of people, vehicles, and objects in the city under risk disturbance, modeling the multi-scale unimodal data, and constructing a spatiotemporal coupling relationship diagram; the risk disturbance refers to a sudden, periodic, or long-term trend disturbance phenomenon encountered; based on the spatiotemporal coupling relationship diagram, processing the spatiotemporal coupling relationship diagram using a feature fusion network system to generate interpolated data to form enhanced unimodal data, and coupling multiple enhanced unimodal data into cross-modal data;

[0011] The data compensation steps of the cross-modal association network are as follows: using feature alignment technology to adjust the feature dimensions and scales in the cross-modal data, eliminating the differences between the cross-modal data, and obtaining the cross-modal data after feature alignment; using feature migration technology on the cross-modal data after feature alignment to obtain the feature association relationship of each single modal data between the cross-modal data; using graph fusion feature construction technology to process the feature association relationship of each single modal data between the cross-modal data, and obtain the correlation and complementarity between the cross-modal data; finally, based on enhancing the correlation and complementarity between the cross-modal data, multi-scale cross-modal data enhancement is achieved to complete cross-modal data compensation.

[0012] In particular, in the data enhancement step based on multi-spatiotemporal scale coupling, the multi-scale single-modal data is modeled and a spatiotemporal coupling relationship diagram is constructed as follows:

[0013] The multi-scale single-modal data of a person-vehicle-object movement behavior is recorded as a data source , the data source The relationship diagram is recorded as a graph , multiple data sources at the same time Composition diagram A collection of nodes , data source A diagram of the relationship between urban space and the surrounding environment The edge set , which reflects the data source Spatial correlation between data sources The historical data at different times are different, so The historical data is recorded as , then one The collection can reflect the time relevance of the data source;

[0014] Leveraging data sources and pictures Complete cross-modal complex network modeling of multi-scale single-modal data and construct a spatiotemporal coupling relationship diagram, which is implemented as follows: Read each pair of adjacent nodes in the connected edges ,definition is the traffic flow characteristic of a node, is the traffic flow characteristic of another node, which refers to the quantitative change pattern of pedestrian and vehicle flows under different conditions and their mutual relationship. is a sliding time window; When the value is short enough, the graph The same behavior states between adjacent nodes at the same time are considered to be linearly correlated. The Pearson correlation coefficient is used to define the relationship between adjacent nodes. The Pearson correlation of traffic flow characteristics between adjacent nodes is calculated, and the absolute value of the correlation is taken. Based on the obtained correlation, the space-time coupling relationship graph is constructed by the threshold Gaussian kernel function, and the adjacency matrix of the space-time coupling relationship graph is calculated and obtained. , the adjacency matrix changes with the sliding time window The movement changes dynamically over time.

[0015] In particular, in the data enhancement step based on multi-spatiotemporal scale coupling, the feature fusion network system includes a graph convolutional neural network (GCN) module, a bidirectional long short-term memory neural network (BiLSTM) module, a feature fusion module, and an autoencoder network module;

[0016] First, the graph convolutional neural network module extracts the topological relationships in the spatiotemporal coupling relationship graph to obtain the spatial dependencies between each person-vehicle-object movement behavior and other person-vehicle-object movement behaviors. Based on these spatial dependencies, the spatial features of the person-vehicle-object movement behaviors are generated.

[0017] The bidirectional long short-term memory neural network then extracts the temporal relationship in the spatiotemporal coupling relationship graph to obtain the temporal dependency of each person-vehicle-object movement behavior on other person-vehicle-object movement behaviors. Based on this temporal dependency, the temporal features of the person-vehicle-object movement behaviors are generated.

[0018] Afterwards, the feature fusion module fuses the spatial and temporal features of each person-vehicle-object movement behavior to obtain the fused cross-modal features;

[0019] The fused cross-modal features and the original multi-scale single-modal data of human-vehicle-object movement behaviors in time and space are input into the autoencoder network module in sequence form;

[0020] A temporal and spatial attention layer is inserted into the neural network layer between the encoder and the decoder in the autoencoder network module to implement an attention mechanism. The attention mechanism calculates the attention weight of the data with strong temporal and spatial correlation according to the data of the input sequence and its context data, determines the location of the missing data, and interpolates the multi-scale unimodal data.

[0021] In particular, in the data compensation step of the cross-modal association network, the feature alignment technology adopts an unsupervised explicit cross-modal feature alignment technology, which is implemented as follows: an unsupervised method is adopted to process cross-modal data using an explicit alignment algorithm based on a dynamic Bayesian network to achieve feature alignment of cross-modal data; the explicit alignment algorithm based on the dynamic Bayesian network uses multiple time slices that constitute the dynamic Bayesian network to store cross-modal data, and then uses the joint probability of the Bayesian network to automatically obtain the conditional probability of the current time slice state, thereby achieving feature alignment of cross-modal data.

[0022] In particular, in the data compensation step of the cross-modal association network, the feature transfer technology is implemented as follows:

[0023] (1) First, data complementation is performed on the cross-modal data after feature alignment to complete the missing modal data in each group of cross-modal data. The data complementation refers to reading the missing modal data in the current group of cross-modal data from other groups of cross-modal data and completing the modality of the current group of cross-modal data to achieve data complementation.

[0024] (2) Secondly, data enhancement is performed on the cross-modal data after data complementation. The robust features of a certain modality in the cross-modal data are extracted and learned from each group of cross-modal data, and the learned single modality features are applied to the other modality.

[0025] (3) Finally, the feature dimensionality reduction mapping function is implemented to extract the key information of each single-modal data in each set of cross-modal data, map the key information into the shared feature space, and perform feature fusion to obtain the correlation relationship between each single-modal feature in the cross-modal data;

[0026] In (1)-(3), the maximum mean discrepancy (MMD) value is used for normalization to achieve the unification of different but related data, so as to facilitate comparison between data and further feature migration and fusion.

[0027] In particular, in the data compensation step of the cross-modal association network, the cross-modal feature construction technology based on graph fusion is:

[0028] First, a feature graph is created for each single-modal data. Each node in the graph represents the features contained in the modality, and the edge represents the connection or similarity between features. The weight of the edge is determined by the similarity between the features. These weights form a weight matrix.

[0029] Then, the nodes and edges in each unimodal feature map are merged, and the feature maps of different modalities are integrated into a cross-modal feature map. A graph attention network is constructed in this map, so that each node can aggregate information from other modalities to realize information transmission, further update the nodes and the weights in the weight matrix, and finally obtain the correlation and complementarity between cross-modal data.

[0030] In a second aspect, the present invention provides a multi-scale cross-modal data enhancement system under risk disturbance, comprising a data enhancement module based on multi-spatiotemporal scale coupling and a data compensation module based on a cross-modal association network, wherein:

[0031] The data enhancement modality module based on multi-spatiotemporal scale coupling: under risk disturbance, collects multi-scale unimodal data on the temporal and spatial movement behaviors of people, vehicles, and objects in the city, models the multi-scale unimodal data, and constructs a spatiotemporal coupling relationship diagram; the risk disturbance refers to the sudden, periodic, or long-term trend disturbance phenomenon encountered; based on the spatiotemporal coupling relationship diagram, the spatiotemporal coupling relationship diagram is processed using a feature fusion network system to generate interpolated data to form enhanced unimodal data, and multiple enhanced unimodal data are coupled into cross-modal data;

[0032] The data compensation module of the cross-modal association network: uses feature alignment technology to adjust the feature dimensions and scales in the cross-modal data, eliminates the differences between the cross-modal data, and obtains the cross-modal data after feature alignment; uses feature migration technology to obtain the feature association relationship of each single modal data between the cross-modal data; uses graph fusion feature construction technology to process the feature association relationship of each single modal data between the cross-modal data, and obtains the correlation and complementarity between the cross-modal data; finally, based on enhancing the correlation and complementarity between the cross-modal data, multi-scale cross-modal data enhancement is achieved to complete cross-modal data compensation.

[0033] In a third aspect, the present invention further provides an electronic device, comprising a processor and a memory, wherein:

[0034] Memory for storing computer programs;

[0035] The processor is used to execute the computer program stored in the memory, and when executed, implements the method described in the first aspect or the system described in the second aspect.

[0036] In a third aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed, implements the method described in the first aspect or the system described in the second aspect.

[0037] The advantages of the present invention compared with the prior art are:

[0038] The present invention provides a multi-scale cross-modal data enhancement method under risk disturbance, including two parts: data enhancement based on multi-spatiotemporal scale coupling and data compensation based on cross-modal association network. The data enhancement part based on multi-spatiotemporal scale coupling can comprehensively and effectively process the relationship between data of different scales and modalities, comprehensively utilize risk disturbance factors, and improve the diversity and expression ability of data; the data compensation part based on cross-modal association network adopts feature fusion, alignment, migration and other technologies to realize the integrated processing of data enhancement and compensation, making data processing more comprehensive and efficient. By modeling and processing the feature correlation relationship of cross-modal data, the correlation between data is clearer, the robustness and stability of data enhancement are improved, and the data enhancement is more comprehensive and accurate, which can better cope with the challenges of risk disturbance and provide a better foundation for further data analysis and application. In addition, the present invention has strong practicality and innovation in processing multi-scale cross-modal data enhancement under risk disturbance, can effectively cope with complex data environments, and provide a richer and more reliable information basis for data analysis and decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0040] Figure 1 This is a flow chart of a multi-scale cross-modal data enhancement method under risk disturbance according to an embodiment of the present invention;

[0041] Figure 2 A flowchart for constructing a spatiotemporal coupling relationship diagram in an embodiment of the present invention;

[0042] Figure 3 This is a schematic diagram of the structure of a feature fusion network system in an embodiment of the present invention;

[0043] Figure 4 Flowchart for implementing the data compensation step of the cross-modal association network in an embodiment of the present invention;

[0044] Figure 5 This is a flowchart of the feature alignment technology implementation in an embodiment of the present invention;

[0045] Figure 6 This is a flowchart of the feature migration technology implementation in an embodiment of the present invention;

[0046] Figure 7 This is a flowchart for implementing a cross-modal feature construction technology based on graph fusion in an embodiment of the present invention;

[0047] Figure 8 This is a block diagram of a multi-scale cross-modal data enhancement system under risk disturbance according to an embodiment of the present invention. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.

[0049] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but merely represents selected embodiments of the invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.

[0050] The present invention is described in detail below with reference to the embodiments.

[0051] like Figure 1 As described above, an embodiment of the present invention provides a multi-scale cross-modal data enhancement method under risk disturbance, which takes risk disturbance factors into account during the data enhancement process. Compared with traditional methods, it solves or alleviates the problem of risk disturbance during the data enhancement process. This technology uses the cross-connection and interaction of multi-source data and multimodal data to enhance the ability to resist risk disturbance. The robustness and stability of data enhancement can be improved by fusing and associating multimodal data. The comprehensive use of this multimodal data makes data enhancement more comprehensive and accurate, and can better cope with the challenges of risk disturbance.

[0052] The method provided by the embodiment of the present invention includes a data enhancement step based on multi-spatiotemporal scale coupling and a data compensation step based on a cross-modal association network, wherein:

[0053] (1) Data enhancement steps based on multi-spatiotemporal scale coupling

[0054] Currently, urban agglomeration traffic data is subject to a range of issues, including localized omissions, sparse sampling, noise perturbations, and unclear patterns, influenced by various factors, including data collection equipment and environmental factors. The collected data also suffers from sparsity in both temporal and regional data on the movement of people, vehicles, and objects within urban agglomerations. However, these movements also exhibit certain spatiotemporal correlations, which can be exploited to achieve data enhancement through multi-scale coupling.

[0055] Under risk disturbances, multi-scale unimodal data on the temporal and spatial movement of people, vehicles, and objects in the city are collected, modeled, and a spatiotemporal coupling relationship diagram is constructed. The risk disturbance refers to a sudden, periodic, or long-term trend disturbance phenomenon encountered. Based on the spatiotemporal coupling relationship diagram, a feature fusion network system is used to process the spatiotemporal coupling relationship diagram to generate interpolated data, forming enhanced unimodal data. Multiple enhanced unimodal data are coupled into cross-modal data.

[0056] Cross-modal data includes images, videos, audio, text, voice, video, GPS, radar, Equal cross-modal data.

[0057] like Figure 2 As shown in the figure, multi-scale single-modal data is modeled and a spatiotemporal coupling relationship diagram is constructed. Figure 2 In [1], each slice represents the spatial information of each node at a certain moment. The spatiotemporal information of nodes with missing data is supplemented by the spatiotemporal dependency of nodes normally collected in each slice. t is the current moment, and T is the sliding time window. The specific implementation is as follows:

[0058] To model multi-scale unimodal data, it is necessary to first abstract the multi-scale unimodal data and obtain the mathematical connections within and between the data.

[0059] Therefore, the first step of the data enhancement step based on multi-spatiotemporal-scale coupling is to construct a spatiotemporal coupling relationship graph.

[0060] The multi-scale single-modal data of a person-vehicle-object movement behavior is recorded as a data source , the data source The relationship diagram is recorded as a graph , multiple data sources at the same time Composition diagram A collection of nodes , data source A diagram of the relationship between urban space and the surrounding environment The edge set ,in is the node representing the data source, is the number of nodes, is the edge representing the data source association relationship, is the number of edges, which reflects the data source Spatial correlation between data sources The historical data at different times are different. time The historical data is recorded as , then one The collection can reflect the time relevance of the data source;

[0061] Considering that the human-vehicle-object mobility behavior data has strong spatial and temporal correlations, in the real-world spatial network, the human-vehicle-object mobility behavior of a node is not only spatially correlated with the human-vehicle-object mobility behavior of other nodes at the current moment, but also temporally correlated with the historical data of the node.

[0062] Leveraging data sources and pictures Complete cross-modal complex network modeling of multi-scale single-modal data and construct a spatiotemporal coupling relationship diagram, which is implemented as follows: Read each pair of adjacent nodes in the connected edges ,definition is the traffic flow characteristic of a node, where Indicates that the node is The characteristics of the moment, Indicates that the node is Characteristics of a moment, and so on; definition is the traffic flow characteristic of another node, where Indicates that the node is The characteristics of the moment, Indicates that the node is The traffic flow characteristics refer to the quantitative changes of pedestrian and vehicle flows under different conditions and their mutual relationships. is a sliding time window; When the value is short enough, the graph The same behavior states between adjacent nodes at the same time are considered to be linearly related. The Pearson correlation coefficient is used to define the relationship between adjacent nodes. and another node The Pearson correlation of traffic flow characteristics between is:

[0063] ;

[0064] in, The value range is If there is no edge connection between two nodes, the node correlation is weak. Therefore, only the correlation between neighboring nodes is considered, and the negative correlation is not considered. The correlation between nodes after taking the absolute value is defined as:

[0065] ;

[0066] After taking the absolute value of the correlation, the space-time coupling relationship graph is constructed based on the obtained correlation through the threshold Gaussian kernel function, and its adjacency matrix is ​​defined as:

[0067] :

[0068] in, Adjacency matrix representing the spatiotemporal coupling graph The Rank Elements of the column, is the standard deviation of node correlation, To control the threshold, is the set of edges, For nodes The edge between them (if any).

[0069] Calculate and obtain the adjacency matrix of the spatiotemporal coupling relationship graph , with the sliding time window The movement of the node features input each time is constantly changing, which causes the correlation between nodes to change accordingly. Therefore, the adjacency matrix changes with the sliding time window. The movement changes dynamically over time.

[0070] The innovations and advantages of modeling multi-scale unimodal data and constructing a spatiotemporal coupling relationship graph are as follows: First, it utilizes multi-scale spatiotemporal coupling modeling. By establishing a spatiotemporal coupling relationship graph, cross-modal complex network modeling of multi-scale unimodal data is achieved, thereby enabling more comprehensive analysis of inter-data correlations. By comprehensively considering spatial and temporal correlations, comprehensive modeling and enhancement of multi-scale unimodal data are achieved, which is not possible with traditional methods. Second, it utilizes traffic flow characteristics. By analyzing traffic flow characteristics using the Pearson correlation coefficient and Gaussian kernel function, the changing patterns of pedestrian and vehicle flows under different conditions and their interrelationships are described. This helps to better understand the correlations between data in subsequent steps and improve the accuracy of data analysis. Third, it addresses dynamic spatiotemporal correlations by using a dynamically changing adjacency matrix to store information. This matrix can more accurately reflect the spatiotemporal correlations between data over time, making data enhancement more accurate and effective. The used adjacency matrix changes dynamically over time, enabling this technology to better cope with data changes at different time scales and provide greater real-time and adaptability.

[0071] After obtaining the spatiotemporal coupling relationship graph represented by an adjacency matrix, deep learning is used to extract features from it, thereby revealing the deeper information within the data. Deep learning networks can extract complex feature representations from raw data through multiple layers of nonlinear transformations. The spatiotemporal coupling relationship graph contains rich spatiotemporal information, from which deep learning networks can extract more representative features.

[0072] like Figure 3 As shown, the feature fusion network system includes a graph convolutional neural network (GCN) module, a bidirectional long short-term memory neural network (BiLSTM) module, a feature fusion module, and an autoencoder network module;

[0073] First, the graph convolutional neural network (GCN) module extracts the topological relationship in the spatiotemporal coupling relationship graph to obtain the spatial dependency between each person-vehicle-object movement behavior and other person-vehicle-object movement behaviors, and then generates the spatial features of the person-vehicle-object movement behaviors based on the spatial dependency. and its characteristic matrix , the extracted spatial feature vector is:

[0074] ;

[0075] The Bidirectional Long Short-Term Memory Neural Network (BiLSTM) module then extracts the temporal relationship in the spatiotemporal coupling relationship graph to obtain the temporal dependency of each person-vehicle-object movement behavior and other person-vehicle-object movement behaviors, and generates the temporal features of the person-vehicle-object movement behavior based on the temporal dependency. and the feature matrix , the extracted time feature vector is:

[0076] ;

[0077] Afterwards, if Figure 3 As shown in the figure, the feature fusion module is used to fuse the spatial and temporal features of each person-vehicle-object movement behavior obtained by the above two modules to obtain the fused cross-modal features:

[0078] ;

[0079] The fused cross-modal features and the original multi-scale single-modal data of human-vehicle-object movement behaviors in time and space are input into the autoencoder network module in sequence form;

[0080] A temporal and spatial attention layer is inserted into the neural network layer between the encoder and the decoder in the autoencoder network module to implement an attention mechanism. The attention mechanism calculates the attention weight of the data with strong temporal and spatial correlation according to the data of the input sequence and its context data, determines the location of the missing data, and interpolates the multi-scale unimodal data.

[0081] The innovations and advantages of the aforementioned feature fusion network system lie in the following: First, the GCN and BiLSTM modules can be used to extract topological relationships in the spatiotemporal coupling graph, helping to capture the spatial and temporal dependencies between human, vehicle, and object movements. Through the GCN and BiLSTM modules, the feature fusion network system can more effectively model spatiotemporal coupling relationships and extract spatial and temporal dependencies, making data processing more comprehensive and consistent. Second, the feature fusion module fuses spatial and temporal features to generate fused cross-modal features, which helps comprehensively consider spatiotemporal information. Finally, the autoencoder network module reconstructs features through an encoder and decoder structure, while simultaneously inserting temporal and spatial attention layers to calculate attention weights and interpolate data with strong spatiotemporal correlations, further improving data integrity and accuracy. This feature network system utilizes the feature fusion and autoencoder modules to achieve feature fusion of data at different scales and interpolation of missing data, thereby improving data integrity and quality. This feature network system also incorporates an attention mechanism for processing data with strong spatiotemporal correlations, better capturing the relevance and importance of data, further enhancing data processing effectiveness. Compared with existing technologies, this feature network system has better performance in data enhancement and cross-modal data processing, can more effectively process multi-scale data under risk disturbances, and improve the robustness and generalization ability of the model.

[0082] (2) Data compensation step of cross-modal association network

[0083] In order to achieve the collaborative work between multi-source and cross-modal data, it is necessary to propose a data compensation technology of cross-modal association network. This technology extracts cross-modal features of heterogeneous data to establish a cross-modal association network, so that data of different modalities can work together and provide a more comprehensive and accurate data foundation for data-driven urban agglomeration applications. The interpolated multi-scale single-modal data obtained in the previous step is used to realize cross-modal association in this step.

[0084] like Figure 4 As shown in the figure, two different local areas within a region can each collect multimodal data including video, GPS, radar, and 5G. However, the data collected in each location may be missing data from a certain modality. The data compensation step of the cross-modal association network performs feature alignment, feature transfer, and feature fusion on this multimodal data to obtain fused features for subsequent data analysis of transportation, population, commercial and residential needs, and so on.

[0085] Using feature alignment technology, the feature dimensions and scales in the cross-modal data are adjusted to eliminate the differences between the cross-modal data and obtain the feature-aligned cross-modal data; the feature migration technology is used on the feature-aligned cross-modal data to obtain the feature correlation relationship of each single modal data between the cross-modal data; the feature construction technology of graph fusion is used to process the feature correlation relationship of each single modal data between the cross-modal data to obtain the correlation and complementarity between the cross-modal data; finally, based on the enhanced correlation and complementarity between the cross-modal data, multi-scale cross-modal data enhancement is achieved and cross-modal data compensation is completed.

[0086] like Figure 5 As shown in the figure, during cross-modal fusion, the features of data from different modalities are not uniform. Aligning these features is necessary to complete and unify the data and facilitate subsequent data supplementation and unified operations. Cross-modal feature alignment is a key technology in cross-modal fusion and is widely used in cross-modal tasks. This technology can adjust the dimension and scale of features from different data sources.

[0087] The feature alignment technology uses unsupervised display cross-modal feature alignment technology, which is implemented as follows:

[0088] like Figure 5 As shown in the process, this technology uses an unsupervised method to process the input cross-modal data, and then inputs the data into the dynamic Bayesian network. Unsupervised methods do not require labeled data, but rely on algorithms to automatically discover the correlation between data. They are suitable for situations where the data volume is large, calibration is difficult, and there is no direct alignment of supervised labels. Dynamic Bayesian networks are a widely used unsupervised explicit alignment method. By measuring the correlation between two sequences, they find the best match and achieve voice, video, GPS, radar, Cross-modal alignment. Use an explicit alignment algorithm based on dynamic Bayesian networks to process cross-modal data and achieve feature alignment of cross-modal data. Dynamic Bayesian networks are usually composed of multiple time slices, each of which contains a set of random variables. These random variables are represented as , in Represents the index of the time slice. Within each time slice, the dependency between random variables is represented by a Bayesian network. The joint probability distribution of the dynamic Bayesian network is shown in the formula.

[0089] ;

[0090] in, Indicates from time slice 1 to time slice All random variables of represents the probability distribution of the initial state, Represents the conditional probability distribution of the current time slice state given the state of the previous time slice.

[0091] The explicit alignment algorithm based on dynamic Bayesian network uses multiple time slices that constitute the dynamic Bayesian network to store cross-modal data, and then uses the joint probability of the Bayesian network to automatically obtain the conditional probability of the current time slice state, thereby achieving feature alignment of cross-modal data. Figure 5 As shown in the process, the dynamic Bayesian network obtains aligned data after storing and aligning the features.

[0092] Compared to existing technologies, the feature alignment technology employed above offers the following innovations and advantages: First, it utilizes an unsupervised explicit cross-modal feature alignment technique. This unsupervised approach to processing cross-modal data eliminates the need for large amounts of labeled data, reducing the cost and complexity of data labeling. Unsupervised feature alignment also reduces the need for large amounts of labeled data, saving time and costs. Second, it utilizes an explicit alignment algorithm based on a dynamic Bayesian network. By processing cross-modal data using a dynamic Bayesian network, data can be stored in multiple time slices, automatically deriving the conditional probability of the current time slice state, thereby achieving feature alignment across the modal data. This method better captures the correlations and dynamic changes between data, improving the accuracy of data alignment. The explicit alignment algorithm based on the dynamic Bayesian network can better handle the correlations between cross-modal data, improving the accuracy of data alignment and thus enhancing model performance. Third, it utilizes dynamic data processing. By processing data across multiple time slices using a dynamic Bayesian network, dynamic data changes can be better captured, enabling the model to more quickly adapt to data changes in real applications.

[0093] like Figure 6 As shown in the figure, feature transfer is a key technology in cross-modal fusion. It can help effectively combine information from different modalities to improve the performance and representation of cross-modal tasks. Cross-modal data usually contains rich complementary information. Through feature transfer, the key information in each modality can be extracted and mapped to a shared feature space, so as to make full use of the information of different modalities during fusion, improve and supplement the missing information, and achieve information complementarity. At the same time, it can map high-dimensional data to a low-dimensional shared feature space, thereby alleviating the curse of dimensionality and improving the efficiency and generalization ability of the model. In addition, feature transfer can help the model learn robust features from one modality and then apply them to another modality, thereby enhancing the robustness of the cross-modal system, making it more resistant to data changes and noise, and better able to cope with noisy situations.

[0094] Combine Figure 6 The above feature migration technology is implemented as follows:

[0095] (1) The first step is to perform data complementation on the cross-modal data after feature alignment to fill in the missing modal data in each group of cross-modal data; the data complementation refers to reading the missing modal data in the current group of cross-modal data from other groups of cross-modal data and filling in the modalities of the current group of cross-modal data to achieve data complementation and obtain the complemented cross-modal data;

[0096] (2) In the second step, data enhancement is performed on the cross-modal data after data complementation. The robust features of a certain modality in the cross-modal data are extracted and learned from each group of cross-modal data. The learned single modality features are applied to the other modality to obtain enhanced single modality data in the cross-modal data.

[0097] (3) The third step is to realize the feature dimensionality reduction mapping function, extract the key information of each single modal data in each set of cross-modal data, map the key information into the shared feature space, and perform feature fusion to obtain the correlation relationship between each single modal feature in the cross-modal data;

[0098] In (1)-(3), the maximum mean discrepancy (MMD) value is used for normalization to achieve the unification of different but related data, so as to facilitate comparison between data and further feature migration and fusion. The MMD is mainly used to measure the distance between two different but related data distributions. Its basic definition is shown in the formula:

[0099] ;

[0100] Among them, sup represents the upper bound of the function. and Respectively expressed in the distribution and expectations, represents the mapping function between samples, To limit The value is , such normalization helps unify the data and facilitates comparison. Through a series of mathematical operations, the mathematical simplified expression of MMD can be obtained, as shown in the following formula.

[0101] ;

[0102] in, is the kernel function, and is obtained by summing the functions in the original sample space. The result of calculation is equal to The inner product in the feature space. Through MMD, the distribution difference between the source domain and the target domain can be measured, thereby helping to achieve domain adaptation and feature migration, helping to achieve effective data fusion and comparison, and improving the performance of the target task. is a set of functions, There are 2 distributions. For distribution The sample set of is the number of samples in the sample set; For distribution The sample set of is the number of samples in the sample set.

[0103] like Figure 7 As shown in the figure, based on feature migration, a cross-modal feature construction technology based on graph fusion is implemented. This technology analyzes feature consistency and complementarity, and uses fused features to describe the spatiotemporal operation status of people, vehicles, and objects in the city.

[0104] The cross-modal feature construction technology based on graph fusion is:

[0105] (1) First, a feature graph is created for each single-modal data. Each node in the graph represents each feature contained in the modality, and the edge represents the connection or similarity between features. The weight of the edge is determined by the similarity between features. These weights form a weight matrix. The weight of the edge is calculated based on the similarity between features, as shown in the following formula:

[0106] ;

[0107] in, For the picture, is a node set, is an edge set, which is used to represent the relationship structure in the feature fusion process. Represents the edge weight matrix, which is used to represent the strength of association between different nodes. for The number of midpoints, are different feature points of the same mode, Express The characteristic gap is calculated by summing up.

[0108] (2) Then merge the nodes and edges in each single-modal feature graph, and integrate the feature graphs of different modalities into a cross-modal feature graph. In this graph, a graph attention network is constructed so that each node gathers information from other modalities to achieve information transmission, further update the nodes and weights in the weight matrix, and finally obtain the correlation and complementarity between cross-modal data. This is shown in the following formula:

[0109] ;

[0110] It further adjusts the weights on the original basis. is the updated edge weight matrix, is the original edge weight matrix, is a scalar coefficient used to adjust the influence of node differences. This increases the degree of correlation between nodes and meets the needs of feature fusion. During the information transfer process, the weights of nodes and edges are updated during the information transfer process to achieve the mutual fusion of different modal features. are different feature points of the same mode, .

[0111] Compared to existing technologies, the above-mentioned cross-modal feature construction technology based on graph fusion offers the following innovations and advantages: First, it utilizes graph fusion technology to construct a cross-modal feature graph, integrating features from different modalities. This technology effectively integrates multimodal data and captures inter-data relationships, achieving comprehensive feature representation. Through graph fusion, features from different modalities can be fused into a single cross-modal feature graph, avoiding the information loss between single-modal features in traditional methods. Second, it utilizes a graph attention network to achieve information transfer and updates. Through the graph attention mechanism, each node aggregates information from nodes in other modalities, thereby updating the weights in the node and weight matrix, achieving information complementation and transfer. This information transfer facilitates cross-modal data complementation and transfer, facilitating the integration of features from various modalities, improving data connectivity and representation capabilities. Third, it integrates global features. By integrating single-modal feature graphs to construct a cross-modal feature graph, feature information from various modalities can be more comprehensively integrated, resulting in a more comprehensive data representation. Compared with the existing technology, it comprehensively considers the connection and similarity between different modal features, and further combines the graph attention network to realize information transmission and update, so as to globally capture the correlation and complementarity between cross-modal data, which helps to realize information transmission and communication between multimodal data and improve the overall data representation effect.

[0112] like Figure 8 As shown, an embodiment of the present invention further provides a multi-scale cross-modal data enhancement system under risk disturbance, including a data enhancement module based on multi-spatiotemporal scale coupling and a data compensation module based on a cross-modal association network, wherein:

[0113] The data enhancement modality module based on multi-spatiotemporal scale coupling: under risk disturbance, collects multi-scale unimodal data on the temporal and spatial movement behaviors of people, vehicles, and objects in the city, models the multi-scale unimodal data, and constructs a spatiotemporal coupling relationship diagram; the risk disturbance refers to the sudden, periodic, or long-term trend disturbance phenomenon encountered; based on the spatiotemporal coupling relationship diagram, the spatiotemporal coupling relationship diagram is processed using a feature fusion network system to generate interpolated data to form enhanced unimodal data, and multiple enhanced unimodal data are coupled into cross-modal data;

[0114] The data compensation module of the cross-modal association network: uses feature alignment technology to adjust the feature dimensions and scales in the cross-modal data, eliminates the differences between the cross-modal data, and obtains the cross-modal data after feature alignment; uses feature migration technology to obtain the feature association relationship of each single modal data between the cross-modal data; uses graph fusion feature construction technology to process the feature association relationship of each single modal data between the cross-modal data, and obtains the correlation and complementarity between the cross-modal data; finally, based on enhancing the correlation and complementarity between the cross-modal data, multi-scale cross-modal data enhancement is achieved to complete cross-modal data compensation.

[0115] The execution process of various modules in the above system is similar to the method process provided in the embodiment of the present invention.

[0116] Based on the same inventive concept, another embodiment of the present invention also provides an electronic device (including a computer, a server, a smart phone, etc.), which includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes a module for executing each step and system in the method provided above in the embodiment of the present invention.

[0117] Based on the same inventive concept, another embodiment of the present invention further provides a computer-readable storage medium (such as ROM / RAM, disk, or CD), which stores a computer program. When the computer program is executed by a computer, it implements the various steps of the method and system modules provided in the above embodiment of the present invention.

[0118] The above embodiments are provided for the purpose of describing the present invention only and are not intended to limit the scope of the present invention. The scope of the present invention is defined by the appended claims. Various equivalent substitutions and modifications made without departing from the spirit and principles of the present invention are intended to be within the scope of the present invention.

Claims

1. A multi-scale cross-modal data augmentation method under risk disturbance, characterized by: It includes a data enhancement step based on multi-spatiotemporal scale coupling and a data compensation step based on a cross-modal association network, where: The data enhancement step based on multi-spatiotemporal scale coupling is as follows: under risk disturbance, multi-scale single-modal data of the movement behavior of people, vehicles and objects in the city in time and space are collected, the multi-scale single-modal data are modeled, and a spatiotemporal coupling relationship diagram is constructed; the risk disturbance refers to the sudden, periodic or long-term trend disturbance phenomenon encountered; based on the spatiotemporal coupling relationship diagram, the spatiotemporal coupling relationship diagram is processed by a feature fusion network system to generate interpolated data to form enhanced single-modal data, and multiple enhanced single-modal data are coupled into cross-modal data; the feature fusion network system extracts the topological relationship in the spatiotemporal coupling relationship diagram, generates the spatial features of the movement behavior of people, vehicles and objects, and extracts the temporal relationship to obtain the temporal features; the spatial features and temporal features of each person's movement behavior of vehicles and objects are then fused to obtain the fused cross-modal features; the fused cross-modal features and the original multi-scale single-modal data of the movement behavior of people, vehicles and objects in time and space are input into the autoencoder network module for interpolation; The data compensation step of the cross-modal association network is as follows: using feature alignment technology to adjust the feature dimensions and scales in the cross-modal data, eliminating the differences between the cross-modal data, and obtaining the feature-aligned cross-modal data; using feature migration technology on the feature-aligned cross-modal data to obtain the feature association relationship of each single modal data between the cross-modal data; using graph fusion feature construction technology to process the feature association relationship of each single modal data between the cross-modal data, and obtain the correlation and complementarity between the cross-modal data; finally, based on enhancing the correlation and complementarity between the cross-modal data, multi-scale cross-modal data enhancement is achieved to complete cross-modal data compensation; The feature alignment technology adopts an unsupervised dynamic Bayesian network-based display cross-modal feature alignment technology; The feature migration technology is implemented by performing data complementation on the cross-modal data after feature alignment, performing data enhancement on the cross-modal data after data complementation, and then obtaining the correlation relationship between each single-modal feature in the cross-modal data through feature dimensionality reduction mapping; The cross-modal feature construction technology based on graph fusion creates a feature map for each single modal data, integrates the feature maps of different modalities into a cross-modal feature map, and constructs a graph attention network in the feature map, ultimately obtaining the correlation and complementarity between cross-modal data; The cross-modal data includes images, videos, audio, text, voice, video, GPS and radar.

2. The multi-scale cross-modal data enhancement method under risk disturbance according to claim 1, characterized in that: In the data enhancement step based on multi-spatiotemporal scale coupling, the multi-scale single-modal data is modeled and a spatiotemporal coupling relationship diagram is constructed as follows: The multi-scale single-modal data of a person-vehicle-object movement behavior is recorded as a data source , the data source The relationship diagram is recorded as a graph , multiple data sources at the same time Composition diagram A collection of nodes , data source A diagram of the relationship between urban space and the surrounding environment The edge set , which reflects the data source Spatial correlation between data sources The historical data at different times are different, so The historical data is recorded as , then one The collection can reflect the time relevance of the data source; Leveraging data sources and pictures Complete cross-modal complex network modeling of multi-scale single-modal data and construct a spatiotemporal coupling relationship diagram, which is implemented as follows: Read each pair of adjacent nodes in the connected edges ,definition is the traffic flow characteristic of a node, is the traffic flow characteristic of another node, which refers to the quantitative change pattern of pedestrian and vehicle flows under different conditions and their mutual relationship. is a sliding time window; When the value is short enough, the graph The same behavior states between adjacent nodes at the same time are considered to be linearly correlated. The Pearson correlation coefficient is used to define the relationship between adjacent nodes. The Pearson correlation of traffic flow characteristics between adjacent nodes is calculated. Based on the obtained correlation, the spatiotemporal coupling relationship graph is constructed through the threshold Gaussian kernel function, and the adjacency matrix of the spatiotemporal coupling relationship graph is calculated and obtained. , the adjacency matrix changes with the sliding time window The movement changes dynamically over time.

3. The multi-scale cross-modal data enhancement method under risk disturbance according to claim 1, characterized in that: In the data enhancement step based on multi-spatiotemporal scale coupling, the feature fusion network system includes a graph convolutional neural network module, a bidirectional long short-term memory neural network module, a feature fusion module and an autoencoder network module; First, the graph convolutional neural network module extracts the topological relationships in the spatiotemporal coupling relationship graph to obtain the spatial dependencies between each person-vehicle-object movement behavior and other person-vehicle-object movement behaviors. Based on these spatial dependencies, the spatial features of the person-vehicle-object movement behaviors are generated. The bidirectional long short-term memory neural network then extracts the temporal relationship in the spatiotemporal coupling relationship graph to obtain the temporal dependency of each person-vehicle-object movement behavior on other person-vehicle-object movement behaviors. Based on this temporal dependency, the temporal features of the person-vehicle-object movement behaviors are generated. Afterwards, the feature fusion module fuses the spatial and temporal features of each person-vehicle-object movement behavior to obtain the fused cross-modal features; The fused cross-modal features and the original multi-scale single-modal data of human-vehicle-object movement behaviors in time and space are input into the autoencoder network module in sequence form; A temporal and spatial attention layer is inserted into the neural network layer between the encoder and the decoder in the autoencoder network module to implement an attention mechanism. The attention mechanism calculates the attention weight of the data with strong temporal and spatial correlation according to the data of the input sequence and its context data, determines the location of the missing data, and interpolates the multi-scale unimodal data.

4. The multi-scale cross-modal data enhancement method under risk disturbance according to claim 1, characterized in that: In the data compensation step of the cross-modal association network, the feature alignment technology adopts an unsupervised explicit cross-modal feature alignment technology, which is implemented as follows: an unsupervised method is used to process cross-modal data using an explicit alignment algorithm based on a dynamic Bayesian network to achieve feature alignment of cross-modal data; the explicit alignment algorithm based on the dynamic Bayesian network uses multiple time slices that constitute the dynamic Bayesian network to store cross-modal data, and then uses the joint probability of the dynamic Bayesian network to automatically obtain the conditional probability of the current time slice state, thereby achieving feature alignment of cross-modal data; The distribution of the joint probability of the dynamic Bayesian network is shown in the following formula: ; in, Indicates from time slice 1 to time slice All random variables of represents the probability distribution of the initial state, Represents the conditional probability distribution of the current time slice state given the state of the previous time slice.

5. The multi-scale cross-modal data enhancement method under risk disturbance according to claim 1, characterized in that: In the data compensation step of the cross-modal association network, the feature migration technology is implemented as follows: (1) First, data complementation is performed on the cross-modal data after feature alignment to complete the missing modal data in each group of cross-modal data. The data complementation refers to reading the missing modal data in the current group of cross-modal data from other groups of cross-modal data and completing the modality of the current group of cross-modal data to achieve data complementation. (2) Secondly, data enhancement is performed on the cross-modal data after data complementation. The robust features of a certain modality in the cross-modal data are extracted and learned from each group of cross-modal data, and the learned single modality features are applied to the other modality. (3) Finally, the feature dimensionality reduction mapping function is implemented to extract the key information of each single-modal data in each set of cross-modal data, map the key information into the shared feature space, and perform feature fusion to obtain the correlation relationship between each single-modal feature in the cross-modal data; In (1)-(3) above, the maximum mean difference value is used for normalization to achieve the unification of different but related data, so as to facilitate comparison between data and further feature migration and fusion.

6. The multi-scale cross-modal data enhancement method under risk disturbance according to claim 1, characterized in that: In the data compensation step of the cross-modal association network, the cross-modal feature construction technology based on graph fusion is: (1) First, a feature graph is created for each single-modal data. Each node in the graph represents each feature contained in the modality, and the edge represents the connection or similarity between features. The weight of the edge is determined by the similarity between features. These weights form a weight matrix. The weight of the edge is calculated based on the similarity between features, as shown in the following formula: ; in, For the picture, is a node set, is an edge set, which is used to represent the relationship structure in the feature fusion process. Represents the edge weight matrix, which is used to represent the strength of association between different nodes. for The number of midpoints, are different feature points of the same mode, Express The characteristic gap is summed up; (2) The nodes and edges in each single-modal feature graph are then merged, and the feature graphs of different modalities are integrated into a cross-modal feature graph. A graph attention network is constructed in this graph, so that each node aggregates information from other modalities to achieve information transmission, and the weights in the nodes and weight matrix are further updated. Finally, the correlation and complementarity between cross-modal data are obtained, as shown in the following formula: ; is the updated edge weight matrix, is the original edge weight matrix, is a scalar coefficient used to adjust the influence of node differences, thereby increasing the degree of correlation between nodes and adapting to the needs of feature fusion; in the process of information transmission, the weights of nodes and edges are updated during the information transmission process to achieve the fusion of different modal features. are different feature points of the same mode, .

7. A multi-scale cross-modal data augmentation system under risk disturbance, characterized by: It includes a data enhancement module based on multi-spatiotemporal scale coupling and a data compensation module based on a cross-modal association network, where: The data enhancement modal module based on multi-spatiotemporal scale coupling: under risk disturbance, collects multi-scale single-modal data of the movement behavior of people, vehicles and objects in the city in time and space, models the multi-scale single-modal data, and constructs a spatiotemporal coupling relationship diagram; the risk disturbance refers to the sudden, periodic or long-term trend disturbance phenomenon encountered; based on the spatiotemporal coupling relationship diagram, uses the feature fusion network system to process the spatiotemporal coupling relationship diagram, generates interpolated data, forms enhanced single-modal data, and couples multiple enhanced single-modal data into cross-modal data; the feature fusion network system extracts the topological relationship in the spatiotemporal coupling relationship diagram, generates the spatial features of the movement behavior of people, vehicles and objects, extracts the time series relationship to obtain the time series features; then fuses the spatial features and time series features of each person-vehicle-object movement behavior to obtain the fused cross-modal features; the fused cross-modal features and the original multi-scale single-modal data of the movement behavior of people, vehicles and objects in time and space are input into the autoencoder network module for interpolation data; The data compensation module of the cross-modal association network: uses feature alignment technology to adjust the feature dimensions and scales in the cross-modal data, eliminates the differences between the cross-modal data, and obtains the feature-aligned cross-modal data; uses feature migration technology to obtain the feature association relationship of each single modal data between the cross-modal data; uses graph fusion feature construction technology to process the feature association relationship of each single modal data between the cross-modal data, and obtains the correlation and complementarity between the cross-modal data; finally, based on the enhanced correlation and complementarity between the cross-modal data, multi-scale cross-modal data enhancement is achieved to complete cross-modal data compensation; The feature alignment technology adopts an unsupervised dynamic Bayesian network-based display cross-modal feature alignment technology; The feature migration technology is implemented by performing data complementation on the cross-modal data after feature alignment, performing data enhancement on the cross-modal data after data complementation, and then obtaining the correlation relationship between each single-modal feature in the cross-modal data through feature dimensionality reduction mapping; The cross-modal feature construction technology based on graph fusion creates a feature map for each single modal data, integrates the feature maps of different modalities into a cross-modal feature map, and constructs a graph attention network in the feature map, ultimately obtaining the correlation and complementarity between cross-modal data; The cross-modal data includes images, videos, audio, text, voice, video, GPS and radar.

8. An electronic device, characterized in that: comprising a processor and a memory, wherein: Memory for storing computer programs; A processor, configured to execute a computer program stored in a memory, and to implement the method of any one of claims 1 to 6, or the system of claim 7, when executed.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the method according to any one of claims 1 to 6 or the system according to claim 7 is implemented.

Citation Information

Patent Citations

  • Cross-modal pedestrian re-identification method based on modal specific memory network

    CN114882525A

  • Space-time interaction prediction method and system for risk disturbance and individual behaviors

    CN118313513A