Multi-modal data intelligent clustering method, system and device based on spatial-temporal characteristics and medium

By constructing multimodal data vectors and introducing space-time constraint optimization clustering centers, dynamically adjusting density thresholds, the problems of inaccurate clustering and insufficient adaptability in multimodal data analysis are solved, and efficient and accurate multi-dimensional data clustering and hot spot discovery are achieved.

CN120541555APending Publication Date: 2025-08-26浪潮智慧城市科技有限公司

Patent Information

Application Number
CN202510586338.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The prior art is difficult to fully integrate multi-dimensional information such as text, geography and time in multi-modal data analysis, resulting in inaccurate clustering results and inability to discover hot issues in time, and unable to dynamically adapt to data changes.

Method used

By constructing a combination of text semantic vectors, geocoding vectors, and time-decaying vectors, multimodal data processing is performed using the improved Transformer architecture, and spatial and temporal constraint optimization cluster centers are introduced to dynamically adjust the density threshold to identify multi-grained hotspots.

Benefits of technology

It improves the accuracy and adaptability of clustering, can better mine the spatial and temporal correlation and potential patterns of data, and is suitable for big data analysis in a variety of fields, especially in scenarios where spatiotemporal related data are processed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541555A_ABST
    Figure CN120541555A_ABST
Patent Text Reader

Abstract

The invention discloses a spatio-temporal feature-based multi-modal data intelligent clustering method, system and device and a medium, belongs to the field of data processing and analysis technologies and artificial intelligence cross technologies, and aims to solve the technical problem of how to improve the clustering efficiency, accuracy and adaptability and better meet the actual requirements of multi-modal data analysis. According to the technical scheme, the method comprises the steps of obtaining a multi-modal event, wherein the multi-modal event is obtained, and a text semantic vector, a geocoding vector and a time decay vector are constructed according to the multi-modal event; constructing a multi-modal data vector: representing the multi-modal data as a combination of a text semantic vector, a geocoding vector and a time decay vector; processing of an improved Transform architecture: multi-modal data representation is generated based on the improved Transform architecture, multi-modal data vectors are processed, and multi-dimensional information of text semantics, geography and time is fully fused; optimizing a clustering center by space-time constraint; and adjusting a dynamic density threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of data processing and analysis technology and the cross-technical field of artificial intelligence, and specifically to a method, system, device and medium for intelligent clustering of multimodal data based on spatiotemporal features. Background Art

[0002] With the rapid development of information technology, the generation and dissemination of various types of multimodal data are accelerating. These data come from a wide range of sources and in diverse forms, encompassing multiple modalities such as text, geographic information, and time series. Traditional clustering methods have numerous limitations when analyzing and processing this complex data. On the one hand, clustering based solely on textual semantics fails to fully account for the spatiotemporal properties of the data, resulting in inaccurate and incomplete clustering results and an inability to effectively mine hotspot information with spatiotemporal correlations. On the other hand, faced with massive amounts of multimodal data, existing methods struggle to dynamically adapt to data development and changes, hindering the timely identification of hotspot issues at different granularities. Furthermore, traditional clustering algorithms often fail to fully integrate multidimensional information such as text, geography, and time when processing multimodal data, hindering clustering effectiveness and in-depth understanding of the data.

[0003] Therefore, how to improve the efficiency, accuracy and adaptability of clustering to better meet the actual needs of multimodal data analysis is a technical problem that needs to be solved urgently. Summary of the Invention

[0004] The technical task of the present invention is to provide a method, system, device and medium for intelligent clustering of multimodal data based on spatiotemporal characteristics to solve the problem of how to improve the efficiency, accuracy and adaptability of clustering and better meet the actual needs of multimodal data analysis.

[0005] The technical task of the present invention is achieved in the following manner: a multimodal data intelligent clustering method based on spatiotemporal features, the method is specifically as follows:

[0006] Acquiring multimodal events: Acquire multimodal events and construct text semantic vectors, geocoding vectors, and time decay vectors based on them. The text semantic vector extracts key semantic features of text data using natural language processing techniques. The geocoding vector is generated by encoding the geographic coordinates of the data's location, using the GeoHash tool to convert the geographic coordinates into a fixed-length encoding vector. The time decay vector is calculated based on the interval between the data's occurrence time and the current time using a specific decay function (such as an exponential decay function) to reflect the data's timeliness.

[0007] Constructing a multimodal data vector: Representing multimodal data as a combination of a text semantic vector, a geocoding vector, and a time decay vector, i.e., multimodal data vector = text semantic vector ⊕ geocoding vector ⊕ time decay vector;

[0008] Improved Transformer Architecture Processing: Generates multimodal data representations based on the improved Transformer architecture, processes multimodal data vectors, and fully integrates multi-dimensional information such as textual semantics, geography, and time.

[0009] Optimize cluster centers with spatiotemporal constraints: Optimize the distribution of cluster centers through spatiotemporal constraints, and adjust the cluster centers based on the geographic spatial distribution and time series characteristics of multimodal data;

[0010] Dynamic density threshold adjustment: Dynamically adjust the density threshold to achieve multi-granularity hotspot discovery. Based on the distribution characteristics of multimodal data and real-time feedback information during the clustering process, the density threshold is automatically adjusted to identify hotspot events at different levels.

[0011] As a preferred approach, the improved Transformer architecture achieves efficient processing and fusion of multimodal data by optimizing the network structure of the encoder and decoder, as well as the input embedding layer and self-attention mechanism; the details are as follows:

[0012] Encoder Structure: The encoder is composed of multiple identical encoder layers stacked together. Multimodal fusion and feature extraction are repeated in each encoder layer, deepening the understanding and fusion of the data layer by layer. Each encoder layer includes a multimodal fusion attention layer and an encoding feedforward neural network layer. The multimodal fusion attention layer is used to capture the complex relationships between textual, geographic, and temporal features in the input data. The encoding feedforward neural network layer performs nonlinear transformations on the contextual representation of each data point to further extract high-level representations of the features.

[0013] Decoder structure: The decoder is composed of multiple identical decoder layers stacked together; each decoder layer includes a self-attention layer, a cross-attention layer, and a decoding feedforward neural network layer; among them, the self-attention layer is used to capture the feature relationship in the decoder input data; the cross-attention layer allows the decoder to access the encoder output, thereby realizing information interaction between the encoder and decoder. By calculating the correlation weight between the decoder input and the encoder output, the cross-attention layer can combine the multimodal features extracted by the encoder with the current state of the decoder to further fuse the multimodal information; the decoding feedforward neural network layer performs nonlinear transformation on the fused features to generate the final output representation.

[0014] Input embedding and multimodal fusion: In the input embedding layer, text semantic vectors, geocoding vectors, and time decay vectors are respectively embedded into the corresponding dimensional spaces to form a unified multimodal input representation. This process lays the foundation for subsequent feature fusion, enabling the model to process data from different modalities simultaneously.

[0015] More optimally, the multimodal fusion attention layer dynamically adjusts the weight distribution between features through the self-attention mechanism; specifically, the self-attention mechanism calculates the correlation weight between each data point and other data points, and performs weighted summation of the features according to the correlation weight, thereby generating a contextual representation of each data point.

[0016] Preferably, spatiotemporal constraint optimization of cluster centers refers to introducing spatiotemporal constraints to optimize the distribution of cluster centers during the clustering process; specifically, as follows:

[0017] The geographic coordinate information of the data points is obtained through the geocoding vector, and the geographic distance (Euclidean distance) between the data points is calculated. At the same time, the time information of the data points is obtained through the time decay vector, and the time interval between the data points is calculated;

[0018] In the cluster initialization stage, the initial cluster centers are selected based on the geographic spatial distribution and time series characteristics to make them representative in the spatiotemporal dimension;

[0019] During the clustering iteration process, the update rules of the cluster centers are constrained according to the geographical distance and time interval. For example, the movement range of the cluster centers in the geographical space is limited to prevent the cluster centers from deviating too much from the geographical distribution center of the data points. At the same time, the time weight of the cluster centers is adjusted according to the time series characteristics to enable it to better reflect the time trend of the data points.

[0020] By introducing spatiotemporal constraints, the cluster centers can better represent the characteristics of the corresponding data groups in the spatiotemporal dimensions, thereby improving the rationality and accuracy of the clustering results and avoiding clustering bias caused by ignoring spatiotemporal factors.

[0021] As a preference, the dynamic density threshold is adjusted as follows:

[0022] The initial density threshold is set according to the initial distribution characteristics of the multimodal data and is determined by statistical methods (such as the average density of data points) or empirical rules;

[0023] During the clustering process, the density threshold is dynamically adjusted according to the density distribution of the current clustering results:

[0024] If there are high-density areas in the current clustering results, it indicates that there may be hotspot events. Appropriately lower the density threshold to further subdivide the hotspot areas.

[0025] If the hot spots in the current clustering results are not obvious, increase the density threshold appropriately to merge smaller hot spots.

[0026] By dynamically adjusting the density threshold, hot events with different importance and impact ranges can be identified at different levels. For example, large-scale hot events can be discovered at the global level, and hot issues within a small range or a specific time period can be discovered at the local level, thereby improving the flexibility and practicality of multimodal data clustering analysis.

[0027] A multimodal data intelligent clustering system based on spatiotemporal features, the system is as follows:

[0028] The multimodal event acquisition module is used to acquire multimodal events and construct text semantic vectors, geocoding vectors, and time decay vectors based on these events. The text semantic vector refers to the key semantic features of text data extracted through natural language processing technology. The geocoding vector is generated by encoding the geographic coordinates of the data's location, using the GeoHash tool to convert the geographic coordinates into a fixed-length encoding vector. The time decay vector is calculated based on the interval between the data's occurrence time and the current time using a specific decay function (such as an exponential decay function) to reflect the timeliness of the data.

[0029] A multimodal data vector construction module is used to represent multimodal data as a combination of a text semantic vector, a geocoding vector, and a time decay vector, i.e., multimodal data vector = text semantic vector ⊕ geocoding vector ⊕ time decay vector;

[0030] The multimodal data processing module is used to generate multimodal data representations based on the improved Transformer architecture, process multimodal data vectors, and fully integrate multi-dimensional information such as textual semantics, geography, and time;

[0031] The optimization module is used to optimize the distribution of cluster centers through spatiotemporal constraints and adjust the cluster centers according to the geographic spatial distribution and time series characteristics of multimodal data;

[0032] The dynamic density threshold adjustment module is used to dynamically adjust the density threshold to achieve multi-granularity hotspot discovery. According to the distribution characteristics of multimodal data and real-time feedback information in the clustering process, the density threshold is automatically adjusted to identify hot events at different levels.

[0033] As a preferred approach, the improved Transformer architecture achieves efficient processing and fusion of multimodal data by optimizing the network structure of the encoder and decoder, as well as the input embedding layer and self-attention mechanism; the details are as follows:

[0034] Encoder structure: The encoder is composed of multiple identical encoder layers stacked together. Multimodal fusion and feature extraction are repeated in each encoder layer, deepening the understanding and fusion of the data layer by layer. Each encoder layer includes a multimodal fusion attention layer and an encoding feedforward neural network layer. The multimodal fusion attention layer is used to capture the complex relationship between text, geography, and time features in the input data. The multimodal fusion attention layer dynamically adjusts the weight distribution between features through the self-attention mechanism. Specifically, the self-attention mechanism is used to calculate the correlation weight between each data point and other data points, and the features are weighted and summed according to the correlation weight to generate a contextual representation of each data point. The encoding feedforward neural network layer is used to perform nonlinear transformations on the contextual representation of each data point to further extract high-level representations of the features.

[0035] Decoder structure: The decoder is composed of multiple identical decoder layers stacked together; each decoder layer includes a self-attention layer, a cross-attention layer, and a decoding feedforward neural network layer; among them, the self-attention layer is used to capture the feature relationship in the decoder input data; the cross-attention layer allows the decoder to access the encoder output, thereby realizing information interaction between the encoder and decoder. By calculating the correlation weight between the decoder input and the encoder output, the cross-attention layer can combine the multimodal features extracted by the encoder with the current state of the decoder to further fuse the multimodal information; the decoding feedforward neural network layer performs nonlinear transformation on the fused features to generate the final output representation.

[0036] Input embedding and multimodal fusion: In the input embedding layer, text semantic vectors, geocoding vectors, and time decay vectors are respectively embedded into the corresponding dimensional spaces to form a unified multimodal input representation. This process lays the foundation for subsequent feature fusion, enabling the model to process data from different modalities simultaneously.

[0037] Preferably, the optimization module obtains the geographic coordinate information of the data points through the geocoding vector and calculates the geographic distance (Euclidean distance) between the data points. At the same time, it obtains the time information of the data points through the time decay vector and calculates the time interval between the data points. In the clustering initialization stage, the initial cluster center is selected according to the geographic spatial distribution and time series characteristics to make it representative in the spatiotemporal dimension. In the clustering iteration process, the update rule of the cluster center is constrained according to the geographic distance and time interval, for example, the movement range of the cluster center in the geographic space is restricted to avoid the cluster center from deviating too much from the geographic distribution center of the data point. At the same time, the time weight of the cluster center is adjusted according to the time series characteristics to enable it to better reflect the time trend of the data point. By introducing spatiotemporal constraints, the cluster center can better represent the characteristics of the corresponding data group in the spatiotemporal dimension, thereby improving the rationality and accuracy of the clustering results and avoiding clustering bias caused by ignoring spatiotemporal factors.

[0038] The initial density threshold of the dynamic density threshold adjustment module is set according to the initial distribution characteristics of the multimodal data and determined by statistical methods (such as the average density of data points) or empirical rules. During the clustering process, the density threshold is dynamically adjusted according to the density distribution of the current clustering result: if there is a high-density area in the current clustering result, it indicates that there may be a hot event, and the density threshold is appropriately lowered to further subdivide the hot area. If the hot spot is not obvious in the current clustering result, the density threshold is appropriately increased to merge smaller hot areas. By dynamically adjusting the density threshold, hot events with different importance and impact ranges can be identified at different levels. For example, large-scale hot events can be discovered at the global level, and hot issues within a small range or a specific time period can be discovered at the local level, thereby improving the flexibility and practicality of multimodal data clustering analysis.

[0039] An electronic device comprising: a memory and at least one processor;

[0040] Wherein, the memory stores a computer program;

[0041] The at least one processor executes the computer program stored in the memory, so that the at least one processor performs the multimodal data intelligent clustering method based on spatiotemporal features as described above.

[0042] A computer-readable storage medium stores a computer program, which can be executed by a processor to implement the multimodal data intelligent clustering method based on spatiotemporal features as described above.

[0043] The multimodal data intelligent clustering method, system, device, and medium based on spatiotemporal features of the present invention have the following advantages:

[0044] (1) This invention achieves efficient and accurate cluster analysis of complex data by comprehensively considering multi-dimensional features such as text semantics, geographic space, and time decay. By constructing multimodal data vectors (including text semantic vectors, geographic coding vectors, and time decay vectors), using an improved Transformer architecture to generate high-quality multimodal data representations, and introducing spatiotemporal constraints to optimize the distribution of cluster centers, the density threshold is dynamically adjusted to achieve multi-granularity hotspot discovery.

[0045] (2) This invention can significantly improve clustering accuracy, spatiotemporal correlation mining capabilities, and the flexibility of hotspot discovery. It is applicable to big data analysis in a variety of fields, and performs particularly well in scenarios that require processing spatiotemporal related data, providing a powerful technical means for data management and decision support.

[0046] (3) This invention is mainly used to perform efficient and accurate cluster analysis on data containing multi-dimensional information such as text, geography, and time, in order to mine potential patterns and hot spots in the data. It is applicable to big data analysis in various fields, and has significant advantages in scenarios where time-space related data needs to be processed.

[0047] (4) The present invention constructs multimodal data vectors by comprehensively considering multi-dimensional features such as text semantics, geographic space, and time decay, and uses an improved Transformer architecture to generate high-quality multimodal data representations. At the same time, it introduces spatiotemporal constraints to optimize the distribution of cluster centers and dynamically adjusts the density threshold to achieve multi-granular hotspot discovery, thereby improving the accuracy, efficiency, and adaptability of clustering, better meeting the actual needs of multimodal data analysis, and is particularly suitable for scenarios that require processing spatiotemporal related data.

[0048] (5) This invention improves clustering accuracy: By comprehensively considering multi-dimensional features such as text semantics, geography, and time to construct multimodal data vectors, and using an improved Transformer architecture to generate more accurate multimodal data representations, the clustering results can more realistically reflect the similarities and correlations between data. Compared with traditional clustering methods, this invention significantly improves clustering accuracy and can more effectively identify potential patterns and hot spots in the data;

[0049] (6) This invention enhances spatiotemporal correlation mining: by introducing spatiotemporal constraints to optimize the distribution of cluster centers, it fully explores the spatiotemporal correlation characteristics of data, helps to discover hot events with geographical clustering and temporal trends, and provides a more comprehensive and in-depth perspective for data analysis and decision-making, enabling a better grasp of the development context and spatial distribution patterns of data. This is particularly important for scenarios that require processing spatiotemporal related data, and can significantly enhance the value of analysis results.

[0050] (7) The present invention achieves multi-granularity hotspot discovery: The mechanism of dynamically adjusting the density threshold enables this method to adapt to different levels of hotspot discovery needs. It can identify both large-scale, widely impactful hotspot events and small-scale, short-term hotspot issues. This multi-granularity hotspot discovery capability provides stronger support for targeted response measures, improves the level of refined data management, and is particularly suitable for complex and changing data environments.

[0051] (8) The present invention improves clustering efficiency and adaptability: The improved Transformer architecture and optimized clustering algorithm have higher computational efficiency and better scalability when processing massive multimodal data. They can quickly respond to dynamic changes in data and update clustering results in a timely manner. This provides an effective technical means for real-time data monitoring and analysis, and enhances the adaptability and practicality of data analysis systems in complex and changing data environments. It is particularly suitable for scenarios that require real-time processing and analysis of large amounts of multimodal data.

[0052] (9) The present invention optimizes the self-attention mechanism through an improved Transformer architecture to adapt to the characteristics of multimodal data and better capture the complex relationship between text semantics, geographic and temporal features;

[0053] (10) In the clustering process, the present invention introduces spatiotemporal constraints to enable the cluster centers to better represent the characteristics of the corresponding data group in the spatiotemporal dimensions, thereby improving the rationality and accuracy of the clustering results;

[0054] (11) The strategy of dynamically adjusting the density threshold of the present invention can automatically adjust the density threshold used to divide high-density areas and low-density areas according to the distribution characteristics of multimodal data and real-time feedback information during the clustering process, thereby realizing multi-granularity hotspot discovery. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The present invention will be further described below with reference to the accompanying drawings.

[0056] Attachment Figure 1 This is a flowchart of the intelligent clustering method of multimodal data based on spatiotemporal features. DETAILED DESCRIPTION

[0057] The multimodal data intelligent clustering method, system, device and medium based on spatiotemporal features of the present invention are described in detail below with reference to the accompanying drawings and specific embodiments.

[0058] Example 1:

[0059] As attached Figure 1 As shown, this embodiment provides a multimodal data intelligent clustering method based on spatiotemporal features, the method is specifically as follows:

[0060] S1. Obtain multimodal events: Obtain multimodal events and construct text semantic vectors, geocoding vectors, and time decay vectors based on the multimodal events. The text semantic vector refers to the key semantic features of text data extracted through natural language processing technology. The geocoding vector is generated by encoding the geographic coordinates of the data's location, using the GeoHash tool to convert the geographic coordinates into a fixed-length encoding vector. The time decay vector is calculated based on a specific decay function (such as an exponential decay function) based on the interval between the data's occurrence time and the current time, and is used to reflect the timeliness of the data. The above-mentioned method of constructing multimodal data vectors can fully capture the multi-dimensional characteristics of the data, providing a richer information basis for subsequent clustering analysis.

[0061] S2. Construct a multimodal data vector: Represent the multimodal data as a combination of a text semantic vector, a geocoding vector, and a time decay vector, i.e., multimodal data vector = text semantic vector ⊕ geocoding vector ⊕ time decay vector;

[0062] S3. Improved Transformer Architecture Processing: Generate multimodal data representation based on the improved Transformer architecture, process multimodal data vectors, and fully integrate multi-dimensional information such as textual semantics, geography, and time;

[0063] S4. Optimize cluster centers with spatiotemporal constraints: Optimize the distribution of cluster centers through spatiotemporal constraints, and adjust the cluster centers according to the geographic spatial distribution and time series characteristics of multimodal data;

[0064] S5. Dynamic density threshold adjustment: Dynamically adjust the density threshold to achieve multi-granularity hotspot discovery. According to the distribution characteristics of multimodal data and real-time feedback information in the clustering process, the density threshold is automatically adjusted to identify hot events at different levels.

[0065] The improved Transformer architecture in step S3 of this embodiment achieves efficient processing and fusion of multimodal data by optimizing the network structure of the encoder and decoder, as well as the input embedding layer and the self-attention mechanism; specifically, as follows:

[0066] S301. Encoder structure: The encoder is composed of multiple identical encoder layers stacked together. Multimodal fusion and feature extraction are repeated in each encoder layer, deepening the understanding and fusion of the data layer by layer. Each encoder layer includes a multimodal fusion attention layer and an encoding feedforward neural network layer. The multimodal fusion attention layer is used to capture the complex relationship between text, geographic, and time features in the input data. The multimodal fusion attention layer dynamically adjusts the weight distribution between features through a self-attention mechanism. Specifically, the self-attention mechanism calculates the correlation weight between each data point and other data points, and performs a weighted summation of the features based on the correlation weights to generate a contextual representation of each data point. The encoding feedforward neural network layer is used to perform a nonlinear transformation on the contextual representation of each data point to further extract a high-level representation of the features.

[0067] S302. Decoder structure: The decoder is composed of multiple identical decoder layers stacked together; each decoder layer includes a self-attention layer, a cross-attention layer, and a decoding feedforward neural network layer; among them, the self-attention layer is used to capture the feature relationship in the decoder input data; the cross-attention layer allows the decoder to access the encoder output, thereby realizing information interaction between the encoder and decoder. By calculating the correlation weight between the decoder input and the encoder output, the cross-attention layer can combine the multimodal features extracted by the encoder with the current state of the decoder to further fuse the multimodal information; the decoding feedforward neural network layer performs nonlinear transformation on the fused features to generate the final output representation.

[0068] S303, Input Embedding and Multimodal Fusion: In the input embedding layer, text semantic vectors, geocoding vectors, and time decay vectors are respectively embedded into the corresponding dimensional spaces to form a unified multimodal input representation. This process lays the foundation for subsequent feature fusion, enabling the model to process data from different modalities simultaneously.

[0069] Among them, the encoder's feature extraction and fusion: The encoder captures the complex relationship between different features in the input data through a multimodal fusion attention layer. In each layer of the encoder, the multimodal fusion attention layer calculates the correlation weights between each data point and other data points, and performs a weighted summation of the features based on these weights to generate a contextual representation of each data point. Subsequently, the feedforward neural network layer performs a nonlinear transformation on the contextual representation of each data point to further extract a high-level representation of the features. This process is repeated in each layer of the encoder, deepening the understanding and fusion of the data layer by layer, and ultimately generating the encoder's output representation, which fully integrates multi-dimensional information such as text, geography, and time.

[0070] Information interaction and fusion in the decoder: The decoder interacts with the encoder's output via a cross-attention layer. Within each decoder layer, the cross-attention layer calculates the correlation weight between the decoder input and the encoder output, combining the multimodal features extracted by the encoder with the decoder's current state. This process enables the decoder to fully utilize the rich feature information extracted by the encoder and further fuse multimodal data. Subsequently, the feedforward neural network layer performs a nonlinear transformation on the fused features to generate the final output representation. Through the decoder's multi-layer processing, the model is able to generate a more accurate and comprehensive representation of multimodal data, providing a high-quality data foundation for subsequent clustering analysis.

[0071] Through this improved Transformer architecture, the model can effectively process multimodal data, capture the complex relationships between different features, and generate a multimodal data representation that fully integrates multidimensional information. This representation not only incorporates textual semantics but also incorporates geographic space and time decay features, providing richer and more accurate data support for subsequent clustering analysis and improving the model's ability to understand complex data relationships.

[0072] The spatiotemporal constraint optimization of cluster centers in step S4 of this embodiment refers to introducing spatiotemporal constraints to optimize the distribution of cluster centers during the clustering process; the details are as follows:

[0073] S401, obtaining geographic coordinate information of the data points through the geocoding vector, and calculating the geographic distance (Euclidean distance) between the data points. At the same time, obtaining time information of the data points through the time decay vector, and calculating the time interval between the data points;

[0074] S402, in the cluster initialization stage, selecting the initial cluster center based on the geographic spatial distribution and time series characteristics to make it representative in the spatiotemporal dimension;

[0075] S403. During the clustering iteration process, constrain the update rules of the cluster centers based on geographic distance and time interval. For example, the movement range of the cluster centers in geographic space is limited to prevent the cluster centers from deviating too much from the geographic distribution center of the data points. At the same time, the time weight of the cluster centers is adjusted based on the time series characteristics to better reflect the time trend of the data points.

[0076] S404. By introducing spatiotemporal constraints, the cluster centers can better represent the characteristics of the corresponding data groups in the spatiotemporal dimensions, thereby improving the rationality and accuracy of the clustering results and avoiding clustering bias caused by ignoring spatiotemporal factors.

[0077] The dynamic density threshold adjustment in step S5 of this embodiment is specifically as follows:

[0078] S501: An initial density threshold is set according to the initial distribution characteristics of the multimodal data and is determined by statistical methods (such as the average density of data points) or empirical rules.

[0079] S502: During the clustering process, dynamically adjust the density threshold according to the density distribution of the current clustering result:

[0080] If there are high-density areas in the current clustering results, it indicates that there may be hotspot events. Appropriately lower the density threshold to further subdivide the hotspot areas.

[0081] If the hot spots in the current clustering results are not obvious, increase the density threshold appropriately to merge smaller hot spots.

[0082] S503. By dynamically adjusting the density threshold, hot events with different importance and impact ranges are identified at different levels. For example, large-scale hot events are discovered at the global level, and hot issues within a small range or a specific time period are discovered at the local level, thereby improving the flexibility and practicality of multimodal data clustering analysis.

[0083] Example 2:

[0084] This embodiment provides a multimodal data intelligent clustering system based on spatiotemporal features. The system is specifically as follows:

[0085] The multimodal event acquisition module is used to acquire multimodal events and construct text semantic vectors, geocoding vectors, and time decay vectors based on these events. The text semantic vector refers to the key semantic features of text data extracted through natural language processing technology. The geocoding vector is generated by encoding the geographic coordinates of the data's location, using the GeoHash tool to convert the geographic coordinates into a fixed-length encoding vector. The time decay vector is calculated based on the interval between the data's occurrence time and the current time using a specific decay function (such as an exponential decay function) to reflect the timeliness of the data.

[0086] A multimodal data vector construction module is used to represent multimodal data as a combination of a text semantic vector, a geocoding vector, and a time decay vector, i.e., multimodal data vector = text semantic vector ⊕ geocoding vector ⊕ time decay vector;

[0087] The multimodal data processing module is used to generate multimodal data representations based on the improved Transformer architecture, process multimodal data vectors, and fully integrate multi-dimensional information such as textual semantics, geography, and time;

[0088] The optimization module is used to optimize the distribution of cluster centers through spatiotemporal constraints and adjust the cluster centers according to the geographic spatial distribution and time series characteristics of multimodal data;

[0089] The dynamic density threshold adjustment module is used to dynamically adjust the density threshold to achieve multi-granularity hotspot discovery. According to the distribution characteristics of multimodal data and real-time feedback information in the clustering process, the density threshold is automatically adjusted to identify hot events at different levels.

[0090] The improved Transformer architecture in this embodiment achieves efficient processing and fusion of multimodal data by optimizing the network structure of the encoder and decoder, as well as the input embedding layer and self-attention mechanism; the details are as follows:

[0091] Encoder structure: The encoder is composed of multiple identical encoder layers stacked together. Multimodal fusion and feature extraction are repeated in each encoder layer, deepening the understanding and fusion of the data layer by layer. Each encoder layer includes a multimodal fusion attention layer and an encoding feedforward neural network layer. The multimodal fusion attention layer is used to capture the complex relationship between text, geography, and time features in the input data. The multimodal fusion attention layer dynamically adjusts the weight distribution between features through the self-attention mechanism. Specifically, the self-attention mechanism is used to calculate the correlation weight between each data point and other data points, and the features are weighted and summed according to the correlation weight to generate a contextual representation of each data point. The encoding feedforward neural network layer is used to perform nonlinear transformations on the contextual representation of each data point to further extract high-level representations of the features.

[0092] Decoder structure: The decoder is composed of multiple identical decoder layers stacked together; each decoder layer includes a self-attention layer, a cross-attention layer, and a decoding feedforward neural network layer; among them, the self-attention layer is used to capture the feature relationship in the decoder input data; the cross-attention layer allows the decoder to access the encoder output, thereby realizing information interaction between the encoder and decoder. By calculating the correlation weight between the decoder input and the encoder output, the cross-attention layer can combine the multimodal features extracted by the encoder with the current state of the decoder to further fuse the multimodal information; the decoding feedforward neural network layer performs nonlinear transformation on the fused features to generate the final output representation.

[0093] Input embedding and multimodal fusion: In the input embedding layer, text semantic vectors, geocoding vectors, and time decay vectors are respectively embedded into the corresponding dimensional spaces to form a unified multimodal input representation. This process lays the foundation for subsequent feature fusion, enabling the model to process data from different modalities simultaneously.

[0094] The optimization module in this embodiment obtains the geographic coordinate information of the data points through the geocoding vector and calculates the geographic distance (Euclidean distance) between the data points. At the same time, it obtains the time information of the data points through the time decay vector and calculates the time interval between the data points. In the clustering initialization stage, the initial cluster center is selected according to the geographic spatial distribution and time series characteristics to make it representative in the time and space dimensions. In the clustering iteration process, the update rules of the cluster center are constrained according to the geographic distance and time interval, for example, the movement range of the cluster center in the geographic space is limited to avoid the cluster center from deviating too much from the geographic distribution center of the data point. At the same time, the time weight of the cluster center is adjusted according to the time series characteristics to enable it to better reflect the time trend of the data point. By introducing time and space constraints, the cluster center can better represent the characteristics of the corresponding data group in the time and space dimensions, thereby improving the rationality and accuracy of the clustering results and avoiding clustering bias caused by ignoring time and space factors.

[0095] The initial density threshold of the dynamic density threshold adjustment module in this embodiment is set according to the initial distribution characteristics of the multimodal data and is determined by statistical methods (such as the average density of data points) or empirical rules; during the clustering process, the density threshold is dynamically adjusted according to the density distribution of the current clustering result: if there is a high-density area in the current clustering result, it means that there may be a hot event, and the density threshold is appropriately lowered to further subdivide the hot area; if the hot spot is not obvious in the current clustering result, the density threshold is appropriately increased to merge smaller hot areas; by dynamically adjusting the density threshold, hot events with different importance and impact ranges are identified at different levels, for example, large-scale hot events are discovered at the global level, and hot issues within a small range or a specific time period are excavated at the local level, thereby improving the flexibility and practicality of multimodal data clustering analysis.

[0096] Example 3:

[0097] This embodiment also provides an electronic device, including: a memory and a processor;

[0098] wherein the memory stores computer-executable instructions;

[0099] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the multimodal data intelligent clustering method based on spatiotemporal features in any embodiment of the present invention.

[0100] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or any conventional processor, etc.

[0101] The memory can be used to store computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area. The program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, the memory can also include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, at least one disk storage period, a flash memory device, or other volatile solid-state memory devices.

[0102] Example 4:

[0103] This embodiment also provides a computer-readable storage medium storing a plurality of instructions, which are loaded by a processor to cause the processor to execute the multimodal data intelligent clustering method based on spatiotemporal features in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided, wherein the storage medium stores software program code that implements the functions of any of the above embodiments, and the computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.

[0104] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute part of the present invention.

[0105] Examples of storage media for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (e.g., CD-ROMs, CD-Rs, CD-RWs, DVD-ROMs, DVD-RYMs, DVD-RWs, DVD+RWs), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code may be downloaded from a server computer via a communications network.

[0106] In addition, it should be clear that the functions of any of the above embodiments can be achieved not only by executing the program code read by the computer, but also by enabling the operating system operating on the computer to complete part or all of the actual operations based on the instructions of the program code.

[0107] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU installed on the expansion board or expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above embodiments.

[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multimodal data intelligent clustering method based on spatiotemporal features, characterized in that: The method is as follows: Obtain multimodal events: Obtain multimodal events and construct text semantic vectors, geocoding vectors, and time decay vectors based on the multimodal events; Among them, the text semantic vector refers to the key semantic features of text data extracted through natural language processing technology; the geocoding vector refers to the encoding generated based on the geographic coordinates of the data generation location, and the geographic coordinates are converted into a fixed-length encoding vector using the GeoHash tool; the time decay vector refers to the interval between the data generation time and the current time, calculated according to a specific decay function, and is used to reflect the timeliness of the data; Constructing a multimodal data vector: Representing multimodal data as a combination of a text semantic vector, a geocoding vector, and a time decay vector, i.e., multimodal data vector = text semantic vector ⊕ geocoding vector ⊕ time decay vector; Improved Transformer Architecture Processing: Generates multimodal data representations based on the improved Transformer architecture, processes multimodal data vectors, and fully integrates multi-dimensional information such as textual semantics, geography, and time. Optimize cluster centers with spatiotemporal constraints: Optimize the distribution of cluster centers through spatiotemporal constraints, and adjust the cluster centers based on the geographic spatial distribution and time series characteristics of multimodal data; Dynamic density threshold adjustment: Dynamically adjust the density threshold to achieve multi-granularity hotspot discovery. Based on the distribution characteristics of multimodal data and real-time feedback information during the clustering process, the density threshold is automatically adjusted to identify hotspot events at different levels.

2. The multimodal data intelligent clustering method based on spatiotemporal features according to claim 1 is characterized in that: The improved Transformer architecture achieves efficient processing and fusion of multimodal data by optimizing the network structure of the encoder and decoder, as well as the input embedding layer and self-attention mechanism; the details are as follows: Encoder Structure: The encoder is composed of multiple identical encoder layers stacked together. Multimodal fusion and feature extraction are repeated in each encoder layer, deepening the understanding and fusion of the data layer by layer. Each encoder layer includes a multimodal fusion attention layer and an encoding feedforward neural network layer. The multimodal fusion attention layer is used to capture the complex relationships between textual, geographic, and temporal features in the input data. The encoding feedforward neural network layer performs nonlinear transformations on the contextual representation of each data point to further extract high-level representations of the features. Decoder structure: The decoder is composed of multiple identical decoder layers stacked together; each decoder layer includes a self-attention layer, a cross-attention layer, and a decoding feedforward neural network layer; among them, the self-attention layer is used to capture the feature relationship in the decoder input data; the cross-attention layer allows the decoder to access the encoder output, thereby realizing information interaction between the encoder and decoder. By calculating the correlation weight between the decoder input and the encoder output, the cross-attention layer can combine the multimodal features extracted by the encoder with the current state of the decoder to further fuse the multimodal information; the decoding feedforward neural network layer performs nonlinear transformation on the fused features to generate the final output representation. Input embedding and multimodal fusion: In the input embedding layer, the text semantic vector, geocoding vector, and time decay vector are embedded into the corresponding dimensional space to form a unified multimodal input representation.

3. The multimodal data intelligent clustering method based on spatiotemporal features according to claim 2 is characterized in that: The multimodal fusion attention layer dynamically adjusts the weight distribution between features through the self-attention mechanism; specifically: the self-attention mechanism calculates the correlation weight between each data point and other data points, and performs weighted summation of the features according to the correlation weight to generate a contextual representation of each data point.

4. The multimodal data intelligent clustering method based on spatiotemporal features according to claim 1 is characterized in that: Spatiotemporal constraint optimization of cluster centers refers to the introduction of spatiotemporal constraints to optimize the distribution of cluster centers during the clustering process. The details are as follows: The geographic coordinate information of the data points is obtained through the geocoding vector, and the geographic distance between the data points is calculated. At the same time, the time information of the data points is obtained through the time decay vector, and the time interval between the data points is calculated. In the cluster initialization stage, the initial cluster centers are selected based on the geographic spatial distribution and time series characteristics to make them representative in the spatiotemporal dimension; During the clustering iteration process, the update rules of the cluster centers are constrained according to the geographical distance and time interval. At the same time, the time weight of the cluster centers is adjusted according to the time series characteristics so that it can better reflect the time trend of the data points. By introducing spatiotemporal constraints, the cluster centers can better represent the characteristics of the corresponding data groups in the spatiotemporal dimensions.

5. The multimodal data intelligent clustering method based on spatiotemporal features according to claim 1 is characterized in that: The dynamic density threshold adjustment is as follows: The initial density threshold is set according to the initial distribution characteristics of the multimodal data and is determined by statistical methods or empirical rules; During the clustering process, the density threshold is dynamically adjusted according to the density distribution of the current clustering results: If there are high-density areas in the current clustering results, it indicates that there may be hotspot events. Appropriately lower the density threshold to further subdivide the hotspot areas. If the hot spots in the current clustering results are not obvious, increase the density threshold appropriately to merge smaller hot spots. By dynamically adjusting the density threshold, hot events with different importance and impact ranges are identified at different levels.

6. A multimodal data intelligent clustering system based on spatiotemporal features, characterized by: The system is as follows: Multimodal event acquisition module, used to acquire multimodal events and construct text semantic vectors, geocoding vectors and time decay vectors based on multimodal events; Among them, the text semantic vector refers to the key semantic features of text data extracted through natural language processing technology; the geocoding vector refers to the encoding generated based on the geographic coordinates of the data generation location, and the geographic coordinates are converted into a fixed-length encoding vector using the GeoHash tool; the time decay vector refers to the interval between the data generation time and the current time, calculated according to a specific decay function, and is used to reflect the timeliness of the data; A multimodal data vector construction module is used to represent multimodal data as a combination of a text semantic vector, a geocoding vector, and a time decay vector, i.e., multimodal data vector = text semantic vector ⊕ geocoding vector ⊕ time decay vector; The multimodal data processing module is used to generate multimodal data representations based on the improved Transformer architecture, process multimodal data vectors, and fully integrate multi-dimensional information such as textual semantics, geography, and time; The optimization module is used to optimize the distribution of cluster centers through spatiotemporal constraints and adjust the cluster centers according to the geographic spatial distribution and time series characteristics of multimodal data; The dynamic density threshold adjustment module is used to dynamically adjust the density threshold to achieve multi-granularity hotspot discovery. According to the distribution characteristics of multimodal data and real-time feedback information in the clustering process, the density threshold is automatically adjusted to identify hot events at different levels.

7. The multimodal data intelligent clustering system based on spatiotemporal features according to claim 6, characterized in that: The improved Transformer architecture achieves efficient processing and fusion of multimodal data by optimizing the network structure of the encoder and decoder, as well as the input embedding layer and self-attention mechanism; the details are as follows: Encoder structure: The encoder is composed of multiple identical encoder layers stacked together. Multimodal fusion and feature extraction are repeated in each encoder layer, deepening the understanding and fusion of the data layer by layer. Each encoder layer includes a multimodal fusion attention layer and an encoding feedforward neural network layer. The multimodal fusion attention layer is used to capture the complex relationship between text, geography, and time features in the input data. The multimodal fusion attention layer dynamically adjusts the weight distribution between features through the self-attention mechanism. Specifically, the self-attention mechanism is used to calculate the correlation weight between each data point and other data points, and the features are weighted and summed according to the correlation weight to generate a contextual representation of each data point. The encoding feedforward neural network layer is used to perform nonlinear transformations on the contextual representation of each data point to further extract high-level representations of the features. Decoder structure: The decoder is composed of multiple identical decoder layers stacked together; each decoder layer includes a self-attention layer, a cross-attention layer, and a decoding feedforward neural network layer; among them, the self-attention layer is used to capture the feature relationship in the decoder input data; the cross-attention layer allows the decoder to access the encoder output, thereby realizing information interaction between the encoder and decoder. By calculating the correlation weight between the decoder input and the encoder output, the cross-attention layer can combine the multimodal features extracted by the encoder with the current state of the decoder to further fuse the multimodal information; the decoding feedforward neural network layer performs nonlinear transformation on the fused features to generate the final output representation. Input embedding and multimodal fusion: In the input embedding layer, the text semantic vector, geocoding vector, and time decay vector are embedded into the corresponding dimensional space to form a unified multimodal input representation.

8. The multimodal data intelligent clustering system based on spatiotemporal features according to claim 6 or 7, characterized in that: The optimization module obtains the geographic coordinate information of the data points through the geocoding vector and calculates the geographic distance between the data points. At the same time, it obtains the time information of the data points through the time decay vector and calculates the time interval between the data points. In the clustering initialization phase, the initial cluster centers are selected based on the geographic spatial distribution and time series characteristics to make them representative in the spatiotemporal dimension. During the clustering iteration process, the update rules of the cluster centers are constrained according to the geographic distance and time interval. At the same time, the time weights of the cluster centers are adjusted according to the time series characteristics to better reflect the temporal trend of the data points. By introducing spatiotemporal constraints, the cluster centers can better represent the characteristics of the corresponding data group in the spatiotemporal dimension. The initial density threshold of the dynamic density threshold adjustment module is set according to the initial distribution characteristics of the multimodal data and determined by statistical methods or empirical rules. During the clustering process, the density threshold is dynamically adjusted according to the density distribution of the current clustering result. If there is a high-density area in the current clustering result, it indicates that there may be a hotspot event. The density threshold is appropriately lowered to further subdivide the hotspot area. If the hotspot is not obvious in the current clustering result, the density threshold is appropriately increased to merge smaller hotspot areas. By dynamically adjusting the density threshold, hotspot events with different importance and impact ranges can be identified at different levels.

9. An electronic device, characterized in that: include: memory and at least one processor; Wherein, the memory stores a computer program; The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the multimodal data intelligent clustering method based on spatiotemporal features as described in any one of claims 1 to 5.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which can be executed by a processor to implement the multimodal data intelligent clustering method based on spatiotemporal features as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Space-time data clustering method

    CN113886667A

  • Large-scale spatio-temporal stream data real-time spatial connection query method and related equipment

    CN115809360A

  • Social grouping method for large-scale scene video

    CN116403286A

  • Long-term sequence dependency model optimization method and system based on context analysis

    CN118468936A

  • Video timestamp event identification and reasoning method based on multi-modal large model

    CN119723431A

Cited By

  • Complex table structure understanding and information extraction method, system and device and medium

    CN121354146A