Self-supervised spatio-temporal data completion method and system based on time sequence semantic attention

Through the time-series semantic attention self-supervision method, the timing segmented modeling and cross-attention mechanism are used, combined with the graph convolution network, the nonlinear spatiotemporal correlation capture problem of traditional methods in data loss scenarios is solved, and the high robust timing embedding is achieved, which improves the accuracy of data completion and model generalization ability.

CN120336736AActive Publication Date: 2025-07-18CHANGSHU INSTITUTE OF TECHNOLOGY

Patent Information

Application Number
CN202510841249.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-18
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

Traditional methods are difficult to effectively capture nonlinear spatiotemporal correlations in urban spatiotemporal data, especially in the data loss scenarios caused by sensor failure or communication interruption, which affects traffic management and environmental monitoring tasks in smart cities.

Method used

The self-supervised method based on timing semantic attention is adopted, through timing segment modeling, cross-attention mechanism and graph convolution network, the potential timing dependence patterns between segments are captured, and the generalization of the model is enhanced through self-supervised learning.

Benefits of technology

In the data loss scenario, high-rootability timing embedding is achieved, the accuracy of data completion and the generalization ability of the model are improved, and it is suitable for space-time node-level tasks in smart cities and the Internet of Things.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336736A_ABST
    Figure CN120336736A_ABST
Patent Text Reader

Abstract

The invention discloses a time sequence semantic attention-based self-supervised spatio-temporal data completion method and system, and the method comprises the steps: obtaining spatio-temporal data, and dividing an incomplete time sequence into semantic segment units through time sequence segmentation modeling; mining an inter-segment potential time sequence dependence mode by using a cross attention mechanism to obtain a segmented attention calculation result; constructing a time sequence comparison model based on a segmented attention calculation result, and obtaining strong generalization embedding of time sequence modeling through a self-supervision method; and carrying out spatial domain modeling through the graph convolutional network according to the time sequence characteristics which are obtained by time sequence modeling and are embedded as each node, so as to obtain a spatial domain modeling result. According to the method, strong dependence of a traditional model on data integrity is broken through, nonlinear space-time correlation can be effectively captured in a data missing scene, and a high-robustness time sequence embedding basis is provided for subsequent space-time node level tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of urban spatiotemporal data completion, and the present invention relates to a self-supervised spatiotemporal data completion method and system based on temporal semantic attention. Background Art

[0002] In urban traffic monitoring, sensors may lose data due to malfunctions, communication interruptions, power failures, etc., which has a negative impact on the effectiveness of subsequent applications.

[0003] Complete spatiotemporal data is the basis for tasks such as traffic management and environmental monitoring in smart cities. For example, PM2.5 concentration interpolation requires modeling long-term temporal trends and cross-regional spatial associations, while traffic data completion requires processing dynamically changing spatiotemporal semantic relationships in large-scale road networks.

[0004] Urban spatiotemporal data (such as traffic flow, environmental monitoring, sensor networks, etc.) usually have multidimensionality (such as time, space, semantic attributes), high missing rate (sensor failure, transmission interruption) and complex spatiotemporal dependencies (such as spatiotemporal propagation of traffic flow, dynamic associations between regions). For example, random missing and non-random missing (such as sensor failure for multiple consecutive days) coexist in traffic data, making it difficult for traditional interpolation methods to capture nonlinear spatiotemporal associations. Summary of the invention

[0005] The purpose of the present invention is to provide a self-supervised spatiotemporal data completion method and system based on temporal semantic attention. This method breaks through the strong dependence of traditional models on data integrity, can effectively capture nonlinear spatiotemporal correlations in data missing scenarios, and provides a highly robust temporal embedding foundation for subsequent spatiotemporal node-level tasks.

[0006] The technical solution to achieve the purpose of the present invention is: A self-supervised spatiotemporal data completion method based on temporal semantic attention includes the following steps: S01: Obtain spatiotemporal data and divide the incomplete time series into semantic segment units through time series segmentation modeling; S02: Use the cross-attention mechanism to mine the potential temporal dependency patterns between segments and obtain the segmented attention calculation results; S03: Construct a temporal contrast model based on the segmented attention calculation results, and obtain a strong generalization embedding for temporal modeling through self-supervision methods; S04: The embedding obtained through time series modeling is used as the time series feature of each node, and spatial domain modeling is performed through the graph convolutional network to obtain the spatial domain modeling result.

[0007] In the preferred technical solution, after step S04, it also includes designing a segment-to-segment reconstruction loss function to constrain the consistency of local segment features and global temporal semantics.

[0008] In the preferred technical solution, the method of dividing the incomplete time series into semantic segment units by time series segment modeling in step S01 includes: S11: Perform long-short time dependency encoding. The encoding result contains global time information. For the input time series data ,in, is a real number, Represents the number of nodes, and Represent the dimension and length of the time series features respectively. After splicing the time information, we get data containing global time information. , Greater than , indicating the new time series feature dimension: ;

[0009] in, Represents a splicing operation, Indicates time information including global time information; S12: Perform short-term time-dependent mapping coding. The mapping coding only considers the time information of a day and uses position coding to input the time information of each time series. S13: In the timing segmentation stage, each input channel First, it is normalized to zero mean and unit standard deviation through a reversible instance; then, Split into contiguous non-overlapping blocks ,in is the number of blocks, is the original dimension of each block, and then they are embedded as , the embedding method uses a simple linear layer as a block embedder with dimension .

[0010] In the preferred technical solution, step S02 utilizes the cross-attention mechanism to mine the potential temporal dependency patterns between segments, including: S21: Inter-segment attention uses the temporal segment as the basic computational unit and represents the input of each segment-related module as ,in and are the length and dimension of the input, the initial hidden state , for single-head attention, the input sequence Projected by three projection matrices to obtain the query ,key Sum ; Query segment and key segments Correlation measure between is: ;

[0011] Among them, is the dot product operation between two matrices of the same size; Each query segment The relevance measure with all key segments will be normalized by a function to obtain the aggregated weights : ;

[0012] The output at the th segment position is the weighted sum of all value segments ;

[0013] The output of the inter-segment attention module is obtained by concatenating all outputs as follows: ;

[0014] The final output result after passing through the linear layer is , achieving multi-scale segment attention modeling; S22: Calculate the segmented temporal reconstruction loss function is: ;

[0015] Among them, represents the total number of reconstructed segments, represents the true value of the reconstructed segment, represents the temporal cross-modeling result of the reconstructed segment; The goal of the segmented temporal reconstruction loss function is to minimize the overall loss of each segment.

[0016] In the preferred technical solution, the method for obtaining the strongly generalized embedding of temporal modeling in step S03 by the self-supervised method includes: Taking the segmented cross-modeling intermediate representation as the input spatio-temporal data , generating strong samples and weak samples through the augmentation strategy in contrastive learning, that is, representing its strong augmented view as , and representing its weak augmented view as ; then encoding the augmented views. First, the encoder maps them to a high-dimensional latent representation , using a one-layer neural network for encoding, and the encoding result is represented as , two corresponding views are obtained and strong enhancement encoding and weak enhancement encoding ; Through the autoregressive model all the time series data before the moment in the segment is encoded and summarized into a context vector , , , where is the hidden dimension; Using the log bilinear model retains the mutual information between the input and , and the mutual information obtained is , where is the exponential function, is the linear transformation parameter.

[0017] In the preferred technical solution, step S03 further includes: The autoregressive model adopts Transformer, stacks L identical layers to generate the final features, and adds a token to the input, and its state acts as a representative context vector in the output; Apply the features to a linear projection layer, which maps the features to the hidden dimension, that is ; then, send the output of this linear projection to Transformer to obtain the output result , and the input features become after including the token, where the subscript 0 represents the input of the first layer; next, is passed to the Transformer layer, as shown in the following formula: ;

[0018] ;

[0019] where , represents the result of the th layer, is the output of the attention calculation and is an intermediate result, represents the multi-head attention calculation, represents the normalization of the values, represents the multi-layer perceptron; Reattach the token value from the final output, so that , as the input of the time series contrast model.

[0020] In the preferred technical solution, the spatial domain modeling by the graph convolutional network in step S04 includes: Calculating the learned adaptive adjacency matrix : ;

[0021] Wherein, is the activation function; Obtaining the predefined adjacency matrix , combining the adaptive adjacency matrix and to define the multi-directional multi-graph spatial convolution, including the forward spatial adjacency matrix and the backward adjacency matrix , and the calculation process of the spatial domain modeling is: ;

[0022] Wherein, is the hidden layer state of the layer, , and are the linear transformation parameters; Using residual connection and batch normalization to obtain the final enhanced result ; The loss function in the spatial domain modeling stage is: ;

[0023] Wherein, represents all the points participating in the spatial domain loss calculation, is the output result of the spatial domain modeling, is the label value corresponding to the spatial domain modeling.

[0024] In the preferred technical solution, the segment-to-segment reconstruction loss function includes: the piecewise reconstruction loss function, the time-domain cross loss function, and the maximum similarity loss function; In the time-domain cross loss function, the strong augmented context vector / weak augmented context vector is used to cross-predict the weak / strong augmented future time steps , and the time-domain cross loss function minimizes the dot product between the prediction result and the true result of the same sample, while maximizing the dot product with other samples in the same augmented batch. The two loss functions and formulas are: ;

[0025] ;

[0026] wherein, represents the th time step, represents the total number of future time steps, represents the th moment, represents a linear transformation, represents an exponential function, represents the prediction result of the future moment with weak enhancement ; represents the prediction result of the future moment with strong enhancement ; For the maximum similarity loss function, under the temporal contrast model, given input samples, obtain contexts from two enhanced views. For the context , represent as the positive sample of , and take as the positive pair. At the same time, the remaining contexts from other inputs in the same batch are regarded as negative samples, and a context contrast loss is obtained to maximize the similarity between positive pairs and minimize the similarity between negative pairs, resulting in the context contrast loss function : ;

[0027] wherein, , represents the normalization of the dot product between two vectors and , is an indicator function that takes the value of 1 if and only if , is the temperature parameter; The total loss function of the temporal cross-contrast learning module is: ;

[0028] wherein, , and are hyperparameters representing the relative weights of each loss, is the piecewise reconstruction loss function, and are the temporal cross losses, is the maximum similarity loss function.

[0029] The present invention also discloses a self-supervised spatio-temporal data completion system based on temporal semantic attention, including: The time-series segmentation modeling module obtains spatio-temporal data and divides incomplete time series into semantic segment units through time-series segmentation modeling; The segmented attention calculation module uses the cross-attention mechanism to mine the potential time-series dependence patterns between segments and obtains the segmented attention calculation results; The time-series contrast model construction module constructs a time-series contrast model based on the segmented attention calculation results and obtains strong generalization embeddings for time-series modeling through self-supervised methods; The spatial domain modeling enhancement module uses the embeddings obtained through time-series modeling as the features of each node and performs spatial domain modeling through a graph convolutional network to obtain enhanced results.

[0030] The present invention also discloses a computer storage medium, on which a computer program is stored, and when the computer program is executed, the above-mentioned self-supervised spatio-temporal data completion method based on time-series semantic attention is implemented.

[0031] Compared with the prior art, the present invention has the following significant advantages: The present invention adopts node-level time-series cross-attention and self-supervised modeling. Time-series data is first segmented. Segmentation not only makes the modeling of time-series semantics more complete compared to single time points, but also is more reasonable for the data completion problem in terms of block time-series modeling. Using the surrounding mean to complete single time point reconstruction can achieve good results, and segment time-series reconstruction can fully consider long time-series relationships. Through segmented cross-attention for time-series inter-segment modeling, different time-series cycles are considered in time-series modeling. During the calculation process of attention, the original point-to-point attention calculation is replaced by segment-to-segment calculation, greatly saving the calculation amount. And through the self-supervision of time-series data between nodes, the generalization of time-series modeling is enhanced.

[0032] The present invention proposes a strong generalization feature embedding framework for time-series dependence breaks, aiming to achieve robust representation of missing data through segmented modeling and self-supervised learning. This method breaks through the strong dependence of traditional models on data integrity, provides a high-robustness time-series embedding basis for subsequent spatio-temporal node-level tasks, and has important application value in scenarios where data loss occurs frequently, such as smart cities and the Internet of Things. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a flowchart of the self-supervised spatio-temporal data completion method based on time-series semantic attention of the present invention; Figure 2 is the self-supervision principle diagram of node-level time-series cross-attention in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0034] The principle of the present invention is as follows: Through time-series segmentation modeling, the present invention divides incomplete time series into semantic segment units, uses the cross-attention mechanism to mine potential temporal dependence patterns between segments, and alleviates the impact of data breaks on feature integrity. Secondly, based on the segmentation results, a multi-scale self-supervised task is constructed, and by explicitly modeling sudden events, periodic trends, and long-term dependence relationships, the model's ability to capture complex temporal patterns is enhanced. Finally, a segment-to-segment reconstruction loss function is designed to synchronously improve the feature reconstruction accuracy and the model generalization performance by constraining the consistency between local segment features and global temporal semantics.

[0035] Embodiment 1: A self-supervised spatio-temporal data completion method based on temporal semantic attention, comprising the following steps: S01: Obtain spatio-temporal data, and divide the incomplete time series into semantic segment units through time-series segmentation modeling; S02: Use the cross-attention mechanism to mine potential temporal dependence patterns between segments to obtain the segmentation attention calculation result; S03: Build a temporal contrast model based on the segmentation attention calculation result, and obtain a strong generalization embedding for temporal modeling through a self-supervised method; S04: Use the embedding obtained through temporal modeling as the feature of each node, and perform spatial domain modeling through a graph convolutional network to obtain an enhanced result.

[0036] The following takes a preferred embodiment as an example to illustrate in detail the self-supervised spatio-temporal data completion method based on temporal semantic attention, which specifically includes the following steps, as Figure 1 、 Figure 2 shown: S1: Input the spatio-temporal data with missing values into the model. First, the time series data is segmented and divided into semantic segment units according to the time series to ensure semantic integrity during the modeling process.

[0037] After the time series data is segmented, the cross-attention mechanism is used to mine potential temporal dependence patterns for the time series segments.

[0038] S3: Build a temporal contrast model based on the segmentation attention calculation result, and obtain a strong generalization embedding for temporal modeling through a self-supervised method.

[0039] S4: The embedding obtained through temporal modeling is the feature of each node, and spatial domain modeling is further realized through a graph convolutional network.

[0040] S5: Design a segment-to-segment reconstruction loss function to constrain the consistency between local segment features and global temporal semantics.

[0041] Specifically, S1: First, the embedding of the node time series information is performed, and this process mainly includes three stages.

[0042] S11: In the long short-term memory encoding stage, most existing deep learning methods split spatio-temporal data through batch processing techniques, destroying the long-term temporal dependencies in each time series. To regain long-term dependencies over more than one sequence length, a new encoding mechanism is designed, and the encoding result contains global time information (temporal dependencies over more than one sequence length). For example, "a certain hour of the day" and "a certain day of the week" are encoded to enhance the ability to capture long-term temporal dependencies. Different from the positional encoding in the transformer that only processes the temporal dependencies within each sequence, the proposed module processes global temporal dependencies. For the input traffic data , where represents the number of nodes, and represent the dimension and length of the time series features respectively. The concatenation process is shown in Equation 1. After concatenating the time information, we get , greater than , indicating the new time series feature dimension.

[0043] (1) where represents the concatenation operation, represents the time information, such as periodic information like "a certain hour of the day", "a certain day of the week", etc.

[0044] S12: In the short-term temporal dependency mapping encoding stage, the present invention intends to use mapping encoding to encode time to ensure order agnosticism. The mapped time encoding only considers the time information within a day. In this paper, positional encoding is used to input the time information of each time series, and the calculation process is shown in Equation 2.

[0045] (2) where, is the index of the th time point in the time series, is the time mapping in day units. represents the sine function, represents the cosine function. represents the concatenation operation. is the sine encoding of the time series, is the cosine encoding of the time series. The present invention concatenates and as the final time encoding . The calculation of the temporal dependency is shown in Equation 3.

[0046] (3) S13: In the time series segmentation stage, for each input channel First, it is normalized to zero mean and unit standard deviation through invertible instance normalization. Then, is segmented into consecutive non-overlapping blocks , where is the number of blocks, is the original dimension of each block, the time series length is , the block length is , so the total number of input blocks is , which is calculated as shown in Equation 4, where represents the horizontal sliding step size.

[0047] (4) The present invention embeds as , and uses a simple linear layer as the block embedder to create a dimension . Since the complexity of attention is quadratic with the number of time series segments. If each segment represents an attention calculation point instead of each time point representing a calculation point, this obviously reduces the number of calculation points and overall reduces the computational amount of attention.

[0048] S2: The cross-attention between segments mainly includes two stages. First, the attention calculation between time series segments, and then the calculation of the reconstruction loss function for the segment time series.

[0049] S21: The cross-attention between segments is a key module for time series modeling. The attention between segments uses time series segments as the basic calculation unit instead of single time points, and models semantics more fully. The present invention intends to represent the input of each segment-related module as , where and are the length and dimension of the input respectively, and the initial hidden state is . Formally, for the case of single-head attention, the input sequence will be projected by three projection matrices to obtain the query, key, and value, that is, , , , , , are the projection matrices.

[0050] The correlation measure between the query segment and the key segment can be calculated by Equation 5.

[0051] (5) Among them, It is the dot product operation between two matrices of the same size.

[0052] For each query segment , its relevance measurement with all key segments will be normalized by a function to obtain the aggregated weights , and the calculation process is shown in Equation 6.

[0053] (6) The output at the th segment position is the weighted sum of all value segments

[0054] (7) Finally, by concatenating all the outputs the output of the inter-segment attention module is obtained, and the calculation process is shown in Equation 8.

[0055] (8) The final output result after passing through the linear layer is . For the segment length, it can be customized according to the temporal modeling semantics to achieve multi-scale segment attention modeling.

[0056] S22: In the part of the segmented temporal reconstruction loss function, for the entire temporal data of the segmented cross-modeling, the present invention intends to adopt the reconstruction loss. The segmented temporal reconstruction loss function is calculated as shown in Equation 9, and its goal is to minimize the overall loss of each segment.

[0057] (9) Among them, represents the total number of reconstructed segments, represents the true value of the reconstructed segment, represents the temporal cross-modeling result of the reconstructed segment.

[0058] S3: In the temporal contrast learning stage, although the present invention adopts segmented cross-modeling and segment reconstruction, the supervision information under data missing still makes the model lack generalization. The present invention intends to increase temporal contrast learning to enable the model to model macroscopic information. The specific means is to generate strong and weak samples from the segmented cross-modeling results, and then train the model through cross-contrast learning. The generalization of temporal modeling is improved through co-training.

[0059] Taking the intermediate representation of the segmented cross-modeling as the input spatio-temporal data Generate strong samples and weak samples through the augmentation strategy in contrastive learning. And represent its strongly augmented view as , and represent its weakly augmented view as . For the input spatio-temporal data , the encoder will map to a high-dimensional latent representation . Adopt a one-layer neural network for encoding, define , and obtain the strongly augmented encodings and of the two corresponding views and the weakly augmented encoding .

[0060] Context vector generation stage: Through the autoregressive model encode and summarize all the temporal data before the moment in the segment into the context vector , , , , where is 's hidden dimension.

[0061] The mutual information stage depends on the context vector generated in the previous step. To predict future time steps, use the log bilinear model to preserve the mutual information between the input and , making . Among them, is the exponential function, is the linear transformation parameter.

[0062] The present invention intends to adopt Transformer as the autoregressive model, stack L identical layers to generate the final features. Add a token to the input, and its state acts as the representative context vector in the output. First, apply the feature to the linear projection layer, which maps the feature to the hidden dimension, that is . Then, send the output of this linear projection to Transformer to obtain the output result . After the input feature contains the token, it becomes , where the subscript 0 indicates that it is the input of the first layer. Next, pass to the Transformer layer, as shown in formula 13.

[0063] (13) Finally, reattach the token value to the final output, making This context vector will be the input to the contrast module.

[0064] S4: In the spatial domain modeling stage, calculate the learned adaptive adjacency matrix : ;

[0065] Where is the activation function; Since the predefined adjacency matrix is constructed based on expert experience (topology-based or distance-based) and cannot fully reflect the complex correlations between nodes. Therefore, combine the learned adaptive adjacency matrix and to define multi-directional multi-graph spatial convolution, including the forward spatial adjacency matrix and the backward adjacency matrix . The calculation process of spatial domain modeling is shown in Equation 10.

[0066] (10) Where is the hidden layer state of the layer, 、 and are linear transformation parameters. Then use residual connection and batch normalization to obtain the final enhanced result . Take and the ground truth The mean absolute error between as the loss function for the spatio-temporal enhancement task.

[0067] Multi-directional multi-graph spatial convolution includes graph convolutional networks in the forward, backward, and multiple spatial graphs.

[0068] The loss function in the spatial domain modeling stage is: ;

[0069] Where represents the number of points involved in the spatial domain loss calculation, is the output result of the spatial domain modeling, is the label value corresponding to the spatial domain modeling.

[0070] S5: In the process of calculating the loss function, relying on strong enhancement , weak enhancement , the present invention predicts the future time steps of weak enhancement by using the context of strong enhancement ​ Conversely, the contrastive loss attempts to minimize the dot product between the predicted result and the true result of the same sample, while maximizing the dot product with other samples within the same augmented batch. Therefore, the two loss functions and are as shown in Equations (11) and (12).

[0071] (11) (12) Where represents the th order, represents the th moment. represents the linear transformation of , represents the exponential function.

[0072] Under the temporal cross-contrast model, a context contrast module is designed. Given input samples, contexts will be obtained from its two augmented views. For context , is represented as the positive sample of . Therefore, is regarded as a positive pair. At the same time, the remaining contexts from other inputs in the same batch are regarded as 's negative samples. Therefore, a context contrast loss can be derived to maximize the similarity between positive pairs and minimize the similarity between negative pairs. defines the context contrast loss function, and the calculation process is as shown in Equation (14). Given a context , divide its similarity with the positive sample by its similarity with all other samples (including positive pairs and negative pairs) to normalize the loss.

[0073] (14) Where , represents the normalization of the dot product between two vectors and (i.e., cosine similarity), is the indicator function, which takes the value of 1 if and only if , is the temperature parameter. The overall self-supervised loss is a combination of two temporal contrast losses and the context contrast loss, and the total loss function of the temporal cross-contrast learning module is as shown in Equation (15).

[0074] (15) Among them, , and are hyperparameters representing the relative weights of each loss. is a piecewise reconstruction loss function, and is a time-domain cross loss, is a maximum similarity loss function.

[0075] Another embodiment is a computer storage medium storing a computer program, which when executed implements the above-mentioned self-supervised spatio-temporal data completion method based on temporal semantic attention. The specific implementation is the same as the above method and will not be elaborated here.

[0076] Another embodiment is a self-supervised spatio-temporal data completion system based on temporal semantic attention, including: A temporal segmentation modeling module that obtains spatio-temporal data and divides incomplete time series into semantic segment units through temporal segmentation modeling; A segmented attention calculation module that uses a cross-attention mechanism to mine potential temporal dependence patterns between segments to obtain a segmented attention calculation result; A temporal contrast model construction module that constructs a temporal contrast model based on the segmented attention calculation result and obtains strong generalization embeddings for temporal modeling through a self-supervised method; An airspace modeling enhancement module that uses the embeddings obtained through temporal modeling as the features of each node and performs airspace modeling through a graph convolutional network to obtain an enhancement result.

[0077] The specific implementation is the same as the above method and will not be elaborated here.

[0078] The missing completion accuracy of a method for completing urban spatio-temporal data based on sequential segment attention and spatio-temporal contrast learning of the present invention is verified under different datasets and different missing rates. The datasets mainly include traffic speed time series data of fixed points on the road network, and there are certain sparsity and missing in the monitoring data. The datasets used include: PEMS-Bay consists of traffic speed time series of 325 sensors in the Bay Area from January 1, 2017 to May 31, 2017, containing 52,116 time monitoring points; METR-LA is a four-month traffic speed dataset from 207 sensors in Los Angeles, with 1,515 edges, including 119 days and containing 34,272 = 12 * 24 * 119 time monitoring points, and its own missing rate is 8.1%; PEMS04 contains a total of 307 monitoring detectors, and each detector collects data with 3-dimensional features each time, with a total of 16,992 time monitoring points, 16,992 = 59 × 24 × 12 (data is collected every 5 minutes, so it can be collected 12 times within an hour, and there are 24 hours in a day, and it is collected for 59 days, so it is 59 × 24 × 12 = 16,992); in addition, there are also datasets PEMS03, PEMS07, and PEMS08. Examples and data characteristics of each dataset are shown in Table 1.

[0079] Table 1 Statistical Information of Datasets

[0080] For the project evaluation metrics, the present project intends to adopt MSE, RMSE, and R2 for the reconstruction loss function; for contrast learning, the maximum similarity is used as the loss measurement. Among them, the baseline contrast methods mainly include relevant mainstream methods such as M-RNN, GAIN, E2GAN, NAOMI, KCN, and IGNNK.

[0081] The experimental results are shown in Table 2: Table 2 Experimental Results

[0082] The experimental results verify the overall performance of the present invention under various continuous missing conditions; the invention is superior to relevant mainstream methods respectively in terms of the repair accuracy of continuous missing data.

[0083] The above embodiments are the preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A self-supervised spatio-temporal data completion method based on temporal semantic attention, characterized in that It includes the following steps: S01: Obtain spatio-temporal data, and divide the incomplete time series into semantic segment units through time series segmentation modeling; S02: Utilize the cross-attention mechanism to mine the potential temporal dependence patterns between segments, and obtain the segmented attention calculation results; S03: Construct a temporal contrast model based on the segmented attention calculation results, and obtain strong generalization embeddings for temporal modeling through self-supervised methods; S04: Use the embeddings obtained through temporal modeling as the temporal features of each node, and perform spatial domain modeling through a graph convolutional network to obtain the spatial domain modeling results.

2. The self-supervised spatio-temporal data completion method based on temporal semantic attention according to claim 1, wherein After step S04, it also includes designing a segment-to-segment reconstruction loss function to constrain the consistency between local segment features and global temporal semantics.

3. The self-supervised spatio-temporal data completion method based on temporal semantic attention according to claim 1, wherein The method for dividing the incomplete time series into semantic segment units through time series segmentation modeling in step S01 includes: S11: Perform long-short term dependence encoding. The encoding result contains global time information for the input time-series data , where is a real number, represents the number of nodes, and represent the dimension and length of the time-series features respectively. After splicing the time information, data with global time information is obtained , is greater than , indicating the new time-series feature dimension: , Among them, represents a splicing operation, represents time information including global time information; S12: Perform short-term time dependence mapping encoding. The mapping encoding only considers the time information in a day, and uses positional encoding to input the time information of each time series; S13: In the time series segmentation stage, for each input channel is first normalized to zero mean and unit standard deviation by invertible instance normalization; then, is segmented into consecutive non-overlapping blocks , where is the number of blocks, is the original dimension of each block, and then they are embedded into . The embedding method uses a simple linear layer as the block embedder, and its dimension is .

4. The self-supervised spatio-temporal data completion method based on temporal semantic attention according to claim 3, characterized in that The method for utilizing the cross-attention mechanism to mine the potential temporal dependence patterns between segments in step S02 includes: S21: The segment - level attention takes time segments as the basic computational unit, and represents the input of each segment - related module as , where and are the length and dimension of the input respectively, and the initial hidden state . For single - head attention, the input sequence is projected by three projection matrices to obtain the query , the key and the value ; Query segment and key segment The correlation metric between is as follows: , Among them, is the dot product operation between two matrices of the same size; Each query segment Relevance measures with all key segments will be normalized by a function to obtain the aggregated weights : , Output of the th segment position is the weighted sum of all value segments and the calculation process is as follows: , By splicing all outputs Obtain the output of the inter-segment attention module. The calculation process is as follows: , The final output result after passing through the linear layer is , achieving multi-scale segment attention modeling; S22: Calculate the piecewise temporal reconstruction loss function It is as follows: , Among them, represents the total number of reconstructed segments, represents the true value of the reconstructed segment, represents the time-series cross-modeling result of the reconstructed segment; The objective of the segmented temporal reconstruction loss function is to minimize the overall loss of each segment.

5. The self-supervised spatio-temporal data completion method based on temporal semantic attention according to claim 4, wherein The method for obtaining strong generalization embeddings for temporal modeling through self-supervised methods in step S03 includes: Segmented cross-modeling intermediate representation Spatio-temporal data as input , generate strong and weak samples through the augmentation strategy in contrastive learning, that is, represent its strong augmented view as , represent its weak augmented view as ; then encode the augmented views. First, the encoder maps them to a high-dimensional latent representation , Use a one-layer neural network for encoding, and the encoding result is represented as , to obtain the strong augmented encodings and of the corresponding two views and the weak augmented encoding ; Through an autoregressive model Encode and summarize all the time-series data before the time point in the segment into a context vector , , , where is the hidden dimension of Using a logarithmic bilinear model Retain the input and The mutual information between them is obtained as , where is an exponential function is a linear transformation parameter 6. The self-supervised spatio-temporal data completion method based on temporal semantic attention according to claim 5, wherein Step S03 also includes: Autoregressive model Using a Transformer, stack L identical layers to generate the final features, and add a token to the input whose state serves as a representative context vector in the output; Apply the feature to a linear projection layer that maps the feature to a hidden dimension, i.e., ; then, send the output of this linear projection to the Transformer to obtain the output result , and the input feature becomes after including the token, where the subscript 0 represents the input of the first layer; next, send to the Transformer layer as shown in the following formula: , , Among them, , represents the result of the th layer. is the output of the attention calculation and is an intermediate result. represents the multi-head attention calculation. represents the normalization of the values. represents the multi-layer perceptron; Re-attach the token value from the final output so that is used as the input to the time series comparison model.

7. The self-supervised spatio-temporal data completion method based on temporal semantic attention according to claim 5, wherein The method for performing spatial domain modeling through a graph convolutional network in step S04 includes: Compute the learned adaptive adjacency matrix : , Among them, is an activation function; Obtain a predefined adjacency matrix , combine the adaptive adjacency matrix and to define multi-directional multi-graph spatial convolution, including the forward spatial adjacency matrix and the backward adjacency matrix , and the calculation process of spatial domain modeling is as follows: , Among them, is the hidden layer state of the layer, is the result of airspace modeling, , and are linear transformation parameters; The final enhanced result is obtained by using residual connections and batch normalization ; Loss function in the airspace modeling stage is as follows: , Among them, represents all the points participating in the airspace loss calculation, is the output result of airspace modeling, is the label value corresponding to the airspace modeling.

8. The self-supervised spatio-temporal data completion method based on temporal semantic attention according to claim 2, wherein The segment-to-segment reconstruction loss function includes: a segment reconstruction loss function, a temporal domain cross loss function, and a maximum similarity loss function; In the time-domain cross-loss function, strong augmented context vectors / weak augmented context vectors are used to cross-predict weak / strong augmented future time steps . The time-domain cross-loss function minimizes the dot product between the prediction result and the ground truth of the same sample, while maximizing the dot product with other samples in the same augmented batch. The two loss functions and are formulated as follows: , , Among them, represents the time step, represents the total number of future time steps, represents the th moment, represents a linear transformation, represents an exponential function, represents the prediction result of the weakly enhanced future moment , represents the prediction result of the strongly enhanced future moment ; For the maximum similarity loss function, under the temporal contrast model, given input samples, obtain contexts from two augmented views. For context , denote as the positive sample of . Take as the positive pair. At the same time, the remaining contexts from other inputs in the same batch are regarded as negative samples, and a context contrast loss is derived to maximize the similarity between positive pairs and minimize the similarity between negative pairs, obtaining the context contrast loss function : , Among them, , represents the normalization of the dot product between two vectors and , is an indicator function that takes the value of 1 if and only if , is the temperature parameter; The total loss function of the temporal cross-contrast learning module is: , Among them, , and are hyperparameters representing the relative weights of each loss, is the piecewise reconstruction loss function, and is the time-domain cross loss, is the maximum similarity loss function.

9. A self-supervised spatio-temporal data completion system based on temporal semantic attention, characterized in that, It includes: A temporal segmentation modeling module that obtains spatio-temporal data and divides the incomplete time series into semantic segment units through time series segmentation modeling; A segmented attention calculation module that utilizes the cross-attention mechanism to mine the potential temporal dependence patterns between segments and obtains the segmented attention calculation results; A temporal contrast model construction module that constructs a temporal contrast model based on the segmented attention calculation results and obtains strong generalization embeddings for temporal modeling through self-supervised methods; A spatial domain modeling enhancement module that uses the embeddings obtained through temporal modeling as the features of each node and performs spatial domain modeling through a graph convolutional network to obtain enhanced results.

10. A computer storage medium, on which a computer program is stored, characterized in that, When the computer program is executed, it implements the self-supervised spatio-temporal data completion method based on temporal semantic attention described in any one of claims 1-8.

Citation Information

Patent Citations

  • Time-series-data completion method based on distance matrix

    CN108228832A

  • Time series data completion method and device, storage medium and electronic equipment

    CN117807380A

Cited By

  • UPS (Uninterrupted Power Supply) data complementation method for fusing bidirectional space-time diagram convolution with MLP (Multiplayer Library

    CN121188366A