A spatiotemporal large model training method and system based on spatiotemporal information elements, an electronic device, and a readable storage medium

CN122886705APending Publication Date: 2026-10-09ZHEJIANG SHIZIZHIZI BIG DATA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611014279.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-10-09

AI Technical Summary

Technical Problem

[0007]本发明的目的在于提供一种面向各类时空监测场景的通用大模型训练方法,克服现有方案无法在输入端深度融合多源异构传感器数据、难以处理动态置信度差异、时序归一化中插补值导致统计量崩塌及信息泄露,以及稀疏事件学习困难的缺陷

Benefits of technology

1. 动态置信度驱动的信息无损前端融合:通过引入动态置信度掩码,为Transformer架构提供了一种能够精细化区分异构传感器动态观测质量的标准化数据结构,显著减少噪声引入。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122886705A_ABST
    Figure CN122886705A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on space-time information element space-time big model training method, system, electronic equipment and readable storage medium, including to the time alignment processing of the space-time data collected, obtain equidistant time series data;For each time step of each entity, construct effective state vector, and generate dynamic confidence mask, based on the time difference of adjacent standard time step constructs time interval matrix;Generation entity's spatial representation vector;According to the feature distribution before current time step, the adaptive normalization of current input feature is carried out;Extract standardized atomic event tuple, construct two-stage auxiliary supervision task, fuse entity type, spatial representation vector and normalized state vector into information element, form information element sequence along time step and entity dimension, and simultaneously give the dynamic confidence mask and the time interval matrix;Initial model based on Transformer encoder is constructed, and the flattened information element sequence is used as input to train.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of spatiotemporal data mining, deep learning and multimodal perception technology, and specifically relates to a spatiotemporal large model training method, system, electronic device and readable storage medium based on spatiotemporal information elements. Background Technology

[0002] In modern urban, industrial, and natural environmental monitoring systems, widely deployed sensing devices such as sensors, cameras, navigation and positioning terminals, and event recording systems continuously generate massive amounts of spatiotemporal data. While these data inherently possess spatiotemporal attributes, they generally face the following technical challenges: 1. Non-uniform and asynchronous sampling: Video data is usually at a fixed frame rate, while the trajectory sampling interval of mobile terminals varies with the motion state. The reporting frequency of environmental monitoring sensors is affected by network or sleep strategies, and event records are completely sparse and random. Simply aligning different data sources by timestamp will cause a dramatic increase in sequence length or a large amount of information loss.

[0003] 2. Dynamic Confidence Differences: Different data sources have varying confidence levels for the same physical quantity, and the confidence level of the same data source can change dynamically with the physical environment (e.g., GPS signal strength drops sharply when entering a tunnel, and gas sensor confidence decreases under extreme temperature drift). Traditional methods use static weights or simply ignore them, making it difficult to adaptively utilize this heterogeneous and dynamically changing information within a unified model framework.

[0004] 3. Difficulty in learning long-tail events: In real-world monitoring, truly valuable abnormal events (such as traffic accidents, sudden equipment failures, and pollution exceeding standards) occur very infrequently. Data-driven models rely solely on dense prediction tasks and struggle to learn effective sparse event representations.

[0005] Existing Transformer-based spatiotemporal prediction schemes typically encode different modalities separately before performing cross-attention fusion, or simply input all timestamp data directly as the information element sequence. The former results in the loss of fine-grained information at the lower levels before cross-modal interaction; the latter cannot effectively handle non-uniform sampling, and all information elements are treated equally, making it difficult for the model to distinguish between true observations and imputed values. Furthermore, conventional normalization methods (such as LayerNorm and BatchNorm) inevitably introduce future information leakage in time-series scenarios, and in spatiotemporal streams with a large number of missing imputed values, directly using imputed values ​​to update normalized statistics can easily lead to statistical collapse in the early stages of training; while encoding methods based on fixed spatial grids or static graphs cannot adapt to the dynamic changes in the motion state of entities; and purely learnable position encoding cannot be extrapolated to longer sequences not seen during training.

[0006] Therefore, how to construct a unified input representation for multi-source, non-uniform physical sensing data streams with dynamic confidence differences, which can both preserve the original fine-grained information and embed physical priors, and is suitable for efficient consumption by Transformers, and on this basis, design robust normalization and pre-training tasks without information leakage, is a technical problem that urgently needs to be solved in this field. This invention is proposed to solve this problem. Summary of the Invention

[0007] The purpose of this invention is to provide a general large-scale model training method for various spatiotemporal monitoring scenarios, overcoming the shortcomings of existing solutions such as the inability to deeply fuse multi-source heterogeneous sensor data at the input end, difficulty in handling dynamic confidence differences, statistical collapse and information leakage caused by interpolation values ​​during time-series normalization, and difficulty in learning sparse events. By constructing a standardized spatiotemporal information element sequence containing a dynamic confidence mask, and introducing adaptive time-series normalization based on confidence-scaled historical moments and prior-guided event contrast learning, a spatiotemporal pre-trained model with fine-grained perception and long-tail event capture capabilities is finally trained.

[0008] To address the aforementioned technical problems, this invention provides a method for training a large spatiotemporal model based on spatiotemporal information elements, comprising the following steps: The collected spatiotemporal data is time-aligned to obtain equally spaced time-series data, and missing observations are interpolated. Construct an effective state vector for each entity at each time step and generate a dynamic confidence mask. Construct a time interval matrix based on the time difference between adjacent standard time steps. A multi-level geographic grid partitioning system is constructed, and contextual representations at multiple spatial scales are obtained based on the type and motion characteristics of entities. Spatial representation vectors of entities are generated by dynamically selecting the dominant scale. Adaptive normalization is performed on the current input features based on the feature distribution prior to the current time step; Standardized atomic event tuples are extracted to construct a two-stage auxiliary supervision task. The first stage is to freeze the parameters in the event association strength score calculation layer and use the calculated association strength score to perform threshold binarization into pseudo-labels to train the main network. The second stage is to unfreeze the event association strength score calculation layer, jointly optimize it with the main network at a learning rate lower than that of the main network, and introduce self-supervised contrastive loss to optimize the event representation space. The entity type embedding, spatial representation vector, and normalized state vector are fused into information elements, arranged along the time step and entity dimension to form an information element sequence, and a dynamic confidence mask and the time interval matrix are given simultaneously. An initial model based on a Transformer encoder is constructed and trained using a flattened sequence of information elements as input. The model pre-training tasks include masked state recovery prediction and cross-modal event consistency discrimination.

[0009] Furthermore, the spatiotemporal data originates from multi-source heterogeneous sensor data in the target scene. The multi-source heterogeneous sensor data includes at least two of the following: gridded continuous observation, discrete trajectory recording, and sparse event recording. The physical quantities include position, velocity, or environmental parameters.

[0010] Furthermore, the dynamic confidence mask contains continuous values ​​between 0 and 1. The dynamic confidence mask indicates the physical quantity state at the corresponding time step. The physical quantity state comes from at least one of high-confidence true observations, low-confidence observations, and imputation missing values.

[0011] Furthermore, the methods for generating dynamic confidence masks include: For the position parameter, the horizontal precision factor HDOP is obtained. When HDOP is greater than the preset threshold or the entity enters the preset signal blocking area, the initial confidence is reduced according to the exponential decay function that is positively correlated with the HDOP value. For environmental parameters, when the environmental parameters exceed the calibration range or the time series variance changes abruptly, the initial confidence level is linearly decayed based on the relative deviation rate between the current value and the historical mean. For missing state vectors generated through interpolation, their dynamic confidence mask is set to 0.

[0012] Furthermore, the spatial representation vector generation steps include: Construct a multi-level geographic grid partitioning system and dynamically assign a learnable embedding vector to each specific grid unit that appears during training; Obtain the mesh cell embedding of each entity at each time step at each level; The query vector is a concatenation of the type embedding of each entity and its recent short-term motion features. Using the mesh embeddings corresponding to each level of each entity as keys and values, a multi-scale context embedding is computed through scaled dot product attention; Select a dominant scale mesh embedding based on each entity type and instantaneous velocity; The spatial representation vector is formed by concatenating the dominant scale embedding and the multi-scale context embedding.

[0013] Furthermore, the recent short-term motion characteristics are formed by splicing the average velocity vector and the rate of change of heading from at least three recent time steps.

[0014] Furthermore, the masked state recovery prediction includes: randomly masking the state sub-vectors in some information elements, calculating the loss only for positions where the dynamic confidence mask value is greater than 0, and the loss weight is determined by the dynamic confidence mask value at the corresponding position, enabling the model to learn to use effective observations to predict missing or future information. The formula for calculating the loss of a single sample is: ; Confidence weighting further reduces the interference of low-quality observations on reconstruction loss.

[0015] Cross-modal event consistency determination includes: using pseudo-labels as supervision signals to predict whether event pairs belong to the same real event, forcing the model to learn event-level representations.

[0016] Furthermore, specific implementations of adaptive timing normalization include: Maintain two exponential moving average shadow variables: historical first moment The squared expectation of the uncentered second-order moment of history .

[0017] Based on the input features of the previous time step and its corresponding dynamic confidence mask Calculate the effective update step size and perform a stopping gradient operation to update the shadow variables: in This indicates stopping gradient operations to cut off the backpropagation path of variance collapse, with α being the base smoothing coefficient. The historical variance is then calculated from this:

[0018] Input the current time step Standardization is performed using statistics that have been updated based on historical information: Finally, the normalized state representation is obtained through the learnable affine parameters γ and β.

[0019] Furthermore, the step of constructing the information element sequence includes: generating a spatiotemporal information element (ST-TOKEN) for each entity at each standard time step, the vector structure of which is as follows:

[0020] in, Embedding for entity types, For spatial representation of vectors, For the normalized state representation, all information elements are arranged into a sequence in (time step, entity) order, and the dynamic confidence mask matrix and the time interval matrix are synchronously attached.

[0021] Furthermore, the Transformer encoder employs time-interval-aware hybrid position encoding. The construction method of time-interval-aware hybrid position encoding includes mapping the values ​​of the time interval matrix to the same dimension as the position encoding through a learnable linear projection matrix, then adding them element-wise to the basic sinusoidal position encoding, and then appending them to the information element representation of the corresponding time step. The self-attention module of the Transformer model uses a dynamic confidence mask to ignore information elements at missing positions and uses the inverse of the time interval matrix to scale the attention weights.

[0022] This invention also provides a spatiotemporal pre-training model training system comprising: The data alignment and dynamic confidence module is used to perform the data alignment steps; The spatial encoding module is used to perform spatial encoding steps; The information element generation module is used to execute the information element generation steps. The adaptive normalization module is used to execute the adaptive temporal normalization method; The pre-training module is used to perform the model pre-training steps.

[0023] The present invention also provides an electronic device comprising a processor and a memory, wherein the memory stores a computer program, and the processor executes the computer program to implement the spatiotemporal large model training method based on spatiotemporal information elements of the present invention.

[0024] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the spatiotemporal large model training method based on spatiotemporal information elements of the present invention.

[0025] The beneficial effects of this invention are: 1. Lossless front-end fusion driven by dynamic confidence: By introducing a dynamic confidence mask, a standardized data structure is provided for the Transformer architecture that can finely distinguish the dynamic observation quality of heterogeneous sensors, significantly reducing the introduction of noise.

[0026] 2. Solving the problems of information leakage and statistical collapse in time series normalization: The proposed AHN based on confidence scaling of historical moments not only fundamentally eliminates the inflow of future information, but also avoids statistical pollution caused by low-quality imputation values ​​by scaling the effective step size of confidence. Stopping gradients blocks the backpropagation path of variance collapse, which significantly improves the prediction stability in long series and high missing rate scenarios.

[0027] 3. Dynamic spatial representation with embedded physical priors: By using the kinematic features of entities as queries for dynamic aggregation and multi-scale grid embedding, the model is embedded with the physical priors of "the faster the movement, the more macroscopic the scale of concern".

[0028] 4. Solving the cold start difficulty of long-tail event learning: Two-stage event comparison learning overcomes the cold start difficulty of sparse event representation learning through the freezing and unfreezing mechanism of the network layer.

[0029] 5. Unification of length extrapolation and non-uniform modeling: Time interval-aware hybrid position encoding enables the model to perfectly adapt to non-uniform sampling and generalize to ultra-long sequences. Attached Figure Description

[0030] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0031] Figure 1: Overall flowchart of an embodiment of the present invention.

[0032] Figure 2 : A schematic diagram of dynamic multi-scale adaptive spatial coding according to an embodiment of the present invention.

[0033] Figure 3 An embodiment of the present invention describes a cross-modal event association and a two-stage training process.

[0034] Figure 4 A schematic diagram of the structure of a spatiotemporal element sequence and a dynamic confidence mask matrix according to an embodiment of the present invention.

[0035] Figure 5 The following is an overall architecture diagram of an embodiment of the present invention, highlighting the AHN layer and hybrid positional encoding. Detailed Implementation

[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0037] The following description, in conjunction with specific implementation methods, provides an explanation.

[0038] like Figure 1 - Figure 5 As shown, a preferred embodiment of the present invention specifically includes the following steps: Step A: Data preprocessing and alignment, timestamp calibration / resampling This step performs time alignment processing on the collected spatiotemporal data to obtain equally spaced time-series data and interpolates missing observations; constructs an effective state vector for each entity at each time step and generates a dynamic confidence mask; and constructs a time interval matrix based on the time difference between adjacent standard time steps.

[0039] Preferably, the system accesses multi-source heterogeneous sensor data from the target scene. The sensor data is preferably trajectory / video / environmental event records, including at least two of the following: rasterized continuous observations, discrete trajectory records, and sparse event records. Physical quantities include position, velocity, or environmental parameters. After uniformly calibrating all data timestamps, resampling is performed at preset equal-interval standard time steps. For each data source, the observation is assigned to the nearest standard time step; if no observation is found, it is marked as missing and imputed.

[0040] The key is to generate an effective state vector for each entity at each standard time step, and simultaneously generate a dynamic confidence mask. This mask is a continuous value between 0 and 1, used to indicate whether the physical quantity state at that time step comes from a high-confidence true observation, a low-confidence observation, or imputation of missing data. The true time difference between adjacent standard time steps is recorded to form a time interval matrix.

[0041] Step B: Adaptive selection of multi-scale adaptive spatial coding grid layers.

[0042] A multi-level geographic grid partitioning system is constructed, and contextual representations at multiple spatial scales are obtained based on the type and motion characteristics of entities. By dynamically selecting the dominant scale, spatial representation vectors of entities are generated.

[0043] like Figure 2 As shown, a multi-level geographic grid partitioning system is constructed, and a learnable embedding vector is dynamically assigned to each specific grid unit that appears during training.

[0044] For each entity at each time step: Obtain its mesh cell embedding at each level.

[0045] The query vector is a concatenation of the entity's type embedding and recent short-term motion features. Preferably, the short-term motion features are formed by concatenating the average velocity vector and the rate of change of heading over at least three recent time steps.

[0046] The entity's various levels of mesh are embedded as keys and values, and a multi-scale context embedding is computed through scaled dot product attention, enabling the entity to dynamically focus on the most relevant spatial scales based on its own motion state.

[0047] Simultaneously, a dominant scale grid is selected for embedding based on the entity type and instantaneous velocity, where fast-moving entities correspond to the macroscopic scale and slow-moving entities correspond to the microscopic scale.

[0048] The final spatial representation vector is formed by concatenating the dominant scale embedding and the multi-scale context embedding.

[0049] Step C: AHN time series normalization and low-confidence observation weighting Adaptive normalization is performed on the current input features based on the feature distribution prior to the current time step.

[0050] To address the inherent problems of non-uniform time-series streams with missing interpolation caused by asynchronous sensor sampling and signal occlusion, this invention designs an adaptive time-series normalization method to solve the issues of traditional normalization layers leaking future information and low-quality interpolated values ​​participating in statistical calculations leading to early variance collapse. Its core principle is to normalize the current input using only historical input features from before the current time step, with the influence of historical features dynamically adjusted by their confidence level. Specific implementation: Maintain two exponential moving average shadow variables: historical first moment And the second moment of history (uncentered squared expectation) .

[0051] Before processing time step t, the input features from the previous time step are used. and its corresponding dynamic confidence mask Calculate the effective update step size (so that the effective step size approaches zero when the previous time step contains imputed missing values, thus preventing low-quality imputed values ​​from contaminating historical statistics), and perform a stopping gradient operation to update the shadow variables: in This indicates stopping gradient operations to cut off the backpropagation path of variance collapse, with α being the base smoothing coefficient. The historical variance is then calculated as follows: Input the current time step Standardization is performed using statistics that have been updated based on historical information: Finally, learnable affine parameters γ, β The normalized state representation is obtained.

[0052] Step D: Event extraction and association to generate pseudo-tags We extract standardized atomic event tuples and construct a two-stage auxiliary supervision task. The first stage is to freeze the parameters in the event association strength score calculation layer and use the calculated association strength score to perform threshold binarization into pseudo-labels to train the main network. The second stage is to unfreeze the event association strength score calculation layer, jointly optimize it with the main network at a learning rate lower than that of the main network, and introduce self-supervised contrastive loss to optimize the event representation space. To address the difficulty of learning sparse, long-tailed events, an evolvable auxiliary task is introduced.

[0053] like Figure 3 As shown, event extraction involves extracting standardized atomic event tuples from each source data.

[0054] Two-stage training strategy: Phase 1 (Prior Guidance Phase): Freeze the parameters in the event association strength score calculation layer, and use threshold binarization to convert the calculated association strength scores into pseudo-labels to train the main network, thus solving the cold start problem.

[0055] The second stage (joint fine-tuning stage) involves unfreezing the event association strength score calculation layer, jointly optimizing it with the main network at a learning rate lower than that of the main network, and introducing a self-supervised contrastive loss to optimize the event representation space.

[0056] Step E: Information Element Construction For each entity, a spatiotemporal information element (ST-TOKEN) is generated at each standard time step, and its vector structure is as follows:

[0057] in, Embedding for entity types, The spatial representation generated in step B, This is the state representation generated in step C. All information elements are arranged in sequence according to (time step, entity), and the dynamic confidence mask matrix and time interval matrix generated in the previous step are appended synchronously.

[0058] Step F: Hybrid position encoding fusion time interval matrix Step G: Transformer pre-training state recovery + event consistency An initial model based on a Transformer encoder is constructed. The model input is a flattened sequence of information elements, employing a time-interval-aware hybrid positional encoding: the values ​​of the time interval matrix are mapped to the same dimension as the positional encoding through a learnable linear projection matrix, then added element-wise to the basic sinusoidal positional encoding, and finally appended to the information element representation at the corresponding time step. The self-attention module uses a dynamic confidence mask to ignore information elements at missing positions and scales the attention weights using the reciprocal of the time interval matrix.

[0059] Pre-training includes two core tasks: Task 1: Masked State Recovery Prediction. The state sub-vectors in a portion of the information elements are randomly masked. Loss is calculated only for positions where the dynamic confidence mask value is greater than 0, and the loss weight is determined by the dynamic confidence mask value at the corresponding position. This forces the model to learn to use valid observations to predict missing or future information. The formula for calculating the loss per sample is:

[0060] Confidence weighting further reduces the interference of low-quality observations on reconstruction loss.

[0061] Task 2: Cross-modal event consistency determination. Using the pseudo-labels generated in the first stage of step D as supervision signals, predict whether event pairs belong to the same real event, forcing the model to learn event-level representations.

[0062] The total loss function is a weighted sum of the two losses:

[0063] All learnable parameters of the model are updated end-to-end through backpropagation.

[0064] The output spatiotemporal pre-trained model supports downstream tasks.

[0065] Embodiments of the present invention also provide a spatiotemporal pre-training model training system, comprising: a data alignment and dynamic confidence module for performing the aforementioned data temporal alignment and confidence mask generation; a spatial encoding module for generating spatial representation vectors; an information element generation module for fusing features and constructing an information element sequence; an adaptive normalization module for performing adaptive temporal normalization; and a pre-training module for performing model pre-training tasks. This system can be deployed on an electronic device including a processor and memory, whereby the processor executes the computer program to implement the aforementioned method steps. Alternatively, the program can be stored in a computer-readable storage medium for the processor to call and execute.

[0066] The following provides methods for training large models in specific application scenarios. I. Spatiotemporal Large Model Training Methods in Intelligent Transportation Scenarios 1. Data preprocessing and dynamic confidence alignment Access video cameras, vehicle GPS trajectories, and traffic incident records in the target intersection area. Set the standard time step Δt = 5 seconds. The dynamic confidence mask generation rule is illustrated using a vehicle at time step t as an example: The confidence level for video speed testing is fixed at 1.0; The initial confidence level for GPS speed measurement is 0.8. The horizontal accuracy factor HDOP is obtained. When HDOP is greater than a preset threshold of 2 or an entity enters a predefined signal-blocking area such as a tunnel, the initial confidence level is reduced (down to a minimum of 0.3) according to an exponential decay function that is positively correlated with the HDOP value. For environmental sensor data, when environmental parameters exceed the calibration range or there is a sudden change in time series variance, the initial confidence level is linearly decayed based on the relative deviation rate between the current value and the historical mean. If there is no sensor data at this time step, and historical mean interpolation is used (e.g., speed 15m / s), then the dynamic confidence mask of this state vector is set to 0.

[0067] Record the actual time difference between each adjacent standard time step to form a time interval matrix.

[0068] 2. Adaptive History Normalization (AHN) Implementation and Deduction Initialize first-order moments Second moment Basic smoothing coefficient Assuming Step input features Its confidence level (i.e., imputing missing values).

[0069] Calculate the effective step size: At this point, the effective step size approaches zero, preventing low-quality interpolation values ​​from contaminating historical statistics.

[0070] Update the first moment: ; Update the second moment: .

[0071] As can be seen, when encountering imputed missing values, the statistic remains completely unchanged from the previous step, thoroughly preventing variance collapse. Assume t steps of input... Confidence level Standardize it using the newly updated statistics: .

[0072] During model training, a one-dimensional causal convolution kernel slides along the time dimension to compute all time-step moments based on history in parallel at once, ensuring strict consistency between training and inference and preventing the leakage of future information. At the same time, stopping gradient operations cuts off the backpropagation path of the normalized statistics to the current input gradient, avoiding variance collapse.

[0073] 3. Implementation of Dynamic Multi-Scale Spatial Coding An H3 grid with 8-12 layers is used. For each entity, the average velocity vector and heading change rate of its most recent 3 time steps are maintained, and after being concatenated with the type embedding, they are mapped to a query vector.

[0074] When the speed of the vehicle is >15m / s, the dominant scale is increased to level 8 (macro grid, adapted to rapid movement). When the speed is <2m / s or for pedestrians, the dominant scale is fixed at 12 levels (microgrid).

[0075] Using the mesh embeddings at each level as keys and values, multi-scale contextual embeddings are computed through scaled dot product attention, and then concatenated with the dominant scale embeddings to generate the final spatial representation. .

[0076] 4. Generation of Spatiotemporal Information Elements Embed entity types (64-dimensional) and spatial representation. (Dimension 256), State Representation (After AHN normalization and projection to dimension 192) The concatenation generates a standardized ST-TOKEN (spatial information element) with dimension 512. All information elements are arranged in time step and entity order, and a dynamic confidence mask matrix and time interval matrix are added synchronously.

[0077] 5. Structure and pre-training methods of large spatiotemporal model instances (1) Model structure parameters A large-scale spatiotemporal model based on a Transformer encoder is constructed, comprising 12 encoder layers, 8 attention heads, a feedforward network with 2048 hidden layers, and a total information element dimension of 512. A time-interval-aware hybrid positional encoding is employed: the values ​​of the time interval matrix ΔT are mapped to the same dimension as the positional encoding using a 512×512 learnable linear projection matrix, and then element-wise added to the basic sinusoidal positional encoding. The self-attention module scales the attention weights using the reciprocal of the time interval matrix.

[0078] (2) Batch data construction Each training iteration extracts trajectory segments of 256 entities over 512 consecutive standard time steps (approximately 42 minutes) to form an information element input tensor of shape (256, 512, 512). Simultaneously, a dynamic confidence mask tensor and a time interval matrix of the same shape are generated.

[0079] (3) Pre-training task 1: Prediction of recovery from cover For the input information sequence, its state subvector is randomly masked with a 15% probability. The masked locations are replaced with zero vectors. The model outputs the predicted value for the corresponding location. When calculating the MSE loss, it is only calculated for the true observation locations where the dynamic confidence mask c > 0, and the loss weight is determined by the dynamic confidence mask value at the corresponding location. The formula for calculating the loss of a single sample is: Confidence weighting further reduces the interference of low-quality observations on reconstruction loss.

[0080] (4) Pre-training Task 2: Cross-modal event consistency discrimination (two-stage training) "Vehicle sudden stop" and "traffic accident event" are extracted from the batch data to form positive and negative sample pairs.

[0081] Phase 1 (Prior Guidance Phase, First 500 Iterations): Set initial weight values ​​for spatial overlap, entity co-occurrence, and semantic similarity. Spatial scale parameters The parameters mentioned above in the freeze event association strength score calculation layer are used to calculate the association strength score. and will Pseudo-labels are generated through threshold binarization. This is used to train the main network. The threshold at this stage is dynamically and adaptively adjusted based on the average correlation strength of the current batch of event pairs, blocking immature event representations from interfering with the main network and solving the cold start problem.

[0082] Phase Two (Joint Fine-Tuning Phase, after 500 steps): Unfreeze the event association strength score calculation layer, and... and Set as a learnable parameter, to The learning rate (the main network learning rate) Joint optimization is performed. Simultaneously, a self-supervised contrastive loss is introduced to narrow the distance between positive sample pairs and event representations, while widening the distance between negative sample pairs.

[0083] (5) Total Loss and Optimization Strategy The total loss is:

[0084] The AdamW optimizer was used, with a warm-up step count of 1000, and a cosine annealing learning rate scheduling strategy was employed. All model parameters were updated end-to-end through backpropagation until convergence, resulting in the final spatiotemporal pre-trained model.

[0085] 6. Performance Verification On a test set at a core intersection in a city, compared to the baseline, the speed prediction MAE decreased by 18.0%, and the abnormal parking F1 improved by 11.7%. Ablation experiments demonstrate that: Replacing AHN with regular BN increased MAE by 7%; If AHN does not use confidence scaling step size (i.e., the imputed value is updated with the normal step size), variance collapse occurs in subsets with a missing rate of 40%, and MAE increases by 12%. Disabling dynamic spatial coding reduced F1 by 13 percentage points.

[0086] The effectiveness of the present invention has been fully verified.

[0087] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can also be implemented in other ways. The system embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0088] Furthermore, the functional modules in the various embodiments of this invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part. If the function is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0089] Based on the above-described preferred embodiments of the present invention, and through the foregoing description, those skilled in the art can make various changes and modifications without departing from the inventive concept. The technical scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.

Claims

1. A method for training a large spatiotemporal model based on spatiotemporal information elements, characterized in that, Includes the following steps: The collected spatiotemporal data is time-aligned to obtain equally spaced time-series data, and missing observations are interpolated. Construct an effective state vector for each entity at each time step and generate a dynamic confidence mask. Construct a time interval matrix based on the time difference between adjacent standard time steps. A multi-level geographic grid partitioning system is constructed, and contextual representations at multiple spatial scales are obtained based on the type and motion characteristics of entities. Spatial representation vectors of entities are generated by dynamically selecting the dominant scale. Adaptive normalization is performed on the current input features based on the feature distribution prior to the current time step; Standardized atomic event tuples are extracted to construct a two-stage auxiliary supervision task. The first stage is to freeze the parameters in the event association strength score calculation layer and use the calculated association strength score to perform threshold binarization into pseudo-labels to train the main network. The second stage is to unfreeze the event association strength score calculation layer, jointly optimize it with the main network at a learning rate lower than that of the main network, and introduce self-supervised contrastive loss to optimize the event representation space. The entity type embedding, spatial representation vector, and normalized state vector are fused into information elements, arranged along the time step and entity dimension to form an information element sequence, and the dynamic confidence mask and the time interval matrix are given simultaneously. An initial model based on a Transformer encoder is constructed and trained using the flattened information element sequence as input. The model pre-training tasks include masked state recovery prediction and cross-modal event consistency discrimination.

2. The spatiotemporal large-scale model training method based on spatiotemporal information elements according to claim 1, characterized in that, The spatiotemporal data originates from multi-source heterogeneous sensor data in the target scene. The multi-source heterogeneous sensor data includes at least two types: gridded continuous observation, discrete trajectory recording, and sparse event recording. The physical quantities include position, velocity, or environmental parameters.

3. The spatiotemporal large-scale model training method based on spatiotemporal information elements according to claim 2, characterized in that, The dynamic confidence mask contains a continuous value between 0 and 1, and the dynamic confidence mask indicates the physical quantity state at the corresponding time step. The physical quantity state comes from at least one of high-confidence true observations, low-confidence observations, and imputation missing values.

4. The spatiotemporal large model training method based on spatiotemporal information elements according to claim 3, characterized in that, The dynamic confidence mask is generated in the following ways: For the position parameter, the horizontal precision factor HDOP is obtained. When HDOP is greater than a preset threshold or the entity enters a preset signal blocking area, the initial confidence level is reduced according to an exponential decay function that is positively correlated with the HDOP value. For environmental parameters, when the environmental parameters exceed the calibration range or the time series variance changes abruptly, the initial confidence level is linearly decayed based on the relative deviation rate between the current value and the historical mean. For missing state vectors generated through interpolation, their dynamic confidence mask is set to 0.

5. The method for training a large spatiotemporal model based on spatiotemporal information elements according to claim 1, characterized in that, The spatial representation vector generation step includes: Construct a multi-level geographic grid partitioning system and dynamically assign a learnable embedding vector to each specific grid unit that appears during training; Obtain the mesh cell embedding of each entity at each time step at each level; The query vector is a concatenation of the type embedding of each entity and its recent short-term motion features. Using the mesh embeddings corresponding to each level of each entity as keys and values, a multi-scale context embedding is computed through scaled dot product attention; Select a dominant scale mesh embedding based on each entity type and instantaneous velocity; The spatial representation vector is formed by concatenating the dominant scale embedding and the multi-scale context embedding.

6. The spatiotemporal large model training method based on spatiotemporal information elements according to claim 5, characterized in that, The recent short-term motion characteristics are formed by splicing the average velocity vector and the rate of change of heading from at least three recent time steps.

7. The spatiotemporal large-scale model training method based on spatiotemporal information elements according to claim 1, characterized in that, The masking state recovery prediction includes: partial state information in random masking information elements, and recovery prediction is only performed at the positions where the confidence mask satisfies the valid conditions; The cross-modal event consistency determination: using the pseudo-label as a monitoring signal, it determines whether different event pairs belong to the same real event and have cross-modal consistency.

8. The spatiotemporal large model training method based on spatiotemporal information elements according to claim 1, characterized in that, The specific implementation methods of the adaptive time series normalization include: Maintain two exponential moving average shadow variables: historical first moment The squared expectation of the uncentered second-order moment of history , Based on the input features of the previous time step and its corresponding dynamic confidence mask Calculate the effective update step size and perform a stopping gradient operation to update the shadow variables: in This indicates stopping gradient operations to cut off the backpropagation path of variance collapse, with α being the base smoothing coefficient. The historical variance is then calculated as follows: Input the current time step Standardization is performed using statistics that have been updated based on historical information: Finally, learnable affine parameters γ, β The normalized state representation is obtained.

9. The spatiotemporal large model training method based on spatiotemporal information elements according to claim 1, characterized in that: The steps for constructing the information element sequence include: generating a spatiotemporal information element for each entity at each standard time step, denoted by ST-TOKEN, whose vector structure is as follows: in, Embedding for entity types, For spatial representation of vectors, For the normalized state representation, all information elements are arranged into a sequence in (time step, entity) order, and a dynamic confidence mask matrix and the time interval matrix are synchronously attached.

10. The method for training a large spatiotemporal model based on spatiotemporal information elements according to claim 1, characterized in that, The Transformer encoder employs time-interval-aware hybrid position encoding, and the construction method of the time-interval-aware hybrid position encoding includes: The values ​​of the time interval matrix are mapped to the same dimension as the position code through a learnable linear projection matrix, then added element-wise to the basic sinusoidal position code, and then appended to the information element representation of the corresponding time step. The self-attention module of the Transformer model uses the dynamic confidence mask to ignore information elements at missing positions and uses the reciprocal of the time interval matrix to scale the attention weights.

11. A spatiotemporal pre-training model training system, characterized in that, include: The data alignment and dynamic confidence module is used to perform the data alignment steps; The spatial encoding module is used to perform spatial encoding steps; The information element generation module is used to execute the information element generation steps. The adaptive normalization module is used to execute the adaptive temporal normalization method; The pre-training module is used to perform the model pre-training steps.

12. An electronic device comprising a processor and a memory, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the spatiotemporal large model training method based on spatiotemporal information elements as described in any one of claims 1-10.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the spatiotemporal large model training method based on spatiotemporal information elements as described in any one of claims 1-10.