Target spatial-temporal trajectory pre-training model construction method and device, equipment and medium
By establishing a spatiotemporal trajectory pre-training model based on transformer decoder, the problem that the existing technology is difficult to deal with spatiotemporal trajectory data is solved, and the direct processing and powerful generalization capabilities of spatiotemporal trajectory data are realized, which is suitable for a variety of downstream tasks.
Patent Information
- Application Number
- CN202510265056.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to directly process spatiotemporal trajectory data, and it is impossible to effectively capture complex relationships and patterns in spatiotemporal trajectory data, especially when labeled data is scarce.
Based on the transformer decoder architecture, a spatiotemporal trajectory pre-training model is established. By mixing the embedding layer, multi-layer multi-head masking self-attention layer and fully connected feedforward neural network layer, a spatiotemporal trajectory pre-training data set is constructed, and a composite cross-entropy loss function is used for unsupervised pre-training.
It realizes direct processing of spatiotemporal trajectory data, can show strong generalization capabilities in various spatiotemporal learning scenarios, and is suitable for downstream tasks such as clustering analysis, trajectory classification, behavior recognition, and abnormal detection.
Smart Images

Figure CN120180905A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target spatio-temporal trajectory processing, and more specifically, to a method, device, equipment and medium for constructing a pre-trained model for target spatio-temporal trajectories. Background Art
[0002] A spatio-temporal trajectory is a curve in a multi-dimensional space formed by adding a time axis to a geographical space, which can represent the position changes of a moving object over a relatively long period of time and consists of a series of trajectory points. Currently, the academic community has carried out a large amount of research work on spatio-temporal trajectories, including motion pattern mining, clustering analysis, trajectory classification, behavior recognition, anomaly detection, etc. However, many methods rely heavily on sufficient labeled data to generate accurate spatio-temporal trajectory representations. In actual application scenarios, the problem of data sparsity is widespread. In some cases, it becomes challenging to collect any labeled data from downstream scenarios, further exacerbating the problem. Therefore, it becomes necessary to establish a spatio-temporal model that can exhibit strong generalization ability in various spatio-temporal learning scenarios.
[0003] In recent years, with the development of large model technology, intelligent computing has developed to the fourth stage - large model computing systems, and AI has shifted from "small models + discriminative" to "large models + generative". Large models were first developed based on language large models, constructing a training data set based on a large amount of natural language description texts to achieve the training of basic language large models, and supplemented by supervised fine-tuning, reward modeling, and reinforcement learning techniques to construct generative artificial intelligence products based on conversational interactions, which have been widely applied. An important direction for the development of large models is multi-modal large models. From the perspective of AI, vision, audition, etc. can also be modeled as sequences of tokens, and the same methods as those for large models can be adopted for learning and further aligned with the semantics in language models to form intelligent capabilities with multi-modal alignment.
[0004] Currently, multi-modal large models mainly focus on two modalities: natural language + vision (images, videos). However, in fields such as smart cities, intelligent transportation, and marine transportation, a large amount of target information exists in the form of spatio-temporal trajectories, which is an important structured data modality. Currently, the typical way to integrate spatio-temporal trajectory data into large models is to convert it into text descriptions and then use existing large-scale language models for analysis and processing. Since it is unable to directly process the original spatio-temporal trajectory data, it is difficult to capture the complex relationships and patterns in spatio-temporal trajectory data. How to construct a spatio-temporal trajectory pre-trained model with the target trajectory as the direct input is an urgent problem to be solved for introducing this structured data modality of spatio-temporal trajectories into multi-modal large models. Summary of the Invention
[0005] In view of the above problems, inspired by the remarkable achievements of large language models, the present invention provides a method, apparatus, device and medium for constructing a target spatio-temporal trajectory pre-training model. Referring to the architecture of the generative language-based large model based on transformer, a spatio-temporal trajectory pre-training model is established, and the spatio-temporal trajectory pre-training model is subjected to generative unsupervised pre-training based on a large amount of historical spatio-temporal trajectory data.
[0006] In a first aspect, a method for constructing a target spatio-temporal trajectory pre-training model provided by the present invention includes:
[0007] Establish a spatio-temporal trajectory pre-training model based on a transformer decoder; the spatio-temporal trajectory pre-training model includes a hybrid embedding layer, a multi-layer multi-head masked self-attention layer, and a fully connected feed-forward neural network layer connected in sequence;
[0008] Construct a spatio-temporal trajectory pre-training data set, the spatio-temporal trajectory pre-training data set includes a target spatio-temporal trajectory token sequence;
[0009] Based on the spatio-temporal trajectory pre-training data set, train the spatio-temporal trajectory pre-training model using a composite cross-entropy loss function.
[0010] In some embodiments, the hybrid embedding layer includes a channel embedding module, an embedding vector splicing module, and a linear transformation network module connected in sequence.
[0011] In some embodiments, the hybrid embedding layer is used to convert the input target spatio-temporal trajectory token sequence into an embedding tensor, specifically including: first calculating the embedding tensor of each feature separately, then splicing the embedding tensors of all features, and finally converting the spliced tensor into an embedding tensor with a specified embedding dimension through a non-linear transformation network.
[0012] In some embodiments, the multi-layer multi-head masked self-attention layer includes N transformer decoders connected in sequence; each transformer decoder includes a masked multi-head self-attention module, a first residual connection and layer normalization module, a per-position feed-forward neural network module, and a second residual connection and layer normalization module connected in sequence.
[0013] In some embodiments, the construction of the spatio-temporal trajectory pre-training data set includes:
[0014] Calculate the trajectory resampling time interval;
[0015] According to the trajectory resampling time interval, resample the trajectory data; the trajectory data includes target time, longitude, and latitude information;
[0016] According to the resampled trajectory data, calculate the target speed and heading information;
[0017] Tokenize the features of the target spatio-temporal trajectory to obtain a target spatio-temporal trajectory token sequence; the features of the target spatio-temporal trajectory include resampled trajectory data and target speed and heading information.
[0018] In some embodiments, calculating the trajectory resampling time interval includes:
[0019] Calculate the trajectory resampling time interval based on the target average speed and the maximum spatial distance represented by the longitude and latitude feature encoding lengths of adjacent trajectory points.
[0020] In some embodiments, the composite cross-entropy loss function is defined as: a function for calculating the weighted average of each feature cross-loss value, and its calculation steps are as follows: first, split the logical probability vector output by the spatio-temporal trajectory pre-training model according to the feature dimension, then calculate the cross-entropy loss of each feature respectively, and finally perform a weighted sum of the cross-entropy losses of multiple features.
[0021] In a second aspect, the present invention provides an apparatus for constructing a pre-training model of a target spatio-temporal trajectory, including:
[0022] A model construction unit for establishing a spatio-temporal trajectory pre-training model based on a Transformer decoder; the spatio-temporal trajectory pre-training model includes a hybrid embedding layer, a multi-layer multi-head masked self-attention layer, and a fully connected feed-forward neural network layer connected in sequence;
[0023] A dataset construction unit for constructing a spatio-temporal trajectory pre-training dataset, where the spatio-temporal trajectory pre-training dataset includes a target spatio-temporal trajectory token sequence;
[0024] A model training unit for training the spatio-temporal trajectory pre-training model based on the spatio-temporal trajectory pre-training dataset using the composite cross-entropy loss function.
[0025] In a third aspect, the present invention provides an electronic device, including:
[0026] At least one processor; and a memory communicatively connected to the at least one processor;
[0027] Wherein, the memory stores instructions executable by the at least one processor, and the at least one processor, by executing the instructions stored in the memory, causes the at least one processor to execute the above method.
[0028] In a fourth aspect, the present invention provides a computer-readable storage medium for storing instructions, which, when executed, implement the above method.
[0029] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are as follows:
[0030] 1. The target spatio-temporal trajectory pre-training model of the present invention realizes the joint embedding of multi-attribute information in the target trajectory by introducing a hybrid embedding layer, enabling the target spatio-temporal trajectory pre-training model to directly process spatio-temporal trajectory data formats.
[0031] 2. During the training process of the present invention, a composite cross-entropy loss function is constructed, enabling the target spatio-temporal trajectory pre-training model to perform backpropagation training.
[0032] 3. The target spatio-temporal trajectory pre-training model of the present invention can be used as a basic model in clustering analysis, trajectory classification, behavior recognition, anomaly detection, etc., enabling a wide range of downstream tasks to share its generalization ability to better cope with the scarcity of labeled data. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a flowchart of a method for constructing a target spatio-temporal trajectory pre-training model provided by an embodiment of the present invention.
[0034] Figure 2 is a schematic structural diagram of the target spatio-temporal trajectory pre-training model in an embodiment of the present invention.
[0035] Figure 3 is a flowchart of constructing a target spatio-temporal trajectory pre-training dataset in an embodiment of the present invention.
[0036] Figure 4 is a flowchart of training the target spatio-temporal trajectory pre-training model in an embodiment of the present invention.
[0037] Figure 5 is a schematic structural diagram of a spatio-temporal trajectory pre-training model construction device provided by an embodiment of the present invention.
[0038] Figure 6 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0040] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0041] Referring to Figure 1 , an embodiment of the present invention provides a method for constructing a target spatio-temporal trajectory pre-training model, including the following steps:
[0042] Step 1, establish a spatio-temporal trajectory pre-training model based on a Transformer decoder;
[0043] In some embodiments, referring to the architecture of a large-scale generative language foundation model based on Transformer, a spatio-temporal trajectory pre-training model is established. Referring to Figure 2 , the spatio-temporal trajectory pre-training model includes a hybrid embedding layer, a multi-layer multi-head masked self-attention layer, and a fully connected feed-forward neural network layer connected in sequence.
[0044] In some embodiments, the hybrid embedding layer includes a channel embedding module, an embedding vector splicing module, and a linear transformation network module connected in sequence, and is used to convert the target spatio-temporal trajectory token sequence into an embedding tensor where S represents the number of trajectory points included in the target spatio-temporal trajectory, F represents the number of features of the trajectory points, D represents the feature embedding dimension of the trajectory points, and its size can be set according to actual needs. The target spatio-temporal trajectory token sequence X is defined as:
[0045]
[0046] where represents the token value corresponding to the trajectory point at the i-th moment, and this token value is composed of the feature encodings of multiple features. a ij represents the feature encoding corresponding to the j-th feature of the trajectory point at the i-th moment of the target spatio-temporal trajectory, and the value range of the feature encoding is an integer value starting from 0.
[0047] The specific steps for converting the target spatio-temporal trajectory token sequence into an embedding tensor are: first, calculate the embedding tensor of each feature separately, then splice the embedding tensors of all features, and finally convert the spliced tensor into an embedding tensor with a specified embedding dimension through a non-linear transformation network.
[0048] In some embodiments, the multi-layer multi-head masked self-attention layer includes N sequentially connected Transformer decoders, and each Transformer decoder includes a masked multi-head self-attention module, a first residual connection and layer normalization module, a position-wise feed-forward neural network module, and a second residual connection and layer normalization module connected in sequence. The number N of the Transformer decoders can be set according to actual needs. This multi-layer multi-head masked self-attention layer mainly performs multiple non-linear transformations on the embedding tensor E based on the self-attention mechanism to obtain the output tensor
[0049] In some embodiments, the fully connected feed-forward neural network layer is used to convert the output tensor Z into an output probability where V represents the sum of the maximum values of the feature encodings of all features describing the trajectory points.
[0050] Step 2, construct a spatio-temporal trajectory pre-training dataset, and the spatio-temporal trajectory pre-training dataset includes a target spatio-temporal trajectory token sequence;
[0051] The target spatio-temporal trajectory consists of a series of trajectory points, and the features of each trajectory point include numerical features and categorical features;
[0052] The numerical features include time, longitude, latitude, speed, altitude, heading, etc.;
[0053] The categorical features include country, type, etc.
[0054] Refer to Figure 3 , in some embodiments, Step 2 includes the following sub-steps:
[0055] Step 2.1, calculate the trajectory resampling time interval;
[0056] The setting of the trajectory resampling time interval should follow the following criterion: in most cases, the target can cross to different spatial grids within this time to avoid the occurrence of repeated patterns in the longitude and latitude feature encoding combinations of adjacent trajectory points. Therefore, the trajectory resampling time interval t r depends on the target average speed and the maximum spatial distance d represented by the longitude and latitude feature encoding lengths of adjacent trajectory points. The calculation formula is as follows:
[0057]
[0058] where the average speed is in m / s, and the spatial distance d is in m.
[0059] Step 2.2, resample the trajectory data according to the trajectory resampling time interval;
[0060] According to the set resampling time interval, for each target spatio-temporal trajectory, an interpolation algorithm, such as linear interpolation, Spline interpolation, etc., is used to resample the trajectory data of the target spatio-temporal trajectory, and the trajectory data includes target time, longitude, and latitude information.
[0061] Step 2.3, calculate the target speed and heading information according to the resampled trajectory data;
[0062] In some embodiments, the calculated target speed and heading information need to be filtered, such as median filtering, mean filtering, etc.
[0063] Step 2.4, perform feature tokenization on the features of the target spatio-temporal trajectory to obtain a target spatio-temporal trajectory token sequence; the features of the target spatio-temporal trajectory include the resampled trajectory data and the target speed and heading information;
[0064] In some embodiments, for each target spatio-temporal trajectory after the foregoing preprocessing, the feature encodings of features such as time, longitude, latitude, speed, and heading of each trajectory point are calculated respectively. Among them, the time feature can be split into 3 features of hour, minute, and second, or the feature encoding can be performed according to the total number of seconds; for numerical features, an equal-width discretization method is used to map them to a continuous integer starting from 0; for categorical features, a continuous integer starting from 0 is directly assigned to each value.
[0065] Based on the feature encodings corresponding to the features included in each trajectory point, they are combined into a set in a certain order as an element in the target spatio-temporal trajectory token sequence, so as to obtain the target spatio-temporal trajectory token sequence, and thus construct a spatio-temporal trajectory pre-training dataset.
[0066] Step 3, based on the spatio-temporal trajectory pre-training dataset, train the spatio-temporal trajectory pre-training model using the composite cross-entropy loss function.
[0067] Refer to Figure 4 , in some embodiments, Step 3 includes the following sub-steps:
[0068] Step 3.1, construct a composite cross-entropy loss function;
[0069] In some embodiments, the composite cross-entropy loss function is defined as: a function for calculating the weighted average of each feature cross-loss value, and its calculation steps are: first, the logical probability vector output by the spatio-temporal trajectory pre-training model is segmented according to the feature dimension, then the cross-entropy loss of each feature is calculated respectively, and finally the cross-entropy losses of multiple features are weighted and summed. The calculation formula is as follows:
[0070]
[0071] Among them, is the logical probability vector at a certain moment output by the spatio-temporal trajectory pre-training model. After being segmented according to the feature dimension, it can be obtained indicating the output logical probability vector corresponding to the i-th feature; indicating the j-th element of indicating the logical probability output corresponding to the position of the true label under this feature; w i indicating the cross-entropy loss weight of the i-th feature, and there is
[0072] Step 3.2, set the training parameters of the spatio-temporal trajectory pre-training model;
[0073] According to the actual situation, set the training parameters required for training the spatio-temporal trajectory pre-training model, including the number of training epochs and the batch size of training data.
[0074] Step 3.3, based on the spatio-temporal trajectory pre-training dataset, perform autoregressive training on the spatio-temporal trajectory pre-training model;
[0075] In some embodiments, based on the constructed spatio-temporal trajectory pre-training dataset and the set training parameters, using the composite cross-entropy loss function, the autoregressive training of the spatio-temporal trajectory pre-training model is performed by using the backpropagation algorithm to obtain a trained spatio-temporal trajectory pre-training model.
[0076] Based on the same technical concept, corresponding to the method for constructing the target spatio-temporal trajectory pre-training model described in the above embodiments, Figure 5 a structural block diagram of a spatio-temporal trajectory pre-training model construction device provided by an embodiment of the present invention is given. The spatio-temporal trajectory pre-training model construction device may be a software unit or a unit combining software and hardware. For the sake of convenience of description, only the parts related to this embodiment are shown.
[0077] Referring to Figure 5 , the spatio-temporal trajectory pre-training model construction device includes:
[0078] A model construction unit for establishing a spatio-temporal trajectory pre-training model based on a transformer decoder;
[0079] A dataset construction unit for constructing a spatio-temporal trajectory pre-training dataset, and the spatio-temporal trajectory pre-training dataset includes a target spatio-temporal trajectory token sequence;
[0080] A model training unit for training the spatio-temporal trajectory pre-training model based on the spatio-temporal trajectory pre-training dataset by using the composite cross-entropy loss function.
[0081] The specific working principles of the functional units in the above device can be implemented with reference to the methods in the foregoing embodiments, and will not be elaborated herein.
[0082] Based on the same inventive concept, an embodiment of the present invention further provides an electronic device, which can implement the method flow of collaborative resource configuration optimization based on network nodes and software provided in the foregoing embodiments of the present invention. In one embodiment, the electronic device may be a server, or a terminal device or other electronic devices. As Figure 6 shown, the electronic device may include:
[0083] At least one processor, and a memory connected to at least one processor. In the embodiments of the present invention, the specific connection medium between the processor and the memory is not limited. Figure 6 In [the figure], it is taken as an example that the processor and the memory are connected through a bus. The bus is Figure 6 shown as a thick line in [the figure]. The connection manners between other components are only for illustrative purposes and are not to be construed as limiting. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 6 it is only shown as a thick line in [the figure], but it does not mean that there is only one bus or one type of bus. Alternatively, the processor may also be referred to as a controller, and the name is not limited.
[0084] In the embodiments of the present invention, the memory stores instructions executable by at least one processor. By executing the instructions stored in the memory, at least one processor can execute a method of collaborative resource configuration optimization based on network nodes and software described above. The processor can implement Figure 6 the functions of each module in the device shown in [the figure].
[0085] Among them, the processor is the control center of the device, and can connect various parts of the entire control device through various interfaces and lines. By running or executing the instructions stored in the memory and calling the data stored in the memory, various functions of the device and process data, so as to monitor the device as a whole.
[0086] In an alternative design, the processor may include one or more processing units. The processor may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor. In some embodiments, the processor and the memory may be implemented on the same chip, and in some embodiments, they may also be separately implemented on independent chips.
[0087] The processor can be a general-purpose processor, such as a CPU, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of a method for constructing a target spatio-temporal trajectory pre-training model disclosed in combination with the embodiments of the present invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0088] As a non-volatile computer-readable storage medium, the memory can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory can include at least one type of storage medium, for example, it can include flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disk, and so on. The memory is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory in the embodiments of the present invention can also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.
[0089] By designing and programming the processor, the code corresponding to the method for constructing a target spatio-temporal trajectory pre-training model introduced in the foregoing embodiments can be solidified into the chip, so that the chip can execute the steps of the method in the above embodiments when running. How to design and program the processor is a well-known technology to those skilled in the art and will not be elaborated here.
[0090] Based on the same inventive concept, the embodiments of the present invention also provide a storage medium storing computer instructions, which when run on a computer, cause the computer to execute a method for constructing a target spatio-temporal trajectory pre-training model discussed above.
[0091] In some alternative embodiments, various aspects of a method for constructing a target spatio-temporal trajectory pre-training model according to the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a device, the program code is used to cause the control device to execute the steps in a method for constructing a target spatio-temporal trajectory pre-training model according to various exemplary embodiments of the present invention described above in this specification.
[0092] It should be noted that although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present invention, the features and functions of two or more of the above-described units can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units. In addition, although the operations of the method of the present invention are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.
[0093] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0094] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a server, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0095] Program code for performing the operations of the present invention may be written using any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0096] In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network including a local area network (LAN) or a wide area network (WAN), or alternatively, may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).
[0097] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one process Figure 1 one process or multiple processes and / or boxes Figure 1 specified in one box or multiple boxes.
[0098] These computer program instructions may also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or boxes Figure 1 specified in one box or multiple boxes.
[0099] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for constructing a target spatiotemporal trajectory pre-training model, characterized in that: include: Establishing a spatiotemporal trajectory pre-training model based on a transformer decoder; the spatiotemporal trajectory pre-training model includes a sequentially connected hybrid embedding layer, a multi-layer multi-head masked self-attention layer, and a fully connected feedforward neural network layer; Constructing a spatiotemporal trajectory pre-training dataset, wherein the spatiotemporal trajectory pre-training dataset includes a target spatiotemporal trajectory token sequence; Based on the spatiotemporal trajectory pre-training dataset, the spatiotemporal trajectory pre-training model is trained using a composite cross entropy loss function.
2. The method for constructing a target spatiotemporal trajectory pre-training model according to claim 1, characterized in that: The hybrid embedding layer includes a channel embedding module, an embedding vector concatenation module and a linear transformation network module which are connected in sequence.
3. The target spatiotemporal trajectory pre-training model construction method according to claim 2 is characterized in that: The hybrid embedding layer is used to convert the input target spatiotemporal trajectory token sequence into an embedding tensor, specifically including: firstly calculating the embedding tensor of each feature separately, then concatenating the embedding tensors of all features, and finally converting the concatenated tensor into an embedding tensor of the specified embedding dimension through a nonlinear transformation network.
4. The method for constructing a target spatiotemporal trajectory pre-training model according to claim 1, characterized in that: The multi-layer multi-head masked self-attention layer includes N transformer decoders connected in sequence; each transformer decoder includes a masked multi-head self-attention module, a first residual connection and layer normalization module, a position-by-position feedforward neural network module and a second residual connection and layer normalization module connected in sequence.
5. The method for constructing a target spatiotemporal trajectory pre-training model according to claim 1, characterized in that: The step of constructing a spatiotemporal trajectory pre-training dataset includes: Calculate trajectory resampling time interval; Resampling the trajectory data according to the trajectory resampling time interval; the trajectory data includes target time, longitude and latitude information; Calculate the target speed and heading information based on the resampled trajectory data; The features of the target spatiotemporal trajectory are tokenized to obtain a target spatiotemporal trajectory token sequence; the features of the target spatiotemporal trajectory include resampled trajectory data and target speed and heading information.
6. The method for constructing a target spatiotemporal trajectory pre-training model according to claim 5, characterized in that: The calculating trajectory resampling time interval comprises: The trajectory resampling time interval is calculated based on the target average speed and the maximum spatial distance represented by the length of the longitude and latitude feature codes of adjacent trajectory points.
7. The method for constructing a target spatiotemporal trajectory pre-training model according to claim 1, characterized in that: The composite cross entropy loss function is defined as: a function that calculates the weighted average of the cross loss values of each feature, and its calculation steps are: first, the logical probability vector output by the spatiotemporal trajectory pre-training model is divided according to the feature dimension, and then the cross entropy loss of each feature is calculated separately, and finally the cross entropy losses of multiple features are weighted summed.
8. A target spatiotemporal trajectory pre-training model construction device, characterized in that: include: A model building unit, used to establish a spatiotemporal trajectory pre-training model based on a transformer decoder; the spatiotemporal trajectory pre-training model includes a sequentially connected hybrid embedding layer, a multi-layer multi-head masked self-attention layer, and a fully connected feedforward neural network layer; A data set construction unit, used to construct a spatiotemporal trajectory pre-training data set, wherein the spatiotemporal trajectory pre-training data set includes a target spatiotemporal trajectory token sequence; The model training unit is used to train the spatiotemporal trajectory pre-training model based on the spatiotemporal trajectory pre-training data set using a composite cross entropy loss function.
9. An electronic device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the at least one processor executes the method as described in any one of claims 1 to 7 by executing the instructions stored in the memory.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is implemented.