Task prediction method, device and equipment based on multi-scale space-time pre-training large model, medium and product

By converting space-time trajectory points into space-time trajectory text and building a space-time pre-training big model, the problem that traditional trajectory modeling methods fail to fully learn user trajectory characteristics is solved, and more accurate user representation and downstream task prediction are achieved.

CN120011766APending Publication Date: 2025-05-16CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510104128.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The traditional trajectory modeling method does not fully learn user trajectory characteristics, making it difficult to apply user representation to user portrait and travel mode tasks.

Method used

The task prediction method based on multi-scale space-time pre-training large model is adopted. By converting space-time trajectory points into space-time trajectory text, and using pre-trained language models to build a space-time pre-training large model, obtaining the representation of the target user, and inputting it into the downstream task prediction model to output prediction information.

Benefits of technology

By unifying the space-time business logic, avoid the loss of user trajectory details, learn the preferences and rules of user temporal travel activities, improve the accuracy of user representation, and ensure accurate prediction of downstream tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011766A_ABST
    Figure CN120011766A_ABST
Patent Text Reader

Abstract

The invention discloses a task prediction method and device based on a multi-scale space-time pre-training large model, equipment, a medium and a product, and relates to the technical field of artificial intelligence, and the method comprises the steps: converting a space-time trajectory point into a space-time trajectory text based on a multi-scale space grid division result; based on the spatio-temporal trajectory text and a pre-training language model, a spatio-temporal pre-training large model is constructed, the pre-training language model comprises a plurality of coding units, and each coding unit comprises a multi-head attention module, a residual module, a standardization module and a feedforward neural network module; inputting a target spatio-temporal trajectory text of a target user into the large spatio-temporal pre-training model to obtain a target user representation; and inputting the target user representation into a downstream task prediction model, and outputting downstream task prediction information. The technical problem that it is difficult to apply user characterization to user portrait type and travel mode type tasks can be solved, and therefore accurate prediction of downstream tasks can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a task prediction method, device, equipment, medium and product based on a multi-scale spatiotemporal pre-trained large model. Background Art

[0002] Traditional trajectory modeling methods mostly extract key activity points or intersection points, highly simplify user trajectories, and then model them. This method discards a large amount of user trajectory details and is unable to capture the user's rich and diverse trajectory activity information, such as stopping (shopping in shopping malls, leisure at home), traveling (subway, bus, cycling, walking), etc., resulting in insufficient learning of user trajectory characteristics, which makes it difficult to apply user representation to user profiling and travel mode tasks.

[0003] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention

[0004] The main purpose of this application is to provide a task prediction method, device, equipment, medium and product based on a multi-scale spatiotemporal pre-trained large model, aiming to solve the technical problem that the traditional trajectory modeling method does not fully learn the user trajectory characteristics, making it difficult to apply user representation to user profiling and travel mode tasks.

[0005] To achieve the above objectives, the present application proposes a task prediction method based on a multi-scale spatiotemporal pre-trained large model, and the task prediction method based on the multi-scale spatiotemporal pre-trained large model includes:

[0006] Based on the multi-scale spatial grid division results, the space-time trajectory points are converted into space-time trajectory texts;

[0007] Based on the spatiotemporal trajectory text and the pre-trained language model, a spatiotemporal pre-trained large model is constructed, wherein the pre-trained language model includes a plurality of encoding units, and the encoding units include a multi-head attention module, a residual module, a normalization module, and a feedforward neural network module;

[0008] Inputting the target spatiotemporal trajectory text of the target user into the spatiotemporal pre-trained large model to obtain a target user representation;

[0009] The target user representation is input into a downstream task prediction model, and downstream task prediction information is output.

[0010] In one embodiment, the step of converting the space-time trajectory points into space-time trajectory text based on the multi-scale spatial grid division result includes:

[0011] Based on the administrative division and grid division methods, determine the grids at all levels across the country;

[0012] Preset characters are used to represent the grids at all levels across the country, and a multi-scale spatial grid division result is obtained;

[0013] Based on the multi-scale spatial grid division result, the space-time trajectory points are converted into original space-time trajectory text;

[0014] Segment the national spatiotemporal trajectory text to obtain segmentation results, and determine high-frequency byte pairs according to the segmentation results;

[0015] Merging the high-frequency byte pairs into new subwords, and constructing a trajectory word list according to the new subwords;

[0016] Based on the trajectory vocabulary, the original spatiotemporal trajectory text is converted into a spatiotemporal trajectory text.

[0017] In one embodiment, the step of constructing a spatiotemporal pre-trained large model based on the spatiotemporal trajectory text and the pre-trained language model includes:

[0018] After embedding the spatiotemporal trajectory text, an input vector is obtained, wherein the embedding process includes text embedding, position embedding and segment embedding;

[0019] Determine a first cross entropy and a second cross entropy of a pre-training task based on the input vector and the pre-trained language model;

[0020] Determining a model iteration error based on the first cross entropy and the second cross entropy;

[0021] After updating the model parameters of the pre-trained language model according to the model iteration error, a spatiotemporal pre-trained large model is obtained.

[0022] In one embodiment, the step of determining a first cross entropy and a second cross entropy of a pre-training task based on the input vector and the pre-trained language model comprises:

[0023] The input vector is randomly processed using a mask flag to obtain a masked input vector;

[0024] Obtaining a first cross entropy of a pre-training task based on the masked input vector, the input vector, and the pre-trained language model;

[0025] Randomly obtain a first vector and a second vector from the input vector, and determine user information between the first vector and the second vector, wherein the user information includes one of belonging to the same user and not belonging to the same user;

[0026] The first vector and the second vector are input into the pre-trained language model to obtain inference information, and a second cross entropy of the pre-training task is determined according to the user information and the inference information.

[0027] In one embodiment, the spatiotemporal pre-trained large model includes a national scale large model, an urban agglomeration scale large model, and a city scale large model; wherein the step of inputting the target spatiotemporal trajectory text of the target user into the spatiotemporal pre-trained large model to obtain the target user representation includes:

[0028] Inputting the target spatiotemporal trajectory text into the national scale large model, the urban agglomeration scale large model and the national scale large model respectively, to obtain a national scale representation, an urban agglomeration scale representation and a city scale representation;

[0029] The national scale representation, the urban agglomeration scale representation and the city scale representation are taken as optional representations;

[0030] According to the task type of the target user, a target optional representation is determined from the optional representations, and the target optional representations are spliced ​​into a target user representation.

[0031] In one embodiment, before the step of inputting the target user representation into the downstream task prediction model and outputting the downstream task prediction information, the step further includes:

[0032] Acquire a training sample, wherein the training sample includes a user representation and a true value label;

[0033] Inputting the user representation into the initial task prediction model and outputting the prediction probability;

[0034] Determining a cross entropy loss value based on the true value label and the predicted probability;

[0035] Based on the cross entropy loss value, after updating the model parameters of the initial task prediction model, a downstream task prediction model is obtained.

[0036] In addition, to achieve the above purpose, the present application also proposes a task prediction device based on a multi-scale spatiotemporal pre-trained large model, and the task prediction device based on the multi-scale spatiotemporal pre-trained large model includes:

[0037] A conversion module is used to convert the space-time trajectory points into space-time trajectory text based on the multi-scale spatial grid division results;

[0038] A construction module, used to construct a spatiotemporal pre-trained large model based on the spatiotemporal trajectory text and the pre-trained language model, wherein the pre-trained language model includes a plurality of encoding units, and the encoding units include a multi-head attention module, a residual module, a normalization module, and a feedforward neural network module;

[0039] An input module, used to input the target spatiotemporal trajectory text of the target user into the spatiotemporal pre-trained large model to obtain a target user representation;

[0040] The output module is used to input the target user representation into the downstream task prediction model and output the downstream task prediction information.

[0041] In addition, to achieve the above-mentioned purpose, the present application also proposes a task prediction device based on a multi-scale spatiotemporal pre-trained large model, the device comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the task prediction method based on the multi-scale spatiotemporal pre-trained large model as described above.

[0042] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the task prediction method based on the multi-scale spatiotemporal pre-trained large model as described above are implemented.

[0043] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the task prediction method based on the multi-scale spatiotemporal pre-trained large model as described above.

[0044] One or more technical solutions proposed in this application have at least the following technical effects:

[0045] The task prediction method, device, equipment, medium and product based on the multi-scale spatiotemporal pre-trained large model proposed in the present application convert the spatiotemporal trajectory points into spatiotemporal trajectory text based on the multi-scale spatial grid division result; construct the spatiotemporal pre-trained large model based on the spatiotemporal trajectory text and the pre-trained language model, wherein the pre-trained language model includes multiple encoding units, and the encoding unit includes a multi-head attention module, a residual module, a normalization module and a feedforward neural network module; input the target spatiotemporal trajectory text of the target user into the spatiotemporal pre-trained large model to obtain the target user representation; input the target user representation into the downstream task prediction model, and output the downstream task prediction information, which solves the technical problem that the traditional trajectory modeling method does not fully learn the user trajectory characteristics, making it difficult to apply the user representation to user portrait and travel mode tasks. Compared with the prior art, the present application avoids the loss of user trajectory detail information by unifying the spatiotemporal business logic, and then learns the preferences and rules of the user's spatiotemporal travel activities through the spatiotemporal pre-trained large model, which can improve the accuracy of the user representation, thereby ensuring the accurate prediction of downstream tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0047] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0048] Figure 1 A flowchart diagram of the first embodiment of the task prediction method based on a multi-scale spatiotemporal pre-trained large model provided in this application;

[0049] Figure 2 A schematic diagram of the national multi-level grid division provided in Example 1 of the task prediction method based on the multi-scale spatiotemporal pre-trained large model of this application;

[0050] Figure 3 A schematic diagram of the representation splicing and task prediction process provided in Example 1 of the task prediction method based on a multi-scale spatiotemporal pre-trained large model of this application;

[0051] Figure 4 A flowchart diagram of the second embodiment of the task prediction method based on a multi-scale spatiotemporal pre-trained large model provided in this application;

[0052] Figure 5 A schematic diagram of the first pre-training process provided in Example 2 of the task prediction method based on a multi-scale spatiotemporal pre-training large model of the present application;

[0053] Figure 6 A schematic diagram of a second pre-training process provided in Example 2 of the task prediction method based on a multi-scale spatiotemporal pre-training large model of the present application;

[0054] Figure 7 This is a schematic diagram of the module structure of a task prediction device based on a multi-scale spatiotemporal pre-trained large model according to an embodiment of the present application;

[0055] Figure 8 Schematic diagram of the device structure of the hardware operating environment involved in the task prediction method based on the multi-scale spatiotemporal pre-trained large model in the embodiment of the present application.

[0056] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0057] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0058] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0059] The main solution of the embodiment of the present application is: based on the multi-scale spatial grid division result, converting the space-time trajectory points into space-time trajectory text; constructing a space-time pre-trained large model based on the space-time trajectory text and the pre-trained language model, wherein the pre-trained language model includes multiple encoding units, and the encoding unit includes a multi-head attention module, a residual module, a normalization module and a feedforward neural network module; inputting the target space-time trajectory text of the target user into the space-time pre-trained large model to obtain the target user representation; inputting the target user representation into the downstream task prediction model, and outputting the downstream task prediction information.

[0060] As can be seen from the above embodiments, the present application converts the spatiotemporal trajectory points into spatiotemporal trajectory texts based on the multi-scale spatial grid division results; constructs a spatiotemporal pre-trained large model based on the spatiotemporal trajectory text and the pre-trained language model, wherein the pre-trained language model includes multiple encoding units, and the encoding units include a multi-head attention module, a residual module, a normalization module, and a feedforward neural network module; inputs the target spatiotemporal trajectory text of the target user into the spatiotemporal pre-trained large model to obtain the target user representation; inputs the target user representation into the downstream task prediction model, and outputs the downstream task prediction information. The technical problem that the traditional trajectory modeling method does not fully learn the user trajectory characteristics and makes it difficult to apply the user representation to user portrait and travel mode tasks is solved. Compared with the prior art, the present application avoids the loss of user trajectory details by unifying the spatiotemporal business logic, and then learns the preferences and rules of the user's spatiotemporal travel activities through the spatiotemporal pre-trained large model, which can improve the accuracy of the user representation, thereby ensuring the accurate prediction of downstream tasks.

[0061] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of realizing the above functions, a task prediction device based on a multi-scale spatiotemporal pre-trained large model, etc. The following takes the task prediction based on the multi-scale spatiotemporal pre-trained large model as an example to illustrate this embodiment and the following embodiments.

[0062] Based on this, the embodiment of the present application provides a task prediction method based on a multi-scale spatiotemporal pre-trained large model, referring to Figure 1 , Figure 1This is a flow chart of the first embodiment of the task prediction method based on the multi-scale spatiotemporal pre-trained large model of the present application.

[0063] In this embodiment, the task prediction method based on the multi-scale spatiotemporal pre-trained large model includes steps S10 to S40:

[0064] Step S10, based on the multi-scale spatial grid division result, converting the spatiotemporal trajectory points into spatiotemporal trajectory text;

[0065] It should be noted that if Figure 2 As shown in the figure, the administrative division method and the Google S2 grid division method can be combined to realize the multi-level fine grid division of the whole country, and then the preset characters are used to represent the grids of each level in the whole country, so as to obtain the multi-scale spatial grid division result. Specifically, for a certain trajectory, the spatial units of each level where a certain trajectory point is located are calculated. Taking the coordinate point of Beijing with longitude and latitude (40.12619460, 116.40709218) as an example, it can be expressed as a five-level string similar to Б / A / あ / コ / a, where Б, A, あ, コ, and a represent the characters corresponding to the province, city, first-level grid, second-level grid, and third-level grid of the point respectively. Then, the trajectory with the longitude and latitude point string (i.e., the space-time trajectory point string) of (40.12619460, 116.40709218; 40.24778745, 116.50626848; 39.96384436, 116.10956329) can be converted into a text similar to the following: Б / A / あ / コ / a, Б / A / あ / コ / b, Б / A / あ / ツ / c. In this way, the longitude and latitude trajectories of different users can be converted into space-time trajectory texts with consistent lengths at each point.

[0066] In one embodiment, the step of converting the space-time trajectory points into space-time trajectory text based on the multi-scale spatial grid division result includes: determining the grids at all levels across the country based on the administrative division method and the grid division method; using preset characters to represent the grids at all levels across the country to obtain the multi-scale spatial grid division result; based on the multi-scale spatial grid division result, converting the space-time trajectory points into original space-time trajectory text; performing word segmentation on the national space-time trajectory text to obtain word segmentation results, and determining high-frequency byte pairs based on the word segmentation results; merging the high-frequency byte pairs into new sub-words, and constructing a trajectory word list based on the new sub-words; based on the trajectory word list, converting the original space-time trajectory text into space-time trajectory text.

[0067] In the specific implementation, the country is first divided into different administrative units according to multiple levels such as provinces and prefecture-level cities. Then, for each prefecture-level city, the Google S2 grid division method is used to divide the administrative area into n grid levels (the Google S2 algorithm is a map projection and spatial index algorithm invented by Google based on the unit sphere. The algorithm divides the world into 30 levels of grids, with grid sizes ranging from 7800*7800 kilometers to 9*9 millimeters). Taking Beijing as an example, Beijing can be divided into three levels of grids, and the size of each level of grid can be set to 40*40 kilometers, 5*5 kilometers, and 500*500 meters respectively. Then different characters can be used to represent the grids at various levels across the country. For example, provincial units are represented by Cyrillic letters Б, Г, Д, etc., each city in each provincial unit is represented by uppercase English letters A, B, C, etc., each first-level grid in each city unit is represented by Hiragana letters such as あ, お, き, etc., each second-level grid in each first-level grid is represented by Katakana letters such as コ, ツ, ク, etc., and each third-level grid in each second-level grid is represented by lowercase English letters a, b, c or numbers 1, 2, 3, etc. If there is a situation where there are not enough characters, two measures can be taken to solve the problem. The first is to increase the level of spatial division units, for example, from five-level spatial division to six or seven-level spatial unit division, so that the spatial units corresponding to each level will be rapidly reduced; the second is to use more character libraries to express spatial units at the same level, such as using uppercase English letters A, B, C, lowercase English letters a, b, c, numbers 1, 2, 3, etc. to express spatial levels with a large number of grids within the city.

[0068] It should be noted that after obtaining the multi-scale spatial grid division results, all the longitude and latitude points (i.e., the spatiotemporal trajectory points) corresponding to the user's spatiotemporal trajectory are converted into a string of the same length (i.e., the original spatiotemporal trajectory text) to realize the conversion of the trajectory point string to the trajectory text, wherein the spatiotemporal trajectory is obtained from the user's spatiotemporal location data recorded by a variety of communication or location services, including Internet map location services such as navigation and positioning, communication, and call operator communication services. The spatiotemporal trajectory points are composed of multiple longitude and latitude points. The attributes of each longitude and latitude point usually include the longitude, latitude, and recording time of the user's location. The spatiotemporal trajectory of the user is obtained by sorting the longitude and latitude points of the same user in time. In the specific implementation, the national spatiotemporal trajectory data of one week or longer is used to screen the trajectory with low time missing rate. The national spatiotemporal trajectory data can be obtained from channels such as anonymous spatiotemporal location information recorded by the background of Internet software and anonymous signaling data recorded by the background of communication operators. Among them, the spatiotemporal location information recorded by Internet software is only generated when the user uses the Internet software and initiates positioning or navigation related services; signaling data is the spatiotemporal location information generated when the user interacts with the communication base station when using the mobile phone to make calls, send text messages or search for Internet information. These spatiotemporal location information are associated through a unique de-privacy user ID, and are sorted according to the time when the records are generated to form the user's spatiotemporal trajectory. Among them, the method for screening trajectories with low time loss rate is: divide a week or longer into N equally spaced time intervals (such as every 15 minutes), calculate the number of intervals where the spatiotemporal trajectory exists, set it as M, then the spatiotemporal trajectory loss rate If the ratio exceeds a threshold (e.g., 70%), the trajectory is considered to have a low temporal missing rate, and these trajectories are sampled at equal intervals (e.g., every 15 minutes). The trajectory point with the smallest starting time difference is selected for each time interval. For example, starting from 0 o'clock, 0 o'clock-0 o'clock 15 minutes is the first time interval, and the trajectory point closest to 0 o'clock in time is selected as the sampling point of this interval. Then, the sampled trajectory points are gridded and converted into trajectory text, so that the text character length of different user trajectories is consistent.

[0069] Then, a double-byte encoding algorithm is used to perform word segmentation calculation and word list construction on the national spatiotemporal trajectory text of the national spatiotemporal trajectory data. First, the subword word list size is set (such as 300,000), and then the text of each trajectory point in the national spatiotemporal trajectory text (such as Б / A / あ / コ / a, Б / A / あ / コ / b, Б / A / あ / ツ / c) is split into character sequences to count the frequency of occurrence of each continuous byte pair, and the highest frequency ones are selected to merge into new subwords (that is, the national spatiotemporal trajectory text is segmented to obtain the segmentation result, and the high-frequency byte pairs are determined according to the segmentation result, and the high-frequency byte pairs are merged into new subwords, and the trajectory word list is constructed according to the mapping relationship between the high-frequency byte pairs and the new subwords). For example, here Б / A / あ has the highest frequency, so it is merged into a subword. Finally, the process is repeated until the subword word list size is set or the frequency of the next highest frequency byte pair is 1, and the subwords are sorted from high to low according to the length of the subwords. For example, the subword vocabulary after splitting here is Б / A / あ / コ, Б / A / あ, ツ / c, a, b.

[0070] In the specific implementation, the method of using the trajectory vocabulary is that in the model pre-training and prediction stage, for a certain trajectory point string, the trajectory point string is first converted into the original trajectory text according to the spatial grid conversion method (i.e., the result of multi-scale spatial grid division), and then the trajectory vocabulary is loaded. By searching the trajectory vocabulary, the original trajectory text is segmented and converted into the final trajectory text (i.e., the original spatiotemporal trajectory text). For example, for the longitude and latitude point (40.12619460, 116.40709218), it is converted into the original spatiotemporal trajectory text (i.e., "Б / A / あ / コ / a") after spatial mapping (i.e., the result of multi-scale spatial grid division), and the subword corresponding to "Б / A / あ / コ" in the original trajectory text is determined by searching the trajectory vocabulary, and then the original spatiotemporal trajectory text is converted into the final spatiotemporal trajectory text (i.e., the subword corresponding to "Б / A / あ / コ", a).

[0071] In this embodiment, by dividing the whole country into multi-scale spatial grids, a unified national spatial location mapping and indexing is established according to the multi-level spatial grid division method of provinces, cities, districts, counties, and grids, and unique characters are assigned to spatial units at different levels. Each longitude and latitude point is converted into a string of consistent length, and the longitude and latitude trajectories of different users are converted into texts of consistent length at each point, which provides a basis for subsequent modeling to be compatible with different spatial scales.

[0072] Step S20, constructing a spatiotemporal pre-trained large model based on the spatiotemporal trajectory text and the pre-trained language model, wherein the pre-trained language model includes a plurality of encoding units, and the encoding units include a multi-head attention module, a residual module, a normalization module, and a feedforward neural network module;

[0073] It should be noted that the core architecture of the Bert model (Bidirectional Encoder Representations from Transformers, pre-trained language model) adopts the Decoder-Only Transformers architecture. The encoding unit of a transformer is generated by the superposition of multi-head attention layers, layer normalization, feedforword neural network, and layer normalization. Each layer of the Bert model consists of such an encoding unit. In the larger Bert model, there are 24 encoding layers, 16 attention units in each layer, and the dimension of the word vector is 1024. In the smaller Bert model, there are 12 encoding layers, 12 attention units in each layer, and the dimension of the word vector is 768. This Decoder-Only Transformers structure can use context to predict the masked input vector, thereby capturing the bidirectional relationship of the trajectory text. Here, the bidirectional relationship means that the Bert model will simultaneously consider the contextual relationship of the trajectory text.

[0074] In the specific implementation, the spatiotemporal pre-trained large model (i.e., the multi-scale national large model) includes a national-scale large model, an urban agglomeration-scale large model, and a city-scale large model. The spatiotemporal pre-trained model (i.e., the multi-scale national large model) can be constructed by using spatiotemporal trajectory text and the Bert model (i.e., the pre-trained language model).

[0075] First, select the scale of model construction. The spatial resolution of the national-scale large model can be set to the district and county level, and the text characters of a single track point are relatively few, for example, it can be set to 3 characters (province, prefecture-level city, district and county); the spatial resolution of the urban agglomeration-scale large model grid can be set to 500 meters * 500 meters grid, and the text characters of a single track point are relatively large, for example, it can be set to 5 characters (province, prefecture-level city, 40 * 40 kilometers, 5 * 5 kilometers, 500 * 500 meters); the spatial resolution of the city-scale large model grid can be set to 100 meters * 100 meters grid, and the text characters of a single track point are relatively large, for example, it can be set to 6 characters ( Province, prefecture-level city, 40*40 km, 5*5 km, 500*500 m, 100 m*100 m); then, the Bert model is used to construct a large spatiotemporal pre-training model. The input of the Bert model starts with the special symbol CLS. The vector representation of the special character CLS obtained by BERT is usually used as the current sentence representation. Then, the longitude and latitude points of a user's sampled trajectory are input, and the original trajectory text is quickly segmented by searching the vocabulary and converted into the final trajectory text. The final trajectory text is then embedded in text, embedded in position, and embedded in segments to obtain the input vector of the Bert model.

[0076] Step S30, inputting the target spatiotemporal trajectory text of the target user into the spatiotemporal pre-trained large model to obtain the target user representation;

[0077] It should be noted that the target space-time trajectory text is obtained by obtaining the target space-time trajectory data of the target user, and then converting the target space-time trajectory data into the original target space-time trajectory text based on the multi-scale spatial grid division result, and then converting the original target space-time trajectory text into the target space-time trajectory text based on the trajectory vocabulary.

[0078] It should be noted that the target user representation is the output vector corresponding to the CLS special mark in the pre-trained language model, which condenses the characteristics of the user's entire trajectory. CLS is a special mark in the BERT model, which is used to instruct the model how to process the content of the entire trajectory text. In the BERT model structure, after multiple attention mechanism calculations, the output corresponding to the CLS mark is a vector generated by the weighted average of all words, which can better represent the semantics of the entire sentence than the representation of a single word, thereby providing a global trajectory semantic representation. Figure 3 For example, the user in Figure 1 works in city a and often travels to city b for business and tourism. First, three models are trained using the Bert model at the grid scale of city a, the grid scale of city b, and the national county scale. User representations at different scales are obtained through the output vectors corresponding to the CLS special flags.

[0079] In a feasible implementation manner, the spatiotemporal pre-trained large model includes a national-scale large model, an urban agglomeration-scale large model and a city-scale large model; wherein, the step of inputting the target spatiotemporal trajectory text of the target user into the spatiotemporal pre-trained large model to obtain the target user representation includes: inputting the target spatiotemporal trajectory text into the national-scale large model, the urban agglomeration-scale large model and the national-scale large model respectively to obtain a national-scale representation, an urban agglomeration-scale representation and a city-scale representation; using the national-scale representation, the urban agglomeration-scale representation and the city-scale representation as optional representations; determining a target optional representation from the optional representations according to the task type of the target user, and splicing the target optional representations into a target user representation.

[0080] It should be noted that the spatiotemporal pre-trained large model (i.e., multi-scale national large model) includes a national-scale large model, an urban agglomeration-scale large model, and a city-scale large model. Among them, the national-scale large model mainly learns the rules of users' long-distance cross-city travel behavior. Typical usage scenarios include large-scale population mobility monitoring during major holidays such as the Spring Festival and National Day, cross-city tourism and business trip behavior analysis, and winter and summer vacation student return home and school monitoring. The input data of the national-scale large model is mainly the spatiotemporal trajectory of the user's cross-city travel part. The starting point is usually the high-speed rail station, airport, bus station, highway toll station and other transportation hubs in the user's permanent city, and the end point is usually the high-speed rail station, airport, bus station, highway toll station and other transportation hubs at the user's destination. The method for identifying the user's permanent city is the city where the user's spatiotemporal trajectory exceeds the time threshold (such as 3 months or 6 months) within the administrative area of ​​a certain city. The scale of the smallest analysis unit is very coarse, for example, the district and county unit can be used. The purpose of setting the minimum analysis unit here is mainly to limit the size of the vocabulary, so as to avoid the situation where the spatial resolution of the minimum analysis unit is too fine at the national scale, resulting in a vocabulary that is too large and the model is difficult to train and converge. When the minimum analysis unit is set to the county scale, the trajectory only needs to be accurate to the county unit when performing text conversion and vocabulary construction, and does not need to be accurate to the grid scale, so as to achieve the purpose of controlling the size of the vocabulary; an urban agglomeration refers to a large, multi-core, multi-level urban group formed by the aggregation of several megacities and large cities that are closely connected in terms of economy, society, culture, etc. It usually takes one or several central cities as the core and radiates to the surrounding areas to form an urban aggregate centered on these cities. Typical urban agglomerations include the Yangtze River Delta Urban Agglomeration, the Pearl River Delta Urban Agglomeration, the Beijing-Tianjin-Hebei Urban Agglomeration, etc. The large model at the urban agglomeration scale mainly learns the daily activity patterns of users within the urban agglomeration. Typical usage scenarios include cross-city daily commuting analysis, urban agglomeration population and economic connection intensity analysis, etc. The input data of the large-scale model at the city agglomeration scale is the spatiotemporal trajectory data of all users within the city agglomeration, and the scale of the smallest grid is relatively coarse, for example, a 500m*500m grid can be used; the large-scale model at the city scale mainly studies the daily activity patterns of users within the city, and typical usage scenarios include analysis of commuting patterns within the city, quantification of traffic activities, and monitoring of commercial tourism emergency behaviors. The input data of the large-scale model at the city scale is the spatiotemporal trajectory data of all users within the city, and the scale of the smallest grid is very fine, for example, a 100m*100m grid can be used.

[0081] It should be noted that for different types of prediction tasks (i.e., task types), user representations of different scales are selected or spliced ​​separately to obtain the input of downstream tasks (i.e., target user representation). For example, to predict the characteristics of users' daily commuting, work-residence life, etc., the city-scale representation can be selected as the target user representation; to predict the characteristics of users' travel, homecoming, business trips, etc., the city-scale representation, the city cluster-scale representation, and the national-scale representation can be spliced ​​to obtain the target user representation; to predict the characteristics of users' cross-city travel by high-speed rail, airplane, highway, etc., the national-scale representation can be selected as the target user representation.

[0082] In this embodiment, user representations are predicted through models of different scales, and the user's common representations at multiple scales such as city, city cluster, and country are spliced ​​according to the task type. Combined with the fully connected layer model, prediction of multiple user attributes is achieved, which can cope with downstream tasks of different spatial scales, greatly reduce the workload of model building, and unify the spatiotemporal business logic.

[0083] Step S40: input the target user representation into a downstream task prediction model, and output downstream task prediction information.

[0084] It should be noted that the downstream task prediction model uses a fully connected neural network to predict user attributes (i.e., downstream task prediction information), where the number of neuron nodes in the input layer of the fully connected neural network is consistent with the dimension of the input target user representation, the number of neuron nodes in the output layer of the fully connected neural network is consistent with the dimension of the downstream task prediction information, and the number of hidden layers and the number of neuron nodes in each layer of the fully connected neural network can be adjusted according to the complexity of the task. Specifically, Figure 3 As shown in the figure, the target user representation dimension is 9, the number of neuron nodes in the input layer is set to 9, the number of neuron nodes in each hidden layer is set to 11, 13, 13, 11, 9, 5, and the number of neuron nodes in the output layer is set to 1.

[0085] It can be understood that the downstream task prediction model refers to a trained model, and after inputting the target user representation, the downstream task prediction value (ie, downstream task prediction information) will be obtained.

[0086] In a feasible implementation manner, before the step of inputting the target user representation into the downstream task prediction model and outputting the downstream task prediction information, the step also includes: before the step of inputting the target user representation into the downstream task prediction model and outputting the downstream task prediction information, the step also includes: obtaining training samples, wherein the training samples include user representations and true value labels; inputting the user representation into the initial task prediction model and outputting the predicted probability; determining the cross entropy loss value based on the true value label and the predicted probability; and obtaining the downstream task prediction model after updating the model parameters of the initial task prediction model based on the cross entropy loss value.

[0087] It should be noted that in the training stage of the downstream task prediction model, the input is the concatenated representation of the user, and the output is the true value label of the specific task. Taking the occupational profiling task as an example, when predicting whether the user is an online car-hailing driver, the true value label is 0 or 1, 0 represents that the user is not an online car-hailing driver, and 1 represents that the user is an online car-hailing driver. The model loss function is the binary cross entropy loss function L, which is as follows:

[0088]

[0089] Among them, y i represents the label (true value label) of training sample i, the positive class is 1, the negative class is 0, and p i It represents the probability that training sample i is predicted to be a positive class (i.e., predicted probability).

[0090] It is understandable that after updating the initial task prediction model through multiple rounds of iterations based on the training samples, a downstream task prediction model with optimal parameters can be obtained.

[0091] This embodiment converts the spatiotemporal trajectory points into spatiotemporal trajectory texts based on the multi-scale spatial grid division results; constructs a spatiotemporal pre-trained large model based on the spatiotemporal trajectory text and the pre-trained language model, wherein the pre-trained language model includes multiple encoding units, and the encoding units include a multi-head attention module, a residual module, a normalization module, and a feedforward neural network module; inputs the target spatiotemporal trajectory text of the target user into the spatiotemporal pre-trained large model to obtain the target user representation; inputs the target user representation into the downstream task prediction model, and outputs the downstream task prediction information. This solves the technical problem that the traditional trajectory modeling method does not fully learn the user trajectory characteristics, making it difficult to apply the user representation to user portrait and travel mode tasks. Compared with the prior art, this application avoids the loss of user trajectory details by unifying the spatiotemporal business logic, and then learns the preferences and rules of the user's spatiotemporal travel activities through the spatiotemporal pre-trained large model, which can improve the accuracy of the user representation, thereby ensuring the accurate prediction of downstream tasks.

[0092] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated in the following. Figure 4 , step S20 also includes steps S201-S202:

[0093] Step S201, after embedding the spatiotemporal trajectory text, an input vector is obtained, wherein the embedding process includes text embedding, position embedding and segment embedding;

[0094] Step S202, determining a first cross entropy and a second cross entropy of a pre-training task based on the input vector and the pre-trained language model;

[0095] It should be noted that the pre-training tasks of the pre-trained language model are divided into two. The first pre-training task is to predict the text of the masked part, and the second pre-training task is to predict whether two trajectory texts belong to the same person.

[0096] In a feasible implementation, the step of determining the first cross entropy and the second cross entropy of the pre-training task based on the input vector and the pre-trained language model includes: randomly processing the input vector using a mask flag to obtain a masked input vector; obtaining the first cross entropy of the pre-training task based on the masked input vector, the input vector and the pre-trained language model; randomly obtaining a first vector and a second vector from the input vector, and determining user information between the first vector and the second vector, wherein the user information includes one of belonging to the same user and not belonging to the same user; inputting the first vector and the second vector into the pre-trained language model to obtain inference information, and determining the second cross entropy of the pre-training task based on the user information and the inference information.

[0097] In the specific implementation, Figure 5 As shown in the figure, for the first pre-training task, the MASK flag is used to randomly mask the input vector, allowing the model to autonomously learn the trajectory features before and after the masked point, and then predict the masked point text (i.e., mask the input vector). The first pre-training task is a multi-classification task, and cross entropy is used as the loss function. The function calculation formula of the first cross entropy is:

[0098]

[0099] Among them, θ is the parameter of the encoding part of the Bert model, θ1 is the parameter in the output layer of the encoding part for the task of predicting the masked text, M is the dictionary set, and V is the dictionary size.

[0100] In the specific implementation, Figure 6 As shown in the figure, the second pre-training task is to predict whether two spatiotemporal trajectory texts belong to the same person. For example, if two users’ spatiotemporal trajectory texts are input on the same day, the model can infer whether they belong to the same person’s behavior trajectory by learning the deep features of the two trajectories. This task is a binary classification task, and cross entropy is used as the loss function. The function calculation formula of the second cross entropy is:

[0101]

[0102] Among them, θ is the parameter of the encoding part of the Bert model, θ2 is the classifier parameter connected to the encoding part in the sentence prediction task, and N is a logical judgment value 0 or 1. 0 represents that the two trajectory texts of the input model do not belong to the same person, and 1 represents that the two trajectory texts of the input model do not belong to the same person.

[0103] It should be noted that if Figure 5 as well as Figure 6 As shown in the figure, FFNN (Feedforward Fully Connected Neural Network) is a fully connected neural network, Softmax is a normalized exponential function, Self-Attention, Add, Norm, and Feed Forward are the self-attention module, residual module, normalization module, and feedforward neural network module in the Transformers structure respectively. Masked is the masked part of the trajectory text.

[0104] Step S203, determining a model iteration error based on the first cross entropy and the second cross entropy;

[0105] It can be understood that the final model iteration error can be obtained by adding the first cross entropy and the second cross entropy.

[0106] Step S204, after updating the model parameters of the pre-trained language model according to the model iteration error, a spatiotemporal pre-trained large model is obtained.

[0107] It is understandable that the model iteration error can be obtained through multiple rounds of iterations, and then the model parameters of the pre-trained language model can be updated according to the model error.

[0108] This embodiment will enable the model to fully learn the user trajectory features by enabling the model to predict covered trajectory text, predict whether two trajectory texts belong to the same person, and other types of pre-training tasks, thereby avoiding the loss of trajectory information and improving the accuracy of user representation.

[0109] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the task prediction method of the present application based on a multi-scale spatiotemporal pre-trained large model. More forms of simple transformations based on this technical concept are all within the scope of protection of the present application.

[0110] This application also provides a task prediction device based on a multi-scale spatiotemporal pre-trained large model, please refer to Figure 7 , the task prediction device based on the multi-scale spatiotemporal pre-trained large model includes:

[0111] A conversion module 10, for converting the spatiotemporal trajectory points into spatiotemporal trajectory texts based on the multi-scale spatial grid division results;

[0112] A construction module 20 is used to construct a spatiotemporal pre-trained large model based on the spatiotemporal trajectory text and the pre-trained language model, wherein the pre-trained language model includes a plurality of encoding units, and the encoding unit includes a multi-head attention module, a residual module, a normalization module, and a feedforward neural network module;

[0113] An input module 30, used to input the target spatiotemporal trajectory text of the target user into the spatiotemporal pre-trained large model to obtain a target user representation;

[0114] The output module 40 is used to input the target user representation into the downstream task prediction model and output downstream task prediction information.

[0115] The task prediction device based on the multi-scale spatiotemporal pre-trained large model provided in the present application adopts the task prediction method based on the multi-scale spatiotemporal pre-trained large model in the above-mentioned embodiment, which can solve the technical problem that the traditional trajectory modeling method does not fully learn the user trajectory characteristics, making it difficult to apply the user representation to user portrait and travel mode tasks. Compared with the prior art, the beneficial effects of the task prediction device based on the multi-scale spatiotemporal pre-trained large model provided in the present application are the same as the beneficial effects of the task prediction method based on the multi-scale spatiotemporal pre-trained large model provided in the above-mentioned embodiment, and the other technical features of the task prediction device based on the multi-scale spatiotemporal pre-trained large model are the same as the features disclosed in the above-mentioned embodiment method, which will not be repeated here.

[0116] In one embodiment, the conversion module 10 is further used to: determine the grids at all levels across the country based on the administrative division method and the grid division method; use preset characters to represent the grids at all levels across the country to obtain a multi-scale spatial grid division result; based on the multi-scale spatial grid division result, convert the space-time trajectory points into original space-time trajectory text; perform word segmentation on the national space-time trajectory text to obtain a word segmentation result, and determine high-frequency byte pairs based on the word segmentation result; merge the high-frequency byte pairs into new sub-words, and construct a trajectory word table based on the new sub-words; based on the trajectory word table, convert the original space-time trajectory text into space-time trajectory text.

[0117] In one embodiment, the construction module 20 is also used to: obtain an input vector after embedding the spatiotemporal trajectory text, wherein the embedding processing includes text embedding, position embedding and segment embedding; determine a first cross entropy and a second cross entropy of the pre-training task based on the input vector and the pre-trained language model; determine a model iteration error based on the first cross entropy and the second cross entropy; and obtain a spatiotemporal pre-training large model after updating the model parameters of the pre-trained language model according to the model iteration error.

[0118] In one embodiment, the construction module 20 is further used to: randomly process the input vector using a mask flag to obtain a masked input vector; obtain a first cross entropy of the pre-training task based on the masked input vector, the input vector and the pre-trained language model; randomly obtain a first vector and a second vector from the input vector, and determine user information between the first vector and the second vector, wherein the user information includes one of belonging to the same user and not belonging to the same user; input the first vector and the second vector into the pre-trained language model to obtain inference information, and determine the second cross entropy of the pre-training task based on the user information and the inference information.

[0119] In one embodiment, the spatiotemporal pre-trained large model includes a national-scale large model, an urban agglomeration-scale large model and a city-scale large model; wherein the input module 30 is further used to: input the target spatiotemporal trajectory text into the national-scale large model, the urban agglomeration-scale large model and the national-scale large model respectively to obtain a national-scale representation, an urban agglomeration-scale representation and a city-scale representation; use the national-scale representation, the urban agglomeration-scale representation and the city-scale representation as optional representations; determine a target optional representation from the optional representations according to the task type of the target user, and splice the target optional representations into a target user representation.

[0120] In one embodiment, the output module 40 is further used to: obtain training samples, wherein the training samples include user representations and true value labels; input the user representations into the initial task prediction model and output prediction probabilities; determine a cross entropy loss value based on the true value labels and the prediction probabilities; and obtain a downstream task prediction model after updating the model parameters of the initial task prediction model based on the cross entropy loss value.

[0121] The present application provides a task prediction device based on a multi-scale spatiotemporal pre-trained large model, and the task prediction device based on the multi-scale spatiotemporal pre-trained large model includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the task prediction method based on the multi-scale spatiotemporal pre-trained large model in the above-mentioned embodiment one.

[0122] Reference below Figure 8 , which shows a schematic diagram of the structure of a task prediction device based on a multi-scale spatiotemporal pre-trained large model suitable for implementing an embodiment of the present application. The task prediction device based on a multi-scale spatiotemporal pre-trained large model in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The task prediction device based on the multi-scale spatiotemporal pre-trained large model shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0123] like Figure 8As shown, the task prediction device based on the multi-scale spatiotemporal pre-trained large model may include a processing device 1001 (such as a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 to the random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the task prediction device based on the multi-scale spatiotemporal pre-trained large model are also stored. The processing device 1001, ROM1002 and RAM1004 are connected to each other via a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 1009. The communication device 1009 can allow the task prediction device based on the multi-scale spatiotemporal pre-trained large model to communicate wirelessly or wired with other devices to exchange data. Although the figure shows a task prediction device based on a multi-scale spatiotemporal pre-trained large model with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or provided alternatively.

[0124] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0125] The task prediction device based on the multi-scale spatiotemporal pre-trained large model provided by the present application adopts the task prediction method based on the multi-scale spatiotemporal pre-trained large model in the above-mentioned embodiment, which can solve the technical problem that the traditional trajectory modeling method does not fully learn the user trajectory characteristics, making it difficult to apply the user representation to user portrait and travel mode tasks. Compared with the prior art, the beneficial effects of the task prediction device based on the multi-scale spatiotemporal pre-trained large model provided by the present application are the same as the beneficial effects of the task prediction method based on the multi-scale spatiotemporal pre-trained large model provided by the above-mentioned embodiment, and the other technical features of the task prediction device based on the multi-scale spatiotemporal pre-trained large model are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0126] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0127] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0128] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the task prediction method based on the multi-scale spatiotemporal pre-trained large model in the above-mentioned embodiment.

[0129] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0130] The above-mentioned computer-readable storage medium may be included in the task prediction device based on the multi-scale spatiotemporal pre-trained large model; or it may exist independently without being assembled into the task prediction device based on the multi-scale spatiotemporal pre-trained large model.

[0131] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by a task prediction device based on a multi-scale spatiotemporal pre-trained large model, the task prediction device based on the multi-scale spatiotemporal pre-trained large model: based on the multi-scale spatial grid division result, converts the spatiotemporal trajectory points into spatiotemporal trajectory text; based on the spatiotemporal trajectory text and the pre-trained language model, constructs a spatiotemporal pre-trained large model, wherein the pre-trained language model includes multiple encoding units, and the encoding units include a multi-head attention module, a residual module, a normalization module and a feedforward neural network module; inputs the target spatiotemporal trajectory text of the target user into the spatiotemporal pre-trained large model to obtain a target user representation; inputs the target user representation into a downstream task prediction model, and outputs downstream task prediction information.

[0132] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0133] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0134] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0135] The readable storage medium provided in the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned task prediction method based on the multi-scale spatiotemporal pre-trained large model, which can solve the technical problem that the traditional trajectory modeling method does not fully learn the user trajectory characteristics, making it difficult to apply user representation to user portrait and travel mode tasks. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in the present application are the same as the beneficial effects of the task prediction method based on the multi-scale spatiotemporal pre-trained large model provided in the above-mentioned embodiment, and will not be repeated here.

[0136] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the task prediction method based on the multi-scale spatiotemporal pre-trained large model as described above.

[0137] The computer program product provided by the present application can solve the technical problem that the traditional trajectory modeling method does not fully learn the user trajectory characteristics, making it difficult to apply user representation to user portrait and travel mode tasks. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as the beneficial effects of the task prediction method based on the multi-scale spatiotemporal pre-trained large model provided in the above embodiment, which will not be repeated here.

[0138] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A task prediction method based on a multi-scale spatiotemporal pre-trained large model, characterized in that: The method comprises: Based on the multi-scale spatial grid division results, the space-time trajectory points are converted into space-time trajectory texts; Based on the spatiotemporal trajectory text and the pre-trained language model, a spatiotemporal pre-trained large model is constructed, wherein the pre-trained language model includes a plurality of encoding units, and the encoding units include a multi-head attention module, a residual module, a normalization module, and a feedforward neural network module; Inputting the target spatiotemporal trajectory text of the target user into the spatiotemporal pre-trained large model to obtain a target user representation; The target user representation is input into a downstream task prediction model, and downstream task prediction information is output.

2. The method according to claim 1, characterized in that The step of converting the space-time trajectory points into space-time trajectory text based on the multi-scale spatial grid division result comprises: Based on the administrative division and grid division methods, determine the grids at all levels across the country; Preset characters are used to represent the grids at all levels across the country, and a multi-scale spatial grid division result is obtained; Based on the multi-scale spatial grid division result, the space-time trajectory points are converted into original space-time trajectory text; Segment the national spatiotemporal trajectory text to obtain segmentation results, and determine high-frequency byte pairs according to the segmentation results; Merging the high-frequency byte pairs into new subwords, and constructing a trajectory word list according to the new subwords; Based on the trajectory vocabulary, the original spatiotemporal trajectory text is converted into a spatiotemporal trajectory text.

3. The method according to claim 1, characterized in that The step of constructing a spatiotemporal pre-trained large model based on the spatiotemporal trajectory text and the pre-trained language model comprises: After embedding the spatiotemporal trajectory text, an input vector is obtained, wherein the embedding process includes text embedding, position embedding and segment embedding; Determine a first cross entropy and a second cross entropy of a pre-training task based on the input vector and the pre-trained language model; Determining a model iteration error based on the first cross entropy and the second cross entropy; After updating the model parameters of the pre-trained language model according to the model iteration error, a spatiotemporal pre-trained large model is obtained.

4. The method according to claim 3, characterized in that The step of determining a first cross entropy and a second cross entropy of a pre-training task based on the input vector and the pre-trained language model comprises: The input vector is randomly processed using a mask flag to obtain a masked input vector; Obtaining a first cross entropy of a pre-training task based on the masked input vector, the input vector, and the pre-trained language model; Randomly obtain a first vector and a second vector from the input vector, and determine user information between the first vector and the second vector, wherein the user information includes one of belonging to the same user and not belonging to the same user; The first vector and the second vector are input into the pre-trained language model to obtain inference information, and a second cross entropy of the pre-training task is determined according to the user information and the inference information.

5. The method according to claim 1, characterized in that The spatiotemporal pre-trained large model includes a national scale large model, an urban agglomeration scale large model and a city scale large model; wherein the step of inputting the target spatiotemporal trajectory text of the target user into the spatiotemporal pre-trained large model to obtain the target user representation includes: Inputting the target spatiotemporal trajectory text into the national scale large model, the urban agglomeration scale large model and the national scale large model respectively, to obtain a national scale representation, an urban agglomeration scale representation and a city scale representation; The national scale representation, the urban agglomeration scale representation and the city scale representation are taken as optional representations; According to the task type of the target user, a target optional representation is determined from the optional representations, and the target optional representations are spliced ​​into a target user representation.

6. The method according to claim 1, characterized in that Before the step of inputting the target user representation into the downstream task prediction model and outputting the downstream task prediction information, the step further includes: Acquire a training sample, wherein the training sample includes a user representation and a true value label; Inputting the user representation into the initial task prediction model and outputting the prediction probability; Determining a cross entropy loss value based on the true value label and the predicted probability; Based on the cross entropy loss value, after updating the model parameters of the initial task prediction model, a downstream task prediction model is obtained.

7. A task prediction device based on a multi-scale spatiotemporal pre-trained large model, characterized in that: The device comprises: A conversion module is used to convert the space-time trajectory points into space-time trajectory text based on the multi-scale spatial grid division results; A construction module, used to construct a spatiotemporal pre-trained large model based on the spatiotemporal trajectory text and the pre-trained language model, wherein the pre-trained language model includes a plurality of encoding units, and the encoding units include a multi-head attention module, a residual module, a normalization module, and a feedforward neural network module; An input module, used to input the target spatiotemporal trajectory text of the target user into the spatiotemporal pre-trained large model to obtain a target user representation; The output module is used to input the target user representation into the downstream task prediction model and output the downstream task prediction information.

8. A task prediction device based on a multi-scale spatiotemporal pre-trained large model, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the task prediction method based on a multi-scale spatiotemporal pre-trained large model as described in any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the task prediction method based on a multi-scale spatiotemporal pre-trained large model as described in any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the steps of the task prediction method based on a multi-scale spatiotemporal pre-trained large model as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Community multi-task prediction method based on spatial-temporal characteristics and related equipment

    CN121502666A