A sequence data modeling method and system

CN120874904BActive Publication Date: 2026-09-04SHANGHAI QINIU INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511009371.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2026-09-04
Estimated Expiration
2045-07-22

AI Technical Summary

Technical Problem

[0005]对物理时空上下文变化的感知不够直接和精细:标准RNN模型的状态更新缺乏对连续事件之间物理空间距离和时间间隔的具体量化感知

Benefits of technology

[0050] 1. Enhanced sensitivity of state representation to spatiotemporal dynamics: The generated sequential state vector s_t can more accurately reflect the continuous, changing or jumping state of the user's physical spatiotemporal environment, improving the authenticity and immediacy of state representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120874904B_ABST
    Figure CN120874904B_ABST
Patent Text Reader

Abstract

The application discloses a sequence data modeling method and system. The method comprises the following steps: S1, converting a multi-source heterogeneous original feature set at a current moment into an embedded vector of a unified dimension through a learnable mapping function, and integrating the embedded vector into a comprehensive feature embedded vector E_t; S2, calculating a space-time difference between the current event and a previous event, inputting the space-time difference into a learnable network after processing as a structured feature to generate a space-time correlation score vector R_st; and S3, inputting the R_st into a gated recurrent neural network as an adjusting signal to adjust a gate and update a cell state, so as to obtain a sequence state vector s_t. The enhanced state representation is sensitive to space-time dynamics, and the space-time correlation mode is improved in interpretability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data analysis technology, and in particular to a method and system for modeling sequence data. Background Technology

[0002] With the rapid development of mobile internet and IoT technologies, users generate a large amount of data sequences containing rich spatiotemporal information in their daily lives, such as user movement trajectories, Points of Interest (POI) query records, and online activity logs. These data have a certain temporal order or logical dependency. Modeling these spatiotemporal sequence data and conducting in-depth analysis and understanding are of great significance for predicting users' future spatiotemporal query intentions or behaviors, improving the quality of personalized services, optimizing resource allocation, and enhancing user experience.

[0003] Existing sequence data modeling methods have achieved significant success in fields such as natural language processing and speech recognition, and have been widely applied to modeling user behavior sequences. These methods can effectively capture the temporal dependencies and some semantic information in the sequence.

[0004] However, when processing user behavior sequences with strong physical spatiotemporal attributes, traditional RNN models mainly rely on embedding and learning spatiotemporal features (such as latitude and longitude coordinates and timestamps) as ordinary input features. Their state update mechanism focuses more on the semantic coherence of the input content and the order of the sequence. This approach has the following problems and limitations:

[0005] The perception of changes in physical spatiotemporal context is not direct or refined enough: the state updates of standard RNN models lack a specific quantitative perception of the physical spatial distance and time intervals between consecutive events. The model struggles to distinguish the essential difference between a small location move and a large-scale migration across cities in terms of their impact on the user's subsequent intentions, and it also struggles to accurately measure the impact of time intervals of different lengths on the continuity of context.

[0006] The blindness of information filtering and updating: Due to the lack of clear judgment on the degree of spatiotemporal leaps, RNNs may not be intelligent enough in deciding which historical information to retain, which outdated information to forget, and how much new information to absorb. For example, when a user undergoes a significant scene change (such as from home to work, or from one city to another), the model may still retain a large amount of invalid information related to the previous scene, or fail to fully absorb the key information of the new scene, resulting in contamination or lag in state representation.

[0007] The state representation is not adaptable to spatiotemporal dynamics: the generated hidden state vector may not accurately reflect the real dynamic changes in the spatiotemporal environment in which the user is located, resulting in poor performance in downstream tasks that require precise spatiotemporal context for decision-making (such as near-field POI recommendation and service alerts in specific areas).

[0008] Learning efficiency and interpretability issues: Using only raw spatiotemporal data as input features and relying on complex neural network structures to implicitly learn spatiotemporal correlation patterns from a large amount of data may not only require more training data and computational resources, but the learned spatiotemporal correlation patterns often lack intuitive physical interpretations.

[0009] Therefore, how to design a method that can explicitly perceive and quantify the physical spatiotemporal correlation between continuous events, and dynamically and intelligently adjust the sequence state update process accordingly, so as to generate a state representation that is more sensitive to and more adaptive to spatiotemporal dynamic changes, is a technical problem that urgently needs to be solved in the field of spatiotemporal sequence data modeling. Summary of the Invention

[0010] This application provides a sequence data modeling method to address the technical problems existing in the prior art, comprising the following steps:

[0011] Step S1: Convert the multi-source heterogeneous original feature set at the current time into a uniform-dimensional embedding vector through a learnable mapping function, and integrate it into a comprehensive feature embedding vector E_t;

[0012] Step S2: Calculate the spatiotemporal difference between the current event and the preceding event, process it into a structured feature, and then input it into a learnable network to generate a spatiotemporal correlation score vector R_st;

[0013] Step S3: Use R_st as the adjustment signal, input it into the gated recurrent neural network to adjust the gating, update the cell state, and obtain the sequence state vector s_t.

[0014] Further, in step S2, the spatiotemporal difference between the current event and the preceding event is calculated, processed into a structured feature, and then input into a learnable network to generate a spatiotemporal correlation score vector R_st. Specifically, this includes:

[0015] S21: Obtain original spatiotemporal features: Extract the original spatiotemporal features of the current time step t and the original spatiotemporal features of the immediately preceding time step t-1;

[0016] S22: Calculate the original spatiotemporal differences: Based on the above two sets of original spatiotemporal characteristics, calculate the physical spatial differences between them to form a spatial difference vector ΔS_t, and a physical time difference ΔT_t;

[0017] S23: Difference feature transformation and integration: The calculated original spatiotemporal differences ΔS_t and ΔT_t are processed and finally aggregated into a unified, structured difference feature vector F_diff_t;

[0018] S24: Generate a spatiotemporal correlation score vector: Input the unified difference feature vector F_diff_t into a learnable multilayer perceptron, and generate a spatiotemporal correlation score vector R_st with dimension d_r through nonlinear mapping.

[0019] Furthermore, step S3 includes:

[0020] S31: Input the current feature embedding vector E_t, the hidden state s_{t-1} of the previous time step, the cell state C_{t-1}, and the spatiotemporal correlation score vector R_st;

[0021] S32: Input into an LSTM-like structure, integrating the spatiotemporal correlation score vector R_st into the forget gate f_t, input gate i_t, and candidate new information vector. In the calculation of the output gate o_t, the information forgetting, new information absorption, content generation and state output are dynamically adjusted to generate the final hidden state vector s_t.

[0022] Furthermore, step S32 specifically includes:

[0023] S321: Calculate the spatiotemporally adjusted forgetting gate f_t, the formula for which is:

[0024] f_t=σ(W_fE_t+U_fs_{t-1}+V_fR_{st}+b_f)

[0025] S322: Calculate the spatiotemporally adjusted input gate i_t, the formula for which is:

[0026] i_t=σ(W_i E_t+U_i s_{t-1}+V_i R_{st}+b_i)

[0027] S323: Calculate the spatiotemporally adjusted candidate new information vector The calculation formula is as follows:

[0028]

[0029] S324: Update the cell state C_t, which is calculated using the following formula:

[0030]

[0031] S325: Calculate the spatiotemporally adjusted output gate o_t, the calculation formula is as follows:

[0032] o_t=σ(W_o E_t+U_o s_{t-1}+V_o R_{st}+b_o)

[0033] S326: Calculate the current hidden state s_t, the formula for which is:

[0034] s_t=o_t⊙tanh(C_t).

[0035] Furthermore, the gated recurrent neural network is either LSTM or GRU.

[0036] Furthermore, step S1 specifically includes:

[0037] S11: POI type feature embedding;

[0038] S12: Geographic coordinate feature embedding;

[0039] S13: Embedding of user activity state features;

[0040] S14: Concatenate the above features to form a comprehensive feature embedding vector.

[0041] Furthermore, in step S12: geographic coordinate feature embedding, the embedding method is any one of MLP embedding, Geohash encoding embedding, embedding combined with geographic knowledge graph information, or node embedding based on graph neural network.

[0042] Furthermore, the concatenation method in step S14 is: vector concatenation or weighted summation, or dynamic fusion through a small attention network, or dimensionality reduction or feature crossing of the concatenated vectors through an MLP.

[0043] The present invention also provides a sequence data modeling system, comprising: a data embedding processing unit, a spatiotemporal correlation calculation unit, and an update unit, wherein:

[0044] The data embedding processing unit is used to convert the multi-source heterogeneous original feature set at the current time into a uniform-dimensional embedding vector through a learnable mapping function, and integrate it into a comprehensive feature embedding vector E_t;

[0045] The spatiotemporal correlation calculation unit is used to calculate the spatiotemporal difference between the current event and the preceding event, process it into structured features, and then input it into a learnable network to generate a spatiotemporal correlation score vector R_st;

[0046] The update unit uses R_st as an adjustment signal, inputs it into the gated recurrent neural network to adjust the gating, update the cell state, and obtain the sequence state vector s_t.

[0047] The present invention also provides an electronic device, comprising: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the sequence data modeling method as described above.

[0048] In practical applications, the modules described in the methods and systems disclosed in this application can be deployed on a single target server, or each module can be deployed independently on different target servers. In particular, as needed, to provide more powerful computing capabilities, the modules can also be deployed on a cluster of target servers.

[0049] Therefore, the technical effects achieved by the technical approach adopted in this application are as follows:

[0050] 1. Enhanced sensitivity of state representation to spatiotemporal dynamics: The generated sequential state vector s_t can more accurately reflect the continuous, changing or jumping state of the user's physical spatiotemporal environment, improving the authenticity and immediacy of state representation.

[0051] Second: More intelligent information filtering and updating: When a significant spatiotemporal scene transition is detected (reflected by R_st), the model can more effectively forget outdated information related to the old scene and more actively absorb new information related to the new scene, avoiding state contamination and improving the efficiency and accuracy of memory updates.

[0052] Third: Improve the performance of downstream tasks: Since the state vector s_t captures the dynamic evolution of the spatiotemporal context more accurately, it is expected to improve the performance when applied to downstream tasks that rely on spatiotemporal information (such as user future behavior prediction, personalized recommendation, location-aware services, etc.).

[0053] IV. Enhanced interpretability of spatiotemporal correlation patterns: Through the design of R_st, the model's response to different spatiotemporal transition patterns (such as short time at close range, long time at long distance, etc.) can be observed and understood more clearly, providing a certain degree of interpretability for model behavior analysis.

[0054] Fifth, enhanced adaptability: The model can adaptively judge the importance of different degrees and types of spatiotemporal changes on the user's state based on learned experience, and adjust information processing strategies accordingly.

[0055] To provide a clearer and more comprehensive understanding of this application, the specific embodiments of this application will be described in detail below with reference to the accompanying drawings. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0057] Figure 1 This is a flowchart illustrating a sequence data modeling method according to an embodiment of this application.

[0058] Figure 2 This is a schematic diagram of feature embedding mapping in an embodiment of this application. Detailed Implementation

[0059] Please see Figure 1 The technical solution of this application provides a method for modeling sequence data, including the following steps:

[0060] Step S1: Convert the multi-source heterogeneous original feature set at the current time into a uniform-dimensional embedding vector through a learnable mapping function, and integrate it into a comprehensive feature embedding vector E_t;

[0061] In this step, at each discrete time step t, a multi-source heterogeneous raw feature set X_raw_t containing the user's current activity is received (e.g., containing geographic coordinates x_raw_loc_t, POI category x_raw_type_t, timestamp x_raw_time_t, user's current activity state x_raw_useractivity_t, etc.). These raw features x_raw_j_t are transformed into a unified-dimensional embedding vector e_j_t using their respective learnable mapping functions φ_j (such as embedding matrices, multilayer perceptrons, etc.). Subsequently, all individual feature embedding vectors e_j_t are integrated (e.g., concatenated) to form a comprehensive feature embedding vector E_t, serving as a comprehensive numerical representation of the user event and its context at the current moment.

[0062] Step S2: Calculate the spatiotemporal difference between the current event and the preceding event, process it into a structured feature, and then input it into a learnable network to generate a spatiotemporal correlation score vector R_st;

[0063] This step is based on the current integrated feature embedding vector E_t (from stage 1) and the state vector from the previous time step. A state update function Ψ_update is used to generate and output the current state vector. Subsequently, by calculating the spatiotemporal correlation, the inherent correlation strength and pattern between the current input event and the immediately preceding input event in the user activity sequence can be quantitatively evaluated in the physical space and time dimensions.

[0064] Step S3: Use R_st as the adjustment signal, input it into the gated recurrent neural network to adjust the gating, update the cell state, and obtain the sequence state vector s_t.

[0065] After calculating the score vector R_st, which represents the spatiotemporal correlation pattern between the current event and the preceding event, this R_st will be used as a key adjustment signal and deeply integrated into the gating logic of the dynamic state update module.

[0066] All learnable parameters involved (including the parameters θ_φ_j of the feature embedding module, the parameters θ_{SRC} of MLP_SRC, and the weight matrices V_f, Vi, V_c, V_o that interact with R_st in the gated loop structure) are obtained by end-to-end training of the entire method and optimization learning of the loss function L_total for specific downstream tasks (such as future query intent prediction).

[0067] The technical solutions of this application are described below with reference to various embodiments.

[0068] Step S1: Convert the multi-source heterogeneous original feature set at the current time into a uniform-dimensional embedding vector through a learnable mapping function, and integrate it into a comprehensive feature embedding vector E_t;

[0069] The first stage of this application is the embedding processing of the input raw features. At each discrete time step t, the system receives the raw feature set X_raw_t. For example, X_raw_t = {x_raw_loc_t, x_raw_type_t, x_raw_time_t, x_raw_useractivity_t, ...}, where x_raw_loc_t is the geographic coordinates (lon, lat), x_raw_type_t is the POI category (such as "restaurant", "office"), x_raw_time_t is the timestamp information, and x_raw_useractivity_t is the user's current activity state.

[0070] The goal of this stage is to transform these heterogeneous raw features x_raw_j_t (where j represents the index of the feature) into embedding vectors of a uniform dimension (e.g., d_emb dimension) using their respective learnable mapping functions φ_j.

[0071] e_j_t=φ_j(x_raw_j_t; θ_φ_j)

[0072] Where θ_φ_j are learnable parameters of the mapping function φ_j. These parameters, such as the embedding matrices W_type_emb and W_activity_emb described later, and the weights and biases of the multilayer perceptron MLP_loc, can be learned by performing end-to-end training on the method described in this invention or by referring to other existing technologies, and by optimizing the loss function for a specific downstream task (such as future query intent prediction).

[0073] Suppose that at time t, a user queries "coffee shop". The original features include: x_raw_type_t = "coffee shop" (discrete category), x_raw_loc_t = (lon_A, lat_A) (continuous coordinates), and x_raw_useractivity_t = "walking" (discrete user state). An example of the specific feature embedding and mapping process is as follows (e.g.) Figure 2 As shown):

[0074] As a preferred embodiment, the specific steps include the following S11-S14:

[0075] S11: POI type feature embedding, its embedding representation is as follows:

[0076] e_type_t=φ_type("coffee shop"; θ_φ_type)

[0077] φ_type is an embedding matrix V_type refers to the set of all POI categories. For example, if there are 1000 different POI categories (such as "coffee shop", "library", "bank", "park", etc.), then |V_type| = 1000. d_emb represents the set of real numbers, indicating that all elements in the matrix are real numbers; d_emb is the embedding dimension, which is a preset positive integer indicating how many dimensions of vector are desired to represent each POI category.

[0078] Therefore, W_type_emb is a matrix with |V_type| rows and d_emb columns, and each element of the matrix is ​​a real number. The obtained embedding vector e_type_t is the row vector corresponding to "coffee shop" in this matrix.

[0079] like Figure 2 A specific example regarding W_type_emb and e_type_t is given. Suppose we only care about the following POI categories and choose the embedding dimension d_emb = 4, the vocabulary V_type {"coffee shop", "library", "tea house", "bank", "office"}, then |V_type| = 5, and the embedding matrix W_type_emb will be a 5x4 matrix.

[0080] Before training begins, the elements of this matrix are typically initialized randomly, for example, by sampling randomly from a standard normal or uniform distribution, appearing as follows: Figure 2 As shown.

[0081] Through learning, this embedding vector can capture the semantic information of POI categories, making the representations of semantically similar categories (such as "coffee shop" and "tea house") in the embedding space potentially more similar.

[0082] When you need to obtain the embedding vector of a certain POI type (such as "coffee shop"), you are looking up which one in this matrix corresponds to "coffee shop", that is, e_coffee shop = W_type_emb["coffee shop"] = [0.12, -0.34, 0.56, -0.08].

[0083] S12: Geographic coordinate feature embedding, represented as:

[0084] e_loc_t=MLP_loc((lon_A,lat_A);θ_MLP_loc)

[0085] For the original feature x_raw_loc_t = (lon_A, lat_A) (a continuous pair of longitude and latitude values), its mapping function φ_loc can be implemented using a small multilayer perceptron (MLP_loc). This MLP_loc receives two-dimensional coordinate input, passes it through a series of hidden layers with non-linear activation functions, and finally outputs a d_emb-dimensional embedding vector e_loc_t. Its learnable parameters θ_MLP_loc include the weight matrices and bias vectors of each layer.

[0086] For example, Through learning, e_loc_t is not just a simple transformation of the original coordinates, but can also capture the abstract geographic semantics or regional functional characteristics of the geographic location in a specific task context (e.g., whether it is located in a business district, a place frequently visited by users, etc.).

[0087] Besides using MLP for embedding geographic coordinates, other methods such as Geohash encoding or combining geographic knowledge graph information can be considered for richer representation. For specific regions (such as road networks), node embedding methods based on graph neural networks can also be used.

[0088] S13: Embedding User Activity State Features

[0089] For the original feature x_raw_useractivity_t = "Walking" (a discrete category), its mapping function φ_activity is implemented similarly to POI-type feature embedding, and is also an embedding matrix.

[0090] Where |V_activity| is the total number of all possible user activity states (such as "walking", "driving", "stationary", "working", etc.).

[0091] e_useractivity_t = W_activity_emb["Walking"]

[0092] The structure of W_activity_emb is similar to that of W_type_emb. Each row of W_activity_emb represents a d_emb-dimensional embedding vector for a user activity state. The embedding vector e_useractivity_t is the row of vectors in W_activity_emb corresponding to "walking". For example,

[0093] Through learning, e_useractivity_t is able to capture the potential impact of different activity states on user intent.

[0094] S14: Concatenate the above features to form a comprehensive feature embedding vector.

[0095] At time t, the embedding vectors e_j_t obtained after transforming all individual original features x_raw_j_t through their respective embedding functions φ_j are concatenated and integrated into a single, higher-dimensional comprehensive feature embedding vector E_t:

[0096] E_t=[e_type_t; e_loc_t; e_useractivity_t;...]

[0097] Where [;] denotes the (e.g., vertical) concatenation of vectors. If each e_j_t is d_emb-dimensional and a total of N_feat original features are embedded, then the dimension D_E of E_t is N_feat×d_emb.

[0098] This E_t constitutes a comprehensive, numerical, high-dimensional representation of the user query event at time t and its related context, with the relationships between features already preliminarily encoded. It will serve as the direct input for subsequent dynamic state update processing.

[0099] Besides vector concatenation, E_t can also be generated by weighted summation (with learnable weights), dynamic fusion through a small attention network, or by passing the concatenated vector through an MLP for dimensionality reduction or feature crossing.

[0100] Step S2: Calculate the spatiotemporal difference between the current event and the preceding event, process it into a structured feature, and then input it into a learnable network to generate a spatiotemporal correlation score vector R_st.

[0101] This step is based on the current integrated feature embedding vector E_t (from stage 1) and the state vector from the previous time step. A state update function Ψ_update is used to generate and output the current state vector. The formula is as follows:

[0102] s_t=Ψ_update(E_t,s_{t-1}; θ_Ψ)

[0103] Here, θ_Ψ represents the learnable parameters of the internal structure of the state update function Ψ_update (e.g., the weights and biases of the gating units, described later). The purpose of Ψ_update is to effectively capture the temporal information flow, determining how historical information is retained or forgotten, and how new information is incorporated.

[0104] This application introduces a "spatiotemporal context-sensitive gating mechanism" for Ψ_update, which makes the update of the state not only dependent on the semantic similarity of the content, but also explicitly and quantitatively considers the correlation patterns of continuous events in the physical spatiotemporal dimension (such as continuity and jump).

[0105] Subsequently, through spatiotemporal correlation calculation, the goal is to quantitatively assess the inherent correlation strength and pattern between the current input event and the immediately preceding input event in the physical space and time dimensions within a user activity sequence.

[0106] This step calculates the spatiotemporal correlation degree, and the calculation result generates a clear spatiotemporal correlation score vector R_st, which provides a key adjustment signal for the subsequent gating mechanism.

[0107] As a preferred embodiment, it specifically includes:

[0108] S21, Obtain the original spatiotemporal features: Extract the original spatiotemporal features (x_raw_loc_t, x_raw_time_t) of the current time step t and the original spatiotemporal features (x_raw_loc_{t-1}, x_raw_time_{t-1}) of the immediately preceding time step t-1.

[0109] The spatiotemporal feature input for the current time step t: X_raw_t is the raw feature set generated by the user at discrete time step t. The main features to focus on are as follows:

[0110] The original geospatial coordinates representing the current event occurring at time t are usually represented as (lon_t, lat_t), where lon_t is the longitude and lat_t is the latitude.

[0111] The original timestamp representing the current event at time t is a numerical value that precisely represents absolute time.

[0112] For example, in the coffee shop query example, a user queries "coffee shop" at time t while walking in "area B" some time after ending a meeting. Then, x_raw_loc_t = (lon_B, lat_B) (the coordinates of the coffee shop in area B), and x_raw_time_t = T_B (the exact time of the coffee shop query).

[0113] The spatiotemporal feature input from the previous time step t-1: Accordingly, the spatiotemporal correlation calculation requires the following features from the raw feature set X_raw_{t-1} of the immediately preceding time step t-1:

[0114] The original geospatial coordinates (lon_{t-1}, lat_{t-1}) represent the previous event occurring at time t-1.

[0115] This represents the original timestamp of the previous event occurring at time t-1.

[0116] Example of a coffee shop query (time t-1): A user ended a meeting lasting several hours in "Region A" at time t-1. This information is encoded in s_{t-1}, and the raw spatiotemporal inputs associated with the generation of s_{t-1} are: x_raw_loc_{t-1} = (lon_A, lat_A) (the location where the meeting ended in Region A), x_raw_time_{t-1} = T_A (the precise time the meeting ended).

[0117] S22, Calculate the original spatiotemporal differences: Based on the above two sets of original spatiotemporal features, calculate the physical spatial differences between them (such as latitude and longitude difference, geographical distance d_geo_t) to form a spatial difference vector ΔS_t, and the physical time differences (such as the difference in timestamps ΔT_t).

[0118] Spatial difference vector The calculation (k_s is the dimension of the spatial difference feature): ΔS_t can contain multiple

[0119] Spatial difference information in dimensions, such as: longitude difference: Δlon_t = lon_t - lon_{t-1}, latitude difference: Δlat_t = lat_t - lat_{t-1}, geographic distance: d_geo_t = GeoDistance((lon_t, lat_t), (lon_{t-1}, lat_{t-1})). GeoDistance is a function that calculates the actual geographic distance between two points. For example, in the coffee shop query example, Δlon_t = lon_B - lon_A, Δlat_t = lat_B - lat_A, d_geo_t = GeoDistance((lon_B, lat_B), (lon_A, lat_A)), assuming the distance d_geo_t from region A to region B is 2.5 kilometers.

[0120] Besides Euclidean distance or Haversine distance, other distance metrics (such as Manhattan distance or network distance) can be selected depending on the application scenario. Richer kinematic features, such as direction vectors, velocities, and accelerations, can be incorporated as part of the spatial difference.

[0121] Time difference scalar The calculation is: ΔT_t = x_raw_time_t - x_raw_time_{t-1} (e.g., the time difference in seconds). In the coffee shop query example, ΔT_t = T_B - T_A. Assume the user starts walking and queries for a coffee shop (T_B) 30 minutes after the meeting ended (T_A) (ΔT_t = 1800 seconds).

[0122] S23, Differential Feature Transformation and Integration: The calculated original spatiotemporal differences ΔS_t and ΔT_t are processed, including normalization, discretization (e.g., dividing continuous geographical distances and time differences into predefined interval bins), and encoding (e.g., performing one-hot encoding on the discretized intervals), and finally aggregated into a unified, structured differential feature vector F_diff_t.

[0123] The original discrepancy measures (such as d_geo_t and ΔT_t) have different physical units and numerical ranges. To facilitate subsequent neural network processing, they need to be transformed (e.g., normalized, discretized) and integrated.

[0124] P_S(.) and P_T(.) represent a set of processing functions (normalization function N(.), discretization function D(.), and subsequent encoding function Enc(., such as one-hot encoding) for the spatial difference vector ΔS_t and the temporal difference scalar ΔT_t, respectively). Then, the unified difference feature vector F_diff_t is represented as an aggregation of these processed features:

[0125] F_diff_t=Aggregate(P_S(ΔS_t),P_T(ΔT_t))

[0126] in:

[0127] Normalization N(.): Maps d_geo_t and ΔT_t to a specific numerical range by standardizing them using min-max normalization or Z-score, respectively.

[0128] Discretization D(.): Divides the normalized continuous values ​​into several predefined intervals (bins).

[0129] Coffee shop query example (discretization): d_geo_t = 2.5 km. If the interval is defined as: bin_S1 (<0.5 km), bin_S2 (0.5-2 km), bin_S3 (2-5 km), bin_S4 (>5 km), then d_geo_t belongs to bin_S3; ΔT_t = 30 minutes. If the interval is defined as: bin_T1 (<5 min), bin_T2 (5-20 min), bin_T3 (20-60 min), bin_T4 (>60 min), then ΔT_t belongs to bin_T3.

[0130] Enc(.): Converts the discretized interval indices into a numerical vector form using one-hot encoding. For a feature with k possible intervals, its one-hot encoding is a k-dimensional vector, where the dimension corresponding to its interval is 1, and the other dimensions are 0.

[0131] Taking a coffee shop search as an example, the calculation process is as follows:

[0132] Calculate the one-hot encoding of the spatial distance d_geo_t: d_geo_t = 2.5 km, belonging to bin_S3 in the predefined four spatial distance intervals bin_S1 (<0.5 km), bin_S2 (0.5-2 km), bin_S3 (2-5 km), and bin_S4 (>5 km). Therefore, Enc(D(d_geo_t)) (i.e., the one-hot encoding of the interval bin_S3 to which d_geo_t belongs) is a 4-dimensional vector, where the third dimension is 1 and the rest are 0. The meaning of this vector Enc(D(d_geo_t)) = [0,0,1,0] is: the first dimension represents whether bin_S1 is active [no], the second dimension represents whether bin_S2 is active [no], the third dimension represents whether bin_S3 is active [yes], and the fourth dimension represents whether bin_S4 is active [no].

[0133] Calculate the one-hot encoding of the time interval ΔT_t: ΔT_t = 30 minutes, belonging to bin_T3 in the four predefined time intervals bin_T1 (<5min), bin_T2 (5-20min), bin_T3 (20-60min), and bin_T4 (>60min). Therefore, Enc(D(ΔT_t)) (i.e., the one-hot encoding of the interval bin_T3 to which ΔT_t belongs) is also a 4-dimensional vector, where the third dimension is 1 and the rest are 0. The vector Enc(D(ΔT_t)) = [0,0,0,1,0] means: the first dimension represents whether bin_T1 is active [no], the second dimension represents whether bin_T2 is active [no], the third dimension represents whether bin_T3 is active [yes], and the fourth dimension represents whether bin_T4 is active [no].

[0134] Aggregate(.): By concatenating the processed and encoded feature vectors, a final unified difference feature vector F_diff_t is formed.

[0135] Concatenation: The spatial distance one-hot encoded vector and the time interval one-hot encoded vector obtained above are concatenated: F_diff_t = Concatenate(Enc(D(d_geo_t)), Enc(D(ΔT_t))) = [0,0,1,0,0,0,1,0]. The concatenated result F_diff_t is an 8-dimensional vector. The third bit of the vector is 1, explicitly indicating that the spatial difference belongs to bin_S3 (medium-distance transition). The seventh bit of the vector (i.e., the third bit of the second half of the time encoding) is 1, explicitly indicating that the temporal difference belongs to bin_T3 (medium-length interval). Therefore, this 8-dimensional vector F_diff_t now represents the specific spatiotemporal difference pattern of "the current event, compared to the previous event, belongs to a medium-distance transition in space (bin_S3 is activated) and a medium-length interval in time (bin_T3 is activated)" in a clear and numerical way. This vector will serve as the input for subsequent MLP_SRC.

[0136] For the transformation and integration of differential features, for continuous geographical distances d_geo_t and time differences ΔT_t, in addition to discretization and binning, explicit discretization can be avoided. Instead, the values ​​can be directly normalized and input into the MLP_SRC, allowing the MLP_SRC's nonlinear capabilities to autonomously learn its influence range. Alternatively, methods such as Gaussian radial basis functions (RBF) can be used to encode continuous differential values.

[0137] Furthermore, F_diff_t, as a variable implementation method, can also be used with different combinations of the original difference features, or some cross features can be introduced.

[0138] For the structure of MLP_SRC, the number of layers, the number of neurons per layer, and the type of activation function can be adjusted according to the complexity and dimension of F_diff_t. It is even possible to consider using other more complex network structures (such as small networks with internal attention mechanisms) to replace the standard MLP in order to enhance its ability to extract abstract spatiotemporal correlation patterns from F_diff_t.

[0139] The dimension d_r of R_st can also be adjusted according to task requirements. In addition to learning abstract semantic dimensions, some dimensions of R_st can be designed to have more explicit physical or logical meanings, such as directly corresponding to "spatial proximity" or "temporal continuity".

[0140] S24, Generate the spatiotemporal correlation score vector: Input the unified difference feature vector F_diff_t into a learnable multilayer perceptron (MLP_SRC), and generate a spatiotemporal correlation score vector R_st with dimension d_r through nonlinear mapping. Different dimensions of R_st can learn to represent spatiotemporal transition patterns of different types or intensities, such as high spatiotemporal locality, significant spatiotemporal scene transitions, and extreme spatiotemporal discontinuities.

[0141] F_diff_t provides a structured description of spatiotemporal differences. To obtain a more abstract and discriminative spatiotemporal correlation representation, F_diff_t is input into a multilayer perceptron MLP_SRC with parameter θ_{SRC}.

[0142] R_{st}=MLP_{SRC}(F_diff_t; θ_{SRC})

[0143] in, d_r is the dimension of the predefined spatiotemporal correlation score vector (e.g., d_r can be an integer between 3 and 10). MLP_SRC maps the input F_diff_t to a d_r-dimensional embedding space through its internal nonlinear transformation layer. In this space, different dimensions of R_st can capture and characterize spatiotemporal correlation patterns of different types or different "semantics".

[0144] Example of a coffee shop query (R_st interpretation, assuming d_r = 3):

[0145] For F_diff_t (representing "medium spatial distance, medium time interval"), the possible output values ​​of MLP_SRC are R_st = [0.1, 0.85, 0.2]. The three dimensions of R_st may correspond to: R_st = 0.1 (low activation): possibly representing "high spatiotemporal locality / continuity" (because the current scene is not happening in place or at a very close distance, or instantaneously). R_st = 0.85 (high activation): possibly representing "significant spatiotemporal scene transition / activity phase change" (because the user has transitioned from a meeting state in area A to a walking store-finding state in area B over a certain spatiotemporal span). R_st = 0.2 (medium-low activation): possibly representing "extreme spatiotemporal discontinuity / complete context reset" (although the current scene has transitioned, it may still be related to the overall daytime plan and is not a completely random jump).

[0146] Comparative scenario: If a user queries the restroom next to area A within 1 minute after the meeting in area A ends (d_geo_t is minimal, ΔT_t is minimal), then the corresponding F_diff_t will be different, and MLP_SRC may output R_st = [0.9, 0.05, 0.02]. In this case, the high activation value of the first dimension clearly indicates "high spatiotemporal locality".

[0147] The parameters θ_{SRC} of MLP_SRC are similar to those of the feature embedding module. They are learned through backpropagation by performing end-to-end training on the method described in this invention (or a larger system incorporating this method) and optimizing the loss function L_total for a specific downstream task. This means that MLP_SRC is shaped to generate the kind of representation that best helps subsequent gating units make correct information choices, thereby improving the accuracy of the entire system's prediction of the R_st vector. If a certain representation of R_st can effectively reduce L_total, the configuration of the θ_{SRC} parameters that generate that R_st will be strengthened.

[0148] Through the above steps, the original, physical spatiotemporal differences are transformed into a fixed-dimensional numerical vector R_st that contains an abstract understanding of the current spatiotemporal transition mode. This R_st provides a crucial and specialized basis for spatiotemporal context adjustment in the gating mechanism of the subsequent dynamic state update module.

[0149] Therefore, this application, based on an explicit, physical difference-based spatiotemporal correlation calculation mechanism (SRC), can directly calculate the physical spatial differences (such as distance) and temporal differences (such as duration) between consecutive original events, and form a unified difference feature vector F_diff_t through structured processing (such as normalization, discretization, and encoding). Furthermore, this F_diff_t is transformed into a low-dimensional, dense "spatiotemporal correlation score vector R_st" through a learnable mapping network (MLP_SRC), whose different dimensions can learn to represent spatiotemporal transition patterns with different semantics.

[0150] This technical approach enables the model to directly and quantitatively perceive and understand "the extent of spatiotemporal changes between events" in the physical world, rather than through indirect guesswork. R_st provides a clear and interpretable basis for subsequent gating adjustments, making the model's response to spatiotemporal context more accurate and robust. For example, the model can clearly distinguish between two distinct spatiotemporal modes: "short-term stationary stay" and "rapid long-distance movement," and make differentiated state updates accordingly.

[0151] Step S3: Use R_st as the adjustment signal, input it into the gated recurrent neural network to adjust the gating, update the cell state, and obtain the sequence state vector s_t, which specifically includes:

[0152] S31: Input the current feature embedding vector E_t, the hidden state s_{t-1} of the previous time step, the cell state C_{t-1}, and the spatiotemporal correlation score vector R_st, where:

[0153] The current original input X_raw_t (including x_raw_loc_t, x_raw_time_t, etc.) is processed by the "learnable feature embedding module" in Unit 1 (U1) to obtain the comprehensive feature embedding vector.

[0154] For example, in the coffee shop query example, E_t encodes all the relevant semantic and contextual features of "the user is walking and has queried a coffee shop located in area B".

[0155] The hidden state vector at the previous time step t-1. In the coffee shop query example, s_{t-1} encodes the comprehensive state information that "the user has just finished a meeting that lasted for several hours, may be feeling tired, and is located in area A".

[0156] The cell state vector at the previous time step t-1 is the backbone of the information flow.

[0157] The spatiotemporal correlation score vector is calculated by SRC for the current input (x_raw_loc_t, x_raw_time_t) and the previous input (x_raw_loc_{t-1}, x_raw_time_{t-1}). For example, in the coffee shop query example, based on the aforementioned SRC calculation, the current scenario (from the end of the meeting in area A to walking to find a shop in area B, a spatial distance of 2.5km and a time interval of 30min) might yield R_st = [0.1, 0.85, 0.2]. Here, R_st (the first dimension, with a value of 0.1) represents low activation due to "high spatiotemporal locality"; R_st (the second dimension, with a value of 0.85) represents high activation due to "significant spatiotemporal scene transitions"; and R_st (the third dimension, with a value of 0.2) represents low activation due to "extreme spatiotemporal discontinuity".

[0158] S32: Input into an LSTM-like structure, integrating the spatiotemporal correlation score vector R_st into the forget gate (f_t), input gate (i_t), and candidate new information vector. In the calculation of the output gate (o_t) (using the learnable weight matrices V_f, V_i, V_c and V_o respectively), the information forgetting, new information absorption, content generation and state output are dynamically adjusted to generate the final hidden state vector s_t.

[0159] Specifically, it includes:

[0160] S321: Calculate the spatiotemporally regulated forgetting gate

[0161] The forget gate f_t determines which information in the cell state C_{t-1} from the previous time step should be retained and which should be forgotten.

[0162] f_t=σ(W_f E_t+U_f s_{t-1}+V_f R_{st}+b_f)

[0163] in, These are the weight matrix and bias vector of the standard forget gate. This is a newly added learnable weight matrix used to integrate the spatiotemporal correlation score R_st. θ_f={W_f,U_f,V_f,b_f} is the parameter set of the forget gate. σ is the sigmoid activation function.

[0164] The mechanism by which R_st affects f_t lies in the term V_f R_{st}. It linearly transforms the d_r-dimensional R_st into a d_s-dimensional vector, where each element influences the final activation value of the corresponding dimension of the forget gate. In the coffee shop query example, R_st = [0.1, 0.85, 0.2], where R_st = 0.85 indicates a "significant spatiotemporal scene transition". Each row of the weight matrix V_f, V_f[k,:] (corresponding to the k-th element of f_t), learns how to respond to different dimensions of R_st. Assume that some dimensions of C_{t-1} (e.g., dimensions j to j+m) encode information related to "specific location details in region A" or "meeting room environment". Through learning, the rows in V_f corresponding to these dimensions (i.e., V_f[j,:] to V_f[j+m,:]) may learn large negative weights in the columns corresponding to R_st (scene transition signal) (i.e., V_f[j,1] to V_f[j+m,1]).

[0165] Therefore, when R_st = 0.85, the term V_f R_{st} will contribute a significant negative value in the j to j+m dimensions. This will make the activation values ​​of f_t in these dimensions closer to 0 (because σ (a large negative number) ≈ 0).

[0166] As a result, due to the detection of significant spatiotemporal scene transitions, the forgetting gate more strongly "forgets" old information that is strongly correlated with the previous "Area A meeting," is fine-grained, and no longer applicable to the current "Area B walking shop search" scenario. Meanwhile, some more generalized information in C_{t-1}, less correlated with specific spatiotemporal contexts (such as the user's "fatigue"), may be retained if V_f is insensitive to R_st in the corresponding dimension or is positive (encouraging retention). The forgetting decision is no longer based solely on the content of E_t and s_{t-1}, but is explicitly and learnably modulated by the physical spatiotemporal change pattern represented by R_st.

[0167] S322: Calculate the spatiotemporally adjusted input gate

[0168] The input gate i_t determines the candidate new information obtained in the current computation. (See step 2.2.2.c) which parts should be incorporated into the new cell state C_t?

[0169] i_t=σ(W_i E_t+U_i s_{t-1}+V_i R_{st}+b_i)

[0170] in This is the newly added weight matrix. θ_i={W_i,U_i,V_i,b_i}.

[0171] The mechanism by which R_st influences i_t, such as in the coffee shop query example where R_st = 0.85 (scene transition signal), remains dominant. The learning objective of V_i is to enable V_i R_{st} to appropriately adjust the input of new information. For information dimensions highly relevant to the new scene "Coffee Shop in Area B" and the new activity "Walking" (this information will primarily be reflected in...),... In the context of the model, the responses of V_i to R_st in these dimensions (i.e., V_i[k,1]) may learn larger positive weights. Therefore, when a significant scene transition is detected, the input gate will be more "open" to candidate information that matches the new scene and activity. It encourages the model to quickly adapt to the new spatiotemporal context and focus its attention on the new content brought about by the current E_t.

[0172] S323: Calculate the spatiotemporally adjusted candidate new information vector

[0173] It is generated based on the current input E_t and the previous state s_{t-1}, and represents the new content that is "intended" to be added to the cell state.

[0174]

[0175] in This is the newly added weight matrix, θ_c = {W_c, U_c, V_c, b_c}. The tanh activation function compresses the output to the interval [-1, 1].

[0176] R_st pairs The mechanism of influence is illustrated in the coffee shop query example where R_st = 0.85 (scene transition signal). Learning V_c might enable V_c R_{st} to guide the process when a scene transition is detected. The generation of V_c R_{st} focuses more on reinforcing features that are "novel," "exploratory," or "significantly different from previous states" in E_t (the coffee shop in area B). For example, if E_t contains the signal "first visit to the coffee shop," V_c R_{st} might amplify this signal. The expression in the text.

[0177] S324: Update cell status

[0178] Cellular state C_t is the core carrier of the model's long-term memory. Its updates integrate the forgetting history and the process of absorbing new information.

[0179]

[0180] Where ⊙ represents the element-wise product (Hadamard product).

[0181] Operating Mechanism: This formula has the same structure as the standard LSTM, but its underlying principles have undergone profound changes. This is because f_t (forget gate), i_t (input gate), and... (Candidate new information) has been received and integrated with explicit quantitative information R_st from SRC regarding the current spatiotemporal transition mode through its respective V_f, Vi, V_c weight matrices. In the coffee shop query example, the fine-grained information related to "Region A Meeting" in C_{t-1} is significantly weakened due to the strong forgetting effect of f_t (affected by R_st = 0.85). Meanwhile, New information related to "Coffee Shop in Area B" and "Walking" is effectively written into C_t due to the active absorption of i_t (also influenced by R_st = 0.85). The new cell state C_t is thus able to update memories more intelligently and adaptably to the current spatiotemporal context. It clears old memories that are no longer relevant due to significant scene changes and efficiently builds memories about the new scene.

[0182] S325: Calculate the time-space regulated output gate

[0183] The output gate o_t controls which information in the cell state C_t should be "read" out and used as the hidden state s_t (i.e., the output of U1) at the current time step.

[0184] o_t=σ(W_o E_t+U_o s_{t-1}+V_o R_{st}+b_o)

[0185] in It is a newly added weight matrix, θ_o={W_o,U_o,V_o,b_o}.

[0186] The mechanism by which R_st affects o_t: R_st can adjust which information is suitable as the "instantaneous state summary" output for the current moment. In the coffee shop query example, R_st = 0.85 (scene transition signal). The learning of V_o may cause o_t to be more inclined to output those state features in C_t that best represent "the user has entered a new scene and started a new activity" when a scene transition is detected, such as dimensions related to "exploring area B" and "having a need for caffeine / leisure". While more generalized historical information that is retained in C_t but is less relevant to the current instantaneous state (such as "the user is generally tired"), its output may be appropriately suppressed by o_t.

[0187] S326: Calculate the current hidden state

[0188] s_t is the final output of Unit 1 (U1) at time step t. It will be used as s_{t-1} for the next time step t+1 and input into Unit 2 (U2) for subsequent pattern mining.

[0189] s_t=o_t⊙tanh(C_t)

[0190] Characteristics of s_t: Due to the gating of the entire information flow (f_t, i_t, o_t) and the generation of candidate new information. All are modulated by the explicitly quantified spatiotemporal correlation score vector R_st calculated by SRC, and the resulting hidden state s_t is no longer merely a memory update result based on content and simple temporal sequence. It is a deep state representation with a stronger ability to perceive and adapt to spatiotemporal dynamic evolution patterns.

[0191] In the coffee shop query example, s_t would be a vector that clearly reflects the complex dynamic of "the user is currently focused on the coffee shop after a significant scene transition from a meeting in area A to walking in area B." It has effectively "forgotten" irrelevant details of area A and "focused" on new activities in area B, while possibly retaining some persistent states across scenes (such as fatigue, which the model needs to retain if it learns this information).

[0192] Thus, the efficient generation of sequence state vectors that are highly sensitive to spatiotemporal dynamics from heterogeneous raw inputs has been achieved.

[0193] The spatiotemporal correlation score vector R_st obtained by SRC calculation is uniquely introduced as a "meta-control signal" and directly introduced into the forget gate, input gate, output gate and candidate cell state calculation formula of RNN (such as LSTM) (for example, through a learnable linear transformation term of the form V_x R_{st}, where V_x is the corresponding weight matrix).

[0194] This demonstrates that the direct, learnable adjustment of the core gating of the recurrent neural network through the spatiotemporal correlation score vector R_st endows the model with an unprecedented ability to finely regulate its internal information flow based on physical spatiotemporal changes. The model can dynamically and on demand decide how much history to forget, how much current information to absorb, and what to output as a summary of the current state, based on the "severity" or "pattern type" of the current spatiotemporal transition indicated by R_st. This makes the state update process highly adaptive and context-sensitive, enabling it to more intelligently handle non-stationarity and abrupt changes in the spatiotemporal sequence, thereby generating higher-quality state representations. For example, when R_st indicates a "significant spatiotemporal scene transition," the forget gate can be strongly activated to clear old scene memories, while the input gate focuses more on absorbing new scene information.

[0195] Based on the above embodiments, this application also provides a sequence data modeling system, including: a data embedding processing unit, a spatiotemporal correlation calculation unit, and an update unit, wherein:

[0196] The data embedding processing unit is used to convert the multi-source heterogeneous original feature set at the current time into a uniform-dimensional embedding vector through a learnable mapping function, and integrate it into a comprehensive feature embedding vector E_t;

[0197] The spatiotemporal correlation calculation unit is used to calculate the spatiotemporal difference between the current event and the preceding event, process it into structured features, and then input it into a learnable network to generate a spatiotemporal correlation score vector R_st;

[0198] The update unit uses R_st as an adjustment signal, inputs it into the gated recurrent neural network to adjust the gating, update the cell state, and obtain the sequence state vector s_t.

[0199] This application embodiment also provides a storage medium storing a computer program, which is executed by a processor to perform the sequence data modeling method as described above.

[0200] This application also provides an electronic device, including: a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the sequence data modeling method as described above.

Claims

1. A method for modeling sequence data, characterized in that, Includes the following steps: Step S1: Convert the multi-source heterogeneous original feature set at the current time into a uniform-dimensional embedding vector through a learnable mapping function, and integrate it into a comprehensive feature embedding vector E_t, specifically including: S11: POI type feature embedding; S12: Geographic coordinate feature embedding; S13: Embedding of user activity state features; S14: Concatenate the above features to form a comprehensive feature embedding vector E_t; Step S2: Calculate the spatiotemporal difference between the current event and its predecessor events, process it into a structured feature, and input it into a learnable network to generate a spatiotemporal correlation score vector R_st. Specifically, this includes: S21: Obtain original spatiotemporal features: Extract the original spatiotemporal features of the current time step t and the original spatiotemporal features of the immediately preceding time step t-1; S22: Calculate the original spatiotemporal differences: Based on the above two sets of original spatiotemporal characteristics, calculate the physical spatial differences between them to form a spatial difference vector ΔS_t, and a physical time difference ΔT_t; S23: Difference feature transformation and integration: The calculated original spatiotemporal differences ΔS_t and ΔT_t are processed and finally aggregated into a unified, structured difference feature vector F_diff_t; S24: Generate a spatiotemporal correlation score vector: Input the unified difference feature vector F_diff_t into a learnable multilayer perceptron, and generate a spatiotemporal correlation score vector R_st with dimension d_r through nonlinear mapping; Step S3: Use R_st as an adjustment signal, input it into the gated recurrent neural network to adjust the gating, update the cell state, obtain the sequence state vector s_t, and apply it to downstream tasks that depend on spatiotemporal information. These downstream tasks include user future behavior prediction, personalized recommendation, and location-aware services, specifically including: S31: Input the current feature embedding vector E_t, the hidden state s_{t-1} of the previous time step, the cell state C_{t-1}, and the spatiotemporal correlation score vector R_st; S32: Input into an LSTM-like structure, integrating the spatiotemporal correlation score vector R_st into the forget gate f_t, input gate i_t, and candidate new information vector. In the calculation of _t and output gate o_t, the information forgetting, new information absorption, content generation, and state output are dynamically adjusted to generate the final hidden state vector s_t, which specifically includes: S321: Calculate the spatiotemporally adjusted forgetting gate f_t, the formula for which is: f_t =σ(W_f E_t + U_f s_{t-1} + V_f R_{st} + b_f) S322: Calculate the spatiotemporally adjusted input gate i_t, the formula for which is: i_t = σ(W_i E_t + U_i s_{t-1} + V_i R_{st} + b_i) S323: Calculate the spatiotemporally adjusted candidate new information vector _t, its calculation formula is: _t = tanh(W_c E_t + U_c s_{t-1} + V_c R_{st} + b_c) S324: Update the cell state C_t, which is calculated using the following formula: C_t = f_t ⊙ C_{t-1} + i_t ⊙ _t S325: Calculate the spatiotemporally adjusted output gate o_t, the calculation formula is as follows: o_t = σ(W_o E_t + U_o s_{t-1} + V_o R_{st} + b_o) S326: Calculate the current hidden state s_t, the formula for which is: s_t = o_t ⊙ tanh(C_t).

2. The sequence data modeling method as described in claim 1, characterized in that, The gated recurrent neural network is either LSTM or GRU.

3. The sequence data modeling method as described in claim 1, characterized in that, Step S12: In the geographic coordinate feature embedding, the embedding method is any one of MLP embedding, Geohash encoding and embedding, combining geographic knowledge graph information, or node embedding based on graph neural network.

4. The sequence data modeling method as described in claim 1, characterized in that, The concatenation method in step S14 is: vector concatenation or weighted summation, or dynamic fusion through a small attention network, or dimensionality reduction or feature crossing of the concatenated vectors through an MLP.

5. A sequence data modeling system, applied to the sequence data modeling method according to any one of claims 1-4, characterized in that, include: The data embedding processing unit, the spatiotemporal correlation calculation unit, and the update unit are as follows: The data embedding processing unit is used to convert the multi-source heterogeneous original feature set at the current time into a uniform-dimensional embedding vector through a learnable mapping function, and integrate it into a comprehensive feature embedding vector E_t; The spatiotemporal correlation calculation unit is used to calculate the spatiotemporal difference between the current event and the preceding event, process it into structured features, and then input it into a learnable network to generate a spatiotemporal correlation score vector R_st; The update unit uses R_st as an adjustment signal, inputs it into the gated recurrent neural network to adjust the gating, update the cell state, and obtain the sequence state vector s_t.

6. An electronic device, characterized in that it comprises: The device includes a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is in operation, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the sequence data modeling method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Multi-scale space-time fusion image feature extraction method based on traffic flow

    CN120336804A

  • Method for realizing a multi-channel convolutional recurrent neural network EEG emotion recognition model using transfer learning

    US20230039900A1