Data processing method and data processing apparatus
Patent Information
- Application Number
- CN202610967732.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-18
AI Technical Summary
但在部分情况下,网约车业务的供需数据的预测准确性不高,因此可能会导致部分用户长时间无法匹配到合适的网约车司机,进而降低用户的使用体验
[0009] This invention, after obtaining the topological adjacency matrix, semantic adjacency matrix, and supply and demand time-series vectors of multiple geographical regions over past time periods, uses this data and a pre-trained data prediction model to obtain the supply and demand time-series prediction vectors for each geographical region in future time periods. In this invention, the topological adjacency matrix is determined based on the geographical location relationships of each geographical region, the semantic adjacency matrix is determined based on the similarity of the semantic vectors of each geographical region, and the data prediction model includes multiple semantic spatiotemporal modules. Therefore, this invention enhances the feature representation capability of multiple geographical regions by leveraging their geographical location relationships and semantic similarity, and improves the accuracy of supply and demand data prediction by repeatedly fusing the spatial correlation patterns and long-term time-series patterns of multiple geographical regions through semantic spatiotemporal modules.
Smart Images

Figure CN122779906A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and more specifically to a data processing method and a data processing apparatus. Background Technology
[0002] With the continuous development of internet technology, supply and demand data forecasting has become increasingly important in many fields. Taking the ride-hailing industry as an example, user demand for ride-hailing services is often volatile and random due to factors such as weather, holidays, and traffic conditions. Therefore, to ensure that users' travel needs are met, ride-hailing platforms forecast the supply and demand data for their services. However, in some cases, the accuracy of these forecasts is not high, which may result in some users being unable to find suitable drivers for extended periods, thus reducing their user experience. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a data processing method and a data processing apparatus to improve the feature representation capability of multiple geographical regions by leveraging the geographical location relationships and semantic similarity of multiple geographical regions, and to improve the accuracy of supply and demand data prediction by repeatedly fusing the spatial correlation patterns and long-term time series patterns of multiple geographical regions through a semantic spatiotemporal module.
[0004] In a first aspect, embodiments of the present invention provide a data processing method, the method comprising: Obtain supply and demand time series vectors for multiple geographical regions over a predetermined time period, a topological adjacency matrix for each geographical region, and a semantic adjacency matrix. The topological adjacency matrix is determined based on the geographical location relationships of each geographical region, and the semantic adjacency matrix is determined based on the similarity of the semantic vectors of each geographical region. The semantic vectors are determined based on the supply and demand statistics of the corresponding geographical regions. Based on the supply and demand time series vectors, the topological adjacency matrix, and the semantic adjacency matrix, and using a pre-trained data prediction model, supply and demand time series prediction vectors for each geographical region in the target time period are obtained. The data prediction model includes multiple semantic spatiotemporal modules.
[0005] Secondly, embodiments of the present invention provide a data processing apparatus, the apparatus comprising: The data acquisition unit is used to acquire supply and demand time series vectors of multiple geographical regions over a predetermined time period, topological adjacency matrices and semantic adjacency matrices of each geographical region. The topological adjacency matrix is determined based on the geographical location relationships of each geographical region, and the semantic adjacency matrix is determined based on the similarity of the semantic vectors of each geographical region. The semantic vectors are determined based on the supply and demand statistics of the corresponding geographical region. The data prediction unit is used to obtain the supply and demand time series prediction vectors of each geographical region in the target time period based on the supply and demand time series vectors, the topological adjacency matrix and the semantic adjacency matrix, and a pre-trained data prediction model. The data prediction model includes multiple semantic spatiotemporal modules.
[0006] Thirdly, embodiments of the present invention provide a computer-readable storage medium storing computer program instructions thereon, which, when executed by a processor, implement the method described in the first aspect.
[0007] Fourthly, embodiments of the present invention provide an electronic device, the device comprising: Memory is used to store one or more computer program instructions; A processor, wherein the one or more computer program instructions are executed by the processor to implement the method as described in the first aspect.
[0008] Fifthly, embodiments of the present invention provide a computer program product that, when run on a computer, causes the computer to perform the method described in the first aspect.
[0009] This invention, after obtaining the topological adjacency matrix, semantic adjacency matrix, and supply and demand time-series vectors of multiple geographical regions over past time periods, uses this data and a pre-trained data prediction model to obtain the supply and demand time-series prediction vectors for each geographical region in future time periods. In this invention, the topological adjacency matrix is determined based on the geographical location relationships of each geographical region, the semantic adjacency matrix is determined based on the similarity of the semantic vectors of each geographical region, and the data prediction model includes multiple semantic spatiotemporal modules. Therefore, this invention enhances the feature representation capability of multiple geographical regions by leveraging their geographical location relationships and semantic similarity, and improves the accuracy of supply and demand data prediction by repeatedly fusing the spatial correlation patterns and long-term time-series patterns of multiple geographical regions through semantic spatiotemporal modules. Attached Figure Description
[0010] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which: Figure 1 This is a flowchart of the data processing method according to an embodiment of the present invention; Figure 2 This is a flowchart of the data processing method according to an embodiment of the present invention; Figure 3 This is a flowchart of the data processing method according to an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the process of generating an adjacency matrix of multiple geographical regions in an embodiment of the present invention; Figure 5 This is a flowchart of the data processing method according to an embodiment of the present invention; Figure 6 This is a flowchart of the data processing method according to an embodiment of the present invention; Figure 7 This is a schematic diagram of the spatiotemporal feature encoding module according to an embodiment of the present invention; Figure 8 This is a flowchart of the data processing method according to an embodiment of the present invention; Figure 9 This is a schematic diagram of the feature convolution fusion module according to an embodiment of the present invention; Figure 10 This is a flowchart of the data processing method according to an embodiment of the present invention; Figure 11 This is a schematic diagram of the structure of the gate control correction module according to an embodiment of the present invention; Figure 12 This is a flowchart of the data processing method according to an embodiment of the present invention; Figure 13 This is a schematic diagram of the output module according to an embodiment of the present invention; Figure 14 This is a flowchart of the data processing method according to an embodiment of the present invention; Figure 15 This is a data flow diagram of the data processing method according to an embodiment of the present invention; Figure 16 This is a schematic diagram of a data processing device according to an embodiment of the present invention; Figure 17 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0011] The present application is described below based on embodiments, but it is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without these details. To avoid obscuring the substance of the present application, well-known methods, processes, flows, elements, and circuits are not described in detail.
[0012] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.
[0013] Unless the context explicitly requires it, words such as "including" or "contains" throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to".
[0014] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0015] The solutions described in this specification and embodiments, if involving the processing of personal information, will be processed only on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be processed within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.
[0016] Existing supply and demand data forecasting methods mainly employ the following approaches: The first method is the traditional statistical method, which predicts supply and demand data based on past supply and demand data for a geographical region, such as average supply and demand statistics, long-term statistical characteristics, and regional profiles. The second method is the deep learning method, which uses deep learning models such as recurrent neural networks, convolutional neural networks, and transformers to mine the temporal evolution patterns of supply and demand data, thereby predicting future supply and demand data. However, these methods primarily focus on changes in supply and demand data over time, neglecting the locational correlations between different geographical regions and the mutual influence of supply and demand data, thus lacking accuracy. The third method is based on graph neural networks for supply and demand data forecasting. However, this method can only characterize the locational correlations between adjacent geographical regions, similarly neglecting the mutual influence of supply and demand data between geographical regions, and therefore also lacking accuracy.
[0017] To address the aforementioned technical problems, this invention proposes a data processing method and a data processing device to enhance the feature representation capability of multiple geographical regions by leveraging their geographical location relationships and semantic similarity. Furthermore, it utilizes a semantic spatiotemporal module to repeatedly fuse the spatial correlation patterns and long-term time-series patterns of multiple geographical regions, thereby improving the accuracy of supply and demand data prediction.
[0018] This invention primarily describes the ride-hailing application scenario as an example. It is readily understood that this embodiment is not limited thereto; any application scenario that can support the corresponding function, or that will be able to support the corresponding function as technology develops in the future, is within the protection scope of this invention.
[0019] The following describes the method through specific examples. Figure 1 This is a flowchart of a data processing method according to an embodiment of the present invention. Figure 1 As shown, the method in this embodiment may include the following steps: Step S100: Obtain the supply and demand time series vectors of multiple geographical regions during a predetermined time period, the topological adjacency matrix of each geographical region, and the semantic adjacency matrix.
[0020] In this embodiment, a supply and demand time series vector can be determined based on the supply and demand time series data of each geographical region within a predetermined time period, so as to reflect the volatility of supply and demand data through the supply and demand time series vector. A semantic adjacency matrix of each geographical region can be determined based on the supply and demand statistics of each geographical region, so as to reflect the long-term temporal correlation of supply and demand data through the semantic adjacency matrix. At the same time, the spatial correlation between each geographical region can be reflected through the topological adjacency matrix of each geographical region.
[0021] The supply and demand time-series data for each geographical region during the predetermined time period may include at least one of the following: the number of requests initiated, the number of requests responded, and the request response rate within multiple consecutive sampling intervals (i.e., time intervals formed by adjacent time steps) of each geographical region during the predetermined time period. Specifically, the number of requests initiated can be the number of ride-hailing orders submitted by users within the predetermined time period, reflecting users' travel demand; the number of requests responded can be the number of ride-hailing orders accepted and responded to within the predetermined time period, reflecting the supply of the ride-hailing platform; and the request response rate can be the ratio of the number of requests responded to to the number of requests initiated within the predetermined time period, reflecting the ride-hailing platform's capacity to meet travel demand. Optionally, the supply and demand data may also include other relevant data, such as the number of drivers online and the duration of drivers online, which is not limited in this embodiment.
[0022] In some embodiments, the predetermined time period may include one or more time intervals. For example, if the predetermined time period is from January 1 to June 30, then the predetermined time period may include the time interval from January 1 to June 30, or it may include multiple time intervals such as January 1 to January 31, February 1 to February 28, ..., June 1 to June 30.
[0023] Furthermore, when the predetermined time period includes multiple time intervals, the period type of each time interval can be different. For example, it can include time intervals with a daily period type, time intervals with a weekly period type, and time intervals with a monthly period type. Taking a time interval with a daily period type as an example, a daily period means that the supply and demand data generated in each sampling interval within a day is treated as a single supply and demand time series data. For example, if the predetermined time period is from January 1st to June 30th, then the predetermined time period can include time intervals with a daily period type, such as January 1st, January 2nd, ..., June 30th, or it can include time intervals with a monthly period type, such as January, February, ..., June.
[0024] In some embodiments, to improve the ability of supply and demand time series vectors to represent the supply and demand situation of a geographic region, environmental data of each predetermined geographic region during a predetermined time period can be obtained, and then the supply and demand time series vector can be determined based on the supply and demand time series data and the environmental data. Optionally, the environmental data may include at least one of the following: time period type (e.g., whether it is a morning peak or evening peak), weekday type (e.g., Monday or Tuesday), holiday type (e.g., whether it is a holiday), and weather type (e.g., sunny, windy, rainy, etc.).
[0025] Figure 2 This is a flowchart of a data processing method according to an embodiment of the present invention. Figure 2 As shown, in some embodiments, the supply and demand time series vectors can be determined through the following steps: S100A determines relative and absolute numerical data based on the supply and demand time series data of each geographical region within a predetermined time period.
[0026] In some embodiments, supply and demand time-series data can be divided into two types: relative numerical data, such as request response rate, and absolute numerical data, such as the number of requests initiated. Depending on the data type, this embodiment can employ different processing methods for data processing.
[0027] Step S200A: Perform logarithmic processing on the absolute numerical data to obtain the first type of time series data.
[0028] For absolute numerical data, to mitigate the negative impact of extreme values, such as an abnormal increase in the number of requests due to extreme weather, on the accuracy of the data prediction model, this step can perform logarithmic processing on the absolute numerical data to obtain the first type of time series data. Optionally, the first type of time series data x' can be represented by the following formula:
[0029] Where x represents the original absolute numerical data.
[0030] Step S300A: Determine the second type of time series data based on the relative numerical data.
[0031] For relatively numerical data, their distribution is usually relatively concentrated, so this step can determine the second type of time series data in various ways. In some embodiments, the relatively numerical data can be directly identified as the second type of time series data. In some embodiments, the relatively numerical data can also be normalized or standardized to further reduce the differences between the relatively numerical data.
[0032] Step S400A: Determine the supply and demand time series vectors based on the first type of time series data and the second type of time series data.
[0033] In this step, the supply and demand time series vectors for each geographic region can be determined based on the first and second type time series data, or the corresponding environmental data.
[0034] In step S100, the supply and demand statistics for each geographical region may include at least one of the statistical values of supply and demand data such as order price, demand quantity, and transaction amount for the corresponding geographical region. Further, the statistical values may include at least one of the following: average, sum, mode, minimum, and maximum. Optionally, the supply and demand statistics may also include other data, such as statistical values of task transaction volume, etc., which are not limited in this embodiment.
[0035] In some embodiments, supply and demand statistics for each geographic region can be obtained according to at least one of the following: week type, time period type, business cycle type, etc. Taking the week type as an example, supply and demand data for the same week number, such as Monday, can be obtained for each geographic region over multiple weeks, and the statistical values of each supply and demand data corresponding to the same week number can be calculated as supply and demand statistics.
[0036] This embodiment can determine the semantic adjacency matrix of multiple geographic regions in various ways. In some embodiments, the supply and demand statistics of each geographic region can be directly determined as the semantic vector of each geographic region, and the similarity of the semantic vectors of each geographic region can be calculated. Then, if the similarity between the semantic vectors of any geographic region and another geographic region meets a preset similarity condition, such as the cosine value being greater than or equal to a preset similarity threshold, the weights of the two geographic regions are determined based on the similarity. If the preset similarity condition is not met, the weights of the two similarities are set to 0, thereby obtaining the semantic adjacency matrix corresponding to each geographic region.
[0037] In some embodiments, the semantic adjacency matrix corresponding to each geographic region can also be determined based on the importance of each feature dimension in the semantic statistics. Figure 3 This is a flowchart of a data processing method according to an embodiment of the present invention. Figure 3 As shown, in some embodiments, the semantic adjacency matrix can be determined in the following way: Step S100B: Determine the semantic vector of the corresponding geographical region under each feature dimension based on supply and demand statistics.
[0038] In practical applications, the interaction between supply and demand data in different geographical regions may vary under different feature dimensions. Therefore, in this step, for any geographical region, the semantic vector of that geographical region under each feature dimension can be determined based on the supply and demand statistics of that geographical region.
[0039] For example, when determining the semantic vector of geographic region A1, the semantic vector of geographic region A1 in the order price dimension can be determined based on the average order price, total order price, minimum order price, and maximum order price of geographic region A1.
[0040] Optionally, the supply and demand statistics of each geographical region can be normalized, and the semantic vector of each geographical region under each feature dimension can be determined based on the normalized supply and demand statistics, so as to reduce the impact of the difference in dimensions on the similarity calculation.
[0041] Step S200B: Determine the similarity of each geographical region under the same feature dimension based on the semantic vectors under the same feature dimension.
[0042] In this step, semantic vectors of any two geographical regions under the same feature dimension can be obtained, and the similarity between these two semantic vectors can be calculated to obtain the similarity between the geographical regions under the same feature dimension. Optionally, this embodiment can use various similarity calculation methods, such as cosine similarity, Euclidean distance, Jaccard similarity, etc.
[0043] Step S300B: Determine the semantic similarity of each geographical region based on the similarity under each feature dimension.
[0044] In this step, the weighted sum of the similarities of any two geographic regions under each feature dimension can be calculated as the semantic similarity between the two geographic regions, thereby determining the semantic similarity of each geographic region.
[0045] The weights corresponding to each feature dimension are used to characterize the importance of that feature dimension and can be determined in various ways. In some embodiments, they can be preset according to actual needs, for example, setting the weight corresponding to the order price to 0.2, the weight corresponding to the transaction amount to 0.3, and the weight corresponding to the demand quantity to 0.5. In some embodiments, the weights corresponding to each feature dimension can be dynamically generated through dynamic weight networks, large language models, etc.
[0046] For example, the similarity between geographic region A1 and geographic region A2 is 0.8 in the order price dimension, 0.85 in the transaction amount dimension, and 0.75 in the demand dimension. The weights corresponding to order price, transaction amount, and demand are 0.2, 0.3, and 0.5 respectively. Therefore, the semantic similarity between geographic region A and geographic region A2 can be determined to be 0.79.
[0047] Step S400B: Determine the semantic adjacency matrix based on the semantic similarity of each semantic.
[0048] In this step, the semantic neighbors of each geographic region can be determined based on the semantic similarity of each geographic region, and then the semantic adjacency matrix of each geographic region can be determined based on the semantic similarity of the semantic neighbors corresponding to each geographic region.
[0049] In some embodiments, for any geographic region, it can be determined whether the semantic similarity between the geographic region and other geographic regions meets a preset similarity condition. If the semantic similarity with any other geographic region meets the preset similarity condition, the other geographic region can be identified as a semantic neighbor of the geographic region. Then, after determining the semantic neighbors of each geographic region, the corresponding weights are determined based on the semantic similarity between each geographic region and its semantic neighbors, and the weights with other geographic regions are set to 0, thereby obtaining a semantic adjacency matrix corresponding to multiple geographic regions.
[0050] Depending on the similarity calculation method, this embodiment can use different preset similarity conditions. For example, if the similarity calculation method is Euclidean distance, the preset similarity condition can be that the Euclidean distance is less than (or less than or equal to) a first distance threshold, or it can be that the Euclidean distance is ranked in the first preset number of positions with the smallest value.
[0051] For example, multiple geographical regions are included, from region A1 to region A6. The semantic similarity between region A1 and region A2 is 0.72, with region A3 it is 0.93, with region A4 it is 0.45, with region A5 it is 0.57, and with region A6 it is 0.68. The semantic similarity between region A2 and region A3 is 0.51, with region A4 it is 0.86, with region A5 it is 0.91, and with region A6 it is 0.72. The semantic similarity between region A3 and region A4 is 0.41, with region A5 it is 0.88, and with region A6 it is 0.79. The semantic similarity between region A4 and region A5 is 0.74, and with region A6 it is 0.63. The semantic similarity between region A5 and region A6 is 0.33. Therefore, we can determine that the semantic neighbors of region A1 are regions A2 and A3, the semantic neighbors of region A2 are regions A4 and A5, the semantic neighbors of region A3 are regions A1 and A5, the semantic neighbors of region A4 are regions A2 and A5, the semantic neighbors of region A5 are all regions A2 and A3, and the semantic neighbors of region A6 are all regions A2 and A3. Thus, based on the semantic similarity between each geographic region and its semantic neighbors, we can obtain the semantic adjacency matrix for regions A1 to A6 as follows:
[0052] Figure 4 This is a schematic diagram illustrating the process of generating an adjacency matrix of multiple geographical regions in an embodiment of the present invention. For example... Figure 4 As shown, after obtaining the semantic vectors 41 (demand dimension), 42 (order price dimension), and 43 (transaction amount dimension) for each geographic region, the similarity of semantic vector 41 of any geographic region with the semantic vectors 41 of other geographic regions, the similarity of semantic vector 42 of any geographic region with the semantic vectors 42 of other geographic regions, and the similarity of semantic vector 43 of any geographic region with the semantic vectors 43 of other geographic regions can be calculated to obtain the similarity under each feature dimension. The similarity is then weighted to obtain the semantic similarity of each geographic region. Based on the semantic similarity of each geographic region with other geographic regions, a top-K (ranked in the top k positions) semantic neighbor selection is performed to obtain the semantic neighbors corresponding to each geographic region. For example, if the geographic region is region 44, then the top two geographic regions with the highest semantic similarity to region 44, i.e., top1 and top2 geographic regions, can be identified as the semantic neighbors of region 44. Therefore, an adjacency matrix can be generated based on the semantic similarity between each geographic region and its corresponding semantic neighbors.
[0053] Optionally, if the semantic similarity between geographical regions is greater than 1, the semantic similarity can be normalized, and the semantic adjacency matrix of each geographical region can be determined based on the normalized semantic similarity.
[0054] In step S100, the topological adjacency matrix of each geographical region can be determined in a variety of ways. Figure 5 This is a flowchart of a data processing method according to an embodiment of the present invention. Figure 5 As shown, in some embodiments, the topological adjacency matrix can be determined by the following steps: Step S100C: Obtain at least one of the following for each geographical region: boundary adjacency, distance relationship, road network connectivity, and spatial adjacency.
[0055] In some embodiments, the geographical location relationships of different geographical regions can be represented in multiple ways. Optionally, the geographical location relationships of different geographical regions may include at least one of the following: boundary adjacency relationships, distance relationships, road network connectivity relationships, and spatial adjacency relationships.
[0056] Among them, the boundary adjacency relationship is used to characterize whether the boundaries of the corresponding geographical areas are adjacent, that is, whether the corresponding geographical areas are directly adjacent; the distance relationship is used to characterize the distance between the corresponding geographical areas, such as whether the straight-line distance between the center points meets the preset distance conditions, which can be whether the distance between the corresponding geographical areas is less than (or less than or equal to) a second distance threshold; the road network connectivity relationship is used to characterize whether there are passable paths between the corresponding geographical areas; and the spatial adjacency relationship is used to characterize whether the geographical spaces of the corresponding geographical areas are adjacent.
[0057] Step S200C: Determine the topological adjacency matrix based on at least one of the following: boundary adjacency relationship, distance relationship, road network connectivity relationship, and spatial adjacency relationship.
[0058] Depending on the way geographic location relationships are represented, this step can determine the topological adjacency matrix in various ways. In some embodiments, when the geographic location relationship is a boundary adjacency relationship, if the boundary adjacency relationship between any geographic region and other geographic regions indicates that the geographic region is adjacent to the boundary of other geographic regions, then the corresponding weight can be determined to be 1; otherwise, it can be determined to be 0.
[0059] In some embodiments, when the geographic location relationship is a distance relationship, if the distance relationship between any geographic region and other geographic regions indicates that the distance between the geographic region and other geographic regions meets a preset distance condition, then the corresponding weight can be determined as the corresponding distance (or the reciprocal of the distance); otherwise, it is determined as 0.
[0060] In some embodiments, when the geographical location relationship is a road network connectivity relationship, if the road network connectivity relationship between any geographical region and other geographical regions indicates that there is a passable path between the geographical region and other geographical regions, optionally, the corresponding weight can be determined to be 1, or the corresponding weight can be determined to be the path length (or the reciprocal of the path length), otherwise it can be determined to be 0.
[0061] In some embodiments, when the geographic location relationship is a spatial adjacency relationship, if the spatial adjacency relationship between any geographic region and other geographic regions indicates that the geographic space to which the geographic region belongs is adjacent to the boundary of the geographic space of other geographic regions, then the corresponding weight can be determined to be 1; otherwise, it is determined to be 0.
[0062] In some embodiments, the weights of each type of geographic location relationship can be determined based on their importance, and the weighted sum of each type of geographic location relationship can be calculated as the weight of each geographic region in the topological adjacency matrix. The method for setting the weights can refer to the method for setting the weights corresponding to each feature dimension in the semantic statistics data, and will not be repeated here.
[0063] Step S200: Based on each supply and demand time series vector, topological adjacency matrix, and semantic adjacency matrix, and using a pre-trained data prediction model, obtain the supply and demand time series prediction vector for each geographical region in the target time period.
[0064] In this step, the data of the data prediction model can be determined based on the supply and demand time series vectors of each geographic region, the corresponding topological adjacency matrix and semantic adjacency matrix, so as to obtain the supply and demand prediction data of each geographic region in the future one or more time steps, such as the time series vector formed by the number of requests initiated (that is, the supply and demand time series prediction vector).
[0065] In this embodiment, the data prediction model may include multiple semantic spatiotemporal modules. The semantic spatiotemporal modules are used to learn the feature similarities and differences between different geographical regions by fusing supply and demand data, semantic correlations and geographical location correlations of geographical regions. Therefore, feature representation capabilities can be improved, thereby improving the accuracy of supply and demand data prediction.
[0066] The data prediction model can be trained based on a training sample set. Each training sample in the training sample set can include a first supply and demand time-series vector, a sample topological adjacency matrix, a sample semantic adjacency matrix, and a second supply and demand time-series vector for each sample geographical region in the second sample time period. In this embodiment, the first supply and demand time-series vector, the sample topological adjacency matrix, and the sample semantic adjacency matrix can be obtained in the same way as the supply and demand time-series vector, topological adjacency matrix, and semantic adjacency matrix, which will not be elaborated here. When training the data prediction model, the input of the data prediction model can be determined based on the first supply and demand time-series vector, the sample topological adjacency matrix, and the sample semantic adjacency matrix. The corresponding second supply and demand time-series vector is used as the training target to iterate the data prediction model multiple times until the training termination condition is met, such as the convergence of the loss function of the data prediction model or the accuracy of the data prediction model reaching a preset accuracy threshold. The data processing process of the data prediction model is described in detail below.
[0067] In some embodiments, each semantic spatiotemporal module of the data prediction model may include a spatiotemporal feature encoding module, a feature fusion convolution module, a gating correction module, and an output module. The data processing procedures of each semantic spatiotemporal module are described below. Figure 6 This is a flowchart of a data processing method according to an embodiment of the present invention. Figure 6 As shown, in some embodiments, step S200 may include the following steps: Step S210: Determine each supply and demand time series vector as the input vector of the semantic spatiotemporal module that is ranked first.
[0068] In this step, the supply and demand time series vectors of each geographical region can be determined as the input vectors of the semantic spatiotemporal module ranked first, so that each supply and demand time series vector can be processed sequentially by multiple semantic spatiotemporal modules.
[0069] Step S220: Input each input vector into the spatiotemporal feature encoding module of the current semantic spatiotemporal module in an iterative manner to obtain the corresponding spatiotemporal latent state representation vector.
[0070] In this step, each semantic spatiotemporal module can be determined as the current semantic spatiotemporal module according to the order of their arrangement. The supply and demand time series vectors of each geographical region are simultaneously input into the spatiotemporal feature encoding module of the current semantic spatiotemporal module to obtain the spatiotemporal latent state representation vectors corresponding to each geographical region. This removes noise and redundant information from the supply and demand time series vectors and retains information that can represent the core evolutionary laws.
[0071] Figure 7 This is a schematic diagram of the spatiotemporal feature encoding module according to an embodiment of the present invention. Figure 7 As shown, in some embodiments, the spatiotemporal feature encoding module 70 includes a linear projection layer 71 and a spatial attention encoding module 72. In this step, each input vector can be used as the input to the linear projection layer 71, and the output of the linear projection layer 71 can be used as the input to the spatial attention encoding module 72 to obtain the corresponding spatiotemporal latent state representation vector.
[0072] Figure 8 This is a flowchart of a data processing method according to an embodiment of the present invention. Figure 8 As shown, in some embodiments, step S220 may include the following steps: Step S221: Input each input vector into the linear projection layer to obtain the corresponding hidden state vector.
[0073] In this step, each supply and demand time series vector can be input into a linear projection layer to map each supply and demand time series vector to a feature space of the same dimension, and the obtained linear projection result is used as a hidden state vector.
[0074] Step S222: Input each hidden state vector into the spatial attention encoding module to obtain the corresponding spatiotemporal hidden state representation vector.
[0075] In this step, the hidden state vectors can be input into the spatial attention encoding module to dynamically capture the importance relationships of different geographical regions at different time steps, so that the data prediction model can obtain dynamic spatial dependency information in addition to the fixed graph structure.
[0076] Step S230: Input the spatiotemporal latent state representation vectors, topological adjacency matrices and semantic adjacency matrices into the feature convolution fusion module of the current semantic spatiotemporal module to obtain the corresponding fused spatiotemporal representation vectors.
[0077] In this step, the spatiotemporal latent state representation vectors, topological adjacency matrices, and semantic adjacency matrices can be simultaneously input into the feature convolution fusion module of the current semantic spatiotemporal module to obtain the fused spatiotemporal representation vectors corresponding to each geographic region. This allows for the fusion of the mutual influence of supply and demand data from multiple geographic regions in the spatiotemporal dimension, thereby improving the representational capability of the fused spatiotemporal representation vectors.
[0078] Figure 9 This is a schematic diagram of the feature convolution fusion module according to an embodiment of the present invention. Figure 9 As shown, in some embodiments, the feature convolutional fusion module 90 includes a topological graph convolution module 91, a semantic graph convolution module 92, a first normalization layer 93, a second normalization layer 94, and an adaptive fusion module 95. In this step, each spatiotemporal hidden state representation vector and the topological adjacency matrix can be used as the input to the topological graph convolution module 91, each spatiotemporal hidden state representation vector and the semantic adjacency matrix can be used as the input to the semantic graph convolution module 92, the output of the topological graph convolution module 91 can be used as the input to the first normalization layer 93, the output of the semantic graph convolution module 92 can be used as the input to the second normalization layer 94, and then the outputs of the first normalization layer 93 and the second normalization layer 94 can be used simultaneously as the input to the adaptive fusion module 95 to obtain the corresponding fused spatiotemporal representation vector.
[0079] Figure 10 This is a flowchart of a data processing method according to an embodiment of the present invention. Figure 10 As shown, in some embodiments, step S230 may include the following steps: Step S231: Input the spatiotemporal latent state representation vectors and topological adjacency matrices into the topological graph convolution module to obtain the corresponding topological space feature vectors.
[0080] In some embodiments, the topological graph convolution module can be a multi-order Chebyshev spectral graph convolution module. A multi-order Chebyshev spectral graph convolution module is a type of spectral graph convolution module capable of extracting topological neighborhood information of each geographic region within a first-order or multi-order neighborhood through frequency domain filtering. Therefore, in this step, after inputting the spatiotemporal latent state representation vectors and the topological adjacency matrix into the topological graph convolution module, the module can transform the topological adjacency matrix into the corresponding graph Laplacian matrix to extract the basis and frequency values of the graph Laplacian matrix in the frequency domain. Then, based on the spatiotemporal latent state representation vectors and the basis and frequency values of the graph Laplacian matrix in the frequency domain, the corresponding topological spatial feature vector is obtained.
[0081] Alternatively, the topological graph convolution module can also be a graph convolutional neural network (GCN), a graph sampling and aggregation network (GraphSAGE), etc.
[0082] Step S232: Input the spatiotemporal latent state representation vectors and semantic adjacency matrices into the semantic graph convolution module to obtain the corresponding semantic space feature vectors.
[0083] In some embodiments, the semantic graph convolution module can be a first-order or multi-order diffusion convolution, which enables feature propagation and aggregation between geographically disadvantaged regions with similar supply and demand patterns, forming supply and demand semantic space features. Therefore, in this step, after inputting the spatiotemporal latent state representation vectors and semantic adjacency matrices into the semantic graph convolution module, the semantic graph convolution module can perform feature aggregation on the spatiotemporal latent state representation vectors according to the semantic adjacency matrix to obtain the corresponding semantic space feature vectors.
[0084] Alternatively, the semantic graph convolution module can also be a graph attention network (GAT), etc.
[0085] Step S233: Input each topological space feature vector into the first normalization layer to obtain the normalized topological space feature vector.
[0086] In this step, each topological space feature vector can be input into the first normalization layer to eliminate the dimensional differences between the topological space feature vectors and obtain the normalized topological space feature vectors.
[0087] Step S234: Input each semantic space feature vector into the second normalization layer to obtain the normalized semantic space feature vector.
[0088] In this step, each semantic space feature vector can be input into the second normalization layer to eliminate the dimensional differences between semantic space feature vectors and obtain normalized semantic space feature vectors.
[0089] Step S235: Input the normalized topological space feature vectors and the normalized semantic space feature vectors into the adaptive fusion module to obtain the corresponding fused spatiotemporal representation vector.
[0090] In this step, the normalized topological space feature vectors and the normalized semantic space feature vectors can be simultaneously input into the adaptive fusion module.
[0091] In some embodiments, the adaptive fusion module can be a gated fusion module, an attention fusion module, a hybrid expert network, a multilayer perceptron, etc. It can generate adaptive weights corresponding to topological space feature vectors and semantic space feature vectors, and perform weighted calculations on the topological space feature vectors and semantic space feature vectors corresponding to each geographical region to obtain the fused spatiotemporal representation vector corresponding to each geographical region.
[0092] Optionally, the adaptive fusion module can also be implemented in other ways. For example, it can consist of a feature stitching layer and a linear projection layer, that is, the normalized topological space feature vectors and normalized semantic space feature vectors corresponding to each geographic space are stitched together and then linearly projected.
[0093] Step S240: Input the supply and demand time series vectors and the fused spatiotemporal representation vectors after linear projection into the gating correction module of the current semantic spatiotemporal module to obtain the corresponding corrected spatiotemporal representation vector.
[0094] In practical applications, while graph convolution can utilize spatial correlation information, it may weaken the peak features of geographical regions in high-demand scenarios. Therefore, in this step, the supply and demand time-series vectors after linear projection and the fused spatiotemporal representation vectors can be simultaneously input into the gating correction module of the current semantic spatiotemporal module to obtain the corrected spatiotemporal representation vectors corresponding to each geographical region. This allows for the adaptive determination of the original time-series information to be retained and the spatial information of the topological and semantic neighborhoods to be fused, based on different geographical regions, time steps, and supply and demand patterns. This improves the likelihood of peak features being retained and alleviates the prediction smoothing problem caused by graph convolution.
[0095] Figure 11 This is a schematic diagram of the gate control correction module according to an embodiment of the present invention. Figure 11 As shown, in some embodiments, the gating correction module 110 includes a feature projection layer 111, a weight generation module 112, and a gating fusion module 113. In this step, each supply and demand time series vector after linear projection can be used as the input to the feature projection layer 111 and the weight generation module 112, respectively, and each fused spatiotemporal representation vector, the output of the feature projection layer 111, and the output of the weight generation module 112 can be used as the input to the gating fusion module 113 to obtain the corresponding corrected spatiotemporal representation vector.
[0096] Figure 12 This is a flowchart of a data processing method according to an embodiment of the present invention. Figure 12 As shown, in some embodiments, step S240 may include the following steps: Step S241: Input each supply and demand time series vector after linear projection into the feature projection layer to obtain the corresponding feature mapping vector.
[0097] In this step, the supply and demand time series vectors after linear projection can be input into the feature projection layer. The feature projection layer aligns the supply and demand time series vectors with the fused spatiotemporal representation vector in terms of dimensions and increases the nonlinear representation capability of the supply and demand time series vectors, thereby obtaining the corresponding feature mapping vector.
[0098] Optionally, the feature projection layer can be a linear projection layer.
[0099] Step S242: Input each supply and demand time series vector after linear projection into the weight generation module to obtain the corresponding gating weights.
[0100] In this step, the supply and demand time series vectors after linear projection can also be input into the gating generation module to adaptively adjust the fusion ratio of multi-path features (i.e., feature mapping vectors and fused spatiotemporal representation vectors) through gating weights.
[0101] In some embodiments, the gated generation module may include a gated projection layer and a sigmoid layer, and the gated projection layer may include multiple one-dimensional convolutional kernels to introduce nonlinear transformations in the channel dimension by combining the gated projection layer and the sigmoid layer, thereby enhancing the representational capability of the gated weights.
[0102] Alternatively, the gating generation module can be implemented in other ways. For example, it can be a fully connected layer, a gated recurrent unit (GRU), or an attention network, etc.
[0103] Step S243: Input each gating weight, each feature mapping vector, and each fused spatiotemporal representation vector into the gating fusion module to obtain the corresponding corrected spatiotemporal representation vector.
[0104] In this step, the gating weights, feature mapping vectors, and fused spatiotemporal representation vectors corresponding to each geographical region can be input into the gating fusion module. This allows the gating fusion module to perform weighted calculations on the linearly projected supply and demand time-series vectors and the fused spatiotemporal representation vectors based on the gating weights, obtaining the corresponding corrected spatiotemporal representation vector. The mapping function f of the gating fusion module can be expressed by the following formula:
[0105] Where g represents the gating weight, Z represents the fused spatiotemporal representation vector, H represents the supply and demand time series vector after linear projection, and proj(H) represents the feature mapping vector.
[0106] Step S250: Input each corrected spatiotemporal representation vector and each supply and demand time series vector into the output module of the current semantic spatiotemporal module to obtain the corresponding input vector.
[0107] In this step, the modified spatiotemporal representation vector and supply-demand time series vector of each geographic region can be input into the output module of the current semantic spatiotemporal module, and the vector output by the output module can be used as the output vector corresponding to each geographic region.
[0108] Figure 13 This is a schematic diagram of the output module according to an embodiment of the present invention. Figure 13As shown, in some embodiments, the output module 130 includes a time series prediction module 131, a residual connection layer 132, a third normalization layer 133, and an output mapping layer 134. In this step, each modified spatiotemporal representation vector can be used as the input to the time series prediction module 131, and each supply and demand time series vector and the output of the time series prediction module 131 can be used as the input to the residual connection layer 132. Then, the output of the residual connection layer 132 can be used as the input to the third normalization layer 133, and the output of the third normalization layer 133 can be used as the input to the output mapping layer 134 to obtain the corresponding output vector.
[0109] Figure 14 This is a flowchart of a data processing method according to an embodiment of the present invention. Figure 14 As shown, in some embodiments, step S250 may include the following steps: Step S251: Input each corrected spatiotemporal representation vector into the time series prediction module to obtain the corresponding time-enhanced feature vector.
[0110] In this step, the modified spatiotemporal representation vectors of each geographical region can be input into the time series prediction module to extract the temporal evolution patterns and periodic fluctuation trends of supply and demand data, thereby obtaining the corresponding time-enhanced feature vectors.
[0111] In some embodiments, the time series prediction module can be a Temporal Convolutional Network (TCN), a Recurrent Neural Network, a Transformer, PatchTST (Patch Time Series Transformer), etc.
[0112] Step S252: Input the time-enhanced feature vectors and the supply-demand time-series vectors into the residual connection layer to obtain the corresponding hybrid feature vectors.
[0113] In this step, the enhanced feature vectors corresponding to each geographical region and each supply and demand time series vector can be simultaneously input into the residual connection layer to solve the gradient vanishing problem. By learning the difference between the original input (i.e., the supply and demand time series vector) and the output (i.e., the enhanced feature vector), the convergence speed of the data prediction model can be improved, while alleviating the prediction smoothing problem caused by graph convolution, and the corresponding hybrid feature vector is obtained.
[0114] Step S253: Input each mixed feature vector into the third normalization layer to obtain the corresponding normalized mixed feature vector.
[0115] In this step, the mixed feature vectors of each geographical region can be input into the third normalization layer to reduce the dimensional differences and obtain the normalized mixed feature vectors.
[0116] Step S254: Input the normalized mixed feature vectors into the output mapping layer to obtain the corresponding output vectors.
[0117] In this step, the normalized hybrid feature vectors corresponding to each geographical region can be input into the output mapping layer to obtain the output vector of the current semantic spatiotemporal module.
[0118] Step S260: Determine whether the current semantic spatiotemporal module is ranked last.
[0119] In this step, it can be determined whether the current semantic spatiotemporal module is the last semantic spatiotemporal module in the data prediction model. If yes, step S270 can be executed; if no, step S280 can be executed.
[0120] Step S270: Determine each output vector as a supply and demand time series prediction vector.
[0121] In this step, the output vector corresponding to each geographical region can be determined as the corresponding supply and demand time series prediction vector to achieve the prediction of supply and demand data.
[0122] Step S280: Update each output vector to the input vector of the next semantic spatiotemporal module.
[0123] In this step, the output vector corresponding to each geographical region can be determined as the input vector of the next semantic spatiotemporal module to continue the data processing of supply and demand data.
[0124] It is easy to understand that after obtaining the input vector of the next semantic spatiotemporal module, we can return to execute step S220.
[0125] Figure 15 This is a data flow diagram of the data processing method according to an embodiment of the present invention. Figure 15As shown, supply and demand time-series data generated in multiple geographic regions from T-1008 (i.e., the time corresponding to 1008 time steps prior to the current time) to T-1000, supply and demand time-series data generated in the T-144-T-136 time period, and time-series data generated in the T-9-T-1 time period can be obtained as supply and demand time-series data 151. A corresponding supply and demand time-series vector 154 is determined based on the supply and demand time-series data of each geographic region. Furthermore, supply and demand statistics data 152 can be obtained for each geographic region, and a corresponding semantic adjacency matrix 155 can be obtained based on the supply and demand statistics data 152. Simultaneously, geographic location relationships 153 of each geographic region can be obtained, and a corresponding topological adjacency matrix 156 can be obtained based on the geographic location relationships 153. Then, the supply and demand time-series vector 154, semantic adjacency matrix 155, and topological adjacency matrix 156 of each geographic region are input into the data prediction model 150. The data prediction model 150 may include two semantic spatiotemporal modules (i.e., stacked blocks) 157. Each stacked block 157 includes a spatiotemporal feature encoding module 70, a feature convolution fusion module 90, a gating correction module (i.e., spatiotemporal correction gate) 110, and an output module 130. Therefore, the linear projection layer of the spatiotemporal feature encoding module 70, the spatial attention (i.e., the spatial attention encoding module), the semantic convolution (i.e., the semantic graph convolution module), the Chebyshev convolution (i.e., the topological graph convolution module) of the feature convolution fusion module 90, the first normalization layer (not shown in the figure), the second normalization layer (not shown in the figure), the adaptive fusion (i.e., the adaptive fusion module), the feature projection layer, the gated projection layer, the sigmoid layer, the gated fusion module of the gated correction module 110, the temporal convolution (i.e., the temporal convolutional network), the residual connection layer, the layer normalization and Dropout (i.e., the third normalization layer and the output mapping layer) of the output module 130 are used to process at least one type of data in the supply and demand time series vector 154, the semantic adjacency matrix 155 and the topological adjacency matrix 156 of each geographic region, thereby outputting the supply and demand prediction time series vector 158 corresponding to the time steps from T to T+3, T to T+6 and T to T+9 of each geographic region.
[0126] This invention, after obtaining the topological adjacency matrix, semantic adjacency matrix, and supply and demand time-series vectors of multiple geographical regions over past time periods, uses this data and a pre-trained data prediction model to obtain the supply and demand time-series prediction vectors for each geographical region in future time periods. In this invention, the topological adjacency matrix is determined based on the geographical location relationships of each geographical region, the semantic adjacency matrix is determined based on the similarity of the semantic vectors of each geographical region, and the data prediction model includes multiple semantic spatiotemporal modules. Therefore, this invention enhances the feature representation capability of multiple geographical regions by leveraging their geographical location relationships and semantic similarity, and improves the accuracy of supply and demand data prediction by repeatedly fusing the spatial correlation patterns and long-term time-series patterns of multiple geographical regions through semantic spatiotemporal modules.
[0127] Figure 16 This is a schematic diagram of a data processing apparatus according to an embodiment of the present invention. Figure 16 As shown, the data processing device in this embodiment of the invention includes a data acquisition unit 1601 and a data prediction unit 1602.
[0128] The data acquisition unit 1601 is used to acquire supply and demand time-series vectors of multiple geographical regions during a predetermined time period, topological adjacency matrices and semantic adjacency matrices of each geographical region. The topological adjacency matrices are determined based on the geographical location relationships of each geographical region, and the semantic adjacency matrices are determined based on the similarity of the semantic vectors of each geographical region. The semantic vectors are determined based on the supply and demand statistics of the corresponding geographical region. The data prediction unit 1602 is used to acquire supply and demand time-series prediction vectors of each geographical region during a target time period based on the supply and demand time-series vectors, the topological adjacency matrices and the semantic adjacency matrices, and a pre-trained data prediction model. The data prediction model includes multiple semantic spatiotemporal modules.
[0129] This invention, after obtaining the topological adjacency matrix, semantic adjacency matrix, and supply and demand time-series vectors of multiple geographical regions over past time periods, uses this data and a pre-trained data prediction model to obtain the supply and demand time-series prediction vectors for each geographical region in future time periods. In this invention, the topological adjacency matrix is determined based on the geographical location relationships of each geographical region, the semantic adjacency matrix is determined based on the similarity of the semantic vectors of each geographical region, and the data prediction model includes multiple semantic spatiotemporal modules. Therefore, this invention enhances the feature representation capability of multiple geographical regions by leveraging their geographical location relationships and semantic similarity, and improves the accuracy of supply and demand data prediction by repeatedly fusing the spatial correlation patterns and long-term time-series patterns of multiple geographical regions through semantic spatiotemporal modules.
[0130] Figure 17 This is a schematic diagram of an electronic device according to an embodiment of the present invention. (For example...) Figure 17As shown, the electronic device is a general-purpose data processing device, which includes a general-purpose computer hardware structure, including at least a processor 1701 and a memory 1702. The processor 1701 and memory 1702 are connected via a bus 1703. The memory 1702 is adapted to store instructions or programs executable by the processor 1701. The processor 1701 can be a standalone microprocessor or a collection of one or more microprocessors. Thus, the processor 1701 executes the instructions stored in the memory 1702, thereby performing the method flow of the embodiments of the present invention as described above to process data and control other devices. The bus 1703 connects the aforementioned components together, and also connects the aforementioned components to a display controller 1704, a display device, and an input / output (I / O) device 1705. The input / output (I / O) device 1705 can be a mouse, keyboard, modem, network interface, touch input device, motion-sensing input device, printer, and other devices known in the art. Typically, the input / output device 1705 is connected to the system via an input / output (I / O) controller 1706.
[0131] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus (devices), or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0132] This application is described with reference to flowchart illustrations of methods, apparatus (devices), and computer program products according to embodiments of this application. It should be understood that each step in the flowchart can be implemented by computer program instructions.
[0133] These computer program instructions may be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction means, the implementation process of which is described in the instruction means. Figure 1 The function specified in one or more processes.
[0134] These computer program instructions may also be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, produce instructions for implementing processes. Figure 1 A device for a function specified in one or more processes.
[0135] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program for use by a computer to execute some or all of the above-described method embodiments.
[0136] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program specifying the relevant hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0137] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. For those skilled in the art, the present invention can be modified and varied in various ways. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of protection of the present invention.
Claims
1. A data processing method, characterized in that, The method includes: Obtain supply and demand time series vectors for multiple geographical regions over a predetermined time period, a topological adjacency matrix for each geographical region, and a semantic adjacency matrix. The topological adjacency matrix is determined based on the geographical location relationships of each geographical region, and the semantic adjacency matrix is determined based on the similarity of the semantic vectors of each geographical region. The semantic vectors are determined based on the supply and demand statistics of the corresponding geographical regions. Based on the supply and demand time series vectors, the topological adjacency matrix, and the semantic adjacency matrix, and using a pre-trained data prediction model, supply and demand time series prediction vectors for each geographical region in the target time period are obtained. The data prediction model includes multiple semantic spatiotemporal modules.
2. The method according to claim 1, characterized in that, The topological adjacency matrix is determined in the following way: Obtain at least one of the following for each of the geographic regions: boundary adjacency relationship, distance relationship, road network connectivity relationship, and spatial adjacency relationship. The boundary adjacency relationship is used to characterize whether the boundaries of each of the geographic regions are adjacent. The road network connectivity relationship is used to characterize whether the road networks of each of the geographic regions are connected. The spatial adjacency relationship is used to characterize whether the geographic spaces to which each of the geographic regions belong are adjacent. The topological adjacency matrix is determined based on at least one of the boundary adjacency relationship, the distance relationship, the road network connectivity relationship, and the spatial adjacency relationship.
3. The method according to claim 1, characterized in that, The semantic adjacency matrix is determined in the following way: Based on the supply and demand statistics, determine the semantic vector of the corresponding geographical region under each feature dimension; The similarity of each geographical region under the same feature dimension is determined based on the semantic vectors under the same feature dimension. The semantic similarity of each geographical region is determined based on the similarity under each feature dimension; The semantic adjacency matrix is determined based on the semantic similarity of each property.
4. The method according to claim 1, characterized in that, The supply and demand time series vector is determined in the following way: Relative and absolute numerical data are determined based on the supply and demand time series data of each geographical region within a predetermined time period. Logarithmic processing is performed on the absolute numerical data to obtain the first type of time series data; The second type of time series data is determined based on the relative numerical data; The supply and demand time series vector is determined based on the first type of time series data and the second type of time series data.
5. The method according to claim 4, characterized in that, The step of determining the second type of time series data based on the relative numerical data includes: The relative numerical data is normalized or standardized to obtain the second type of time series data.
6. The method according to claim 1, characterized in that, The semantic spatiotemporal module includes a spatiotemporal feature encoding module, a feature convolution fusion module, a gating correction module, and an output module.
7. The method according to claim 6, characterized in that, The step of obtaining the supply and demand time-series prediction vector for each geographical region in the target time period based on the supply and demand time-series vector, the topological adjacency matrix, and the semantic adjacency matrix, and using a pre-trained data prediction model, includes: Each of the aforementioned supply and demand time-series vectors is determined as the input vector of the semantic spatiotemporal module that is ranked first; Each input vector is input into the spatiotemporal feature encoding module of the current semantic spatiotemporal module in an iterative manner to obtain the corresponding spatiotemporal latent state representation vector; Each of the spatiotemporal latent state representation vectors, the topological adjacency matrix, and the semantic adjacency matrix are input into the feature convolution fusion module of the current semantic spatiotemporal module to obtain the corresponding fused spatiotemporal representation vector; The supply and demand time series vectors and the fused spatiotemporal representation vectors after linear projection are input into the gating correction module of the current semantic spatiotemporal module to obtain the corresponding corrected spatiotemporal representation vector. Each of the modified spatiotemporal representation vectors and each of the supply and demand time series vectors are input into the output module of the current semantic spatiotemporal module to obtain the corresponding output vector; In response to the fact that the current semantic spatiotemporal module is at the end of the order, each of the output vectors is determined as the supply and demand time series prediction vector; In response to the current semantic spatiotemporal module not being sorted at the end, each of the output vectors is updated to the input vector of the next semantic spatiotemporal module.
8. The method according to claim 7, characterized in that, The spatiotemporal feature encoding module includes a linear projection layer and a spatial attention encoding module; The step of inputting each of the input vectors into the spatiotemporal feature encoding module of the current semantic spatiotemporal module to obtain the corresponding spatiotemporal latent state representation vector includes: Each input vector is input into the linear projection layer to obtain the corresponding hidden state vector; Each of the hidden state vectors is input into the spatial attention encoding module to obtain the corresponding spatiotemporal hidden state representation vector.
9. The method according to claim 7, characterized in that, The feature convolutional fusion module includes a semantic graph convolutional module, a topological graph convolutional module, a first normalization layer, a second normalization layer, and an adaptive fusion module. The step of inputting each of the spatiotemporal latent state representation vectors, the topological adjacency matrix, and the semantic adjacency matrix into the feature convolutional fusion module of the current semantic spatiotemporal module to obtain the corresponding fused spatiotemporal representation vector includes: The spatiotemporal latent state representation vectors and the topological adjacency matrix are input into the topological graph convolution module to obtain the corresponding topological space feature vectors. The spatiotemporal latent state representation vectors and the semantic adjacency matrix are input into the semantic graph convolution module to obtain the corresponding semantic space feature vectors; Each of the topological space feature vectors is input into the first normalization layer to obtain the normalized topological space feature vectors; Each of the semantic space feature vectors is input into the second normalization layer to obtain the normalized semantic space feature vectors; The normalized topological space feature vectors and the normalized semantic space feature vectors are input into the adaptive fusion module to obtain the corresponding fused spatiotemporal representation vector.
10. The method according to claim 7, characterized in that, The gating correction module includes a feature projection layer, a weight generation module, and a gating fusion module; The step of inputting the linearly projected supply and demand time series vectors and the fused spatiotemporal representation vectors into the gating correction module of the current semantic spatiotemporal module to obtain the corresponding corrected spatiotemporal representation vector includes: The supply and demand time series vectors after linear projection are input into the feature projection layer to obtain the corresponding feature mapping vectors. The supply and demand time series vectors after linear projection are input into the weight generation module to obtain the corresponding gating weights; Each of the gating weights, each of the feature mapping vectors, and each of the fused spatiotemporal representation vectors are input into the gating fusion module to obtain the corresponding corrected spatiotemporal representation vector.
11. The method according to claim 7, characterized in that, The output module includes a time series prediction module, a residual connection layer, a third normalization layer, and an output mapping layer; The step of inputting each of the modified spatiotemporal representation vectors and each of the supply and demand time series vectors into the output module of the current semantic spatiotemporal module to obtain the corresponding output vector includes: Each of the modified spatiotemporal representation vectors is input into the time series prediction module to obtain the corresponding time-enhanced feature vectors. The time-enhanced feature vectors and the supply-demand time-series vectors are input into the residual connection layer to obtain the corresponding hybrid feature vectors. Each of the aforementioned hybrid feature vectors is input into the third normalization layer to obtain the corresponding normalized hybrid feature vector; The normalized hybrid feature vectors are input into the output mapping layer to obtain the corresponding output vectors.
12. A data processing apparatus, characterized in that, The device includes: The data acquisition unit is used to acquire supply and demand time series vectors of multiple geographical regions over a predetermined time period, topological adjacency matrices and semantic adjacency matrices of each geographical region. The topological adjacency matrix is determined based on the geographical location relationships of each geographical region, and the semantic adjacency matrix is determined based on the similarity of the semantic vectors of each geographical region. The semantic vectors are determined based on the supply and demand statistics of the corresponding geographical region. The data prediction unit is used to obtain the supply and demand time series prediction vectors of each geographical region in the target time period based on the supply and demand time series vectors, the topological adjacency matrix and the semantic adjacency matrix, and a pre-trained data prediction model. The data prediction model includes multiple semantic spatiotemporal modules.
13. A computer-readable storage medium storing computer program instructions thereon, characterized in that, The computer program instructions, when executed by a processor, implement the method as described in any one of claims 1-11.
14. An electronic device, characterized in that, The device includes: Memory is used to store one or more computer program instructions; A processor, wherein the one or more computer program instructions are executed by the processor to implement the method as described in any one of claims 1-11.
15. A computer program product, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1-11.