Data processing method, system, apparatus, device, and medium

CN116720588BActive Publication Date: 2026-08-21ALIBABA CLOUD COMPUTING CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310574685.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-18
Publication Date
2026-08-21
Estimated Expiration
2043-05-18

AI Technical Summary

Technical Problem

在执行任务过程中需要进行大量数据传输,数据处理效率低,消耗大量带宽和存储资源

Benefits of technology

[0024]本申请实施例提供的技术方案,为了便于对数据管理,有的数据库中会按照时间线进行时序数据存储。在一些数据库系统中,会利用机器学习模型来满足其数据处理需求。因此,本方案中,根据针对目标属性的数据处理请求,将数据库中内置的机器学习模型的目标算子发送到存储有目标属性对应的时间线的存储节点中,而不需要将存储节点中的目标时序数据传输到计算节点,能够有效减少节点之间的数据传输量,提升计算效率。同时,利用机器学习模型处理时序数据时,只需要根据数据处理请求选择对应的目标算子,并将该目标算子下推到对应的存储节点中,而不需要传输过于复杂的机器学习模型就能够满足数据处理需求。同时,通过将同一机器学习模型中目标算子分别下推到不同存储节点中,并由不同存储节点执行相应数据处理任务,实现多存储节点并行处理,有效提高数据处理效率和数据处理能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116720588B_ABST
    Figure CN116720588B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a data processing method, system, device, equipment and medium. The method comprises: in response to a data processing request for a target attribute, determining a target operator of a machine learning model built-in in a database; determining a target timeline corresponding to the time sequence identifier corresponding to the target attribute, and a storage node storing the target timeline; pushing the target operator to the storage node, so that the storage node executes the data processing request for the target time sequence data in the target timeline by using the target operator. According to the data processing request for the target attribute, the target operator of the machine learning model built-in in the database is sent to the storage node storing the timeline corresponding to the target attribute, without the need to transmit the target time sequence data in the storage node to the computing node, which can effectively reduce the data transmission amount between nodes and improve the computing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to data processing methods, systems, apparatus, devices and media. Background Technology

[0002] With the development of artificial intelligence technology, machine learning models are being used in more and more scenarios to perform related learning and prediction tasks, thereby effectively reducing the workload of human staff.

[0003] In practical applications, when using data from a database, it is usually necessary to retrieve the required data from the database. For example, when using a machine learning model to perform a task, data needs to be retrieved from the database, and then the machine learning model can be used to perform data processing tasks. This process involves a large amount of data transfer, resulting in low data processing efficiency and consuming significant bandwidth and storage resources. Summary of the Invention

[0004] To address or improve the problems existing in the prior art, various embodiments of this application provide data processing methods, systems, apparatuses, devices, and media.

[0005] In a first aspect, one embodiment of this application provides a data processing method. Applied to a computing node, the method includes:

[0006] In response to a data processing request for a target attribute, determine the target operator of the machine learning model built into the database;

[0007] Based on the time sequence identifier corresponding to the target attribute, determine the target timeline corresponding to the time sequence identifier, and the storage node for storing the target timeline;

[0008] The target operator is pushed down to the storage node so that the storage node can use the target operator to execute the data processing request for the target time series data in the target timeline.

[0009] Secondly, in one embodiment of this application, a data processing method is provided. Applied to a storage node, the method includes:

[0010] The computing node receives the target operator in the machine learning model pushed down by the computing node for executing data processing requests; the computing node is used to determine the corresponding target timeline and the storage node for storing the target timeline based on the time sequence identifier of the target attribute corresponding to the data processing request.

[0011] Based on the target timeline, determine the target time-series data used to satisfy the data processing request;

[0012] The target operator is used to perform data processing on the target time series data, and the data processing results are sent to the computing node.

[0013] Thirdly, in one embodiment of this application, a data processing apparatus is provided, the apparatus comprising:

[0014] The first determining module is used to determine the target operator of the machine learning model built into the database in response to a data processing request for the target attribute.

[0015] The second determining module is used to determine the target timeline corresponding to the timeline identifier and the storage node storing the target timeline based on the timeline identifier corresponding to the target attribute.

[0016] The push-down module is used to push the target operator down to the storage node so that the storage node can use the target operator to execute the data processing request for the target time series data in the target timeline.

[0017] Fourthly, in one embodiment of this application, a data processing system is provided, comprising:

[0018] A computing node for performing the method described in any one of the first aspects;

[0019] Storage nodes are used to execute the methods described in the second aspect.

[0020] Fifthly, in one embodiment of this application, an electronic device is provided, including a memory and a processor; wherein,

[0021] The memory is used to store programs;

[0022] The processor, coupled to the memory, is configured to execute the program stored in the memory for implementing the method of the first aspect or for implementing the method of the second aspect.

[0023] In a sixth aspect, in one embodiment of this application, a non-transitory machine-readable storage medium is provided, wherein executable code is stored on the non-transitory machine-readable storage medium, and when the executable code is executed by a processor of an electronic device, the processor performs the method as described in the first aspect, or performs the method as described in the second aspect.

[0024] The technical solution provided in this application addresses the issue that, for ease of data management, some databases store time-series data according to a timeline. In some database systems, machine learning models are used to meet their data processing needs. Therefore, in this solution, based on the data processing request for the target attribute, the target operator of the machine learning model built into the database is sent to the storage node containing the timeline corresponding to the target attribute, without needing to transfer the target time-series data from the storage node to the computing node. This effectively reduces the amount of data transfer between nodes and improves computational efficiency. Furthermore, when processing time-series data using a machine learning model, it is only necessary to select the corresponding target operator according to the data processing request and push the target operator down to the corresponding storage node, without needing to transfer an overly complex machine learning model to meet the data processing requirements. Moreover, by pushing the target operators from the same machine learning model down to different storage nodes and having different storage nodes execute the corresponding data processing tasks, parallel processing across multiple storage nodes is achieved, effectively improving data processing efficiency and capabilities. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 A flowchart illustrating the data processing method provided in an embodiment of this application;

[0027] Figure 2 A flowchart illustrating another data processing method provided in an embodiment of this application;

[0028] Figure 3 A schematic diagram of a data processing apparatus provided in an embodiment of this application;

[0029] Figure 4 A schematic diagram of another data processing apparatus provided in an embodiment of this application;

[0030] Figure 5 This is a schematic diagram of the system structure provided in the embodiments of this application;

[0031] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0032] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0033] In some processes described in the specification, claims, and accompanying drawings of this application, multiple operations appearing in a specific order are included. These operations may be executed out of order or in parallel. Operation numbers such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the terms "first," "second," etc., used herein are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types. Moreover, the embodiments described below are only a part of the embodiments of this application, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0034] In existing databases, machine learning models are often used to achieve better data processing results. There are two main approaches: one is to use an external machine learning model to perform the data processing task, and the other is to use a machine learning model built into the database. Specifically, when using an external machine learning model, data needs to be retrieved from the database. The larger the amount of data retrieved, the greater the bandwidth consumption and the lower the data processing efficiency; it also makes the database system more complex. Alternatively, the machine learning model can be built into the database. Specifically, in relational databases, when processing data based on a timeline, the data in a timeline is flattened into multiple rows. Because the time-series data is fragmented—meaning each time-series data point is stored on different storage nodes—the time-series data needs to be transferred to the computing nodes to perform the corresponding computational tasks. This data transfer between nodes consumes a significant amount of bandwidth.

[0035] Terminology Explanation:

[0036] TimeSeries: A timeline is formed by the change of a specific metric from a data source over time. When storing data, databases cluster data from the same timeline to improve the efficiency of accessing time-series data. This also better supports time-series data compression. In a time-series data table, a series of data rows with the same Tag column value constitutes a timeline.

[0037] Operator: A type of operator used in databases to perform calculations on data.

[0038] The technical solution implemented in this application will be explained and described below with reference to specific embodiments.

[0039] like Figure 1 This is a flowchart illustrating the data processing method provided in an embodiment of this application. The execution entity of this method can be a computing node. From Figure 1 The specific steps can be seen as follows:

[0040] 101: In response to a data processing request for a target attribute, determine the target operator of the machine learning model built into the database.

[0041] 102: Based on the time sequence identifier corresponding to the target attribute, determine the target timeline corresponding to the time sequence identifier and the storage node for storing the target timeline.

[0042] 103: Push the target operator down to the storage node so that the storage node can use the target operator to execute the data processing request for the target time series data in the target timeline.

[0043] It should be noted that in the technical solution of this application, the data stored in the database is time-series data; specifically, the data in the database is time-series data stored according to a timeline. Time-series data of the same target attribute are often stored in the same timeline, and time-series data within a continuous period of a timeline are stored in the same storage node, thus facilitating the management of time-series data. When a data processing task needs to be performed on a specific target attribute, it is often not necessary to collect time-series data across multiple storage nodes.

[0044] In this application, the machine learning model is built into the database. When using the machine learning model, there is no need to export the data in the database to an external platform, which can effectively improve data processing efficiency and ensure data security in the database.

[0045] The target operators required for processing the same data with different target attributes differ, the target operators required for processing different data with the same target attribute differ, and the target operators required for processing different data with different target attributes differ. It's important to note that different target attributes can be understood as attributes of different products. These attributes generate time-series related data; for example, a target attribute could be the customer repurchase rate of a certain garment or the peak electricity consumption time in a certain region. Therefore, in practical applications, to ensure more accurate processing results from machine learning models, different machine learning models need to be trained for different target attributes to meet the data processing needs of those attributes. The different data processing requests mentioned here could be, for example, time-series prediction requests or time-series anomaly detection requests.

[0046] Machine learning models offer a variety of operators; in other words, different target operators are selected for different data processing needs based on different target attributes. For example, operators required for time series prediction requests could include: the DeepAR algorithm, a deep neural network algorithm based on a recurrent neural network (RNN), or the attention-based Temporal Fusion Transformer (TFT) algorithm, a deep neural network algorithm based on the Transformer mechanism. Operators required for time series anomaly detection requests could include: nsigma, which has a simple principle and is convenient for analyzing the causes of time series data anomalies.

[0047] It's important to note that in time-series data, each timeline has a unique time-series identifier. Time-series data storage defines a timeline as a combination of a metric and a set of tags. Within a timeline, sampled data at consecutive time points constitutes time-series data. Continuous time-series data (timestamp, value) is stored within the same timeline. During storage, the same timeline is stored on the same storage node. However, if the storage node is expanded, the expanded time-series data may be stored on different storage nodes than the original data. In this case, the distribution of the required time-series data across storage nodes needs to be determined based on the routing table to decide on the appropriate data processing method (including: the compute node pushing the target operator down to the storage node, or the storage node transmitting time-series data to the compute node). The specific selection process will be explained in detail in the following embodiments and will not be repeated here.

[0048] When machine learning models execute data processing tasks, they select target time-series data within a specific time period and determine the storage node where this target time-series data resides, based on requirements. In database systems, one computing node corresponds to multiple storage nodes. When processing time-series data using target operators, the data typically needs to be continuous. Therefore, the target operator is pushed down to the storage node storing continuous timelines. The completion of the data processing request may involve multiple different target timelines corresponding to the same target attribute working together. In this case, the same target operator can be sent to the corresponding storage nodes, and each storage node can execute its respective data processing task, enabling multiple storage nodes to execute data processing tasks in parallel. Furthermore, multiple different target timelines may be stored in the same storage node. In this case, the target operator can be sent to that storage node, and the storage node can execute data processing tasks for different target timelines sequentially, or the storage node can execute data processing for multiple different target timelines in parallel. If multiple storage nodes send time-series data to the compute nodes separately, the limited computing power of each compute node necessitates sequentially executing data processing tasks on the time-series data provided by each storage node. This makes data processing by a single compute node significantly less efficient than parallel processing. It should be noted that in the database system of this application, each storage node possesses sufficient computing power to meet the requirements of the target operator in executing data processing tasks. The amount of data output after executing the data processing tasks is significantly smaller than the amount of time-series data in the timeline, thus effectively reducing the bandwidth consumed by data transmission.

[0049] The above approach embeds the machine learning model into the database and selects the appropriate target operator from the model based on the data processing request and target attributes during data processing tasks. Simultaneously, it determines the storage node containing the corresponding timeline based on the time series identifier corresponding to the target attribute. Then, the target operator is pushed down to the corresponding storage node without needing to transfer the time series data from the storage node to the computing node, effectively reducing data transfer pressure. Furthermore, processing time series data according to the timeline allows data processing tasks to be distributed to various storage nodes for parallel execution, effectively improving data processing efficiency.

[0050] In one or more embodiments of this application, determining the target timeline corresponding to the timeline identifier and the storage node storing the target timeline based on the timeline identifier corresponding to the target attribute includes: determining the timeline identifier corresponding to the target attribute and determining the corresponding target timeline based on the timeline identifier; determining the routing table corresponding to the target timeline according to the timeline identifier; and determining the storage node corresponding to the target timeline and the number of storage nodes according to the routing information stored in the routing table.

[0051] In practical applications, when accessing data, the corresponding storage node is located through the routing table to access the data. Specifically, based on the target attributes, the corresponding time sequence identifier and target timeline are determined. Furthermore, the routing table corresponding to the target timeline is determined using the time sequence identifier. The routing path in the routing table reveals the storage node corresponding to that timeline and the number of storage nodes. In other words, when the routing table contains more than one storage node, it indicates that the time-series data corresponding to that timeline is distributed and stored across different storage nodes. When the routing table contains only one storage node, it means that the time-series data corresponding to that timeline is stored in the same storage node.

[0052] For example, a timeline {"metric": "cpu", "tags": "site": "et2", "ip": "1.1.1.1"} is obtained. The combination of metric and tag forms a timeline, which contains the CPU information corresponding to the metric, and the site and IP information corresponding to the tags. For example, the site is et2, and the IP is 1.1.1.1. Continuous time-series data (timestamp, value) is stored under the same timeline. For example, timestamp: 12345, value: 1; timestamp: 12346, value: 2.

[0053] Because time-series data in a timeline is time-related, when storage nodes are expanded, the routing table is updated. The routing table allows for accurate identification of the storage node corresponding to the timeline. Consequently, different data processing methods are employed for different storage methods of the timeline (stored on the same storage node or stored on multiple different storage nodes), thereby maximizing data processing efficiency while meeting data processing requirements.

[0054] In one or more embodiments of this application, the step of pushing the target operator down to the storage node includes: if the storage node corresponding to the target timeline is found to be a unique node, then the target operator is pushed down to the storage node.

[0055] In practical applications, if the routing table determines that all time-series data for a given timeline is stored in the same storage node, it means that when using a machine learning model to process the time-series data, the target operator can be pushed down to the corresponding storage node. This eliminates the need to upload the time-series data to the compute node, allowing the data processing task to be completed. Since machine learning models often require selecting continuous data within a certain time period for processing time-series data, the target operator can only be pushed down to that storage node and the corresponding data processing task executed if the time-series data is in the same storage node. The time-series data in the storage node meets the requirements of the data processing task, eliminating the need to retrieve data from other storage nodes. In other words, the corresponding data processing results can be achieved within the storage node, avoiding data retrieval from storage nodes and reducing network bandwidth consumption caused by data transmission.

[0056] In one or more embodiments of this application, if multiple storage nodes are found to correspond to the target timeline, the target time series data in the multiple storage nodes corresponding to the same target timeline is received.

[0057] In practical applications, if storage space is insufficient, storage nodes may be expanded, meaning the number of storage nodes will be increased. This will increase the number of routing paths in the routing table, allowing multiple storage nodes to be found based on time-series identifiers. This implies that the time-series data required by the machine learning model is stored on different storage nodes. If the target operator is sent to multiple storage nodes, the data processing requirements cannot be met because each storage node does not store complete time-series data. Therefore, multiple storage nodes can send time-series data from the same timeline to the compute nodes, which then perform the data processing tasks.

[0058] For example, consider storage nodes A and B. Storage node A stores time-series data for timeline A1, while storage node B stores time-series data for timeline B1. Suppose that storage node B runs out of storage space, requiring the addition of storage node C, which will then store the time-series data for timeline B1. If a data processing request for timeline B1 is received, a routing table lookup reveals that timeline B1 is stored in both storage nodes B and C. Storage nodes B and C can then send the corresponding time-series data for timeline B1 to the compute nodes, which will then execute the appropriate data processing tasks.

[0059] In one or more embodiments of this application, determining the target operator of the machine learning model built into the database in response to a data processing request for a target attribute includes:

[0060] In response to a data processing request for a target attribute, determine the target operator type and the target operator required for the data processing request; or, determine the target operator based on the function name carried in the data processing request.

[0061] In practical applications, operators from machine learning models can be stored as database functions in a database function set. This database function set can contain various machine learning model operators, such as time-series prediction operators and time-series anomaly detection operators. Furthermore, different specialized machine learning models are used for different target attributes, resulting in variations in training sample selection and training, as well as the data used for data processing. Therefore, upon receiving a data processing request, the target operator type matching the request type and target attribute is determined, and then the corresponding target operator is retrieved from the database functions based on that type. Alternatively, users can specify the target operator by function name when initiating a data processing request. When the computing node receives the request, it can parse the request and accurately locate the appropriate target operator based on the function name.

[0062] When a data processing request is a comprehensive request involving multiple different processing methods, it needs to be split into tasks. For example, it might be split into two data processing requests: one for prediction and the other for anomaly detection. Since different data processing requests select different target operator types, the corresponding target operators are selected and sent to the storage node, and the corresponding data processing tasks are executed in sequence. Of course, if two data processing tasks are related, the result of the first task can be used as input for the second task, allowing the data processing tasks to be completed without data transfer.

[0063] In one or more embodiments of this application, determining the storage node corresponding to the target timeline and the number of storage nodes based on the routing information stored in the routing table includes:

[0064] Based on the time range specified in the data processing request, determine the routing information corresponding to the time range in the routing table;

[0065] Based on the routing information, the storage nodes corresponding to the target timeline and the number of storage nodes are determined.

[0066] In practical applications, when using machine learning models to perform data processing tasks on time-series data, the processing is often performed on consecutive time-series data. Although the time-series data of some timelines are stored on different storage nodes, the time-series data required by the machine learning model to perform data processing tasks may happen to be concentrated in the same storage node. That is, it is not necessary to collect time-series data across multiple storage nodes. Therefore, based on the time range and routing table, the corresponding routing information can be determined (in other words, the routing information includes the time label range of the time-series data and the corresponding storage node information, for example, the time-series data corresponding to time A1 to time A99 is stored in storage node D1, and the time-series data corresponding to time B1 to time B99 is stored in storage node D2). This allows the determination of the storage nodes corresponding to the timelines required to execute the data processing request and the number of storage nodes. When the target timeline determined based on the routing information corresponds to multiple storage nodes and multiple storage nodes, the multiple storage nodes send the time-series data to the computing nodes, and the computing nodes execute the corresponding data processing tasks. If the target timeline determined by the routing information corresponds to only one storage node, then the compute node can send the corresponding target operator to that storage node, and the storage node can then execute the corresponding data processing task. It should be noted that this time range can be a start and end time specified by the data processing request, or a specified time length without a defined start and end time.

[0067] Of course, the timeline to be split, as well as the time range and corresponding storage nodes of the sub-timelines split from the same timeline, can also be determined based on the routing information stored in the routing table. Furthermore, the time-series data in different storage nodes can be used to obtain the corresponding data processing results.

[0068] The above approach leverages the temporal correlation of time-series data. Even with storage node expansion, the feasibility of operator pushdown can be determined based on the actual analysis results within a specified time range. Although the target timeline in the routing table corresponds to routing paths across multiple different storage nodes, if the time-series data required by the machine learning model happens to reside on a specific storage node, the target operator can be pushed down to that node, effectively improving data processing efficiency.

[0069] In one or more embodiments of this application, the training method of the machine learning model includes:

[0070] Push down the target operator in the machine learning model to be trained to the storage node;

[0071] The machine learning model to be trained is trained using the target time-series data stored in the timeline of the storage node.

[0072] In practical applications, machine learning models can also be trained on corresponding storage nodes. Specifically, when training a machine learning model, its purpose is clearly defined, and operators that meet the requirements are determined. Ideally, when using a machine learning model, the target operator needs to be pushed down to the storage node, and the time-series data in the storage node is used to perform the corresponding data processing tasks. Therefore, training the machine learning model also requires using the time-series data stored in the storage node. Since a large number of training samples are needed to obtain relatively accurate training results, the target operator in the machine learning model is pushed down to the storage node that can provide training samples. The storage node provides historical time-series data as training samples to train the machine learning model's operators, thus obtaining the target operator. Alternatively, if the training samples corresponding to the same timeline are distributed across multiple different storage nodes, these different storage nodes need to distribute the training samples up to the computing node, where the computing node performs the training. Using the above method, when training samples from the same timeline are in the same storage node, it is not necessary to extract and transmit a large number of training samples from the storage node to the computing node to train the machine learning model. This can effectively reduce the network bandwidth consumption of data transmission and improve training efficiency.

[0073] Furthermore, since the training samples provided by multiple timelines are stored in different storage nodes, the operators to be trained can be sent to different storage nodes to execute the corresponding training tasks, thereby achieving parallel training and improving training efficiency.

[0074] Using the above method, it is not necessary to send a large number of training samples (i.e., time-series data stored in storage nodes) to training nodes (e.g., compute nodes). Instead, a large amount of time-series data can be used as training samples in the storage nodes, which can effectively improve training efficiency. The trained machine learning model or target operator is then sent to the compute node.

[0075] In one or more embodiments of this application, determining the target timeline corresponding to the timeline identifier based on the timeline identifier corresponding to the target attribute, and the storage node storing the target timeline, includes:

[0076] Based on the time sequence identifier corresponding to at least one of the target attributes, determine the target timeline corresponding to each of the time sequence identifiers;

[0077] Determine the storage nodes corresponding to each of the target timelines;

[0078] The step of pushing the target operator down to the storage node includes: pushing the target operator down to each of the storage nodes.

[0079] In practical applications, if a user-initiated data processing request targets multiple different attributes, it's necessary to determine the corresponding time series identifier based on at least one target attribute, and then determine the target timeline corresponding to each time series identifier. It's important to note that the target timelines corresponding to each time series identifier may be stored on completely different storage nodes or on the same storage node. If stored on different storage nodes, the compute node can push down the corresponding target operators to the respective storage nodes. If stored on the same storage node, due to the limited computing power of the storage node, it may be necessary to execute data processing tasks in batches. This means sending the same or different target operators in batches to the same storage node for data processing. Alternatively, the compute node and storage node can execute data processing tasks for different target attributes separately. For example, the timelines corresponding to target attribute 1 and target attribute 2 are both stored on the same storage node. After comparison, it's found that the amount of time series data corresponding to target attribute 1 is less than the amount of time series data corresponding to target attribute 2. Therefore, the target operator corresponding to target attribute 2 can be pushed down to the storage node, and the storage node can then transmit the time series data corresponding to target attribute 1 to the compute node. Thus, while meeting the needs of parallel computing and improving data processing efficiency, it can minimize the network bandwidth resource consumption of data transmission.

[0080] In one or more embodiments of this application, at least one of the storage nodes receives a timing prediction result or anomaly detection result based on the data processing request.

[0081] In practical applications, depending on the type of data processing request, corresponding processing results are returned. These results can be time-series prediction results or anomaly detection results. When storage nodes send data processing results to compute nodes, the amount of data that needs to be sent is far less than the amount of time-series data involved in the data processing, thus meeting data processing requirements while reducing data transmission volume. It should be noted that the data processing method is not limited to time-series prediction results or anomaly detection results mentioned above; data processing requests can also be time-series data analysis, etc., and the corresponding feedback results will be time-series analysis results, etc.

[0082] Based on the same idea, embodiments of this application also provide a data processing method. For example... Figure 2 This is a flowchart illustrating another data processing method provided in an embodiment of this application. The method is applied to a storage node, from... Figure 2 The method described herein includes the following steps:

[0083] 201: Receive the target operator in the machine learning model pushed down by the computing node for executing the data processing request; the computing node is used to determine the corresponding target timeline and the storage node storing the target timeline according to the time sequence identifier of the target attribute corresponding to the data processing request.

[0084] 202: Determine the target time series data to satisfy the data processing request based on the target timeline.

[0085] 203: Perform data processing on the target time series data using the target operator, and send the data processing results to the computing node.

[0086] In practical applications, one compute node corresponds to multiple storage nodes. Machine learning models are embedded in a database, where time-series data is stored chronologically. Generally, time-series data from the same timeline are stored on the same storage node. When a data processing task needs to be performed using the machine learning model, the corresponding timeline and the storage node storing that timeline are determined based on the data processing request. The target operators capable of executing the data processing request are then pushed down to the corresponding storage nodes to perform the relevant data processing task. This approach satisfies the data processing request while avoiding sending large amounts of data from storage nodes to compute nodes, reducing network bandwidth consumption. Furthermore, sending operators from the machine learning model specifically to storage nodes effectively improves data processing efficiency. Additionally, if data is sent to multiple storage nodes, these nodes can execute data processing tasks in parallel, resulting in higher data processing efficiency compared to having only one compute node execute multiple tasks.

[0087] Based on the same idea, embodiments of this application also provide a data processing apparatus. For example... Figure 3 This is a schematic diagram of a data processing apparatus provided in an embodiment of this application. Figure 3 As can be seen from the image, the device includes:

[0088] The first determining module 31 is used to determine the target operator of the machine learning model built into the database in response to a data processing request for the target attribute.

[0089] The second determining module 32 is used to determine the target timeline corresponding to the timeline identifier and the storage node storing the target timeline based on the timeline identifier corresponding to the target attribute.

[0090] The push-down module 33 is used to push the target operator down to the storage node so that the storage node can use the target operator to execute the data processing request for the target time series data in the target timeline.

[0091] Optionally, the second determining module 32 is used to determine the time sequence identifier corresponding to the target attribute, and to determine the corresponding target timeline based on the time sequence identifier;

[0092] Based on the time sequence identifier, determine the routing table corresponding to the target timeline;

[0093] Based on the routing information stored in the routing table, determine the storage node corresponding to the target timeline and the number of storage nodes.

[0094] Optionally, the push-down module 33 is used to push the target operator down to the storage node if the storage node corresponding to the target timeline is found to be a unique node.

[0095] Optionally, it also includes a receiving module 34, which is used to receive the target time series data from the multiple storage nodes corresponding to the same target timeline if multiple storage nodes are found to correspond to the target timeline.

[0096] Optionally, the first determining module 31 is configured to, in response to a data processing request for a target attribute, determine the target operator type and the target operator required by the data processing request;

[0097] Alternatively, the target operator can be determined based on the function name carried in the data processing request.

[0098] Optionally, the second determining module 32 is used to determine the routing information corresponding to the time range specified in the data processing request in the routing table;

[0099] Based on the routing information, the storage nodes corresponding to the target timeline and the number of storage nodes are determined.

[0100] Optionally, it also includes a training module 35 for pushing down the target operator in the machine learning model to be trained to the storage node;

[0101] The machine learning model to be trained is trained using the target time-series data stored in the timeline of the storage node.

[0102] The second determining module 32 is used to determine the target timeline corresponding to each of the time series identifiers based on the time series identifiers corresponding to at least one of the target attributes.

[0103] Determine the storage nodes corresponding to each of the target timelines;

[0104] The step of pushing the target operator down to the storage node includes: pushing the target operator down to each of the storage nodes.

[0105] Optionally, the receiving module 34 is configured to receive timing prediction results or anomaly detection results fed back by at least one of the storage nodes based on the data processing request.

[0106] Based on the same idea, embodiments of this application also provide another data processing apparatus. For example... Figure 4 This is a schematic diagram of another data processing apparatus provided in an embodiment of this application. Figure 4 As can be seen from the image, the device includes:

[0107] The receiving module 41 is used to receive the target operator in the machine learning model pushed down by the computing node for executing the data processing request; the computing node is used to determine the corresponding target timeline and the storage node storing the target timeline according to the time sequence identifier of the target attribute corresponding to the data processing request.

[0108] The determination module 42 is used to determine the target time series data to satisfy the data processing request based on the target timeline.

[0109] The sending module 43 is used to perform data processing on the target time series data using the target operator and send the data processing result to the computing node.

[0110] Based on the same idea, embodiments of this application also provide a data processing system. For example... Figure 5 This is a schematic diagram of the system structure provided for an embodiment of this application. Figure 5 As can be seen, the system includes:

[0111] Computation node 51 corresponds to at least one storage node 52. Computation node 51 can be divided into two functional layers: an SQL layer to support user-provided SQL language and a data processing layer to perform data processing tasks. The data processing layer contains built-in machine learning model operators (e.g., time series prediction operators or anomaly detection operators). Storage node 52 can be divided into a data processing layer and a data storage layer. The data processing layer contains built-in target operators that are pushed down (e.g., time series prediction operators or anomaly detection operators), and the data storage layer stores time-series data chronologically. Figure 5 As can be seen, timeline 1 and timeline 2 are stored in storage node 1, and timeline 3 is stored in storage node 2.

[0112] For example, suppose the data processing request is for time series prediction or anomaly detection. The data processing layer is primarily responsible for accessing and processing time series data. It has the capability to pull data from the data storage layer along the timeline dimension and perform timeline-related processing (such as downsampling and some aggregation calculations), and supports extending processing operators for the timeline dimension. The data processing layer exists on both compute nodes and storage nodes. Therefore, some existing database time series engines can push some time series data processing operators from compute nodes to storage nodes when conditions permit. On the one hand, this allows for parallel computation optimization using multiple storage nodes; on the other hand, it enables optimization by placing computation closer to the data, thereby reducing data transfer between nodes and improving computational efficiency.

[0113] Both time series prediction and time series anomaly detection are algorithms that process data along the timeline dimension. Time series prediction predicts a future continuous window based on a segment of historical window data on the timeline. Similarly, time series anomaly detection requires detecting newly generated data based on historical data on the timeline. In implementation, data is usually processed in batches along the timeline dimension to improve detection efficiency.

[0114] This solution supports inference functionality for time-series prediction / anomaly detection algorithms based on machine learning. On the one hand, it leverages the time-series engine's characteristics in processing time-line dimensions to achieve batch processing optimization. On the other hand, it utilizes the time-series engine's data processing pushdown capability to push prediction / anomaly detection algorithms down to storage nodes, thereby achieving parallel and data-centric inference optimization.

[0115] One embodiment of this application also provides an electronic device. This electronic device is a master node electronic device in a computing unit. For example... Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of this application. The electronic device includes a memory 601, a processor 602, and a communication component 603; wherein,

[0116] The memory 601 is used to store programs;

[0117] The processor 602, coupled to the memory, is configured to execute the program stored in the memory for:

[0118] In response to a data processing request for a target attribute, determine the target operator of the machine learning model built into the database;

[0119] Based on the time sequence identifier corresponding to the target attribute, determine the target timeline corresponding to the time sequence identifier, and the storage node for storing the target timeline;

[0120] The target operator is pushed down to the storage node so that the storage node can use the target operator to execute the data processing request for the target time series data in the target timeline.

[0121] The processor 602 is used to determine the timing identifier corresponding to the target attribute, and to determine the corresponding target timeline based on the timing identifier;

[0122] Based on the time sequence identifier, determine the routing table corresponding to the target timeline;

[0123] Based on the routing information stored in the routing table, determine the storage node corresponding to the target timeline and the number of storage nodes.

[0124] The processor 602 is used to push the target operator down to the storage node if the storage node corresponding to the target timeline is found to be a unique node.

[0125] The processor 602 is configured to receive the target time-series data from the multiple storage nodes corresponding to the same target time-series if multiple storage nodes are found to correspond to the target time-series.

[0126] Processor 602 is configured to, in response to a data processing request for a target attribute, determine the target operator type and the target operator required by the data processing request;

[0127] Alternatively, the target operator can be determined based on the function name carried in the data processing request.

[0128] The processor 602 is used to determine the routing information corresponding to the time range specified in the data processing request in the routing table;

[0129] Based on the routing information, the storage nodes corresponding to the target timeline and the number of storage nodes are determined.

[0130] The processor 602 is used to push down the target operator in the machine learning model to be trained to the storage node;

[0131] The machine learning model to be trained is trained using the target time-series data stored in the timeline of the storage node.

[0132] Processor 602 is configured to determine a target timeline corresponding to each of the timeline identifiers based on a timeline identifier corresponding to at least one of the target attributes.

[0133] Determine the storage nodes corresponding to each of the target timelines;

[0134] The step of pushing the target operator down to the storage node includes: pushing the target operator down to each of the storage nodes.

[0135] The processor 602 is used to receive timing prediction results or anomaly detection results fed back by at least one of the storage nodes based on the data processing request.

[0136] The processor 602 is used to receive the target operator in the machine learning model pushed down by the computing node for executing the data processing request; the computing node is used to determine the corresponding target timeline and the storage node storing the target timeline according to the time sequence identifier of the target attribute corresponding to the data processing request.

[0137] Based on the target timeline, determine the target time-series data used to satisfy the data processing request;

[0138] The target operator is used to perform data processing on the target time series data, and the data processing results are sent to the computing node.

[0139] The aforementioned memory 601 can be configured to store various other data to support operation on the electronic device. Examples of such data include instructions for any application or method used to operate on the electronic device. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0140] Furthermore, the processor 602 in this embodiment may specifically be a programmable switching processing chip, which is configured with a data copying engine and can copy the received data.

[0141] When the processor 602 executes the program in memory, in addition to the functions described above, it can also perform other functions, as detailed in the descriptions of the preceding embodiments. Furthermore, as... Figure 6 As shown, the electronic device also includes other components such as the power supply component 604.

[0142] This application also provides a non-transitory machine-readable storage medium storing executable code. When the executable code is executed by a processor of an electronic device, the processor performs... Figure 1 or Figure 2 The method described in the corresponding embodiment.

[0143] Based on the above embodiments, some databases store time-series data according to a timeline for easier data management. In some database systems, machine learning models are used to meet their data processing needs. Therefore, in this solution, based on the data processing request for the target attribute, the target operator of the machine learning model built into the database is sent to the storage node containing the timeline corresponding to the target attribute, without needing to transfer the target time-series data from the storage node to the computing node. This effectively reduces the amount of data transfer between nodes and improves computational efficiency. Furthermore, when processing time-series data using the machine learning model, it is only necessary to select the corresponding target operator according to the data processing request and push that target operator down to the corresponding storage node, without needing to transfer an overly complex machine learning model to meet the data processing requirements. Moreover, by pushing the target operators from the same machine learning model down to different storage nodes and having different storage nodes execute the corresponding data processing tasks, parallel processing across multiple storage nodes is achieved, effectively improving data processing efficiency and capabilities.

[0144] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0145] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0146] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A data processing method applied to a database comprising computing nodes and corresponding storage nodes, the method being specifically applied to the computing nodes, the method comprising: In response to a data processing request for a target attribute, determine the target operator of the machine learning model built into the database; Based on the time sequence identifier corresponding to the target attribute, determine the target timeline corresponding to the time sequence identifier, and the storage node for storing the target timeline; If the target timeline stored in the same storage node meets the requirement of executing the target operator in the same storage node, the target operator is pushed down to the storage node so that the storage node can use the target operator to execute the data processing request for the target time series data in the target timeline.

2. The method according to claim 1, wherein determining the target timeline corresponding to the timeline identifier based on the timeline identifier corresponding to the target attribute, and the storage node for storing the target timeline, comprises: Determine the time sequence identifier corresponding to the target attribute, and determine the corresponding target timeline based on the time sequence identifier; Based on the time sequence identifier, determine the routing table corresponding to the target timeline; Based on the routing information stored in the routing table, determine the storage node corresponding to the target timeline and the number of storage nodes.

3. The method according to claim 2, wherein if the target timeline stored in the same storage node satisfies the requirement of executing the target operator in the same storage node, the target operator is pushed down to the storage node, comprising: If the storage node corresponding to the target timeline is found to be unique, then the target operator is pushed down to the storage node.

4. The method according to claim 2, further comprising: If multiple storage nodes are found to correspond to the target timeline, the target time series data from the multiple storage nodes corresponding to the same target timeline is received.

5. The method according to claim 1, wherein determining the target operator of the machine learning model embedded in the database in response to a data processing request for a target attribute includes: In response to a data processing request for a target attribute, determine the target operator type and the target operator required for the data processing request; Alternatively, the target operator can be determined based on the function name carried in the data processing request.

6. The method according to claim 2, wherein determining the storage node corresponding to the target timeline and the number of storage nodes based on the routing information stored in the routing table includes: Based on the time range specified in the data processing request, determine the routing information corresponding to the time range in the routing table; Based on the routing information, the storage nodes corresponding to the target timeline and the number of storage nodes are determined.

7. The method according to claim 1, wherein the training method of the machine learning model includes: Push down the target operator in the machine learning model to be trained to the storage node; The machine learning model to be trained is trained using the target time-series data stored in the timeline of the storage node.

8. The method according to claim 1, wherein determining the target timeline corresponding to the timeline identifier based on the timeline identifier corresponding to the target attribute, and the storage node for storing the target timeline, comprises: Based on the time sequence identifier corresponding to at least one of the target attributes, determine the target timeline corresponding to each of the time sequence identifiers; Determine the storage nodes corresponding to each of the target timelines; The step of pushing the target operator down to the storage node includes: pushing the target operator down to each of the storage nodes.

9. The method according to claim 1, further comprising: Receive timing prediction results or anomaly detection results from at least one of the storage nodes based on the data processing request.

10. A data processing method applied to a database comprising computing nodes and storage nodes corresponding to the computing nodes, the method being specifically applied to the storage nodes, the method comprising: The target operator in the machine learning model that is pushed down by the computing node to execute the data processing request; The computing node is used to determine the corresponding target timeline and the storage node storing the target timeline based on the time sequence identifier of the target attribute corresponding to the data processing request. If the target timeline stored in the same storage node meets the requirement of executing the target operator in the same storage node, the target operator is pushed down to the storage node so that the storage node can use the target operator to execute the data processing request for the target time series data in the target timeline. Based on the target timeline, determine the target time-series data used to satisfy the data processing request; The target operator is used to perform data processing on the target time series data, and the data processing results are sent to the computing node.

11. A data processing apparatus, characterized in that, An apparatus for use with a database comprising compute nodes and corresponding storage nodes, the apparatus being specifically applied to the compute nodes, the apparatus comprising: The first determining module is used to determine the target operator of the machine learning model built into the database in response to a data processing request for the target attribute. The second determining module is used to determine the target timeline corresponding to the timeline identifier and the storage node storing the target timeline based on the timeline identifier corresponding to the target attribute. The push-down module is used to push down the target operator to the storage node if the target timeline stored in the same storage node meets the requirement of executing the target operator in the same storage node, so that the storage node can use the target operator to execute the data processing request for the target time series data in the target timeline.

12. A data processing system, the system comprising: A computing node for performing the method according to any one of claims 1 to 9; A storage node for performing the method of claim 10.

13. An electronic device, comprising a memory and a processor; wherein, The memory is used to store programs; The processor, coupled to the memory, is configured to execute the program stored in the memory to implement the method of any one of claims 1 to 9, or to implement the method of claim 10.

14. A non-transitory machine-readable storage medium storing executable code that, when executed by a processor of an electronic device, causes the processor to perform the method as claimed in any one of claims 1 to 9, or to perform the method as claimed in claim 10.

Citation Information

Patent Citations

  • Distributed machine learning system, model training method, node equipment and medium

    CN112508067A