Data stream processing method, apparatus, device, and computer storage medium

CN117850877BActive Publication Date: 2026-09-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211204736.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2026-09-25
Estimated Expiration
2042-09-29

AI Technical Summary

Technical Problem

[0004]本申请实施例提供一种数据流处理方法、装置、设备及计算机存储介质,用于解决超自然数据和延迟数据所带来的数据统计不准确的技术问题

Benefits of technology

[0044]本申请实施例中,接收实时数据流,并针对实时数据流中的待处理数据,会根据待处理数据的数据产生时刻与数据到达时刻,来确定待处理数据的目标处理类型,并基于数据产生时刻以及目标处理类型,获得为待处理数据分配窗口所需的分配基准时刻,从而基于该分配基准时刻,来确定待处理数据所属的目标时间窗口的目标窗口标识,并基于目标窗口标识以及待处理数据的数据标签,确定待处理数据的数据聚合标识,最终将待处理数据发送至数据聚合标识对应的计算设备,使得计算设备基于待处理数据,更新数据标签相对于目标时间窗口的关联数据。也就是说,本申请实施例通过结合待处理数据的目标处理类型,来获得为待处理数据分配窗口所需的分配基准时刻,而不是直接根据数据产生时刻来进行分配,有效的避免超自然数据和延迟数据无法正确分配窗口的问题,并且在本申请实施例中,也无需等待每个窗口中的数据全部到达后进行计算,而是将待处理数据的数据标签和窗口标识进行组合来得到数据聚合标识,并发送该数据聚合标识对应的计算设备进行处理,则相同数据聚合标识的待处理数据都会被发往同一计算设备,也就是同一窗口内相同数据标签的待处理数据只会通过相同的计算设备来进行计算,不管数据是否延迟到达,还是出现了超自然数据,计算设备收到数据即可进行计算,无需等待同一窗口的数据全部到达,提升了数据处理效率,并且不会对数据的处理过程造成影响。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117850877B_ABST
    Figure CN117850877B_ABST
Patent Text Reader

Abstract

The application discloses a data stream processing method and device, equipment and computer storage medium, and relates to the technical field of data stream processing. The method obtains an allocation reference moment required for allocating a window for to-be-processed data by combining a target processing type of the to-be-processed data, effectively avoids the problem that supernatural data and delayed data cannot be correctly allocated a window, and does not need to wait for all data in each window to arrive before performing calculation. Instead, a data label of the to-be-processed data and a window identifier are combined to obtain a data aggregation identifier, and a corresponding calculation device of the data aggregation identifier is sent for processing. To-be-processed data with the same data label in the same window is only calculated by the same calculation device, regardless of whether the data arrives with delay or supernatural data appears. The calculation device does not need to wait for all data in the same window to arrive, the data processing efficiency is improved, and the processing process of the data is not affected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more particularly to the field of data stream processing technology, providing a data stream processing method, apparatus, device, and computer storage medium. Background Technology

[0002] Real-time stream computing, also known as streaming computation, refers to a computing method where data streams are processed and used immediately upon arrival, with only a very small portion of the data being persistently stored and the majority being discarded. In real-time stream computing, statistical calculations are performed on a continuous stream of data within a window, each window having a time range. However, in practice, situations such as out-of-time data and delayed data may occur. Out-of-time data is caused by inaccurate server clocks, resulting in the event time of the data being in the future, exceeding the latest time of the current computing device. Delayed data refers to data arriving beyond the latency tolerance limit.

[0003] Supernatural data occurs before the actual window, causing the window to close prematurely. Subsequent data that originally belonged to that window will not be counted because the window has already closed and the corresponding window cannot be found. Similarly, delayed data arrives too late and cannot be counted because the window has already closed. Both of these situations lead to inaccurate data statistics. Summary of the Invention

[0004] This application provides a data stream processing method, apparatus, device, and computer storage medium to solve the technical problem of inaccurate data statistics caused by supernatural data and delayed data.

[0005] On the one hand, a data stream processing method is provided, the method comprising:

[0006] Receive real-time data streams and determine the target processing type of the data to be processed based on the data generation time and data arrival time corresponding to the data to be processed in the real-time data streams;

[0007] Based on the data generation time and the target processing type, obtain the allocation reference time required to allocate a window for the data to be processed;

[0008] Based on the allocation reference time, determine the target window identifier of the target time window to which the data to be processed belongs;

[0009] Based on the target window identifier and the data tag of the data to be processed, the data aggregation identifier of the data to be processed is determined;

[0010] The data to be processed is sent to the computing device corresponding to the data aggregation identifier, so that the computing device updates the associated data of the data tag relative to the target time window based on the data to be processed.

[0011] On one hand, a data stream processing apparatus is provided, the apparatus comprising:

[0012] A receiving and processing unit is used to receive a real-time data stream and determine the target processing type of the data to be processed based on the data generation time and data arrival time corresponding to the data to be processed in the real-time data stream.

[0013] The allocation benchmark determination unit is used to obtain the allocation benchmark time required to allocate a window for the data to be processed based on the data generation time and the target processing type;

[0014] A window allocation unit is used to determine the target window identifier of the target time window to which the data to be processed belongs, based on the allocation reference time.

[0015] An aggregation determination unit is used to determine the data aggregation identifier of the data to be processed based on the target window identifier and the data tag of the data to be processed;

[0016] The sending unit is used to send the data to be processed to the computing device corresponding to the data aggregation identifier, so that the computing device updates the associated data of the data tag relative to the target time window based on the data to be processed.

[0017] In one possible implementation, the receiving and processing unit is specifically used for:

[0018] If the time when the data is generated is later than the time when the data arrives, then the target data type is supernatural data;

[0019] If the data generation time is earlier than the data arrival time, and the difference between the data generation time and the data arrival time is greater than a preset duration threshold, then the target data type is delayed data.

[0020] If the data generation time is earlier than the data arrival time, and the difference between the data generation time and the data arrival time is less than a preset duration threshold, then the target data type is normal data.

[0021] In one possible implementation, the allocation reference determination unit is specifically used for:

[0022] If the target data type is delayed data or normal data, then the data generation time is determined as the allocation reference time;

[0023] If the target data type is supernatural data, then the time of data generation is corrected, and the corrected time of data generation is determined as the allocation reference time.

[0024] In one possible implementation, the allocation reference determination unit is specifically used for:

[0025] The current time is determined as the time when the corrected data was generated; or...

[0026] The arrival time of the data is determined as the generation time of the corrected data.

[0027] In one possible implementation, the allocation reference determination unit is specifically used for:

[0028] Obtain the time difference between the data production equipment and the current equipment corresponding to the data to be processed;

[0029] Based on the device time difference, the time of data generation is corrected to obtain the corrected time of data generation.

[0030] In one possible implementation, the window allocation unit is specifically used for:

[0031] Based on the service type to which the real-time data stream belongs, determine the size of the time window and the data stream processing cycle corresponding to the real-time data stream;

[0032] Determine the time difference between the allocation reference time and the start time of the current data stream processing cycle, and the method of taking the value of the time difference is matched with the service type;

[0033] The target window identifier is determined based on the time difference and the window size.

[0034] In one possible implementation, the aggregation determining unit is specifically used for:

[0035] Based on the current data stream processing cycle and the allocation reference time, the cycle identifier corresponding to the data to be processed is determined, and each cycle identifier uniquely corresponds to a data stream processing cycle.

[0036] The target window identifier, the period identifier, and the data label are concatenated to obtain the data aggregation identifier.

[0037] In one possible implementation, the device further includes a grouping unit for:

[0038] Based on the data aggregation identifiers of the data included in the target time window, the data are grouped and a corresponding computing device is allocated to each group; wherein, each group includes at least one data with the same data aggregation identifier;

[0039] The sending unit is specifically used for:

[0040] Based on the data aggregation identifier of the data to be processed, the data to be processed is divided into corresponding groups, and the data to be processed is sent to the computing device corresponding to the group.

[0041] On one hand, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any of the above methods.

[0042] On the one hand, a computer storage medium is provided that stores computer program instructions thereon, which, when executed by a processor, implement the steps of any of the above methods.

[0043] On one hand, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of any of the methods described above.

[0044] In this embodiment, a real-time data stream is received, and for the data to be processed in the real-time data stream, the target processing type of the data to be processed is determined according to the data generation time and data arrival time of the data to be processed. Based on the data generation time and the target processing type, the allocation reference time required to allocate a window for the data to be processed is obtained. Based on the allocation reference time, the target window identifier of the target time window to which the data to be processed belongs is determined. Based on the target window identifier and the data tag of the data to be processed, the data aggregation identifier of the data to be processed is determined. Finally, the data to be processed is sent to the computing device corresponding to the data aggregation identifier, so that the computing device updates the associated data of the data tag relative to the target time window based on the data to be processed. In other words, this application embodiment obtains the allocation reference time required to allocate windows for the data to be processed by combining the target processing type of the data to be processed, rather than directly allocating based on the data generation time. This effectively avoids the problem of incorrect window allocation for supernatural data and delayed data. Furthermore, in this application embodiment, it is not necessary to wait for all the data in each window to arrive before calculation. Instead, the data label and window identifier of the data to be processed are combined to obtain a data aggregation identifier, which is then sent to the computing device corresponding to the data aggregation identifier for processing. Thus, data to be processed with the same data aggregation identifier will be sent to the same computing device. That is, data to be processed with the same data label within the same window will only be calculated by the same computing device, regardless of whether the data arrives late or supernatural data occurs. The computing device can start calculation as soon as it receives the data, without waiting for all the data in the same window to arrive, thereby improving data processing efficiency and not affecting the data processing process. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0046] Figure 1 This is a schematic diagram of a real-time window calculation scheme in related technologies;

[0047] Figure 2 This is a schematic diagram of an application scenario provided by an embodiment of this application;

[0048] Figure 3 A flowchart illustrating the data stream processing method provided in an embodiment of this application;

[0049] Figure 4 This is a flowchart illustrating the process of determining the processing type of the data to be processed, provided in an embodiment of this application.

[0050] Figure 5 A schematic diagram illustrating the determination of a time window as provided in an embodiment of this application;

[0051] Figure 6 Example diagram of the window representation corresponding to the allocation reference time provided in the embodiments of this application;

[0052] Figure 7 A schematic diagram illustrating a combination of data aggregation identifiers provided in an embodiment of this application;

[0053] Figure 8 A schematic diagram illustrating another combination of data aggregation identifiers provided in an embodiment of this application;

[0054] Figures 9a to 9d A schematic diagram illustrating the allocation of time windows for data to be processed in an embodiment of this application;

[0055] Figure 10 A schematic diagram of a data stream processing apparatus provided in an embodiment of this application;

[0056] Figure 11 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0058] It is understood that, in the following specific embodiments of this application, the data involved requires relevant licenses or consents when the embodiments of this application are applied to specific products or technologies, and the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0059] To facilitate understanding of the technical solutions provided in the embodiments of this application, some key terms used in the embodiments of this application will be explained below:

[0060] Watermark: In related technologies, a timestamp is used to mark each piece of data in a real-time data stream, indicating that all data before the watermark time has arrived. Watermarking ensures the resolution of out-of-order issues, because in stream processing, it is impossible to wait for data to arrive without performing any operations on it, especially for operations such as aggregation.

[0061] EventTime: This refers to the time when the data was generated. This time is usually generated by the data producer based on the server time. Each piece of data in the real-time data stream carries a timestamp representing the time when the data was generated.

[0062] A window is the smallest unit of time for statistical data in real-time stream computing. A window is a time range defined by the start and end times. In practical applications, the size of a window can be determined based on the actual business requirements.

[0063] EndTime: Each Window is a time range consisting of a start time and an end time, and it includes the start time but excludes the end time. The EndTime of a Window is the end time of the Window.

[0064] Maximum EventTime: Each piece of data has an EventTime. The maximum EventTime refers to the maximum value of EventTime in the real-time data stream within a certain period of time. The later the time represented by EventTime, the larger the EventTime will be.

[0065] Data tag: also known as data calculation key, is the smallest granularity used for calculation and statistics. For example, a topic, an article ID, or an article ID plus channel can be used as a calculation key.

[0066] Window identifier: Divide a time period into fixed window sizes with no overlap or gaps. Take the end time of each interval as a window time. Each window corresponds to a different time period. The window identifier is used to uniquely identify a window. For example, you can take the EndTime of each window and calculate a unique identifier representing the current window based on the algorithm.

[0067] The design concept of the embodiments of this application will be briefly introduced below.

[0068] In real-time stream computing scenarios, due to limitations of actual conditions, supernatural data and delayed data will inevitably appear.

[0069] In related technologies, real-time window calculation schemes are commonly used. Real-time window calculation is implemented through the window functions built into the real-time stream computing component. Each window has a time range, and its execution is triggered by a timer. The timer executes when the current latest time is greater than the timer's EndTime, at which point the calculation for that window is triggered. The current latest time is calculated using a Watermark. The Watermark calculation process can be understood as adding a marker to the real-time data stream. When the marker is encountered, the Watermark is triggered to calculate the current latest time. The calculation formula is as follows:

[0070] Watermark time = Maximum EventTime – Delay duration

[0071] The delay duration is the approximate delay duration of the data estimated based on the overall data arrival delay duration. The calculation of the current window is triggered when the Watermark time is greater than or equal to the EndTime of the current window.

[0072] See Figure 1 The diagram shows a real-time window calculation scheme in related technologies. w(4) and w(9) are Watermark identifiers. When these identifiers are encountered, a window calculation will be triggered. Taking w(4) calculation as an example, the maximum EventTime in the period before this is 7, and the delay time is, for example, 3. Then w(4) = 7 - 3 = 4, which means that the latest time is 4. The window before this time has reached its time, that is, it has reached the EndTime of the first window, so the window calculation will be triggered, that is, the calculation of the windows corresponding to T1 to T4 on the right is triggered.

[0073] However, due to the differences in clocks across machines, the actual time of each data entry can be abnormal, such as unpredictable timing. This can cause the Watermark calculated using the above method to close prematurely, preventing subsequent data that should belong to that window from participating in the calculation and resulting in an underestimation of the window's value. Furthermore, if the data latency is too high, the data will be discarded since the window has already closed, also leading to inaccurate calculations. Additionally, when the time window is large, the inability to close the window promptly results in a large number of windows remaining in memory. Each window corresponds to a timer, leading to a large number of timers in memory, which consumes significant memory resources. As the number of windows increases, the stability of the device will continuously decrease.

[0074] In view of this, this application provides a data stream processing method. In this method, the allocation reference time required for allocating windows to the data to be processed is obtained by combining the target processing type of the data to be processed, instead of directly allocating based on the data generation time. This effectively avoids the problem of incorrect window allocation for superfluous or delayed data. Furthermore, in this application embodiment, it is not necessary to wait for all data in each window to arrive before calculation. Instead, the data label and window identifier of the data to be processed are combined to obtain a data aggregation identifier, which is then sent to the computing device corresponding to the data aggregation identifier for processing. Thus, data to be processed with the same data aggregation identifier will be sent to the same computing device. That is, data to be processed with the same data label within the same window will only be calculated through the same computing device, regardless of whether the data arrives late or superfluous data occurs. The computing device can start calculation as soon as it receives the data, without waiting for all data in the same window to arrive, improving data processing efficiency and without affecting the data processing process. This application embodiment, through the design of the data aggregation identifier, abandons the original window calculation method and greatly improves the stability of calculation in the case of large windows.

[0075] In addition, considering that a large number of expired data aggregation identifiers will occur during the calculation process, which will seriously affect the calculation efficiency over time, this application embodiment introduces a strict expiration cleanup mechanism for data aggregation identifiers to prevent the problem of insufficient memory and excessive computing power as the number of data aggregation identifiers increases.

[0076] The following is a brief introduction to the application scenarios to which the technical solutions of the embodiments of this application are applicable. It should be noted that the application scenarios described below are only for illustrating the embodiments of this application and are not intended to limit the scope. In specific implementation, the technical solutions provided by the embodiments of this application can be flexibly applied according to actual needs.

[0077] The solutions provided in this application can be applied to scenarios involving real-time stream computing, such as real-time business decision-making, online feature engineering, rule engine early warning, and other scenarios with high requirements for data timeliness and accuracy, as well as online log analysis, online machine learning, online graph computing, and online recommendation algorithm applications. Figure 2 The diagram shown is an application scenario provided by an embodiment of this application. In this scenario, a data production device 101, a data stream processing device 102, and a computing cluster 103 may be included.

[0078] The data production device 101 collects data, generates a data stream, and sends it to the data stream processing device 102 for processing. The data production device 101 can be implemented using terminal devices such as mobile phones or personal computers (PCs), or it can be implemented using a server. Data collection can be performed on the data to be collected in the data production device 101 to collect the corresponding data, and a timestamp can be added to the collected data according to the time of the data production device 101, so that the time of data generation can be obtained during data stream processing. In practical applications, the data production device 101 can be any device in the data source system. For example, the data source system can be a website system, application system, etc. The data involved in these data source systems is usually stored on a server or a server-related storage device; therefore, the data production device 101 can refer to a server or a storage device.

[0079] In a real-world scenario, there can be multiple data production devices 101, and these multiple data production devices 101 can belong to the same data source system. In this case, the embodiments of this application can perform statistics on some data indicators for a specific data source system. Alternatively, the multiple data production devices 101 can also belong to different data source systems. In this case, the embodiments of this application can also perform statistics on some data indicators for each data source system separately.

[0080] Data stream processing device 102 and computing cluster 103 can both belong to the real-time stream computing system. In some cases, if the physical device has sufficient computing resources, data stream processing device 102 and computing cluster 103 can also be deployed in the same physical device. Data stream processing device 102 can receive real-time data streams sent by data production device 101, and through the steps of the data stream processing method provided in this application embodiment, distribute data with the same data label in the same time window to the same computing device in computing cluster 103 for computation.

[0081] The servers mentioned above can be, for example, independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms, but are not limited to these.

[0082] The data stream processing device 102 may include one or more processors, memory, and I / O interfaces for interacting with terminals. The memory of the data stream processing device 102 may also store program instructions for the data stream processing method provided in the embodiments of this application. When these program instructions are executed by the processor, they can be used to implement the steps of the data stream processing method provided in the embodiments of this application.

[0083] In one possible implementation, the technical solution of this application embodiment can be applied to a news scenario. The window calculation involved can be the number of views of a statistical object in each window of a news website. Taking an article as an example, in order to collect the number of views of each article in the news website, it is necessary to retrieve the corresponding browsing record (i.e., a piece of data in this application embodiment) from the backend of the news website and send it to the data stream processing device 102. For each browsing record, the data stream processing device will determine the target processing type of each browsing record based on the corresponding data production time, and determine the allocation reference time of the allocation window based on the data production time and the target processing type, thereby obtaining the corresponding target window identifier. The target window identifier is then assembled with a data tag (e.g., article ID) to obtain a data aggregation identifier. The browsing record is then sent to the corresponding computing device for processing, that is, the number of views of the corresponding article is calculated. The obtained statistical results can be used in downstream application scenarios.

[0084] In one possible implementation, the technical solution of this application embodiment can be applied to online feature engineering scenarios. The window calculation involved can be the feature values ​​of each feature dimension required for statistical model training or model forward calculation. In order to realize the calculation of feature values, after obtaining the real-time data stream from the data production device 101, for each piece of data, the target processing type to which it belongs can be determined according to the corresponding data production time, and the allocation reference time of the allocation window can be determined based on the data production time and the target processing type. Then, the corresponding target window identifier can be obtained, and the target window identifier can be assembled with the data label (e.g., feature dimension name) to obtain the data aggregation identifier. The browsing record is then sent to the corresponding computing device for processing, thereby calculating the feature value of the feature dimension.

[0085] The data production device 101, data stream processing device 102, and computing cluster 103 described above can be directly or indirectly connected via one or more networks. This network can be a wired network or a wireless network; for example, a wireless network can be a mobile cellular network or a Wireless-Fidelity (WIFI) network, or any other possible network. This embodiment of the invention does not limit the types of networks that can be used.

[0086] It should be noted that, in this embodiment of the application, there is no limitation on the number of the data production equipment 101, the data stream processing equipment 102, and the computing cluster 103.

[0087] In one possible application scenario, the real-time stream computing involved in this application embodiment can be implemented using cloud computing technology. Cloud computing is a computing model that distributes computing tasks across a resource pool composed of a large number of computers, enabling various application systems to obtain computing power, storage space, and information services as needed. The network providing these resources is called the "cloud." From the user's perspective, the resources in the "cloud" are infinitely scalable, readily available, on-demand, expandable, and pay-as-you-go. Cloud computing is a product of the convergence and development of traditional computer and network technologies such as grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balancing.

[0088] As a provider of fundamental cloud computing capabilities, a cloud resource pool (referred to as a cloud platform, generally called an IaaS (Infrastructure as a Service) platform) is established. Various types of virtual resources are deployed in the resource pool for external customers to choose from. The cloud resource pool mainly includes: computing devices (virtualized machines containing operating systems), storage devices, and network devices. For example, the data stream processing device 102 and computing cluster 103 mentioned above can be implemented using resources from the cloud platform.

[0089] Based on logical function, a PaaS (Platform as a Service) layer can be deployed on top of the IaaS (Infrastructure as a Service) layer, and a SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. Alternatively, SaaS can be deployed directly on top of IaaS. PaaS is a platform for running software, such as databases and web containers. SaaS refers to various types of business software, such as web portals and bulk SMS senders. Generally speaking, SaaS and PaaS are upper layers compared to IaaS.

[0090] Of course, the methods provided in the embodiments of this application are not limited to... Figure 2 The application scenarios shown can also be used in other possible scenarios, and this application embodiment does not impose any limitations. Figure 1 The functions that each device in the application scenario shown can achieve will be described in subsequent method embodiments, and will not be elaborated on here.

[0091] The method flows provided in the various embodiments of this application can be used... Figure 2 The data stream processing device 102 can be used to execute the operation, or it can be executed jointly by the data stream processing device 102 and the computing cluster 103. See also Figure 3 The diagram shown is a flowchart illustrating the data stream processing method provided in an embodiment of this application.

[0092] Step 301: Receive real-time data stream.

[0093] In this embodiment, to achieve data stream acquisition, data tracking can be performed on the data production equipment, or data acquisition software can be installed on the data production equipment. This allows the required data to be collected from the data production equipment and sent to the data stream processing equipment via a real-time data stream. The real-time data stream refers to the fact that the data production equipment is a continuous data source, constantly generating new data; correspondingly, data can be continuously collected and sent to the data stream processing equipment.

[0094] Specifically, after data collection, a timestamp can be added to the data to indicate the time the data was generated. This timestamp can be generated based on the time of the data production device or the time of the server corresponding to the data production device. In real-world scenarios, there can be multiple data production devices. Real-time data streams can then be obtained from these devices, and each data production device can be considered a separate real-time data stream. Alternatively, when multiple data production devices belong to the same data source system, their real-time data streams can use different stream identifiers for separate processing, or they can use the same stream identifier, thus being considered a single real-time data stream.

[0095] In this embodiment of the application, the real-time data stream can be collected according to the actual business needs of the scenario. For example, if it is necessary to count the amount of access data, then it is necessary to collect access record data; or if it is necessary to count the feature value of a certain feature dimension, then it is necessary to collect data related to that feature dimension. This embodiment of the application does not limit the data contained in the real-time data stream.

[0096] Step 302: Based on the data generation time and data arrival time of the data to be processed in the real-time data stream, determine the target processing type of the data to be processed.

[0097] It should be noted that in this embodiment, the real-time data stream continuously transmits data to be processed, and the processing procedure for each piece of data to be processed by the data stream processing device is similar. Therefore, the following description will mainly use a single piece of data to be processed as an example. In practical applications, the data to be processed in a real-time data stream can be stored in a message queue, and then each piece of data to be processed can be processed sequentially. That is, the data to be processed can be read serially from the message queue for subsequent processing. Alternatively, to improve processing efficiency, multiple threads can work together to process each piece of data to be processed in the message queue in parallel.

[0098] In this embodiment of the application, the purpose of classifying the data to be processed is mainly to provide targeted processing for different types of data, so that the overall steps are universal for all data to be processed, but different processing methods are used for different types of data to be processed in each step.

[0099] The data processing types involved can include at least the following:

[0100] (1) Supernatural data: Supernatural data arises from clock asynchrony between different data production devices, or between data production devices and data stream processing devices. For example, the time of the data production device may be later than that of the data stream processing device. Consequently, when the data arrives at the data stream processing device, its generation time is still later than the arrival time of the data in the data stream processing device. This type of data is called supernatural data. Therefore, for a given set of data to be processed, if its generation time is later than its arrival time, then the target data type of the data to be processed is supernatural data.

[0101] (2) Delayed data: In real-world scenarios, data undergoes multiple processing steps from upstream collection to downstream computation. Therefore, data may arrive at the downstream location with a delay. If the delay is small, it is normal. However, if the delay exceeds a certain tolerance level, such data is called delayed data. Therefore, for a given set of data to be processed, if the data generation time is earlier than the data arrival time, and the difference between the data generation time and the data arrival time is greater than a preset time threshold (i.e., exceeding the tolerance for delay), then the target data type for this set of data is delayed data.

[0102] (3) Normal data. Compared with the two types mentioned above, normal data refers to data whose generation time is not later than the arrival time of the data stream processing device, and whose delay time is within a tolerable range. Therefore, for a piece of data to be processed, if the data generation time is earlier than the data arrival time, and the difference between the data generation time and the data arrival time is less than the preset time threshold, then the target data type of the data to be processed is normal data.

[0103] Of course, other possible types may also be included, and this application embodiment does not limit this.

[0104] In one possible implementation, after receiving the real-time data stream, the data stream processing device can classify each piece of data to be processed by combining the data generation time and the data arrival time (i.e., the data arrival time), determining that each piece of data to be processed belongs to one of the aforementioned types, thereby enabling the corresponding processing for each type of data to be processed. However, it should be noted that in some cases, the data arrival time and the processing time are not necessarily the same. That is, after the data arrives, it can be buffered first, and the buffered data can be processed sequentially. In this case, after receiving the real-time data stream, the data stream processing device can add an arrival timestamp to each piece of data to be processed by combining the data stream processing device's clock to characterize the arrival time of each piece of data to be processed. Thus, when processing the data to be processed, the data processing type of each piece of data can be determined by combining the data generation time and the data arrival time.

[0105] See Figure 4 The diagram shown is a flowchart illustrating the process of determining the processing type of data to be processed according to an embodiment of this application, which includes the following steps:

[0106] Step 3021: Obtain the data generation time and data arrival time of the data to be processed.

[0107] Step 3022: Determine whether the data generation time is earlier than the data arrival time. If not, proceed to step 3023; otherwise, if yes, proceed to step 3024.

[0108] Step 3023: Determine the target processing type as supernatural data.

[0109] Step 3024: Determine the delay time of the data to be processed. The delay time is the time difference between the time the data is generated and the time the data arrives.

[0110] Step 3025: Determine whether the delayed arrival time is greater than the preset time threshold. If yes, proceed to step 3026; otherwise, proceed to step 3027.

[0111] Step 3026: Determine the target processing type as delayed data.

[0112] Step 3027: Determine the target processing type as normal data.

[0113] In this embodiment, the preset duration threshold refers to the tolerable delay duration, which can be set based on the delay of historical data. For example, the average delay duration of historical data to be processed can be statistically analyzed, as the data to be processed usually arrives within this delay duration, and therefore can be used as the preset duration threshold. Alternatively, a certain amount of redundancy can be added to this threshold. Furthermore, since different services have different latency requirements, a delay tolerance duration can be defined based on the service, and a preset duration threshold can be set accordingly. Of course, a combination of the above two methods can also be used, and this embodiment does not impose any limitations on this approach.

[0114] Step 303: Based on the data generation time and the target processing type, obtain the allocation reference time required to allocate a window for the data to be processed.

[0115] In this embodiment, the real-time window-based calculation method is abandoned. Instead of calculating whether the calculation time of a window has arrived based on the data in the data stream, the received data to be processed is divided into the correct window and distributed to the correct computing device for processing by combining the window identifier and the data label of the data. As long as the data to be processed belongs to this window and data label, it will be sent to the computing device. Therefore, the computing device can process the data to be processed as soon as it receives it, without waiting or considering whether the window has been triggered.

[0116] To achieve the above effect, the data to be processed must first be assigned to the correct time window. Generally speaking, a time window refers to a time range. In real-world scenarios, it is often necessary to statistically analyze the value of a certain data indicator within a time range. Therefore, the accurate allocation of data to its corresponding time window is a prerequisite for accurate statistical results.

[0117] Specifically, if the time of data generation is used as the basis for dividing the time window, then the time of data generation of the data to be processed needs to be as accurate as possible. Of course, in addition to the time of data generation, other times can also be used as the basis for dividing the time window, such as the time of data collection, the time of arrival, etc., which can be adjusted according to the actual application scenario. This application embodiment does not limit this.

[0118] Generally, data stream processing is based on the clock of the data stream processing device. Therefore, for data to be processed that is either normal or delayed, the data generation time is definitely before the current time of the data stream processing device, which is considered normal. Thus, the data generation time can be considered accurate and can be used as the basis for dividing the window. That is, when the data to be processed is normal or delayed, its generation time can be considered as the allocation reference time required for allocating a window to that data, resulting in the following formula:

[0119]

[0120] That is, allocating the reference time. The moment the data was generated

[0121] For data to be processed that is classified as supernatural data, if the data generation time is later than the data arrival time, or later than the current time of the data stream processing device, then there is obviously a time difference between the data production device and the data stream processing device. Therefore, it is necessary to correct the data generation time and use the corrected data generation time as the allocation reference time required for the allocation window of the data to be processed.

[0122] In one possible implementation, if the target data type is supernatural data, when correcting the time of its data generation, the current time of the data stream processing device can be determined as the corrected time of data generation. That is, the data stream processing device can obtain the corresponding time information from its own clock and use it as the corrected time of data generation. Alternatively, the data stream processing device can also determine the time of receiving the data to be processed, that is, the data arrival time mentioned above, as the corrected time of data generation.

[0123] In one possible implementation, if the target data type is supernatural data, when correcting the time of data generation, the time difference between devices can also be used to correct the time of data generation. That is, the data stream processing device can obtain the device time difference between the data production device corresponding to the data to be processed and the current device. Then, when correcting the time of data generation, the time of data generation can be corrected based on the device time difference to obtain the corrected time of data generation.

[0124] Specifically, data stream processing equipment can maintain the time difference between itself and different data production equipment, and then find the corresponding time difference based on the data production equipment identifier carried in the data to be processed. Alternatively, the data stream processing equipment can also measure the time difference between itself and the data production equipment corresponding to the data to be processed to obtain the corresponding time difference.

[0125] In practical applications, if the target data type is supernatural data, it can be used without calibration, and the time of its data generation can be directly used as the allocation reference time for window allocation; or, the deviation of its data generation time can be evaluated. If the deviation from the correct time is not high, the data generation time can be directly used as the allocation reference time. If the deviation is high, it can be corrected and the corrected data generation time can be used as the allocation reference time.

[0126] Step 304: Based on the allocation reference time, determine the target window identifier of the target time window to which the data to be processed belongs.

[0127] In this embodiment of the application, the allocation reference time is used to allocate a corresponding time window for the data to be processed. After obtaining the allocation reference time, the target time window to which the data to be processed belongs can be determined sequentially.

[0128] In one possible implementation, once the size of each time window is determined, the time range corresponding to each time window is also determined. Then, the allocation reference time can be matched with the time range of each time window. If the allocation reference time is within the time range, then the time window is the target time window to which the data to be processed belongs. In this way, the window identifier corresponding to the determined target time window is determined as the target window identifier.

[0129] See Figure 5 The diagram illustrates the determination of a time window. Dividing the window into four parts, with a window size of 4, results in multiple time windows, as shown below. Figure 5 The time range of window1 shown is 1 to 4, the time range of window1 is 4 to 8, and so on. The allocation reference time of the data to be processed is 2. After matching, it can be found that it should belong to window1. Thus, the data to be processed is allocated to window1. The target window identifier of the data to be processed is window1.

[0130] It should be noted that the above window identifier examples are for illustrative purposes only and do not limit the method of window identifiers. In actual scenarios, other forms of window identifiers can also be used, as long as they can distinguish different time windows.

[0131] In one possible implementation, considering that the allocation reference time essentially represents time, while the time window represents a time range, the time windows can be numbered according to their temporal order. Then, the target window identifier of the target time window to which the data to be processed belongs can be calculated using a certain function formula.

[0132] Specifically, the size of the time window is usually determined based on the business type of the real-time data stream. For example, if the business requirement is to count hourly impressions, the window size can be 1 hour, or 1 hour can be divided into multiple parts, such as a 10-minute window. Alternatively, if the business requirement is to count daily pageviews, the window size can be 1 day, or 1 day can be divided into multiple parts, such as a 10-minute window. Furthermore, the data stream processing cycle can be set according to the business type. The data stream processing cycle can typically be set according to calendar days, weeks, months, etc., or it can be set to other values ​​according to specific needs.

[0133] Furthermore, for a real-time data stream, the size of the corresponding time window and the data stream processing cycle can be determined based on the business type to which the real-time data stream belongs. This allows us to determine the time difference between the allocated baseline time and the start time of the current data stream processing cycle. The value of this time difference is matched to the business type. For example, when the business requirement is to perform hourly statistics, the time difference can be expressed in minutes. Alternatively, it can be expressed in minutes or hours. For instance, if the allocated baseline time is 1:00, and the data stream processing cycle is 1 day with a start time of 0:00, the time difference would be 60 minutes if expressed in minutes, or 1 hour if expressed in hours.

[0134] Furthermore, the target window identifier can be determined based on the time difference and window size. One possible calculation formula is as follows:

[0135]

[0136] Among them, L i The target window identifier for the data to be processed; T i This represents the time difference based on the allocated baseline time, such as the number of minutes in days as shown above; 'a' represents the window size. Indicates to The result is then rounded down.

[0137] See Figure 6 The diagram shows examples of window representations corresponding to several allocation reference times. Here, a window size of 5 minutes is used as an example.

[0138] See also Figure 6 For data with an allocation reference time of "00:02", the corresponding number of minutes in days, T i If the result is 2, the rounded result is 0. Multiplying by the window size still results in 0. Therefore, the data to be processed belongs to the window corresponding to 0, which is a window between 0 and 5. Its window identifier is represented by "0".

[0139] See also Figure 6 For data with an allocation reference time of "00:08", the corresponding number of minutes in days, T i If the value is 8, the rounded result is 1. Multiplying this by the window size of 5, the data to be processed belongs to the window corresponding to 5, which is a window between 5 and 10. The window identifier is represented by "5".

[0140] See also Figure 6 For data with an allocation reference time of "00:10", the corresponding number of minutes in days, T i If the value is 10, the rounded result is 2. Multiplying this by the window size of 10, the data to be processed belongs to the window corresponding to 10, which is a window between 10 and 15. The window identifier is represented by "15".

[0141] Based on the above process, a unique window identifier can be generated for each window, and the corresponding target window identifier can be determined according to the allocation reference time of each piece of data to be processed.

[0142] In this embodiment, normal data can be directly placed into the correct window. Delayed data, since it is only late but its generation time is correct, can be placed into the corresponding time window based on the window identifier. As for paranormal data, since its generation time has been corrected, the correct time window can also be found based on the window marker.

[0143] In one possible implementation, since the corresponding time window may not be found directly after the arrival of supernatural data, a supernatural window can be generated first, the supernatural data can be placed in the supernatural window, and then each piece of supernatural data can be placed into the corresponding time window according to the window identifier calculated by the above process.

[0144] Step 305: Determine the data aggregation identifier of the data to be processed based on the target window identifier and the data label of the data to be processed.

[0145] In this embodiment, a data tag refers to the calculation key when the data to be processed is calculated, that is, the indicator data of which statistical object it is counted into. For example, when counting the number of views of an article, the article ID can be used as the data tag, or the article ID combined with the channel, such as the article ID published on website A, can be used as the data tag; or, when it is necessary to count the feature value of a feature dimension, the feature dimension can be used as the data tag. The specific setting needs to be based on the actual business needs, and this embodiment does not impose any restrictions on this.

[0146] Specifically, data tags can be carried in the data to be processed, and the corresponding data tags can be obtained from the data to be processed.

[0147] This application's embodiments cleverly design the keys of the data to be processed in the real-time data stream to obtain a data aggregation identifier (or aggregation key) that can represent a certain calculation key within a certain time window. All data to be processed are grouped according to the value of the aggregation key, and then calculations are performed within each group for aggregation keys with the same value. The aggregation key is the smallest unit for downstream computation. Each record in the real-time stream data is marked with its key in a specific column. Data with the same key value is distributed to the same computing device for processing, while data with different key values ​​is distributed to different computing devices for processing.

[0148] In one possible implementation, data labels can be combined with window identifiers to obtain a data aggregation identifier. See also Figure 7 The diagram illustrates one possible combination of data aggregation identifiers. For each piece of data to be processed, its data label can be obtained, and its corresponding target window identifier can be calculated. Then, concatenating the two yields a data aggregation identifier, as shown below. Figure 7 As shown, with Figure 6 Taking the window identifiers shown as an example, the data label of the data corresponding to "00:02" is 10000, and the corresponding window identifier is 0. After concatenation, the data aggregation identifier can be obtained as 010000. Similarly, the data label of the data corresponding to "00:08" is 10000, and the corresponding window identifier is 5. After concatenation, the data aggregation identifier can be obtained as 510000.

[0149] In one possible implementation, in addition to combining data tags and window identifiers, other content can be added to obtain a comprehensive data aggregation identifier. For example, it can be combined with the current data stream processing cycle. Of course, other content can also be added, and this application embodiment does not limit this.

[0150] Specifically, based on the current data stream processing cycle and the allocation baseline time, the cycle identifier corresponding to the data to be processed can be determined. Each cycle identifier uniquely corresponds to a data stream processing cycle. Then, the target window identifier, cycle identifier, and data label are concatenated to obtain the data aggregation identifier. Taking a day as an example, the cycle identifier can be the date of each day, uniquely identifying each day. Taking adding a date as a specific example, when the data to be processed is delayed or normal data, its date is considered normal because delayed data is simply late, but the data generation date has not changed, so the normal date is still used in the calculation. However, when there is unnatural data, since the data generation date may have changed, the corrected data generation time needs to be used to generate the data aggregation identifier when calculating the aggregation key.

[0151] Furthermore, for each piece of data to be processed, the cycle identifier corresponding to the data can be determined based on the current data stream processing cycle and the allocation reference time. Each cycle identifier uniquely corresponds to a data stream processing cycle. The target window identifier, cycle identifier, and data label are then concatenated to obtain a data aggregation identifier. In other words, its data label can be obtained, and its corresponding target window identifier and the date of the data to be processed can be calculated. These can then be concatenated to obtain a data aggregation identifier. (See [link to relevant documentation]). Figure 8 The diagram shown is a combination illustration, also using... Figure 6 Taking the window identifiers shown as an example, the data label for "00:02" is 10000, and the window identifier for the date "0603" is 0. After concatenation, the resulting data aggregation identifier is 0060310000. Similarly, the data label for "00:08" is 10000, and the corresponding window identifier is 5. After concatenation, the resulting data aggregation identifier is 5060310000.

[0152] Step 306: Send the data to be processed to the computing device corresponding to the data aggregation identifier, so that the computing device updates the associated data of the data label relative to the target time window based on the data to be processed.

[0153] In this embodiment, the same data aggregation identifier indicates that the time window and data label are the same. The data to be processed is forwarded according to the data aggregation identifier. That is, the data to be processed with the same data aggregation identifier will be sent to the same computing device for calculation. Then, when the same data aggregation identifier appears again, it will be sent to the same computing device. This ensures that the data to be processed with the same data label within the same time window will be sent to the same computing device. In this way, the computing device does not need to wait for the endtime of a time window to be triggered. Once it receives the data to be processed, it can start the calculation, which improves the calculation efficiency and enables the statistical data to be obtained in real time.

[0154] In one possible implementation, after calculating the target window identifier of the data to be processed, the data to be processed can be added to the corresponding time window. This allows for grouping the data within a time window based on their data aggregation identifiers, and assigning a corresponding computing device to each group. Each group includes at least one data item with the same data aggregation identifier, ensuring that data items in the same group are sent to the same computing device. However, it should be noted that while grouping here theoretically means that data items with the same data aggregation identifier belong to the same group, in practice, to improve data processing efficiency, after obtaining the data aggregation identifier, it can be determined whether a computing device has already been assigned to that data aggregation identifier. If it has, the data to be processed can be divided into corresponding groups based on the data aggregation identifier, and the data can be sent to the corresponding computing device for processing. If no computing device has been assigned, it indicates that the data aggregation identifier is new and has not appeared before; in this case, a new computing device needs to be assigned to it before sending the data to the computing device for processing.

[0155] In one possible implementation, the data stream processing device may process multiple real-time data streams simultaneously. In this case, the above method can be used to generate an aggregate key representing a certain calculation key in a certain time window. Then, all real-time data streams are grouped according to the value of the aggregate key, and the calculation of aggregate keys with the same value is performed within each group.

[0156] In this embodiment of the application, considering that a large number of aggregate keys will expire over time in real-world scenarios, which will seriously affect our computing efficiency, it is necessary to introduce an expiration cleanup mechanism for aggregate keys to prevent problems such as insufficient memory and excessive computing power as the number of aggregate keys increases.

[0157] Once all aggregate keys within a time window arrive and trigger calculations, the aggregate key expires. To prevent the number of aggregate keys from increasing and consuming excessive memory, aggregate keys need to be evicted. This can be achieved by setting expiration times for aggregate keys and periodically cleaning up expired ones, ensuring that the number of aggregate keys in memory doesn't keep increasing and preventing performance degradation over time. Therefore, after generating an aggregate key from the data to be processed, if the aggregate key already exists, its unupdated duration needs to be updated. The unupdated duration represents the time difference between the time the most recent data to be processed associated with the aggregate key was generated and the current time; in other words, it's the length of time the data associated with that aggregate key hasn't been updated. If an aggregate key hasn't been updated for a long time, it indicates that it has been processed and expired, and should be cleaned up.

[0158] Specifically, after generating an aggregation key from the data to be processed, the non-updated duration corresponding to the data aggregation identifier can be updated based on the time difference between the data generation time and the current time. When the non-updated duration exceeds the effective duration threshold set for that aggregation key, the associated data of the aggregation key is cleaned up, such as clearing the aggregation key cached in memory. In practical scenarios, the effective duration threshold can be set according to the actual business requirements.

[0159] The technical solutions of the embodiments of this application are described below with specific examples. See also Figures 9a-9c The diagram illustrates the allocation of time windows for the data to be processed. Here, a 10-minute time window is used as an example. (See attached image.) Figures 9a-9c As shown, this involves three time windows: "00:00~00:10", "00:10~00:20" and "00:20~00:30".

[0160] See Figure 9a As shown, the data arriving at this time are in the order of "3, 1, 7, 3, 15, 11, 16, 23". Here, each value corresponds to a piece of data to be processed, and the value represents the time when the data was generated. For example, "3" represents "00:03". Based on the generation and arrival times of each piece of data to be processed, we can see that... Figure 9aAll the data to be processed in the table are normal data. Taking one of the data to be processed, "3", as an example, the window identifier can be calculated as "0" using the above window identifier calculation formula, corresponding to the first window "00:00~00:10", so it is assigned to window "00:00~00:10". In this way, other data to be processed can also be assigned to the corresponding windows, that is, "3, 7, 1, 3" are assigned to window "00:00~00:10", "16, 11, 15" are assigned to window "00:10~00:20", and "23" is assigned to "00:20~00:30".

[0161] See Figure 9b As shown, the data queue that arrives at this time contains delayed data "4". According to the above window identifier calculation formula, the window identifier can be obtained as "0", corresponding to the first window "00:00~00:10". Therefore, it can be correctly assigned to the window "00:00~00:10". In other words, the delayed data is only delayed in arrival time. However, according to the solution of this application embodiment, even delayed data can be correctly found in the corresponding time window.

[0162] See Figure 9c and Figure 9d As shown, the data queue arriving at this point contains the supernatural data "48". There are two possible processing methods, see [link / reference]. Figure 9c As shown, one approach is to not correct for the data generation time. Based on the window identifier calculation formula mentioned above, the window identifier is "40," corresponding to a window "00:40~00:50." However, according to the current time window, this time window is not yet open. Therefore, a new supernatural window can be created, and the data to be processed can be placed within this supernatural window. See also... Figure 9d As shown, another approach is to correct the data generation time. For example, if the corrected data generation time is "5", the window identifier can be calculated as "0" according to the above window identifier calculation formula, corresponding to the first window "00:00~00:10". This allows the data to be correctly assigned to the window "00:00~00:10". This method can improve the accuracy of the statistical results for each time window, because supernatural time itself does not follow a pattern. Assigning it to a supernatural window is equivalent to statistically assigning the data of a certain time window to a future time window. By correcting the data generation time, it can be assigned to the correct time window as much as possible.

[0163] Furthermore, in this embodiment of the application, the time window identifier and the calculation key are combined to obtain an aggregate key, which is used to group the data and send it to the corresponding computing device for calculation, thereby realizing the statistics of window data, and eliminating expired aggregate keys and data that has already participated in the calculation in the window.

[0164] In summary, this application transforms window computation into a key-based classification and aggregation problem. By cleverly designing the keys to be included in the aggregation calculation, all computation keys within the same window are sent to the same location, achieving aggregated statistics. Specifically, the EventTime of the real-time data stream is used to calculate the window identifier needed to generate the aggregation key. The computation keys to be aggregated and the window identifier are then assembled to generate an aggregation key representing a specific computation key within a specific window. All data is grouped based on the value of this aggregation key, thus collecting all data for the aggregation key within the current window. Aggregation or more complex computational logic is then implemented within each group. Since data is processed as soon as it is received, delays or unexpected data occurrences do not affect the computation, greatly ensuring that unexpected or delayed data can be processed, significantly increasing the accuracy and stability of data processing. This solution can be applied in scenarios with extremely high requirements for data timeliness and accuracy, such as real-time business decision-making, online feature engineering, and rule engine alerts. It can efficiently calculate statistical values ​​of content within a window, obtaining highly consistent and accurate data, and improving user experience. It can also be applied to business scenarios with extremely high data throughput requirements, such as online log analysis, online machine learning, online graph computation, and online recommendation algorithms.

[0165] Please see Figure 10 Based on the same inventive concept, embodiments of this application also provide a data stream processing apparatus 100, which includes:

[0166] The receiving and processing unit 1001 is used to receive real-time data streams and determine the target processing type of the data to be processed based on the data generation time and data arrival time corresponding to the data to be processed in the real-time data stream.

[0167] The allocation reference determination unit 1002 is used to obtain the allocation reference time required to allocate a window to the data to be processed based on the data generation time and the target processing type.

[0168] The window allocation unit 1003 is used to determine the target window identifier of the target time window to which the data to be processed belongs, based on the allocation reference time.

[0169] The aggregation determination unit 1004 is used to determine the data aggregation identifier of the data to be processed based on the target window identifier and the data label of the data to be processed;

[0170] The sending unit 1005 is used to send the data to be processed to the computing device corresponding to the data aggregation identifier, so that the computing device updates the associated data of the data tag relative to the target time window based on the data to be processed.

[0171] In one possible implementation, the device further includes a cleaning unit 1006 for:

[0172] Based on the time difference between the data generation time and the current time, update the non-updated duration corresponding to the data aggregation identifier. The non-updated duration is used to represent the time difference between the data generation time of the most recent data to be processed corresponding to the data aggregation identifier and the current time.

[0173] If the duration of the update is not greater than the effective duration threshold corresponding to the data aggregation identifier, then the associated data of the data aggregation identifier will be cleaned up.

[0174] In one possible implementation, the receiving and processing unit 1001 is specifically used for:

[0175] If the data is generated later than the data arrives, then the target data type is supernatural data;

[0176] If the data generation time is earlier than the data arrival time, and the difference between the data generation time and the data arrival time is greater than the preset duration threshold, then the target data type is delayed data.

[0177] If the data generation time is earlier than the data arrival time, and the difference between the data generation time and the data arrival time is less than the preset duration threshold, then the target data type is normal data.

[0178] In one possible implementation, the allocation reference determination unit 1002 is specifically used for:

[0179] If the target data type is delayed data or normal data, then the data generation time is determined as the allocation reference time;

[0180] If the target data type is supernatural data, then the time of data generation is corrected, and the corrected time of data generation is determined as the allocation reference time.

[0181] In one possible implementation, the allocation reference determination unit 1002 is specifically used for:

[0182] Determine the current time as the corrected data generation time; or...

[0183] The arrival time of the data is determined as the corrected data generation time.

[0184] In one possible implementation, the allocation reference determination unit 1002 is specifically used for:

[0185] Obtain the time difference between the data production equipment and the current equipment corresponding to the data to be processed;

[0186] Based on the device time difference, the time of data generation is corrected to obtain the corrected data generation time.

[0187] In one possible implementation, the window allocation unit 1003 is specifically used for:

[0188] Based on the business type to which the real-time data stream belongs, determine the size of the time window and the data stream processing cycle corresponding to the real-time data stream.

[0189] Determine the time difference between the allocation reference time and the start time of the current data stream processing cycle, and the method of determining the time difference should match the business type;

[0190] The target window identifier is determined based on the time difference and window size.

[0191] In one possible implementation, the aggregation determination unit 1004 is specifically used for:

[0192] Based on the current data stream processing cycle and the allocation reference time, determine the cycle identifier corresponding to the data to be processed. Each cycle identifier uniquely corresponds to a data stream processing cycle.

[0193] The target window identifier, period identifier, and data label are concatenated to obtain the data aggregation identifier.

[0194] In one possible implementation, the device further includes a grouping unit 1007, for:

[0195] Based on the data aggregation identifiers of the data included in the target time window, the data are grouped and a corresponding computing device is allocated to each group; wherein each group includes at least one data with the same data aggregation identifier;

[0196] Then the sending unit 1005 is specifically used for:

[0197] Based on the data aggregation identifier of the data to be processed, the data to be processed is divided into corresponding groups, and the data to be processed is sent to the computing device corresponding to the group.

[0198] The aforementioned device allows for the determination of the allocation reference time required to allocate windows for the data to be processed by combining the target processing type of the data, rather than directly allocating based on the data generation time. This effectively avoids the problem of incorrect window allocation for supernatural or delayed data. Furthermore, in this embodiment, it eliminates the need to wait for all data in each window to arrive before calculation. Instead, the data label and window identifier of the data to be processed are combined to obtain a data aggregation identifier, which is then sent to the computing device corresponding to the data aggregation identifier for processing. Thus, regardless of whether the data arrives late or supernatural data is present, the computing device can perform calculations upon receiving the data without waiting for all data in the same window to arrive, improving data processing efficiency and avoiding any impact on the data processing process.

[0199] This device can be used to execute the methods shown in the various embodiments of this application. Therefore, the functions that each functional module of this device can achieve can be referred to the description of the foregoing embodiments, and will not be repeated here.

[0200] Please see Figure 11 Based on the same technical concept, embodiments of this application also provide a computer device. This computer device, as... Figure 11 As shown, it includes a memory 1101, a communication module 1103, and one or more processors 1102.

[0201] The memory 1101 is used to store computer programs executed by the processor 1102. The memory 1101 may mainly include a program storage area and a data storage area. The program storage area may store the operating system and programs required to run instant messaging functions, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.

[0202] Memory 1101 may be volatile memory, such as random-access memory (RAM); memory 1101 may also be non-volatile memory, such as read-only memory, flash memory, hard disk drive (HDD), or solid-state drive (SSD); or memory 1101 may be any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 1101 may be a combination of the above-described memories.

[0203] Processor 1102 may include one or more central processing units (CPUs) or digital processing units, etc. Processor 1102 is used to implement the above-described search result reordering method when it calls a computer program stored in memory 1101.

[0204] The communication module 1103 is used to communicate with terminal devices and other servers.

[0205] This application embodiment does not limit the specific connection medium between the memory 1101, communication module 1103, and processor 1102. This application embodiment... Figure 11 The memory 1101 and the processor 1102 are connected via a bus 1104, and the bus 1104 is in Figure 11 The diagram uses thick lines to describe the connections between other components; these are for illustrative purposes only and should not be considered limiting. Bus 1104 can be divided into address bus, data bus, control bus, etc. For ease of description, Figure 11 It is described using only a thick line, but does not indicate that there is only one bus or one type of bus.

[0206] The memory 1101 stores a computer storage medium, which stores computer-executable instructions. The computer-executable instructions are used to implement the search result reordering method of the embodiments of this application, and the processor 1102 is used to execute the search result reordering methods of the above embodiments.

[0207] Based on the same inventive concept, embodiments of this application also provide a storage medium storing a computer program that, when run on a computer, causes the computer to perform the steps in the search result reordering method according to various exemplary embodiments of this application described above.

[0208] In some possible implementations, various aspects of the search result reordering method provided in this application may also be implemented in the form of a computer program product, which includes a computer program that, when run on a computer device, causes the computer device to perform the steps in the search result reordering method according to various exemplary embodiments of this application described above. For example, the computer device may perform the steps of the various embodiments.

[0209] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0210] The program product of the embodiments of this application may employ a portable compact disc read-only memory (CD-ROM) and include a computer program, and may run on a computer device. However, the program product of this application is not limited thereto. In this application, the readable storage medium may be any tangible medium that contains or stores a program, and the computer program included therein may be used by or in conjunction with a command execution system, apparatus, or device.

[0211] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a readable computer program. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting a program for use by or in conjunction with a command execution system, apparatus, or device.

[0212] Computer programs contained on readable media may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0213] Computer programs for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages.

[0214] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0215] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0216] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0217] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0218] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A data stream processing method, characterized in that, The method includes: Receive real-time data streams and determine the target processing type of the data to be processed based on the data generation time and data arrival time corresponding to the data to be processed in the real-time data streams; If the target processing type is delayed data or normal data, then the data generation time is determined as the allocation reference time; If the target processing type is supernatural data, then the current time or the data arrival time is determined as the corrected data generation time, or the data generation time is corrected based on the device time difference between the data production device corresponding to the data to be processed and the current device to obtain the corrected data generation time; and the corrected data generation time is determined as the allocation reference time. Based on the allocation reference time, determine the target window identifier of the target time window to which the data to be processed belongs; Based on the target window identifier and the data tag of the data to be processed, the data aggregation identifier of the data to be processed is determined; The data to be processed is sent to the computing device corresponding to the data aggregation identifier, so that the computing device updates the associated data of the data tag relative to the target time window based on the data to be processed.

2. The method as described in claim 1, characterized in that, After determining the data aggregation identifier of the data to be processed based on the target window identifier and the data label of the data to be processed, the method further includes: Based on the time difference between the data generation time and the current time, the unupdated duration corresponding to the data aggregation identifier is updated, wherein the unupdated duration is used to characterize the time difference between the data generation time of the most recent data to be processed corresponding to the data aggregation identifier and the current time; If the duration of the unupdated data is greater than the effective duration threshold corresponding to the data aggregation identifier, then the associated data of the data aggregation identifier will be cleaned up.

3. The method as described in claim 1, characterized in that, Based on the data generation time and data arrival time corresponding to the data to be processed in the real-time data stream, the target processing type of the data to be processed is determined, including: If the data is generated later than the data arrives, then the target processing type is supernatural data; If the data generation time is earlier than the data arrival time, and the difference between the data generation time and the data arrival time is greater than a preset duration threshold, then the target processing type is delayed data. If the data generation time is earlier than the data arrival time, and the difference between the data generation time and the data arrival time is less than a preset duration threshold, then the target processing type is normal data.

4. The method as described in claim 1, characterized in that, Based on the allocation reference time, the target window identifier of the target time window to which the data to be processed belongs is determined, including: Based on the service type to which the real-time data stream belongs, determine the size of the time window and the data stream processing cycle corresponding to the real-time data stream; Determine the time difference between the allocation reference time and the start time of the current data stream processing cycle, and the method of taking the value of the time difference is matched with the service type; The target window identifier is determined based on the time difference and the window size.

5. The method as described in claim 4, characterized in that, Based on the target window identifier and the data tag of the data to be processed, the data aggregation identifier of the data to be processed is determined, including: Based on the current data stream processing cycle and the allocation reference time, the cycle identifier corresponding to the data to be processed is determined, and each cycle identifier uniquely corresponds to a data stream processing cycle. The target window identifier, the period identifier, and the data label are concatenated to obtain the data aggregation identifier.

6. The method as described in claim 1, characterized in that, Before sending the data to be processed to the computing device corresponding to the data aggregation identifier, the method further includes: Based on the data aggregation identifiers of the data included in the target time window, the data are grouped and a corresponding computing device is allocated to each group; wherein, each group includes at least one data with the same data aggregation identifier; Then, the data to be processed is sent to the computing device corresponding to the data aggregation identifier, including: Based on the data aggregation identifier of the data to be processed, the data to be processed is divided into corresponding groups, and the data to be processed is sent to the computing device corresponding to the group.

7. A data stream processing apparatus, characterized in that, The device includes: A receiving and processing unit is used to receive a real-time data stream and determine the target processing type of the data to be processed based on the data generation time and data arrival time corresponding to the data to be processed in the real-time data stream. The allocation reference determination unit is used to determine the data generation time as the allocation reference time if the target processing type is delayed data or normal data; if the target processing type is supernatural data, it determines the current time or the data arrival time as the corrected data generation time, or it performs time correction on the data generation time based on the device time difference between the data production device corresponding to the data to be processed and the current device to obtain the corrected data generation time; and determines the corrected data generation time as the allocation reference time. A window allocation unit is used to determine the target window identifier of the target time window to which the data to be processed belongs, based on the allocation reference time. An aggregation determination unit is used to determine the data aggregation identifier of the data to be processed based on the target window identifier and the data tag of the data to be processed; The sending unit is used to send the data to be processed to the computing device corresponding to the data aggregation identifier, so that the computing device updates the associated data of the data tag relative to the target time window based on the data to be processed.

8. The apparatus as claimed in claim 7, characterized in that, The device further includes a cleaning unit for: Based on the time difference between the data generation time and the current time, the unupdated duration corresponding to the data aggregation identifier is updated, wherein the unupdated duration is used to characterize the time difference between the data generation time of the most recent data to be processed corresponding to the data aggregation identifier and the current time; If the duration of the unupdated data is greater than the effective duration threshold corresponding to the data aggregation identifier, then the associated data of the data aggregation identifier will be cleaned up.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer storage medium storing computer program instructions thereon, characterized in that, When executed by a processor, the computer program instructions implement the steps of the method according to any one of claims 1 to 6.

11. A computer program product comprising computer program instructions, characterized in that, When executed by a processor, the computer program instructions implement the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-thread data processing method and device based on streaming computing framework and medium

    CN112286582A

  • Processing method and device for data timeout in real-time calculation

    CN113204387A