Big data stream processing method, device and storage medium

By judging the activity and persistence of big data streams, and labeling and processing different types of data streams, the problems of wasted computing resources and low data filtering efficiency are solved, enabling efficient data processing and scientific business decision-making.

CN119544547BActive Publication Date: 2025-10-24HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411613983.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-10-24
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

Existing big data stream processing methods suffer from wasted computing resources in terms of processing speed and efficiency, and they struggle to identify data elements that have a long-term impact on business operations, affecting the quality of decision-making and analysis.

Method used

By judging the activity and persistence of data streams, inactive data streams are marked, temporarily stored or filtered, and persistent active data streams are stored. Dynamic scheduling strategies are adopted to reduce the waste of computing resources and improve data processing efficiency and decision-making.

Benefits of technology

Effectively identify and filter inactive data streams, reduce computing resource consumption, improve the efficiency of initial data processing screening and the scientific nature of business decisions, and ensure the depth and accuracy of key data items.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119544547B_ABST
    Figure CN119544547B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of big data, and provides a big data stream processing method, which comprises the following steps: determining whether a target data stream is an active data stream in a big data stream; if the target data stream is an inactive data stream in the big data stream, marking the inactive data stream; if the target data stream is an active data stream in the big data stream, determining whether the active data stream is a persistent active data stream; if the active data stream is the persistent active data stream, storing and marking the persistent active data stream; and if the active data stream is a non-persistent active data stream, temporarily storing and marking the non-persistent active data stream as a temporary active data stream. The technical scheme can reduce invalid consumption of computing resources, improve the preliminary screening efficiency of data processing, and improve the scientificity of business decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of big data, and particularly relates to a big data stream processing method, device and storage medium. BACKGROUND

[0002] In the field of data stream processing, especially in the field of big data and real-time data analysis, various technologies have been developed to monitor and analyze continuously flowing data. These technologies aim to quickly capture and process data to support immediate and long-term decisions. For example, one solution is to help detect persistent behavior in network security by estimating the persistence of data items, and another solution is to focus on discovering those persistent but infrequent data streams, which are particularly important for identifying security threats such as advanced persistent threats (APTs).

[0003] Although the above-mentioned big data stream processing method has advantages in processing speed and efficiency, there are still disadvantages such as wasting computing resources and reducing processing efficiency due to indiscriminate processing of all data, and it is difficult to filter out data elements that have a long-term impact on business operations, thereby affecting long-term decision-making and analysis quality. SUMMARY

[0004] The present application provides a big data stream processing method, device and storage medium, which can reduce the invalid consumption of computing resources and improve the preliminary screening efficiency of data processing and the scientificity of business decision-making.

[0005] In one aspect, the present application provides a big data stream processing method, which comprises:

[0006] determining whether a target data stream is an active data stream in a big data stream;

[0007] if the target data stream is an inactive data stream in the big data stream, marking the inactive data stream;

[0008] if the target data stream is an active data stream in the big data stream, determining whether the active data stream is a persistent active data stream;

[0009] if the active data stream is a persistent active data stream, storing and marking the persistent active data stream;

[0010] if the active data stream is a non-persistent active data stream, temporarily storing and marking the non-persistent active data stream as a temporary active data stream.

[0011] In another aspect, the present application provides a big data stream processing device, which comprises:

[0012] a first determining module configured to determine whether a target data stream is an active data stream in a big data stream;

[0013] A first marking module, configured to mark the inactive data flow if the target data flow is an inactive data flow in the large data flow;

[0014] a second determining module, configured to determine whether the target data stream is an active data stream in the large data stream, if the target data stream is an active data stream;

[0015] a second marking module, configured to store and mark the persistent active data flow if the active data flow is a persistent active data flow;

[0016] The third marking module is configured to temporarily store and mark the active data flow as a temporary active data flow if the active data flow is a non-persistent active data flow.

[0017] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the technical solution of the above-mentioned large data stream processing method when executing the computer program.

[0018] In a fourth aspect, the present application provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the technical solution of the above-mentioned large data stream processing method.

[0019] From the technical solution provided by the present application, it can be seen that, on the one hand, after determining that the target data stream is an inactive data stream in the large data stream, by marking these inactive data streams, the inactive elements in the small stream data can be effectively identified and filtered out, significantly reducing the ineffective consumption of computing resources and improving the initial screening efficiency of data processing; on the other hand, after further determining that the active data stream is a persistent active data stream, by storing and marking these persistent active data streams, the depth and accuracy of the large data stream processing are ensured, and the system can also identify and focus on key data items that have a long-term impact on business decisions, greatly improving the scientific nature of business decisions. In summary, the technical solution of the present application can reduce the ineffective consumption of computing resources, improve the initial screening efficiency of data processing and the scientific nature of business decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0021] Figure 1 is a flow chart of a big data stream processing method provided by an embodiment of the present application;

[0022] Figure 2 is a structural schematic diagram of a big data stream processing device provided by an embodiment of the present application;

[0023] Figure 3 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0024] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0025] In this specification, adjectives such as first and second can be used merely to distinguish one element or action from another element or action, without necessarily requiring or implying any actual such relationship or order. Where the context permits, reference to elements or components or steps (etc.) can be interpreted as referring to one or more elements, components, or steps (etc.).

[0026] In this specification, for the convenience of description, the sizes of the various parts shown in the drawings are not drawn in accordance with the actual proportional relationship.

[0027] In the field of data stream processing, especially in the aspects of big data and real-time data analysis, various technologies have been developed to monitor and analyze continuously flowing data. These technologies aim to quickly capture and process data to support immediate and long-term decisions. For example, one solution is to help detect persistent behaviors in network security by estimating the persistence of data items, and another solution is to focus on finding those persistent but infrequent data streams, which are particularly important for identifying security threats such as advanced persistent threats (APTs). Although the above-mentioned big data stream processing methods have advantages in processing speed and efficiency, there are still disadvantages such as wasting computing resources and reducing processing efficiency due to indiscriminate processing of all data, and there are also defects such as difficulty in screening out data elements that have a long-term impact on business operations, thereby affecting long-term decision-making and analysis quality.

[0028] In view of the above problems of the prior art, the present application provides a big data stream processing method, the flow chart of which is shown in FIG. 1, mainly comprising steps S101-S105, which are described in detail as follows: Figure 1

[0029] ​Step S101: judging whether the target data flow is an active data flow in the big data flow.

[0030] In the field of big data flow processing, not every data flow or every kind of data flow in the big data flow plays a role in decision-making. In fact, those small flow data or data flows with low activity in the big data flow not only do not help business decision-making, but also may form invalid consumption of computing resources. Therefore, in order to reduce the invalid consumption of computing resources and improve the preliminary screening efficiency of data processing, the technical scheme of the present application can first judge whether the target data flow is an active data flow in the big data flow. It should be noted that the target data flow refers to the data flow in the big data flow which needs to be monitored by the embodiments of the present application.

[0031] As an embodiment of the present application, whether the target data flow is an active data flow in the big data flow can be realized by steps S1011 to S1013, which are described in detail as follows:

[0032] Step S1011: determining the dimension for evaluating whether the target data flow is an active data flow in the big data flow.

[0033] In the embodiments of the present application, the dimension for evaluating whether the target data flow is an active data flow in the big data flow includes the update frequency of the target data flow per unit time, the update frequency of the big data flow per unit time, and the diversity or uncertainty of the target data flow, etc., wherein the update frequency of the target data flow per unit time can be the number of events of the target data flow per unit time, and the update frequency of the big data flow per unit time is the number of events of all big data flows.

[0034] Step S1012: calculating the activity index of the target data flow in the time window w j according to the dimension for evaluating whether the target data flow is an active data flow in the big data flow.

[0035] If the update frequency of the target data flow per unit time is denoted as f target , the update frequency of the big data flow per unit time is denoted as total_flow_rate, and the diversity or uncertainty of the target data flow is denoted as H target , then the activity index of the target data flow in the time window w j can be calculated according to the dimension for evaluating whether the target data flow is an active data flow in the big data flow by the formula The activity index A target of the target data flow in the time window w j can be calculated according to the dimension for evaluating whether the target data flow is an active data flow in the big data flow by the formula target . In the above embodiments, the diversity or uncertainty H target of the target data flow can be determined by the following way: obtaining a set S composed of events contained in the target data flow; calculating the diversity or uncertainty H jAny event s in the inner set S i The probability of presence in the target data stream p(s i ); According to the formula Calculate the diversity or uncertainty H of the target data stream target , where k is the number of types of events contained in the set S. In the above embodiment, if the time window w j Any event s in the inner set S i The number of times it is presented in the target data stream is count(s i ), the total number of events in the big data stream is N, then in the time window w j Any event s in the inner set S i The probability of presence in the target data stream p(s i ) can be calculated according to the formula To calculate.

[0036] Step S1013: If the target data stream is within the time window w j If the activity index within exceeds a first preset threshold, the target data stream is determined to be an active data stream in the large data stream.

[0037] If the target data stream is in the time window w j If the activity index in the data stream exceeds the first preset threshold, it means that the target data stream is highly active, and the target data stream can be determined to be an active data stream in the big data stream. The first preset threshold is an important parameter that determines when the target data stream is considered "active". Therefore, the method of the above embodiment may also include: extracting the activity index from the historical big data stream and calculating the activity of all big data streams; obtaining the statistical characteristics of the activity of all big data streams, calculating their mean μ and standard deviation σ; according to the formula θ active =μ+α1*σ to calculate θ active , and θ active Determine as the first preset threshold, α1 is the first adjustment coefficient. Usually, the value of the first adjustment coefficient α1 can be 1 or 2. Of course, the value of the first adjustment coefficient α1 can be adjusted dynamically. As an embodiment of the present application, the value of the first adjustment coefficient α1 can be dynamically adjusted: by analyzing the volatility of the big data stream through the standard deviation σ of the big data stream, that is, the larger the standard deviation σ, the greater the volatility of the big data stream, and vice versa; according to the volatility of the data, select a reasonable α1 range, that is, if the standard deviation σ of the big data stream is large, select α1 in the range of [1, 2], so that the first preset threshold θ active Relatively stable and not easily affected by small fluctuations. If the standard deviation σ of the large data stream is small, α1 can be selected in the range of [2, 5] so that the first preset threshold θ activeMore sensitive to changes; or dynamically adjust α1 according to the periodicity, trend and volatility of the big data stream, that is, if the big data stream is in a high volatility period, increase the value of the first adjustment coefficient α1 so that the first preset threshold θ active More flexible, if the large data flow is in a low fluctuation period, the value of the first adjustment coefficient α1 is reduced so that the first preset threshold θ active More stable.

[0038] As another embodiment of the present application, the first preset threshold θ aetive Alternatively, the weighted average of historical activity data can be used for setting the value. The specific scheme is as follows: Steps S1 to S3:

[0039] Step S1: Collect the activity data {A1, A2, ..., A i ,...,A n}, where A i Indicates the activity value of the target data stream in the i-th time window.

[0040] Step S2: Calculate the weighted average activity A of the target data stream according to the following formula: weighted :

[0041]

[0042] Among them, ω i It is the weight associated with the time window, usually set according to the distance between the time window and the current time; a linear attenuation method can be selected, that is:

[0043]

[0044] In this way, the closer the target data stream is to the current time window, the higher its weight ω i The bigger.

[0045] Step S3: The weighted average activity A of the target data stream weighted Multiply by a threshold ratio γ, for example γ = 90%, and the product A weighted *γ is the first preset threshold θ active , that is, θ active =A weighted *γ.

[0046] Step S102: If the target data flow is an inactive data flow in the large data flow, the inactive data flow is marked.

[0047] When a target data stream is an inactive data stream in a large data stream, marking it as an inactive data stream means that attention can be reduced to the target data stream and it will not be considered as an influencing factor for business decisions.

[0048] Step S103: If the target data stream is an active data stream in the large data stream, it is determined whether the active data stream is a persistent active data stream.

[0049] In the large data stream, the activity level of a data stream is dynamically changing, some data streams can only be active for a short time, while other data streams can be continuously active. In this case, simply treating all "active" data streams equally can lead to resource waste or low processing efficiency. In other words, for each data stream, especially the large data stream, a certain amount of computing resources are needed for processing, storage, indexing and retrieval. If a data stream is determined to be active (i.e. a certain event or data is generated) for a short time, but it stops being active soon after, it is a waste to continuously retain computing resources for it. While the persistent active data stream usually contains valuable or high-priority data, it needs to be continuously tracked and analyzed. For example, financial transaction stream, social media active user stream or system log stream, which can reflect the long-term trend of system status or user behavior. Based on the above facts, when it is determined that the target data stream is an active data stream in the large data stream, it is further determined whether the active data stream is a persistent active data stream.

[0050] As an embodiment of the present application, whether the active data stream is a persistent active data stream can be achieved by steps S1031 to S1033, which are described in detail as follows:

[0051] Step S1031: Determine the dimension for evaluating whether the target data stream is an active data stream in the large data stream.

[0052] In the embodiment of the present application, the dimension for evaluating whether the target data stream is an active data stream in the large data stream can be the active change rate and the persistence of activity, the former represents the change of the activity level of the data stream in different time windows, and the latter can be measured by the average activity level of the target data stream in multiple time windows.

[0053] Step S1032: According to the dimension for evaluating whether the target data stream is an active data stream in the large data stream, the activity index of the target data stream in multiple time windows is calculated.

[0054] If the activity index of the target data stream in the time window t i is denoted as A target (t i ), the activity index P target of the target data stream in n time windows can be calculated using the formula . It should be noted that the activity index of the target data stream in the time window ti The activity index of the target data stream within the time window w target i The calculation scheme of the activity index of the target data stream within the time window w j is the same as the technical scheme of calculating the activity index of the target data stream within the time window w

[0055] Step S1033: If the activity index of the target data stream within the multiple time windows exceeds the second preset threshold, it is determined that the target data stream is a persistent active data stream.

[0056] If the second preset threshold is denoted as θ persistent , when the activity index of the target data stream within the multiple time windows exceeds the second preset threshold, i.e., P target > θ persistent , it can be determined that the target data stream is a persistent active data stream. It should be noted that, similar to the first preset threshold θ active in the foregoing embodiments, in the embodiments of the present application, the second preset threshold θ persistent is also an important parameter for determining whether an active data stream is a persistent active data stream. As an embodiment of the present application, the method for determining the second preset threshold θ persistent may be: obtaining the activity of the target data stream in T historical time windows; calculating the weighted average activity A weighted of the activity of the target data stream in the T historical time windows; calculating θ persistent according to the formula θ weighted = A persistent + α2* β, and determining θ persistent as the second preset threshold, where α2 is a second adjustment coefficient, and β is the rate of change of activity with time, i.e., the trend slope, obtained through linear regression. If the activity of the target data stream in the T historical time windows is denoted as {A1, A2,..., A t ,..., A T}, where 4 represents the activity of the target data stream in the tth historical time window, the weighted average activity A weighted of the activity of the target data stream in the T historical time windows can be calculated according to the formula , where ω t represents the weight of the data stream corresponding to the activity A t . In the time window, the data stream closer to the current time contributes more to the activity evaluation, and the weight ω t can be set through a linear function, i.e., the closer to the current time, the greater the weight, and the farther from the current time, the smaller the weight. Specifically, if it is assumed that the current time is t​now , the length of the time window is T, and the data streams in the time window are in chronological order from t0to t T (where t0is the earliest time and t T is the latest time), the weight ω t may be set in the following manner:

[0057]

[0058] where t is the timestamp of a data stream in the time window, t now is the current time, and T is the total length of the time window. This means that, the closer to the current time t now , i.e., t now , the smaller ω t is; the farther away from the current time t now , i.e., t now -t, the larger ω t is; and the value of ω t ranges between 0 and 1.

[0059] In addition to setting the weight ω t in the above manner, the weight ω t may also be set in an exponentially weighted manner, i.e.:

[0060]

[0061] where α is a decay factor (usually 0<α<1) that determines the rate of time decay, t is the timestamp of a data point in the time window, and t now is the current time. The exponentially weighted manner of setting the weight ω t has the advantage of being more sensitive to changes in the short term, and the decay rate can be controlled by adjusting α. Generally, α takes a small value (e.g., 0.9 or 0.95), so that the more recent data dominates.

[0062] In the above embodiments, the calculation formula of the second preset threshold θ persistent is θ persistent =A weighted +α2*β, and α2 is a hyperparameter, where β can be obtained by linear regression, for example, by fitting a linear regression model using the least squares method, and α2 is a hyperparameter that controls the influence of the trend on the second preset threshold θ persistent . The selection of α2 will affect the sensitivity of the threshold. In the embodiments of the present application, one method of determining α2 is to define a loss function L(α2) for evaluating the performance of the model in predicting long-term activity, i.e., for a set of real long-term activity labels θ real , the loss function can be set as:

[0063]

[0064] wherein A weighted,i is the weighted average activity of the i-th data stream, β i is its corresponding trend slope, θ real,i is the actual persistent activity threshold; then, an optimization algorithm such as gradient descent is used to minimize the loss function L(a2) to obtain the optimal second adjustment coefficient a0.

[0065] Step S104: If the active data stream is a persistent active data stream, store and mark the persistent active data stream.

[0066] The persistent active data stream usually contains valuable or high-priority data, which may be associated with real-time monitoring, early warning mechanisms or business decisions. They not only require more computing resources, but also need to be processed continuously. Therefore, if the active data stream is a persistent active data stream, these persistent active data streams need to be stored and marked, i.e. using persistent storage, to ensure that they can be fully processed throughout their life cycle. For example, dynamically adjusting the scheduling strategy of computing and storage to achieve more efficient load balancing and resource allocation.

[0067] Step S105: If the active data stream is a non-persistent active data stream, temporarily store and mark the non-persistent active data stream as a temporary active data stream.

[0068] The non-persistent active data stream has relatively low timeliness and importance, and may only be a fluctuation at a certain moment, which will not have a significant impact on the overall operation or decision of the system. For example, a sudden event in a small range may cause the data stream to increase for a short time, but it will return to normal after a period of time. For non-persistent active data streams, they can be temporarily stored and marked as temporary active data streams for on-demand processing, such as processing data during the active period of the data stream, and then clearing the cache or reducing the processing frequency. The advantage of this strategy is that unnecessary computing and storage can be avoided when the activity of the stream is low, thereby improving system efficiency and reducing resource waste.

[0069] From the above description of the preferred embodiments of the present application, it can be seen that the present application has the following advantages: Figure 1The example big data stream processing method can know that, on one hand, after judging that the target data stream is a non-active data stream in the big data stream, the non-active elements in the small stream data can be effectively identified and filtered by marking the non-active data stream, the invalid consumption of the computing resource is significantly reduced, and the preliminary screening efficiency of the data processing is improved; on the other hand, after further judging that the active data stream is a persistent active data stream, the depth and accuracy of the big data stream processing are ensured by storing and marking the persistent active data stream, and the system can identify and process the key data item which has a long-term influence on the business decision, so that the scientificity of the business decision is greatly improved. In conclusion, the technical scheme of the application can reduce the invalid consumption of the computing resource, improve the preliminary screening efficiency of the data processing and the scientificity of the business decision.

[0070] Please refer to the accompanying drawings Figure 2 The device for processing big data stream provided by the embodiment of the application can include a first judging module 201, a first marking module 202, a second judging module 203, a second marking module 204 and a third marking module 205, which are described as follows:

[0071] The first judging module 201 is used for judging whether the target data stream is an active data stream in the big data stream.

[0072] The first marking module 202 is used for marking the non-active data stream if the target data stream is a non-active data stream in the big data stream.

[0073] The second judging module 203 is used for judging whether the active data stream is a persistent active data stream if the target data stream is an active data stream in the big data stream.

[0074] The second marking module 204 is used for storing and marking the persistent active data stream if the active data stream is a persistent active data stream.

[0075] The third marking module 205 is used for temporarily storing and marking the non-persistent active data stream as a temporary active data stream if the active data stream is a non-persistent active data stream.

[0076] From the above description of the accompanying drawings Figure 2The example big data stream processing apparatus can know that, on one hand, after judging that the target data stream is a non-active data stream in the big data stream, by marking the non-active data streams, the non-active elements in the small stream data can be effectively identified and filtered, the invalid consumption of the computing resources is significantly reduced, and the preliminary screening efficiency of the data processing is improved; on the other hand, after further judging that the active data stream is a persistent active data stream, by storing and marking the persistent active data streams, the depth and accuracy of the big data stream processing are ensured, and the system can identify and process the key data items which have long-term influence on the business decision, so that the scientificity of the business decision is greatly improved. In conclusion, the technical scheme of the present application can reduce the invalid consumption of the computing resources, improve the preliminary screening efficiency of the data processing and the scientificity of the business decision.

[0077] Figure 3 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. As shown in the figure, the electronic device 3 of the embodiment mainly comprises a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, for example, a program of a big data stream processing method. The processor 30 implements the steps in the above-mentioned embodiment of the big data stream processing method when executing the computer program 32, for example, the steps S101 to S105 shown in the figure. Alternatively, the processor 30 implements the functions of the modules / units in the above-mentioned embodiment of each apparatus when executing the computer program 32, for example, the functions of the first judging module 201, the first marking module 202, the second judging module 203, the second marking module 204, and the third marking module 205 shown in the figure. Figure 3 Figure 1 Figure 2

[0078] ​​​Exemplarily, the computer program 32 of the big data stream processing method mainly includes: judging whether a target data stream is an active data stream in a big data stream; if the target data stream is an inactive data stream in the big data stream, marking the inactive data stream; if the target data stream is an active data stream in the big data stream, judging whether the active data stream is a persistent active data stream; if the active data stream is the persistent active data stream, storing and marking the persistent active data stream; if the active data stream is a non-persistent active data stream, temporarily storing and marking the non-persistent active data stream as a temporary active data stream. The computer program 32 can be divided into one or more modules / units, one or more modules / units are stored in the memory 31 and executed by the processor 30 to complete the present application. One or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program 32 in the electronic device 3. For example, the computer program 32 can be divided into the functions of the first judging module 201, the first marking module 202, the second judging module 203, the second marking module 204 and the third marking module 205 (modules in a virtual device), and the specific functions of each module are as follows: the first judging module 201 is used for judging whether a target data stream is an active data stream in a big data stream; the first marking module 202 is used for marking an inactive data stream if the target data stream is an inactive data stream in the big data stream; the second judging module 203 is used for judging whether an active data stream is a persistent active data stream if the target data stream is an active data stream in the big data stream; the second marking module 204 is used for storing and marking a persistent active data stream if the active data stream is the persistent active data stream; and the third marking module 205 is used for temporarily storing and marking a non-persistent active data stream as a temporary active data stream if the active data stream is the non-persistent active data stream.

[0079] The electronic device 3 can include but is not limited to the processor 30 and the memory 31. Those skilled in the art can understand that, Figure 3 The electronic device 3 is only an example and does not constitute a limitation on the electronic device 3, and can include more or fewer components than the illustration, or combine certain components, or different components, for example, the electronic device can also include an input / output device, a network access device, a bus, etc.

[0080] The processor 30 can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0081] The memory 31 can be an internal storage unit of the electronic device 3, such as a hard disk or a memory of the electronic device 3. The memory 31 can also be an external storage device of the electronic device 3, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 3. Further, the memory 31 can include both the internal storage unit and the external storage device of the electronic device 3. The memory 31 is used to store computer programs and other programs and data required by the electronic device. The memory 31 can also be used to temporarily store data that has been output or will be output.

[0082] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is exemplified, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the above described functions. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit, and the integrated unit can be realized in the form of hardware or in the form of software. In addition, the specific names of each functional unit and module are only for easy distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the above device can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0083] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0084] Those skilled in the art can understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0085] In the embodiments provided in the present application, it should be understood that the disclosed apparatuses / devices and methods can be implemented in other ways. For example, the above-described apparatus / device embodiments are merely illustrative, for example, the division of modules or units is merely a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0086] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.

[0087] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0088] The integrated modules / units, if implemented in the form of software function units and sold or used as independent products, can be stored in a storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware. The computer program of the big data stream processing method can be stored in a storage medium. When the computer program is executed by a processor, the steps of each method embodiment can be implemented, that is, it is determined whether a target data stream is an active data stream in a big data stream; if the target data stream is an inactive data stream in the big data stream, the inactive data stream is marked; if the target data stream is an active data stream in the big data stream, it is determined whether the active data stream is a persistent active data stream; if the active data stream is a persistent active data stream, the persistent active data stream is stored and marked; and if the active data stream is a non-persistent active data stream, the non-persistent active data stream is temporarily stored and marked as a temporary active data stream. The computer program includes computer program code, which can be in the form of source code, object code, an executable file, or some intermediate form. The storage medium can include any entity or device capable of carrying computer program code, a recording medium, a U disk, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the storage medium does not include electrical carrier signals and telecommunication signals.

[0089] The above embodiments are only used to illustrate the technical solutions of the present application, but not limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application. The above specific embodiments further illustrate the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application, and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A method for processing large data streams, characterized by, The method comprises: The method comprises: determining a dimension for evaluating whether the target data stream is an active data stream in the big data stream; calculating an activity index of the target data stream in a time window w j according to the dimension for evaluating whether the target data stream is an active data stream in the big data stream; and determining that the target data stream is an active data stream in the big data stream if the activity index of the target data stream in the time window w j exceeds a first preset threshold; wherein the dimension for evaluating whether the target data stream is an active data stream in the big data stream comprises an update frequency of the target data stream per unit time, an update frequency of the big data stream per unit time, and diversity or uncertainty of the target data stream. if the target data stream is an inactive data stream in the large data stream, marking the inactive data stream; if the target data stream is an active data stream in the large data stream, determining whether the active data stream is a persistent active data stream, the determination of whether the active data stream is a persistent active data stream comprising: determining a dimension for evaluating whether the target data stream is an active data stream in the large data stream; calculating an activity index of the target data stream in multiple time windows according to the dimension for evaluating whether the target data stream is an active data stream in the large data stream; if the activity index of the target data stream in the multiple time windows exceeds a second preset threshold, determining that the target data stream is the persistent active data stream; if the active data stream is a persistent active data stream, storing and marking the persistent active data stream; if the active data stream is a non-persistent active data stream, temporarily storing and marking the non-persistent active data stream as a temporary active data stream; The method further includes: extracting an activity index from historical big data streams and calculating the activity of all big data streams; obtaining statistical characteristics of the activity of all big data streams, calculating their mean μ and standard deviation σ; and calculating the activity index according to the formula θ active =μ+α1*σ to calculate θ active , the θ active Determine as the first preset threshold, the α1 is the first adjustment coefficient; obtain the activity of the target data stream in T historical time windows; calculate the weighted average activity A of the activity of the target data stream in T historical time windows weighted According to the formula θ persistent =A weighted +α2*β to calculate θ persistent , the θ persistent is determined as the second preset threshold, α2 is the second adjustment coefficient, and β is the rate of change of activity over time obtained by linear regression.

2. The method of claim 1, wherein, the diversity or uncertainty of the target data stream is determined by: obtaining a set S composed of events contained in the target data stream; calculating a presence probability p(s j in the set S in the time window w i in the target data stream i ); According to the formula The diversity or uncertainty H of the target data stream is calculated target , where k is the number of categories of events contained in the set S.

3. A big data stream processing apparatus, characterized by, The device comprises: The first judging module is used for judging whether the target data stream is an active data stream in a large data stream, and specifically comprises: determining a dimension for evaluating whether the target data stream is an active data stream in a large data stream; calculating an activity index of the target data stream in a time window w j according to the dimension for evaluating whether the target data stream is an active data stream in a large data stream; if the activity index of the target data stream in the time window w j exceeds a first preset threshold, determining that the target data stream is an active data stream in the large data stream; and the dimension for evaluating whether the target data stream is an active data stream in a large data stream comprises an update frequency of the target data stream per unit time, an update frequency of the large data stream per unit time, and diversity or uncertainty of the target data stream. a first marking module configured to mark the inactive data stream if the target data stream is an inactive data stream in the large data stream; a second determining module configured to determine whether the active data stream is a persistent active data stream if the target data stream is an active data stream in the large data stream, the determination of whether the active data stream is a persistent active data stream comprising: determining a dimension for evaluating whether the target data stream is an active data stream in the large data stream; calculating an activity index of the target data stream in multiple time windows according to the dimension for evaluating whether the target data stream is an active data stream in the large data stream; if the activity index of the target data stream in the multiple time windows exceeds a second preset threshold, determining that the target data stream is the persistent active data stream; a second marking module configured to store and mark the persistent active data stream if the active data stream is a persistent active data stream; a third marking module configured to temporarily store and mark the non-persistent active data stream as a temporary active data stream if the active data stream is a non-persistent active data stream; The device further comprises a module for extracting the activity index from historical big data streams and calculating the activity of all big data streams; obtaining statistical characteristics of the activity of all big data streams, calculating the mean μ and the standard deviation σ; calculating θ according to the formula θ = μ + α1*σ, determining the θ as the first preset threshold, wherein the α1 is a first adjustment coefficient; obtaining the activity of the target data stream in T historical time windows; calculating the weighted average activity A of the activity of the target data stream in T historical time windows; calculating θ according to the formula θ = A + α2*β, determining the θ as the second preset threshold, wherein the α2 is a second adjustment coefficient, and the β is the rate of change of the activity over time obtained by linear regression. active active active weighted persistent weighted persistent persistent ​​​​​​​​ 4. An electronic device, the device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to realize the steps of the method of any one of claims 1 to 2.

5. A storage medium storing a computer program, characterized by The computer program is executed by the processor to realize the steps of the method of any one of claims 1 to 2.

Citation Information

Patent Citations

  • Collaborative flow identification method, system and server using said method

    CN107181724A

  • Data flow management method, network equipment and storage medium

    CN112838989A