Data processing inspection method and related product
By injecting simulated data into the data stream and utilizing the processing strategy of the online processing system, the problem of low efficiency of real-time online data processing quality inspection in the existing technology is solved, and efficient and accurate data processing inspection is achieved, which is suitable for a wide range of data flow scenarios.
Patent Information
- Application Number
- CN202510388462.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-24
AI Technical Summary
The existing technology lacks effective real-time online data processing quality inspection solutions, resulting in low inspection efficiency, limited applicable scenarios, large resource consumption, complex implementation and low accuracy.
By injecting simulation data into the data stream regularly, the simulation data and real data are processed by the online processing system to obtain the processing results of the simulation data and the processing results of the real data, and determine the reference processing results based on the processing strategy and the number of simulation data. If the processing results of the simulation data are inconsistent with the reference results, it is judged that the processing results of the real data are abnormal.
It realizes continuous inspection of online data real-time processing results with high efficiency, low consumption and accurate without introducing additional technical frameworks, improves detection accuracy and is suitable for a wide range of data flow scenarios.
Smart Images

Figure CN120196662A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of computer technology, and particularly to a data processing inspection method and related products. Background Art
[0002] In some business scenarios, a set of real-time online data processing services is usually required. Through continuous data collection and analysis, business characteristics, business accumulation, and change trends over a period of time are mined, so as to provide a basis for real-time business decisions. Online real-time data processing is a double-edged sword. If errors occur during the real-time data processing process, it will lead to deviations in business decisions. Therefore, real-time online data processing requires a quality inspection ability to be able to detect data processing errors in a timely manner and issue relevant warnings.
[0003] Currently, there is no mature quality inspection scheme for real-time online data processing. Generally, an offline data quality assurance scheme is adopted, that is, an additional set of offline data processing processes is introduced to compare the results of the two sets of data processing.
[0004] However, such schemes have problems such as low inspection efficiency, limited applicable scenarios, high resource consumption, complex implementation, and low accuracy. Summary of the Invention
[0005] The purpose of the embodiments of this specification is to provide a data processing inspection method and related products, so as to realize continuous inspection of the real-time processing results of online data in a lower-cost, more efficient, and more accurate manner without introducing an additional technical framework and with almost no intrusion into the existing link.
[0006] To achieve the above purpose, the embodiments of this specification adopt the following technical solutions: In the first aspect, a data processing inspection method is provided, including: Regularly obtain the data stream within the current time window, and inject simulated data with the same data structure as the real data in the data stream into the data stream; Process the data stream injected with the simulated data through an online processing system, and obtain the processing strategy used by the online processing system and the processing results of the simulated data; Based on the processing strategy and the quantity of the simulated data, determine the reference processing results of the simulated data; If the processing results of the simulated data are inconsistent with the reference processing results, it is determined that the processing results of the real data are abnormal.
[0007] In the second aspect, a data processing inspection device is provided, including: An injection module, configured to periodically obtain data streams within a current time window and inject simulated data having the same data structure as the real data in the data streams into the data streams; A processing module, configured to process the data streams injected with the simulated data through an online processing system and obtain the processing strategy used by the online processing system and the processing results of the simulated data; A first determination module, configured to determine the reference processing results of the simulated data based on the processing strategy and the quantity of the simulated data; A second determination module, configured to determine that the processing results of the real data are different if the processing results of the simulated data are inconsistent with the reference processing results.
[0008] In a third aspect, there is provided an electronic device, including: A processor; A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the instructions to implement the data processing and inspection method provided in the first aspect.
[0009] In a fourth aspect, there is provided a computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the data processing and inspection method provided in the first aspect.
[0010] In a fifth aspect, there is provided a computer program product, the computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute some or all of the steps in the data processing and inspection method provided in the first aspect.
[0011] The solution of the embodiments of this specification takes into account that for any online data real-time processing product, the object of real-time processing is the data stream. By injecting simulated data into the data stream at regular intervals, the simulated data and the data stream are processed together by the online processing system to obtain the processing results of the simulated data and the real data in the data stream. Further, the reference processing result of the simulated data is determined according to the processing strategy used by the online processing system and the quantity of the simulated data. Since both the simulated data and the real data are processed by the online processing system and the processing strategy is the same, whether the processing result of the simulated data is consistent with the reference processing result can reflect whether there are problems in the production and processing process. Based on this, by comparing the reference processing result with the processing result of the simulated data, if they are inconsistent, it indicates that there are problems in the production and processing process, and then it is determined that the processing result of the real data is abnormal, thereby improving the detection accuracy. It can be seen that this method completes data inspection based on the existing online data processing framework, without introducing other technical frameworks or other independent processing processes, has almost no intrusion into the existing link, is simple to implement, has linearly controllable complexity, and reduces resource consumption. In addition, the entire process completes the tracking and detection of the data processing process by injecting simulated data into the data stream at regular intervals. It has high timeliness, can reach the minute level or even higher timeliness, has higher inspection efficiency, and has no requirements for the window period of the data stream, and can support rolling windows and sliding windows, so it is applicable to a wider range of scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings described herein are used to provide a further understanding of this specification, form a part of this specification, and the illustrative embodiments and descriptions thereof of this specification are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings: Figure 1 is a schematic diagram of the offline data quality assurance method currently adopted; Figure 2 is a schematic flowchart of a data processing inspection method provided by an embodiment of this specification; Figure 3 is a schematic diagram of a data stream injected with simulated data provided by an embodiment of this specification; Figure 4 is a schematic flowchart of a method for injecting simulated data provided by an embodiment of this specification; Figure 5 is a schematic flowchart of a data processing inspection method provided by another embodiment of this specification; Figure 6 is a schematic flowchart of a data processing inspection method provided by yet another embodiment of this specification; Figure 7Schematic structural diagram of a data processing inspection device provided for an embodiment of this specification; Figure 8 Schematic structural diagram of an electronic device provided for an embodiment of this specification. Detailed implementation manners
[0013] To make the objectives, technical solutions and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part rather than all of the embodiments of this specification. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this document.
[0014] The term "including" and its variations used in this document are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. The term "in response to" is used to indicate the conditions or states on which the operations performed depend. When the dependent conditions or states are met, one or more of the operations performed may be real-time or may have a set delay. Without special instructions, there is no limit on the order of the multiple operations performed.
[0015] It should be noted that the concepts such as "first" and "second" mentioned in this document are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependent relationships.
[0016] It should be noted that the modifiers such as "one" and "multiple" mentioned in this document are illustrative rather than restrictive. Those skilled in the art should understand that unless clearly stated otherwise in the context, it should be understood as "one or more".
[0017] The names of the messages or information exchanged between multiple devices in the implementation manners of this document are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0018] As mentioned above, there is currently no mature real-time online data processing quality inspection solution. Generally, an offline data quality assurance solution is adopted, that is, an additional offline data processing process is introduced to compare the results of the two data processing processes. As Figure 1 shown, the process is as follows: First, not only is the real-time data stream input into the online processing system for processing, but also the real-time data stream is input into the offline data processing system in a bypass manner to perform logical processing on the data offline.
[0019] Then, the data processed offline is output to the data warehouse, and the online data processing results are exported to the data warehouse at regular intervals.
[0020] Finally, based on the offline processing results, the offline processing results and the online processing results of the same data are compared to determine whether the online processing results of the data are abnormal.
[0021] However, the above solution has the following problems, which limit its practical application: (1) Low inspection efficiency: For a large amount of data, the offline processing usually takes more than an hour, and the result comparison process also takes at least an hour. Each inspection takes a long time and can only cover the data from an hour ago.
[0022] (2) Limited applicable scenarios: The online processing logic has a clear periodic attribute, and the above solution is only applicable to rolling windows. Since the data within the sliding window is updated slidingly rather than fixed, if the data accuracy and the window accuracy do not match, the window data statistically obtained is inaccurate.
[0023] (3) High resource consumption: Another set of offline processing logic needs to be introduced to verify the accuracy of the online processing results, resulting in at least doubling the consumed resources.
[0024] (4) Complex implementation: It is impossible to strictly ensure the consistency between the offline processing logic and the online processing logic. This is because the offline processing logic and the online processing logic usually belong to different technology stacks, and it is very difficult to ensure complete consistency in the implementation of all business semantic operators, such as statistical operators like TopN.
[0025] (5) Unable to cover the query link: The data processing results will be stored in the memory, and the online data processing system extracts the online data from the memory through the supporting query service or SDK (Software Development Kit) to assist in making decisions. The above solution can only cover the processing process of the production link and cannot cover the stability problems of the query service or the SDK link.
[0026] (6) Weak standard accuracy: When the online processing results are inconsistent with the offline processing results, it is possible that there is a problem with the offline processing results used as the standard.
[0027] In view of this, an embodiment of this specification proposes a data processing inspection scheme. Considering that for any online data real-time processing product, the object of its real-time processing is the data stream. By injecting simulated data into the data stream at regular intervals, the simulated data and the data stream are processed together by the online processing system to obtain the processing results of the simulated data and the processing results of the real data in the data stream. Further, according to the processing strategy used by the online processing system and the quantity of the simulated data, the reference processing result of the simulated data is determined. Since both the simulated data and the real data are processed by the online processing system and the processing strategy is the same, whether the processing result of the simulated data is consistent with the reference processing result can reflect whether there are problems in the production processing process. Based on this, by comparing the reference processing result with the processing result of the simulated data, if they are inconsistent, it indicates that there are problems in the production processing process, and then it is determined that the processing result of the real data is abnormal, thereby improving the detection accuracy. It can be seen that this method is based on the existing online data processing framework to complete data inspection, without introducing other technical frameworks or other independent processing processes, with almost no intrusion into the existing link, simple implementation, linearly controllable complexity, and reduced resource consumption. In addition, the entire process completes the tracking and detection of the data processing process by injecting simulated data into the data stream at regular intervals, with high timeliness, which can reach the minute level or even higher timeliness, higher inspection efficiency, and no requirement for the window period of the data stream, and can support rolling windows and sliding windows, so the applicable scenarios are wider.
[0028] Further, considering that in an online processing system, the processing results of the real data in the data stream are usually extracted from the memory through a supporting query service to assist in completing business decisions, and the accuracy of the query results will also affect the business decision effect. For this reason, after obtaining the processing results of the simulated data, the query service is also called to query the processing results of the simulated data from the memory. If the query results are inconsistent with the processing results of the simulated data, it is determined that the query service is abnormal. In this way, the data processing inspection method provided by the embodiment of this specification simultaneously covers the quality inspection of two links, namely the production link and the query link, and can cover a wider range of stability problems.
[0029] The following will describe in detail the technical solutions provided by each embodiment of this specification with reference to the accompanying drawings.
[0030] Please refer to Figure 2 , which is a schematic flowchart of a data processing inspection method provided by an embodiment of this specification. The method includes the following steps: S202, regularly obtain the data stream within the current time window, and inject simulated data with the same data structure as the real data in the data stream into the data stream.
[0031] Specifically, according to the specified time window, a certain length of data stream is intercepted from the real-time data stream at regular intervals. This data stream is the data stream within the current time window. The time window here can be a sliding window or a rolling window. The embodiments of this specification do not limit the type of the time window. The length of the time window and the interval duration of the regular time points can be adjusted according to actual needs. For example, every 10 minutes, the data stream received in the past 15 minutes is intercepted from the real-time data stream. In this case, the past 15 minutes is the current time window, and 10 minutes is the interval duration of the regular time points.
[0032] The data in the data stream is called real data. Simulated data is a kind of data with special marks, also called "buoys". Simulated data has the same data structure as real data, only the data content is different. Figure 3 A data stream into which simulated data is injected is shown.
[0033] By injecting simulated data into the data stream within the current time window, the simulated data and the data stream are processed together, and thus can play a role in dyeing the data stream. In this way, according to the simulated data, the process of data processing can be tracked, and the inspection of data quality can be completed.
[0034] In the above S202, the simulated data can be constructed and injected into the data stream in various appropriate ways, and the embodiments of this specification do not limit this.
[0035] In one implementation, the data structure of the real data includes at least one field. In this regard, a candidate data including the same fields is constructed; then, based on the preset simulated data content, the field content of each field is filled, and thus the simulated data with the same data structure as the real data is obtained.
[0036] In another implementation, the injection of the simulated data with the same data structure as the real data in the data stream into the data stream includes the following steps: S221, count the events from which the real data in the data stream comes, and obtain an event set.
[0037] An event of a real data source refers to an event that triggers the generation of real data. For example, taking real transaction data as an example, assume that there are two real transaction data belonging to user A and the data content is the same. However, one of the real transaction data is generated after user a makes an order payment through payment tool 1. Then, the event related to the source of this real transaction data is: making a payment through payment tool 1. The other real transaction data is generated after user A makes an order payment through payment tool 2. Then, the event related to the source of this real transaction data is: making a payment through payment tool 2. It can be seen that although these two real transaction data both belong to user A and have the same data content, the events of their sources are different, and they can be distinguished by the events of their sources.
[0038] S222. For each event in the event set, construct simulated data that originates from this event and has the same data structure as the real data.
[0039] Exemplarily, the data structure of the real data includes at least a first field and a second field. The field content of the first field represents the event of the real data source, and the field content of the second field represents the data content of the real data.
[0040] In this case, first, construct candidate data that includes at least a first field and a second field. Then, for each event in the event set, fill the field content of the first field in the candidate data based on this event, and fill the field content of the second field based on the preset simulated data content, to obtain simulated data that originates from this event and has the same data structure as the real data.
[0041] For the first field in the candidate data, the event information of the event (such as event identifier, event name, basic attributes, and description) can be used as the field content of this first field.
[0042] For the second field in the candidate data, as an example, the simulated data content can be used as the field content of the second field. As another example, in order to distinguish the simulated data from the real data to avoid interference and affect the subsequent inspection accuracy, the simulated data can be subjected to a hash operation to obtain a first hash value, and the first hash value can be used as the field content of the second field; or, the simulated data content can be concatenated with a preset identifier and used as the field content of the second field. Here, the identifier refers to a symbol that can be distinguished from the data content of the real data, and it can be a special string. The identifier can be concatenated with the simulated data content either as a prefix or as a suffix.
[0043] It should be noted that the data structure of the real data may further include more fields, such as a third field representing the data identifier, a fourth field representing the generation time, etc. The embodiments of this specification do not limit this. In this case, the constructed simulated data also includes these fields.
[0044] S223, inject the constructed simulated data into the data stream at a preset frequency.
[0045] The preset frequency can be set according to actual needs, and the embodiments of this specification do not limit this. For example, assuming that the length of the current time window is 1 second and the data stream within the current time window contains 100,000 real data, then the preset frequency can be 1 per second, which is equivalent to adding a simulated data set among 100,000 real data per second. The resource consumption is negligible, and the amount of consumed resources will not change drastically as the processing strategies are continuously added.
[0046] In practical applications, there are various processing strategies for the data stream, such as counting the number of user requests within a certain window (i.e., counting), the trading volume of a store (i.e., summing up), the categories of transactions (i.e., classifying and counting according to specified dimensions), and so on. These processing strategies are usually associated with the event dimension, that is, based on the events from which the real data in the data stream comes, the real data in the data stream is statistically analyzed. Therefore, it is more appropriate to construct simulated data at the event granularity. That is to say, when constructing simulated data, there is no need to care about the specific details of the real data and the processing strategies, but to construct simulated data from the data stream dimension, only need to count how many events the real data in the data stream comes from, and construct a simulated data for each event. In this way, when processing the simulated data injected into the data stream, the simulated data will also be statistically analyzed according to the events from which the simulated data comes, so that the inspection results based on the simulated data are more in line with the processing process of the real data, and the inspection accuracy is higher.
[0047] In practical applications, the above operations of obtaining the data stream, constructing simulated data, and injecting simulated data can be executed through a scheduled task. Specifically, as Figure 4 shown, start a scheduled task, and perform the following operations through this scheduled task: First, obtain the data stream within the current time window and count the events from which the real data in the data stream comes; then, loop and perform the following operations: construct corresponding simulated data with the same data structure as the real data for each event, inject the simulated data into the data stream, and retry after the injection fails.
[0048] S204, process the data stream injected with simulated data through an online processing system, and obtain the processing strategies used by the online processing system and the processing results of the simulated data.
[0049] The online processing system is pre - configured with processing strategies. The processing strategies are used to describe the processing methods for data streams. In applications, the processing strategies can, for example, include but are not limited to: counting (such as counting the number of user requests within the current time window), summing up (such as counting the trading volume of stores within the current time window), classifying and counting according to specified dimensions (such as counting the transaction categories within the current time window), etc.
[0050] Send the data stream to which simulated data will be injected to the online processing system. The online processing system processes such data streams based on the processing strategies and outputs the processing results of real data and the processing results of simulated data. In addition, pull the processing strategies used within the current time window from the configuration information of the online processing system.
[0051] Exemplarily, if the processing strategy is to count the number of data within the current time window, then count the number of real data in the data stream to obtain the processing result of real data, and count the number of simulated data injected into the data stream to obtain the processing result of simulated data.
[0052] If the processing strategy is to sum the number of data within the current time window and then multiply by a value and sum up, then count the number of real data in the data stream, multiply this number by the value to obtain the processing result of real data, and count the number of simulated data injected into the data stream, multiply this data by the value to obtain the processing result of simulated data.
[0053] In the case where the processing strategy is to classify and count according to specified dimensions, if the data content of the simulated data does not contain information about the specified dimension, then the simulated data cannot fully cover the processing strategy. This will lead to inaccurate processing results of the simulated data, and further affect the inspection accuracy. For example, each transaction data has a corresponding payment method. A processing strategy is to count the number of transaction data corresponding to each payment method. Since there are multiple payment methods corresponding to transaction data and they cannot be enumerated, the constructed simulated transaction data cannot cover all payment methods, resulting in inaccurate statistical results for the simulated transaction data, and further unable to accurately determine whether the processing results of real transaction data under this processing strategy are accurate.
[0054] Therefore, in the case where the simulated data cannot fully cover the processing strategy, by identifying and specially processing the simulated data injected into the data stream, the inspection of this processing strategy is achieved. Specifically, each simulated data has a corresponding generation timestamp. In one implementation, if the processing strategy is to classify and count according to specified dimensions and the data content of the simulated data does not contain information about the specified dimension, then the simulated data with the same generation timestamp are divided into the same category; determine the number of simulated data in each category as the processing result of the simulated data.
[0055] It can be seen that processing the simulation data in this processing manner is equivalent to replacing the specified dimension with the generation timestamp, so as to classify and count the simulation data according to the generation timestamp, which can effectively avoid the problem of inaccurate detection results caused by the inability of the simulation data to cover all the information of the specified dimension. In the application, the generation timestamp of the simulation data can be the timestamp when the simulation data is injected into the data stream. In this way, the generation timestamps of the simulation data injected into the data stream are different, and the processing results of the simulation data obtained therefrom can more accurately reflect whether the processing process based on the above processing strategy is abnormal, and further more accurately determine whether the processing results of the real data in the data stream under this processing strategy are abnormal.
[0056] S206. Determine the reference processing result of the simulation data based on the processing strategy and the quantity of the simulation data.
[0057] In the above processing strategy, whether it is counting, summarizing and summing up, or classifying and counting according to the specified dimension, the processing result of the simulation data is related to the quantity of the simulation data. Therefore, based on the processing strategy and the quantity of the simulation data, the reference processing result of the simulation data can be determined.
[0058] In one implementation, the above S206 includes the following steps: determine the operator coefficient matching the processing strategy; determine the reference processing result of the simulation data based on the product between the operator coefficient and the quantity of the simulation data.
[0059] In the above processing strategy, whether it is counting, summarizing and summing up, or classifying and counting according to the specified dimension, the processing of the simulation data can be regarded as multiplying the quantity of the simulation data by a value. Therefore, this value can be determined as the operator coefficient matching the processing strategy, and the product between the operator coefficient and the quantity of the simulation data is the reference processing result of the simulation data. The reference processing result of the simulation data is also called the target value, denoted as f(target value)=quantity of the simulation data * operator coefficient.
[0060] For example, if the processing strategy is counting, the processing of the simulation data can be regarded as multiplying the quantity of the simulation data by 1. Therefore, the operator coefficient matching this processing strategy is 1, that is, the reference processing result of the simulation data = quantity of the simulation data.
[0061] If the processing strategy is to sum up the quantity of data in the current time window and then multiply by a value and sum up, the processing of the simulation data can be regarded as multiplying the quantity of the simulation data by this value. Therefore, this value is the operator coefficient matching this processing strategy.
[0062] If the processing strategy is to classify and count according to a specified dimension, and the data content of the simulated data does not contain information about the specified dimension, then the specified dimension is replaced by the generation timestamp, that is, the simulated data is classified and counted according to the generation timestamp. Since the simulated data is periodically injected into the data stream, the processing of the simulated data can be regarded as multiplying the reciprocal of the number of generation timestamps by the number of simulated data. In this case, count the number of generation timestamps of the simulated data injected into the data stream, and determine the operator coefficient matching the processing strategy based on this number. Specifically, determine the reciprocal of this number as the operator coefficient matching the processing strategy.
[0063] S208. If the processing result of the simulated data is inconsistent with the reference processing result, determine that the processing result of the real data is abnormal.
[0064] If the processing result of the simulated data is inconsistent with the reference processing result, it indicates that the processing result of the simulated data is abnormal. Since the simulated data and the real data in the data stream are processed by the online processing system together, and the online processing system has the same processing strategy for both, it can be determined that the processing result of the real data is also abnormal. In this case, the first warning message can be output to indicate that the processing result of the real data is abnormal.
[0065] It can be seen that the entire patrol inspection process is completed based on the existing online processing framework, without the need to introduce other technical frameworks or other independent processing processes, with almost no intrusion into the existing processing link, simple implementation, linearly controllable complexity, and reduced resource consumption. Secondly, since both the simulated data and the real data are processed by the online processing system and have the same processing strategy, whether the processing result of the simulated data is consistent with the reference processing result can accurately reflect whether the processing result of the real data is abnormal, thus improving the detection accuracy. In addition, the entire process completes the tracking and detection of the data processing process by periodically injecting simulated data into the data stream. It has high timeliness, can reach the minute level or even higher timeliness, higher patrol inspection efficiency, has no requirement for the window period of the data stream, and can support rolling windows and sliding windows, so it is applicable to a wider range of scenarios.
[0066] In another embodiment of this specification, considering that the processing strategy usually requires an idempotency mechanism to prevent duplicate processing of data, that is, processing the same data based on the same processing strategy will not result in different processing results. For this reason, when constructing the simulated data, the same simulated data can be injected into the data stream with a certain probability, so as to verify whether the idempotency mechanism is effective.
[0067] Specifically, the number of pieces of simulated data injected into the data stream is multiple, and some of the simulated data among these simulated data can be the same. In this case, after the above S208, the following steps may further be included: comparing the processing results of the same simulated data; if the comparison is inconsistent, it is determined that there is an idempotency anomaly in the processing strategy used by the online processing system. Among them, the same simulated data refers to the simulated data that is the same in multiple dimensions such as the same source event, the same data content, and the same generated timestamp.
[0068] Further, a second warning message may be output to indicate that there is an idempotency anomaly in the processing strategy.
[0069] In another embodiment of this specification, after the above S208, the following steps may further be included: writing the processing result of the simulated data into a memory; invoking a query service to query the processing result of the simulated data from the memory; if the queried processing result is inconsistent with the processing result of the simulated data, it is determined that the query service is abnormal.
[0070] Further, a third warning message may also be output to indicate that the query service is abnormal.
[0071] The query service here refers to the query service supporting the online processing system. In some business scenarios, in the online processing system, the processing results of the real data in the data stream are usually extracted from the memory through the supporting query service to assist in completing business decisions, and the accuracy of the query results will also affect the business decision-making effect. After obtaining the processing result of the simulated data, by writing the processing result into the memory and invoking the query service to query the processing result of the simulated data from the memory, if the queried processing result is inconsistent with the processing result of the simulated data, it indicates that there is a problem in the query process, and then it is determined that the query service is abnormal. In this way, the data processing inspection method provided by the embodiments of this specification simultaneously covers the quality inspection of both the production link and the query link, and can cover a wider range of stability problems.
[0072] The data processing inspection method provided by the embodiments of this specification can be applied to various business scenarios with online data processing inspection requirements. Taking risk control as an example, risk control plays a crucial role in the fields of Internet finance and e-commerce. In order to better conduct real-time confrontation with black and gray production, the risk control system usually has a set of real-time online data processing services (such as data accumulation, profiling, black and white lists). Through continuous data collection and analysis, risk characteristics and the accumulation and change trends of risks within a period of time are mined, and based on this, vouchers are provided for real-time risk control decisions.
[0073] Through the data processing inspection method provided by the embodiments of this specification, it is possible to continuously inspect the real-time processing results of online data in a lower-cost, more efficient, and more accurate manner without introducing an additional technical framework and with almost no intrusion into the existing link. The inspection results can provide reliable data support for the decision-making of the risk control system.
[0074] To facilitate the understanding of the above data processing inspection method, the following combines Figure 5 and Figure 6 , and uses a specific embodiment to illustrate the data processing inspection process.
[0075] As Figure 5 shown, in the data injection link, simulated data is constructed and injected into the data stream. In the data processing link, the data stream injected with simulated data is processed by the online processing system to obtain the processing results of real data and the processing results of simulated data, and these processing results are written into the memory.
[0076] In the data inspection link, according to the processing strategy used by the online processing system, the reference processing result of the simulated data is calculated. Further, on the one hand, the reference processing result of the simulated data is compared with the processing result output by the online processing system. If they are inconsistent, it indicates that there is a problem in the processing link, and then it is determined that the processing result of the real data is abnormal; on the other hand, the query service is also called to query the processing result of the simulated data in the memory, and the query result is compared with the processing result output by the online processing system. If they are inconsistent, it indicates that the query service is abnormal.
[0077] As Figure 6 shown, in the data inspection stage, through a timed task, the data stream within the current time window is obtained at regular intervals, and simulated data with the same data structure as the real data in the data stream is injected into the data stream.
[0078] Then, the data stream injected with simulated data is processed by the online processing system.
[0079] Further, the following operations are executed concurrently: pull the processing strategy used by the online processing system, calculate the reference processing result of the simulated data, obtain the processing result of the simulated data, and compare the processing result of the simulated data with the reference processing result.
[0080] When calculating the reference processing result of the simulated data, the number of simulated data injected into the data stream is counted, and based on the processing strategy and the number of simulated data, the reference processing result of the simulated data is determined.
[0081] When obtaining the processing result of the simulated data, not only the processing result of the simulated data output by the online processing system is obtained, but also the query service is called to query the processing result of the simulated data.
[0082] When comparing the processing results of the simulation data with the reference processing results, the processing results output by the online processing system are compared with the reference processing results of the simulation data. If they are consistent, it is skipped; if they are inconsistent, it is determined that the processing results of the real data are abnormal, and the corresponding warning information is output. In addition, the processing results output by the online processing system are also compared with the queried processing results. If they are consistent, it is skipped; if they are inconsistent, it is determined that the query service is abnormal, and the corresponding warning information is output.
[0083] In addition, corresponding to the data processing inspection method described above Figure 2 An embodiment of this specification also provides a data processing inspection device. Figure 7 FIG. 8 is a schematic structural diagram of a data processing inspection device 700 provided by an embodiment of this specification, including: an injection module 710, a processing module 720, a first determination module 730, and a second determination module 740.
[0084] The injection module 710 is configured to periodically obtain a data stream within a current time window, and inject simulation data having the same data structure as the real data in the data stream into the data stream.
[0085] The processing module 720 is configured to process the data stream injected with the simulation data through an online processing system, and obtain the processing strategy used by the online processing system and the processing results of the simulation data.
[0086] The first determination module 730 is configured to determine the reference processing results of the simulation data based on the processing strategy and the quantity of the simulation data.
[0087] The second determination module 740 is configured to determine that the processing results of the real data are abnormal if the processing results of the simulation data are inconsistent with the reference processing results.
[0088] The data processing inspection device provided by the embodiments of this specification, considering that no matter what kind of online data real-time processing product is, the object of its real-time processing is the data stream. By injecting simulated data into the data stream at regular intervals, the simulated data and the data stream are processed by the online processing system together, and the processing results of the simulated data and the processing results of the real data in the data stream are obtained. Further, the reference processing result of the simulated data is determined according to the processing strategy used by the online processing system and the quantity of the simulated data. Since both the simulated data and the real data are processed by the online processing system and the processing strategy is the same, whether the processing result of the simulated data is consistent with the reference processing result can reflect whether there are problems in the generation and processing process. Based on this, by comparing the reference processing result with the processing result of the simulated data, if they are inconsistent, it indicates that there are problems in the production and processing process, and then it is determined that the processing result of the real data is abnormal, thereby improving the detection accuracy. It can be seen that this method completes data inspection based on the existing online data processing framework, without introducing other technical frameworks or other independent processing processes, has almost no intrusion into the existing link, is simple to implement, has linearly controllable complexity, and reduces resource consumption. In addition, the entire process completes the tracking and detection of the data processing process by injecting simulated data into the data stream at regular intervals. Its timeliness is high, which can reach the minute level or even higher timeliness, the inspection efficiency is higher, and there is no requirement for the window period of the data stream, and it can support rolling windows and sliding windows, so the applicable scenarios are wider.
[0089] In another embodiment, the injection module includes: The first statistical sub-module is used to count the events from which the real data in the data stream originates, and obtain an event set; The construction sub-module is used to construct, for each event in the event set, simulated data that originates from the event and has the same data structure as the real data; The injection sub-module is used to inject the constructed simulated data into the data stream at a preset frequency.
[0090] In another embodiment, the data structure of the real data includes at least a first field and a second field. The field content of the first field represents the event from which the real data originates, and the field content of the second field represents the data content of the real data; The construction sub-module is used for: Construct candidate data including at least the first field and the second field; For each event in the event set, fill the field content of the first field in the candidate data based on the event, and fill the field content of the second field based on the preset simulated data content, to obtain simulated data that originates from the event and has the same data structure as the real data.
[0091] In another embodiment, the construction sub-module fills the field content of the second field in the following manner: Performing a hash operation on the analog data content to obtain a first hash value, and using the first hash value as the field content of the second field; or, Concatenating the analog data content with a preset identifier and using the result as the field content of the second field.
[0092] In another embodiment, the analog data has a corresponding generation timestamp; The processing result of the analog data is obtained by processing the analog data in the following manner: If the processing strategy is to perform classification and statistics according to a specified dimension, and the data content of the analog data does not contain information about the specified dimension, then the analog data with the same generation timestamp is divided into the same category; Determining the quantity of the analog data in each category as the processing result of the analog data.
[0093] In another embodiment, the first determination module includes: A first determination sub-module, configured to determine an operator coefficient matching the processing strategy; A second determination sub-module, configured to determine a reference processing result of the analog data based on the product of the operator coefficient and the quantity of the analog data.
[0094] In another embodiment, the analog data has a corresponding generation timestamp; The first determination sub-module is configured to: If the processing strategy is to perform classification and statistics according to a specified dimension, and the data content of the analog data does not contain information about the specified dimension, then count the quantity of the generation timestamps of the analog data injected into the data stream; Based on the quantity of the generation timestamps, determine an operator coefficient matching the processing strategy.
[0095] In another embodiment, the quantity of the analog data is multiple, and some of the multiple analog data are the same; The data processing inspection device 700 further includes: A comparison module, configured to compare the processing results of the same analog data; A third determination module, configured to determine that there is an idempotency anomaly in the processing strategy if the comparison is inconsistent.
[0096] In another embodiment, the data processing inspection device 700 further includes: A storage module for writing the processing result of the analog data into a memory; A query module for calling a query service to query the processing result of the analog data from the memory; A fourth determination module for determining that the query service is abnormal if the queried processing result is inconsistent with the processing result of the analog data.
[0097] Obviously, the data processing inspection device in the embodiments of this specification can be used as the execution subject of the data processing inspection method shown above. Figure 2 Therefore, it can implement the functions achieved by the data processing inspection method. Since the principle is the same, it will not be elaborated here. Figure 2
[0098] Figure 8 Figure 8 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of this specification. Please refer to , at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.
[0099] The processor, the network interface, and the memory can be interconnected through an internal bus, and the internal bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 8 only a bidirectional arrow is used in [FIGURE], but it does not mean that there is only one bus or one type of bus.
[0100] The memory is used to store a program. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory can include a memory and a non-volatile memory, and provide instructions and data to the processor.
[0101] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a data processing inspection device logically. The processor executes the program stored in the memory and is specifically used to perform the following operations: Regularly obtain the data stream within the current time window, and inject simulated data with the same data structure as the real data in the data stream into the data stream; Process the data stream injected with the simulated data through an online processing system, and obtain the processing strategy used by the online processing system and the processing result of the simulated data; Based on the processing strategy and the quantity of the simulated data, determine the reference processing result of the simulated data; If the processing result of the simulated data is inconsistent with the reference processing result, determine that the processing result of the real data is abnormal.
[0102] The method executed by the data processing inspection device disclosed in the above embodiments of this specification Figure 2 can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software. The above processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this specification. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of this specification can be directly implemented by the hardware decoding processor, or completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0103] It should be understood that the electronic device in the embodiments of this specification can implement the functions of the data processing inspection device in Figure 2 the embodiments shown. Due to the same principle, it will not be elaborated herein in the embodiments of this specification.
[0104] Of course, in addition to the software implementation, the electronic devices described in this specification do not exclude other implementation manners, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and may also be hardware or logic devices.
[0105] An embodiment of this specification also provides a computer-readable storage medium. The computer-readable storage medium stores one or more programs. The one or more programs include instructions that, when executed by an electronic device including multiple application programs, can enable the electronic device to execute Figure 2 the method of the illustrated embodiment, and specifically used to perform the following operations: Regularly obtain the data stream within the current time window, and inject analog data having the same data structure as the real data in the data stream into the data stream; Process the data stream injected with the analog data through an online processing system, and obtain the processing strategy used by the online processing system and the processing result of the analog data; Based on the processing strategy and the quantity of the analog data, determine the reference processing result of the analog data; If the processing result of the analog data is inconsistent with the reference processing result, determine that the processing result of the real data is abnormal.
[0106] An embodiment of this specification also provides a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program. The computer program is operable to cause a computer to execute some or all of the steps in the data processing inspection method provided by the embodiment of this specification.
[0107] The specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0108] In summary, the above description is only a preferred embodiment of this specification and is not intended to limit the protection scope of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included in the protection scope of this specification.
[0109] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0110] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0111] It should also be noted that the term "comprises," "comprising," or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, commodity, or device that comprises a list of elements does not include only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, commodity, or device. Without further limitation, an element defined by the statement "comprising an..." does not preclude the presence of additional identical elements in the process, method, commodity, or device that comprises the element.
[0112] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the corresponding part of the method embodiment for the relevant content.
Claims
1. A data processing inspection method, characterized in that: include: Obtaining the data stream in the current time window at a fixed time, and injecting simulated data having the same data structure as the real data in the data stream into the data stream; Processing the data stream injected into the simulation data through an online processing system, and obtaining a processing strategy used by the online processing system and a processing result of the simulation data; Determining a reference processing result of the simulation data based on the processing strategy and the amount of the simulation data; If the processing result of the simulation data is inconsistent with the reference processing result, it is determined that the processing result of the real data is abnormal.
2. The method according to claim 1, characterized in that The step of injecting simulated data having the same data structure as real data in the data stream into the data stream comprises: Counting events of real data sources in the data stream to obtain an event set; For each event in the event set, constructing simulated data derived from the event and having the same data structure as the real data; The constructed simulation data is injected into the data stream at a preset frequency.
3. The method according to claim 2, characterized in that The data structure of the real data at least includes a first field and a second field, the field content of the first field represents the event of the source of the real data, and the field content of the second field represents the data content of the real data; The step of constructing, for each event in the event set, simulated data derived from the event and having the same data structure as the real data, comprises: constructing candidate data including at least the first field and the second field; For each event in the event set, the field content of the first field in the candidate data is filled based on the event, and the field content of the second field is filled based on the preset simulation data content, so as to obtain simulation data derived from the event and having the same data structure as the real data.
4. The method according to claim 3, characterized in that The filling the field content of the second field based on the preset simulation data content includes: Performing a hash operation on the simulated data content to obtain a first hash value, and using the first hash value as the field content of the second field; or, The simulated data content and the preset identifier are concatenated to serve as the field content of the second field.
5. The method according to claim 1, characterized in that The simulation data has a corresponding generation timestamp; The processing result of the simulation data is obtained by processing the simulation data in the following manner: If the processing strategy is to classify and count according to a specified dimension, and the data content of the simulation data does not contain information of the specified dimension, the simulation data with the same generation timestamp are classified into the same category; The amount of simulation data of each category is determined as a processing result of the simulation data.
6. The method according to claim 1, characterized in that The step of determining a reference processing result of the simulation data based on the processing strategy and the amount of the simulation data includes: determining operator coefficients matching the machining strategy; Based on the product between the operator coefficient and the amount of the simulation data, a reference processing result of the simulation data is determined.
7. The method according to claim 6, characterized in that The simulation data has a corresponding generation timestamp; The determining of the operator coefficients matching the processing strategy comprises: If the processing strategy is to perform classification statistics according to a specified dimension, and the data content of the simulation data does not contain information of the specified dimension, then the number of generation timestamps of the simulation data injected into the data stream is counted; Based on the number of generated timestamps, operator coefficients matching the machining strategy are determined.
8. The method according to any one of claims 1 to 7, characterized in that There are multiple simulation data, and some of the multiple simulation data are the same; After determining the reference processing result of the simulation data based on the processing strategy and the amount of the simulation data, the method further includes: Compare the processing results of the same simulation data; If the comparison is inconsistent, it is determined that the processing strategy has an idempotency anomaly.
9. The method according to any one of claims 1 to 7, characterized in that After obtaining the processing strategy used by the online processing system and the processing result of the simulation data, the method further includes: Writing the processing result of the simulation data into a memory; Invoking a query service to query the processing result of the simulation data from the memory; If the queried processing result is inconsistent with the processing result of the simulation data, it is determined that the query service is abnormal.
10. A data processing inspection device, characterized in that: include: An injection module is used to periodically obtain a data stream in a current time window and inject simulated data having the same data structure as real data in the data stream into the data stream; A processing module, used for processing the data stream injected into the simulation data through an online processing system, and obtaining a processing strategy used by the online processing system and a processing result of the simulation data; A first determination module, configured to determine a reference processing result of the simulation data based on the processing strategy and the amount of the simulation data; The second determining module is configured to determine that the processing result of the real data is different if the processing result of the simulation data is inconsistent with the reference processing result.
11. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the data processing and inspection method as described in any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that: When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the data processing and inspection method as described in any one of claims 1 to 9.
13. A computer program product, characterized in that The computer program product includes a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to execute part or all of the steps in the data processing inspection method according to any one of claims 1 to 9.