Streaming data real-time calculation tracking method, abnormal data troubleshooting method and system thereof
By obtaining and recording the full-link calculation related data of stream data in streaming calculation, the problem of inefficient abnormal data detection in streaming calculation is solved, and efficient and accurate abnormal data positioning and problem positioning are achieved.
Patent Information
- Application Number
- CN202510451122.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-15
AI Technical Summary
In streaming calculation, the troubleshooting of abnormal data is inefficient, and the cause of abnormal data cannot be determined, which is time-consuming and labor-intensive, and due to the lack of intermediate processing process data, it is difficult to accurately locate the source of abnormal data.
It provides a real-time calculation and tracking method for streaming data, which can obtain tracking configuration information in the task tracking configuration table, analyze the calculation task identification, and record the full-link calculation related data of the calculation task when the conditions are met, including source data, intermediate calculation results and final result data, so as to accurately locate the abnormal data when it occurs.
It realizes full-link tracking of computing tasks, saves manpower, improves the efficiency of abnormal data detection, can accurately locate the location of problems, and simplifies the logic improvement of computing tasks.
Smart Images

Figure CN120492257A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data real-time stream computing technology, and in particular to a stream data real-time computing tracking method, an abnormal data troubleshooting method and a system thereof. Background Art
[0002] Stream computing is a typical and commonly used computing model in the big data field, primarily used in real-time scenarios with high timeliness requirements, such as real-time recommendations and business monitoring. The data processed by stream computing is called stream data (or data stream). Stream data is a dynamic collection of data that is infinite in time distribution and volume. The value of this data decreases over time, necessitating real-time computing. In stream computing, to improve timeliness, each computing task, from data collection to program processing, is handled within a single program, using a black-box approach. This means that only the source data and the result of the computation are recorded during the computation, without saving any intermediate processing steps or data. While this black-box approach improves timeliness, it also makes troubleshooting abnormal data such as outliers. First, without intermediate data, it's impossible to determine whether abnormal data is caused by a data error or a problem with intermediate program processing. Second, stream computing is characterized by continuous computing tasks and massive data volumes, such as processing billions of data daily. Given these computing tasks and data volumes, troubleshooting the cause of an outlier is like looking for a needle in a haystack. Current methods for troubleshooting abnormal data are typically manual, time-consuming, labor-intensive, and inefficient. Summary of the Invention
[0003] In response to the technical problems existing in the prior art, the present invention proposes a real-time calculation and tracking method for stream data, an abnormal data troubleshooting method and a system thereof, so as to improve the efficiency of monitoring and troubleshooting abnormal data.
[0004] In order to solve the above technical problems, according to one aspect of the present invention, a method for real-time computing and tracking of streaming data is provided, comprising:
[0005] Acquire tracking configuration information in a task tracking configuration table during the real-time computing process of the stream data, wherein the tracking configuration information at least includes an identifier of a computing task to be tracked;
[0006] Parse the tracking configuration information to obtain the computing task identifier that needs to be tracked;
[0007] Monitoring whether the current computing task is a computing task that needs to be tracked in the tracking configuration information; and
[0008] In response to the current computing task being a computing task that needs to be tracked, full-link computing-related data of the computing task is recorded.
[0009] Optionally, the full-link computing-related data of the computing task includes the source data used by the computing task, the intermediate computing result data obtained from each intermediate computing step, and the result data obtained when the computing task is completed.
[0010] Optionally, the tracking configuration information further includes a tracking condition corresponding to the computing task identifier. Correspondingly, when parsing the tracking configuration information, the tracking condition corresponding to the computing task identifier is also obtained;
[0011] When the current computing task is determined to be a computing task that needs to be tracked, the method further includes:
[0012] Determine whether the current computing task meets the tracking conditions in the tracking configuration information; and
[0013] In response to the current computing task satisfying the tracking conditions in the tracking configuration information, full-link computing-related data of the computing task is recorded.
[0014] Optionally, the tracking condition is a start time and a stop time of tracking, and the step of determining whether the current computing task satisfies the tracking condition in the tracking configuration information includes:
[0015] Determine whether the current time is within the time range determined by the tracking start time and stop time; and
[0016] In response to the current time being within a time range determined by the tracking start time and the stop time, it is determined that the current computing task satisfies the tracking condition.
[0017] Optionally, the tracking configuration information further includes status information of a computing task to be tracked, where the status information is activation status information or dormant status information. Correspondingly, when parsing the tracking configuration information, status information corresponding to the computing task identifier to be tracked is also obtained. The step of determining whether the current computing task satisfies the tracking condition in the tracking configuration information includes:
[0018] Determining whether the status information corresponding to the current computing task obtained by parsing the tracking configuration information is in an activated state; and
[0019] In response to the state information corresponding to the current computing task being in an activated state, it is determined that the current computing task satisfies a tracking condition.
[0020] Optionally, the task tracking configuration table is stored in a database. Correspondingly, during the real-time calculation of stream data, the tracking configuration information in the task tracking configuration table is obtained from the database and written into the memory; when parsing the tracking configuration information, the tracking configuration information is read from the memory.
[0021] Optionally, the streaming data real-time calculation and tracking method further comprises, after the calculation task is completed:
[0022] Clears the tracking configuration information of completed computation tasks in memory; and
[0023] Clears or marks the tracking configuration information of completed computing tasks in the task tracking configuration table.
[0024] Optionally, the step of obtaining the tracking configuration information in the task tracking configuration table from the database and writing it into the memory includes:
[0025] Scan the task tracking configuration table at preset time intervals;
[0026] Determining whether there is any untracked tracking configuration information for the computing task; and
[0027] In response to the task tracking configuration table having tracking configuration information of the untracked computing task, the tracking configuration information of the untracked computing task is read and written into the memory.
[0028] Optionally, the step of obtaining the tracking configuration information in the task tracking configuration table from the database and writing it into the memory includes:
[0029] Monitor whether new tracking configuration information for computing tasks has been added to the task tracking configuration table; and
[0030] In response to the tracking configuration information of the computing task being newly added in the task tracking configuration table, the tracking configuration information of the newly added computing task is read and written into the memory.
[0031] Optionally, the streaming data real-time computing and tracking method further includes: creating a task tracking configuration table, and storing the task tracking configuration table in a database.
[0032] Optionally, when recording the full-link calculation-related data of the computing task, the source data used by the computing task, the intermediate calculation result data obtained in each intermediate calculation step, and the result data obtained after the computing task is completed are recorded in different data tables respectively.
[0033] According to another aspect of the present invention, the present invention further provides a method for troubleshooting abnormal data based on the aforementioned streaming data real-time computing and tracking method, wherein the abnormal data is the output data of the computing result of the streaming data computing task, and the method comprises:
[0034] Determining, based on the target abnormal data, a computing task that generates the target abnormal data;
[0035] Obtaining source data and full-link tracking data of the computing task; and
[0036] The source data of the computing task and the full-link tracking data are used as the data investigation scope, and the location where the target abnormal data is generated is located within the data investigation scope;
[0037] The target abnormal data is generated at the source data acquisition stage of the computing task, the processing stage corresponding to the intermediate computing step, or the result data output stage.
[0038] According to another aspect of the present invention, the present invention further provides a streaming data real-time computing and tracking system, comprising:
[0039] A tracking configuration information acquisition module is configured to acquire tracking configuration information in a task tracking configuration table during a real-time computing process of stream data, wherein the tracking configuration information at least includes an identifier of a computing task to be tracked;
[0040] a parsing module configured to parse the tracking configuration information to obtain an identifier of a computing task to be tracked;
[0041] a monitoring module configured to monitor whether the current computing task is a computing task that needs to be tracked in the tracking configuration information; and
[0042] The data recording module is configured to record the full-link computing-related data of the computing task in response to the current computing task being a computing task that needs to be tracked.
[0043] According to another aspect of the present invention, the present invention further provides an abnormal data troubleshooting system, comprising:
[0044] a computing task locating module configured to determine, based on target abnormal data, a computing task that generates the target abnormal data;
[0045] a data acquisition module configured to acquire source data and full-link tracking data of the computing task; and
[0046] The troubleshooting module is configured to use the source data of the computing task and the full-link tracking data as the data troubleshooting scope, and locates the location where the target abnormal data is generated within the data troubleshooting scope;
[0047] The target abnormal data is generated at the source data acquisition stage of the computing task, the processing stage corresponding to the computing step, or the result data output stage.
[0048] According to another aspect of the present invention, the present invention also provides an electronic device, including a processor and a memory, wherein a computer program instruction set is stored on the memory, and when the processor executes the computer program instruction set on the memory, the aforementioned streaming data real-time calculation and tracking method or the aforementioned abnormal data troubleshooting method is implemented.
[0049] According to another aspect of the present invention, the present invention also provides a computer-readable storage medium, on which a computer program instruction set is stored. When the computer program instruction set is executed by a processor, it implements the aforementioned real-time calculation and tracking method for streaming data or the aforementioned abnormal data troubleshooting method.
[0050] According to another aspect of the present invention, the present invention also provides a computer program product, including a computer program instruction set, which, when executed by a processor, implements the aforementioned streaming data real-time calculation and tracking method or the aforementioned abnormal data troubleshooting method.
[0051] The present invention can track the entire link of a computing task, and no longer relies on manual investigation during the investigation of abnormal data, saving manpower and improving investigation efficiency; by tracking the computing process data, the location where the problem occurs can be accurately located, facilitating the improvement of the computing task logic. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Below, the preferred embodiments of the present invention will be further described in detail with reference to the accompanying drawings, in which:
[0053] Figure 1 is a flow chart of a method for real-time computing and tracking of streaming data according to an embodiment of the present invention;
[0054] Figure 2 is a flow chart of a method for real-time computing and tracking of streaming data according to another embodiment of the present invention;
[0055] Figure 3 is a flow chart of a method for troubleshooting abnormal data according to one embodiment of the present invention;
[0056] Figure 4 This is a principle block diagram of a streaming data real-time computing and tracking system according to an embodiment of the present invention;
[0057] Figure 5 This is a principle block diagram of an abnormal data troubleshooting system according to an embodiment of the present invention;
[0058] Figure 6 is a schematic diagram of a recruitment system framework according to an embodiment of the present invention; and
[0059] Figure 7 FIG. 1 is a schematic diagram of the hardware structure principle of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0061] In the detailed description that follows, reference may be made to the various drawings that form part of this application and illustrate specific embodiments of the present application. In the drawings, similar reference numerals describe substantially similar components in different figures. Each specific embodiment of the present application is described below in sufficient detail to enable a person of ordinary skill in the art to implement the technical solutions of the present application. It should be understood that other embodiments may be utilized or that structural, logical, or electrical changes may be made to the embodiments of the present application.
[0062] To facilitate troubleshooting the source and cause of abnormal data in stream computing, the present invention provides a real-time streaming data tracking method. The real-time streaming data computation in the present invention can be performed by a streaming data computing system provided by any stream computing framework, such as STORM, Spark Streaming, Flink, and others. When executing the real-time streaming data computation, the streaming data computing system initiates a parallel process to implement the tracking method provided by the present invention.
[0063] See also Figure 1 , Figure 1 According to a flow chart of a method for real-time computation and tracking of streaming data according to an embodiment of the present invention, in this embodiment, the method includes the following steps:
[0064] Step S11 : obtaining tracking configuration information in a task tracking configuration table during the real-time computing process of the stream data, wherein the tracking configuration information at least includes an identifier of a computing task that needs to be tracked.
[0065] Step S12: Parse the tracking configuration information to obtain the identifier of the computing task that needs to be tracked.
[0066] Step S13: monitor whether the current computing task is a computing task that needs to be tracked in the tracking configuration information. If the current computing task is a computing task that needs to be tracked, execute step S14. If the current computing task is not a computing task that needs to be tracked, return to step S13.
[0067] Step S14: Record the full-link computing-related data of the computing task.
[0068] In order to enable full-link tracking and recording of certain computing tasks during the stream computing process, a task tracking configuration table is first created, in which the tracking configuration information is recorded, and the task tracking configuration table is stored in a database. Therefore, in step S11, the task tracking configuration table and the tracking configuration information therein can be read based on the storage address of the task tracking configuration table in the database. In order to improve the reading and processing speed of the tracking configuration information due to the real-time nature of stream computing, in one embodiment, after the tracking configuration information is obtained from the database, it is stored in memory, thereby facilitating the reading of the tracking configuration information during stream computing.
[0069] The tracking configuration information in the task tracking configuration table includes at least the identifier of the computing task that needs to be tracked, so that the identifier of the computing task that needs to be tracked can be obtained when parsing the tracking configuration information in step S12, so that in step S13, the identifier of the current computing task is compared with the identifier of the computing task that needs to be tracked during the real-time calculation of the stream data to see if they are consistent. If they are consistent, it means that the current computing task is the computing task that needs to be tracked, then in step S14, during the execution of the current computing task, the full-link computing-related data of the computing task is recorded. In order to distinguish the location of the link where the data is located, the present invention divides the full-link tracking data of a computing task into three stages: the starting stage, the intermediate stage and the ending stage, and uses the source data used by the computing task during the calculation as the tracking data of the starting stage, the intermediate calculation result data of each computing step as the tracking data of the intermediate stage, and the result data obtained when the computing task is completed as the tracking data of the ending stage. Among them, when recording the tracking data, there can be multiple ways, for example, the full-link tracking data of all computing tasks are stored in a data table, each record corresponds to a computing task, and the computing task identifier is used as the record identifier, and the tracking data of the aforementioned three stages are stored by partitioning, columnarization, etc. For another example, the tracking data is stored in three data tables according to the aforementioned stages. The data in each data table uses the computing task identifier as the record identifier. For example, tracking table 1 stores the source data used by all computing tasks during calculation, tracking table 2 stores the calculation result data of all computing tasks in the intermediate stages of the calculation process, and tracking table 3 stores the calculation result data obtained by all computing tasks.
[0070] In another embodiment, the tracking configuration information also includes tracking conditions corresponding to the computing task identifier to be tracked, and the tracking conditions corresponding to the computing task identifier to be tracked are also obtained when parsing the tracking configuration information. When the current computing task is determined to be a computing task to be tracked in step S13, it is further determined whether the current computing task satisfies the tracking conditions in the tracking configuration information. When the current computing task satisfies the tracking conditions in the tracking configuration information, step S14 is executed to record the full-link computing-related data of the computing task.
[0071] In one embodiment, the tracking conditions are, for example, the start time and stop time of tracking. At this time, when judging whether the current computing task meets the tracking conditions in the tracking configuration information, it is judged whether the current time is within the time range determined by the start time and stop time of tracking; when the current time is within the time range determined by the start time and stop time of tracking, it is determined that the current computing task meets the tracking conditions.
[0072] In another embodiment, the tracking configuration information further includes status information of the computing task to be tracked, wherein the status information is activation status information or dormant status information; and the tracking condition is that the computing task to be tracked is in an activation state in the task tracking configuration table. In this embodiment, when parsing the tracking configuration information in step S12, status information corresponding to the computing task identifier to be tracked is also obtained. When determining whether the current computing task satisfies the tracking condition in the tracking configuration information, it is determined whether the status information corresponding to the current computing task parsed from the tracking configuration information is in an activation state; if the status information corresponding to the current computing task is in an activation state, it is determined that the current computing task satisfies the tracking condition.
[0073] For example, when creating a task tracking configuration table (tasks_table_track_status), the structure of the tracking configuration information is as follows:
[0074] (task_name,is_track,start_track_date,end_track_date,etl_time).
[0075] Among them, the field "task_name" represents the computing task identifier, and its field value is a string; the field "is_track" represents the status, and the field value is a binary number "0" or "1". In one embodiment, "0" represents dormant status information, that is, no tracking, and "1" represents activated status information, indicating that the corresponding computing task needs to be tracked; the field "start_track_date" represents the start date of tracking; the field "end_track_date" represents the end date of tracking; the field "etl_time" represents the creation time of this configuration information.
[0076] Each time a new computing task needs to be tracked, the following statement is executed on the database (insert intotable tasks_table_track_status values('project_1','1','2025-01-01 00:00:00','2025-01-12 00:00:00','2025-01-10 09:00:00') to insert the configuration information of a specific computing task ('project_1','1','2025-01-01 00:00:00','2025-01-12 00:00:00','2025-01-10 09:00:00') is added to the task tracking configuration table. "project_1" is the value of the "task_name" field, which identifies a specific computing task; "1" is the value of the "is_track" field, indicating that the computing task needs to be tracked; "2025-01-01 00:00:00" and "2025-01-12 00:00:00" are the values of the "start_track_date" and "end_track_date" fields, indicating the start and end dates of tracking, respectively; and "2025-01-10 09:00:00" is the value of the "etl_time" field, indicating the creation time of this configuration information.
[0077] In a further embodiment, when adding the configuration information of a specific computing task to the task tracking configuration table, it is also checked whether the configuration information of the computing task is correct. For example, the current field value is checked field by field to see if it complies with the regulations, such as whether the data in the computing task identification field complies with the naming rules of the computing task identification. For another example, whether the start date field and the end date field comply with the end date field value being greater than the start date field value, or whether the creation time field is less than the end date field value. If all fields comply with the regulations after verification, the configuration information is added to the task tracking configuration table. Otherwise, the addition fails and the operator is prompted with the error.
[0078] The tracking configuration information in this embodiment includes both the start time and stop time of tracking, and the status information of the computing task that needs to be tracked. Therefore, in this embodiment, the priorities of the two tracking conditions are set, among which the tracking priority of the status of the field "is_track" is greater than the priority of the start time and stop time of tracking.
[0079] See also Figure 2 , Figure 2 This is a flow chart of a method for real-time computation and tracking of streaming data according to another embodiment of the present invention. The method comprises the following steps:
[0080] Step S21: Obtain the tracking configuration information in the task tracking configuration table and store it in the memory.
[0081] Step S22: Parse the tracking configuration information. In this embodiment, the identifier of the computing task to be tracked, status information, and tracking start and stop times can be obtained.
[0082] Step S23, monitor whether the current computing task is a computing task that needs to be tracked in the tracking configuration information. If the current computing task is a computing task that needs to be tracked, execute step S24. If the current computing task is not a computing task that needs to be tracked, return to step S23. For example, check whether the current computing task identifier is
[0083] Step S24: determine whether the status information is active status information, for example, check whether the field value of the "is_track" field is "1". If the status information is active status information, execute step S25. If the status information is dormant status information, ignore it and return to step S23.
[0084] Step S25 determines whether the current time is within the time range determined by the start and stop times of tracking. For example, the current time is compared with the data in the fields "start_track_date" and "end_track_date" in the aforementioned embodiment. If the current time is later than the data in the field "start_track_date" and earlier than the data in the field "start_track_date", then the current time is considered to be within the time range determined by the start and stop times of tracking. Otherwise, it is considered to be outside the time range. If the current time is within the time range determined by the start and stop times of tracking, step S26 is executed. If the current time is not within the time range determined by the start and stop times of tracking, the current time is ignored and the process returns to step S23.
[0085] Step S26, record the full-link calculation-related data of the calculation task. Specifically, in the process of the calculation task being executed, the data related to the calculation at each step is stored. For example, in order to execute the calculation task, the stream data calculation system needs to read the source data according to the source data storage location of the calculation task. When the stream data calculation system reads the source data and performs related calculations based on the source data, the source data used in the calculation is saved to the tracking table 1. In the process of performing related calculations based on the source data, the intermediate calculation result data obtained in each calculation step is saved to the tracking table 2. When the result data is obtained after the calculation task is completed, the result data is saved to the tracking table 3.
[0086] In one embodiment, the tracking table is a primary key table, allowing detailed information to be queried based on the primary key, facilitating subsequent abnormal data investigation. The three aforementioned tracking tables 1, 2, and 3 have the same structure. For example, the tracking table structure is as follows: task_name, primary_id, msg_detail, etl_time. The "task_name" field is the computing task identifier, the "primary_id" field is the primary key identifier, the "msg_detail" field is the detailed content of the tracking data, and the "etl_time" field is the record creation time. The primary key identifier and the detailed content of the tracking data are adapted to the computing task.
[0087] For example, in a scenario where we need to count the number of resumes viewed by company HR personnel, we need to calculate the number of resumes viewed daily. This calculation method is to aggregate and count the resume viewing records of HR personnel from the viewing logs, calculate the number of resumes viewed by each HR personnel each day, and then remove dirty data. Dirty data refers to log records with incomplete information.
[0088] This computing task is implemented by subscribing to the stream data computing system. The stream data computing system calculates the final number of events (resume viewing) that need to be counted in real time and outputs it as numerical data to a specified storage location. In one embodiment, the stream data computing system executes the computing task once every certain period of time, such as 10 seconds. Each time the computing task is executed, its computing process is tracked and full-link tracking data is obtained.
[0089] Specifically, in order to calculate the number of resume views during this calculation process, the stream data computing system reads the source data for statistics from the view log and deletes dirty data from it. After the stream data computing system reads the source data for statistics from the view log, it stores it in Tracking Table 1. In one embodiment, the data in Tracking Table 1 is as follows:
[0090] Tracking table 1 (track_info_detail_1):
[0091] ('project_1','001','{id:001,sex:1,event:view_resume,time:2025-03-18
[0092] 10:00:00}','2025-03-18 10:00:00'),
[0093] ('project_1','001','{id:001,sex:1,event:view_resume,time:2025-03-18
[0094] 10:00:01}','2025-03-18 10:00:00'),
[0095] ('project_1','001','{id:001,sex:0,event:view_resume,time:2025}',
[0096] '2025-03-18 10:00:00').
[0097] The structure of each data in Tracking Table 1 is the same. Take the first data as an example:
[0098] 'project_1' is the field value of the "task_name" field, which is a computing task identifier.
[0099] '001' is the field value of the field "primary_id", which is the primary key ID, and in this embodiment, it is the ID of HR.
[0100] '{id:001,sex:1,event:view_resume,time:2025-03-18 10:00:00}' is the field value of the field "msg_detail", which represents detailed tracking data. In this embodiment, it is the resume viewing record information of an HR.
[0101] '2025-03-18 10:00:00' is the value of the "etl_time" field, indicating the time when the data was created.
[0102] It can be seen from the three pieces of data that these three pieces of data were created at the same time.
[0103] This calculation task has one processing step: removing dirty data from the source data. After the calculation task is processed, the third record in the source data is deleted. At this point, the calculation result data for this step is recorded in Tracking Table 2. The data in Tracking Table 2 is as follows:
[0104] Tracking table 2 (track_info_detail_2):
[0105] ('project_1','001','{id:001,sex:1,event:view_resume,time:2025-03-18
[0106] 10:00:00}','2025-03-18 10:00:00'),
[0107] ('project_1','001','{id:001,sex:1,event:view_resume,time:2025-03-18
[0108] 10:00:00}','2025-03-18 10:00:00')
[0109] After the calculation task obtains the result, the calculation result data is recorded in Tracking Table 3. The data in Tracking Table 3 is as follows:
[0110] Tracking table 3 (track_info_detail_3):
[0111] ('project_1','001','{view_resume_count:2}','2025-03-18 10:00:03')
[0112] After completing the computation task, the stream data computing system stores the result data in the designated location.
[0113] For example, in a scenario where we want to count the number of individual job applicants contacted by company HR, we need to calculate the number of applicants contacted by HR via chat each day. This calculation method involves aggregating and summarizing chat logs, calculating the daily chat volume for each HR, and removing dirty data. Dirty data refers to log records with incomplete information.
[0114] The contents of the three tracking tables obtained by the tracking method provided by the present invention are as follows:
[0115] Tracking table 1 (track_info_detail_1):
[0116] ('project_1','001','{id:001,sex:1,event:chat_person,time:2025-03-18
[0117] 10:00:00}','2025-03-18 10:00:00'),
[0118] ('project_1','001','{id:001,
[0119] chat_id:101,sex:1,event:chat_person,time:2025-03-18 10:00:01}','2025-03-1810:00:00'),
[0120] ('project_1','001','{id:001,chat_id:101,sex:0,event:chat_person,time:2025}','2025-03-18 10:00:00')
[0121] Tracking table 2 (track_info_detail_2)
[0122] ('project_1','001','{id:001,
[0123] chat_id:101,sex:1,event:chat_person,time:2025-03-18 10:00:00}','2025-03-1810:00:00'),
[0124] ('project_1','001','{id:001,
[0125] chat_id:101,sex:1,event:chat_person,time:2025-03-18 10:00:00}','2025-03-1810:00:00')
[0126] Tracking table 3 (track_info_detail_3)
[0127] ('project_1','001','{chat_person_count:2}','2025-03-18 10:00:03')
[0128] For another example, in a scenario where the company's HR's most recent active time is counted for individual job seekers to view the HR's activity, the data stored in the designated location is of text time type, and thus the data type of the abnormal data in this embodiment is of text time type.
[0129] The calculation task in this embodiment is to calculate the last active time of HR. The calculation method is as follows: count the active logs, calculate the last active time of each HR, remove dirty data, and then output the calculated time data.
[0130] The obtained tracking table content is as follows:
[0131] Tracking table 1 (track_info_detail_1):
[0132] ('project_1','001','{id:001,sex:1,event:hr_active_event,time:2025-03-1810:00:00}','2025-03-18 10:00:00'),
[0133] ('project_1','001','{id:001,
[0134] chat_id:101,sex:1,event:hr_active_event,time:2025-03-18 11:00:00}','2025-03-1811:00:00'),
[0135] ('project_1','001','{id:001,
[0136] chat_id:101,sex:111,event:hr_active_event,time:2025-03-18 12:00:00}','2025-03-18 12:00:00')
[0137] Tracking table 2 (track_info_detail_2): Filtered out data for sex:111
[0138] ('project_1','001','{id:001,
[0139] chat_id:101,sex:1,event:hr_active_event,time:2025-03-18 11:00:00}','2025-03-1811:00:00')
[0140] Tracking table 3 (track_info_detail_3):
[0141] ('project_1','001','{hr_active_event_time:2025-03-18 11:00:00}','2025-03-18 11:00:00')
[0142] After a computing task is executed, the tracking configuration information for the completed computing task is cleared from memory, thereby freeing up memory space. However, it should be noted that when a tracking time period is set in the tracking configuration information, the tracking configuration information for the completed computing task in memory is cleared when the current time exceeds the tracking time period set in the tracking configuration information. Furthermore, the tracking configuration information for the completed computing task in the task tracking configuration table is cleared or marked to distinguish it from the tracking configuration information for the newly written computing task.
[0143] Since stream computing is constantly ongoing, in order to obtain tracking configuration information in a timely manner, in one embodiment, the task tracking configuration table is scanned at preset time intervals, such as once every 10 minutes by the stream data computing system, and the tracking configuration information of any untracked computing tasks is determined. If the tracking configuration information of any untracked computing tasks is found, the tracking configuration information of the untracked computing tasks is read and written to the memory. In another embodiment, the task tracking configuration table is monitored to determine whether tracking configuration information of any new computing tasks has been added. If the tracking configuration information of any new computing tasks has been added to the task tracking configuration table, the tracking configuration information of the newly added computing tasks is read and written to the memory.
[0144] See also Figure 3 , Figure 3 This is a flow chart of a method for checking abnormal data according to an embodiment of the present invention. In this embodiment, the method for checking abnormal data includes the following steps:
[0145] Step S31: determining a computing task that generates the target abnormal data based on the target abnormal data.
[0146] Step S32: Obtain source data and full-link tracking data of the computing task.
[0147] In step S33, the source data of the computing task and the full-link tracking data are used as the data screening range to locate the location where the target abnormal data is generated.
[0148] Because the stream data computing system outputs computational result data with detailed information about the result data, such as the computation task ID, computation time, corresponding event, etc., when an abnormal data needs to be checked, in step S31, the computation result output data containing the target abnormal data is queried to obtain the corresponding computation task ID.
[0149] In step S32, the source data required to complete the computing task is determined based on the computing task identifier, and the source data is read from the source address. At the same time, the tracking data table of the computing task is obtained from the database. If the tracking data is stored in a data table, the source data used by the corresponding computing task, the intermediate calculation result data obtained in each intermediate calculation step, and the result data obtained after the computing task is completed are determined based on the structure of the tracking data table. If the tracking data is stored in different data tables in stages, the corresponding data table is obtained based on the computing task identifier and the stage identifier, and the corresponding data is read from it. For example, as mentioned above, a specific data table name is used to correspond to the stage, and the data table name is used as the stage identifier, such as Tracking Table 1, Tracking Table 2, and Tracking Table 3 in the above embodiment.
[0150] In step S33, when the full-link tracking data is stored in Tracking Table 1, Tracking Table 2, and Tracking Table 3 respectively, the data in Tracking Table 1 is compared with the source data of the computing task read from the source address. If the two are consistent, it means that there is no problem with the computing task in the source data reading stage. If the two are inconsistent, it means that there is a problem with the computing task in the source data reading stage. After the computing task has no problem in the source data reading stage, the abnormal data is compared with the data in Tracking Table 3. Since the abnormal data is the output data after the computing task is completed, the data in Tracking Table 3 is the result data after the computing task is completed. If the two are consistent, it means that there is no problem in the result data output stage. If they are consistent, it means that there is a problem in the result data output stage. When the abnormal data is consistent with the data in tracking table 3, each calculation step of the calculation task is executed one by one, and a new calculation result is obtained. The new calculation result is compared with the corresponding tracking data in tracking table 2 to determine whether the two are consistent. When all the new calculation results are consistent with the corresponding tracking data in tracking table 2, it means that the abnormal data is not abnormal data, but correct data. If a new calculation result is inconsistent with the corresponding tracking data in tracking table 2, it is determined that the step is the location where the abnormal data is generated.
[0151] In one embodiment, after the troubleshooting is completed, troubleshooting information is generated, in which the computing task identifier, target abnormal data, and whether it is true abnormal data are recorded. If it is true abnormal data, the location where it is generated is also included.
[0152] Based on the troubleshooting information, staff can determine the target abnormal data and whether it is truly abnormal. If it is, they can analyze the cause based on the location of the abnormal data and make appropriate corrections. For example, in the aforementioned scenario of counting the number of resumes viewed by company HR, the original calculation task's processing logic would treat incomplete log records as dirty data, which may result in a discrepancy between the actual number of views. Staff can modify the calculation task's calculation logic, such as using the current year, month, and day as the occurrence date for log records without a specific year, month, and day. The calculation task can then be tracked, troubleshooted, and the results verified again.
[0153] After modifying the calculation logic of the calculation task, for this embodiment, the third source data item is no longer filtered out, resulting in a final result of 3, and the value "3" is output to the designated storage location. For another example, when calculating the company HR's recent active time, the original calculation logic filtered out records with unusual genders. However, in reality, unusual genders should not affect active time. Therefore, after staff modified the filtering algorithm, the calculated result after re-tracking is "2025-03-18 12:00:00", rather than "2025-03-18 11:00:00" in the aforementioned Tracking Table 3.
[0154] If no abnormal data is obtained after the repair, the computing task will no longer be checked. You can also further update the tracking configuration information in the task tracking configuration table to no longer track the computing task. For example, set the field value of the "is_track" field in the tracking configuration information to "0" to put the status of the computing task into a dormant state. For example, execute the SQL statement (update table tasks_table_track_status set is_track='0'where task_name='project_1') in the database to update the tracking configuration information.
[0155] In one embodiment, a task tracking and troubleshooting configuration table can also be created, in which one or more computing tasks that need to be checked are configured, and the triggering conditions for the troubleshooting are set. When the triggering conditions for the troubleshooting are met, the aforementioned Figure 3 Check using the method described above.
[0156] As can be seen from the aforementioned troubleshooting process, since the present invention can track the entire link of a computing task, in the process of troubleshooting abnormal data, firstly, it no longer relies on manual troubleshooting, which saves manpower and improves efficiency. Secondly, by tracking the computing process data, it is possible to accurately locate the location where the problem occurs, which facilitates the improvement of the computing task logic.
[0157] On the other hand, the present invention also provides a real-time computing and tracking system for streaming data. Figure 4 , Figure 4 It is a principle block diagram of a streaming data real-time computing and tracking system according to an embodiment of the present invention. The streaming data real-time computing and tracking system 10 in this embodiment includes a tracking configuration information acquisition module 11, a parsing module 12, a monitoring module 13 and a data recording module 14, wherein the tracking configuration information acquisition module 11 is configured to obtain the tracking configuration information in the task tracking configuration table during the real-time computing process of streaming data, and the tracking configuration information at least includes the computing task identifier that needs to be tracked; the parsing module 12 is configured to parse the tracking configuration information to obtain the computing task identifier that needs to be tracked; the monitoring module 13 is configured to monitor whether the current computing task is the computing task that needs to be tracked in the tracking configuration information; the data recording module 14 is configured to record the full-link computing-related data of the computing task in response to the current computing task being the computing task that needs to be tracked.
[0158] The tracking configuration information acquisition module 11 acquires the tracking configuration information in the task tracking configuration table from the database and writes it into the memory, and the parsing module 12 reads the tracking configuration information from the memory and performs parsing.
[0159] When the tracking configuration information also includes tracking conditions corresponding to the computing task identifier, the parsing module 12 also obtains the tracking conditions corresponding to the computing task identifier when parsing the tracking configuration information; correspondingly, the monitoring module 13 further determines whether the current computing task meets the tracking conditions in the tracking configuration information when determining that the current computing task is a computing task that needs to be tracked; when the current computing task meets the tracking conditions in the tracking configuration information, the data recording module 14 records the full-link computing-related data of the computing task.
[0160] When the tracking condition is the start time and the stop time of tracking, the monitoring module 13 determines whether the current computing task meets the tracking condition in the tracking configuration information, including:
[0161] Determine whether the current time is within the time range determined by the tracking start time and stop time;
[0162] In response to the current time being within a time range determined by the tracking start time and the stop time, it is determined that the current computing task satisfies the tracking condition.
[0163] The tracking configuration information also includes status information of the computing task to be tracked, and the status information is activation status information or dormancy status information. When parsing the tracking configuration information, the parsing module 12 also obtains status information corresponding to the computing task identifier to be tracked. The monitoring module 13 determines whether the current computing task meets the tracking conditions in the tracking configuration information, including the following steps:
[0164] Determining whether the status information corresponding to the current computing task obtained by parsing the tracking configuration information is in an activated state; and
[0165] In response to the state information corresponding to the current computing task being in an activated state, it is determined that the current computing task satisfies a tracking condition.
[0166] In another aspect, the present invention also provides an abnormal data detection system, see Figure 5 , Figure 5 It is a principle block diagram of an abnormal data investigation system according to an embodiment of the present invention. The abnormal data investigation system 20 in this embodiment includes a computing task positioning module 21, a data acquisition module 22 and an investigation module 23, wherein the computing task positioning module 21 is configured to determine the computing task that generates the target abnormal data based on the target abnormal data; the data acquisition module 22 is configured to obtain the source data and full-link tracking data of the computing task; the investigation module 23 is configured to use the source data and full-link tracking data of the computing task as the data investigation range, and locate the generation location of the target abnormal data from the data investigation range; wherein the generation location of the target abnormal data is the source data acquisition stage of the computing task, the processing stage corresponding to the computing step, or the result data output stage.
[0167] See also Figure 6 , Figure 6 The following is a schematic diagram of a recruitment system framework according to one embodiment of the present invention. The recruitment system comprises a recruitment terminal 1, a platform terminal 2, and a job seeker terminal 3. Recruitment terminal 1 and job seeker terminal 3 are user terminals, corresponding to recruiting users and job seekers, respectively. The business terminal 2 is the recruitment platform, comprising a business terminal 21a used by platform staff and a business backend. The business backend comprises multiple servers 22a and a database 23a. Recruitment terminal 1, business terminal 21a, and job seeker terminal 3 are located on personal computers, laptops 3a and 3c, or smart mobile terminals 3b. The stream data computing system on platform terminal 2 performs various real-time computations for recruitment services. The stream data real-time computation tracking system 10 and data troubleshooting system 20 provided by the present invention are each a component of platform terminal 2 and located on one or more servers. The task tracking configuration table, result data, tracking table, and other components described in the present invention are all stored in database 23a. The stream data real-time computation tracking method and abnormal data troubleshooting method provided by the present invention enable the rapid identification of abnormal values within massive amounts of data when they occur in real-time computations, enabling timely correction and ensuring the accuracy of recruitment system data.
[0168] Figure 71 is a schematic diagram of the hardware structure principle of an electronic device according to an embodiment of the present invention. The electronic device can be implemented as a server or various other terminal devices, such as a desktop personal computer, a tablet computer, a laptop computer, a mobile phone, etc., and includes a processor 601 and a memory 602. The memory 602 stores a program instruction set. When the processor 601 executes the program instruction set on the memory 602, the aforementioned real-time calculation and tracking method of streaming data and the data troubleshooting method are implemented.
[0169] Specifically, the processor 601 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiment of the present invention.
[0170] The memory 602 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 602 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 602 may include removable or non-removable (or fixed) media. Where appropriate, the memory 602 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 602 is a non-volatile solid-state memory.
[0171] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Therefore, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the real-time computing and tracking method for streaming data and the data troubleshooting method provided by the present invention.
[0172] In one example, the electronic device may further include a communication interface 603 and a bus 604. The processor 601, the memory 602, and the communication interface 603 are connected via the bus 604 and communicate with each other.
[0173] The communication interface 603 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiment of the present invention.
[0174] Bus 604 includes hardware, software or both, and the components of online data flow metering equipment are coupled to each other. For example, but not limitation, bus can include accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnect (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 604 can include one or more buses. Although the embodiment of the present invention describes and shows a specific bus, the present invention considers any suitable bus or interconnection.
[0175] The present invention also provides a computer-readable storage medium having computer program instructions stored thereon, which can be executed by a processor to implement any one of the software test case generation methods in the aforementioned embodiments. The computer-readable storage medium can be any tangible medium that contains or stores computer-executable instructions for use by or in combination with an instruction execution system, device, and apparatus. The storage medium can be a transient computer-readable storage medium or a non-transient computer-readable storage medium. Non-transient computer-readable storage media may include, but are not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Examples of such storage devices include, for example, magnetic disks, optical disks based on CD, DVD, or Blu-ray technology, and persistent solid-state memories such as flash memory, solid-state drives, and the like.
[0176] The present invention also provides a computer program product comprising a set of computer program instructions that, when executed by a processor, implement any of the methods for real-time computation and tracking of streaming data and data troubleshooting described in the aforementioned embodiments. The computer program product includes, but is not limited to, an application installation package published on a website or in an app store, an application plug-in, or a mini-program that can be run within certain applications.
[0177] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.
[0178] The above embodiments are only used to illustrate the present invention, and are not intended to limit the present invention. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the scope of the present invention. Therefore, all equivalent technical solutions should also fall within the scope of the present invention.
Claims
1. A method for real-time calculation and tracking of streaming data, characterized in that: include: Acquire tracking configuration information in a task tracking configuration table during the real-time computing process of the stream data, wherein the tracking configuration information at least includes an identifier of a computing task to be tracked; Parse the tracking configuration information to obtain the computing task identifier that needs to be tracked; Monitor whether the current computing task is a computing task that needs to be tracked in the tracking configuration information; as well as In response to the current computing task being a computing task that needs to be tracked, full-link computing-related data of the computing task is recorded.
2. The method for real-time calculation and tracking of streaming data according to claim 1, characterized in that: The full-link computing-related data of the computing task includes the source data used by the computing task, the intermediate computing result data obtained from each intermediate computing step, and the result data obtained after the computing task is completed.
3. The method for real-time calculation and tracking of streaming data according to claim 1, characterized in that: The tracking configuration information also includes a tracking condition corresponding to the computing task identifier. Correspondingly, when parsing the tracking configuration information, the tracking condition corresponding to the computing task identifier is also obtained; When the current computing task is determined to be a computing task that needs to be tracked, the method further includes: Determine whether the current computing task meets the tracking conditions in the tracking configuration information; as well as In response to the current computing task satisfying the tracking conditions in the tracking configuration information, full-link computing-related data of the computing task is recorded.
4. The method for real-time calculation and tracking of streaming data according to claim 3, characterized in that: The tracking condition is the start time and stop time of tracking. The step of determining whether the current computing task meets the tracking condition in the tracking configuration information includes: Determine whether the current time is within the time range determined by the tracking start time and stop time; and In response to the current time being within a time range determined by the tracking start time and the stop time, it is determined that the current computing task satisfies the tracking condition.
5. The method for real-time calculation and tracking of streaming data according to claim 3, characterized in that: The tracking configuration information also includes status information of the computing task to be tracked, where the status information is activation status information or dormant status information; Correspondingly, when parsing the tracking configuration information, status information corresponding to the computing task identifier to be tracked is also obtained. The step of determining whether the current computing task meets the tracking conditions in the tracking configuration information includes: Determine whether the status information corresponding to the current computing task obtained from the tracking configuration information is in an activated state; as well as In response to the state information corresponding to the current computing task being in an activated state, it is determined that the current computing task satisfies a tracking condition.
6. The method for real-time calculation and tracking of streaming data according to claim 1, characterized in that: The task tracking configuration table is stored in a database. Correspondingly, during the real-time calculation of stream data, the tracking configuration information in the task tracking configuration table is obtained from the database and written into the memory; when parsing the tracking configuration information, the tracking configuration information is read from the memory.
7. The method for real-time calculation and tracking of streaming data according to claim 6, characterized in that: Further including: After the calculation task is completed, it further includes: Clear the tracking configuration information of completed computing tasks in memory; as well as Clears or marks the tracking configuration information of completed computing tasks in the task tracking configuration table.
8. The method for real-time calculation and tracking of streaming data according to claim 6, characterized in that: The steps to obtain the tracking configuration information in the task tracking configuration table from the database and write it into memory include: Scan the task tracking configuration table at preset time intervals; Determining whether there is any untracked tracking configuration information for the computing task; and In response to the task tracking configuration table having tracking configuration information of the untracked computing task, the tracking configuration information of the untracked computing task is read and written into the memory.
9. The method for real-time calculation and tracking of streaming data according to claim 6, characterized in that: The steps to obtain the tracking configuration information in the task tracking configuration table from the database and write it into memory include: Monitor whether new tracking configuration information for computing tasks has been added to the task tracking configuration table; and In response to the tracking configuration information of the computing task being newly added in the task tracking configuration table, the tracking configuration information of the newly added computing task is read and written into the memory.
10. The method for real-time calculation and tracking of streaming data according to claim 1, characterized in that: Further including: Create a task tracking configuration table and store it in the database.
11. The method for real-time calculation and tracking of streaming data according to claim 2, characterized in that: When recording the full-link calculation-related data of the calculation task, the source data used by the calculation task, the intermediate calculation result data obtained from each intermediate calculation step, and the result data obtained after the calculation task is completed are recorded in different data tables respectively.
12. A method for troubleshooting abnormal data based on any one of claims 1-11, wherein the abnormal data is output data of a calculation result of a stream data calculation task, characterized in that: include: Determining, based on the target abnormal data, a computing task that generates the target abnormal data; Obtain source data and full-link tracking data of the computing task; as well as Use the source data of the computing task and the full-link tracking data as the data investigation scope to locate the location where the target abnormal data is generated; The target abnormal data is generated at the source data acquisition stage of the computing task, the processing stage corresponding to the intermediate computing step, or the result data output stage.
13. A real-time computing and tracking system for streaming data, characterized in that: include: A tracking configuration information acquisition module is configured to acquire tracking configuration information in a task tracking configuration table during a real-time computing process of stream data, wherein the tracking configuration information at least includes an identifier of a computing task to be tracked; a parsing module configured to parse the tracking configuration information to obtain an identifier of a computing task to be tracked; A monitoring module configured to monitor whether a current computing task is a computing task that needs to be tracked in the tracking configuration information; as well as The data recording module is configured to record the full-link computing-related data of the computing task in response to the current computing task being a computing task that needs to be tracked.
14. An abnormal data investigation system, characterized in that: include: a computing task locating module configured to determine, based on target abnormal data, a computing task that generates the target abnormal data; A data acquisition module, configured to acquire source data and full-link tracking data of the computing task; as well as The troubleshooting module is configured to use the source data of the computing task and the full-link tracking data as the data troubleshooting scope, and locates the location where the target abnormal data is generated within the data troubleshooting scope; The target abnormal data is generated at the source data acquisition stage of the computing task, the processing stage corresponding to the intermediate computing step, or the result data output stage.
15. An electronic device comprising a processor and a memory, characterized in that: The memory stores a computer program instruction set, and when the processor executes the computer program instruction set on the memory, it implements the real-time calculation and tracking method for stream data according to any one of claims 1 to 11 or the abnormal data troubleshooting method according to claim 12.
16. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program instruction set, which, when executed by a processor, implements the real-time calculation and tracking method for stream data according to any one of claims 1 to 11 or the abnormal data troubleshooting method according to claim 12.
17. A computer program product, characterized in that The invention comprises a computer program instruction set, which, when executed by a processor, implements the streaming data real-time calculation and tracking method according to any one of claims 1 to 11 or the abnormal data troubleshooting method according to claim 12.