Dynamic label statistical method and system based on real-time data flow
Through the dynamic tag statistics method of real-time monitoring and idempotent control, the problem of insufficient real-time and accuracy in traditional data processing methods is solved, efficient and accurate statistics of real-time data sources are achieved, and the data processing capabilities of enterprises are improved.
Patent Information
- Application Number
- CN202510310825.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-08
AI Technical Summary
Traditional data processing methods are difficult to meet the efficiency and accuracy of modern enterprises for real-time data. Manual or timed statistical methods are susceptible to human factors, resulting in a decline in data quality and being unable to respond to market changes and customer behavior in a timely manner.
Through the initial configuration steps, access multiple data sources, load predefined data source configuration and statistical strategies, monitor data source changes in real time, adopt idempotent control and locking mechanisms to ensure data accuracy and consistency of statistical tasks, and update statistical results in real time.
It realizes automatic monitoring and efficient statistics of real-time data sources, improves the real-time and accuracy of data, reduces the impact of human errors, and improves statistical efficiency and data processing flexibility.
Smart Images

Figure CN120277126A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, specifically to the dynamic label statistics technology of big data, and particularly to a dynamic label statistics method and system based on real-time data streams. Background Art
[0002] In the context of the big data era, enterprises are facing unprecedented data challenges and opportunities. With the rapid development of information technology, data has become the core resource for enterprise decision-making and operation. In this context, enterprises have put forward more stringent requirements for the real-time and accuracy of data processing. Especially for real-time label statistics with high timeliness, its importance has become increasingly prominent.
[0003] Traditional data processing methods, such as manual statistics or regular batch processing, are difficult to meet the urgent needs of modern enterprises for data real-time. These methods are not only inefficient but also unable to respond promptly to market changes and customer behaviors, thus restricting the ability of enterprises to make rapid and accurate decisions in a highly competitive market environment.
[0004] In addition, manual or regular statistical methods are also vulnerable to human factors, resulting in a decrease in data accuracy. Problems such as human errors, operational mistakes, or inconsistencies in data entry can all have a negative impact on data quality, thereby affecting the effectiveness of business analysis and decision-making based on these data.
[0005] Therefore, the present invention proposes a dynamic label statistics method and system based on real-time data streams. Summary of the Invention
[0006] In view of this, the present invention hopes to provide a dynamic label statistics method and system based on real-time data streams to solve or alleviate the technical problems existing in the prior art, that is, how to automatically monitor real-time data sources, efficiently count real-time labels, while ensuring data accuracy, improving statistical efficiency, and providing at least one beneficial option for this; the technical solution of the present invention is implemented as follows:
[0007] In the first aspect, a dynamic label statistics method based on real-time data streams:
[0008] (I) Overview:
[0009] The present invention aims to solve the above technical problems. By means of an initialization configuration step, various data sources are accessed, and the predefined data source configurations, statistical calibers, and policy configurations are loaded to provide a basis for subsequent data processing. During real-time monitoring, changes in the data source are continuously monitored. During this process, through an idempotency control and a locking mechanism, the accuracy of the data and the consistency of the statistical tasks are ensured. In the real-time statistical calculation and update step, the system calculates the changed data according to the preset statistical calibers and policies, and updates the statistical results in real time to ensure that users can obtain the latest data status in a timely manner.
[0010] (2) Technical solution:
[0011] To solve the above technical problems, the present invention adopts the following steps S1 to S4.
[0012] 2.1 Step S1, Initialization configuration:
[0013] Input the data source, and load the predefined data source configuration, statistical caliber, and policy configuration.
[0014] Establish a monitoring connection according to the policy configuration and the data source (including databases, API interfaces, and message queues).
[0015] 2.1.1 Step S100, Parsing:
[0016] Parse the types of accessed data sources, including relational databases, NoSQL databases, API interfaces, or / and message queues. According to the data source type, configure the corresponding connection parameters, including the database URL, username, password, API key, or / and the link address of the message queue.
[0017] 2.1.2 Step S101, Loading the predefined data source configuration:
[0018] Read the predefined data source configuration information from the configuration file.
[0019] 2.1.3 Step S102, Loading the statistical caliber and policy configuration:
[0020] Read the predefined statistical caliber and policy configuration. The policy configuration verifies the effectiveness and consistency of the statistical caliber to ensure that they match the data source and business requirements.
[0021] 2.1.4 Step S103, Establishing a monitoring connection with the data source:
[0022] According to the loaded data source configuration, configure the listener to capture change events in the data source, including adding, modifying, or deleting data.
[0023] 2.2 Step S2, Real-time monitoring:
[0024] Monitor the data source and wait for data change events to occur (including adding, modifying, or deleting data). When the data source changes, the system automatically captures the change event and triggers the corresponding statistical process.
[0025] Check whether the current statistical task has been executed to avoid duplicate calculations. And lock the real-time tag to ensure that only one calculation task is processing the tag at the same time.
[0026] 2.2.1 Step S200, Monitor the data source:
[0027] Continuously monitor the configured data source and wait for data change events to occur. Utilize the monitoring mechanism provided by the data source (including database triggers, API pushes, and / or message queue subscriptions) to capture change events in real time.
[0028] 2.2.2 Step S201, Capture change events:
[0029] When change events such as adding, modifying, or deleting occur in the data source, the system captures these events, parses the content of the change events, and extracts the real-time data that needs to be statistically calculated.
[0030] 2.2.3 Step S202, Trigger the statistical process:
[0031] Based on the captured change event type and content, trigger the corresponding statistical process. Load the predefined policy configuration and prepare for data statistical calculation.
[0032] 2.2.4 Step S203, Check the execution status of the statistical task:
[0033] If it is found that the same statistical task is already being executed or has been completed, avoid duplicate calculations and directly skip the current statistical process.
[0034] 2.2.5 Step S204, Lock the real-time tag:
[0035] Before executing the statistical task, perform a locking operation on the real-time tag. Ensure that only one calculation task is processing the tag at the same time to avoid data conflicts and inconsistencies.
[0036] 2.3 Step S3, Statistical calculation and update:
[0037] After the data change event occurs, extract the real-time data that needs to be statistically calculated from the data source. According to the preset statistical caliber and strategy, calculate the extracted data to generate the statistical result of the real-time tag. Update the calculated real-time tag statistical result to the corresponding real-time metric. After the statistical calculation is completed, release the lock of the real-time tag so that other calculation tasks can process it.
[0038] 2.3.1 Step S300, Extract real-time data:
[0039] After a data change event occurs, use a data query statement or API call to obtain real-time data from the data source. Determine the data range and fields to be extracted based on the content of the data change event.
[0040] 2.3.2 Step S301, perform statistical calculations:
[0041] Calculate the extracted real-time data according to the preset statistical caliber and strategy. Apply statistical functions to generate the statistical results of real-time tags.
[0042] 2.3.3 Step S302, update real-time metrics:
[0043] Update the calculated statistical results of real-time tags to the corresponding real-time metrics. Determine the storage location and update method of the real-time metrics (including database update or / and cache refresh). Execute the update operation to ensure that the real-time metrics can reflect the latest statistical results.
[0044] 2.3.4 Step S303, release the lock of the real-time tag:
[0045] After the statistical calculation is completed, release the lock of the real-time tag. Clean up relevant resources to prepare for the next statistical calculation.
[0046] 2.4 Step S4, standby:
[0047] After the current statistical task is executed, wait for the trigger of the next data source change event.
[0048] (III) Mechanism for solving technical problems:
[0049] In the initialization configuration step (S1), the technical solution of the present invention first inputs the data source and loads the predefined data source configuration, statistical caliber, and strategy configuration. This step provides clear directions and rules for subsequent data processing and statistics. Then, according to the strategy configuration, the technical solution of the present invention establishes a listening connection with the data source (including databases, API interfaces, message queues, etc.) to ensure that changes in the data source can be captured in real time.
[0050] After entering the real-time listening step (S2), the technical solution of the present invention continuously listens to the data source and waits for a data change event to occur. Once a change such as addition, modification, or deletion occurs in the data source, the system will automatically capture these change events and immediately trigger the corresponding statistical process. During this process, the technical solution of the present invention also checks whether the current statistical task has been executed through an idempotency control mechanism to avoid resource waste and data errors caused by repeated calculations. At the same time, lock the real-time tag to ensure that only one calculation task is processing the tag at the same time, thus ensuring data consistency and accuracy.
[0051] In the real-time statistical calculation and update step (S3), after a data change event occurs, the technical solution of the present invention extracts the real-time data to be statistically calculated from the data source. Then, according to the preset statistical caliber and strategy, the extracted data is calculated to generate the statistical results of the real-time tags. These results will be updated to the corresponding real-time metrics in a timely manner so that users can view the latest statistical status at any time. After the statistical calculation is completed, the technical solution of the present invention releases the lock of the real-time tag so that other calculation tasks can process other tags or subsequent changes of the same tag.
[0052] Finally, in the standby step (S4), the technical solution of the present invention waits for the trigger of the next data source change event. During this period, the system will maintain the listening state of the data source and be ready to respond to new change events at any time and start a new statistical process.
[0053] In summary, the technical solution of the present invention realizes the automatic listening of the real-time data source and the efficient statistics of the real-time tags through an automated and intelligent data processing process. At the same time, through means such as idempotent control and locking mechanism, the accuracy of the data and the consistency of the statistical tasks are ensured. These mechanisms together improve the statistical efficiency and meet the high requirements of enterprises for data real-time and accuracy.
[0054] In the second aspect, a dynamic tag statistical system based on real-time data stream:
[0055] As Figure 5 shown, the system includes a processor and a connection to the processor:
[0056] (1) An initialization configuration module responsible for parsing and identifying the types of accessed data sources:
[0057] The core function of this module is to perform preliminary preparation and configuration work to ensure that the data processing process can be started smoothly, and configure the corresponding connection parameters according to the characteristics of the data source, such as the URL, username, password of the database, or API key, etc. Read the predefined data source configuration information and statistical caliber and strategy configuration from the configuration file to ensure that the data processing work can be carried out according to the predetermined rules.
[0058] (2) A real-time listening module responsible for continuously listening to the configured data source and capturing data change events in real time:
[0059] When changes such as new additions, modifications, or deletions occur in the data source, this module immediately captures these events, parses the event content, and extracts the real-time data that needs to be counted. At the same time, it also triggers the corresponding statistical process based on the captured change events and checks whether the current statistical task has been executed to avoid duplicate calculations. To ensure that only one calculation task is processing a certain label at the same moment, this module also performs a locking operation on the real-time label.
[0060] (3) A real-time statistical calculation and update module for performing real-time statistical calculations and result updates:
[0061] After a data change event occurs, this module extracts the real-time data that needs to be counted from the data source, calculates according to the preset statistical caliber and strategy, and generates the statistical results of the real-time label. Subsequently, it updates these statistical results to the corresponding real-time metrics to ensure that the real-time metrics can reflect the latest statistical results. Finally, to prepare for the next statistical calculation, this module also releases the lock on the real-time label and cleans up the relevant resources.
[0062] (4) A memory connected to the processor, where program instructions are stored. When the program instructions are executed by the processor, the processor executes the dynamic label statistical method according to any one of claims 1-8.
[0063] Compared with the prior art, the beneficial effects of the present invention are:
[0064] I. Improved real-time performance: By automatically monitoring the data source, the present invention can capture data change events in real time and immediately trigger the statistical process, ensuring that users can obtain the latest data statistical results at any time, meeting the high requirements for data real-time performance. An idempotency control and locking mechanism are adopted in the statistical calculation process, effectively avoiding the problems of duplicate calculations and data conflicts, and ensuring the accuracy of data statistics.
[0065] II. Improved statistical efficiency: Through the pre-defined data source configuration, statistical caliber, and strategy configuration, the present invention can efficiently extract, calculate, and update data, greatly improving the statistical efficiency and reducing the data processing cost.
[0066] III. Enhanced flexibility: The real-time updated data statistical results of the present invention can provide users with a more intuitive and accurate data view, helping users better understand the business situation and make more informed decisions. Description of the Drawings
[0067] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0068] Figure 1 It is a schematic flowchart of the method of the present invention;
[0069] Figure 2 It is a schematic flowchart of the method of step S1 of the present invention;
[0070] Figure 3 It is a schematic flowchart of the method of step S2 of the present invention;
[0071] Figure 4 It is a schematic flowchart of the method of step S3 of the present invention;
[0072] Figure 5 It is a schematic diagram of the system composition of the present invention. Specific Embodiments
[0073] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention with reference to the drawings. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein. Those skilled in the art can make similar improvements without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below;
[0074] It should be noted that the various embodiments in this specification are described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0075] Embodiment 1: As Figures 1 to 4 shown, the dynamic label statistics method based on real-time data streams disclosed in this embodiment can be widely applied to various scenarios that require real-time data, such as the driver's home page, the display of real-time indicators such as driver turnover, completed order volume, and driver income. These indicators all put higher requirements on real-time performance and accuracy. Therefore, this embodiment will take the driver turnover as the target and further illustrate the specific implementation of this solution.
[0076] (1) Initialization configuration (step S1):
[0077] Specifically, analyze the data source type (step S100): For driver transaction data, identify its source, including relational databases, NoSQL databases, API interfaces, or message queues. According to the data source type, configure connection parameters such as database URLs, usernames, passwords, etc., to ensure stable connection and data extraction.
[0078] Specifically, load the data source configuration (step S101): Read the specific configuration information of the driver transaction data source from the configuration file, including the data source location, access permissions, etc.
[0079] Specifically, load the statistical caliber and policy configuration (step S102): Read the predefined driver transaction statistical calibers and policies, such as transaction calculation rules, time period divisions, etc. Verify the validity and consistency of the statistical calibers to ensure they match the data source and business requirements.
[0080] Specifically, establish a listening connection (step S103): According to the loaded data source configuration, set up a listener to capture change events in the data source in real time, such as order status updates, payment completions, etc.
[0081] (2) Real-time monitoring (step S2)
[0082] Specifically, it includes the following sub-steps:
[0083] Specifically, monitor the data source (step S200): Continuously monitor the configured driver transaction data source and wait for data change events to occur. Utilize the listening mechanism provided by the data source, such as database triggers, API pushes, etc., to capture change events in real time.
[0084] Specifically, capture change events (step S201): When the data source changes, such as a new order is generated, the order status is updated, etc., the system immediately captures these events. Analyze the content of the change events and extract the real-time data that needs to be statistically calculated, such as order amount, driver ID, etc.
[0085] Specifically, trigger the statistical process (step S202): According to the type and content of the captured change events, trigger the corresponding driver transaction statistical process. Load the predefined policy configuration and prepare for data statistical calculation.
[0086] Specifically, check the execution status of the statistical task (step S203): Before executing the statistical task, check whether there is already the same statistical task in execution or completed. If a duplicate is found, avoid duplicate calculations and directly skip the current statistical process.
[0087] Specifically, lock the real-time tag (step S204): Before executing the statistical task, perform a locking operation on the real-time tag (such as the total driver turnover). Ensure that only one calculation task is processing the tag at the same time to avoid data conflicts and inconsistencies.
[0088] (III) Real-time statistical calculation and update (step S3)
[0089] Specifically, it includes the following sub-steps:
[0090] Specifically, extract real-time data (step S300): After a data change event occurs, use a data query statement or API call to obtain real-time data from the data source. Determine the data range and fields to be extracted according to the content of the data change event, such as the order time range, driver ID, etc.
[0091] Specifically, perform statistical calculations (step S301): Calculate the extracted real-time data according to the preset statistical caliber and strategy. Apply statistical functions to generate statistical results of real-time tags, such as the total driver turnover, the number of orders, etc.
[0092] Specifically, update real-time metrics (step S302): Update the statistical results of the real-time tags calculated to the corresponding real-time metrics. Determine the storage location and update method of the real-time metrics, such as database update or cache refresh. Perform the update operation to ensure that the real-time metrics can reflect the latest statistical results.
[0093] Specifically, release the lock of the real-time tag (step S303): After the statistical calculation is completed, release the lock of the real-time tag. Clean up relevant resources to prepare for the next statistical calculation.
[0094] (IV) Standby (step S4)
[0095] After the current statistical task is executed, the system enters the standby state, waiting for the trigger of the next data source change event. In the standby state, the system will maintain a listening connection to respond to new data change events at any time and trigger the corresponding statistical process. By continuously looping through the above steps, the system can statistically calculate and update real-time metrics such as driver turnover in real time and accurately, providing strong support for business decision-making.
[0096] (V) Application program:
[0097] The application program (python) of the above dynamic tag statistical method based on real-time data stream in the driver turnover scenario is as follows:
[0098]
[0099]
[0100]
[0101] In the above program:
[0102] (5.1) Initialize configuration (step S1):
[0103] In the initialize method, the program first parses the data source type and configures the connection parameters according to the data source type. Load the data source configuration and the statistical caliber and policy configuration. Establish a listening connection to capture change events in the data source in real time.
[0104] (5.2) Listen in real time (step S2):
[0105] In the listen_for_data_changes method, the program continuously listens to the configured driver transaction data source and waits for data change events to occur. When the data source changes, capture these events and trigger the corresponding driver transaction statistics process.
[0106] (5.3) Real-time statistical calculation and update (step S3):
[0107] In the trigger_statistics_workflow method, the program first checks whether there is already the same statistical task being executed or completed to avoid duplicate calculations.
[0108] Lock the real-time tag to ensure that only one calculation task is processing the tag at the same time.
[0109] Extract the real-time data and calculate according to the preset statistical caliber and policy.
[0110] Update the calculated real-time tag statistical results to the corresponding real-time metrics.
[0111] After the statistical calculation is completed, release the lock of the real-time tag and clean up the relevant resources.
[0112] (5.4) Standby (step S4):
[0113] After the current statistical task is completed, the system enters the standby state, waiting for the next data source change event to be triggered. The system will maintain the listening connection so as to respond to new data change events at any time and trigger the corresponding statistical process.
[0114] Embodiment 2: As Figures 1 to 4As shown, this embodiment discloses a dynamic label statistics method based on real-time data streams, which can be widely applied to various scenarios that require real-time data, such as the display of real-time metrics on the driver side, such as driver turnover, completed orders, driver income, etc. These metrics pose higher requirements for real-time performance and accuracy. Therefore, this embodiment will use the driver's completed order volume as the target to further illustrate the specific implementation of this solution.
[0115] (1) Step S1: Initialization configuration:
[0116] Specifically, parse the data source type (step S100): Identify whether the data source is a relational database, a NoSQL database, an API interface, or a message queue. According to the data source type, configure the connection parameters, such as the database URL, username, password, API key, or the link address of the message queue.
[0117] Specifically, load the predefined data source configuration (step S101): Read the specific configuration information of the data source from the configuration file.
[0118] Specifically, load the statistical caliber and policy configuration (step S102): Read the predefined statistical caliber and policy configuration. Verify the effectiveness and consistency of the statistical caliber to ensure they match the data source and business requirements.
[0119] Specifically, establish a listening connection with the data source (step S103): According to the loaded data source configuration, configure the listener to capture change events in the data source.
[0120] (2) Step S2: Real-time listening:
[0121] Specifically, listen to the data source (step S200): Continuously listen to the configured data source and wait for data change events to occur. Utilize the listening mechanism provided by the data source (such as database triggers, API pushes, or message queue subscriptions) to capture change events in real time.
[0122] Specifically, capture change events (step S201): When the data source changes, the system captures these events and parses the content of the change events. Extract the real-time data that needs to be statistically calculated.
[0123] Specifically, trigger the statistical process (step S202): According to the type and content of the captured change events, trigger the corresponding statistical process. Load the predefined policy configuration and prepare for data statistical calculation.
[0124] Specifically, check the execution status of the statistical task (step S203): Check whether there is already the same statistical task being executed or completed to avoid duplicate calculations.
[0125] Specifically, lock the real-time tag (step S204): Before executing the statistical task, perform a locking operation on the real-time tag to ensure that only one calculation task is processing the tag at the same time.
[0126] (III) Step S3: Real-time statistical calculation and update:
[0127] Specifically, extract real-time data (step S300): Use a data query statement or API call to obtain real-time data from the data source. Determine the data range and fields to be extracted according to the content of the data change event.
[0128] Specifically, perform statistical calculation (step S301): Calculate the extracted real-time data according to the preset statistical caliber and strategy. Apply statistical functions to generate the statistical results of the real-time tag.
[0129] Specifically, update real-time metrics (step S302): Update the statistical results of the real-time tag calculated to the corresponding real-time metrics. Determine the storage location and update method of the real-time metrics, and perform the update operation.
[0130] Specifically, release the lock of the real-time tag (step S303): After the statistical calculation is completed, release the lock of the real-time tag and clean up the relevant resources.
[0131] (IV) Step S4: Standby:
[0132] After the current statistical task is executed, the system enters the standby state. Wait for the trigger of the next data source change event so as to respond and trigger the corresponding statistical process at any time.
[0133] By continuously looping and executing the above steps, the system can statistically calculate and update real-time metrics such as the number of completed orders of drivers in real time and accurately, providing strong support for business decisions.
[0134] (V) Application program:
[0135] The application program (python) of the above dynamic tag statistical method based on real-time data stream in the scenario of the number of completed orders of drivers is as follows, where pandas is used for data processing, sqlalchemy is used for database operations, and redis is used for distributed locks:
[0136]
[0137]
[0138]
[0139]
[0140] In the above program:
[0141] (5.1) Initialization Configuration:
[0142] Parse the data source type and configure the connection parameters according to the type. Load the configurations of the data source, statistical caliber, and strategy from the configuration file (or hard-coded). Establish a listening connection with the data source to prepare for capturing change events.
[0143] (5.2) Real-time Listening:
[0144] Listen to the data source and wait for data change events to occur. When the data source changes, capture these events and parse their contents.
[0145] Trigger the corresponding statistical process and check whether the same statistical task is already in execution or has been completed to avoid duplicate calculations.
[0146] Before executing the statistical task, lock the real-time tag to ensure that only one calculation task is processing the tag at the same time.
[0147] (5.3) Real-time Statistical Calculation and Update:
[0148] Extract the real-time data to be statistically calculated from the data source. Calculate the extracted data according to the preset statistical caliber and strategy.
[0149] Update the statistical results of the real-time tag calculated to the corresponding real-time metrics. After the statistical calculation is completed, release the lock of the real-time tag and clean up the relevant resources.
[0150] (5.4) Standby:
[0151] After the current statistical task is completed, the system enters the standby state. Wait for the next data source change event to be triggered so as to respond and trigger the corresponding statistical process at any time.
[0152] Embodiment 3: As Figure 5 shown, this embodiment discloses a dynamic tag statistical system based on real-time data streams: The system includes a processor, and connected to the processor are:
[0153] (1) An initialization configuration module responsible for parsing and identifying the accessed data source type:
[0154] The core function of this module is to perform preliminary preparation and configuration work to ensure that the data processing process can be started smoothly, and configure the corresponding connection parameters according to the characteristics of the data source, such as the URL, username, password of the database, or API key, etc. Read the predefined data source configuration information, statistical caliber, and strategy configuration from the configuration file to ensure that the data processing work can be carried out according to the predetermined rules.
[0155] Its sub - modules include:
[0156] (1.1) Parsing sub - module (S100):
[0157] Function: Responsible for parsing the accessed data source type and configuring corresponding connection parameters according to the data source type.
[0158] Input: Data source type information.
[0159] Output: Configured data source connection parameters.
[0160] (1.2) Loading data source configuration sub - module (S101):
[0161] Function: Read the predefined data source configuration information from the configuration file.
[0162] Input: Configuration file path.
[0163] Output: Data source configuration information.
[0164] (1.3) Loading statistical caliber and policy configuration sub - module (S102):
[0165] Function: Read the predefined statistical caliber and policy configuration and perform validity verification.
[0166] Input: Statistical caliber and policy configuration file path.
[0167] Output: Verified statistical caliber and policy configuration.
[0168] (1.4) Establishing listening connection sub - module (S103):
[0169] Function: Establish a listening connection with the data source according to the loaded data source configuration.
[0170] Input: Data source configuration information.
[0171] Output: Listening connection with the data source.
[0172] (2) Real - time listening module responsible for continuously listening to the configured data sources and capturing data change events in real - time:
[0173] When changes such as addition, modification, or deletion occur to the data source, this module will immediately capture these events, parse the event content, and extract the real - time data that needs to be statistically analyzed. At the same time, it will trigger the corresponding statistical process according to the captured change events and check whether the current statistical task has been executed to avoid duplicate calculations. To ensure that only one calculation task is processing a certain tag at the same moment, this module will also perform a locking operation on the real - time tag. Its sub - modules include:
[0174] (2.1) Monitoring Data Source Sub-module (S200):
[0175] Function: Continuously monitor the configured data source and wait for the occurrence of data change events.
[0176] Input: Monitoring connection with the data source.
[0177] Output: Captured data change events.
[0178] (2.2) Capturing Change Events Sub-module (S201):
[0179] Function: Capture change events in the data source and parse the event content.
[0180] Input: Monitored data change events.
[0181] Output: Parsed real-time data.
[0182] (2.3) Triggering Statistical Process Sub-module (S202):
[0183] Function: Trigger corresponding statistical processes based on the captured change events.
[0184] Input: Parsed real-time data.
[0185] Output: Triggered statistical process information.
[0186] (2.4) Checking the Execution Status of Statistical Tasks Sub-module (S203):
[0187] Function: Check whether the current statistical task has been executed to avoid repeated calculations.
[0188] Input: Triggered statistical process information.
[0189] Output: Check result (whether it has been executed).
[0190] (2.5) Real-time Tag Locking Sub-module (S204):
[0191] Function: Lock the real-time tags before executing the statistical tasks.
[0192] Input: Real-time tag information to be locked.
[0193] Output: Locking status.
[0194] (3) Real-time Statistical Calculation and Update Module for Performing Real-time Statistical Calculations and Result Updates:
[0195] After a data change event occurs, this module extracts the real-time data to be counted from the data source, calculates it according to the preset statistical caliber and strategy, and generates the statistical results of real-time tags. Subsequently, it updates these statistical results to the corresponding real-time metrics to ensure that the real-time metrics can reflect the latest statistical results. Finally, to prepare for the next statistical calculation, this module also releases the lock of the real-time tag and clears the relevant resources. Its sub-modules include:
[0196] (3.1) Real-time data extraction sub-module (S300):
[0197] Function: Extract the real-time data to be counted from the data source.
[0198] Input: The content of the data change event.
[0199] Output: The extracted real-time data.
[0200] (3.1) Statistical calculation execution sub-module (S301):
[0201] Function: Calculate the extracted data according to the preset statistical caliber and strategy.
[0202] Input: The extracted real-time data.
[0203] Output: The calculated statistical results of real-time tags.
[0204] (3.3) Real-time metric update sub-module (S302):
[0205] Function: Update the calculated statistical results of real-time tags to the corresponding real-time metrics.
[0206] Input: The calculated statistical results of real-time tags.
[0207] Output: The updated real-time metrics.
[0208] (3.4) Real-time tag lock release sub-module (S303):
[0209] Function: Release the lock of the real-time tag after the statistical calculation is completed.
[0210] Input: The real-time tag information for which the lock needs to be released.
[0211] Output: The status of the released lock.
[0212] (4) A memory connected to the processor, wherein program instructions are stored in the memory, and when the program instructions are executed by the processor, the processor executes the dynamic tag statistical method according to any one of claims 1-8.
[0213] All of the above embodiments merely represent the implementation manners of the relevant actual applications of the present invention. The descriptions thereof are relatively specific and detailed, but should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent shall be subject to the appended claims.
[0214] For those skilled in the art, it can be further realized that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0215] Meanwhile, those skilled in the art can understand that all or part of the processes of implementing the methods of all the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
Claims
1. A dynamic label statistics method based on real-time data streams, characterized in that, After the data source generates new data, the following steps are executed: S1. Initialize the configuration: Input the data source, load the predefined data source configuration, statistical caliber, and policy configuration; establish a listening connection with the data source according to the policy configuration; S2. Monitor in real time: Capture the change event when the data source changes, and trigger the corresponding statistical process; S3. Statistical calculation and update: Extract the real-time data to be statistically calculated from the data source, calculate the extracted data according to the preset statistical caliber and policy, generate the statistical result of the real-time label; update the calculated real-time label statistical result to the corresponding real-time metric; After the statistical calculation is completed, release the lock of the real-time label; S4. Standby: After the current statistical task is executed, wait for the trigger of the next data source change event.
2. The dynamic label statistics method according to claim 1, wherein: The implementation method of S1 is as follows: S100. Analyze the type of the accessed data source, including relational database, NoSQL database, API interface, or / and message queue; configure the corresponding connection parameters according to the data source type, including database URL, username, password, API key, or / and the link address of the message queue; S101. Read the predefined data source configuration information from the configuration file; S102. Read the predefined statistical caliber and policy configuration; S103. Configure the listener to capture the change event in the data source according to the loaded data source configuration.
3. The dynamic label statistical method according to claim 1, wherein: The implementation method of S2 is as follows: S200. Continuously monitor the configured data source and wait for the occurrence of the data change event; use the listening mechanism provided by the data source to capture the change event in real time; S201. Capture the event and parse the content of the change event to extract the real-time data to be statistically calculated; S202. Load the predefined policy configuration.
4. The dynamic label statistics method according to claim 3, wherein: The listening mechanism includes database trigger, API push, or / and message queue subscription.
5. The dynamic label statistics method according to claim 3, wherein: In S2, it also includes: S203. If it is found that the same statistical task is being executed or has been completed, directly skip the current statistical process; S204. Perform a locking operation on the real-time label.
6. The dynamic label statistical method according to any one of claims 1 to 3, characterized in that: The implementation method of S3 is as follows: S300. When the data change event occurs, obtain the real-time data from the data source by using a data query statement or API call; S301. Calculate the extracted real-time data according to the preset statistical caliber and policy; apply statistical functions to generate the statistical result of the real-time label; S302. Update the calculated real-time label statistical result to the corresponding real-time metric S303. Release the lock of the real-time label.
7. The dynamic label statistics method according to claim 6, wherein: In S300, determine the data range and fields to be extracted according to the content of the data change event.
8. The dynamic label statistics method according to claim 6, wherein: In S302, determine the storage location and update method of the real-time metric, including database update or / and cache refresh; then perform the update operation.
9. A dynamic label statistics system based on real-time data streams, characterized in that: The system includes a processor and a memory connected to the processor. Program instructions are stored in the memory. When the program instructions are executed by the processor, the processor executes the dynamic label statistical method described in any one of claims 1-8.
10. The dynamic label statistics system according to claim 9, wherein: The processor is connected to: An initialization configuration module responsible for parsing and identifying the type of the accessed data source; A real-time monitoring module responsible for continuously monitoring the configured data sources and capturing data change events in real time; A real-time statistical calculation and update module for performing real-time statistical calculations and result updates.