Real-time burying point data distribution method and device based on Flink and medium

By using a single Flink program in Kafka to associate with the tracking write configuration table, the issues of high resource consumption and storage costs in a multi-Flink program architecture are resolved, efficient tracking data distribution and query optimization are achieved, and real-time performance is guaranteed in high-concurrency scenarios.

CN120763233APending Publication Date: 2025-10-10QIAN JIN NETWORK INFORMATION TECH SHANGHAI LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510843132.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

In the existing technology, the distribution of tracking data under the multi-Flink program architecture has problems such as high resource consumption, surge in network I/O pressure, high storage costs and low query efficiency. It is particularly difficult to ensure real-time performance in high-concurrency scenarios.

Method used

A single Flink program is used in conjunction with the tracking data writing configuration table. Tracking data is filtered and distributed from Kafka through associations to achieve dynamic routing distribution, reduce resource consumption, optimize storage structure, and avoid duplicate consumption and full table scans.

Benefits of technology

It effectively reduces Flink resource consumption, reduces Kafka network I/O pressure, reduces storage costs, improves query efficiency, and ensures real-time performance in high-concurrency scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763233A_ABST
    Figure CN120763233A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time burying point data distribution method and device based on Flink, a medium and a product. Comprising the following steps: establishing an association relationship between each burying point event and a burying point data table based on a service line and / or a service scene corresponding to each burying point event and the burying point data table associated with each service line and each service scene, and establishing a burying point write-in configuration table based on the association relationship and storing the burying point write-in configuration table in a real-time data bin; calling a single Flink program to read a burying point write-in configuration table from the real-time data bin, and consuming real-time burying point data in each Topic of the Kafka according to the burying point write-in configuration table; in the process of consuming the real-time buried point data by using a single Flink program, screening buried points from the real-time buried point data, and writing the buried points into the real-time buried point data corresponding to the event name contained in the configuration table to obtain batch buried point data; and distributing and writing the batch burying point data into the corresponding burying point data table according to the incidence relation between each burying point event and the burying point data table.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a Flink-based real-time tracking data distribution method, device, medium, and product. Background Art

[0002] With the rapid development of big data applications, real-time analysis of user behavior tracking data has become a core requirement for precise enterprise operations. Tracking data is often collected in high-throughput, real-time fashion using message queues (such as Kafka). Apache Flink, with its low-latency, highly reliable stream processing capabilities, has become a leading technology solution for consuming Kafka data in real time and generating analytical metrics. By using Flink programs to clean, distribute, and store tracking data, businesses can quickly respond to real-time monitoring of user behavior logs, provide anomaly alerts, and perform funnel analysis, thus supporting data-driven, immediate decision-making.

[0003] In related technologies, the distribution of tracking data in multi-business line scenarios is mainly achieved through two methods:

[0004] 1) Multiple Flink programs for independent consumption: Create a separate Flink program for each business requirement. Each program filters target events from the full Kafka tracking data and writes them to a separate data table. For example, activity conversion analysis and user path tracking are performed by two separate Flink programs, both of which independently subscribe to the complete Kafka data stream.

[0005] 2) Full storage followed by downstream distribution: A single Flink program stores all Kafka tracking data in a unified wide table, allowing downstream businesses to filter required data through queries. For example, all tracking events are stored in a single table, and downstream tasks perform a full table scan to filter out related tracking events.

[0006] However, the above solution has the following drawbacks: In a multi-Flink program architecture, repeated consumption of similar data causes Kafka bandwidth resources to be duplicated by multiple consumers, increasing network I / O pressure and even causing transmission delays. Furthermore, running multiple Flink jobs in parallel consumes a large amount of computing resources, and the complexity of program deployment and monitoring increases dramatically as business demand grows. While the full storage solution simplifies upstream links, the expansion of detailed data tables significantly increases storage costs. Repeated full table scans and filtering downstream also lead to inefficient queries, especially in highly concurrent analysis scenarios. Computing resources are further wasted, making real-time performance difficult to guarantee. Therefore, balancing resource consumption and data distribution efficiency has become a pressing technical challenge. Summary of the Invention

[0007] In view of this, the embodiments of the present application provide a real-time tracking data distribution method, device, equipment, medium and product based on Flink, which can reduce Flink resource consumption while ensuring the transmission efficiency of tracking data.

[0008] On the first aspect, an embodiment of the present application provides a real-time point-of-sale data distribution method based on Flink, which includes: in the real-time point-of-sale configuration interface of the big data platform, configuring the point-of-sale events corresponding to each business line and the point-of-sale events corresponding to each business scenario under each business line based on business needs; establishing an association relationship between each point-of-sale event and the point-of-sale data table based on the business line and / or business scenario corresponding to each point-of-sale event, and creating a point-of-sale write configuration table based on the association relationship and storing it in the real-time data warehouse, wherein the field of the point-of-sale write configuration table includes the event name of the point-of-sale event and the table name of the point-of-sale data table, and each point-of-sale data table is stored in the detailed data DWD layer of the real-time data warehouse; when allocating a single point-of-sale event to the consumer message queue system Kafka, In the case of the stream processing engine Flink program, a single Flink program is called to read the tracking write configuration table from the real-time data warehouse, and a single Flink program is used to consume the real-time tracking data in each data topic of Kafka based on the tracking write configuration table; in the process of using a single Flink program to consume real-time tracking data, the real-time tracking data corresponding to the event name in the tracking write configuration table is filtered from the real-time tracking data to obtain the batch tracking data corresponding to the business needs; based on the association between each tracking event and the tracking data table in the tracking write configuration table, the batch tracking data is distributed and written to the corresponding tracking data table, so that each downstream data processing node can read different tracking data tables based on its own real-time tasks, and obtain the target tracking data required for real-time calculation.

[0009] On the second aspect, an embodiment of the present application provides a real-time point-of-sale data distribution device based on Flink, which includes: a configuration module for configuring the point-of-sale events corresponding to each business line and the point-of-sale events corresponding to each business scenario under each business line based on business needs in the real-time point-of-sale configuration interface of the big data platform; an association module for establishing an association relationship between each point-of-sale event and the point-of-sale data table based on the business line and / or business scenario corresponding to each point-of-sale event, and the point-of-sale data table associated with each business line and business scenario, and creating a point-of-sale write configuration table based on the association relationship and storing it in the real-time data warehouse, wherein the fields of the point-of-sale write configuration table include the event name of the point-of-sale event and the table name of the point-of-sale data table, and each point-of-sale data table is stored in the detailed data DWD layer of the real-time data warehouse; a consumption module for distributing the point-of-sale events in the Kafka distribution system for consuming message queues. When equipped with a single stream processing engine Flink program, a single Flink program is called to read the point-of-sale write configuration table from the real-time data warehouse, and a single Flink program is used to consume the real-time point-of-sale data in each data topic of Kafka based on the point-of-sale write configuration table; the filtering module is used to filter the real-time point-of-sale data corresponding to the event name in the point-of-sale write configuration table from the real-time point-of-sale data in the process of using a single Flink program to consume real-time point-of-sale data, and obtain the batch point-of-sale data corresponding to the business needs; the writing module is used to distribute and write the batch point-of-sale data to the corresponding point-of-sale data table based on the association between each point-of-sale event in the point-of-sale write configuration table, so that each downstream data processing node can read different point-of-sale data tables based on its own real-time tasks, and obtain the target point-of-sale data required for real-time calculation.

[0010] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the steps of the real-time point-of-sale data distribution method based on Flink as in the first aspect are implemented.

[0011] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the steps of the real-time tracking data distribution method based on Flink as in the first aspect are implemented.

[0012] In a fifth aspect, an embodiment of the present application provides a computer program product, which is stored in a non-volatile storage medium. When the computer program product is executed by a processor, it implements the steps of the real-time point-of-sale data distribution method based on Flink as in the first aspect.

[0013] In a sixth aspect, an embodiment of the present application provides a chip comprising a processor and a communication interface, wherein the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the steps of the real-time tracking data distribution method based on Flink as in the first aspect.

[0014] The present application provides a method, device, equipment, medium and product for distributing real-time tracking data based on Flink. By allocating a single Flink program to Kafka and implementing dynamic routing distribution based on the tracking write configuration table, Kafka's consumption logic is rebuilt. Specifically, in a single Flink program, by reading the event name and target data table associated with the tracking write configuration table, the real-time tracking data of each business line / scenario is directly filtered out from each Kafka Topic, and distributed and written to the pre-associated independent tracking data table. This transforms the traditional mode of multiple Flink programs repeatedly consuming the entire amount of data into a single program consuming once and distributing on demand, reducing the I / O pressure of the Kafka network from O(n) of multiple programs superimposed to O(1) of a single program, and eliminating the waste of computing resources and operation and maintenance complexity caused by multiple jobs running in parallel. At the same time, through pre-distributed business isolation storage (instead of full-width tables), the scale of the tracking data table is compressed from full expansion to lightweight storage by business, reducing the Hologres storage cost. Downstream businesses do not need to scan and filter the entire table, but can directly access the corresponding independent tracking data table to complete the query, effectively shortening the query response time, improving query efficiency, and ensuring real-time performance in high-concurrency scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings in the embodiments of the present application.

[0016] Figure 1 This is a flowchart of a real-time tracking data distribution method based on Flink provided in one embodiment of the present application;

[0017] Figure 2 This is an exemplary schematic diagram of a real-time tracking configuration interface provided by an embodiment of the present application;

[0018] Figure 3 This is a flowchart of a real-time tracking data distribution method based on Flink provided by another embodiment of the present application;

[0019] Figure 4 This is a flowchart of a real-time tracking data distribution method based on Flink provided in another embodiment of the present application;

[0020] Figure 5 This is a structural diagram of a real-time tracking data distribution device based on Flink provided in an embodiment of the present application;

[0021] Figure 6 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] The principles and spirit of the present application will be described below with reference to several exemplary embodiments. It should be understood that the purpose of providing these embodiments is to make the principles and spirit of the present application clearer and more thorough, so that those skilled in the art can better understand and implement the principles and spirit of the present application. The exemplary embodiments provided herein are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments herein, all other embodiments obtained by those of ordinary skill in the art without creative work are within the scope of protection of this application.

[0023] In this document, terms such as first, second, and third are only used to distinguish one entity (or operation) from another entity (or operation), and are not intended to require or imply any order or relationship between these entities (or operations).

[0024] The following is a brief description of the concepts, technical terms and other related contents that may be involved in the embodiments of this application.

[0025] Tracking data or tracking data analysis is a commonly used data collection method for website analysis. It refers to attaching data collection program code to the functional program code at the "operation node" where data needs to be collected, and capturing, processing and sending related technologies and implementation processes for user behaviors or events on the operation node.

[0026] Flink is a mainstream component used in real-time computing and plays a key role in this field. Hologres, as a carrier of real-time data, enables real-time aggregation, analysis, and querying of data. Flink extracts data from upstream sources, such as Kafka and MySQL, cleanses it according to specific rules, and writes it to Hologres for downstream real-time reading and analysis.

[0027] Before describing the technical solutions provided by the embodiments of the present application, in order to facilitate understanding of the embodiments of the present application, the present application first specifically describes the problems existing in the related art:

[0028] In related technologies, multiple Flink programs are created based on multiple business needs, and the name of the tracking event corresponding to the business needs is written in each Flink program. Each Flink program needs to obtain the tracking data required for the business needs from Kafka's full tracking data, and then store it in the data table corresponding to the Flink program.

[0029] In this way, since each Flink program obtains data from the full amount of buried data, multiple Flink programs will consume a lot of bandwidth, which will not only consume a lot of Flink resources but also be difficult to manage. At the same time, there will be many consumers on the Kafka side, causing the network bandwidth to be full, resulting in a slowdown in Kafka data transmission.

[0030] If only one Flink program is used, the Flink program obtains the full amount of tracking data from Kafka and stores it in a Hologres data table. This results in a large amount of data in the data table, increasing the storage cost of Hologres. Then, the downstream data processing nodes filter out the data required for business needs from the full amount of tracking data, making queries more difficult.

[0031] In view of the above research findings of the inventors, the embodiments of the present application provide a real-time tracking data distribution method, device, equipment, medium and product based on Flink, which aims to solve the current problems of high Flink resource consumption and low transmission efficiency during tracking data transmission.

[0032] The following, in conjunction with the accompanying drawings, describes in detail the Flink-based real-time tracking data distribution method provided in the embodiments of the present application through specific embodiments and their application scenarios.

[0033] Figure 1 This is a flow chart of a real-time tracking data distribution method based on Flink provided by an embodiment of the present application. The execution entity of the real-time tracking data distribution method based on Flink can be a big data platform.

[0034] The following uses the execution subject of the Flink-based real-time tracking data distribution method as an example to illustrate the Flink-based real-time tracking data distribution method of this application. It should be noted that the above execution subject and application scenario do not constitute a limitation of this application.

[0035] like Figure 1 As shown, the real-time tracking data distribution method based on Flink provided in an embodiment of the present application may include steps 110 to 150.

[0036] Step 110: In the real-time tracking configuration interface of the big data platform, configure tracking events corresponding to each business line and tracking events corresponding to each business scenario under each business line based on business needs.

[0037] Step 120: Based on the business lines and / or business scenarios corresponding to each tracking event, as well as the tracking data tables associated with each business line and business scenario, establish an association between each tracking event and the tracking data table. Create a tracking write configuration table based on the association and store it in the real-time data warehouse.

[0038] Step 130: When a single stream processing engine Flink program is assigned to the consumer message queue system Kafka, the single Flink program is called to read the tracking write configuration table from the real-time data warehouse, and the single Flink program is used to consume the real-time tracking data in each Kafka data topic according to the tracking write configuration table.

[0039] Step 140: In the process of consuming real-time tracking data using a single Flink program, the real-time tracking data corresponding to the event name in the tracking write configuration table is filtered from the real-time tracking data to obtain batch tracking data corresponding to the business needs;

[0040] Step 150, based on the association between each burying point event and the burying point data table in the burying point writing configuration table, distribute and write the batch burying point data to the corresponding burying point data table, so that each downstream data processing node can read different burying point data tables based on its own real-time task, obtain the target burying point data required by each node for real-time calculation.

[0041] The real-time tracking data distribution method based on Flink provided in the embodiment of the present application reconstructs Kafka's consumption logic by assigning a single Flink program to Kafka and implementing dynamic routing distribution based on the tracking write configuration table. Specifically, in a single Flink program, by reading the event name and target data table associated with the tracking write configuration table, the real-time tracking data of each business line / scenario is directly filtered out from each Kafka Topic, and distributed and written to the pre-associated independent tracking data table. This transforms the traditional mode of multiple Flink programs repeatedly consuming the entire amount of data into a single program consuming once and distributing on demand, reducing the I / O pressure of the Kafka network from O(n) of multiple programs superimposed to O(1) of a single program, and eliminating the waste of computing resources and operation and maintenance complexity caused by multiple jobs running in parallel. At the same time, through pre-distributed business isolation storage (instead of full-width tables), the data table size is compressed from full expansion to lightweight storage by business, reducing Hologres storage costs. Downstream businesses do not need to scan and filter the entire table, but can directly access the corresponding independent data table to complete the query, effectively shortening the query response time, improving query efficiency, and ensuring real-time performance in high-concurrency scenarios.

[0042] The specific implementation of the above steps will be described in detail below with reference to specific embodiments.

[0043] Involving step 110, in the real-time tracking configuration interface of the big data platform, based on the business lines and / or business scenarios corresponding to each tracking event, and the tracking data tables associated with each business line and business scenario, an association relationship is established between each tracking event and the tracking data table, and based on the association relationship, a tracking write configuration table is created and stored in the real-time data warehouse.

[0044] In step 110, the fields of the tracking point writing configuration table include the event name of the tracking point event and the table name of the tracking point data table. Each tracking point data table is stored in the detail data (Data Warehouse Detail, DWD) layer of the real-time data warehouse. The DWD layer is one of the core layers in the data warehouse hierarchical architecture, mainly storing atomic granularity data that has been cleaned and integrated.

[0045] Specifically, in response to the user's burying point configuration operations based on business needs, the business line and / or business scenario corresponding to each burying point event can be determined, and then each burying point event can be associated with the burying point data table of the corresponding business line and the corresponding business scenario, so that the burying point data of the acquired burying point event can be written into the corresponding burying point data table.

[0046] In some embodiments of the present application, in order to solve the problem of slow query speed and data transmission speed caused by excessive data volume in the tracking data table, the above-mentioned step 110 establishes an association relationship between each tracking event and the tracking data table based on the business line and / or business scenario corresponding to each tracking event and the tracking data table associated with each business line and business scenario. The step 110 may specifically include the following steps:

[0047] Obtain the daily average data volume of each tracking event within the historical time period;

[0048] When the daily average data volume of a tracking event is less than or equal to a preset data volume threshold, the tracking data table associated with the business line and / or business scenario corresponding to the tracking event is determined as the tracking data table associated with the tracking event.

[0049] When the daily average data volume of a tracking event is greater than the preset data volume threshold, a new tracking data table is created and associated with the tracking event.

[0050] Specifically, the historical time period can be a preset time period before the current moment. The preset time period and the preset data volume threshold can be set according to specific needs, and this application does not make specific limitations on this.

[0051] As a specific example, the preset data volume threshold is set to 40 million items, then the event classification rule can be: if the daily data volume of a single buried event exceeds 40 million items, the buried data corresponding to the single buried event will be split into a table, for example, the exposure event table name on the C side is: dwd_c_event_exposure; if the daily data volume of a single buried event does not exceed 40 million items, the buried event will be merged into the buried data table associated with the business line to which it belongs, for example, the merged event table name on the C side is: dwd_c_event_merge.

[0052] In this way, when the amount of tracking data for a certain tracking event is too large, a separate tracking data table can be created for this tracking event. By partitioning the table, the amount of data in a single table can be controlled within the threshold, avoiding the problem of slow query speed and data transmission speed caused by the expansion of data in a single table, balancing the storage load, ensuring query efficiency, and adaptively adjusting the storage structure according to data growth, reducing the cost of manual intervention.

[0053] In offline computing, if a new tracking event is added, the existing SQL script must be modified, the new event name must be added, and the task must be republished. In real-time development, if offline computing practices are followed, the Flink task code must be modified and the task restarted. However, real-time tracking tasks are core underlying tasks, and many downstream real-time tasks will use the tracking data produced by these tasks. Furthermore, many downstream real-time tasks cannot experience data interruptions, such as providing activation data effects to third-party suppliers in real time for financial settlement. Therefore, it is undesirable to restart real-time tasks due to event modification, thereby disrupting downstream data.

[0054] Based on this, in some embodiments of the present application, in order to solve the problem of real-time task interruption caused by adding tracking events in the above business requirements, after creating the tracking write configuration table based on the association relationship and storing it in the real-time data warehouse in the above step 110, the present application may further include the following steps:

[0055] In the real-time tracking configuration interface, add, delete, and modify tracking events corresponding to each business line and each business scenario to update the tracking write configuration table in real time.

[0056] When the Flink program is started, the tracking data is loaded and written into the configuration table at regular intervals.

[0057] The global management node in the Flink program compares the current reading with the last reading of the tracking write configuration table. When it is determined that the tracking write configuration table has changed, the change is updated to the cached data and synchronized to each computing node in the Flink program in the form of a broadcast stream. This allows each computing node to consume real-time tracking data based on the latest tracking write configuration table.

[0058] Specifically, Flink's broadcast stream function is used to periodically read tracking data and write it into the configuration table (the interval for periodic reading can be configured). After reading the configuration table data, it is determined whether the data has been updated or changed. The changed data is sent to each thread in the form of a broadcast. After receiving the data, each thread updates its own cached configuration data and re-filters the tracking data.

[0059] For example, the user adds a "Payment Successful" tracking event in the configuration interface and updates the tracking write configuration table; the Flink program pulls the latest tracking write configuration table every 5 minutes and broadcasts the changes to each computing node; the computing node immediately filters the payment event according to the new rules and writes it to the corresponding table.

[0060] In this way, the Flink program periodically detects changes to the tracking write configuration table and uses the broadcast stream feature to synchronize these changes to all Flink compute nodes. This prevents data distribution errors caused by configuration differences between nodes and enables hot rule updates. Flink jobs continue to run during configuration updates, allowing the tracking events required by Flink tasks to be dynamically adjusted without restarting the Flink task, enabling hot processing and hot deployment of Flink program configurations without downtime. Furthermore, Flink tasks can be dynamically adjusted by simply updating the tracking write configuration table, without modifying the code, thus reducing the workload of code modification.

[0061] In some embodiments of the present application, the real-time tracking configuration interface may include Figure 2 The Add, Modify, Delete, and Export controls shown in the figure allow you to add, delete, and modify the tracking events required by your business needs. The Export control allows you to export tracking events into the configuration table with one click.

[0062] In addition, in step 110, the user can also perform tracking configuration operations based on business needs in the real-time tracking configuration interface, such as Figure 2 As shown, the business line corresponding to each burying point event and the target table name (buried point data table to be written) are set. The target table name is the table name of the buried point data table associated with the buried point event.

[0063] In some embodiments of the present application, in order to quickly add, delete, and modify the tracking events corresponding to each business line and business scenario, an enable control can be set for each tracking event in the real-time tracking configuration interface. The enable control is used to control the switching of the tracking event between the enabled state and the disabled state. Only in the enabled state, the tracking write configuration table contains the event name of the tracking event. If in the disabled state, the tracking write configuration table does not contain the event name of the tracking event.

[0064] Based on this, when a certain burying point event needs to be deleted, the burying point event is switched from the enabled state to the disabled state by touching the enabled control; when a certain burying point event needs to be restored, the burying point event is switched from the disabled state to the enabled state by touching the enabled control.

[0065] In this way, when the Flink task does not need a certain burying point event temporarily, the burying point event can be directly switched to the deactivated state through the enabling control, avoiding the need to re-fill burying point event information for the next Flink task due to direct deletion, and reducing the complexity of configuration operations.

[0066] Step 120 is involved. Based on the business lines and / or business scenarios corresponding to each burying point event, and the burying point data tables respectively associated with each business line and business scenario, an association relationship between each burying point event and burying point data table is established, and a burying point writing configuration table is created based on the association relationship and stored in the real-time data warehouse.

[0067] In step 120, each business line and each business scenario has an associated burying point data table. For each burying point event, if the burying point event corresponds to business line A and business scenario A1 under business line A, it is associated with the burying point data table associated with business line A and the burying point data table associated with business scenario A1. The burying point events in the same burying point data table usually involve the same business line or the same business scenario.

[0068] Specifically, taking the big data platform as a recruitment data platform, each business line at least includes a job seeker side and a recruiter side, and the business scenarios under the job seeker side at least include constructing a job seeker user portrait and an IM chat session, and the business scenarios of the recruiter side at least include constructing a recruiter user portrait, value-added services, and job positions.

[0069] Exemplarily, for the business line c side (job seeker side), the business scenario 1 is the c side user portrait, and the burying point events that need to be collected at least include: share_click (click to share), appclick (page click), and rollpagescreen (slide page); the business scenario 2 is a new activated user of the marketing department (determined by burying point whether the user registers after being promoted by the marketing department), and the burying point events that need to be collected at least include: app_regist (registration), app_login (click), and pageview (view page); the business scenario 3 is an IM chat session, and the burying point events that need to be collected at least include: im_chat (click chat), im_enter (enter IM chat session), and im_reply (reply chat).

[0070] For the B side of the business line (HR side), business scenario 1 is the HR user portrait, and the events that need to be collected include at least: resumeprocess (resume processing), app_login (login), resumeprint (resume export); business scenario 2 is the value-added service, and the events that need to be collected include at least: zzfwresult (value-added service processing result), zzfwpurchase (purchase of value-added services), zzfwks (value-added service consumption deduction); business scenario 3 is the position, and the events that need to be collected include at least: job_view (position view), job_post (position delivery), job_pub (position listing).

[0071] For business line AB experiments, business scenario 1 is a UI experiment, and the tracking events that need to be collected include at least: abtest_appclick (AB experiment page click), abtest_appview (AB experiment page view), abtest_rollpagescreen (AB experiment sliding page); business scenario 2 is a C-end experiment, and the tracking events that need to be collected include at least: abtest_postclick (AB experiment delivery click), abtest_shortexposure (AB experiment short exposure), abtest_search (AB experiment search); business scenario 3 is a B-end experiment, and the tracking events that need to be collected include at least: abtest_resumeprocess (AB experiment resume processing), abtest_app_login (AB experiment login), abtest_resumeprint (AB experiment resume export).

[0072] Involving step 130, when a single stream processing engine Flink program is assigned to the consumer message queue system Kafka, a single Flink program is called to read the point-of-sale write configuration table from the real-time data warehouse, and a single Flink program is used to consume the real-time point-of-sale data in each data topic Topic of Kafka based on the point-of-sale write configuration table.

[0073] In step 130, instead of assigning a Flink program to each business to consume all the data from Kafka (the source of the tracking data), a single Flink program is assigned to Kafka, replacing the multiple Flink programs originally corresponding to multiple businesses. Furthermore, because the tracking write configuration table contains the event names of the tracking events required by each business line and business scenario, a single Flink program only needs to consume Kafka once based on the tracking write configuration table to obtain the real-time tracking data required by each business line and business scenario, eliminating the need for multiple consumption.

[0074] Involving step 140, in the process of consuming real-time implant data by using a single Flink program, the real-time implant data corresponding to the event name in the implant write configuration table is filtered from the real-time implant data, and the batch implant data corresponding to the business requirement is obtained.

[0075] In step 140, for real-time computing, not all implant data of implant events is needed, only part of the implant data corresponding to the business requirement is needed, therefore, the Flink program can only filter the real-time implant data corresponding to the event name in the implant write configuration table from Kafka, and obtain the batch implant data corresponding to the business requirement.

[0076] Involving step 150, according to the association between each implant event in the implant write configuration table and the implant data table, the batch implant data is distributed and written to the corresponding implant data table, so that each downstream data processing node reads different implant data tables based on respective real-time tasks to obtain the target implant data required by each node for real-time computing.

[0077] In step 150, the implant events and implant fields required to be obtained can be defined in the real-time task, and each downstream data processing node can correspond to at least one real-time task. For example, the batch implant data contains data p1, and p1 belongs to implant event P1, which can be written to the implant data table associated with P1.

[0078] Specifically, after the Flink program consumes the real-time implant data of the upstream Kafka, the FlinkJar form can be used, that is, the Java code is used to execute the logic cleaning that SQL cannot handle, that is, the real-time implant data is filtered based on the implant write configuration table, only the implant data corresponding to the implant event in the implant write configuration table is retained, the batch implant data is obtained, and each implant data is assigned a corresponding implant event name, and finally stored in the Tuple3(implant data, implant event name, implant data table) tuple format(Flink program temporarily stores data when processing data), each implant data can obtain a Turple3 data.

[0079] Then, using the Put object of the official holo connection tool of holo-client, each Turple3 data is encapsulated in the form of Put entity class, and all data in the Put object is written to Hologres at regular intervals, and the Put object is emptied after writing is completed to wait for the next processing.

[0080] In the related art, the buried data of the buried event is sent in json format. The data format is relatively loose and flexible. The buried fields contained in each buried event are inconsistent. Therefore, it is necessary to configure all the buried fields involved in the buried event in the buried data table. If different buried data tables store the buried data of different buried events, it is also necessary to configure different table fields and table structures for different buried data tables. The configuration steps are relatively cumbersome. And in offline calculations, if a unique field of a certain buried event appears, it is necessary to manually add the unique field to the buried data table, adjust the table structure, and modify the SQL script and redeploy it online to parse the field value of the unique field from the buried data in json format based on the modified SQL script. In this way, the workload of data table configuration and code modification is greatly increased, the efficiency of buried data distribution is reduced, and the program cannot be restarted in real-time calculations.

[0081] In some embodiments of the present application, in order to solve the above problem, before the above step 150 of distributing and writing the batch burial point data to the corresponding burial point data table, the following steps may be further included:

[0082] In the tracking data tables associated with each business line and each business scenario, configure the same N fixed reporting fields and one target tracking field.

[0083] Specifically, N is a positive integer. Fixed reporting fields are required to be reported for all tracking events and are used for basic data analysis. For example, they may include, but are not limited to: tracking event name (event), tracking reporting time (event_time), data write time to the holo table (etl_time), data ID (pk), and user ID (accountid). The target tracking field is used to store the data content of tracking fields in the tracking data other than the N fixed reporting fields. In other words, it stores event-specific fields other than the fixed fields and can be flexibly expanded using formats such as JSON.

[0084] As a specific example, assume that the tracking data table for an e-commerce platform is configured as follows: fixed reporting fields (N = 5): user_id, event_time, device_type, app_version, geo_location; and target tracking field: extended_params (a JSON string storing dynamic fields such as product ID and order amount). When the Flink program processes a click event, only the five fixed fields are parsed and written to the corresponding columns; the remaining fields are directly stored in the extended_params field.

[0085] In this way, the same N fixed reporting fields and a target burying point field can be set in all burying point data tables, the table fields and table structures are the same, there is no need to configure different burying point data tables, and there is no need to configure all burying point fields of each burying point event in the burying point data table, thereby effectively reducing the complexity and workload of the configuration of the burying point data table. Moreover, the target burying point field as a "flexible field container" can store the data content of all burying point fields except the fixed reporting fields, thereby avoiding frequent DDL operations (such as ALTER TABLE) caused by the addition of new burying point fields, reducing the frequency of table structure changes, and ensuring the stability of writing.

[0086] In some embodiments of the present application, the real-time burying point data in Kafka is in json format, and the step 150 of distributing and writing the batch burying point data into the corresponding burying point data table can specifically include the following steps:

[0087] For each piece of json format burying point data, only the data content of the N fixed reporting fields in the burying point data is parsed to obtain N field values corresponding to the N fixed reporting fields;

[0088] Based on the burying point event to which the burying point data belongs, the corresponding burying point data table is matched as a writing object for the burying point data;

[0089] The N field values are respectively written into the corresponding N fixed reporting fields in the burying point data table, and the data content of the burying point fields in the burying point data except the N fixed reporting fields is written into the target burying point field in the burying point data table, and the target burying point field is in jsonb format.

[0090] Specifically, jsonb is a field format specific to Hologres, and ordinary json is only a string, but after upgrading to jsonb format, it supports Generalized Inverted Index (GIN) to speed up query, and binary storage is more compact, allowing partial updates. After storing the data content in json format in the target burying point field in jsonb format, the downstream data processing node can use special syntax to query the field value from the target burying point field.

[0091] For example, the fixed reporting fields include a1 and a2, and the fixed reporting fields a1 and a2 and a target burying point field b1 can be configured in each burying point data table. If the burying point event x includes burying point fields a1, a2, a3, and a4, after obtaining the real-time burying point data of the burying point event x, the data content of a1 and a2 is parsed to obtain the field value, and the json data of a3 and a4 is directly written into the b1 field.

[0092] Through the hierarchical design of fields (fixed reporting fields + target tracking fields) and the targeted parsing and jsonb storage of json format data, the efficiency and flexibility of tracking data processing have been significantly improved. Specifically, the same N high-frequency fixed reporting fields (such as user ID, event time, etc.) and a target tracking field in jsonb format for storing dynamic content are pre-configured in various tracking data tables. This allows the system to only parse the key N fixed reporting fields when processing real-time tracking data in json format in Kafka (instead of parsing all tracking fields in traditional solutions), reducing processing overhead. At the same time, writing the data of non-fixed fields as a json object into the target tracking field in jsonb format can avoid the need to change the table structure. When adding a new tracking field, there is no need to trigger the ALTER TABLE operation, ensuring write continuity. Storing the original loose json format string into a jsonb format field for downstream on-demand query can also improve query speed.

[0093] In the offline calculation process provided by the relevant technology, all tracking fields for each tracking event must be pre-configured in the offline table, and the obtained Kafka tracking data in JSON format must be fully parsed.

[0094] As a specific example, Figure 3 As shown in the figure, for the shortexposure tracking event, all fields (event, uuid, version, domain, content, initiactive) must be configured in advance in the offline table (dwd_c_event_merge_d) used to store the tracking event. Afterwards, when the Kafka tracking data (JSON format) for the tracking event is obtained, the Kafka tracking data must be fully parsed to extract the data values ​​of the six tracking fields in the tracking data and store them in the corresponding fields of the offline table for use by downstream SQL processing nodes.

[0095] Different from the table configuration and json parsing solution, continue to refer to Figure 3 The real-time table (i.e., the tracking data table) of this application only contains the fixed reporting field event and the target tracking field properties in jsonb format. In this way, when parsing the Kafka tracking data in json format, only the data value of event needs to be parsed. The remaining tracking fields do not need to be parsed and are directly written to the properties field in the real-time table, and then parsed when used by the downstream processing node.

[0096] When the tracking field of a tracking event changes, such as adding a new tracking field, or when the tracking field of a new tracking event is inconsistent with that of an existing tracking event, in the offline calculation process provided by the relevant technology, it is necessary to manually modify the table structure of the offline table and perform the field addition operation.

[0097] As a specific example, Figure 4 As shown, shortexposure is an existing buried event, and pageexposure is an existing buried event. The buried fields of the two are not completely consistent. The latter lacks the uuid field and adds a new deviceid field. Based on this, in the offline calculation process provided by the relevant technology, it is necessary to manually modify the table structure of the offline table and perform the add field operation: alter table dwd_c_event_merge_d add columns(decviceid string). Afterwards, when two pieces of Kafka buried data (json format) of the two buried events are obtained, each piece of Kafka buried data needs to be fully parsed to parse out the data values ​​of the seven buried fields in the buried data (if there is no corresponding buried field, the data value is empty) and store it in the corresponding field in the offline table for use by the downstream processing node SQL.

[0098] Different from the table configuration and json parsing solution, continue to refer to Figure 4 Even if a new buried event pageexposure is added, the table structure of the real-time table (buried data table) of this application does not need to be changed. When two pieces of Kafka buried data (json format) of two buried events are obtained, it is only necessary to parse the fixed reporting field event for each Kafka buried data to obtain the data value. The data content of the remaining non-fixed reporting fields is stored in properties without the need for json parsing, and will be parsed on demand when used by the downstream processing node. In this way, the target buried field properties acts as a "flexible field container" that can store the data content of the unique buried field decciceid of the newly added buried event, as well as the data content of the newly added buried field of the existing buried event. No matter how the buried field of the buried event changes, the buried data table maintains its original state, keeping the table fields and table structure unchanged.

[0099] In the embodiment of the present application, fixed reporting fields are defined in advance. The Flink program does not need to fully parse the real-time Kafka tracking data in JSON format. It only needs to parse the fixed reporting fields, which effectively reduces CPU consumption. In addition, by separating fixed fields from dynamic fields, it can avoid frequent changes in table structure due to field changes, optimize the storage structure, adapt to the rapid iteration of business, take into account data standardization and scalability, and improve the writing efficiency of tracking data.

[0100] In some embodiments of the present application, after writing the N field values ​​into the corresponding N fixed reporting fields in the buried point data table, and writing the data content of the buried point fields other than the N fixed reporting fields in the buried point data into the target buried point field in the buried point data table, the following steps may also be included:

[0101] Each downstream data processing node reads different tracking data tables based on its own real-time tasks to obtain the target tracking data it needs;

[0102] Based on the target SQL statement, the field value of the target tracking field in the target tracking data is parsed into JSON on demand, and the field value of the tracking field required for the real-time task is obtained and then calculated in real time.

[0103] Among them, the target SQL statement is used to define the required embedding fields for each real-time task, and the required embedding fields include at least part of the embedding fields in the target embedding event required by the real-time task, and the embedding data corresponding to the target embedding event is the target embedding data. The target embedding events defined by different real-time tasks can be the same or different, but even if the same target embedding event is defined, the embedding fields required to be obtained for the same target embedding event may be different, and this application does not make specific restrictions on this. If the embedding fields required by the real-time task include fixed reporting fields, the field value of the fixed reporting field can be directly obtained from the target embedding data.

[0104] As a specific example, after batch tracking data is written to Hologres, the tracking data for the pageexposoure event will appear in the associated tracking data table. When using this data downstream, if developers need the data in the actiontype tracking field in this tracking data, they do not need to worry about whether the actiontype field exists in the tracking data table. Instead, they only need to query the properties of the target tracking field in jsonb format using the specific SQL syntax: selectproperties->>'actiontype'from dwd_c_event_pageexposoure to obtain the data of the 'actiontype' element in json format in the original tracking data.

[0105] In an embodiment of the present application, when the Flink program writes the buried point data into the buried point data table, it only needs to perform json parsing on the fixed reported fields, and the remaining fields are directly stored in the target buried point field, without the need for full json parsing. Afterwards, when the downstream data processing node reads the buried point data table, it parses the json data content stored in the target buried point field on demand according to the assigned real-time task, and only parses the required buried point fields. Similarly, there is no need to parse all the json data content under the target buried point field, which can effectively reduce the amount of parsed data, improve the data acquisition efficiency of the downstream data processing node, and thus improve the execution efficiency of the real-time task. In addition, the binary storage and compression characteristics of jsonb (such as dictionary encoding) reduce the size of a single data entry from 1.2KB of traditional flat storage to 450B, and jsonb supports efficient generalized inverted indexes (Generalized Inverted Index, GIN), so that when the downstream data processing node obtains the field values ​​required for the real-time task, it can reduce delays, speed up query response time, and achieve high-concurrency real-time analysis compared to full table scans.

[0106] Corresponding to the method embodiment of the present application, the present application also provides a real-time tracking data distribution device based on Flink.

[0107] Figure 5 This is a structural diagram of a real-time tracking data distribution device based on Flink provided by an embodiment of the present application. Figure 5 As shown, the Flink-based real-time tracking data distribution device 500 may include: a configuration module 510, an association module 520, a consumption module 530, a screening module 540 and a writing module 550.

[0108] Among them, the configuration module 510 is used to configure the burying events corresponding to each business line and the burying events corresponding to each business scenario under each business line based on business needs in the real-time burying configuration interface of the big data platform; the association module 520 is used to establish the association relationship between each burying event and the burying data table based on the business line and / or business scenario corresponding to each burying event, as well as the burying data table associated with each business line and business scenario, and create a burying write configuration table based on the association relationship and store it in the real-time data warehouse, wherein the fields of the burying write configuration table include the event name of the burying event and the table name of the burying data table, and each burying data table is stored in the detailed data DWD layer of the real-time data warehouse; the consumption module 530 is used to allocate a single stream processing engine Flink program to the consumer message queue system Kafka. In this case, a single Flink program is called to read the point-of-sale write configuration table from the real-time data warehouse, and a single Flink program is used to consume the real-time point-of-sale data in each data topic of Kafka based on the point-of-sale write configuration table; the screening module 540 is used to filter the real-time point-of-sale data corresponding to the event name in the point-of-sale write configuration table from the real-time point-of-sale data in the process of using a single Flink program to consume the real-time point-of-sale data, and obtain the batch point-of-sale data corresponding to the business needs; the writing module 550 is used to distribute and write the batch point-of-sale data to the corresponding point-of-sale data table based on the association relationship between each point-of-sale event and the point-of-sale data table in the point-of-sale write configuration table, so that each downstream data processing node can read different point-of-sale data tables based on its own real-time task, and obtain the target point-of-sale data required by each node for real-time calculation.

[0109] The real-time tracking data distribution device based on Flink provided in the embodiment of the present application reconstructs Kafka's consumption logic by allocating a single Flink program to Kafka and realizing dynamic routing distribution based on the tracking write configuration table. Specifically, in a single Flink program, by reading the event name and target data table associated with the tracking write configuration table, the real-time tracking data of each business line / scenario is directly filtered out from each Kafka Topic, and distributed and written to the pre-associated independent tracking data table. This transforms the traditional mode of multiple Flink programs repeatedly consuming the full amount of data into a single program consuming once and distributing on demand, reducing the I / O pressure of the Kafka network from O(n) of multiple programs superimposed to O(1) of a single program, and eliminating the waste of computing resources and operation and maintenance complexity caused by multiple jobs running in parallel. At the same time, through pre-distributed business isolation storage (instead of full-width tables), the scale of the tracking data table is compressed from full expansion to lightweight storage by business, reducing the Hologres storage cost. Downstream businesses do not need to scan and filter the entire table, but can directly access the corresponding independent tracking data table to complete the query, effectively shortening the query response time, improving query efficiency, and ensuring real-time performance in high-concurrency scenarios.

[0110] In some embodiments of the present application, it also includes: a configuration module, which is also used to configure the same N fixed reporting fields and a target burying word in the burying data table associated with each business line and each business scenario before distributing and writing batch burying data to the corresponding burying data table; wherein, the fixed reporting field is the burying field that must be reported for all burying events, and the target burying field is used to store the data content of the burying fields in the burying data except the N fixed reporting fields.

[0111] In some embodiments of the present application, the real-time buried point data in Kafka is in json format, and the writing module includes: a parsing unit, which is used to parse only the data content of N fixed reporting fields in the buried point data for each json format buried point data, and obtain N field values ​​corresponding to the N fixed reporting fields; a matching unit, which is used to match the corresponding buried point data table as the writing object based on the buried point event to which the buried point data belongs; a writing unit, which is used to write the N field values ​​into the corresponding N fixed reporting fields in the buried point data table respectively, and write the data content of the buried point fields in the buried point data except the N fixed reporting fields into the target buried point field in the buried point data table, and the target buried point field is in jsonb format.

[0112] In some embodiments of the present application, it also includes: a reading module, which is used to write N field values ​​into the corresponding N fixed reporting fields in the burying point data table, and write the data content of the burying point fields other than the N fixed reporting fields in the burying point data into the target burying point field in the burying point data table. Then, each downstream data processing node reads different burying point data tables based on its own real-time task to obtain the target burying point data required by each node; the parsing module is also used to perform json parsing on the field value of the target burying point field in the target burying point data as needed based on the target SQL statement, and perform real-time calculation after obtaining the field value of the burying point field required for the real-time task, wherein the target SQL statement is used to define the burying point field required for each real-time task.

[0113] In some embodiments of the present application, it also includes: an update module, which is used to create a tracking write configuration table based on the association relationship and store it in the real-time data warehouse. In the real-time tracking configuration interface, by adding, deleting, and modifying the tracking events corresponding to each business line and the tracking events corresponding to each business scenario, the tracking write configuration table is updated in real time; a loading module, which is used to load the tracking write configuration table at regular intervals when the Flink program has been started; a synchronization module, which is used to compare the current reading with the last reading of the tracking write configuration table through the global management node in the Flink program. When it is determined that the tracking write configuration table has changed, the change is updated to the cached data, and synchronized to each computing node in the Flink program in the form of a broadcast stream, so that each computing node consumes the real-time tracking data based on the latest tracking write configuration table.

[0114] In some embodiments of the present application, the association module is specifically used to: obtain the average daily data volume of each burial event in a historical time period; when the average daily data volume of the burial event is less than or equal to the preset data volume threshold, the burial data table associated with the business line and / or business scenario corresponding to the burial event is determined as the burial data table associated with the burial event; when the average daily data volume of the burial event is greater than the preset data volume threshold, a new burial data table is created to associate with the burial event.

[0115] In some embodiments of the present application, the big data platform is a recruitment data platform, and each business line includes at least a job seeker side and a recruiter side. The business scenarios on the job seeker side include at least building a job seeker user profile and IM chat sessions, and the business scenarios on the recruiter side include at least building a recruiter user profile, value-added services, and position status.

[0116] The real-time tracking data distribution device based on Flink provided in the embodiment of the present application can achieve Figures 1-4 The various processes implemented by the service platform in the method embodiment can achieve the same technical effect. To avoid repetition, they will not be described here.

[0117] Figure 6 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application.

[0118] like Figure 6 As shown, the electronic device 600 includes a memory 601 , a processor 602 , and a computer program stored in the memory 601 and executable on the processor 602 .

[0119] In an example, the processor 602 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0120] The memory 601 may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Therefore, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the real-time point-of-sale data distribution method based on Flink in the embodiment according to the first aspect of the present application.

[0121] The processor 602 runs a computer program corresponding to the executable program code by reading the executable program code stored in the memory 601, so as to implement the real-time tracking data distribution method based on Flink in the embodiment of the first aspect mentioned above.

[0122] In some examples, the electronic device 600 may further include a communication interface 603 and a bus 610. Figure 6 As shown, the memory 601 , the processor 602 , and the communication interface 603 are connected via a bus 610 and communicate with each other.

[0123] The communication interface 603 is mainly used to implement communication between the modules, devices, units and / or equipment in the embodiment of the present application. Input devices and / or output devices can also be connected through the communication interface 603.

[0124] Bus 610 includes hardware, software, or both, to couple components of electronic device 600 to each other and to couple electronic device 600 to other systems or devices. While Fig. 6 shows bus 610 as a single bus, bus 610 can be composed of multiple buses or separate communication lines, which are arranged to allow communication between various components of electronic device 600. For example, without limitation, bus 610 can include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-E) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or another suitable bus or a combination of two or more of these. Where appropriate, bus 610 can include one or more buses of the same type or buses of different types. Although the example embodiment described and shown herein includes a particular bus, the present application contemplates any suitable bus or interconnect.

[0125] The electronic device provided by the example embodiment of the present application can realize Figures 1-4 the processes realized by the electronic device in the method embodiment, and can realize the same technical effects. To avoid repetition, the same will not be described here.

[0126] In combination with the real-time burying point data distribution method based on Flink in the above-described embodiments, the example embodiment of the present application can provide a computer storage medium to realize. The computer storage medium has computer program instructions stored thereon; the computer program instructions are executed by a processor to realize the steps of any one of the real-time burying point data distribution methods based on Flink in the above-described embodiments.

[0127] In conjunction with the Flink-based real-time tracking data distribution method in the above embodiments, embodiments of the present application may provide a computer program product for implementation. This (computer) program product is stored in a non-volatile storage medium and, when executed by at least one processor, implements the steps of any of the Flink-based real-time tracking data distribution methods in the above embodiments.

[0128] An embodiment of the present application further provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned Flink-based real-time point-of-sale data distribution method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0129] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0130] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.

[0131] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in unit, a function card, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0132] It should also be noted that the example embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps are performed simultaneously.

[0133] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other processing device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other processing device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0134] The above only describes specific implementation manners of the present application. For the convenience and brevity of description, the specific working process of the system, module and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be described herein. It should be understood that the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed in the present application, and these modifications or replacements should be covered by the protection scope of the present application.

Claims

1. A real-time tracking data distribution method based on Flink, characterized by: include: In the real-time tracking configuration interface of the big data platform, configure tracking events for each business line based on business needs, as well as tracking events for each business scenario under each business line. Based on the business lines and / or business scenarios corresponding to each buried point event, as well as the buried point data tables associated with each business line and each business scenario, establish an association relationship between each buried point event and the buried point data table, and create a buried point write configuration table based on the association relationship and store it in the real-time data warehouse. The fields of the buried point write configuration table include the event name of the buried point event and the table name of the buried point data table. Each buried point data table is stored in the detailed data DWD layer of the real-time data warehouse; When a single stream processing engine Flink program is assigned to the consumer message queue system Kafka, the single Flink program is called to read the tracking write configuration table from the real-time data warehouse, and the single Flink program is used to consume the real-time tracking data in each data topic of Kafka according to the tracking write configuration table; In the process of consuming the real-time tracking data using the single Flink program, the real-time tracking data corresponding to the event name in the tracking write configuration table is filtered from the real-time tracking data to obtain batch tracking data corresponding to the business needs; Based on the association between each burying point event and the burying point data table in the burying point writing configuration table, the batch burying point data is distributed and written to the corresponding burying point data table, so that each downstream data processing node reads different burying point data tables based on its own real-time task, obtains the target burying point data required by each node for real-time calculation.

2. The method according to claim 1, characterized in that Before distributing and writing the batch burying point data into the corresponding burying point data table, the method further includes: In the tracking data tables associated with each business line and each business scenario, configure the same N fixed reporting fields and one target tracking word. Among them, the fixed reporting field is the burying point field that needs to be reported for all burying point events, and the target burying point field is used to store the data content of the burying point field in the burying point data except the N fixed reporting fields.

3. The method according to claim 2, characterized in that The real-time tracking data in Kafka is in JSON format. The batch tracking data is distributed and written to the corresponding tracking data table, including: For each piece of buried data in JSON format, only the data content of the N fixed reporting fields in the buried data is parsed to obtain the N field values ​​corresponding to the N fixed reporting fields; Based on the buried point event to which the buried point data belongs, matching the buried point data with the corresponding buried point data table as a write object; The N field values ​​are written into the corresponding N fixed reporting fields in the burying point data table respectively, and the data content of the burying point fields other than the N fixed reporting fields in the burying point data is written into the target burying point field in the burying point data table, and the target burying point field is in jsonb format.

4. The method according to claim 3, characterized in that After writing the N field values ​​into the corresponding N fixed reporting fields in the buried point data table, and writing the data contents of the buried point fields other than the N fixed reporting fields in the buried point data into the target buried point field in the buried point data table, the method further includes: Each downstream data processing node reads different tracking data tables based on its own real-time tasks to obtain the target tracking data it needs; Based on the target SQL statement, the field value of the target tracking field in the target tracking data is parsed in json as needed, and real-time calculation is performed after obtaining the field value of the tracking field required for the real-time task, wherein the target SQL statement is used to define the tracking field required for each real-time task.

5. The method according to claim 1, wherein After creating a tracking point writing configuration table based on the association relationship and storing it in the real-time data warehouse, the following steps are also included: In the real-time tracking configuration interface, the tracking events corresponding to each business line and each business scenario are added, deleted, and modified to update the tracking configuration table in real time. When the Flink program is started, the tracking data is loaded and written into the configuration table at regular intervals. Through the global management node in the Flink program, the current reading is compared with the last reading of the tracking write configuration table. When it is determined that the tracking write configuration table has changed, the change is updated to the cached data and synchronized to each computing node in the Flink program in the form of a broadcast stream, so that each computing node consumes real-time tracking data based on the latest tracking write configuration table.

6. The method according to claim 1, characterized in that Based on the business lines and / or business scenarios corresponding to each tracking event, as well as the tracking data tables associated with each business line and business scenario, establish an association between each tracking event and the tracking data table, including: Obtain the daily average data volume of each tracking event within the historical time period; When the daily average data volume of a tracking event is less than or equal to a preset data volume threshold, the tracking data table associated with the business line and / or business scenario corresponding to the tracking event is determined as the tracking data table associated with the tracking event. When the daily average data volume of a tracking event is greater than the preset data volume threshold, a new tracking data table is created and associated with the tracking event.

7. The method according to claim 1, characterized in that The big data platform is a recruitment data platform, and each business line includes at least a job seeker side and a recruiter side. The business scenarios under the job seeker side include at least building a job seeker user profile and IM chat sessions, and the business scenarios on the recruiter side include at least building a recruiter user profile, value-added services, and position status.

8. An electronic device, characterized in that: The electronic device comprises: a processor and a memory storing computer program instructions; and when the electronic device executes the computer program instructions, the method according to any one of claims 1 to 7 is implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the method according to any one of claims 1 to 7.

10. A computer program product, characterized in that The computer program product is stored in a non-volatile storage medium, and when the computer program product is executed by a processor, the method according to any one of claims 1 to 7 is implemented.