A Business Data Monitoring Method, Device and Storage Medium Based on Flink
Through the Flink-based business data monitoring method, the APP buried points are associated with the ClickHouse table, which solves the problems of high index costs and difficulty in using non-traditional SQL, and realizes high concurrency and easy-to-scaling data storage and monitoring.
Patent Information
- Application Number
- CN202111523688.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-14
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-12-14
AI Technical Summary
In the prior art, with the increase in APP burial points, the indexing cost is getting higher and higher. In addition, SQL that configures charts and alarms is not traditional SQL, so it is difficult for non-R&D personnel to use it.
Through the Flink-based business data monitoring method, data storage rules are configured, APP buried points are associated with ClickHouse table, N buried objects are associated into a structure table, and data filtering and storage are achieved through the log server, Flink and ClickHouse.
It realizes a highly concurrent and easy-to-scaling business data monitoring system, while reducing data storage costs.
Smart Images

Figure CN114186000B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method, device, and storage medium for monitoring business data based on Flink. Background Art
[0002] The acquisition of APP data metrics depends on data logging. In the early stage, the solution for APP data metric statistics was to log data on the APP, and then report it to the cloud. After storing it using ELK or other data storage carriers, a dedicated department would collect and calculate it to obtain the data required by the product team. This solution required the cooperation of a large number of data warehouse personnel, had a high cost, and was not good at supporting complex data logging.
[0003] To address these drawbacks, Alibaba Cloud launched the SLS logging system, which supports unstructured data storage and can solve the above-mentioned drawbacks. Its commercialized charts and alarm configurations can significantly shorten the data metric configuration process for data logging. Compared with the early APP data logging, it greatly reduces the technical threshold and cost. However, its configured metrics require creating indexes for query fields. As the number of APP data logging increases, the cost of indexes also becomes higher and higher. In addition, the SQL for its configured charts and alarms is not traditional SQL, which has a certain learning cost and is difficult for non-developers to use. Summary of the Invention
[0004] The main purpose of the present invention is to provide a method, device, and storage medium for monitoring business data based on Flink, aiming to solve the problems in the prior art that as the number of APP data logging increases, the cost of indexes becomes higher and higher, and in addition, the SQL for its configured charts and alarms is not traditional SQL, which has a certain learning cost and is difficult for non-developers to use.
[0005] To achieve the above purpose, the present invention provides a method for monitoring business data based on Flink, and the method includes the following steps:
[0006] Configure the data storage rules for the first business data;
[0007] Flink obtains the first business data from the log server, filters the first business data according to the data storage rules to obtain the second business data; stores the second business data in the database;
[0008] Query the database to obtain the third business data, and process the third business data.
[0009] Optionally, the configuring the data storage rules for the first business data includes the following steps:
[0010] Set the identifier and / or basic fields and / or custom fields of the first business data;
[0011] Establish an association relationship between the identifier of the first service data and a database table, and establish an association relationship between the identifiers of N (N≥1) pieces of the first service data and one database table;
[0012] Create fields of the database table corresponding to the identifier according to the basic fields and / or the custom fields;
[0013] Establish an association relationship between the basic fields and / or the custom fields and the fields of the database table.
[0014] Optionally, the method further includes:
[0015] Create a first SQL statement for querying field data corresponding to the basic fields and / or the custom fields corresponding to the identifier from the database corresponding to the identifier; the identifier and the first SQL statement are in one-to-one correspondence.
[0016] Optionally, the Flink obtains the first service data from a log server, filters the first service data according to the data storage rule to obtain second service data; stores the second service data in a database, including the following steps:
[0017] Filter the first service data according to the identifier of the first service data in the data storage rule to obtain fourth service data corresponding to the identifier;
[0018] Obtain field data corresponding to the basic fields and / or the custom fields from the fourth service data according to the basic fields and / or the custom fields corresponding to the flag;
[0019] Create a second SQL statement according to the field data and the database table corresponding to the identifier, and execute the second SQL statement to save the field data into the database table corresponding to the identifier.
[0020] Optionally, the method further includes the following steps:
[0021] Create a third SQL statement for batch storing data for field data corresponding to multiple flags and / or database tables corresponding to multiple flags, and execute the third SQL statement to save the field data corresponding to the multiple flags into the database.
[0022] Optionally, the method further includes the following steps:
[0023] The Flink regularly obtains the data storage rule, and then updates the locally saved data storage rule with the latest data storage rule.
[0024] Optionally, querying the database to obtain third service data and processing the third service data includes the following steps:
[0025] Obtain the corresponding first SQL statement according to the identifier;
[0026] Modify the first SQL statement according to the query condition to obtain a fourth SQL statement;
[0027] Execute the fourth SQL statement to query the field data corresponding to the basic field and / or the custom field from the database.
[0028] In addition, to achieve the above object, the present invention also proposes a service data monitoring device based on Flink, and the device includes:
[0029] A data configuration unit for configuring a data storage rule for first service data;
[0030] A data storage unit for Flink to obtain the first service data from a log server, filter the first service data according to the data storage rule to obtain second service data; and store the second service data in a database;
[0031] A data processing unit for querying the database to obtain third service data and processing the third service data.
[0032] In addition, to achieve the above object, the present invention also proposes an electronic device, and the electronic device includes: a memory, a processor, and a Flink-based service data monitoring program stored on the memory and executable on the processor, and the Flink-based service data monitoring program is configured to implement the steps of the Flink-based service data monitoring method as described above.
[0033] In addition, to achieve the above object, the present invention also proposes a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the Flink-based service data monitoring method as described above are implemented.
[0034] Through the embodiments of the present invention, the APP buried point service data is associated with the ClickHouse table, N buried point objects are associated into a structure table, and the buried point fields are corresponding to the fields of the structure table one by one; and the log server, Flink, and ClickHouse are used for collaborative processing. Thus, it is ensured that the service data monitoring system has high concurrency, is easy to expand, and at the same time reduces the data storage cost. Description of the Drawings
[0035] Figure 1 A flowchart of the business data monitoring method based on Flink provided by the present invention.
[0036] Figure 2 A structural diagram of the business data monitoring system based on Flink provided by the present invention.
[0037] Figure 3 A flowchart of the process of configuring data storage rules provided by the present invention.
[0038] Figure 4 A structural diagram of the system for business data warehousing provided by the present invention.
[0039] Figure 5 A flowchart of the process of business data warehousing provided by the present invention.
[0040] Figure 6 A flowchart of the process of querying business data provided by the present invention.
[0041] Figure 7 A structural block diagram of an embodiment of the business data monitoring device based on Flink provided by the present invention.
[0042] Figure 8 A structural diagram of an electronic device provided by an embodiment of the present invention.
[0043] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0044] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not used to limit the present invention.
[0045] In subsequent descriptions, the use of suffixes such as "module", "component" or "unit" for representing elements is only for the convenience of description of the present invention, and they have no specific meaning in themselves. Therefore, "module", "component" or "unit" can be used interchangeably.
[0046] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0047] In one embodiment, as Figure 1 shown, the present invention provides a business data monitoring method based on Flink, and the method includes:
[0048] Step 101: Configure the data storage rules for the first service data.
[0049] In the era of mobile Internet, the APP service has developed rapidly. Obtaining the usage status of the APP in real time is of great significance for subsequent service decision-making and quality monitoring. The traditional method of obtaining APP usage data metrics is to embed points at certain positions in the APP to count data metrics, then report the embedded point data, and then collect the embedded points for data analysis. Then, the trend of a certain embedded point data is presented in the form of a chart to support service decision-making.
[0050] The APP in high-speed iteration has a new version on average every two weeks. Service decision-making highly depends on the data of the embedded points. The product team will adjust the service direction in a timely manner according to the reported data of the embedded points, and the R & D team will also optimize according to the trend of the embedded points. Accurate and real-time embedded points are of great significance for the development of the APP. Large-scale APPs have many embedded points, and new embedded points will be added with each version according to the service scenario. An embedded point system with high scalability and easy creation can greatly improve the efficiency of embedded points.
[0051] The service data monitoring system based on Flink is as Figure 2 shown. It includes data definition, data reporting, data storage, and data parsing and display. Data definition supports the APP to flexibly define embedded point fields. After the SDK in the APP generates log records, they are reported to Alibaba Cloud SLS for caching. Then, Flink is used to consume the data records from Alibaba Cloud SLS and enter them into the corresponding ClickHouse table for data storage. Then, the data metrics are supported to be displayed on the front end in the way of defining SQL.
[0052] When the Flink cluster processes the data in the Alibaba Cloud SLS server, it needs to be processed according to the data storage rules. Therefore, it is necessary to configure the data storage rules first. For the configuration of the data storage rules, see the Figure 3 process shown.
[0053] Step 201: Set the identifier and / or basic fields and / or custom fields of the first service data.
[0054] APP data metrics, from the perspective of usage, can be mainly divided into two categories. One is statistical data metrics, such as PV and UV, which only count the times users have used a certain function, for example, clicked on a certain button or entered a certain page. The content of such data points is simple, and only one record needs to be generated each time a user uses it. It is mainly used to display the usage times and the number of users of a certain function. The other is quality data points, which mainly record the performance and abnormal data during the user's usage process. The data content is relatively complex and needs to include the exception stack, the version of the exception, the model, and the network, etc. Different environments have an impact on the quality of the APP. The fields of such data points are relatively flexible and vary according to different business scenarios, and sufficient support needs to be provided for quality iteration.
[0055] After collecting the operation records of users through the SDK, APP data points generate log files and report them to the Alibaba Cloud SLS server. Each log record has a unique identifier to distinguish the log record; at the same time, each log record contains corresponding basic fields and custom fields. The basic fields include: UUID, userID, app version number, device type, os type, etc.; the custom fields can be customized according to each type of diary record. For example, the custom fields of the log record for recording users browsing products include: log type, product details page, favorite product, purchase product. When analyzing the log file, the system can set the corresponding basic fields and custom fields according to business requirements, and this technical solution does not make specific restrictions.
[0056] Operation and maintenance personnel can dynamically modify the identifiers, basic fields, and custom fields in the data storage rules as needed. When analyzing new data points (i.e., new log records), the corresponding identifier needs to be set according to the new log records; it is also possible to reset the basic fields and custom fields for the already analyzed (i.e., already configured with identifiers) log records. For example, when analyzing a certain user's operation behavior, the corresponding fields can be added. The added fields need to have corresponding records in the data points.
[0057] For example, the set identifiers, basic fields, and custom fields are shown in the following table:
[0058]
[0059] Step 202, establish the association relationship between the identifier of the first business data and the database table, and establish the association relationship between N (N is greater than or equal to 1) identifiers of the first business data and 1 database table.
[0060] The present invention uses a structured data table to generate a general multi-dimensional table, and performs an association configuration between the buried point object and the storage table, so as to support processing the field differences of different business buried points, associate N buried point objects into a structure table, and correspond the buried point fields with the fields of the structure table one by one, thus completing the structural transformation of the buried point definition.
[0061] Log data usually does not change. ClickHouse is used as the storage database, which supports maintaining good query speed in the case of large-scale storage.
[0062] The unique identifier of the buried point is called Section. The relationship between the buried point Section and the ClickHouse table is n:1. The relationship between the buried point and the ClickHouse table is managed by manual configuration. The n buried points associated with the ClickHouse table will all be stored in this ClickHouse table, and then SQL queries are performed on the single table, and the query results in three dimensions (horizontal axis, vertical axis, line name) are provided for the front end, so as to support multi-dimensional data indicators.
[0063] For example, if the log records corresponding to the three buried points A, B, and C are all stored in the Table_ID_001 table of the ClickHouse data, then the identifiers of the log records corresponding to these three buried points are associated and configured with the table name, as configured in the following table:
[0064] Diary record name Log record identifier Table name Buried point A Product browsing log A Table_ID_001 Buried point B Product browsing log B Table_ID_001 Buried point C Product browsing log C Table_ID_001
[0065] Step 203: Create the fields of the database table corresponding to the identifier according to the basic fields and / or the custom fields.
[0066] After associating N buried point objects into a structure table, the buried point fields are corresponded with the fields of the structure table one by one, thus completing the structural transformation of the buried point definition. For example, buried points A, B, and C contain multiple basic fields and custom fields, and then these fields are corresponded to the database table fields one by one. The fields contained in buried points A, B, and C are shown in the following table:
[0067]
[0068] Then the fields contained in the database table Table_ID_001 corresponding to buried points A, B, and C are shown in the following table:
[0069]
[0070] Step 204: Establish an association relationship between the basic fields and / or the custom fields and the fields of the database table.
[0071] The fields recorded in the buried point logs need to correspond one by one to the corresponding database table fields. Subsequently, Flink needs to store the field data in the logs into the corresponding field relationships in the database table according to this correspondence. The configured correspondence is shown in the following table:
[0072]
[0073] Step 205: Create a first SQL statement to query the field data corresponding to the basic field and / or the custom field corresponding to the identifier from the database corresponding to the identifier; the identifier and the first SQL statement correspond one by one.
[0074] Write a SQL statement for database query for the log records of each buried point. When querying the log record data of this buried point subsequently, obtain the corresponding SQL statement according to the log record identifier of this buried point for query. For example, the SQL statement corresponding to the query of the log record data of buried point A is:
[0075] select userID, APP version number, product details page A, favorite product A from Table_ID_001; The specific SQL query statement can be written according to actual requirements, and this technical solution does not make any limitations.
[0076] After configuring the data storage rules for the buried point log records, it is necessary to send the data storage rules to the Flink cluster. The data storage rules of the present invention are configured through the WEB management terminal, saved to the database table after configuration, such as in the section configuration table, and then the Flink cluster regularly obtains the data storage rules from the section configuration table.
[0077] After the Flink cluster obtains the data storage rules from the section configuration table, it saves them locally. If there are corresponding data storage rules already existing locally, such as the data storage rules for the log records of buried point A, then use the latest data storage rules to update the data storage rules saved locally.
[0078] Step 102: Flink obtains the first business data from the log server, filters the first business data according to the data storage rules to obtain second business data; and stores the second business data into the database.
[0079] The present invention uses Alibaba Cloud SLS as the cache for the buried point log record data of the APP. The storage cost of SLS is not high, it can support high concurrency, and can make up for the concurrency limitation problem of ClickHouse. However, the data analysis and chart configuration of SLS require creating indexes, and the cost of indexes is relatively high. The present invention uses ClickHouse with lower cost for storage.
[0080] Flink connects to the corresponding Logstore for data consumption, retrieves the Section configuration (i.e., data storage rules) from the Section configuration table every 5 seconds, filters according to the Section fields to reduce invalid data, and then dynamically generates an inbound statement based on the fields configured in the Section and the associated tables. After splicing, the data is stored in the database. The data storage system structure is as Figure 4 shown; for the detailed steps of the Flink cluster to filter and enter the log record data, see Figure 5 the process shown.
[0081] Step 301: Filter the first service data according to the identifier of the first service data in the data storage rule to obtain the fourth service data corresponding to the identifier.
[0082] Each table for storing data in SLS is called LogStore. The Flink task uses Source to consume data from SLS to obtain data and uses Sink to store the data in ClickHouse. After the Source obtains the log record, according to the identifier of the log record, it looks up whether the corresponding identifier is configured in the locally saved data storage rule. If the corresponding identifier is not configured, the log record is discarded; if the corresponding identifier is configured, the log record data is obtained for subsequent processing. For example, the identifier corresponding to the buried point A is the product browsing log A, and this identifier is configured in the data storage rule, then the Source obtains the log record data corresponding to the buried point A, such as the obtained log record data is the log record_buried point A.
[0083] Step 302: Obtain the field data corresponding to the basic field and / or the custom field from the fourth service data according to the basic field and / or the custom field corresponding to the flag.
[0084] According to the obtained identifier of the log record (such as the product browsing log A), obtain the corresponding basic field and custom field of the identifier. For example, when obtaining the fields of the identifier product browsing log A, the following fields are obtained as shown in the following table:
[0085]
[0086] Then obtain the data content of the fields corresponding to the identifier from the obtained log record, and the obtained results are as shown in the following table:
[0087]
[0088] The log record data may contain other data in addition to the fields corresponding to the identifier. Flink will only obtain the data of the fields corresponding to the identifier, and Flink will not obtain other data.
[0089] Step 303: Create a second SQL statement based on the field data and the database table corresponding to the identifier, and execute the second SQL statement to save the field data into the database table corresponding to the identifier.
[0090] Flink obtains the corresponding ClickHouse database table according to the identifier of the log record. For example, the database table corresponding to the product browsing log A is Table_ID_001. Then, according to the fields of the log record corresponding to the identifier, the corresponding database table fields in the database table Table_ID_001 are obtained. The fields of the product browsing log A and the corresponding database table fields in the database table Table_ID_001 are shown in the following table:
[0091]
[0092] Flink constructs an SQL statement based on the data content of the log record fields, the database table corresponding to the identifier, and the corresponding database table fields. Then, the SQL statement is executed to save the data content corresponding to the fields of the log record into the corresponding database table, such as saving to the table Table_ID_001.
[0093] Step 304: Create a third SQL statement for batch storing data for the field data corresponding to multiple said flags and / or the database tables corresponding to multiple said flags, and execute the third SQL statement to save the field data corresponding to multiple said flags into the database.
[0094] Due to the lack of concurrency in ClickHouse, in the Sink of the present invention, SQL is batch-concatenated according to the fields and tables configured in Section, and SQL is executed for batch storage into the database. After the Flink cluster obtains multiple log records, the field contents that need to be saved into the database are respectively parsed from these log records. The field contents of multiple log records are constructed into 1 SQL statement, and these contents are batch-stored into the database, thereby reducing the concurrent operations on the ClickHouse database.
[0095] Step 103: Query the database to obtain the third service data and process the third service data.
[0096] The most commonly used multi-line chart for data metric analysis only requires three dimensions. By fixing the dimension fields and changing the names in the configured query statement, a general multi-line chart can be obtained.
[0097] Configure the display and alarm of data metrics. The display is an executable SQL statement stored in the Section configuration. The efficiency of the ClickHouse join query is relatively low. When defining the data, generate a single ClickHouse table from N buried points. When querying the chart of the Section, directly query from the single table, which has a relatively high query speed. For the specific query process, see Figure 6 the process shown.
[0098] Step 401: Obtain the corresponding first SQL statement according to the identifier.
[0099] On the WEB side, obtain the corresponding query SQL statement according to the log record identifier to be queried. Each identifier has a corresponding query SQL statement set during the setting stage, which is used to query the content of the corresponding field of the log record. For example, the SQL statement for querying the log record data of buried point A (identifier: commodity browsing log A) is:
[0100] select userID, APP version number, commodity detail page A, favorite commodity A from Table_ID_001;
[0101] Step 402: Modify the first SQL statement according to the query condition to obtain the fourth SQL statement.
[0102] When querying on the WEB segment, query conditions can be added to the corresponding SQL query statement. For example, specify the favorite commodity as the favorite record of Nike sneakers. Then, according to the query condition, modify the corresponding SQL query statement, add the corresponding condition in the query condition, and the modified SQL statement is, for example: select userID, APP version number, commodity detail page A, favorite commodity A fromTable_I D_001where favorite commodity A = Nike sneakers;
[0103] Step 403: Execute the fourth SQL statement to query the field data corresponding to the basic field and / or the custom field from the database.
[0104] On the WEB side, execute the SQL statement with modified query conditions to obtain the log record information corresponding to the buried point identifier (log record identifier), and then present the query result on the WEB page. For example, the query result is presented through the most commonly used multi-line chart for data metric analysis. The specific processing and display of the query data are not limited in this technical solution and can be processed accordingly according to requirements.
[0105] Through the embodiments of the present invention, the APP buried point service data is associated with the ClickHouse table, N buried point objects are associated into a structure table, and the buried point fields are corresponding to the fields of the structure table one by one; and the log server, Flink and ClickHouse are used for collaborative processing. Thus, it is ensured that the service data monitoring system has high concurrency and is easy to expand, while reducing the data storage cost.
[0106] In addition, the embodiments of the present invention also propose a service data monitoring device based on Flink, referring to Figure 7 , the service data monitoring device based on Flink includes:
[0107] A data configuration unit 10, configured to configure the data storage rule of the first service data;
[0108] A data storage unit 20, configured to enable Flink to obtain the first service data from the log server, filter the first service data according to the data storage rule to obtain second service data; and store the second service data in the database;
[0109] A data processing unit 30, configured to query the database to obtain third service data and process the third service data.
[0110] Through the above solution in this embodiment, the APP buried point service data is associated with the ClickHouse table, N buried point objects are associated into a structure table, and the buried point fields are corresponding to the fields of the structure table one by one; and the log server, Flink and ClickHouse are used for collaborative processing. Thus, it is ensured that the service data monitoring system has high concurrency and is easy to expand, while reducing the data storage cost.
[0111] It should be noted that each unit in the above device can be used to implement each step in the above method and achieve the corresponding technical effects, which will not be elaborated in this embodiment.
[0112] Referring to Figure 8 , Figure 8 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention.
[0113] As Figure 8As shown in the figure, the electronic device may include: a processor 1001, such as a CPU, a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as WI-FI, 4G, 5G interfaces). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0114] Those skilled in the art can understand that Figure 8 the structure shown in the figure does not constitute a limitation on the electronic device, and it may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0115] As Figure 8 shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a Flink-based service data monitoring program.
[0116] In Figure 8 the electronic device shown in the figure, the network interface 1004 is mainly used for data communication with an external network; the user interface 1003 is mainly used to receive input instructions from users; the electronic device calls the Flink-based service data monitoring program stored in the memory 1005 through the processor 1001 and performs the following operations:
[0117] Configure the data storage rules for the first service data;
[0118] Flink obtains the first service data from the log server, filters the first service data according to the data storage rules to obtain second service data; stores the second service data in a database;
[0119] Query the database to obtain third service data and process the third service data.
[0120] Optionally, the configuration of the data storage rules for the first service data includes the following steps:
[0121] Set the identifier and / or basic fields and / or custom fields of the first service data;
[0122] Establish the association relationship between the identifier of the first service data and the database table, and establish the association relationship between the identifiers of N (N≥1) pieces of the first service data and 1 database table;
[0123] Create the fields of the database table corresponding to the identifier according to the basic fields and / or the custom fields;
[0124] Establish the association relationship between the basic fields and / or the custom fields and the fields of the database table.
[0125] Optionally, the method further includes:
[0126] Create a first SQL statement for querying the field data corresponding to the basic fields and / or the custom fields corresponding to the identifier from the database corresponding to the identifier; the identifier and the first SQL statement are in one-to-one correspondence.
[0127] Optionally, the Flink obtains the first service data from the log server, filters the first service data according to the data storage rule to obtain the second service data; stores the second service data in the database, including the following steps:
[0128] Filter the first service data according to the identifier of the first service data in the data storage rule to obtain the fourth service data corresponding to the identifier;
[0129] Obtain the field data corresponding to the basic fields and / or the custom fields from the fourth service data according to the basic fields and / or the custom fields corresponding to the flag;
[0130] Create a second SQL statement according to the field data and the database table corresponding to the identifier, and execute the second SQL statement to save the field data into the database table corresponding to the identifier.
[0131] Optionally, the method further includes the following steps:
[0132] Create a third SQL statement for batch storing data for the field data corresponding to multiple flags and / or the database tables corresponding to multiple flags, and execute the third SQL statement to save the field data corresponding to the multiple flags into the database.
[0133] Optionally, the method further includes the following steps:
[0134] The Flink regularly obtains the data storage rule, and then updates the locally saved data storage rule with the latest data storage rule.
[0135] Optionally, querying the database to obtain third service data and processing the third service data includes the following steps:
[0136] Obtain the corresponding first SQL statement according to the identifier;
[0137] Modify the first SQL statement according to the query condition to obtain a fourth SQL statement;
[0138] Execute the fourth SQL statement to query the field data corresponding to the basic field and / or the custom field from the database.
[0139] In this embodiment, through the above solution, the APP buried point service data is associated with the ClickHouse table, N buried point objects are associated into a structure table, and the buried point fields are corresponding to the fields of the structure table one by one; and the log server, Flink and ClickHouse are used for collaborative processing. Thus, it is ensured that the service data monitoring system has high concurrency and is easy to expand, while reducing the data storage cost.
[0140] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which a service data monitoring program based on Flink is stored. When the service data monitoring program based on Flink is executed by a processor, the following operations are implemented:
[0141] Configure the data storage rule of the first service data;
[0142] Flink obtains the first service data from the log server, filters the first service data according to the data storage rule to obtain second service data; and stores the second service data in the database;
[0143] Query the database to obtain third service data and process the third service data.
[0144] Optionally, the configuration of the data storage rule of the first service data includes the following steps:
[0145] Set the identifier and / or basic field and / or custom field of the first service data;
[0146] Establish an association relationship between the identifier of the first service data and the database table, and establish an association relationship between the identifiers of N (N is greater than or equal to 1) first service data and 1 database table;
[0147] Create the fields of the database table corresponding to the identifier according to the basic field and / or the custom field;
[0148] Establish an association relationship between the basic fields and / or the custom fields and the fields of the database table.
[0149] Optionally, the method further includes:
[0150] Create a first SQL statement to query the field data corresponding to the basic fields and / or the custom fields corresponding to the identifier from the database corresponding to the identifier; the identifier and the first SQL statement are in one-to-one correspondence.
[0151] Optionally, the Flink obtains the first service data from a log server, filters the first service data according to the data storage rule to obtain second service data; stores the second service data in a database, including the following steps:
[0152] Filter the first service data according to the identifier of the first service data in the data storage rule to obtain fourth service data corresponding to the identifier;
[0153] Obtain the field data corresponding to the basic fields and / or the custom fields from the fourth service data according to the basic fields and / or the custom fields corresponding to the flag;
[0154] Create a second SQL statement according to the field data and the database table corresponding to the identifier, and execute the second SQL statement to save the field data into the database table corresponding to the identifier.
[0155] Optionally, the method further includes the following steps:
[0156] Create a third SQL statement for batch storing data for the field data corresponding to multiple identifiers and / or the database tables corresponding to multiple identifiers, and execute the third SQL statement to save the field data corresponding to the multiple identifiers into the database.
[0157] Optionally, the method further includes the following steps:
[0158] The Flink regularly obtains the data storage rule, and then updates the locally saved data storage rule with the latest data storage rule.
[0159] Optionally, querying the database to obtain third service data and processing the third service data includes the following steps:
[0160] Obtain the corresponding first SQL statement according to the identifier;
[0161] Modify the first SQL statement according to the query condition to obtain a fourth SQL statement;
[0162] Execute the fourth SQL statement to query the field data corresponding to the basic field and / or the custom field from the database.
[0163] In this embodiment, through the above solution, the APP buried point service data is associated with the ClickHouse table, and N buried point objects are associated into a structure table, and the buried point fields are corresponding to the fields of the structure table one by one; and the log server, Flink, and ClickHouse are coordinated for processing. Thus, it is ensured that the business data monitoring system has high concurrency and is easy to expand, while reducing the data storage cost.
[0164] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or system including that element.
[0165] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0166] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions for causing a terminal device (which may be a mobile phone, a computer, a server, a controller, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0167] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A business data monitoring method based on Flink, characterized in that The method includes the following steps: Configure the data storage rule for the first service data; Flink obtains the first service data from the log server, filters the first service data according to the data storage rule to obtain the second service data, and stores the second service data in the database; Query the database to obtain the third service data and process the third service data; Among them, configuring the data storage rule for the first service data includes the following steps: Set the identifier, basic fields, and custom fields of the first service data; Establish the association relationship between the identifier of the first service data and the database table, and establish the association relationship between the identifiers of N pieces of the first service data and 1 database table; where N is greater than or equal to 1; Create the fields of the database table corresponding to the identifier according to the basic fields and the custom fields; Establish the association relationship between the basic fields and the custom fields and the fields of the database table.
2. The method according to claim 1, wherein The method further includes: Create a first SQL statement for querying the field data corresponding to the basic fields and the custom fields corresponding to the identifier from the database corresponding to the identifier; the identifier and the first SQL statement are in one-to-one correspondence.
3. The method according to claim 1, wherein The process that Flink obtains the first service data from the log server, filters the first service data according to the data storage rule to obtain the second service data, and stores the second service data in the database includes the following steps: Filter the first service data according to the identifier of the first service data in the data storage rule to obtain the fourth service data corresponding to the identifier; Obtain the field data corresponding to the basic fields and the custom fields from the fourth service data according to the basic fields and the custom fields corresponding to the identifier; Create a second SQL statement according to the field data and the database table corresponding to the identifier, and execute the second SQL statement to save the field data into the database table corresponding to the identifier.
4. The method according to claim 3, characterized in that The method further includes the following steps: Create a third SQL statement for batch storing data for the field data corresponding to multiple identifiers and the database tables corresponding to multiple identifiers, and execute the third SQL statement to save the field data corresponding to multiple identifiers into the database.
5. The method according to claim 1, wherein The method further includes the following steps: Flink regularly obtains the data storage rule, and then updates the locally saved data storage rule with the latest data storage rule.
6. The method according to claim 2, characterized in that, The process of querying the database to obtain the third service data and processing the third service data includes the following steps: Obtain the corresponding first SQL statement according to the identifier; Modify the first SQL statement according to the query condition to obtain a fourth SQL statement; Execute the fourth SQL statement to query the field data corresponding to the basic fields and the custom fields from the database.
7. A business data monitoring device based on Flink, characterized in that, The device includes: A data configuration unit for configuring the data storage rule for the first service data; A data storage unit, configured to enable Flink to obtain the first service data from a log server, filter the first service data according to the data storage rule to obtain second service data, and store the second service data in a database; A data processing unit, configured to query the database to obtain third service data and process the third service data; Among them, the data configuration unit is further configured to: Set the identifier, basic fields, and custom fields of the first service data; Establish an association relationship between the identifier of the first service data and a database table, and establish an association relationship between the identifiers of N pieces of the first service data and 1 database table; where N is greater than or equal to 1; Create fields of the database table corresponding to the identifier according to the basic fields and the custom fields; Establish an association relationship between the basic fields and the custom fields and the fields of the database table.
8. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a Flink-based service data monitoring program stored on the memory and executable on the processor. The Flink-based service data monitoring program is configured to implement the steps of the Flink-based service data monitoring method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the Flink-based service data monitoring method according to any one of claims 1 to 6.
Citation Information
Patent Citations
User behavior analysis system and method, storage medium and computing equipment
CN111488261A