A business data analysis method, a data processing method, a data analysis system, and a storage medium
By using Kafka and Flink technologies in online education products and using business identifiers for data partition storage and processing, the accuracy of business indicator analysis and real-time data timing problems in online education products are solved, and efficient and accurate business data analysis is achieved.
Patent Information
- Application Number
- CN202110500884.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-05-08
AI Technical Summary
It is difficult for the prior art to accurately analyze business indicators, such as teaching conversion rate and renewal rate, in online education products, and the timing is disordered in real-time data processing, which affects the accuracy of the analysis results.
The Kafka distributed storage system and Flink streaming processing technology are used to partition data storage using service identifiers as primary keys to ensure data timing, and execute processing logic through Flink, and output it to MySQL or Doris database for verification.
It realizes accurate analysis of online education product business indicators, ensures the timing and accuracy of real-time data processing, and improves the efficiency and reliability of data processing.
Smart Images

Figure CN113220682B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to data warehouse technology, and in particular to an order data analysis system, method, and storage medium. Background Art
[0002] With the rapid development of Internet online services, for online product design, interface optimization, and precision marketing to improve the user experience, it is necessary to output certain metric data based on the requirements of the business side for use as a decision-making reference for business adjustment.
[0003] For example, in the business scenario of online education products, it is necessary to use the feedback of users on business products and user behavior as inputs for model calculation to obtain relevant data metrics. For example, the teaching conversion rate and renewal rate of a certain teacher, the renewal rate of a certain course product, etc. Usually, starting from user orders, some dimensional data of teachers or students is associated from the order details, and then corresponding metrics for guiding the business are obtained through statistical analysis. Summary of the Invention
[0004] The present application provides a business data analysis method. The method includes:
[0005] Obtaining a target information query request;
[0006] Obtaining the processing logic of business data according to the target information query request, and the storage area of the business data related to the processing logic in the distributed storage system;
[0007] Respectively obtaining all the business data in the storage area, where the business data has a business identifier;
[0008] Processing the business data according to the processing logic and the business identifier of the business data to obtain the target information.
[0009] The above method further includes: storing the business data with the same business identifier in the same storage area in the distributed storage system according to a preset policy.
[0010] The above method further includes respectively obtaining historical business data and real-time business data. Specifically, obtaining real-time business data and sending it to the distributed storage system for storage; and before obtaining the real-time business data, obtaining the historical business data stored in the data warehouse and sending it to the distributed storage system.
[0011] In the above method, before the processing logic processes the business data according to the business identifier, it further includes: verifying whether the processing logic is correct by using the business identifier in the processing logic.
[0012] One implementation of the above method is to use Flink to execute the processing logic, process the service data according to the service identifier; and write the processing results with the same service identifier into the same storage area of the distributed storage system.
[0013] In the above method, after writing the processing results with the same service identifier into the same storage area of the distributed storage system, it further includes: the distributed storage system outputs the operation result to a MySQL or Doris database; the MySQL or Doris database displays the operation result as an output unit and outputs the queried target information.
[0014] A service data processing method includes: obtaining service data and obtaining the service identifier of the service data; writing the service data with the same service identifier into the same storage area in the distributed storage system.
[0015] Among them, the distributed storage system is Kafka.
[0016] In the above method, the method for obtaining service data includes: obtaining real-time service data by using Canal;
[0017] Or / and, using Flink to perform operations on the first-layer service data stored in Kafka to obtain second-layer service data; and writing the second-layer service data with the same service identifier into the same storage area in Kafka.
[0018] In the above method, obtaining the service identifier of the service data and writing the service data with the same service identifier into the same storage area in Kafka includes: obtaining the service identifier from the obtained service data packet according to the definition of the service identifier in the preset mapping table; performing partition operation on the service identifier according to the preset policy, and writing the service data into the corresponding storage area in Kafka according to the partition operation result.
[0019] The above method further includes: when the preset conditions are met, when there are two or more service data with the same service identifier stored, retaining the latest obtained service data and deleting the rest of the service data.
[0020] A service data analysis system includes:
[0021] A receiving unit that obtains a target information query request;
[0022] A distributed storage system for storing service data, and the service data with the same service identifier is stored in the same storage area in the distributed storage system;
[0023] The processing unit obtains the processing logic of business data according to the target information query request, and the storage area of the business data related to the processing logic in the distributed storage system; obtains all the business data in the storage area, and executes the processing logic to process the business data according to the business identifier to obtain the target information.
[0024] Among them, the distributed storage system is Kafka;
[0025] And it further includes: an incremental data subscription module for synchronizing the real-time business data of users to Kafka; Kafka stores the business data with the same business identifier in the same storage area.
[0026] Among them, the processing unit is Flink;
[0027] Flink obtains the first-layer business data stored in Kafka according to the target information query request and executes the processing logic to obtain the second business data, and writes the second-layer business data with the same business identifier into the same storage area in Kafka.
[0028] The system further includes: a historical business database for storing historical business data;
[0029] After obtaining the historical business data, Kafka obtains the real-time business data synchronized by the incremental data subscription module.
[0030] In the above system, there are multiple Flinks; the multiple Flinks are connected in series for performing hierarchical calculations.
[0031] The above system further includes: a MySQL or Doris database; Kafka writes the Flink operation results into the MySQL or Doris database.
[0032] The above system further includes a verification unit for verifying whether the processing logic is correct by using the business identifier in the processing logic.
[0033] In the above data analysis method, data processing method and system, the business identifier includes, but is not limited to, one of the following identifiers or a combination of two or more identifiers: order number, class ID, teacher ID, user ID, semester ID, course project ID.
[0034] A non-transitory machine-readable storage medium stores executable code thereon, and when the executable code is executed by a processor of a computing device, it causes the processor to execute the method as described above.
[0035] Since the business status changes in real time, in order to obtain business metrics, the present invention adopts a method of stream data processing, using real-time business data as input to obtain the business metrics required by the business. In particular, the embodiments of the present invention calculate and output the business metrics based on the mechanism of real-time data stream processing of Kafka and Flink.
[0036] By tracking the changes in the business status, the analysis of the business side is realized, and metrics such as the business conversion rate and the renewal rate facing the sales end are obtained. In reality, for a business, its status may change. For example, for an order, behaviors such as user placing an order, withdrawing an order, and post-order payment will all cause changes in the order data status, and the change in the order status will directly affect the subsequent analysis of related businesses.
[0037] Therefore, the present invention provides a business data analysis method and a business data processing method.
[0038] First, for the need of business analysis, the present invention obtains the full-volume historical business data, and then obtains the business status change data by listening to the business incremental data. Thus, the data requirements for business analysis are met in terms of data acquisition methods.
[0039] Secondly, the business data analysis method and system provided by the present invention adopt a Kafka distributed storage system and a Flink streaming processing distributed system to achieve timely and efficient processing and analysis of real-time data.
[0040] Finally, based on the characteristics of the distributed structure of the present invention, in order to obtain accurate changes in the business data status and ensure that the timeliness of the business data stored in the distributed message system does not get confused, the business identifier of the business data is used as the primary key of the Kafka storage table, where each business data has a unique business identifier. Further, the primary key is used as a variable of the allocation strategy for partition calculation to determine the storage area to which the business data should be allocated. Thus, business data with the same primary key can enter the same storage area of the Kafka distributed system, avoiding the situation where the same order data is stored in different storage areas and the timeliness gets confused, and ensuring that the order of processing the order data is consistent with the order of the user's operations on the order.
[0041] In summary, using the link structure and data analysis method of the present invention, the analysis of the order business is more accurate and reliable.
[0042] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. Description of the Drawings
[0043] The above and other objects, features, and advantages of the present application will become more apparent from the following detailed description of exemplary embodiments of the present application in conjunction with the accompanying drawings, wherein in the exemplary embodiments of the present application, the same reference numerals generally represent the same components.
[0044] Figure 1 It is a schematic diagram of a link structure shown in an embodiment of the present invention;
[0045] Figure 2 It is a schematic diagram of the temporal deviation of real-time data generated during Flink calculation in the prior art. Detailed Embodiments
[0046] Preferred embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.
[0047] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0048] It should be understood that although the terms "first", "second", "third", etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, the meaning of "a plurality" is two or more unless otherwise specifically defined.
[0049] The object of the present invention is to obtain certain business indicators of one or more online products.
[0050] The following takes an online education product as an example to illustrate the embodiments of the present invention.
[0051] Taking the online education service industry as an example, the ordering of online courses is to a certain extent highly relevant to aspects such as the teaching ability of service providers and the continuity of course design. Therefore, it is necessary to reflect issues such as the rationality of relevant course design through the analysis of course order status and the changes in order status, so as to guide the improvement of business. The required analysis indicators include, for example, business conversion rate, course renewal rate, etc.
[0052] In reality, the business status of online products will change, thus sequentially generating a series of business data. For example, in the online course order business, when the initial order data is created, its status often changes. For example, user ordering, order withdrawal, post-order payment and other behaviors can all lead to changes in the order data status, and the changes in the order data status will directly affect the subsequent analysis of relevant businesses. In addition to order data, it also includes user data, teacher data, class data, semester data, etc.
[0053] The present invention provides a business data analysis method, a business data processing method, and a business data analysis system. Using the present invention can better meet the requirements of the order business scenario.
[0054] This application provides an order data analysis method. It includes:
[0055] Step 11: Obtain business requirements, that is, target information query requests, such as query requests for certain business indicators. The business indicators can be the current order quantity of a certain teacher, the course renewal rate of the teacher, etc.;
[0056] Step 12: Obtain the processing logic of business data according to the target information query request, and the storage area of the business data related to the processing logic in the distributed storage system;
[0057] According to the target information query request, obtain the corresponding business processing logic. The processing logic records the logical calculation method and the data source for implementing the logical calculation. Furthermore, through the data source described in the processing logic, the storage address of the data required to execute the processing logic can be obtained.
[0058] Step 13: Respectively obtain all the business data in the storage area, and the business data has a business identifier.
[0059] The business data required to implement the processing logic may be stored in multiple areas of the distributed storage system. Respectively obtain all the business data in these storage areas.
[0060] In an embodiment of the present invention, business data is stored in a distributed storage system; and, according to a predefined rule, each piece of business data has a unique business identifier; and, according to a preset policy, by performing operations with the business identifier as a variable, the storage area in the distributed storage system to which the relevant business data is assigned can be determined. According to this method, business data with the same business identifier is stored in the same storage area in the distributed storage system.
[0061] Step 14: Process the business data according to the processing logic and the business identifier of the business data to obtain target information.
[0062] Use a distributed logic processing engine to implement the processing logic, and perform logic processing on all the data obtained for each storage area. The business identifier of the business data is used in the processing logic, and the business data is processed according to the business identifier. For example, based on the business identifier, it is determined whether the business data corresponds to the same order, so as to screen the data of the same order, analyze the change of the order status, etc.
[0063] In the above method, the business data stored in the distributed storage system includes real-time business data and historical business data.
[0064] For example, use Canal to obtain real-time business data from a MySQL database, and send the real-time business data to the distributed storage system for storage. The real-time business data can be incremental data of an order. For example, an order is generated, an order is cancelled, an order is paid, etc. based on user behavior, or the personal information of a teacher changes, the teacher's position changes, or the class time of a course changes, the duration of a course changes, etc. will all generate the real-time business data;
[0065] Before obtaining the real-time business data, the historical business data stored in the data warehouse can be imported first and sent to the distributed storage system.
[0066] The main purpose of Canal is to parse the incremental log based on the MySQL database and provide incremental data subscription and consumption.
[0067] As a preferred implementation, the distributed storage system can be Kafka.
[0068] Kafka is a distributed publish-subscribe messaging system, a high-throughput distributed publish-subscribe messaging system that processes all action stream data of consumers in a website.
[0069] In the above method, Flink can be used to execute the processing logic.
[0070] Flink is a distributed processing engine for distributed data stream processing and batch data processing, providing functions to support both stream processing and batch processing applications. In the Flink framework, all tasks are treated as streams, thus achieving real-time stream processing with lower latency.
[0071] Use Flink to perform logical operations on the business data stored in the distributed storage system; for the business data obtained from the operation results of Flink, by defining the business identifier of the business data, use this business identifier as a variable of the preset policy to determine the storage area of the business data in Kafka. By adopting this method, the operation results with the same business identifier can be written into the same storage area of the distributed storage system.
[0072] In a preferred embodiment, before Flink executes the processing logic operation, it is also possible to verify whether the processing logic is correct.
[0073] For the determined processing logic, pre-define the business identifier that should be used to save the business data obtained from the processing logic to Kafka, so as to ensure that after the processing logic operation is completed, the obtained business data can be correctly stored to meet the business requirements.
[0074] Before Flink executes the processing logic operation, if it is determined that the business identifier for storing the logic processing result data in Kafka is used in the processing logic, it indicates that the business logic is correct; if the business identifier based on which Kafka stores the logic processing result data does not contain the key used in the calculation logic, it indicates that the processing logic is incorrect.
[0075] The method further includes: the distributed storage system outputs the operation result to a MySQL or Doris database. The MySQL or Doris database is used as an output unit to display the operation result, so as to respond to business requirements and output the queried target information. When the calculation result is an intermediate result of multi-layer Flink calculations, the MySQL or Doris database can be used as a tool to verify the Flink calculation result, such as comparing it with the previous layer of Flink calculation result, or outputting it for manual verification.
[0076] An embodiment of the present invention further provides a method for processing business data, including:
[0077] Obtain business data,
[0078] Obtain the business identifier of the business data;
[0079] Write business data with the same business identifier into the same storage area in the distributed storage system.
[0080] As an implementation, the distributed storage system is Kafka. However, the present invention does not limit the application of other distributed storage systems to the present invention.
[0081] In this embodiment, the ways to obtain service data include at least the following two:
[0082] Obtain real-time service data by using Canal; for example, obtain real-time service data by listening to the MySQL database through Canal;
[0083] Use Flink to perform arithmetic operations on the service data (hereinafter referred to as the first-layer service data) stored in Kafka. Based on the operations of Flink, such as join, group by, etc., the obtained service data (hereinafter referred to as the second-layer service data); that is, the second-layer service data is obtained from the first-layer service data through relevant operations. Furthermore, according to the service identifier of the second-layer service data, the storage area of the second-layer service data is determined by performing partition calculation according to a preset policy, and Kafka saves the second-layer service data according to the partition calculation result. Through this method, the second-layer service data with the same service identifier can be written into the same storage area in Kafka.
[0084] The above descriptions of the first layer and the second layer indicate the relationship between the two layers of data, and do not mean that the Flink operation can only perform one operation. By connecting the Flink operation modules in series, multiple layers of calculations can be performed. For example, on the basis of the above embodiment, Flink continues to use the second-layer service data to obtain the third-layer service data through the operation of the processing logic, and stores the third-layer service data with the same service identifier in the same storage area of Kafka.
[0085] In the embodiment of this method, the method of writing service data with the same service identifier into the same storage area in Kafka includes: a mapping table is preset, and the definition of the service identifier is stored in the mapping table. According to this definition, the service identifier can be obtained from the obtained service data packet. For example, a certain field in the service data packet, or a combination of several fields is defined as the service identifier;
[0086] Let the service identifier participate in the operation of the preset policy, that is, the partition operation. After the partition operation, it is obtained which storage area in Kafka the service data should be assigned to. Thus, according to the partition operation result, the service data is written into the corresponding storage area in Kafka.
[0087] The service data processing method in the embodiment of the present invention further includes compressing the data in the storage area. The data compression is based on the service identifier. Specifically, when the preset conditions are met, when two or more service data with the same service identifier are stored, the latest obtained service data is retained, and the rest of the service data is deleted.
[0088] Correspondingly, an embodiment of the present invention further provides a business data analysis system, including:
[0089] A receiving unit that acquires a target information query request;
[0090] A distributed storage system for storing business data, and business data with the same business identifier is stored in the same storage area in the distributed storage system;
[0091] A processing unit that obtains the processing logic of business data according to the target information query request, and the storage area of the business data related to the processing logic in the distributed storage system; obtains all the business data in the storage area, and executes the processing logic to process the business data according to the business identifier to obtain the target information.
[0092] Among them, preferably, the distributed storage system is Kafka;
[0093] And it further includes: an incremental data subscription module for synchronizing the real-time business data of users to Kafka; Kafka stores business data with the same business identifier in the same storage area.
[0094] Among them, preferably, the processing unit is Flink; Flink obtains the first-layer business data stored in Kafka according to the target information query request and executes the processing logic to obtain the second business data, and writes the second-layer business data with the same business identifier into the same storage area of Kafka.
[0095] The system further includes: a historical business database for storing historical business data; after Kafka obtains the historical business data, it obtains the real-time business data synchronized by the incremental data subscription module.
[0096] In the system embodiment of the present invention, there may be multiple Flinks, and the multiple Flinks are connected in series for performing hierarchical calculations.
[0097] The above system further includes: a MySQL or Doris database; Kafka writes the Flink operation result into the MySQL or Doris database.
[0098] The above system further includes: a verification unit for verifying whether the processing logic is correct by using the business identifier in the processing logic.
[0099] In the above data analysis method, data processing method, and system, the service identifier includes, but is not limited to, one of the following identifiers or a combination of two or more identifiers: order number, class ID, teacher ID, user ID, semester ID, and course project ID.
[0100] The following specifically describes the structure of the service data system of the present invention and the execution process of related methods with reference to the accompanying drawings.
[0101] Figure 1 A service data analysis system provided by the present invention is shown. Historical data in the Hive data warehouse is imported into Kafka; and, the real-time data of MySQL is synchronized through Canal; Doris and MySQL are used in the middle layer for data verification; Kafka is used as the traffic data middle layer, and Flink is used to concatenate and implement the hierarchical calculation logic, and finally the results are output to MySQL or Doris to provide service analysis services.
[0102] As Figure 1 shown, the service data analysis system includes: an incremental data subscription module, a Hive data warehouse, a Kafka distributed publish-subscribe message system, and a Flink calculation module.
[0103] The incremental data subscription module synchronizes real-time data from a data source database (such as a MySQL database) using the incremental data subscription module.
[0104] The incremental data subscription module can be implemented by Canal. The main purpose of Canal is to parse the incremental log of the MySQL database to provide incremental data subscription and consumption. In this embodiment, Canal is used to obtain the incremental data of business data from MySQL in real time. For order data, the incremental data includes newly added order data and data for operations on existing orders.
[0105] Specifically, Canal simulates the interaction protocol of a MySQL slave to send a dump protocol to the MySQL Master. The MySQL Master receives the dump request sent by Canal, starts to push the binary log to Canal, and then Canal parses the binary log and sends it to the storage destination, such as the Kafka distributed publish-subscribe message system adopted by the present invention.
[0106] The Hive data warehouse stores the historical data of user orders. The Hive is a data warehouse tool based on Hadoop, used for data extraction, transformation, and loading.
[0107] The Kafka distributed publish-subscribe messaging system first obtains the historical business data from the Hive data warehouse at one time. After obtaining the historical data of user orders, the Kafka distributed publish-subscribe messaging system monitors the incremental data through an incremental data subscription module (such as Canal) and synchronizes the incremental data in real time.
[0108] As described above, Canal is used to collect the log data of the MySQL database and perform real-time incremental data synchronization. As a message provider for Kafka, Canal passes the insert, update, and delete operations on the data in the MySQL database to Kafka in the form of logs.
[0109] The Kafka distributed publish-subscribe messaging system has multiple storage areas, and Kafka stores the obtained business data in different storage areas.
[0110] In the embodiment of the present invention, Kafka maintains a mapping table, which stores the business identification information as the primary key, records which type of business data uses which field or fields as the primary key, so that after Kafka obtains the business data, the content used as the primary key can be obtained from the business data. The corresponding preset allocation strategy uses the primary key as a variable to calculate the storage area to which the business data is allocated. Thus, the business data with the same business identification is sent to the same storage area in Kafka for storage.
[0111] Since Kafka is a distributed storage system, the storage areas exist in the form of clusters. Therefore, the mapping table also stores the storage area description information, through which the corresponding storage area in Kafka can be linked to realize data writing and reading.
[0112] The present invention does not limit the business identification as the primary key. The business identification may include, for example, one of the following identifications or a combination of two or more identifications: order number, class ID (class number, course progress number, or class and course progress number), teacher ID (teacher identity identification), user ID (user identity identification), semester ID (or academic quarter number, used to indicate the course stage), course project ID (indicating the course type, such as Chinese, thinking), etc.
[0113] In the following embodiments, the business identification is the order number, and the order number is used as the primary key in the Kafka table. After obtaining the order incremental data of the MySQL database, the business data with different order numbers is distributed and stored in different areas according to the order strategy. The present invention does not limit the use of other fields or combinations of fields, as long as the field or combination of fields can be obtained from the MySQL database and has a one-to-one correspondence with the order.
[0114] The present invention does not limit the setting of the specific distribution strategy, as long as the order data with the same order number is distributed to the same storage area.
[0115] Specifically, assume that the order data structure in the MySQL database is id and value, and the id can be the order number. Canal monitors the insert operations of data (1, a) and (2, b) in the MySQL database. The primary key set in the table maintained in the Kafka system is the above-mentioned id. Thus, taking the id as a variable, according to the corresponding strategy, it is determined which storage area the data related to the id is distributed to. The data with the same id will be distributed to the same storage area. Taking the insert operations of data (1, a) and (2, b) as an example, (1, a) will be written into Kafka partition 1, and (2, b) will be written into Kafka partition 2.
[0116] The Flink computing module performs preset statistical analysis on the incremental data obtained from each distributed storage area in the Kafka distributed publish-subscribe message system.
[0117] Specifically, Canal transmits the data change log of the database MySQL to Kafka, and Kafka stores the order incremental data in partitions according to the primary key strategy. Thus, according to the change of the order status by the user, the incremental data of the same order is stored in the same Kafka storage area in the order of order change. According to the order of order data change in the insert and delete operations of the order data, the Flink stream computing engine can calculate the state data of the operator in each computing operator according to the order of order data change.
[0118] In the prior art, Kafka can only ensure the order of data in a single partition. However, if multiple data with order are distributed to different storage areas, Kafka cannot ensure that the order among the data with order in different storage areas is not disrupted. Similarly, in the prior art, the default way for Flink to write to Kafka is that one partition corresponds to one parallelism or multiple parallelisms. In this way, the insert, update, delete, and query records of a piece of data may fall into different storage areas, and finally, due to the disruption of the data order, the result of Flink's stream data processing will be incorrect.
[0119] For example, historical data is written into Kafka through Hive. When inputting stream data, one parallelism will only write to one storage area, and there will be change records in the stream data that modify the historical data. When the method of the Kafka primary key in the embodiment of the present invention is not adopted, the change records of the same historical data may be distributed to different storage areas, which may result in the order of the obtained change records being inconsistent with the order in which the user actually changes the order status during processing.
[0120] Therefore, the present invention determines the storage area to which the real-time data should be allocated by defining a primary key. Specifically for Flink, a primary key is set for the mapping table that Flink maps to Kafka. When Flink writes the processed stream data into Kafka, according to the data allocation strategy formulated based on the primary key, Kafka stores the processed data into the corresponding storage area. The primary key has a one-to-one correspondence with a piece of data (such as an order data), for example, the order number is used as the primary key. The data allocation strategy based on the primary key ensures that the change records (status change data) of each piece of data (order data) can be stored in the same storage area. Furthermore, in the same storage area, the orderliness between the change data is ensured. It is ensured that each record obtained by real-time calculation is the latest and correct value, and at the same time, the horizontal expansion ability of the data link is ensured.
[0121] Similar to Kafka, Flink is a distributed system, and the change records of the same piece of data need to be processed within one parallelism. When performing group by or join, the data will be hashed and partitioned according to the key of group by or join, which will cause the problem of out-of-order data in some scenarios.
[0122] Among them, the function of the group by is to group the data according to specific fields or several specific fields according to by. The join is used to associate data based on certain conditions, for example, to associate Table A and Table B, expressed as A inner join B.
[0123] See Figure 2 . Assume that the id is the order number. The three real-time data in Processor 1 (Operator1) in the figure are, in sequence: 1) Insert operation, the order data has an id number of 1 and a value of A; 2) Delete operation, the order data has an id number of 1 and a value of A; 3) Insert operation, the order data has an id number of 1 and a value of B. In the above process, the value (Value) of the order data in Processor 1 should change from A to B.
[0124] The data of Processor 1 decides which storage area the data is allocated to based on the hash algorithm (HASH). When the order data id (assuming the id is the order number) is not used as the primary key, the same order data will be allocated to different storage areas 1 and 2. If the operator operates on the order data by group by id (id is the order number), there will be Figure 2The situation shown on the right. As shown in the figure, due to the deviation in the processing capacity of each operator and the amount of data, it is possible that the data with the value of B in storage partition 2 (partition2) is processed first, that is, the order data with order number 1 and value B is inserted first; while the data with the value of A in storage partition 1 (Parttition1) is finally deleted (delete). In this way, errors will occur in the data at the final output.
[0125] For example, for the process of a user generating an order 1 and making a payment. Assume that value A represents the user's order placement operation, and value B represents the user's order payment operation. The correct real-time data timing sequence is: insert the order placement data for order number 1, delete the order placement data for order number 1, and insert the payment data for order number 1. However, Figure 2 The changed timing between the order data shown will lead to errors in this process.
[0126] Therefore, when Flink performs real-time calculation of data change logs (cdc data streams) and involves operators such as groupby or join that require repartitioning of data, the primary key described in the above embodiments is used for the calculation object to ensure that the timing of the data remains unchanged at the output.
[0127] When Flink writes the data processing and analysis results to Kafka, according to the primary key information maintained in the Kafka mapping table, the fields serving as the primary key in the data are calculated according to the preset strategy, and the data with different field values is written into different storage areas in Kafka.
[0128] Based on the use of the primary key (such as the order number) described in the embodiments of the present invention, in the link structure of the present invention, Kafka is used as the traffic data intermediate layer, and Flink is used in series to implement the hierarchical calculation logic to obtain the hierarchical data metrics for orders. Kafka outputs the calculation results of each layer of stream data to the MySQL or Doris database to provide business services.
[0129] The following is a specific example of the two-layer calculation of Flink in the present invention.
[0130] In the stage of obtaining real-time data, assume that user A places an order to purchase a course from teacher B, and the order number is X; for the real-time data of the order, Kafka allocates the storage area of the order real-time data with the order number as the primary key.
[0131] Storing and allocating according to the order number as the primary key enables the real-time data generated at different times with order number X to be stored in the same storage area of Kafka when the status of order X changes.
[0132] Suppose it is necessary to count which courses Teacher B is responsible for. Then, two - layer calculations need to be performed with the help of Flink as follows.
[0133] According to the above - mentioned business requirements, the processing logic of the first layer is to associate order data with teacher data. In this preset processing logic, the storage areas of order data and teacher data used to implement logical calculations in Kafka are described. Specifically, the Kafka mapping table for describing the storage address of order data and the Kafka mapping table for describing the storage address of teacher data can be respectively referenced in the processing logic.
[0134] For example, the processing logic is X inner join Y, and in this processing logic, the mapping table X for describing the storage address of order data and the mapping table Y for describing the storage address of teacher data are referenced.
[0135] The mapping table records the primary key information, indicating which field or fields of which type of business data are used as the primary key, and the content used as the primary key can be obtained based on the business data. In addition, the mapping table describes how to find a specific storage area in the cluster of the storage area in Kafka, so that the corresponding storage area can be linked through the mapping table to achieve the access of business data.
[0136] For example, the mapping table X records the method for obtaining the order number used as the primary key and the description information of the storage area.
[0137] After obtaining the order data and teacher data, Flink performs the join operation of the first - layer processing logic. The data in the storage area storing order data in Kafka is joined with the data in the storage area storing teacher data in Kafka. Through the join operation, the order data is associated with the teacher data.
[0138] The storage area of the corresponding data is allocated with the string composed of the order number and the teacher ID as the primary key. When Flink completes the join operation and writes the data binding the order and the teacher to Kafka, the string composed of the order number and the teacher ID is used as the variable of the corresponding allocation strategy for the primary key, and the storage area allocated to the associated data is obtained. Thus, the data with the same string composed of the order number and the teacher ID is written into the same Kafka storage area.
[0139] Based on the result of the first - layer processing, the processing logic of the second layer is to perform an aggregation operation with the teacher ID as the condition to obtain the number of orders each teacher is responsible for.
[0140] After obtaining the data of the order bound to the teacher, Flink performs the aggregation operation (group by operation) of the second-layer processing logic. When performing the aggregation calculation group by, using the teacher ID as the condition, the number of orders responsible for each teacher is calculated, including the number of orders of teacher B.
[0141] When writing the result of the aggregation operation to Kafka, using the teacher ID as the primary key, according to the corresponding policy, the storage area for data storage is determined based on the teacher ID.
[0142] Based on the above analysis structure, if the orders responsible for a certain teacher change, for example, the order course is transferred from teacher B to teacher C, then through the above two-layer calculation, the data with the changed number of orders of teacher B and the existing data of teacher B's orders are written to the same storage area in Kafka, and the data with the changed number of orders of teacher C and the existing data of teacher C's orders are written to the same storage area in Kafka.
[0143] To achieve the above partition storage of real-time data, in the embodiment of the present invention, Kafka maintains a mapping table. The primary key information is recorded in the mapping table, recording which data uses which field or fields as the primary key, and obtaining the content used as the primary key from the real-time data. For example, in this embodiment, obtaining real-time order data and obtaining the order number from the order data; when obtaining the result of the first-layer operation of Flink, obtaining the combined string of the order number and the teacher ID from the data; when obtaining the result of the second-layer operation of Flink, obtaining the teacher ID from the data. Thus, using the allocation policy, the primary key is calculated to determine the storage area allocated to the relevant data.
[0144] Based on the above embodiments, before Flink executes the processing logic operation, the present invention can also check the correctness of the processing logic.
[0145] For example, specifically in the above embodiments, before Flink executes the calculation of each layer of processing logic, it can be verified whether the processing logic of this layer is correct.
[0146] Specifically, for the determined processing logic, a Kafka mapping table is predefined, and the mapping table is used to describe which business identifier the business data of the operation result of the processing logic should be allocated to the storage area according to. Before Flink executes the processing logic operation, if the operation key of the processing logic uses the business identifier used as the primary key in the mapping table, that is, the primary key described in the mapping table contains the key in the processing logic calculation, then the processing logic is correct.
[0147] For example, the second-layer processing logic in the above embodiments obtains the number of orders each teacher is responsible for based on the teacher ID. Specifically, an aggregation operation is performed with the teacher ID as the condition, and the key for this logical operation is the teacher ID. Moreover, for the result data of the second-layer operation, the teacher ID is used as the primary key of the mapping table to allocate the storage area for the operation results. Therefore, it can be determined that this operation logic is correct. If the primary key defined in the mapping table is only the teacher ID, and the teacher ID is not used in the second-layer processing logic, it indicates that the processing logic is incorrect.
[0148] For example, in the X inner join Y processing logic of the above embodiments, if the order number and the teacher ID are specifically used as operation keys in the logical operation; in the mapping table, the string combination of the order number and the teacher ID is used as the primary key. Therefore, it can be determined that this processing logic is correct.
[0149] In the above embodiments, the business identifier is used as the primary key to determine the allocation of the data storage area in the Kafka system, ensuring the timeliness of real-time business data in the Kafka system. Based on the link structure of the present invention, the business identifier used as the primary key is also used for data compression in the Kafka system.
[0150] As a message system, Kafka retaining all the changed data will result in a large number of unnecessary change records in Kafka, occupying a large amount of useless storage. It will also cause downstream other tasks to read too much unnecessary data when starting from scratch, slowing down the reading performance and generating a large amount of unnecessary overhead. It is necessary to adopt relevant mechanisms to clean up the data in Kafka.
[0151] In the prior art, a data expiration time is set in the Kafka system to discard the data in the Kafka system whose storage duration is greater than the expiration time.
[0152] In the embodiments of the present invention, the primary key of the Kafka mapping table is used to retain only the latest data among the data with the same primary key.
[0153] Since in the present invention, each time Kafka obtains a real-time data, according to the primary key of this real-time data, this real-time data is allocated to the storage partition storing the real-time data with the same primary key. Since the real-time data with the same primary key enters the same storage partition, and the data stored in the same storage partition of Kafka has timeliness. Therefore, in each partition of Kafka, the latest real-time data among the multiple real-time data with the same primary key can be obtained.
[0154] According to the above method, when compressing the data in Kafka, in the present invention, when there are multiple real-time data with the same primary key in the storage partition, only the latest real-time data is retained.
[0155] When the existing method is used in order-related services, if the order status does not change for a long time, that is, no real-time order data is generated for a long time, then using the existing method will delete all data of the order from the Kafka library, resulting in the loss of order data. On the other hand, for orders that have had multiple status changes recently, multiple real-time data are retained for one order, resulting in data redundancy. The method of the embodiment of the present invention avoids the above situations, has a high compression rate while avoiding data loss, and is more in line with the requirements of business analysis based on changes in order status.
[0156] For example, assume there are 1,000 orders. Since the status of these 1,000 orders may change, there may be a lot of real-time data in the Kafka system, such as 3,500 pieces of real-time data.
[0157] According to the method of the present invention, after data compression, the latest 1,000 real-time data of these 1,000 orders will be saved, that is, the latest order status of these 1,000 orders will be saved. For each order, usually only the latest real-time data is required for business analysis. Therefore, the method of the embodiment of the present invention well meets the requirements of business analysis.
[0158] Assume that according to the prior art, 1,000 real-time data are still retained after data compression, but these 1,000 real-time data may be generated by 800 order status changes. Therefore, the existing compression method has the situation of order loss, and for 800 orders, 1,000 real-time data are saved, and there is still data redundancy after compression. Thus, the existing method does not meet the requirements of order business analysis.
[0159] As a further improvement to the compression method, when the present invention performs data compression, the occupancy ratio of the storage space is used as a variable to determine whether to perform data compression. For example, when the data in the storage area occupies 60% of the storage space, the compression process of the above method is performed. Or other judgment methods for whether to perform data compression are obtained through strategies based on the space occupancy rate as a variable, which is not limited by the present invention. This method is more suitable for the data compression method based on Kafka primary keys adopted by the present invention.
[0160] The above business data analysis system further includes a MySQL or Doris database. Kafka outputs the calculation results written by Flink to the MySQL or Doris database. The MySQL or Doris database is used as an output unit to display the calculation results. When the calculation result is the intermediate result of Flink multi-layer calculation, the MySQL or Doris database can be used as a tool to verify the calculation results of Flink, such as comparing with the calculation results of the previous layer of Flink, or outputting for manual verification.
[0161] Regarding the devices in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments of the related methods, and will not be elaborated herein.
[0162] The solutions of the present application have been described in detail with reference to the accompanying drawings above. In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not described in detail in a certain embodiment, reference may be made to the relevant descriptions of other embodiments. Those skilled in the art should also be aware that the actions and modules involved in the specification are not necessarily essential to the present application. Additionally, it can be understood that the steps in the tone scoring method of the embodiments of the present application can be adjusted, combined, and deleted according to actual needs, and the modules in the devices of the embodiments of the present application can be combined, divided, and deleted according to actual needs.
[0163] The present invention also provides a computing device, including a memory and a processor.
[0164] The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0165] The memory may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, the ROM can store static data or instructions required by the processor or other modules of the computer. The permanent storage device can be a readable and writable storage device. The permanent storage device can be a non-volatile storage device that does not lose the stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a mass storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In some other embodiments, the permanent storage device can be a removable storage device (such as a floppy disk, optical drive). The system memory can be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory can store some or all of the instructions and data required by the processor during operation. In addition, the memory can include any combination of computer-readable storage media, including various types of semiconductor storage chips (DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks can also be used. In some embodiments, the memory can include a removable storage device that is readable and / or writable, such as a compact disc (CD), read-only digital versatile disc (such as DVD-ROM, dual-layer DVD-ROM), read-only Blu-ray disc, super density disc, flash memory card (such as SD card, min SD card, Micro-SD card, etc.), magnetic floppy disk, etc. Computer-readable storage media do not include carrier waves and instantaneous electronic signals transmitted wirelessly or wiredly.
[0166] Executable code is stored on the memory, and when the executable code is processed by the processor, it can cause the processor to execute some or all of the methods described above.
[0167] In addition, the intonation scoring method according to the present application can also be implemented as a computer program or computer program product, which includes computer program code instructions for executing some or all of the steps in the above-mentioned intonation scoring method of the present application.
[0168] Alternatively, the present application can also be implemented as a non-transitory machine-readable storage medium (or computer-readable storage medium, or machine-readable storage medium), on which executable code (or computer program, or computer instruction code) is stored. When the executable code (or computer program, or computer instruction code) is executed by the processor of a computing device (or electronic device, server, etc.), it causes the processor to execute some or all of the steps of the above-mentioned intonation scoring method according to the present application.
[0169] Those skilled in the art will also understand that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the application herein can be implemented as electronic hardware, computer software, or a combination of both.
[0170] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems and intonation scoring methods according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0171] The various embodiments of the present application have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art in the field of this technology without departing from the scope and spirit of the described embodiments. The choice of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the field of this technology to understand the embodiments disclosed herein.
Claims
1. A business data analysis method, characterized in that, Applied to the query of online product business metrics, including: Obtain a target information query request; the target information includes the current order quantity of a certain teacher or the renewal rate of the teacher; Obtain the processing logic of business data according to the target information query request, and the storage area of the business data related to the processing logic in the distributed storage system; the processing logic includes a logical calculation method and a data source for implementing the logical calculation; through the data source of the processing logic, obtain the storage address of the data required to execute the processing logic; Obtain all business data in the storage area respectively, and the business data has a business identifier; Process the business data according to the processing logic and the business identifier of the business data to obtain the target information; The business data stored in the distributed storage system includes real-time business data and historical business data; use Canal to obtain real-time business data from the MySQL database and send the real-time business data to the distributed storage system for storage; the real-time business data is the incremental data of the order.
2. The method according to claim 1, wherein It also includes: Store business data with the same business identifier in the same storage area of the distributed storage system according to a preset policy.
3. The method according to claim 2, wherein It also includes: Obtain real-time business data and send it to the distributed storage system for storage; And, before obtaining the real-time business data, obtain the historical business data stored in the data warehouse and send it to the distributed storage system.
4. The method according to claim 3, characterized in that, Before the processing logic processes the business data according to the business identifier, it also includes: Use the business identifier in the processing logic to verify whether the processing logic is correct.
5. The method according to claim 4, wherein The processing logic processes the business data according to the business identifier, including: Use Flink to execute the processing logic and process the business data according to the business identifier; And write the processing results with the same business identifier into the same storage area of the distributed storage system.
6. The method according to claim 5, wherein After writing the processing results with the same business identifier into the same storage area of the distributed storage system, it also includes: The distributed storage system outputs the operation result to the MySQL or Doris database; the MySQL or Doris database is used as an output unit to display the operation result and output the queried target information.
7. The method according to any one of claims 1 to 6, characterized in that, The business identifier includes, but is not limited to, one of the following identifiers or a combination of two or more identifiers: Order number, class ID, teacher ID, user ID, semester ID, course project ID.
8. A business data processing method, characterized in that, Applied to the query of online product business metrics, including: Obtain business data, Obtain the business identifier of the business data; Write business data with the same business identifier into the same storage area of the distributed storage system; The obtaining of the business data includes: using Flink to perform the operation of the processing logic on the business data already stored in the distributed storage system to obtain the business data; The processing logic includes a logical calculation method and a data source for implementing the logical calculation; through the data source of the processing logic, obtain the storage address of the data required to execute the processing logic; The business data stored in the distributed storage system includes real-time business data and historical business data; Canal is used to obtain real-time business data from the MySQL database and send the real-time business data to the distributed storage system for storage; the real-time business data is the incremental data of orders.
9. The method according to claim 8, wherein: The distributed storage system is Kafka.
10. The method according to claim 9, wherein The obtaining of business data includes: Obtaining real-time business data by using Canal; Or / and, Performing operations on the first-layer business data stored in Kafka by using Flink to obtain second-layer business data; And writing the second-layer business data with the same business identifier into the same storage area in Kafka.
11. The method according to claim 9 or 10, characterized in that: The obtaining of the business identifier of the business data and writing the business data with the same business identifier into the same storage area in Kafka includes: Obtaining the business identifier of the business data according to the definition of the business identifier in the preset mapping table; Performing partition operation on the business identifier according to the preset policy, and writing the business data into the corresponding storage area in Kafka according to the partition operation result.
12. The method according to claim 11, wherein The business identifier includes, but is not limited to, one of the following identifiers or a combination of two or more identifiers: Order number, class ID, teacher ID, user ID, semester ID, course project ID.
13. The method according to claim 12, wherein It further includes: When the preset conditions are met, when two or more business data with the same business identifier are stored, the latest obtained business data is retained, and the rest of the business data is deleted.
14. A business data analysis system, characterized in that, Applied to the query of online product business metrics, including: A receiving unit that obtains a target information query request; the target information includes the current order quantity of a certain teacher or the renewal rate of the teacher; A distributed storage system for storing business data, and the business data with the same business identifier is stored in the same storage area in the distributed storage system; A processing unit that obtains the processing logic of the business data according to the target information query request, and the storage area of the business data related to the processing logic in the distributed storage system; obtains all the business data in the storage area, executes the processing logic, processes the business data according to the business identifier, and obtains the target information; the processing logic includes a logical calculation method and a data source for implementing the logical calculation; through the data source of the processing logic, the storage address of the data required to execute the processing logic is obtained; The business data stored in the distributed storage system includes real-time business data and historical business data; Canal is used to obtain real-time business data from the MySQL database and send the real-time business data to the distributed storage system for storage; the real-time business data is the incremental data of orders.
15. The system according to claim 14, wherein: The distributed storage system is Kafka; And it further includes: An incremental data subscription module for synchronizing the real-time business data of users to Kafka; The Kafka stores the business data with the same business identifier in the same storage area.
16. The system according to claim 15, wherein: The processing unit is Flink; The Flink obtains the first-layer service data stored in Kafka according to the target information query request and executes the processing logic to obtain the second-layer service data, and writes the second-layer service data with the same service identifier into the same storage area in Kafka.
17. The system according to claim 14, characterized in that, It further includes: A historical service database for storing historical service data; Kafka, after obtaining the historical service data, obtains the real-time service data synchronized by the incremental data subscription module.
18. The system according to claim 16, wherein The system includes multiple Flinks; the multiple Flinks are connected in series for performing hierarchical calculation.
19. The system according to claim 18, wherein The system further includes: A MySQL or Doris database; Kafka writes the Flink processing result into the MySQL or Doris database.
20. The system according to claim 18, wherein The service identifier includes, but is not limited to, one of the following identifiers or a combination of two or more identifiers: Order number, class ID, teacher ID, user ID, semester ID, course project ID.
21. The system according to any one of claims 14 to 18, characterized in that, It further includes: A verification unit for verifying whether the processing logic is correct by using the service identifier in the processing logic.
22. A non-transitory machine-readable storage medium having executable code stored thereon, which when executed by a processor of a computing device, causes the processor to execute the method according to any one of claims 1-13.
Citation Information
Patent Citations
Data query method and system
CN110263061A
Data processing method and device for server cluster, computer equipment and medium
CN111046057A