Data analysis method and system and storage medium
By matching fact data and dimensional data groups in the data analysis system and determining the target analysis node for analysis, the problem of low real-time data calculation efficiency in the existing technology is solved, and more efficient data analysis is achieved.
Patent Information
- Application Number
- CN202510168150.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-07-11
AI Technical Summary
Existing real-time data calculation models are inefficient when processing large amounts of factual data, resulting in data analysis extending calculation time and difficult to meet business needs.
By matching factual data with the dimensional data set stored by each data analysis node in the data analysis system, the target analysis node is determined and analyzed at that node, the amount of matching data of each node is reduced, and data analysis is performed using preset analysis methods.
It alleviates the load pressure of data analysis nodes, improves the data analysis efficiency of factual data, and reduces calculation time.
Smart Images

Figure CN120296432A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data analysis, and in particular, to a data analysis method, an electrical system, and a storage medium. Background Art
[0002] A large amount of data is generated by various industries every day. By analyzing and calculating these data, result data with high value density can be obtained, and these result data can provide certain guiding significance for subsequent business planning to assist decision-making. The traditional analysis method is to analyze after storing the data on disk, and this kind of analysis method is called offline analysis. Offline analysis needs to calculate after the data is completely stored, which has a certain time delay. For example, at most, the data of the previous day can be calculated on the same day, and its time delay is T + 1. Due to its long delay, it is difficult to meet the business requirements. Then, storing the data after real-time analysis can solve the current problem of long time delay. Among them, real-time analysis of data can be real-time calculation. Real-time calculation means that the data participates in the calculation immediately after entering the system and generates results, and its time delay is mainly the calculation time. In real-time calculation, there may be multiple types of data participating in the calculation, and logically, there are fact data (fact table) and dimension data (dimension table). Taking e-commerce shopping as an example, the continuous order data is called the fact table, and the data such as product information that does not change frequently is called the dimension table. However, with the development of the Internet, especially the rise of e-commerce, fact data (for example, online network payment transaction volume) has shown explosive growth. Then, the real-time calculation of fact data is usually obtained by using a relatively decentralized calculation model. Calculating a large amount of fact data using the traditional real-time data calculation model will prolong the calculation time, resulting in low data analysis efficiency of fact data.
[0003] Aiming at the existing technical defects, how to provide a solution that can improve the data analysis efficiency of fact data is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0004] This application provides at least one data analysis method, system, and storage medium.
[0005] This application provides a data analysis method, including: a data distribution node distributes fact data related to a target scenario to each data analysis node, and each data analysis node stores a dimension data group, and each dimension data group is obtained by dividing a number of dimension data related to the target scenario; each data analysis node respectively matches the received fact data with the dimension data group stored by itself, and among them, the data analysis node where the dimension data group that matches the fact data successfully is the target analysis node; the target analysis node analyzes the fact data according to the preset analysis method corresponding to the target analysis node to obtain the analysis result of the fact data.
[0006] The present application provides a data analysis system, including several data analysis nodes and a data distribution node. Each data analysis node is communicatively connected to the data distribution node to implement the above data analysis method.
[0007] The present application provides a computer-readable storage medium, on which program instructions are stored. When the program instructions are executed by a processor, the above data analysis method is implemented.
[0008] In the above solution, the data distribution node distributes fact data related to the target scenario to each data analysis node. Each data analysis node stores a dimension data group obtained by dividing several dimension data related to the target scenario. Each data analysis node respectively matches the received fact data with the dimension data group stored therein, which can reduce the amount of data that each data analysis node needs to match. The data analysis node where the dimension data group successfully matched with the fact data is located is the target analysis node. The target analysis node analyzes the fact data according to the preset analysis method corresponding to the target analysis node to obtain the analysis result of the fact data. Compared with the situation where the fact data is only matched with all dimension data on the same data analysis node and then the fact data is analyzed subsequently, resulting in an excessive load on the data analysis node, the present application can relieve the load pressure of the data analysis node, thereby improving the data analysis efficiency of the fact data.
[0009] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings herein are incorporated into the specification and form a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to explain the technical solutions of the present application.
[0011] Figure 1 is a flowchart of an embodiment of the data analysis method of the present application Figure 1 ;
[0012] Figure 2 is a flowchart of an embodiment of the data analysis method of the present application Figure 2 ;
[0013] Figure 3 is a flowchart of an embodiment of the data analysis method of the present application Figure 3 ;
[0014] Figure 4 is a flowchart of an embodiment of the data analysis method of the present application Figure 4 ;
[0015] Figure 5 is a flowchart of an embodiment of the data analysis method of the present application Figure 5 ;
[0016] Figure 6 It is a schematic framework diagram of an embodiment of the data analysis method of the present application;
[0017] Figure 7 It is a schematic structural diagram of an embodiment of the data analysis system of the present application;
[0018] Figure 8 It is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. Detailed Embodiments
[0019] The following will combine the accompanying drawings of the specification to elaborate in detail on the solutions of the embodiments of the present application.
[0020] In the following description, specific details such as specific system structures, interfaces, and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.
[0021] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the front and rear associated objects. In addition, "multiple" in this article means two or more than two. In addition, the term "at least one" in this article represents any one of multiple or any combination of at least two of multiple. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.
[0022] The present application provides some data analysis methods and data analysis devices. The application scenarios of the data analysis method include but are not limited to real-time calculation scenarios for real-time data streams. The execution subject of the data analysis method can be a data analysis device or a data analysis system. For example, the data analysis device can be set in a terminal device, a server, or other processing devices. Among them, the terminal device can be a device for data analysis, a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, etc. In some possible implementation manners, the data analysis method can be implemented by a processor calling computer-readable instructions stored in a memory.
[0023] Please refer to Figure 1 , Figure 1 is a flowchart of an embodiment of the data analysis method of the present application Figure 1. Specifically, the data analysis method is applied to a data analysis system, which includes several data analysis nodes and a data distribution node. The data analysis method may include the following steps:
[0024] Step S11: The data distribution node distributes fact data related to the target scenario to each data analysis node.
[0025] The data distribution node included in the data analysis system is used to distribute the obtained fact data to each data analysis node included in the data analysis system respectively. Each data analysis node included in the data analysis system is used to match the received fact data with the dimension data group stored by the data analysis node itself. For each data analysis node, when the received fact data matches the dimension data group stored by the data analysis node itself, the data analysis node is further used to analyze the fact data according to the preset analysis method corresponding to the data analysis node to obtain the analysis result of the fact data.
[0026] The target scenario can be one of the business scenarios that require data analysis. The business scenarios capable of data analysis include but are not limited to e-commerce transaction scenarios, intelligent transportation scenarios, social media analysis scenarios, target object analysis scenarios in a fixed area, e-commerce recommendation scenarios, and so on. In this application, the target scenario is taken as an e-commerce transaction scenario as an example. Among them, the fact data related to the target scenario can be any piece of fact data in the data stream collected for the target scenario. The data analysis system also includes a data acquisition component. The data acquisition component collects the first data stream for the target scenario in real time. After the collection is completed, each fact data in the data stream can be connected to the data distribution node. After the data source is connected, the data distribution node can be used to distribute the data to each data analysis node. The fact data related to the target scenario can be the fact table in the data stream related to the target scenario. For example, the fact data can be the order data in the e-commerce transaction scenario. The several dimension data can be several dimension tables related to the target scenario. For example, each dimension data can be the product information of each product in the e-commerce transaction scenario. Specifically, each piece of dimension data can contain the complete product information of a product in the target scenario.
[0027] Each data analysis node stores a dimension data group, and each dimension data group is obtained by dividing several dimension data related to the target scenario. Each data analysis node has a built-in dimension data group. After the dimension data groups stored by all data analysis nodes are combined, they form a complete set of several dimension data.
[0028] The above step S11 can be that after a data stream of the target scenario is collected once, the data stream includes several fact data. For any piece of fact data, the data distribution node sends the fact data to each data analysis node in the data analysis system at the same time.
[0029] Step S12: Each data analysis node respectively matches the received factual data with the dimension data groups stored by itself.
[0030] The data analysis node where the successfully matched dimension data group is located is the target analysis node.
[0031] For each data analysis node, the received factual data is matched with the dimension data group stored by the data analysis node. In some application scenarios, for each data analysis node, in response to the successful matching between the received factual data and the dimension data group stored by the data analysis node, the data analysis node is used as the target analysis node. In other application scenarios, for each data analysis node, in response to the failure of the received factual data to match the dimension data group stored by the data analysis node, the data analysis node does not perform data analysis on the factual data.
[0032] The above step S12 may be that for each data analysis node, the received factual data is matched with the information to be matched in the dimension data group stored by the data analysis node to obtain the matching result between the factual data and the dimension data group. The matching result includes successful matching or failed matching.
[0033] The information to be matched in the dimension data group stored by the data analysis node may be each dimension data included in the dimension data group, or an information table synthesized from the same attribute column based on each dimension data. Specifically, the above matching method may be to perform a hash operation on the identification information of the factual data and the identification information in the information to be matched to obtain the matching result. Exemplarily, the identification information of the factual data may be the product identification included in the fact table. The identification information in the information to be matched may be the product identification included in the dimension table corresponding to each dimension data in the dimension data group. If any one of the dimension data in the factual data and the dimension data is successfully matched, it indicates that the received factual data is successfully matched with the dimension data group stored by the data analysis node. If none of the dimension data in the factual data and the dimension data is successfully matched, it indicates that the received factual data fails to match the dimension data group stored by the data analysis node.
[0034] Step S13: The target analysis node analyzes the factual data according to the preset analysis method corresponding to the target analysis node to obtain the analysis result of the factual data.
[0035] The preset analysis method includes a calculation framework corresponding to a dimension data group. Each dimension data group corresponds to a calculation framework. Each calculation framework includes at least one rule framework, and each rule framework includes at least one rule instance. The calculation framework is the calculation logic related to the fact data set according to the requirements of the target scenario. The relevant calculation logic is executed on the fact data to complete the data analysis of the fact data. For the calculation framework corresponding to any dimension data group, the various rule frameworks included in the calculation framework are distinguished by the different attribute information of the various dimension data in the dimension data group. Different attribute information in the various dimension data corresponds to different rule frameworks. For example, the attribute information of the various dimension data includes, but is not limited to, first attribute information, second attribute information, and so on. The attribute value corresponding to each attribute information in the dimension data is the above-mentioned analysis parameter. The calculation framework can be a preset calculation framework or a custom calculation framework. For example, the calculation framework corresponding to any dimension data group includes a first rule framework and a second rule framework, where the first rule framework is used to represent the calculation logic related to the first attribute information in the attribute information of the dimension data group. The second rule framework is used to represent the calculation logic related to the second attribute information in the attribute information of the dimension data group. Specifically, in the case where the target scenario is an e-commerce transaction scenario, the first attribute information can be the selling price of the product in the product attribute information, and the second attribute information can be the cost of the product in the product attribute information, etc. The above-mentioned preset analysis method can be to analyze the fact data according to the calculation framework corresponding to the target analysis node to obtain the analysis result of the fact data. The analysis result of the above-mentioned fact data can be at least one calculation result obtained through the calculation framework corresponding to the target analysis node.
[0036] Exemplarily, in the present application, the data acquisition component collects the first data stream in real time for the target scenario. After the collection is completed, each factual data in the data stream can be connected to the data distribution node. After the data source is connected, the data distribution node can be used to distribute the data to each data analysis node. Each data analysis node loads each rule framework in the specific calculation framework and waits to be filled with the distributed rule instances to obtain a complete calculation framework loading. Each factual data is calculated in the target analysis node through the factual data and the rule instances in the corresponding various rule frameworks to obtain the calculation results corresponding to each rule instance, and the calculation results corresponding to each rule instance are used as the analysis results of the above factual data. Using the calculation results corresponding to each rule instance as the analysis results of the above factual data includes: for each rule framework, aggregating the calculation results corresponding to each rule instance belonging to the rule framework into the calculation result belonging to the rule framework. Aggregating the calculation results of each rule framework into the calculation result belonging to the calculation framework, and using the calculation result belonging to the calculation framework as the analysis result of the above factual data. The analysis results of the above factual data include the calculation results corresponding to each target rule instance, and each target rule instance is each rule instance included in each rule framework in the calculation framework corresponding to the target analysis node.
[0037] The above custom calculation framework can be obtained by distributing rule instances, that is, the user distributes specific rules to the rule framework. Each rule framework will receive rule instances, and the rule instances are also calculation instances. It is determined whether to apply the distributed rule instance to the rule framework according to whether the type of the distributed calculation instance matches the type of the rule framework. Whether the type of the calculation instance matches the type of the rule framework can be whether the corresponding attribute information is consistent. For example, if the attribute information of the distributed rule instance is age and the attribute information corresponding to the rule framework that currently receives the distributed rule instance is price, then the type of the distributed calculation instance does not match the type of the rule framework, and the rule framework that currently receives the distributed rule instance does not need to apply the rule instance to its own rule framework. In the process of the inflow of each factual data in the data stream in the present application, the calculation framework will calculate the data according to the factual data and the rule instance to obtain the calculation result corresponding to the rule instance, and pass the calculation result corresponding to the rule instance downstream to obtain the analysis result of the factual data.
[0038] It can be understood that in some application scenarios, different data analysis nodes can be in different terminal devices. The data distribution node and any one data analysis node can be in the same terminal device. In some other application scenarios, different data analysis nodes can be in the same terminal device and are executed by different threads. In some other application scenarios, the data distribution node and any one data analysis node may not be in the same terminal device. In some other application scenarios, the data distribution node and several data analysis nodes can be in the same terminal device and are executed by different threads.
[0039] In the above solution, the data distribution node distributes the fact data related to the target scenario to each data analysis node. Each data analysis node stores a dimension data group obtained by dividing several dimension data related to the target scenario. Each data analysis node respectively matches the received fact data with the dimension data group stored by itself, which can reduce the amount of data that each data analysis node needs to match. The data analysis node where the dimension data group that successfully matches the fact data is located is the target analysis node. The target analysis node analyzes the fact data according to the preset analysis method corresponding to the target analysis node to obtain the analysis result of the fact data. Compared with the situation where the fact data is only matched with all dimension data on the same data analysis node and then the fact data is analyzed, resulting in an excessive load on the data analysis node, this application can relieve the load pressure of the data analysis node, thereby improving the data analysis efficiency of the fact data.
[0040] Please refer to Figure 2 , Figure 2 which is the process schematic diagram of an embodiment of the data analysis method of this application. Figure 2 .
[0041] In some embodiments, each dimension data carries an analysis parameter. The analysis parameter carried by each dimension data is associated with the computing framework to which the dimension data group corresponding to each dimension data belongs. The dimension data group stored in the target analysis node is the target dimension data group. The analysis result of the fact data includes the calculation result of the fact data.
[0042] The above step S13 may include the following steps: Step S21: Supplement the analysis parameters recorded in the target dimension data group to the fact data to obtain the target fact data. Step S22: Calculate the target fact data according to the computing framework to which the target dimension data group belongs to obtain the calculation result of the fact data.
[0043] The data of each dimension carries analysis parameters. The analysis parameters carried by the data of each dimension are associated with the computing framework to which the dimension data group corresponding to the data of each dimension belongs. Specifically, the analysis parameters carried by the data of each dimension are associated with the attribute information corresponding to each rule framework in the computing framework to which the dimension data group corresponding to the data of each dimension belongs. Each of the several dimension data contains attribute information and attribute values related to the target scenario. Among them, the attribute value corresponding to the attribute information in each dimension data is the analysis parameter. The dimension data group stored in the target analysis node is the target dimension data group. The computing framework corresponding to the target dimension data group is the target computing framework. Each rule framework included in the target computing framework is each target rule framework. Each rule instance in each target rule framework is the target rule instance to which each target rule framework belongs. Based on the above step S12, the target dimension data in the target dimension data group can be determined. Among them, the target dimension data is the dimension data in the target dimension data group that matches the fact data successfully.
[0044] The above step S21 may be to supplement the analysis parameters recorded in the target dimension data to the fact data to obtain the target fact data. In some application scenarios, all the analysis parameters recorded in the target dimension data are supplemented to the fact data to obtain the target fact data. In other application scenarios, the target analysis parameters recorded in the target dimension data are supplemented to the fact data to obtain the target fact data. Among them, the target analysis parameters are the analysis parameters related to the attribute information to which the target computing framework belongs among all the analysis parameters recorded in the target dimension data. The attribute information to which the target computing framework belongs is the attribute information corresponding to each rule framework under this target computing framework.
[0045] The computing framework to which the target dimension data group belongs includes at least one target rule framework. The above step S22 may be to copy the target fact data by the target number of copies, input each copy of the target fact data into each target rule framework simultaneously, obtain the calculation results of each target rule framework, and use the calculation results of each target rule framework as the calculation result of this fact data.
[0046] Exemplarily, the above step S12 is mainly an operation of filtering the factual data in the collected data stream. By each data analysis node respectively colliding the received factual data with the dimension data groups stored by itself, the target analysis node can be determined, and the filtering of the factual data can be realized, that is, not all data analysis nodes will analyze the factual data according to the corresponding preset analysis method to obtain the analysis result of the factual data. The above step S21 is mainly an operation of filling the factual data in the collected data stream. The target dimension table data group stored in the target analysis node contains fixed dimension table data, wherein the target dimension data is the fixed dimension table data. The target dimension table data is mainly used to collide with the received factual data, and the result of the collision is the filled factual data, and the filled factual data is the target factual data.
[0047] It can be understood that the collision between the above factual data and the dimension data in each dimension data group can be to match the factual data with the dimension data groups stored by each data analysis node and supplement the analysis parameters recorded in the target dimension data to the factual data to obtain the target factual data. In this collision process, the factual data will be filtered and filled. The target factual data is sent as the collision result to the downstream process.
[0048] Exemplarily, the filtering and filling of the factual data in this application can be an operation of associating the received factual data with the corresponding dimension table data in each data analysis node's dimension data group. In the actual use of the data analysis system, the data volume of several dimension table data is relatively large. The data analysis system of this application uses a distributed method, and each data analysis node can correspond to one machine for use. In the process of designing the dimension table data, the quantity advantage of the machines is fully utilized, and a part of the dimension table data is stored on each machine. When each data analysis node in the data analysis system starts, the first thing to do is to load the dimension data group stored by this data analysis node, and loading the dimension data group stored by this data analysis node is a prerequisite for filtering and filling the factual data. In each data analysis node's dimension data group, each dimension data has a unique id to determine the unique attribute of this data. In the data stream, that is, in the real-time data stream, each factual data also has this unique id, and the unique id in the factual data and the unique id in each dimension data are in one-to-one correspondence. Whether the unique id in the factual data corresponds to the unique id in any dimension data is used to judge the dimension data group and the target dimension data that are successfully matched in the above step S12. Among them, if the unique id in the factual data corresponds to the unique id in any dimension data, this dimension data is used as the target dimension data, and the dimension data group where this dimension data is located is used as the target dimension data group.
[0049] During the process of loading dimension table data, the number of data analysis nodes currently in use will be obtained. On each data analysis node, a part of the dimension data among several dimension data will be stored. Then, hash processing will be performed based on the unique ID in the fact data and the unique IDs in each dimension data, and each fact data in the data stream will be evenly scattered and distributed to different data analysis nodes for subsequent processing. The unique IDs of the dimension data in each dimension data group of each data analysis node are not repeated and have global uniqueness. That is to say, the unique IDs of each dimension data among several dimension data have global uniqueness, and the attribute information of each dimension data among several dimension data contains identification information. The analysis parameter of the identification information of each dimension data is the value corresponding to the unique ID of the dimension data. The hash operation of this application only guarantees the relative average of the number of dimension table data distributed to each data analysis node, and does not need to guarantee that the number of each dimension data in the dimension data group distributed to each data analysis node is equal.
[0050] The real-time data stream means that when each fact data enters the data analysis system, the same hash operation will be performed based on the unique ID of each fact data and the unique IDs of each dimension data corresponding to each data analysis node, so that the fact data will be sent to the machines related to the target analysis node for subsequent operations related to the computing framework. After the fact data enters the target analysis node, it will be associated with the dimension table data stored in the target analysis node. Here, the operation is to find the only one piece of dimension table data with the same ID in the data loaded on the machine through the unique ID of the real-time data, that is, the target dimension table data, and supplement the required fields in the target dimension data to the fact data to obtain the target fact data. Then, the target fact data will be sent to the downstream process.
[0051] It can be considered that each data analysis node in this application loads each rule framework in the computing framework it belongs to, accepts the upstream data stream, shunts the fact data according to whether the above matching is successful, and fills or supplements the shunted fact data to obtain the target fact data.
[0052] Please refer to Figure 3 , Figure 3 which is the process schematic diagram of an embodiment of the data analysis method of this application. Figure 3 。
[0053] In some embodiments, the computing framework to which the target dimension data group belongs includes at least one target rule framework, and each target rule framework includes at least one target rule instance. The above step S22 may include the following steps: For each target rule framework, execute as Figure 3The following steps are shown: Step S31: Calculate the target fact data according to each target rule instance in the target rule framework to obtain the calculation results corresponding to each target rule instance. Step S32: Use the calculation results corresponding to each target rule instance as the calculation result corresponding to the target rule framework.
[0054] The above step S22 can be: Calculate the target fact data according to each target rule instance in the target rule framework to obtain the calculation results corresponding to each target rule instance. Use the calculation results corresponding to each target rule instance as the calculation result corresponding to the target rule framework. Use the calculation results corresponding to each target rule framework as the calculation result of the fact data.
[0055] Exemplarily, the data analysis node of the present application loads each rule framework in the affiliated calculation framework, receives the upstream data stream, shunts the fact data according to whether the above matching is successful, and fills or supplements the shunted fact data to obtain the target fact data. Each rule framework is a rule model defined by the user. There is the same operator calculation logic between the rule instances in the same rule model, and the rule instances with the same operator calculation logic are attributed to the same rule framework. Classification will be carried out according to different rule frameworks here. For each rule framework, the target fact data obtained from the upstream data stream will be shunted, that is, copied into multiple copies of the target fact data, and each copy of the target fact data will be input into each rule framework respectively. Each rule framework under the same calculation framework accesses the same copied data stream, and the calculation results of each rule instance in each rule framework are aggregated into the calculation result belonging to the rule framework. The calculation results of each rule framework are aggregated into the calculation result belonging to the calculation framework, and the calculation result belonging to the calculation framework is used as the analysis result of the above fact data.
[0056] A rule framework is a classification with the same calculation rules, and the rule framework provides a basis for shunting the distributed calculation of the data stream. A type of rule framework can be understood as a kind of the same data calculation logic. For example, when the attribute information is age, the calculation logic corresponding to the rule framework can classify the fact data according to age or classify the target objects in different age groups. The rule instances under the rule framework are specific calculation rules. When the attribute information is age, the rule framework includes a first rule instance and a second rule instance. Exemplarily, the first rule instance can be to calculate the age greater than 20 as a category. The second rule instance is to calculate the age less than 30 as a category. It can be understood that under this rule framework, all rule instances are calculated according to the analysis parameters corresponding to the attribute information in the target fact data, that is, the value of the age field.
[0057] The rule framework of this application needs to be defined before the data analysis system is started or before each data analysis node filters and fills the factual data, and will not be modified, added to, or deleted midway. Different rule frameworks are independent of each other and do not affect each other. Under one rule framework, multiple fields in the target factual data can be combined for calculation, not limited to the number, type, etc. of the fields, that is, the attribute information corresponding to the rule framework can be multiple. Under one rule framework, all rule instances associated with this framework are related. Rule instances only exist under the rule framework. If the rule instances issued by the user do not have a corresponding rule framework running, these rule instances will be discarded. The rule instances under the rule framework can be preset, and can also be dynamically added, deleted, and modified.
[0058] The rule framework runs on a machine, that is, the rule framework runs on a data analysis node. The target factual data obtained from the upstream data stream will determine how many copies of the target factual data need to be copied according to the number of rule frameworks included in the calculation framework of the target analysis node. Each rule framework included in the calculation framework of the target analysis node will receive a complete target factual data.
[0059] Please refer to Figure 4 , Figure 4 which is the process schematic of an embodiment of the data analysis method of this application Figure 4 .
[0060] In some embodiments, before the above step S22, the data analysis method may further include the following steps: Step S41: Obtain at least one target configuration instruction. Each target configuration instruction represents configuration information related to the rule instances in each rule framework. Step S42: Parse each target configuration instruction to obtain the target configuration information corresponding to each target configuration instruction. Step S43: Configure the initial rule instances in each rule framework based on each target configuration information to obtain each rule instance.
[0061] Each target configuration instruction represents configuration information related to rule instances in each rule framework. Among them, each target configuration instruction can be used to update or delete the rule instances belonging to each rule framework of each data analysis node, and to add new rule instances to each rule framework of each data analysis node. Each target configuration instruction can be instruction information obtained in response to the user's instruction configuration operation related to rule instances on the display interface where the user is located. One target configuration instruction corresponds to one rule instance to be configured. In some application scenarios, the target configuration information includes indication information for indicating the rule framework to which the rule instance to be configured belongs. In some other application scenarios, the target configuration information includes indication information for indicating the configuration type of the rule instance to be configured. In some other application scenarios, the target configuration information includes indication information for indicating the configuration content of the rule instance to be configured.
[0062] The method for obtaining each target configuration instruction in step S41 above can be to pull the target configuration information corresponding to the rule instance to be configured from the database at a preset frequency, or to directly receive the target configuration information corresponding to the rule instance issued by the user. Among them, the database is used to store the target configuration information corresponding to the rule instance issued by the user.
[0063] Step S43 above can be to obtain the initial rule instances in each rule framework, where the initial rule instances are the rule instances in each rule framework before configuration. Based on each target configuration information, determine each rule instance to be configured. In some application scenarios, for each rule instance to be configured, in response to the instance identifier of the rule instance to be configured matching any one of the initial rule instances, update or delete the successfully matched initial rule instance according to the rule instance to be configured. Thus, the rule instances in each rule framework after configuration are obtained. In some other application scenarios, for each rule instance to be configured, in response to the instance identifier of the rule instance to be configured not matching any one of the initial rule instances, determine whether the rule framework to which the rule instance to be configured belongs exists. In response to the rule framework to which the rule instance to be configured belongs not existing, discard the rule instance to be configured and the target configuration information corresponding to the rule instance to be configured. In response to the rule framework to which the rule instance to be configured belongs existing, add the rule instance to be configured under the rule framework to which the rule instance to be configured belongs. Thus, the rule instances in each rule framework after configuration are obtained.
[0064] Please refer to Figure 5 , Figure 5 which is a flowchart illustration of an embodiment of the data analysis method of this application Figure 5 .
[0065] In some embodiments, step S43 above may include the following steps: For each target configuration information, perform as Figure 5The following steps are shown: Step S51: Based on the target configuration information, determine the rule framework indication information, the configuration type, and the configuration content. The rule framework indication information is used to represent the rule framework to which the rule instance to be configured belongs, and the configuration type includes one of addition, deletion, and update. Step S52: In response to the existence of the rule framework corresponding to the rule framework indication information, configure the rule framework corresponding to the rule framework indication information based on the configuration content and the configuration type to obtain the target rule instance. Or, Step S53: In response to the non-existence of the rule framework corresponding to the rule framework indication information, discard the target configuration information.
[0066] The above step S51 may be to determine the rule framework indication information, the configuration type, and the configuration content related to the rule instance to be configured based on the target configuration information.
[0067] In the configuration process of the rule instance, the user needs to send the target configuration instruction and / or the target configuration information to the specified database. The database can store each target configuration instruction and / or target configuration information. The data analysis system also includes a rule broadcaster. The rule broadcaster in the data analysis system will query the latest data in the database every preset time and broadcast the target configuration information of the rule instance to be configured in this latest data to each rule framework.
[0068] The rule instances sent down need to be stored in the specified database, and the rule broadcaster will regularly pull the latest changed data. For example, the time interval for the rule broadcaster to pull the rules in the database is 5s. The rule broadcaster will broadcast and send the changed data to all the rule frameworks. The changed data is at least one rule instance to be configured. Within each rule framework, it will be determined whether to process each rule instance to be configured according to the rule type. Within each rule framework, all the rule instances that need to be calculated under this rule framework will be stored.
[0069] Table 1: Example table of rule instances to be configured
[0070]
[0071] As shown in Table 1, each rule instance has a unique instance identifier (i.e., instance id), which is globally unique. The type of the rule framework to which the rule instance to be configured belongs corresponds to the type listed in the rule frameworks corresponding to each data analysis node in the data analysis system, and is used to confirm the rule framework to which the rule instance to be configured belongs.
[0072] The configuration types of the rule instances to be configured are divided into three types: Add type, which is used to add the rule instances to be configured in any rule framework. Mod type, which is used to update the initial rule instances with the rule instances to be configured in any rule framework to obtain the final rule instances. Del type, which is used to delete the initial rule instances with the instance identifiers corresponding to the rule instances to be configured in any rule framework.
[0073] The rule content of the rule instances to be configured is the above configuration content, specifically the complete and latest rule content corresponding to the rule instances to be configured. Each rule framework will receive the full set of rule instance broadcast data, that is, the rule instance broadcast data includes all the rule instances to be configured. After receiving the rule instance broadcast data, all the rule frameworks corresponding to each data analysis node will perform the following operations: judge the rule framework type to which the rule instance to be configured belongs. If the rule framework type to which the rule instance to be configured belongs does not match the types of all the rule frameworks of each data analysis node before configuration, it will be directly discarded. Then, operate on the rule instance list in the rule framework according to the configuration type of the rule instance to be configured. Specifically, in response to the configuration type being Add type, directly add the rule instance to be configured to the list of the rule framework type to which the rule instance to be configured belongs. In response to the configuration type being Mod type, find the corresponding initial rule instance in the list of the rule framework type to which the rule instance to be configured belongs according to the instance identifier of the rule instance to be configured, and perform a full replacement. In response to the configuration type being Del type, find the corresponding initial rule instance in the list of the rule framework type to which the rule instance to be configured belongs according to the instance identifier of the rule instance to be configured and delete it.
[0074] It can be considered that the present application makes a preliminary judgment based on the rule framework type to which the rule instance to be configured belongs, which can reduce the excessive consumption of computing resources by rule instances that do not meet the requirements. Subsequently, only according to the configuration type of the rule instance to be configured, the initial rule instances in each rule framework can be modified, which can simplify the process of configuring the latest rule instances, thereby improving the efficiency of data analysis.
[0075] In some embodiments, the above data analysis method further includes the following steps: obtain the configuration interval. Determine whether the time interval between the current time and the last configuration event reaches the configuration interval, and the configuration event is the configuration of the rule instances in each rule framework. In response to the time interval between the current time and the last configuration event reaching the configuration interval, execute the step of obtaining at least one target configuration instruction.
[0076] The configuration time interval can be a preset time for periodically configuring rule instances. For example, the configuration interval can be 5s. The last configuration event is the event of the last configuration of the rule instance. The above time interval is used to represent the time difference between the current time and the time when the last configuration of the rule instance occurred.
[0077] In response to the time interval between the current time and the last configuration event reaching the configuration interval, perform the above step S41 until the rule instances in each rule framework are configured and completed, and determine the target rule instances that need to be calculated with the above target fact data among the rule instances in each rule framework. In response to the time interval between the current time and the last configuration event not reaching the configuration interval, determine the target rule instances that need to be calculated with the above target fact data among the initial rule instances in each rule framework.
[0078] Exemplarily, the rule instances sent down need to be stored in a specified database, and the rule broadcaster will regularly pull the latest changed data. For example, the time interval for the rule broadcaster to pull the rules in the database is 5s. The rule broadcaster will broadcast the changed data to all rule frameworks. The changed data is at least one rule instance to be configured. In each rule framework, it will be determined whether to process each rule instance to be configured according to the rule type.
[0079] It can be considered that based on the configuration time interval, periodically modifying the initial rule instances in each rule framework can ensure that the target rule instances determined based on each rule framework later are rule instances that meet the requirements, and can ensure the accuracy of the analysis results of the fact data. In this way, only the rule instances under each fixed rule framework are modified, the operation is simple, and the efficiency of data analysis can be improved.
[0080] In some embodiments, before the above step S11, the above data analysis method further includes the following steps: obtaining at least one dimension data configuration instruction, each dimension data configuration instruction representing configuration information related to the dimension data to which each dimension data group belongs; parsing each dimension data configuration instruction to obtain the dimension data configuration information corresponding to each dimension data configuration instruction; and configuring each dimension data group based on each dimension data configuration information to obtain each dimension data in each dimension data group.
[0081] The configuration instructions for each dimension data represent the configuration information related to the dimension data to which each dimension data group belongs. Among them, the configuration instructions for each dimension data can be used to update or delete each dimension data in the dimension data group stored in each data analysis node, and to add new dimension data to the dimension data group stored in each data analysis node. The configuration instructions for each dimension data can be the instruction information obtained in response to the instruction configuration operation related to the dimension data on the display interface where the user is located. One dimension data configuration instruction corresponds to one dimension data to be configured. In some application scenarios, the dimension data configuration information includes the indication information for indicating the dimension data group to which the dimension data to be configured belongs. In some other application scenarios, the dimension data configuration information includes the indication information for indicating the configuration type of the dimension data to be configured. In some other application scenarios, the dimension data configuration information includes the indication information for indicating the configuration content of the dimension data to be configured.
[0082] The way to obtain the configuration instructions for each dimension data can be to pull the dimension data configuration information corresponding to the dimension data to be configured from the database at a preset frequency, or to directly receive the dimension data configuration information corresponding to the dimension data to be configured sent by the user. Among them, the database is used to store the dimension data configuration information corresponding to the dimension data to be configured sent by the user. Obtain the initial dimension data in each initial dimension data group on each data analysis node. The initial dimension data is the dimension data in the dimension data group stored in each data analysis node before configuration. Based on each dimension data configuration information, determine each dimension data to be configured. In some application scenarios, for each dimension data to be configured, in response to the successful matching of the identification information of the dimension data to be configured with any one of the initial dimension data, update or delete the successfully matched initial dimension data according to the dimension data to be configured. Thus, the dimension data in each dimension data group after configuration is obtained. In some other application scenarios, for each dimension data to be configured, in response to the failure of the identification information of the dimension data to be configured to match any one of the initial dimension data, determine whether the dimension data group to which the dimension data to be configured belongs exists, that is, determine whether the dimension data group to which the dimension data to be configured belongs is one of the initial dimension data groups. In response to the non-existence of the initial dimension data group to which the dimension data to be configured belongs, discard the dimension data to be configured and the dimension data configuration information corresponding to the dimension data to be configured. In response to the existence of the initial dimension data group to which the dimension data to be configured belongs, add the dimension data to be configured under the initial dimension data group to which the dimension data to be configured belongs. Thus, the dimension data in the dimension data group of each data analysis node after configuration is obtained.
[0083] It can be considered that when dimension data needs to be configured, only the dimension data in some dimension data groups is matched. Compared with matching and configuring all dimension data every time dimension data is configured, the present application can reduce the computational amount when dimension data needs to be configured, thereby improving the configuration efficiency of dimension data configuration.
[0084] In some embodiments, before the above step S11, the above data analysis method further includes the following steps: obtaining a plurality of dimension data related to the target scenario; dividing the plurality of dimension data to obtain each dimension data group; and storing each dimension data group to each data analysis node respectively.
[0085] The method of dividing the plurality of dimension data can be randomly dividing or polling dividing the plurality of dimension data based on the total number of each data analysis node. For example, random division can be randomly dividing each dimension data into each dimension data group so that the quantity of the dimension table data distributed to each dimension data group is relatively average. Polling division can be sequentially dividing each dimension data into each dimension data group according to any one attribute information included in the dimension data so that the quantity of the dimension table data distributed to each dimension data group is relatively average. Any one attribute information included in the dimension data can be the fixed value of the field of the identification information, and the dimension data with the same fixed value of the field in the identification information is divided into the same dimension data group. In polling division, through hash operation, it can be ensured that the quantity of the dimension table data distributed to each data analysis node is relatively average, and it is not necessary to ensure that the quantity of each dimension data in the dimension data group distributed to each data analysis node is equal. For each dimension data group, store the dimension data group to any one data analysis node.
[0086] It can be considered that by dividing the plurality of dimension data into each dimension data group and storing them to each data analysis node, compared with storing all dimension data, the present application disperses the dimension data to each data analysis node in advance, which can reduce the matching amount when performing matching between dimension data and fact data subsequently.
[0087] Please refer to Figure 6 , Figure 6 which is a framework schematic diagram of an embodiment of the data analysis method of the present application.
[0088] As Figure 6 shown, each machine corresponds to a data analysis node. For example, machines 1 to 6 respectively correspond to data analysis nodes 1 to 6. Each group of dimension data groups is stored on each machine or data analysis node. As Figure 6As shown in the figure, dimension data groups 1 to 6 are stored on machines 1 to 6 respectively. Each dimension data group includes at least one dimension data. The data distribution node distributes the fact data in the data stream to each machine. After determining in step S12 above that the target analysis node is the data analysis node 3 corresponding to machine 3, the target dimension data in the target dimension data group stored by the target analysis node is used to supplement the fact data to obtain the target fact data. The data analysis node 3, that is, the target analysis node, includes two target rule frameworks in the corresponding computing framework, namely rule framework 1 and rule framework 2. Among them, each target rule framework has a total of three target rule instances. It can be understood that the rule instances belonging to different rule frameworks are different. For example, Figure 6 the rule instance 1 in rule framework 1 in Figure 6 and the rule instance 1 in rule framework 2 are different rule instances. After the data analysis system is started, in response to the completion of the loading of each data analysis node, the target fact data will be copied into two copies and connected to two different rule frameworks. Similarly, each dimension data group on the target analysis node will be copied into two copies and connected to each rule framework under the target analysis node for subsequent use by each rule box when executing the calculation logic.
[0089] After step S13 above, the target fact data obtained from the fact data in the data stream enters each rule framework. The target fact data is calculated with each target rule instance, and the calculation result, that is, the calculation result corresponding to each target rule instance, and the calculation results corresponding to each target rule instance are merged into the result aggregation of the target rule framework to which each target rule instance belongs. The data at each target rule framework aggregation point is aggregated again, and then can be sent to the result collector in the data analysis system.
[0090] In the above solution, the data distribution node distributes the fact data related to the target scenario to each data analysis node. Each data analysis node stores the dimension data groups obtained by dividing a number of dimension data related to the target scenario. Each data analysis node matches the received fact data with the dimension data groups stored by itself, which can reduce the amount of data that each data analysis node needs to match. The data analysis node where the dimension data group that successfully matches the fact data is located is the target analysis node. The target analysis node analyzes the fact data according to the preset analysis method corresponding to the target analysis node to obtain the analysis result of the fact data. Compared with the situation where the fact data is only matched with all dimension data on the same data analysis node and then the fact data is analyzed later, which causes an excessive load on the data analysis node, this application can relieve the load pressure of the data analysis node, thereby improving the data analysis efficiency of the fact data.
[0091] Please refer to Figure 8 , Figure 8It is a schematic structural diagram of an embodiment of the data analysis system of the present application. The data analysis system 70 includes a data distribution node 71 and a plurality of data analysis nodes 72. Each data analysis node 72 is communicatively connected to the data distribution node 71 to implement the steps in the embodiment of the above data analysis method. In some application scenarios, different data analysis nodes 72 may be in different terminal devices. The data distribution node 71 and any one of the data analysis nodes 72 may be in the same terminal device. In some other application scenarios, different data analysis nodes 72 may be in the same terminal device and are executed by different threads. In some other application scenarios, the data distribution node 71 and any one of the data analysis nodes 72 may not be in the same terminal device. In some other application scenarios, the data distribution node 71 and a plurality of data analysis nodes 72 may be in the same terminal device and are executed by different threads.
[0092] In the above solution, the data distribution node distributes fact data related to the target scenario to each data analysis node. Each data analysis node stores a dimension data group obtained by dividing a plurality of dimension data related to the target scenario. Each data analysis node respectively matches the received fact data with the dimension data group stored therein, which can reduce the amount of data that each data analysis node needs to match. The data analysis node where the dimension data group that successfully matches the fact data is located is the target analysis node. The target analysis node analyzes the fact data according to the preset analysis method corresponding to the target analysis node to obtain the analysis result of the fact data. Compared with the situation where the fact data is only matched with all dimension data on the same data analysis node and then the fact data is analyzed subsequently, which causes an excessive load on the data analysis node, the present application can relieve the load pressure of the data analysis node, thereby improving the data analysis efficiency of the fact data.
[0093] Please refer to Figure 8 , Figure 8 It is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. The computer-readable storage medium 80 stores program instructions 801 thereon. When the program instructions 801 are executed by a processor, the steps in any embodiment of the above data analysis method are implemented. Among them, the computer-readable storage medium 80 can be applied to one or more of the above data distribution nodes and the above plurality of data analysis nodes.
[0094] In the above solution, the data distribution node distributes fact data related to the target scenario to each data analysis node. Each data analysis node stores a dimension data group obtained by dividing several dimension data related to the target scenario. Each data analysis node respectively matches the received fact data with the dimension data group stored in it, which can reduce the amount of data that each data analysis node needs to match. The data analysis node where the dimension data group successfully matched with the fact data is located is the target analysis node. The target analysis node analyzes the fact data according to the preset analysis method corresponding to the target analysis node to obtain the analysis result of the fact data. Compared with the situation where the fact data is only matched with all dimension data on the same data analysis node and the subsequent analysis of the fact data causes an excessive load on the data analysis node, this application can relieve the load pressure on the data analysis node, thereby improving the data analysis efficiency of the fact data.
[0095] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the method embodiments above. The specific implementation can refer to the description of the method embodiments above. For the sake of brevity, it will not be repeated here.
[0096] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated in this article.
[0097] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation manners described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0098] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0099] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods according to various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
Claims
1. A data analysis method, characterized in that, The method is applied to a data analysis system, which includes a number of data analysis nodes and a data distribution node. The method includes: The data distribution node distributes fact data related to a target scenario to each of the data analysis nodes. Each data analysis node stores a dimension data group, and each dimension data group is obtained by dividing a number of dimension data related to the target scenario; Each of the data analysis nodes respectively matches the received fact data with the dimension data groups stored in it. Among them, the data analysis node where the dimension data group that successfully matches the fact data is located is the target analysis node; The target analysis node analyzes the fact data according to the preset analysis method corresponding to the target analysis node to obtain the analysis result of the fact data.
2. The method according to claim 1, wherein Each dimension data carries analysis parameters, and the analysis parameters carried by each dimension data are associated with the computing framework to which the dimension data group corresponding to each dimension data belongs. The dimension data group stored in the target analysis node is the target dimension data group. The analysis result of the fact data includes the calculation result of the fact data. The target analysis node analyzes the fact data according to the preset analysis method corresponding to the target analysis node to obtain the analysis result of the fact data, including: Supplement the analysis parameters recorded in the target dimension data group to the fact data to obtain target fact data; Calculate the target fact data according to the computing framework to which the target dimension data group belongs to obtain the calculation result of the fact data.
3. The method according to claim 2, wherein The computing framework to which the target dimension data group belongs includes at least one target rule framework, and each target rule framework includes at least one target rule instance. The calculation result of the fact data includes the calculation results corresponding to each target rule framework. Calculating the target fact data according to the computing framework to which the target dimension data group belongs to obtain the calculation result of the fact data includes: For each of the target rule frameworks, the following steps are performed: Calculate the target fact data according to each target rule instance in the target rule framework to obtain the calculation results corresponding to each target rule instance; Use the calculation results corresponding to each target rule instance as the calculation result corresponding to the target rule framework.
4. The method according to claim 3, wherein Before calculating the target fact data according to the computing framework to which the target dimension data group belongs to obtain the calculation result of the fact data, the method further includes: Obtain at least one target configuration instruction, and each target configuration instruction represents configuration information related to the rule instances in each rule framework; Analyze each target configuration instruction to obtain the target configuration information corresponding to each target configuration instruction; Configure the initial rule instances in each rule framework based on each target configuration information to obtain each rule instance.
5. The method according to claim 4, wherein The configuring the initial rule instances in each rule framework based on each target configuration information to obtain each rule instance includes: For each of the target configuration information, the following steps are performed: Based on the target configuration information, determine the rule framework indication information, configuration type, and configuration content. The rule framework indication information is used to represent the rule framework to which the rule instance to be configured belongs, and the configuration type includes one of addition, deletion, and update; In response to the existence of the rule framework corresponding to the rule framework indication information, configure the rule framework corresponding to the rule framework indication information based on the configuration content and the configuration type to obtain the rule instance; or, In response to the non-existence of the rule framework corresponding to the rule framework indication information, discard the target configuration information.
6. The method according to claim 4, wherein The method further includes: Obtain a configuration interval; Determine whether the time interval between the current time and the last configuration event reaches the configuration interval, where the configuration event is the configuration of the rule instances in each of the rule frameworks; In response to the time interval between the current time and the last configuration event reaching the configuration interval, execute the step of obtaining at least one target configuration instruction.
7. The method according to any one of claims 1 to 6, characterized in that, Before the data distribution node distributes the fact data related to the target scenario to each of the data analysis nodes, the method further includes: Obtain at least one dimension data configuration instruction, and each of the dimension data configuration instructions represents the configuration information related to the dimension data to which each dimension data group belongs; Parse each of the dimension data configuration instructions to obtain the dimension data configuration information corresponding to each of the dimension data configuration instructions; Configure each of the dimension data groups based on each of the dimension data configuration information to obtain each dimension data in each of the dimension data groups.
8. The method according to any one of claims 1 to 6, characterized in that, Before the data distribution node distributes the fact data related to the target scenario to each of the data analysis nodes, the method further includes: Obtain a plurality of dimension data related to the target scenario; Divide the plurality of dimension data to obtain each of the dimension data groups; Store each of the dimension data groups in each of the data analysis nodes respectively.
9. A data analysis system, characterized in that, Includes: The data analysis system includes a plurality of data analysis nodes and a data distribution node, and each of the data analysis nodes is communicatively connected to the data distribution node to implement the method according to any one of claims 1-8.
10. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by a processor, they are used to implement the method according to any one of claims 1-8.