Data analysis method and device, medium, equipment and product
By analyzing the dependencies of database tables and field redundancy, determining the reasons that affect the timeliness of data outputs, solving the problem of difficulty in effectively optimizing the timeliness of data outputs in the existing technology, and achieving more accurate basis and more efficient data processing.
Patent Information
- Application Number
- CN202510169049.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-13
AI Technical Summary
In big data application scenarios, the timeliness of data output directly affects the timeliness and accuracy of business decisions, and the reasons that are difficult to effectively locate and optimize the impact of data output timeliness.
By determining the first database table, the table includes all or part of the database table in the dependent database table, and based on the information of the database table with a dependency relationship with the first database table, the reason information affecting the output time of the first data is determined, indicating that there is a dependency exception and/or field redundancy.
It can accurately locate issues that affect the timeliness of data output, thereby providing an accurate basis for optimizing data output time and improving data output timeliness.
Smart Images

Figure CN119988384A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of database technology, and in particular, to a data analysis method, device, medium, equipment and product. Background Art
[0002] In big data application scenarios, it is necessary to ensure the normal output of business data so that users can understand the business operation status. Since the timeliness of data output directly affects the timeliness and accuracy of business decisions, it is crucial for business operation and maintenance. Therefore, the requirements for the timeliness of data output are getting higher and higher. In order to make data ready and output earlier, it is very important to locate the reasons that affect the timeliness of data output. Summary of the invention
[0003] This summary is provided to introduce concepts in a brief form that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] In a first aspect, the present disclosure provides a data analysis method, the method comprising: In response to receiving an analysis request, determining a first database table, wherein the analysis request is used to request an analysis of a cause affecting the production time of the first data, the first database table includes all or part of the database tables in the dependent database tables, and the dependent database tables include upstream database tables directly dependent on the first data and upstream database tables indirectly dependent on the first data; Based on first information, reason information affecting the production time of the first data is determined, wherein the first information includes information of a database table having a dependency relationship with the first database table, and the reason information is used to indicate that there is a dependency abnormality and / or field redundancy in the first database table.
[0005] In a second aspect, the present disclosure provides a data analysis device, the device comprising: A first determination module is configured to determine a first database table in response to receiving an analysis request, wherein the analysis request is used to request an analysis of a cause affecting the production time of the first data, and the first database table includes all or part of the database tables in the dependent database tables, and the dependent database tables include the upstream database tables directly dependent on the first data and the upstream database tables indirectly dependent on the first data; The second determination module is used to determine reason information affecting the production time of the first data based on the first information, wherein the first information includes information of a database table having a dependency relationship with the first database table, and the reason information is used to indicate that there is a dependency abnormality and / or field redundancy in the first database table.
[0006] In a third aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the data analysis method provided in the first aspect of the present disclosure.
[0007] In a fourth aspect, the present disclosure provides an electronic device, including: a storage device having a computer program stored thereon; A processing device is used to execute the computer program in the storage device to implement the steps of the data analysis method provided in the first aspect of the present disclosure.
[0008] In a fifth aspect, the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the data analysis method provided in the first aspect of the present disclosure.
[0009] Through the above technical solution, the first database table includes all or part of the database tables in the dependent database table, and the dependent database table includes the upstream database table that the first data directly depends on and the upstream database table that it indirectly depends on. Since the output time of the upstream database table directly affects the output time of the downstream database table, the first database table can affect the output time of the first data. According to the information of the database table that has a dependency relationship with the first database table, the cause information affecting the output time of the first data is determined, and the cause information is used to indicate that the first database table has a dependency abnormality and / or field redundancy. In this way, the problem of the first database table that affects the output time of the first data can be located, so as to provide an accurate basis for optimizing the output time of the first data and improving the output timeliness of the first data.
[0010] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 The present invention is a flow chart of a data analysis method according to an exemplary embodiment.
[0012] Figure 2 The diagram is a schematic diagram showing a dependency relationship between database tables according to an exemplary embodiment.
[0013] Figure 3 The diagram is a hierarchical diagram of a data warehouse according to an exemplary embodiment.
[0014] Figure 4 It is a block diagram of a data analysis device according to an exemplary embodiment.
[0015] Figure 5 A schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0016] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0017] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0018] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0019] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0020] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0022] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0023] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0024] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0025] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0026] At the same time, it is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0027] Figure 1 is a flow chart of a data analysis method according to an exemplary embodiment. The method can be applied to an electronic device with processing capabilities, such as a server, Figure 1 As shown, the method may include step 11 and step 12.
[0028] In step 11 , in response to receiving an analysis request, a first database table is determined.
[0029] The analysis request is used to request an analysis of the reasons affecting the production time of the first data. For example, the first data may be sales data such as sales volume and sales volume, or customer data such as customer repurchase rate and number of customers. The present disclosure does not limit the type of the first data.
[0030] In one embodiment, a user may input a task ID and an analysis instruction through a client, where the task ID is the ID of a task for generating first data. The client may send the analysis instruction including the task ID to a server. The server may receive the analysis request and, based on the task ID, identify the analysis request for requesting an analysis of the reasons that affect the production time of the first data.
[0031] After receiving the analysis request, the server may determine a first database table, which includes all or part of the dependent database tables, and the dependent database tables include upstream database tables directly and indirectly dependent on the first data.
[0032] Figure 2 FIG. 1 is a schematic diagram showing a dependency relationship between database tables according to an exemplary embodiment. Figure 2 As shown, the solid-line box can be regarded as a data processing task node. The data processing task generates corresponding data or database tables, and the time information therein indicates the output time. For example, the output time of the first data is 9:30, and the output time of database table A is 8:30. The output time of the database table refers to the time when all field data in the database table have been output. The direction indicated by the arrow is downstream. Specifically, the upstream database tables that the first data directly depends on include database table A and database table M, that is, the first data needs to be calculated based on the data in database table A and the data in database table M. The upstream database tables that database table A directly depends on include database table B and database table J, that is, the data in database table A needs to be calculated based on the data in database table B and the data in database table J. The dependency relationship between other database tables is similar. Among them, the downstream can only use the data in the upstream database table for calculation after the data in the upstream database table is fully ready for output.
[0033] by Figure 2 Taking the dependency relationship shown in FIG. 1 as an example, the dependent database tables include database table A, database table B, database table C, database table D, database table E, database table F, database table G, database table H, database table J, database table K, database table M, and database table N. The upstream database tables that the first data indirectly depends on include database table B, database table C, database table D, database table E, database table F, database table G, database table H, database table J, database table K, and database table N.
[0034] In step 12, cause information affecting the production time of the first data is determined according to the first information.
[0035] The first information includes information of a database table having a dependency relationship with the first database table, and the reason information is used to indicate that there is a dependency abnormality and / or field redundancy in the first database table.
[0036] Figure 2 The link formed by the dependency relationship shown can be regarded as a data production link for generating the first data. In the related art, the data calculation speed is improved by identifying tasks with long execution time in the link, that is, database tables with long calculation time, adjusting the parameters of the CPU and other resources at runtime by adjusting the database engine parameters, or by increasing the parallelism of task execution and the amount of resources. However, the method in the related art mainly increases the speed of data output by adjusting database resources, but it may bring new problems, such as resource competition, system instability and other risks.
[0037] In the present disclosure, considering that the dependency abnormality between database tables will affect the data production time, database table field redundancy, that is, there are too many fields in the database table, and it takes a longer time to generate data for these fields, will also affect the data production time. Therefore, based on the information of the database table that has a dependency relationship with the first database table, the reason information affecting the production time of the first data is determined, and the reason information is used to indicate that the first database table has a dependency abnormality and / or field redundancy.
[0038] In this way, the technician can use the cause information as a reference to optimize the first database table in a targeted manner, so as to achieve the purpose of improving the timeliness of the first data output. For example, if there is an abnormal dependency relationship in the first database table, the first data can be produced earlier by adjusting the dependency relationship of the first database table. If there is field redundancy in the first database table, the first data can be produced earlier by adjusting the fields stored in the first database table, or splitting the first database table into multiple database tables and then adjusting the dependency relationship of the new database tables.
[0039] Through the above technical solution, the first database table includes all or part of the database tables in the dependent database table, and the dependent database table includes the upstream database table that the first data directly depends on and the upstream database table that it indirectly depends on. Since the output time of the upstream database table directly affects the output time of the downstream database table, the first database table can affect the output time of the first data. According to the information of the database table that has a dependency relationship with the first database table, the cause information affecting the output time of the first data is determined, and the cause information is used to indicate that the first database table has a dependency abnormality and / or field redundancy. In this way, the problem of the first database table that affects the output time of the first data can be located, so as to provide an accurate basis for optimizing the output time of the first data and improving the output timeliness of the first data.
[0040] In one implementation, all database tables in the dependent database tables may be used as first database tables.
[0041] In another implementation, the implementation of determining the first database table may be: The upstream database table that the first data directly depends on and that is produced the latest is used as the first database table determined for the first time; Repeat the step of using the upstream database table on which the first database table determined last time directly depends and whose output time is the latest as the first database table determined this time, until the number of the first database tables determined this time reaches a first preset threshold, or the first database table determined this time has no upstream database table on which it is directly dependent.
[0042] by Figure 2 Taking the dependency relationship shown as an example, the output time of the first data is 9:30, and the upstream database tables that the first data directly depends on include database tables A and M, among which the output time of table A is 8:30, and the output time of table M is 0:45. The upstream database table that the first data directly depends on and has the latest output time is database table A, and database table A is used as the first database table determined for the first time. Afterwards, the upstream database tables that database table A directly depends on include database tables B and J, and the output time of database table B is 7:30, and the output time of database table J is 4:00. Database table B is used as the first database table determined this time. Afterwards, the upstream database table that database table B directly depends on is table C, and table C is used as the first database table determined this time. Afterwards, the upstream database table that database table C directly depends on is table D, and table D is used as the first database table determined this time. The database tables that table D directly depends on include tables E and N. The output time of table E is 4:00, and the output time of table N is 4:30. Table N is the first database table determined this time. Table N does not have any upstream database table that it directly depends on. Figure 2According to the dependency relationship shown, the first database tables determined include tables A, B, C, D, and N.
[0043] In one embodiment, if the generation time of database table N is 3:00, the first preset threshold is, for example, 10, and the determined first database table includes tables A, B, C, D, E, F, G, and H. In one embodiment, if the generation time of database table N is 3:00, the first preset threshold is, for example, 5, and the determined first database table includes tables A, B, C, D, and E.
[0044] Through the above technical solution, since the database table with the latest output time directly affects the output time of its downstream, the determined first database table directly affects the output time of the first data, narrowing the scope of cause analysis and making the optimization of the output time of the first data more targeted.
[0045] In one embodiment, the first information includes the layer where the first upstream database table is located in the data warehouse, the first upstream database table is the upstream database table that the first database table directly depends on, the data warehouse includes multiple data layers, and a data layer with a higher layer depends on a data layer with a lower layer; in step 12, determining the reason information affecting the output time of the first data according to the first information may include: If the first database table is located in the first data layer in the data warehouse, the first upstream database table is located in the second data layer in the data warehouse, and the level of the second data layer is higher than that of the first data layer, it is determined that the first database table has a dependency abnormality.
[0046] Figure 3 is a hierarchical diagram of a data warehouse according to an exemplary embodiment. Figure 3As shown in the figure, the data warehouse includes multiple data layers, including: Operational Data Store (ODS), Data Warehouse Detail (DWD), Data Warehouse Middle (DWM), Data Mart (DM) and Application (APP). Among them, the ODS layer is the lowest layer in the data warehouse architecture, which retains the original data in the source system. The DWD layer is used to perform preliminary cleaning, standardization, format conversion, deduplication and other processing on the data in the ODS layer. The DWM layer is used to perform further aggregation operations on the data in the DWD layer, such as data aggregation according to time periods (days, months). The DM layer is used to extract and summarize data from other layers to meet the data analysis of specific business departments (such as sales departments, financial departments, etc.). The APP layer provides users with intuitive data display and analysis, such as reports, dashboards, etc. The data warehouse also includes a dimension layer (DIM), which mainly stores dimension data. For example, the customer dimension includes information such as name and age, and provides dimension reference for the fact table. Figure 2 In the dependency relationship shown, the box also indicates the layer where the database table is located in the data warehouse, for example, database table A is located in the DWM layer in the data warehouse, and database table M is located in the ODS layer in the data warehouse.
[0047] like Figure 3 As shown, among the multiple data layers, the arrows point to the data layers of the higher level, and the data layers of the higher level depend on the data layers of the lower level. Among them, the solid arrows indicate that they can be directly dependent, indicating that the dependency relationship is normal, that is, the DWD layer depends on the ODS layer, the DWM layer depends on the DWD layer, the DM layer depends on the DWM layer, the APP layer depends on the DM layer, and the APP layer depends on the DWM layer. The dotted arrows indicate that they can be dependent under special circumstances, which also indicates that the dependency relationship is normal, that is, the DM layer depends on the DWD layer, and the APP layer depends on the DWD layer. In addition, except for the ODS layer and the APP layer, the database tables in other layers can depend on the same level. This situation also belongs to a normal dependency relationship. For example, the database tables in the DWD layer can depend on other database tables in the DWD layer. For the case of cross-layer dependencies, such as Figure 2 The table M in the ODS layer, on which the first data shown directly depends, generally affects the accuracy of the data, but has little impact on the timeliness of data output.
[0048] If a database table in a lower-level data layer depends on a database table in a higher-level data layer, the dependency relationship is abnormal. This is because the database tables in the higher-level data layer often depend on more downstream database tables and have a later output time, which will cause the output time of the lower-level database tables to be later. This situation can be expressed as reverse dependency, that is, the dependency direction is inconsistent with the normal data flow direction.
[0049] by Figure 2 Taking the dependency relationship shown in the figure as an example, Table A is the first database table, and Table A is located in the DWM layer in the data warehouse, that is, the first data layer is the DWM layer. The upstream database tables that Table A directly depends on include Table B and Table J. Among them, the layer where Table B is located in the data warehouse is the APP layer, that is, the second data layer is the APP layer. Since the level of the APP layer is higher than that of the DWM layer, it can be determined that Table A has a dependency abnormality. Figure 2 It can also be seen from the output time of the database tables shown that database table J was produced at 4:00, while table B is located at a higher level and has a later output time, that is, table B was not produced until 7:30. In order to wait for the data of table B, the output time of database table A is relatively late, thus affecting the output time of the first data.
[0050] If there are multiple first upstream database tables, as long as the level of the layer where one of the first upstream database tables is located is higher than the level of the first data layer, it can be determined that the first database table has a dependency anomaly.
[0051] In one embodiment, in step 12, determining the reason information affecting the production time of the first data according to the first information may also include: If the first database table is located in the dimension layer in the data warehouse, the layer where the first upstream database table is located in the data warehouse is the second data layer, and the second data layer is the data application layer or the data mart layer, it is determined that the first database table has a dependency abnormality.
[0052] like Figure 3 As shown, for the dimension layer, the solid arrow indicates that it can be directly dependent, indicating that the dependency relationship is normal, that is, the DWD layer depends on the DIM layer, the DWM layer depends on the DIM layer, the DM layer depends on the DIM layer, and the DIM layer depends on the ODS layer. The dotted arrow indicates that it can be dependent under special circumstances, which also indicates that the dependency relationship is normal, that is, the APP layer depends on the DIM layer. In addition, the DIM layer depends on the DWM layer, and the DIM layer depends on the DWD layer, which are also normal dependencies.
[0053] The DIM layer depends on the DM layer or the APP layer, which is a dependency abnormality. This is because the data in the DM and APP layers are aggregated data, which is generally used to meet the needs of specific departments or dashboards and is produced later. Therefore, if the first database table is located in the DIM layer and the first upstream database table is located in the DM layer or the APP layer, the production time of the first database table may be affected due to the late production time of the first upstream database table. It can be determined that the first database table has a dependency abnormality. This situation can also be expressed as reverse dependency, that is, the dependency direction is inconsistent with the normal data flow direction.
[0054] If there are multiple first upstream database tables, as long as one of the first upstream database tables is located in the DM layer or the APP layer, it can be determined that there is a dependency abnormality in the first database table.
[0055] For example, the server can output the analysis results to the client, and the analysis results can include the layer where the first database table is located in the data warehouse, the layer where the first upstream database table is located, etc. If the first database table has a dependency abnormality indicating a reverse dependency, the technician can see through the client that the label of the first database table is a dependency abnormality, and indicates that there is a reverse dependency. The technician can build a new database table in the appropriate data layer according to the fields in the upstream database table that the first database table depends on, and adjust the dependency of the first database table so that the first database table depends on the new database table, thereby avoiding the dependency abnormality. For example, Table A in the DWM layer depends on Table B in the APP layer. The technician, for example, builds a new database table in the DWM layer according to the fields in Table B that Table A depends on, and then makes Table A depend on the new database table.
[0056] Through the above technical solution, it is possible to determine whether there is a dependency anomaly in the first database table based on the layer at which the first database table is located in the data warehouse and the layer at which the first upstream database table is located in the data warehouse. If there is a dependency anomaly, it may affect the production time of the first database table, thereby affecting the production time of the first data. In this way, the problem of the first database table that affects the production time of the first data can be accurately located.
[0057] In one embodiment, the first information includes the layer where each dependent database table is located in the data warehouse; in step 12, determining the reason information affecting the output time of the first data according to the first information may include: If the first database table is a database table in the first link, it is determined that there is a dependency abnormality in the first database table, wherein the first link is composed of dependent database tables with continuous dependency relationships, the dependent database tables constituting the first link are located at the same layer in the data warehouse, and the number of dependent database tables in the first link is greater than or equal to a second preset threshold.
[0058] like Figure 2 As shown in the figure, the layers where each dependent database table is located in the data warehouse are marked in the figure. Figure 2 The database tables in the dotted box, Table D, Table E, Table F and Table G, are all located in the DWD layer in the data warehouse, and the dependency relationships of Table D, Table E, Table F and Table G are continuous, that is, Table D directly depends on Table E, Table E directly depends on Table F, and Table F directly depends on Table G. The second preset threshold is, for example, 4, then Table D, Table E, Table F and Table G constitute the first link, and Table D, as the first database table, is located in the first link. Therefore, Table D has a dependency exception.
[0059] Among them, the database tables in the first link all have the problem of abnormal dependencies. The analysis results output by the server include all the database tables in the first link and the layer where the database tables are located in the data warehouse. Taking the first database tables including tables A, B, C, D, and N as an example, the technician can see through the client that the label of the first database table D is abnormal dependency. The technician can also view the analysis results through the client to know that the first link includes tables D, E, F, and G.
[0060] The number of dependent database tables in the first link is greater than or equal to the second preset threshold value, and the dependent database tables constituting the first link are located at the same layer in the data warehouse, which can indicate that the data processing tasks that need to be completed at the same layer are dispersed across multiple database tables, and the dependency depth of database tables at the same layer is too high, resulting in a too long data production link, which affects the timeliness of data output. In the case where the dependency depth at the same layer is too high, technicians can, for example, integrate the data into one or two database tables based on the storage logic of multiple database tables, shorten the data production link, and thus improve the timeliness of data output.
[0061] It should be noted that Figure 2 In the embodiment shown, Table D, Table E, Table F and Table G are all located in the DWD layer in the data warehouse. This is only an example. The present disclosure does not limit the layer where the database table in the first link is located. For example, the database table in the first link may also be located in the DWM layer or the DM layer.
[0062] Through the above technical solution, if the first database table is a database table in the first link, it can indicate that the dependency relationship of the first database table is abnormal, thereby affecting the production time of the first data, and the problem of the first database table that affects the production time of the first data can be accurately located.
[0063] In one embodiment, the first information includes the first field quantity, the first field quantity is the quantity of fields in the first database table on which the first downstream database table depends, and the first downstream database table is a downstream database table that directly depends on the first database table; in step 12, determining the cause information affecting the output time of the first data according to the first information may include: Determine ratio information, where the ratio information is the ratio of the number of first fields to the number of second fields, and the number of second fields is the total number of fields in the first database table; According to the proportion information, it is determined whether there is field redundancy in the first database table.
[0064] Among them, the downstream database table may rely on the data of all fields in the upstream database table, or may only rely on the data of some fields in the upstream database table. According to the proportion information, the field reference rate in the first database table can be quantified. If the number of the first downstream database table is one, then directly determine whether there is field redundancy in the first database table based on the proportion information. For example, when the proportion information is less than the third preset threshold, it means that the field reference rate in the first database table is relatively low, and it can be determined that there is field redundancy in the first database table. If the number of the first downstream database tables is multiple, the average value of the proportion information corresponding to each first downstream database table is used as the average proportion information. If the average proportion information is less than the third preset threshold, it means that the average reference rate of the fields in the first database table is relatively low, and it can be determined that there is field redundancy in the first database table.
[0065] It should be noted that the first downstream database table can be a dependent database table or a database table other than the dependent database table. Taking the first database table C as an example, table B is the first downstream database table. For example, there is another database table Q (not shown in the figure). Table Q directly depends on table C. Table Q is not used to generate the first data and is not a dependent database table, but table Q serves as the first downstream database table.
[0066] If the first database table has field redundancy, because all the field data in the database table must be produced before the database table can be ready and produced, and the downstream can use the data in the first database table for calculation, not only the production time of the first database table is relatively late, but also affects the production time of the downstream database table, thereby affecting the production time of the first data. For example, for example, Table B depends on Field a in Table C, and the data of Field a has been produced at 5:30, but Table B needs to wait until all the fields in Table C have been produced at 6:30 before it can use the data of Field a in Table C for calculation, which causes the production time of Table B to be relatively late, thereby affecting the production time of the first data.
[0067] Among them, if it is determined that there is field redundancy in the first database table, the server can also output to the client which fields in the first database table the first downstream database table specifically depends on, as well as the specific output time of these fields, and the upstream database table on which these fields depend, which can provide a reference for the technician to adjust the dependency relationship of the first database table. For example, if there is field redundancy in the first database table, the technician can, for example, split the first database table into multiple database tables based on the fields of the first database table that the downstream database table depends on, and adjust the dependency relationship of the downstream database table so that the downstream database table depends on the split database table, thereby avoiding the problem of late output time. In addition, for example, in the case where Table B depends on Field a in Table C, the technician can also adjust Table B to directly depend on the upstream database table of Field a, thereby improving the timeliness of data output.
[0068] Through the above technical solution, whether there is field redundancy in the first database table is determined according to the ratio of the number of first fields to the number of second fields, which can provide an accurate basis for optimizing the output time of the first data and improving the timeliness of the output of the first data.
[0069] For example, the first database table may only have a dependency anomaly indicating reverse dependency, such as database table A. The first database table may also only have a dependency anomaly indicating that the same-layer dependency depth is too high, such as database table D. The first database table may also have a normal dependency relationship and only have a field redundancy problem, such as database table C. In addition, the first database table may also have both dependency anomaly and field redundancy problems. For example, database table E is the first database table and has field redundancy. Table E has a dependency anomaly indicating that the same-layer dependency depth is too high, and the first database table E has both dependency anomaly and field redundancy.
[0070] It should be noted that the server may pre-store a data production link for generating the first data, and meta-information of each database table. The meta-information may include the layer of the database table in the data warehouse, which fields are stored in the database table, the number of fields stored in the database table, which fields in the upstream database table the database table depends on, which fields the downstream database table depends on, and other information.
[0071] Based on the same inventive concept, the present disclosure also provides a data analysis device, Figure 4 is a block diagram of a data analysis device according to an exemplary embodiment. Figure 4 As shown, the device 40 may include: A first determination module 41 is configured to determine a first database table in response to receiving an analysis request, wherein the analysis request is used to request an analysis of a cause affecting the production time of the first data, and the first database table includes all or part of the database tables in the dependent database tables, and the dependent database tables include the upstream database tables directly dependent on the first data and the upstream database tables indirectly dependent on the first data; The second determination module 42 is used to determine reason information affecting the production time of the first data based on the first information, wherein the first information includes information of a database table having a dependency relationship with the first database table, and the reason information is used to indicate that there is a dependency abnormality and / or field redundancy in the first database table.
[0072] Optionally, the first determining module 41 is configured to: The upstream database table on which the first data is directly dependent and which is produced the latest is used as the first database table determined for the first time; Repeat the step of using the upstream database table on which the first database table determined last time directly depends and whose output time is the latest as the first database table determined this time, until the number of the first database tables determined this time reaches a first preset threshold, or the first database table determined this time has no upstream database table on which it is directly dependent.
[0073] Optionally, the first information includes a layer where the first upstream database table is located in the data warehouse, the first upstream database table is an upstream database table that the first database table directly depends on, the data warehouse includes multiple data layers, and a data layer at a higher level depends on a data layer at a lower level; The second determining module 42 includes: The first determination submodule is used to determine that the first database table has a dependency abnormality if the first database table is located in the first data layer in the data warehouse, the layer where the first upstream database table is located in the data warehouse is the second data layer, and the level of the second data layer is higher than the level of the first data layer.
[0074] Optionally, the data warehouse further includes a dimension layer, and the multiple data layers include a data mart layer and a data application layer; The second determining module 42 further includes: The second determination submodule is used to determine that the first database table has a dependency abnormality if the first database table is located in the dimensional layer in the data warehouse, the layer where the first upstream database table is located in the data warehouse is the second data layer, and the second data layer is the data application layer or the data mart layer.
[0075] Optionally, the first information includes the layer where each of the dependent database tables is located in the data warehouse; The second determining module 42 includes: The third determination submodule is used to determine that the first database table has a dependency abnormality if the first database table is a database table in a first link, wherein the first link is composed of the dependent database tables with continuous dependency relationships, the dependent database tables constituting the first link are located at the same layer in the data warehouse, and the number of the dependent database tables in the first link is greater than or equal to a second preset threshold.
[0076] Optionally, the first information includes a first field quantity, where the first field quantity is the quantity of fields in the first database table on which the first downstream database table depends, and the first downstream database table is a downstream database table that directly depends on the first database table; The second determining module 42 includes: A fourth determination submodule is used to determine proportion information, where the proportion information is a ratio of the first field quantity to the second field quantity, where the second field quantity is the total quantity of fields in the first database table; The fifth determining submodule is used to determine whether there is field redundancy in the first database table according to the proportion information.
[0077] Optionally, the fifth determining submodule includes: a sixth determining submodule, configured to, if there are multiple first downstream database tables, take an average value of the proportion information corresponding to each of the first downstream database tables as average proportion information; The seventh determination submodule is configured to determine that field redundancy exists in the first database table if the average proportion information is less than a third preset threshold.
[0078] Reference below Figure 5 , which shows a structural schematic diagram of an electronic device 600 suitable for implementing an embodiment of the present disclosure. Figure 5 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0079] like Figure 5As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0080] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 5 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0081] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0082] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than computer readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0083] In some embodiments, the servers may communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0084] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0085] The computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device: in response to receiving an analysis request, determines a first database table, wherein the analysis request is used to request an analysis of a cause affecting the output time of the first data, the first database table includes all or part of the database tables in the dependent database tables, and the dependent database tables include upstream database tables directly dependent on the first data and upstream database tables indirectly dependent on the first data; Based on first information, reason information affecting the production time of the first data is determined, wherein the first information includes information of a database table having a dependency relationship with the first database table, and the reason information is used to indicate that there is a dependency abnormality and / or field redundancy in the first database table.
[0086] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0087] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0088] The modules involved in the embodiments described in the present disclosure may be implemented by software or hardware. The name of a module does not limit the module itself in some cases. For example, the first determination module may also be described as a "module for determining a first database table".
[0089] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0090] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0091] According to one or more embodiments of the present disclosure, Example 1 provides a data analysis method, the method comprising: In response to receiving an analysis request, determining a first database table, wherein the analysis request is used to request an analysis of a cause affecting the production time of the first data, the first database table includes all or part of the database tables in the dependent database tables, and the dependent database tables include upstream database tables directly dependent on the first data and upstream database tables indirectly dependent on the first data; Based on first information, reason information affecting the production time of the first data is determined, wherein the first information includes information of a database table having a dependency relationship with the first database table, and the reason information is used to indicate that there is a dependency abnormality and / or field redundancy in the first database table.
[0092] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1, wherein determining the first database table includes: The upstream database table on which the first data is directly dependent and which is produced the latest is used as the first database table determined for the first time; Repeat the step of using the upstream database table on which the first database table determined last time directly depends and whose output time is the latest as the first database table determined this time, until the number of the first database tables determined this time reaches a first preset threshold, or the first database table determined this time has no upstream database table on which it is directly dependent.
[0093] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 1, wherein the first information includes a layer at which a first upstream database table is located in a data warehouse, the first upstream database table is an upstream database table that the first database table directly depends on, and the data warehouse includes multiple data layers, and a data layer at a higher level depends on a data layer at a lower level; The determining, according to the first information, the cause information affecting the output time of the first data includes: If the first database table is located in the first data layer in the data warehouse, the layer where the first upstream database table is located in the data warehouse is the second data layer, and the level of the second data layer is higher than that of the first data layer, it is determined that there is a dependency abnormality in the first database table.
[0094] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 3, wherein the data warehouse further includes a dimension layer, and the multiple data layers include a data mart layer and a data application layer; The determining, according to the first information, reason information affecting the output time of the first data further includes: If the first database table is located in the dimension layer in the data warehouse, the layer where the first upstream database table is located in the data warehouse is the second data layer, and the second data layer is the data application layer or the data mart layer, it is determined that there is a dependency abnormality in the first database table.
[0095] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 1, wherein the first information includes the layer where each of the dependent database tables is located in the data warehouse; The determining, according to the first information, the cause information affecting the output time of the first data includes: If the first database table is a database table in a first link, it is determined that the first database table has a dependency abnormality, wherein the first link is composed of the dependent database tables with continuous dependency relationships, the dependent database tables constituting the first link are located at the same layer in the data warehouse, and the number of the dependent database tables in the first link is greater than or equal to a second preset threshold.
[0096] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 1, wherein the first information includes a first field quantity, the first field quantity is the quantity of fields in the first database table on which the first downstream database table depends, and the first downstream database table is a downstream database table that directly depends on the first database table; The determining, according to the first information, the cause information affecting the output time of the first data includes: Determine proportion information, where the proportion information is a ratio of the first field quantity to the second field quantity, where the second field quantity is the total quantity of fields in the first database table; According to the proportion information, it is determined whether there is field redundancy in the first database table.
[0097] According to one or more embodiments of the present disclosure, Example 7 provides the method of Example 6, wherein determining whether there is field redundancy in the first database table according to the proportion information includes: If there are multiple first downstream database tables, an average value of the proportion information corresponding to each of the first downstream database tables is used as the average proportion information; If the average proportion information is less than a third preset threshold, it is determined that field redundancy exists in the first database table.
[0098] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.
[0099] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0100] Although the subject matter has been described in language specific to structural features and / or method logic actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims. Regarding the device in the above embodiment, the specific manner in which each module performs the operation has been described in detail in the embodiment related to the method, and will not be elaborated here.
Claims
1. A data analysis method, characterized in that: The method comprises: In response to receiving an analysis request, determining a first database table, wherein the analysis request is used to request an analysis of a cause affecting the production time of the first data, the first database table includes all or part of the database tables in the dependent database tables, and the dependent database tables include upstream database tables directly dependent on the first data and upstream database tables indirectly dependent on the first data; Based on first information, reason information affecting the production time of the first data is determined, wherein the first information includes information of a database table having a dependency relationship with the first database table, and the reason information is used to indicate that there is a dependency abnormality and / or field redundancy in the first database table.
2. The method according to claim 1, characterized in that The determining of the first database table comprises: The upstream database table on which the first data is directly dependent and which is produced the latest is used as the first database table determined for the first time; Repeat the step of using the upstream database table on which the first database table determined last time directly depends and whose output time is the latest as the first database table determined this time, until the number of the first database tables determined this time reaches a first preset threshold, or the first database table determined this time has no upstream database table on which it is directly dependent.
3. The method according to claim 1, characterized in that The first information includes a layer where a first upstream database table is located in a data warehouse, the first upstream database table is an upstream database table that the first database table directly depends on, the data warehouse includes multiple data layers, and a data layer at a higher level depends on a data layer at a lower level; The determining, according to the first information, the cause information affecting the output time of the first data includes: If the first database table is located in the first data layer in the data warehouse, the layer where the first upstream database table is located in the data warehouse is the second data layer, and the level of the second data layer is higher than that of the first data layer, it is determined that there is a dependency abnormality in the first database table.
4. The method according to claim 3, characterized in that: The data warehouse further includes a dimension layer, and the multiple data layers include a data mart layer and a data application layer; The determining, according to the first information, reason information affecting the output time of the first data further includes: If the first database table is located in the dimension layer in the data warehouse, the layer where the first upstream database table is located in the data warehouse is the second data layer, and the second data layer is the data application layer or the data mart layer, it is determined that there is a dependency abnormality in the first database table.
5. The method according to claim 1, characterized in that The first information includes the layer where each of the dependent database tables is located in the data warehouse; The determining, according to the first information, the cause information affecting the output time of the first data includes: If the first database table is a database table in a first link, it is determined that the first database table has a dependency abnormality, wherein the first link is composed of the dependent database tables with continuous dependency relationships, the dependent database tables constituting the first link are located at the same layer in the data warehouse, and the number of the dependent database tables in the first link is greater than or equal to a second preset threshold.
6. The method according to claim 1, characterized in that The first information includes a first field quantity, where the first field quantity is the quantity of fields in the first database table on which the first downstream database table depends, and the first downstream database table is a downstream database table that directly depends on the first database table; The determining, according to the first information, the cause information affecting the output time of the first data includes: Determine proportion information, where the proportion information is a ratio of the first field quantity to the second field quantity, where the second field quantity is the total quantity of fields in the first database table; According to the proportion information, it is determined whether there is field redundancy in the first database table.
7. The method according to claim 6, characterized in that The determining, according to the proportion information, whether there is field redundancy in the first database table includes: If there are multiple first downstream database tables, an average value of the proportion information corresponding to each of the first downstream database tables is used as the average proportion information; If the average proportion information is less than a third preset threshold, it is determined that field redundancy exists in the first database table.
8. A data analysis device, characterized in that: The device comprises: A first determination module is configured to determine a first database table in response to receiving an analysis request, wherein the analysis request is used to request an analysis of a cause affecting the production time of the first data, and the first database table includes all or part of the database tables in the dependent database tables, and the dependent database tables include the upstream database tables directly dependent on the first data and the upstream database tables indirectly dependent on the first data; The second determination module is used to determine reason information affecting the production time of the first data based on the first information, wherein the first information includes information of a database table that has a dependency relationship with the first database table, and the reason information is used to indicate that there is a dependency abnormality and / or field redundancy in the first database table.
9. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processing device, the steps of the method according to any one of claims 1 to 7 are implemented.
10. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.