Artificial Intelligence-Based Data Middle Platform Evaluation Method and Related Devices
Through an artificial intelligence-based method, analyzing the structure and operation log of the data middle platform, building quantitative indicators and calculating weights, the problem of inaccurate evaluation results in the existing technology is solved, and a more accurate data middle platform evaluation is achieved.
Patent Information
- Application Number
- CN202210594836.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-05-27
AI Technical Summary
When evaluating the performance of data middle platform, the prior art ignores the information of key technical nodes and operation logs, resulting in the inaccurate evaluation results.
Using an artificial intelligence-based method, quantitative indicators are constructed by analyzing the structure and operation log of the data middle platform, and the weight of the indicators are calculated to obtain updated indicators and then evaluate them.
It improves the accuracy of data middle platform evaluation, can conduct quantitative evaluations for multiple major issues, and considers the importance of different indicators.
Smart Images

Figure CN114924943B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, device, electronic device, and storage medium for evaluating a data middle platform based on artificial intelligence. Background Art
[0002] A data middle platform is a platform that centrally stores a large amount of data from multiple sources and of multiple types, and can quickly process and analyze the data. To ensure the stability and efficiency of the business logic of the data middle platform, operation and maintenance personnel usually need to evaluate the performance of the data middle platform in real time to ensure that the developers of the data middle platform can respond in a timely manner to the performance degradation of the data middle platform.
[0003] The prior art usually evaluates the performance of the data middle platform based on business results, while ignoring some key technical nodes and operation log information in the data middle platform. This evaluation method is too general, resulting in inaccurate evaluation results. Summary of the Invention
[0004] In view of the above, it is necessary to provide a method and related devices for evaluating a data middle platform based on artificial intelligence to solve the technical problem of how to improve the accuracy of data middle platform evaluation. Among them, the related devices include a data middle platform evaluation device based on artificial intelligence, an electronic device, and a storage medium.
[0005] An embodiment of this application provides a method for evaluating a data middle platform based on artificial intelligence, including:
[0006] Build a data middle platform to obtain a data set, where the data set includes structured data and log data;
[0007] Parse the structured data in the data set to construct structure indicators;
[0008] Calculate the weights of the log data in the data set to construct business indicators;
[0009] Update the structure indicators and the business indicators to obtain updated indicators;
[0010] Interpolate the updated indicators according to the interpolation algorithm to construct an initial training data set;
[0011] Construct a performance evaluation model based on the initial training data set;
[0012] Collect real-time data and input it into the performance evaluation model to obtain the evaluation result of the data middle platform.
[0013] In the above data middle platform evaluation method, quantitative indicators are constructed by analyzing the structure and operation logs of the data middle platform, and the weights of the quantitative indicators are calculated to obtain updated indicators, and then the data middle platform is evaluated based on the updated indicators. In this way, quantitative evaluations can be carried out for multiple main problems of the data middle platform, and the importance of different indicators is considered, improving the accuracy of the evaluation results.
[0014] In some embodiments, the building of the data middle platform to obtain a data set includes:
[0015] Collecting the structure data and log data of the data middle platform according to a preset data sampling time point;
[0016] Jointly storing the structure data and the log data to obtain a data set.
[0017] In this way, the data set is constructed based on the structure data and log data of the data middle platform, which can provide data support for the subsequent construction of quantitative indicators and data analysis steps, and avoid the generalization error caused by the traditional conceptual indicator evaluation method.
[0018] In some embodiments, the structure indicators include:
[0019] The normalization rate indicator, which is the ratio of the number of tables that conform to the table naming specification to the total number of all tables in the data middle platform;
[0020] The reuse rate indicator, which is the ratio of the number of tables with a pre-dependency to the total number of all tables in the data middle platform, where the pre-dependency means that the data in this table is obtained based on the data in other tables;
[0021] The coverage rate indicator, which is the ratio of the number of tables without a pre-dependency but with a post-dependency to the total number of all tables in the data middle platform, where the post-dependency means that the data in other tables is obtained based on the data in this table.
[0022] In this way, the structure indicators of the data middle platform are constructed based on the table naming specification of the data middle platform and the hierarchical structure of the data middle platform. The structure indicators can characterize the structural performance of the data middle platform and provide data support for subsequent indicator evaluations, improving the accuracy of subsequent evaluations.
[0023] In some embodiments, the calculating of the weights of the log data in the data set to construct business indicators includes:
[0024] Calculating the ratio of the number of times an application program in the data middle platform is called to the total number of times all application programs are called, and taking this ratio as the first weight of this application program;
[0025] Calculate the ratio of the total duration that an application in the data middle platform occupies the central processing unit to the total running duration of the application, and use this ratio as the second weight of the application;
[0026] Calculate the ratio of the number of times an application in the data middle platform reports an error to the total number of times the application is called, and calculate the difference between a preset harmonic real number and this ratio, and use this difference as the third weight of the application;
[0027] Construct a business indicator based on the first weight, the second weight, and the third weight, and the business indicator is used to characterize the performance of the application in the data middle platform during operation.
[0028] In this way, by combining the log data during the operation of the data middle platform business logic, the business indicators of the data middle platform are constructed, which can characterize the performance fluctuations of the applications in the data middle platform during operation, and give the business indicators resolvability in time series, thereby improving the performance of the subsequent evaluation model.
[0029] In some embodiments, the updating the structure indicator and the business indicator to obtain an updated indicator includes:
[0030] Construct a covariance matrix based on the structure indicator and the business indicator;
[0031] Calculate the eigenvalues of the covariance matrix to obtain the weights of the structure indicator and the business indicator;
[0032] Update the structure indicator and the business indicator according to the weights to obtain an updated indicator.
[0033] In this way, a covariance matrix is constructed based on the structure indicator and the business indicator, and the eigenvalues of the covariance matrix are calculated. The eigenvalues can characterize the importance of the structure indicator and the business indicator. Using the eigenvalues as weights to update the structure indicator and the business indicator can avoid the error caused by equal weights of the indicators and improve the accuracy of subsequent evaluations.
[0034] In some embodiments, the interpolating the updated indicator according to the interpolation algorithm to construct an initial training data set includes:
[0035] Split the updated indicator to obtain an updated indicator interval set;
[0036] Fit the data in the updated indicator interval set according to the interpolation algorithm to obtain a set of function curves;
[0037] Sort the function curves according to the sampling time points corresponding to the curve starting points in the set of function curves to obtain a sorted set of function curves;
[0038] Connect the curves in the sorting function curve set end to end to obtain the initial training data set.
[0039] In this way, by using the interpolation method to complement the discrete index data in time series into continuous sequence data, the data acquisition result can be expanded and the data reserve can be improved, providing more complete data support for the fitting of the subsequent evaluation model, thereby improving the performance of the evaluation model.
[0040] In some embodiments, the collecting real-time data and inputting it into the performance evaluation model to obtain the evaluation result of the data middle platform includes:
[0041] Collect the relevant data of the data middle platform in real time, where the relevant data includes real-time structure data and real-time log data;
[0042] Construct and update the structure index and business index in real time to obtain the real-time updated index;
[0043] Input the real-time updated index into the performance evaluation model to evaluate the data middle platform and obtain the evaluation result.
[0044] In this way, an error can be reported in time when the structure or business performance of the data middle platform is abnormal to improve the stability of the data middle platform performance.
[0045] The embodiment of the present application also provides a data middle platform evaluation device based on artificial intelligence, and the device includes:
[0046] An acquisition unit, configured to build a data middle platform to obtain a data set, where the data set includes structure data and log data;
[0047] A first construction unit, configured to parse the structure data in the data set to construct a structure index;
[0048] A calculation unit, configured to calculate the weights of the log data in the data set to construct a business index;
[0049] An update unit, configured to update the structure index and the business index to obtain an updated index;
[0050] A second construction unit, configured to perform interpolation on the updated index according to the interpolation algorithm to construct an initial training data set;
[0051] A third construction unit, configured to construct a performance evaluation model according to the initial training data set;
[0052] An evaluation unit, configured to collect real-time data and input it into the performance evaluation model to obtain the evaluation result of the data middle platform.
[0053] The embodiment of the present application also provides an electronic device, and the electronic device includes:
[0054] a memory storing computer readable instructions; and
[0055] A processor executes computer-readable instructions stored in the memory to implement the artificial intelligence-based data middle platform evaluation method.
[0056] An embodiment of the present application also provides a computer-readable storage medium, in which computer-readable instructions are stored. The computer-readable instructions are executed by a processor in an electronic device to implement the artificial intelligence-based data middle platform evaluation method. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a flowchart of a preferred embodiment of the artificial intelligence-based data middle platform evaluation method involved in this application.
[0058] Figure 2 It is a flowchart of a preferred embodiment of calculating the weight of log data to construct business indicators involved in this application.
[0059] Figure 3 It is a flow chart of a preferred embodiment of updating the structural indicator and the business indicator to obtain the updated indicator involved in this application.
[0060] Figure 4 It is a flowchart of a preferred embodiment of the present application for interpolating the update index according to the interpolation algorithm to construct a training data set.
[0061] Figure 5 It is a flowchart of a preferred embodiment of the present application for collecting real-time data, inputting the performance evaluation model and obtaining the evaluation results of the data center.
[0062] Figure 6 It is a functional module diagram of a preferred embodiment of the artificial intelligence-based data middle platform evaluation device involved in this application.
[0063] Figure 7 It is a structural diagram of an electronic device of a preferred embodiment of the artificial intelligence-based data center evaluation method involved in this application. DETAILED DESCRIPTION
[0064] In order to more clearly understand the purpose, features and advantages of the present application, the present application is described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, in the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other. In the following description, many specific details are set forth to facilitate a full understanding of the present application, and the embodiments described are only a part of the embodiments of the present application, rather than all of the embodiments.
[0065] In addition, the terms "first" and "second" are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of this application, "a plurality of" means two or more unless otherwise specifically defined.
[0066] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used in the description of this application herein are for the purpose of describing specific embodiments only and are not intended to limit this application. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0067] An embodiment of this application provides a method for evaluating a data middle platform based on artificial intelligence, which can be applied to one or more electronic devices. The electronic device is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0068] The electronic device can be any electronic product capable of human-computer interaction with a user. For example, a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an Internet protocol television (IPTV), a smart wearable device, etc.
[0069] The electronic device may further include a network device and / or a user device. Among them, the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing.
[0070] The network where the electronic device is located includes but is not limited to the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN), etc.
[0071] Such as Figure 1As shown, it is a flowchart of a preferred embodiment of the data center evaluation method based on artificial intelligence in this application. According to different requirements, the order of steps in this flowchart can be changed, and some steps can be omitted.
[0072] S10. Build a data center to obtain a data set, where the data set includes structured data and log data.
[0073] In an optional embodiment, building a data center to obtain a data set includes:
[0074] S101. Collect the structured data and log data of the data center according to a preset data sampling time point.
[0075] In this optional embodiment, the data center is a platform that centrally stores a large amount of data and can quickly process and analyze the data. The characteristics of the data in the data center include multiple sources and multiple types.
[0076] In this optional embodiment, the sampling duration and sampling frequency can be formulated based on the life cycle of the application program of the enterprise data center. Exemplarily, the sampling duration can be a natural month, that is, 30 natural days; the sampling frequency can be 1 / 3600 Hz, that is, sampling once per hour. Furthermore, the sampling time point can be formulated based on the sampling duration and sampling frequency. The interval between the sampling time points is one hour, and the total number of sampling time points is 720.
[0077] In this optional embodiment, the table naming specification of the data middle platform can be formulated according to the hierarchical structure of the data middle platform. The hierarchical structure of the data middle platform includes the ODS layer, the DWD layer, the DWS layer, the DM layer, and the DIM layer. The full name of the ODS is the Operational Data Store layer, which means the operational data layer. Its main function is to collect, converge, and integrate various business data to add data identifiers and convert unstructured data into structured data. The full name of the DWD layer is Data Warehouse Detail, which means the detailed layer of the data middle platform. Its main function is to store the detailed layer data divided by theme in the data middle platform. The full name of the DWS layer is data warehouse service, which means the summary layer of the data middle platform. Its main function is to store the data after detailed summary, playing the role of reducing the data volume and unifying the index processing. The full name of the DM layer is the Data Market layer, which means the data mart layer. Its main function is to build a local data warehouse starting from a certain business application. The full name of the DIM layer is the Dimension layer, which means the dimension layer. Its main function is to establish a data analysis dimension table, which can reduce the risk of inconsistent data calculation caliber and algorithms. Exemplarily, the table naming specification in the ODS layer is: ODS_business system database name_business system database table name. The table naming specifications in the DWD layer, DWS layer, and DM layer are: hierarchical name_subject domain_business process_description_table splitting rule. The table naming specification in the DIM layer is: DIM_master data domain_description_table splitting rule.
[0078] In this optional embodiment, a preset SQL script can be run based on the sampling time point to collect the dataset Table_name of all table names in the data middle platform, and the table names in the Table_name can be marked according to a preset Python script. Among them, the table names that conform to the table naming specification can be marked as "standard", and the table names that do not conform to the table naming specification are marked as "non-standard". According to the sampling time point, a preset Python script is used for parsing, and the table name data in the Table_name is statistically analyzed to obtain the total number S of tables that conform to the naming specification, where S is a positive integer.
[0079] In this optional embodiment, the number N of all tables in the data center can be counted based on the sampling time points using a custom program, where N is a positive integer. The total number M of post-dependencies of all tables in the data center is counted according to a preset program, where M is a positive integer. The total number K of post-dependencies of all tables in the ODS layer of the data center is counted according to a preset SQL script, where K is a positive integer. At each sampling time point, a piece of data containing S, N, M, and K is collected. Since there are 720 sampling time points, there are a total of 720 pieces of sampling data containing S, N, M, and K. Since M represents the total number of post-dependencies of all tables, and N represents the total number of all tables, M is greater than or equal to N. Since K only represents the total number of post-dependencies of all tables in the ODS layer, and N represents the total number of all tables, K is less than N. Each piece of data containing S, N, M, and K within a sampling duration can be used as the structural data.
[0080] In this optional embodiment, a preset SQL script can be used to extract the data center operation log Log based on the sampling duration. Exemplarily, the first column of the Log can be the timestamp of the log, the second column can be the name of the called application program, and the third column can be the application status / error message. The log data obtained by parsing the information in the Log includes: the total number W of times all application programs are called by the data center within a sampling duration, where W is a positive integer; the number of times Count that each application program is called within a sampling duration, where Count is a positive integer; the total duration Time that each application program runs within a sampling duration, where 0 < Time < 30 days; the number of times Warning_count that each application program reports an error within a sampling duration, where Warning_count is a positive integer; the total duration CPU_time that each application program occupies the CPU within a sampling duration, where 0 < CPU_time < 30 days.
[0081] S102, jointly store the structural data and the log data to obtain the data set.
[0082] In this optional embodiment, the structural data and the log data can be jointly stored as a CSV format document, and further, the CSV format document can be used as the data set.
[0083] In this way, the data set is constructed based on the hierarchical structure and operation log of the data center, providing data support for the subsequent construction of quantitative indicators and data analysis steps, thereby avoiding the generalization error caused by the traditional conceptual indicator evaluation method.
[0084] S11, parse the structural data in the data set to construct structural indicators.
[0085] In an optional embodiment, parsing the structural data in the dataset to construct structural metrics includes:
[0086] In this optional embodiment, the structural data can be parsed according to possible structural problems that may occur in the data middle platform to construct the structural metrics of the data middle platform. The structural problems of the data middle platform at least include: inconsistent data caliber, chimney-style development, and poor source data quality.
[0087] In this optional embodiment, the inconsistent data caliber specifically means that the table names in the data middle platform do not meet the specification requirements, resulting in incorrect data being retrieved when querying data; the chimney-style development specifically means that every time a new requirement is encountered in the data middle platform, the original data in the ODS layer is recalculated, and the data analysis logic is rebuilt for the new requirement, which consumes a large amount of resources and may cause queue blocking; the poor source data quality specifically means that tables other than the ODS layer rely too much on the tables in the ODS layer.
[0088] In this optional embodiment, since the tables in the ODS layer are not data-cleaned, resulting in poor overall data quality in the data middle platform, a standardization rate metric A can be constructed to evaluate the problem of inconsistent data caliber. The calculation method of the standardization rate metric A is:
[0089] A = S / N
[0090] Wherein, S represents the total number of tables that conform to the naming specification in the dataset of the names of all tables in the data middle platform obtained based on a certain sampling time point, and the value of S is a positive integer; N represents the total number of all tables in the data middle platform obtained based on a certain sampling time point, and the value of N is a positive integer; A represents a standardization rate metric corresponding to a certain sampling time point, and the value range of A is (0, 1]; the higher the standardization rate metric, the more standardized the table naming in the data middle platform at that moment, and the lower the probability of reporting an error when calling a data table, indicating that the structure of the data middle platform is more perfect.
[0091] In this optional embodiment, by way of example, when S = 100 and N = 102, the calculation method of the standardization rate metric is:
[0092] A = 100 / 102 = 0.98
[0093] Wherein, the standardization rate metric A = 0.98.
[0094] In this optional embodiment, a reuse rate metric B can be constructed to evaluate the problem of chimney-style development. The calculation method of the reuse rate metric B is:
[0095] B = M / N
[0096] Among them, M represents the total number of post-dependencies of all tables in the data center obtained based on a certain sampling time point, and M is a positive integer; N represents the number of all tables in the data center obtained based on a certain sampling time point, and N is a positive integer; B represents the reuse rate index obtained corresponding to a certain sampling time point, and the value range of B is [1, +∞]; the higher the reuse rate index, the higher the efficiency of the tables in the data center being reused at that moment, indicating that the structure of the data center is more perfect.
[0097] In this optional embodiment, by way of example, when M = 600 and N = 102, the calculation method of the reuse rate index is as follows:
[0098] B = 600 / 102 = 5.88
[0099] Among them, the reuse rate index B = 5.88.
[0100] In this optional embodiment, a coverage rate index C can be constructed to evaluate the problem of poor source data quality. The calculation method of the coverage rate index C is as follows:
[0101] C = 1 - K / N
[0102] Among them, C represents the coverage rate index obtained based on a certain sampling time point, and the value range of C is (0, 1); K represents the total number of post-dependencies of all tables in the ODS layer of the data center obtained based on a certain sampling time point, and the value of K is a positive integer; N represents the number of all tables in the data center obtained based on a certain sampling time point, and the value of N is a positive integer; the higher the coverage rate index, the fewer times the tables in the ODS layer are reused at that moment, indicating that the data quality of the data center is higher.
[0103] In this optional embodiment, by way of example, when K = 50 and N = 102, the calculation method of the coverage rate index is as follows:
[0104] C = 1 - 50 / 102 = 0.51
[0105] Among them, the coverage rate index C = 0.51.
[0106] In this optional embodiment, the indicators A, B, and C can be used as the structure indicators. Since the structure data contains 720 pieces of data, the data center structure indicators A, B, and C also contain 720 pieces of data respectively.
[0107] In this way, according to the possible problems of the data center and combined with the table naming rules of the data center, structure indicators are constructed. The structure indicators can characterize the structure performance of the data center and provide data support for subsequent indicator evaluation, improving the accuracy of subsequent evaluation.
[0108] S12. Calculate the weights of the log data in the dataset to construct business metrics.
[0109] Please refer to Figure 2 , in an optional embodiment, calculating the weights of the log data in the dataset to construct business metrics includes:
[0110] S121. Calculate the ratio of the number of times an application in the data middle platform is called to the total number of times all applications are called, and use this ratio as the first weight of the application.
[0111] In this optional embodiment, the business metrics of the data middle platform can be constructed based on the W, Count, Time, Warning_count, and CPU_time. During the sampling duration, the more times an application is called out of the total number of times all applications are called, the more important the application is for representing the business performance of the data middle platform, and the higher the weight of the application. For this application, the first weight D is constructed, where D = Count / W, and the value range of D is (0, 1). The higher D is, the more important the application is.
[0112] S122. Calculate the ratio of the total duration that an application in the data middle platform occupies the central processing unit to the total running duration of the application, and use this ratio as the second weight of the application.
[0113] During the sampling duration, the higher the ratio of the duration that each application occupies the CPU to the total running duration of the application, the more data interaction the application generates with the data middle platform during operation, and the higher the weight of the application. For this application, the second weight E is constructed, where E = CPU_time / Time, and the value range of E is (0, 1). The higher E is, the more important the application is.
[0114] S123. Calculate the ratio of the number of times an application in the data middle platform reports an error to the total number of times the application is called, and calculate the difference between a preset harmonic real number and this ratio, and use this difference as the third weight of the application.
[0115] During the sampling duration, the lower the ratio of the number of times each application reports an error to the number of times it is called, the more complete the business logic of the application, and the higher the weight of the application. For this application, the third weight F is constructed, where F = R - (Warning_count / Count), where R represents a preset harmonic real number. Exemplarily, R can be 1, and the value range of F is (0, 1). The higher F is, the more important the application is.
[0116] S124. Construct the service metrics based on the first weight, the second weight, and the third weight.
[0117] Construct the service metric G of the data center based on the three weights D, E, and F of the application and the status of the application corresponding to each timestamp in the aforementioned log data. If the status of the application corresponding to the timestamp is normal, the construction method of the service metric of the data center corresponding to this timestamp is: the product of 1 and the first, second, and third weights, denoted as G = 1 * D * E * F; conversely, if the application reports an error, the construction method of the service metric of the data center corresponding to this timestamp is: the product of -1 and the first, second, and third weights, denoted as G = -1 * D * E * F.
[0118] In this alternative embodiment, the sampling duration of the operation log of the data center is 30 days, and the timestamp frequency of the operation log is usually 1 Hz, that is, the application running status is recorded once per second. Therefore, the service metric G of the data center contains 2,592,000 data.
[0119] In this way, the service metrics of the data center are constructed by combining the log data in the operation process of the data center business logic, which can characterize the performance fluctuations in the running process of the application in the data center and endow the service metrics with resolvability in time series, thereby improving the performance of the subsequent evaluation model.
[0120] S13. Update the structure metrics and the service metrics to obtain updated metrics.
[0121] Please refer to Figure 3 , in an alternative embodiment, updating the structure metrics and the service metrics to obtain updated metrics includes:
[0122] S131. Construct a covariance matrix based on the structure metrics and the service metrics.
[0123] In this alternative embodiment, a dataset related to the data center hierarchy structure metrics can be constructed based on the three structure metrics A, B, and C. The metric dataset includes 720 rows (corresponding to 720 sampling time points) and 3 columns (corresponding to three metrics).
[0124] In this alternative embodiment, since the dimension of the service metric G is different from the dimension of the structure metric, 720 service metric data can be screened out from the service metric G according to the sampling time points.
[0125] In this alternative embodiment, the structure metrics A, B, and C and the 720 screened service metrics G can be arranged column by column to construct a dataset, denoted as Data. The covariance matrix can be calculated based on each column of metrics in Data, and the covariance matrix is denoted as Z.
[0126] S132. Calculate the eigenvalues of the covariance matrix to obtain the weights of the structural indicators and business indicators.
[0127] Calculate the eigenvalues of the Z matrix. Since there are four indicators, four eigenvalues can be obtained, denoted as W A 、W B 、W C 、W G .
[0128] In this alternative embodiment, the four eigenvalues can be used as the weights of the structural indicators and business indicators, and the subscripts of the eigenvalues correspond to the corresponding indicators.
[0129] S133. Update the structural indicators and business indicators according to the weights to obtain updated indicators.
[0130] In this alternative embodiment, the structural indicators and business indicators can be updated based on the product of the weights and the indicators. Exemplarily, the updated A indicator is denoted as A new ,A new =A·W A ;The updated B indicator is denoted as B new ,B new =B·W B ;The updated C indicator is denoted as C new ,C new =C·W C ;The updated G indicator is denoted as G new ,G new =G·W G .
[0131] In this way, a covariance matrix is constructed based on the structural indicators and business indicators, and the eigenvalues of the covariance matrix are calculated. The eigenvalues can characterize the importance of the structural indicators and business indicators. Using the eigenvalues as weights to update the structural indicators and business indicators can avoid the error caused by equal weights of indicators and improve the accuracy of subsequent evaluations.
[0132] S14. Interpolate the updated indicators according to the interpolation algorithm to construct an initial training dataset.
[0133] Please refer to Figure 4 , in an alternative embodiment, the interpolating the updated indicators according to the interpolation algorithm to construct an initial training dataset includes:
[0134] S141. Split the updated indicators to obtain an updated indicator interval set.
[0135] In this optional embodiment, for the updated indicator data, the cubic spline interpolation method can be used to obtain the data center structure indicator fluctuation curve and the data center business indicator fluctuation curve with a time span of 720 hours. The cubic spline interpolation method is a model fitting algorithm. Taking indicator A as an example, its main process is to divide the data point set of indicator A into n intervals. For example, since there are 720 sampling time points in total, n=719, and there is an interval between every two sampling time points.
[0136] In this optional embodiment, the 719 interval groups may be used as the interval set.
[0137] S142, fitting the data in the update index interval set according to the interpolation algorithm to obtain a function curve set.
[0138] In this optional embodiment, on each interval, the points between the intervals should satisfy the cubic equation, S(x i )=y i =a i +b i ·x i +c i ·x i 2 +d i ·x i 3 , this equation is called a cubic spline function, where y i Represents the value of the indicator at the i-th sampling time point, x i To represent the value at the i-th time point, the cubic spline function should meet the following three conditions: all points must meet the interpolation condition; the first-order derivative and the second-order derivative of the n-1 internal points should be continuous; the second-order derivative of the function curve at the two endpoints is 0. Based on the three conditions, the cubic spline function can be solved to obtain the coefficient a of the cubic spline equation i , b i 、c i d i , further, the function curve in each interval can be fitted based on the coefficients of the cubic spline equation to obtain the function curve set.
[0139] S143, sorting the function curves according to the sampling time points corresponding to the starting points of the curves in the function curve set to obtain a sorted function curve set.
[0140] In this optional embodiment, the curves in the function curve set may be sorted according to the sampling time points corresponding to the starting points of the function curves to obtain a sorted function curve set. If the sampling time points corresponding to the starting points of the function curves are earlier, the sorting order of the function curves is also earlier.
[0141] S144. Connect the head and tail of the curves in the sorting function curve set to obtain the initial training data set.
[0142] In this optional embodiment, the head and tail of the function curves can be connected according to the order of the function curves in the sorting function curve set to obtain the initial training data set.
[0143] In this way, by using the interpolation algorithm to complement the discrete index data in time series into continuous sequence data, the data acquisition result can be expanded and the data reserve can be improved, providing more complete data support for the fitting of the subsequent evaluation model, thereby improving the performance of the evaluation model.
[0144] S15. Construct a performance evaluation model based on the initial training data set.
[0145] In an optional embodiment, constructing a performance evaluation model based on the initial training data set includes:
[0146] In this optional embodiment, a preset neural network model to be updated can be trained based on the initial training data set to obtain the performance evaluation model.
[0147] In this optional embodiment, the data in the initial training data set can be collected according to a preset frequency to construct an updated training data set, the purpose of which is to convert the continuous time series into discrete data in time series for subsequent training of the preset neural network model to be updated. Exemplarily, the preset frequency can be 3000Hz, then the dimension of the updated training data set is 7.776×10 9 ×4, with a total of 7.776×10 9 rows and 4 columns. Each row of data can be regarded as a 1*4 vector, denoted as v i , i ∈ [1, 7.776×10 9 .
[0148] In this optional embodiment, the neural network model to be updated can be an LSTM network. The LSTM network can be trained based on the updated training data set. The LSTM is a time series neural network, whose full name is Long Short Term Memory, meaning long short-term memory neural network. The LSTM network is composed of neurons, and the neurons are connected in series. Each neuron contains three input data and two output data. The input data includes: C t-1 , the memory information of the previous moment; h t-1 , the output data of the previous moment; x t , the input data of the current moment. The output data includes: C t , the memory information of the current moment; h t the output data of the current moment.
[0149] In this alternative embodiment, each neuron includes three computational parts: a forget gate; an input gate; and an output gate. The main process of the forget gate is to input the output data h at the previous moment t-1 and the input data x at the current moment t into the Sigmoid function together, and the output result is between 0 and 1, denoted as σ 1 . The σ 1 can represent the importance degree of the memory at the previous moment. The main process of the input gate is to input the output data h at the previous moment t-1 and the input data x at the current moment t into the Sigmoid function together to obtain the importance factor of the current input, with a value between [0, 1], denoted as σ 2 ; input the output data h at the previous moment t-1 and the input data x at this moment t into the tanh function together to obtain the centralized value of the current input, with a value between [-1, 1], denoted as tanh 1 . The product between the σ 2 and the tanh 1 is the state information C at the current moment t . The main process of the output gate is to input the output data h at the previous moment t-1 and the input data x at the current moment t into the Sigmoid function together, and the output value is between [0, 1], denoted as σ 3 . The σ 3 is used to evaluate how much output value the state information C at the current moment t has. The higher the σ 3 , the higher the output value of the C t . Input the state information C at the current moment t into the tanh function to obtain a result of [-1, 1], denoted as tanh 2 . The output data h at the current moment t =σ 3 ·tanh 2 .
[0150] In this alternative embodiment, the v 1 is the input value x of the LSTM network at the time point t = 1 1 , C t-1 and h t-1The initial value of can be set to 0. Based on the training method of the LSTM network and the Train_data, a complete LSTM network can be trained. Developers can collect the structural indicators and business indicators of the data center at a certain moment at any time, input the complete LSTM network, and predict and evaluate the quantitative indicators of the data center in the future time period.
[0151] In this optional embodiment, the fully trained LSTM network can be used as the performance evaluation model.
[0152] S16, collecting real-time data and inputting it into the performance evaluation model to obtain the evaluation results of the data center.
[0153] like Figure 5 As shown, in an optional embodiment, collecting real-time data and inputting it into the performance evaluation model to obtain the evaluation result of the data center includes:
[0154] S161, collect relevant data of the data center in real time.
[0155] In this optional embodiment, developers of the data middle platform can collect data related to the data middle platform in real time, and the data related to the data middle platform includes real-time structure data and real-time log data.
[0156] In this optional embodiment, the real-time structure data can be a piece of data containing four dimensions, and the four dimensions include the total number of tables that comply with the naming conventions in the data center at a certain moment, the number of all tables in the data center, the total number of post-dependencies of all tables in the data center, and the total number of post-dependencies of all tables in the ODS layer of the data center.
[0157] In this optional embodiment, the real-time log data includes the total number of times all applications have been called by the middle platform in the past sampling period, the number of times the application recorded at the sampling moment in the operation log of the data middle platform has been called in the past sampling period, the total running time of the application in the past sampling period, the number of times the application reported an error in the past sampling period, and the total time the application occupied the CPU in the past sampling period.
[0158] S162, construct and update the structural indicators and business indicators in real time to obtain real-time updated indicators.
[0159] In this optional embodiment, the real-time structure index and the real-time business index may be constructed based on step S11 and step S12 using the real-time structure data and the real-time log data.
[0160] In this optional embodiment, the real-time structure indicator and the real-time business indicator may be updated based on step S13 to obtain a real-time update indicator.
[0161] S163, input the real-time updated metrics into the performance evaluation model to evaluate the data middle platform and obtain the evaluation result.
[0162] In this optional embodiment, the real-time updated metrics can be input into the performance evaluation model to obtain the volatility of the data middle platform metrics within a certain period of time in the future. Further, the performance volatility of the data middle platform within a certain period of time in the future can be quantitatively evaluated.
[0163] In this optional embodiment, when the cosine distance between the volatility of the data middle platform metrics and the preset threshold is greater than 0.5, the evaluation result is "unqualified"; if the cosine distance between the volatility of the data middle platform metrics and the preset threshold is not greater than 0.5, the evaluation result is "qualified".
[0164] In this way, a performance evaluation model is obtained based on the initial training dataset. The performance evaluation model can evaluate the volatility of the data middle platform performance through real-time metric data, which is more efficient and accurate than the existing methods.
[0165] In the above data middle platform evaluation method, quantitative metrics are constructed by analyzing the structure and operation logs of the data middle platform, and the weights of the quantitative metrics are calculated to obtain the updated metrics. Then, the data middle platform is evaluated based on the updated metrics. In this way, quantitative evaluations can be carried out for multiple main problems of the data middle platform, and the importance of different metrics is considered, improving the accuracy of the evaluation results.
[0166] As Figure 6 shown, it is a functional module diagram of a preferred embodiment of the data middle platform evaluation device based on artificial intelligence provided by an embodiment of the present application. The data middle platform evaluation device 11 based on artificial intelligence includes an acquisition unit 110, a first construction unit 111, a calculation unit 112, an update unit 113, a second construction unit 114, a third construction unit 115, and an evaluation unit 116. The modules / units referred to in the present application refer to a series of computer program segments that can be executed by a processor 13 and can complete fixed functions, and are stored in a memory 12. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0167] In an optional embodiment, the acquisition unit 110 is used to build a data middle platform to obtain a dataset, and the dataset includes structure data and log data.
[0168] In this optional embodiment, the building of the data middle platform to obtain the dataset includes:
[0169] Collect the structure data and log data of the data middle platform according to the preset data sampling time points;
[0170] Jointly store the structure data and the log data to obtain the data set.
[0171] In this optional embodiment, the data middle platform is a platform for centrally storing a large amount of data and capable of quickly processing and analyzing data. The characteristics of the data in the data middle platform include multiple sources and multiple types.
[0172] In this optional embodiment, the sampling duration and sampling frequency can be determined based on the life cycle of the application program of the enterprise data middle platform. Exemplarily, the sampling duration can be one natural month, that is, 30 natural days; the sampling frequency can be 1 / 3600 Hz, that is, sampling once per hour. Furthermore, the sampling time points can be determined based on the sampling duration and sampling frequency. The interval between the sampling time points is one hour, and the total number of sampling time points is 720.
[0173] In this optional embodiment, the table naming specification of the data middle platform can be determined according to the hierarchical structure of the data middle platform. The hierarchical structure of the data middle platform includes the ODS layer, DWD layer, DWS layer, DM layer, and DIM layer. The full name of the ODS is the Operational Data Store layer, which means the operational data layer. Its main function is to collect, converge, and integrate various business data to add data identifiers and convert unstructured data into structured data; the full name of the DWD layer is Data Warehouse Detail, which means the detail layer of the data middle platform. Its main function is to store the detail layer data divided by theme in the data middle platform; the full name of the DWS layer is data warehouse service, which means the summary layer of the data middle platform. Its main function is to store the data after detail summarization, playing the role of reducing the data volume and unifying the index processing; the full name of the DM layer is the Data Market layer, which means the data mart layer. Its main function is to build a local data warehouse starting from a certain business application; the full name of the DIM layer is the Dimension layer, which means the dimension layer. Its main function is to establish a data analysis dimension table, which can reduce the risk of inconsistent data calculation caliber and algorithms. Exemplarily, the table naming specification in the ODS layer is: ODS_business system database name_business system database table name, and the table naming specifications in the DWD layer, DWS layer, and DM layer are: hierarchical name_subject domain_business process_description_table splitting rule, and the table naming specification in the DIM layer is: DIM_master data domain_description_table splitting rule.
[0174] In this optional embodiment, a preset SQL script can be run based on the sampling time point to collect the dataset Table_name of the names of all tables in the data middle platform, and the table names in the Table_name can be marked according to a preset Python script. Among them, the table names that conform to the table naming specification can be marked as "standard", and the table names that do not conform to the table naming specification are marked as "non-standard". Based on the sampling time point, a preset Python script is used for parsing, and the table name data in the Table_name is counted to obtain the total number S of tables that conform to the naming specification, where S is a positive integer.
[0175] In this optional embodiment, a custom program can be used to count the total number N of all tables in the data middle platform based on the sampling time point, where N is a positive integer. The total number M of post-dependencies of all tables in the data middle platform is counted according to a preset program, where M is a positive integer. The total number K of post-dependencies of all tables in the ODS layer of the data middle platform is counted according to a preset SQL script, where K is a positive integer. At each sampling time point, a piece of data containing S, N, M, and K will be collected. Since there are 720 sampling time points, there are a total of 720 pieces of sampling data containing N, M, and K. Since M represents the total number of post-dependencies of all tables and N represents the total number of all tables, M is greater than or equal to N. Since K only represents the total number of post-dependencies of all tables in the ODS layer and N represents the total number of all tables, K is less than N. Each piece of data containing S, N, M, and K within a sampling duration can be used as the structured data.
[0176] In this optional embodiment, a preset SQL script can be used to extract the data middle platform operation log Log based on the sampling duration. Exemplarily, the first column of the Log can be the timestamp of the log, the second column can be the name of the application program called, and the third column can be the application status / error message. The log data obtained by parsing the information in the Log includes: the total number W of all application programs called by the middle platform within a sampling duration, where W is a positive integer; the number of times Count that each application program is called within a sampling duration, where Count is a positive integer; within a sampling duration, the total running duration Time of each application program, 0 < Time < 30 days; within a sampling duration, the number of times Warning_count that each application program reports an error, where Warning_count is a positive integer; within a sampling duration, the total CPU duration CPU_time occupied by each application program, 0 < CPU_time < 30 days.
[0177] In this optional embodiment, the structured data and the log data can be jointly stored as a CSV format document, and further, the CSV format document can be used as the dataset.
[0178] In an optional embodiment, the first construction unit 111 is configured to parse the structural data in the dataset to construct structural metrics.
[0179] In this optional embodiment, the structural metrics include:
[0180] A standardization rate metric, which is the ratio of the number of tables that conform to the table naming standard to the total number of all tables in the data middle platform;
[0181] A reuse rate metric, which is the ratio of the number of tables with pre-dependencies to the total number of all tables in the data middle platform. The pre-dependency means that the data in this table is obtained based on the data in other tables;
[0182] A coverage rate metric, which is the ratio of the number of tables without pre-dependencies but with post-dependencies to the total number of all tables in the data middle platform. The post-dependency means that the data in other tables is obtained based on the data in this table.
[0183] In this optional embodiment, the structural data can be parsed according to possible structural problems that may occur in the data middle platform to construct the structural metrics of the data middle platform. The structural problems of the data middle platform at least include: inconsistent data calibers, chimney-style development, and poor source data quality.
[0184] In this optional embodiment, the inconsistent data calibers specifically refer to that the table names in the data middle platform do not meet the specification requirements, resulting in incorrect data being retrieved when querying data; the chimney-style development specifically refers to that each time a new requirement is encountered in the data middle platform, the original data in the ODS layer is recalculated, and the data analysis logic is rebuilt for the new requirement, which consumes a large amount of resources and may cause queue blockage; the poor source data quality specifically refers to that tables other than those in the ODS layer overly rely on the tables in the ODS layer.
[0185] In this optional embodiment, since the tables in the ODS layer are not data-cleaned, resulting in relatively poor overall data quality in the data middle platform, a standardization rate metric A can be constructed to evaluate the problem of inconsistent data calibers. The calculation method of the standardization rate metric A is:
[0186] A = S / N
[0187] Among them, S represents the total number of tables that conform to the naming convention in the dataset of the names of all tables in the data middle platform obtained based on a certain sampling time point, and the value of S is a positive integer; N represents the total number of all tables in the data middle platform obtained based on a certain sampling time point, and the value of N is a positive integer; A represents a standard rate index obtained corresponding to a certain sampling time point, and the value range of A is (0, 1]; the higher the standard rate index, the more standardized the table naming in the data middle platform at that moment, the lower the probability of reporting an error when calling a data table, indicating that the structure of the data middle platform is more perfect.
[0188] In this optional embodiment, for example, when S = 100 and N = 102, the calculation method of the standard rate index is:
[0189] A = 100 / 102 = 0.98
[0190] Among them, the standard rate index A = 0.98.
[0191] In this optional embodiment, a reuse rate index B can be constructed to evaluate the chimney-style development problem, and the calculation method of the reuse rate index B is:
[0192] B = M / N
[0193] Among them, M represents the total number of post-dependencies of all tables in the data middle platform obtained based on a certain sampling time point, M is a positive integer; N represents the total number of all tables in the data middle platform obtained based on a certain sampling time point, N is a positive integer; B represents a reuse rate index obtained corresponding to a certain sampling time point, and the value range of B is [1, +∞); the higher the reuse rate index, the higher the efficiency of repeated utilization of tables in the data middle platform at that moment, indicating that the structure of the data middle platform is more perfect.
[0194] In this optional embodiment, for example, when M = 600 and N = 102, the calculation method of the reuse rate index is:
[0195] B = 600 / 102 = 5.88
[0196] Among them, the reuse rate index B = 5.88.
[0197] In this optional embodiment, a coverage rate index C can be constructed to evaluate the problem of poor source data quality, and the calculation method of the coverage rate index C is:
[0198] C = 1 - K / N
[0199] Among them, C represents the coverage rate index obtained based on a certain sampling time point, and the value range of C is (0, 1); K represents the total number of post-dependencies of all tables in the ODS layer of the data middle platform obtained based on a certain sampling time point, and the value of K is a positive integer; N represents the number of all tables in the data middle platform obtained based on a certain sampling time point, and the value of N is a positive integer; the higher the coverage rate index, the fewer the number of times the tables in the ODS layer are reused at that moment, indicating that the data quality of the data middle platform is higher.
[0200] In this optional embodiment, for example, when K = 50 and N = 102, the calculation method of the coverage rate index is as follows:
[0201] C = 1 - 50 / 102 = 0.51
[0202] Among them, the coverage rate index C = 0.51.
[0203] In this optional embodiment, the metrics A, B, and C can be used as the structure metrics. Since the structure data contains 720 pieces of data, the data middle platform structure metrics A, B, and C also contain 720 pieces of data respectively.
[0204] In an optional embodiment, the calculation unit 112 is used to calculate the weights of the log data in the dataset to construct business metrics.
[0205] In this optional embodiment, calculating the weights of the log data to construct business metrics includes:
[0206] Calculating the ratio of the number of times an application program in the data middle platform is called to the total number of times all application programs are called, and using this ratio as the first weight of this application program;
[0207] Calculating the ratio of the total duration that an application program in the data middle platform occupies the central processing unit to the total running duration of this application program, and using this ratio as the second weight of this application program;
[0208] Calculating the ratio of the number of times an application program in the data middle platform reports an error to the total number of times this application program is called, and calculating the difference between a preset harmonic real number and this ratio, and using this difference as the third weight of this application program;
[0209] Constructing the business metrics based on the first weight, second weight, and third weight.
[0210] In this optional embodiment, the business metrics of the data middle platform can be constructed based on the W, Count, Time, Warning_count, and CPU_time. During the sampling duration, the higher the ratio of the number of times an application is called to the total number of times all applications are called, the more important the application is in representing the business performance of the data middle platform, and the higher the weight of the application. For this application, the first weight D is constructed, where D = Count / W, and the value range of D is (0, 1). The higher D is, the more important the application is.
[0211] During the sampling duration, the higher the ratio of the CPU time occupied by each application to the total running time of the application, the more data interaction the application has with the data middle platform during operation, and the higher the weight of the application. Then, the second weight E is constructed for this application, where E = CPU_time / Time, and the value range of E is (0, 1). The higher E is, the more important the application is.
[0212] During the sampling duration, the lower the ratio of the number of error reports of each application to the number of times it is called, the more complete the business logic of the application is, and the higher the weight of the application. For this application, the third weight F is constructed, where F = 1 - (Warning_count / Count), where R represents a preset harmonic real number. Exemplarily, R can be 1, and the value range of F is (0, 1). The higher F is, the more important the application is.
[0213] Based on the three weights D, E, F of the application and the status of the application corresponding to each timestamp in the foregoing log data, the business metric G of the data middle platform is constructed. If the status of the application corresponding to the timestamp is normal, the construction method of the business metric of the data middle platform corresponding to this timestamp is: the product of 1 and the first, second, and third weights, denoted as G = 1 * D * E * F; conversely, if the application reports an error, the construction method of the business metric of the data middle platform corresponding to this timestamp is: the product of -1 and the first, second, and third weights, denoted as G = -1 * D * E * F.
[0214] In this optional embodiment, the sampling duration of the operation log of the data middle platform is 30 days, and the timestamp frequency of the operation log is usually 1 Hz, that is, the application running status is recorded once per second. Therefore, the business metric G of the data middle platform contains 2,592,000 data.
[0215] In an optional embodiment, the update unit 113 is used to update the structure metric and the business metric to obtain the updated metric.
[0216] In this optional embodiment, updating the structure indicator and the service indicator to obtain an update indicator includes:
[0217] Constructing a covariance matrix according to the structural indicators and the business indicators;
[0218] Calculating the eigenvalues of the covariance matrix to obtain the weights of the structural indicators and the business indicators;
[0219] The structural indicator and the business indicator are updated according to the weight to obtain an updated indicator.
[0220] In this optional embodiment, a data set of indicators related to the hierarchical structure of the data center can be constructed based on the three structural indicators A, B, and C. The indicator data set includes 720 rows (corresponding to 720 sampling time points) and 3 columns (corresponding to three indicators).
[0221] In this optional embodiment, since the dimension of the business indicator G is different from the dimension of the structural indicator, 720 business indicator data can be screened out from the business indicator G according to the sampling time point.
[0222] In this optional embodiment, the structural indicators A, B, C and the 720 selected business indicators G can be arranged column by column to construct a data set, which is recorded as Data. The covariance matrix can be calculated based on each column of indicators in the Data, and the covariance matrix is recorded as Z.
[0223] Calculate the eigenvalues of the Z matrix. Since the index has four items, four eigenvalues can be obtained, which are recorded as W A , W B , W C , W G .
[0224] In this optional embodiment, the four characteristic values may be used as weights of the structural indicators and business indicators, and the subscripts of the characteristic values correspond to the corresponding indicators.
[0225] In this optional embodiment, the structural index and the business index may be updated based on the product of the weight and the index. For example, the updated A index is marked as A. new , A new =A·W A ; The updated B pointer is marked as B new , B new =B·W B ; The updated C pointer is marked as C new , C new =C.W C ; The updated G index is marked as G new , G new =G·WG .
[0226] In an optional embodiment, the second constructing unit 114 is configured to interpolate the update indicator according to an interpolation algorithm to construct an initial training data set.
[0227] In this optional embodiment, interpolating the update indicator according to the interpolation algorithm to construct the initial training data set includes:
[0228] Splitting the update index to obtain an update index interval set;
[0229] Fitting the data in the update index interval set according to the interpolation algorithm to obtain a function curve set;
[0230] Sorting the function curves according to the sampling time points corresponding to the starting points of the curves in the function curve set to obtain a sorted function curve set;
[0231] The curves in the sorting function curve set are connected end to end to obtain the initial training data set.
[0232] In this optional embodiment, for the updated indicator data, the cubic spline interpolation method can be used to obtain the data center structure indicator fluctuation curve and the data center business indicator fluctuation curve with a time span of 720 hours. The cubic spline interpolation method is a model fitting algorithm. Taking indicator A as an example, its main process is to divide the data point set of indicator A into n intervals. For example, since there are 720 sampling time points in total, n=719, and there is an interval between every two sampling time points.
[0233] In this optional embodiment, the 719 interval groups may be used as the interval set.
[0234] In this optional embodiment, on each interval, the points between the intervals should satisfy the cubic equation, S(x i )=y i =a i +b i ·x i +c i ·x i 2 +d i ·x i 3 , this equation is called a cubic spline function, where yi represents the value of the indicator at the i-th sampling time point, x iTo represent the value at the $i$-th time point, the cubic spline function should satisfy the following three conditions: all points need to satisfy the interpolation condition; the first derivatives and second derivatives of $n - 1$ internal points should be continuous; the second derivatives of the function curve at the two endpoints are 0. Based on these three conditions, the cubic spline function can be solved to obtain the coefficients $a$ i , $b$ i , $c$ i , $d$ i of the cubic spline equation. Further, the function curves in each interval can be fitted based on the coefficients of the cubic spline equation to obtain the set of function curves.
[0235] In this alternative embodiment, the curves in the set of function curves can be sorted according to the sampling time point corresponding to the starting point of the function curve to obtain a sorted set of function curves. If the sampling time point corresponding to the starting point of the function curve is earlier, the sorting position of this function curve is also earlier.
[0236] In this alternative embodiment, the head and tail of the function curves can be connected according to the order of the function curves in the sorted set of function curves to obtain the initial training data set.
[0237] In an alternative embodiment, the third construction unit 115 is used to construct a performance evaluation model based on the initial training data set.
[0238] In this alternative embodiment, a preset neural network model to be updated can be trained based on the initial training data set to obtain the performance evaluation model.
[0239] In this alternative embodiment, the data in the initial training data set can be collected according to a preset frequency to construct an updated training data set, the purpose of which is to convert a continuous time series into discrete data in time series for subsequent training of the preset neural network model to be updated. Exemplarily, the preset frequency can be 3000 Hz, then the dimension of the updated training data set is $7.776\times10$ 9 $\times4$, with a total of $7.776\times10$ 9 rows and 4 columns. Each row of data can be regarded as a $1\times4$ vector, denoted as $v$ i , $i\in[1, 7.776\times10$ 9 .
[0240] In this optional embodiment, the neural network model to be updated may be an LSTM network, and the LSTM network can be trained based on the updated training dataset. The LSTM is a temporal neural network, whose full name is Long Short Term Memory, meaning long short-term memory neural network. The LSTM network is composed of neurons, and the neurons are connected in series. Each neuron contains three input data and two output data. The input data includes: C t-1 , the memory information of the previous moment; h t-1 , the output data of the previous moment; x t , the input data of the current moment. The output data includes: C t , the memory information of the current moment; h t The output data of the current moment.
[0241] In this optional embodiment, each neuron contains three computing parts: forget gate; input gate; output gate. The main process of the forget gate is to input the output data h t-1 of the previous moment and the input data x t of the current moment into the Sigmoid function together, and the output result is between 0 and 1, denoted as σ 1 . The σ 1 can represent the importance of the memory of the previous moment. The main process of the input gate is to input the output data h t-1 of the previous moment and the input data x t of the current moment into the Sigmoid function together to obtain the importance factor of the current input, with a value between [0, 1], denoted as σ 2 ; input the output data h t-1 of the previous moment and the input data x t of this moment into the tanh function together to obtain the centralized value of the current input, with a value between [-1, 1], denoted as tanh 1 . The product between the σ 2 and tanh 1 is the state information C t of the current moment. The main process of the output gate is to input the output data h t-1 of the previous moment and the input data x t of the current moment into the Sigmoid function together, and the output value is between [0, 1], denoted as σ 3 . The σ 3 is used to evaluate how much output value the state information C t has. The higher the σ 3 , the higher the output value of C t . Multiply the state information C tEnter the tanh function and get the result [-1, 1], recorded as tanh 2 The output data h at the current moment t =σ 3 ·tanh 2 .
[0242] In this optional embodiment, the v 1 That is, the input value x of the LSTM network at time point t=1 1 , C t-1 With h t-1 The initial value of can be set to 0. Based on the training method of the LSTM network and the Train_data, a complete LSTM network can be trained. Developers can collect the structural indicators and business indicators of the data center at a certain moment at any time, input the complete LSTM network, and predict and evaluate the quantitative indicators of the data center in the future time period.
[0243] In this optional embodiment, the fully trained LSTM network can be used as the performance evaluation model.
[0244] In an optional embodiment, the evaluation unit 116 is used to collect real-time data and input it into the performance evaluation model to obtain the evaluation result of the data center.
[0245] In this optional embodiment, the collecting real-time data and inputting it into the performance evaluation model to obtain the evaluation result of the data center includes:
[0246] Collect relevant data of the data center in real time;
[0247] Build and update structural indicators and business indicators in real time, and get real-time updated indicators;
[0248] The real-time update index is input into the performance evaluation model to evaluate the data center and obtain the evaluation result.
[0249] In this optional embodiment, developers of the data middle platform can collect data related to the data middle platform in real time, and the data related to the data middle platform includes real-time structure data and real-time log data.
[0250] In this optional embodiment, the real-time structure data can be a piece of data containing four dimensions, and the four dimensions include the total number of tables that comply with the naming conventions in the data center at a certain moment, the number of all tables in the data center, the total number of post-dependencies of all tables in the data center, and the total number of post-dependencies of all tables in the ODS layer of the data center.
[0251] In this optional embodiment, the real-time log data includes the total number of times all application programs are called by the middleware platform within the past sampling duration, the number of times the application program is called within the past sampling duration recorded at the sampling moment in the operation log of the data middleware platform, the total running duration of the application program within the past sampling duration, the number of error reports of the application program within the past sampling duration, and the total CPU occupation duration of the application program within the past sampling duration.
[0252] In this optional embodiment, the real-time structure metrics and real-time business metrics can be constructed based on the first construction unit and the calculation unit using the real-time structure data and the real-time log data.
[0253] In this optional embodiment, the real-time structure metrics and real-time business metrics can be updated based on the update unit to obtain real-time updated metrics.
[0254] In this optional embodiment, the real-time updated metrics can be input into the performance evaluation model to obtain the volatility of the data middleware platform metrics within a certain period in the future, and further, the performance volatility of the data middleware platform within a certain period in the future can be quantitatively evaluated.
[0255] In this optional embodiment, when the cosine distance between the volatility of the data middleware platform metrics and the preset threshold is greater than 0.5, the evaluation result is "unqualified"; if the cosine distance between the volatility of the data middleware platform metrics and the preset threshold is not greater than 0.5, the evaluation result is "qualified".
[0256] As Figure 7 shown, it is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 1 includes a memory 12 and a processor 13. The memory 12 is used to store computer-readable instructions, and the processor 13 is used to execute the computer-readable instructions stored in the memory to implement the data middleware platform evaluation method based on artificial intelligence in any of the above embodiments.
[0257] In an optional embodiment, the electronic device 1 further includes a bus and a computer program stored in the memory 12 and executable on the processor 13, such as a data middleware platform evaluation program based on artificial intelligence.
[0258] Figure 7 Only the electronic device 1 with components 12 - 13 is shown. Those skilled in the art can understand that Figure 7 the shown structure does not constitute a limitation on the electronic device 1, and it may include fewer or more components than shown, or combine some components, or have different component arrangements.
[0259] Combined with Figure 1, the memory 12 in the electronic device 1 stores multiple computer-readable instructions to implement an artificial intelligence-based data middle platform evaluation method, and the processor 13 can execute multiple instructions to implement:
[0260] Build a data middle platform to obtain a data set, where the data set includes structured data and log data;
[0261] Parse the structured data in the data set to construct structure indicators;
[0262] Calculate the weights of the log data in the data set to construct business indicators;
[0263] Update the structure indicators and the business indicators to obtain updated indicators;
[0264] Interpolate the updated indicators according to the interpolation algorithm to construct an initial training data set;
[0265] Construct a performance evaluation model based on the initial training data set;
[0266] Collect real-time data, input it into the performance evaluation model, and obtain the evaluation result of the data middle platform.
[0267] Specifically, the specific implementation method of the processor 13 for the above instructions can refer to Figure 1 the description of the relevant steps in the corresponding embodiments, which will not be elaborated here.
[0268] Those skilled in the art can understand that the schematic diagram is only an example of the electronic device 1, and does not constitute a limitation on the electronic device 1. The electronic device 1 can be either a bus structure or a star structure. The electronic device 1 can also include more or fewer other hardware or software than shown in the figure, or different component arrangements. For example, the electronic device 1 can also include input / output devices, network access devices, etc.
[0269] It should be noted that the electronic device 1 is only an example. Other existing or future possible electronic products that can be adapted to this application should also be included in the protection scope of this application and are included herein by reference.
[0270] Among them, the memory 12 includes at least one type of readable storage medium, which can be non-volatile or volatile. The readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 12 can be an internal storage unit of the electronic device 1, such as the mobile hard disk of the electronic device 1. In other embodiments, the memory 12 can also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the electronic device 1. Further, the memory 12 can also include both the internal storage unit and the external storage device of the electronic device 1. The memory 12 can be used not only to store application software and various types of data installed in the electronic device 1, such as the code of the data middle platform evaluation program based on artificial intelligence, etc., but also to temporarily store the data that has been output or will be output.
[0271] In some embodiments, the processor 13 can be composed of integrated circuits. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions packaged, including the combination of one or more central processing units (CPU), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 13 is the control core (Control Unit) of the electronic device 1, connecting all components of the entire electronic device 1 through various interfaces and circuits. By running or executing the programs or modules stored in the memory 12 (such as executing the data middle platform evaluation program based on artificial intelligence, etc.), and calling the data stored in the memory 12, it can execute various functions of the electronic device 1 and process data.
[0272] The processor 13 executes the operating system of the electronic device 1 and various installed application programs. The processor 13 executes the application programs to implement the steps in the above-mentioned embodiments of various data middle platform evaluation methods based on artificial intelligence, such as Figures 1 - 5 the steps shown.
[0273] Exemplarily, the computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory 12 and executed by the processor 13 to complete the present application. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 1. For example, the computer program may be divided into an acquisition unit 110, a first construction unit 111, a calculation unit 112, an update unit 113, a second construction unit 114, a third construction unit 115, and an evaluation unit 116.
[0274] The integrated units implemented in the form of software function modules as described above may be stored in a computer-readable storage medium. The above-mentioned software function modules stored in a storage medium include several instructions for causing a computer device (which may be a personal computer, a computer device, or a network device, etc.) or a processor to execute a part of the method for evaluating an artificial-intelligence-based data middle platform according to each embodiment of the present application.
[0275] If the modules / units integrated in the electronic device 1 are implemented in the form of software function units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it may also be completed by a computer program instructing relevant hardware devices. The computer program may be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned various method embodiments may be implemented.
[0276] Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory, and other memories, etc.
[0277] Furthermore, the computer-readable storage medium mainly includes a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of blockchain nodes, etc.
[0278] The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. A blockchain, essentially a decentralized database, is a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, an application service layer, etc.
[0279] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, in Figure 7 it is only represented by one arrow, but it does not mean that there is only one bus or one type of bus. The bus is arranged to implement the connection and communication between the memory 12 and at least one processor 13, etc.
[0280] Although not shown, the electronic device 1 may further include a power supply (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the at least one processor 13 through a power management device, so as to implement functions such as charging management, discharging management, and power consumption management through the power management device. The power supply may further include any components such as one or more DC or AC power supplies, a recharge device, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The electronic device 1 may further include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0281] The embodiment of this application also provides a computer-readable storage medium (not shown in the figure). Computer-readable instructions are stored in the computer-readable storage medium, and the computer-readable instructions are executed by a processor in the electronic device to implement the artificial intelligence-based data middle platform evaluation method described in any of the above embodiments.
[0282] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.
[0283] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0284] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical unit, that is, it may be located in one place, or may be distributed across multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0285] In addition, in each embodiment of the present application, the functional modules can be integrated in a processing unit, can also exist separately as individual physical units, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.
[0286] In addition, it is obvious that the term "including" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in the specification can also be implemented by one unit or device through software or hardware. Terms such as first, second, etc. are used to denote names and do not denote any particular order.
[0287] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them. Although the present application has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application.
Claims
1. An artificial intelligence-based data middle platform evaluation method, characterized in that, the method includes: Construct a data middle platform to obtain a data set, where the data set includes structured data and log data; Parse the structured data in the data set to construct structure indicators; Calculate the weights of the log data in the data set to construct business indicators, including: calculating the ratio of the number of times an application in the data middle platform is called to the total number of times all applications are called, and taking this ratio as the first weight of the application; calculating the ratio of the total duration of a central processing unit occupied by an application in the data middle platform to the total running duration of the application, and taking this ratio as the second weight of the application; calculating the ratio of the number of times an application in the data middle platform reports an error to the total number of times the application is called, and calculating the difference between a preset harmonic real number and this ratio, and taking this difference as the third weight of the application; constructing business indicators based on the first weight, the second weight and the third weight, and the business indicators are used to characterize the performance of the applications in the data middle platform during operation; Update the structure indicators and the business indicators to obtain updated indicators; Interpolate the updated indicators according to the interpolation algorithm to construct an initial training data set; Construct a performance evaluation model based on the initial training data set; Collect real-time data and input it into the performance evaluation model to obtain the evaluation result of the data middle platform.
2. The artificial intelligence-based data middle platform evaluation method according to claim 1, characterized in that, the constructing a data middle platform to obtain a data set includes: Collect the structured data and log data of the data middle platform according to preset data sampling time points; Jointly store the structured data and the log data to obtain a data set.
3. The artificial intelligence-based data middle platform evaluation method according to claim 1, characterized in that, the structure indicators include: A normalization rate indicator, which is the ratio of the number of tables that conform to the table naming specification to the total number of all tables in the data middle platform; A reuse rate indicator, which is the ratio of the number of tables with pre-dependencies to the total number of all tables in the data middle platform, and the pre-dependency means that the data in this table is obtained based on the data in other tables; A coverage rate indicator, which is the ratio of the number of tables that do not have pre-dependencies but have post-dependencies to the total number of all tables in the data middle platform, and the post-dependency means that the data in other tables is obtained based on the data in this table.
4. The artificial intelligence-based data middle platform evaluation method according to claim 1, characterized in that, the updating the structure indicators and the business indicators to obtain updated indicators includes: Construct a covariance matrix based on the structure indicators and the business indicators; Calculate the eigenvalues of the covariance matrix to obtain the weights of the structure indicators and the business indicators; Update the structure indicators and the business indicators according to the weights to obtain updated indicators.
5. The artificial intelligence-based data middle platform evaluation method according to claim 1, characterized in that, the interpolating the updated indicators according to the interpolation algorithm to construct an initial training data set includes: Split the update metrics to obtain a set of update metric intervals; Fit the data within the set of update metric intervals according to the interpolation algorithm to obtain a set of function curves; Sort the function curves according to the sampling time points corresponding to the starting points of the curves in the set of function curves to obtain a sorted set of function curves; Connect the curves in the sorted set of function curves end to end to obtain the initial training dataset.
6. The method for evaluating a data middle platform based on artificial intelligence according to claim 1, wherein, collecting the real-time data and inputting the real-time data into the performance evaluation model to obtain the evaluation result of the data middle platform includes: Collecting the data related to the data middle platform in real time, where the related data includes real-time structure data and real-time log data; Constructing and updating the structure metrics and business metrics in real time to obtain real-time update metrics; Inputting the real-time update metrics into the performance evaluation model to evaluate the data middle platform and obtaining the evaluation result.
7. A data middle platform evaluation device based on artificial intelligence, the device includes units for implementing the method according to any one of claims 1 to 6, wherein, the device includes: An obtaining unit, configured to build a data middle platform to obtain a dataset, where the dataset includes structure data and log data; A first construction unit, configured to parse the structure data in the dataset to construct structure metrics; A calculation unit, configured to calculate the weights of the log data in the dataset to construct business metrics; An update unit, configured to update the structure metrics and the business metrics to obtain update metrics; A second construction unit, configured to perform interpolation on the update metrics according to the interpolation algorithm to construct an initial training dataset; A third construction unit, configured to construct a performance evaluation model according to the initial training dataset; An evaluation unit, configured to collect real-time data and input the real-time data into the performance evaluation model to obtain the evaluation result of the data middle platform.
8. An electronic device, wherein, the electronic device includes: A memory, storing computer-readable instructions; and A processor, configured to execute the computer-readable instructions stored in the memory to implement the method for evaluating a data middle platform based on artificial intelligence according to any one of claims 1 to 6.
9. A computer-readable storage medium, wherein: Computer-readable instructions are stored in the computer-readable storage medium, and the computer-readable instructions are executed by a processor in an electronic device to implement the method for evaluating a data middle platform based on artificial intelligence according to any one of claims 1 to 6.
Citation Information
Patent Citations
Server health assessment method, system, equipment and medium
CN113806171A