Data sampling method and device for index log, equipment and medium

By obtaining the target data set in the index log library and using convolutional neural network processing to form a sampling data matrix, the problem of high data dimensions and redundancy in traditional methods is solved, and efficient data sampling is achieved.

CN120296315APending Publication Date: 2025-07-11BEIJING YOUTEJIE INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510372978.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When traditional data sampling methods process index log data with high dimensional, high frequency and real-time characteristics, they have problems such as high data dimensions and redundancy and low sampling efficiency.

Method used

The target data set is obtained based on the data acquisition rules set by the user, the original tensor matrix is formed using preset data sorting rules, and these matrices are processed using a convolutional neural network to obtain the sampled data matrix, and finally an aggregation operation is performed to generate the sampled data results.

Benefits of technology

The data dimension and redundancy of the indicator log are reduced, the efficiency of data sampling is improved, and the real-time requirement is met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296315A_ABST
    Figure CN120296315A_ABST
Patent Text Reader

Abstract

The invention discloses a data sampling method and device for index logs, equipment and a medium. The method comprises the following steps: acquiring at least one target data set in an index log library based on a data acquisition rule set by a user; sorting each target data set according to a preset data sorting rule to obtain an original tensor matrix matched with each target data set; acquiring a data sampling demand of a user, and processing each original tensor matrix based on the data sampling demand by using a preset convolutional neural network to obtain each sampling data matrix matched with each original tensor matrix; and performing aggregation operation on each sampling data matrix to obtain a sampling data result matched with the data sampling demand. By means of the technical scheme, data sampling of the index logs can be achieved, the data dimension and the data redundancy of the index logs are reduced, and the index data sampling efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and particularly to a data sampling method, device, equipment and medium for index logs. Background Art

[0002] At present, with the rapid development of big data and cloud computing technologies, the amount of log data generated by various systems is showing an explosive growth. Among them, index log data, due to its unique nature, poses great challenges to storage, analysis and monitoring.

[0003] Index log data has the remarkable characteristics of high-dimensionality, high-frequency and real-time nature. High-dimensionality means that the data contains numerous feature dimensions, making the data structure extremely complex; high-frequency indicates that the data is generated at a very high frequency, and a huge amount of data will accumulate in a short time; real-time nature requires that these data be processed and analyzed in a timely manner to meet the needs of real-time monitoring and decision-making support for business. Traditional random sampling methods, although simple to operate, due to the randomness of the sampling process, are prone to ignoring those important but low-frequency data points. Periodic sampling seemingly has a certain degree of structure and regularity, but it samples at fixed time intervals and may miss abnormal fluctuations or sudden events occurring between two sampling periods. In addition, neither random sampling nor periodic sampling can effectively cope with the high-dimensional characteristics of index log data. High-dimensional data not only increases the storage burden but also makes the data processing efficiency extremely low. The complex high-dimensional data structure will make subsequent analysis algorithms extremely complex, consume a huge amount of computing resources, and it is difficult to meet the real-time requirements.

[0004] In summary, traditional data sampling methods have problems such as relatively high data dimensions and data redundancy of index logs, and relatively low efficiency of index data sampling when processing index log data with the characteristics of high-dimensionality, high-frequency and real-time nature. Summary of the Invention

[0005] The present invention provides a data sampling method, device, equipment and medium for index logs, which can solve the problems of relatively high data dimensions and data redundancy of index logs, and relatively low efficiency of index data sampling existing in the existing data sampling methods when sampling index log data.

[0006] In a first aspect, an embodiment of the present invention provides a data sampling method for index logs, and the method includes:

[0007] Obtain at least one target data set from an index log library based on a data acquisition rule set by a user;

[0008] Sort out each target data set according to a preset data sorting rule to obtain an original tensor matrix respectively matching each target data set;

[0009] Obtain the data sampling requirements of the user, and use a pre-set convolutional neural network to process each original tensor matrix based on the data sampling requirements to obtain each sampling data matrix respectively matching each original tensor matrix;

[0010] Perform an aggregation operation on each sampling data matrix to obtain a sampling data result matching the data sampling requirements.

[0011] In a second aspect, an embodiment of the present invention provides a data sampling device for metric logs, and the device includes:

[0012] A dataset acquisition module, configured to acquire at least one target dataset in a metric log library based on a data acquisition rule set by a user;

[0013] A data arrangement module, configured to arrange each target dataset according to a preset data arrangement rule to obtain an original tensor matrix respectively matching each target dataset;

[0014] A convolutional neural network module, configured to obtain the data sampling requirements of the user, and use a pre-set convolutional neural network to process each original tensor matrix based on the data sampling requirements to obtain each sampling data matrix respectively matching each original tensor matrix;

[0015] A result generation module, configured to perform an aggregation operation on each sampling data matrix to obtain a sampling data result matching the data sampling requirements.

[0016] In a third aspect, an embodiment of the present invention provides an electronic device, and the electronic device includes:

[0017] At least one processor; and

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute a data sampling method for metric logs according to any embodiment of the present invention.

[0020] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, and the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement a data sampling method for metric logs according to any embodiment of the present invention when executed.

[0021] The technical solution of the embodiment of the present invention obtains at least one target data set in the metric log library based on the data acquisition rule set by the user, then sorts out each target data set according to the preset data sorting rule to obtain the original tensor matrix respectively matching each target data set, then obtains the user's data sampling requirement, and uses the pre-set convolutional neural network to process each original tensor matrix based on the data sampling requirement to obtain each sampling data matrix respectively matching each original tensor matrix, and finally aggregates each sampling data matrix to obtain a sampling data result matching the data sampling requirement, solving the problems existing in the existing data sampling method when sampling metric log data, such as the high data dimension and data redundancy of the metric log, and the low efficiency of metric data sampling, realizing the data sampling of the metric log, reducing the data dimension and data redundancy of the metric log, and improving the efficiency of metric data sampling.

[0022] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Brief Description of the Drawings

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0024] Figure 1 It is a flowchart of a method for sampling metric log data according to Embodiment 1 of the present invention;

[0025] Figure 2 It is a flowchart of a method for sampling metric log data according to Embodiment 2 of the present invention;

[0026] Figure 3 It is a schematic structural diagram of a device for sampling metric log data according to Embodiment 3 of the present invention;

[0027] Figure 4 It is a schematic structural diagram of an electronic device for implementing the method for sampling metric log data of the embodiment of the present invention. Detailed Embodiments

[0028] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solution in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0029] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" any variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0030] Embodiment 1

[0031] Figure 1 FIG. 10 is a flowchart of a data sampling method for index logs provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of data sampling for index logs. This method can be executed by a data sampling device for index logs. The data sampling device for index logs can be implemented in the form of hardware and / or software, and the data sampling device for index logs can be configured in a terminal or server with the data sampling function of index logs.

[0032] As Figure 1 shown, the method includes:

[0033] S110. Obtain at least one target data set from the index log library based on the data acquisition rule set by the user.

[0034] Among them, obtaining at least one target data set from the index log library based on the data acquisition rule set by the user includes: obtaining the index information type of each index log according to the log attributes of each index log in the index log library; if it is determined that the index information type of the index log matches the data acquisition rule, then set the index log as a target index log; extract the numerical data in each target index log, and obtain each target data set based on each numerical data and the target index log respectively matched with each numerical data.

[0035] Specifically, the metric logs record various metrics during the system operation, and each metric log has unique log attributes; the log attributes are a set of information describing the characteristics of the metric logs, such as the timestamp recording the specific time when the metric data is generated, the metric information type, the data type, etc. Through these log attributes, the metric information type of each metric log can be obtained; the metric information type is used to distinguish the information categories represented by different metric logs. For example, the metric information type can be CPU usage rate, disk I / O, memory usage rate, etc.

[0036] In a specific implementation manner of this embodiment, when the user sets a data acquisition rule according to their own business requirements, the system will filter data in the metric log library according to this rule. The data acquisition rule is a filtering condition formulated by the user according to the actual business scenario and analysis purpose, which stipulates which types of metric log data the system should acquire. For example, the user may set to only acquire metric log data related to CPU usage rate and memory usage rate. Then, the CPU usage rate and memory usage rate are the data acquisition rules set by the user. After that, the system will judge one by one whether the metric information type of each metric log in the metric log library matches the data acquisition rule. Exemplarily, if the metric information type of a certain metric log conforms to the data acquisition rule set by the user, for example, the metric information type of this metric log is CPU usage rate, and the data acquisition rule set by the user contains CPU usage rate, then this metric log will be set as the target metric log. After determining the target metric logs, the system will extract the numerical data in these target metric logs. The numerical data is the data with numerical characteristics in the metric logs, such as the specific values "75.3", "80.1", etc. corresponding to the CPU usage rate. These numbers can intuitively reflect the status of the system-related metrics. Based on each numerical data and the target metric logs respectively matched with each numerical data, the system starts to construct the target data set.

[0037] Further, based on each numerical data and the target metric logs respectively matched with each numerical data to obtain each target data set, including: binding and storing each numerical data and the metric information type of the target metric log respectively matched with each numerical data to obtain each original metric data respectively matched with each target metric log; performing a data cleaning operation on each original metric data based on a preset data cleaning rule to obtain each cleaned metric data; aggregating and storing the cleaned metric data with the same metric information type to obtain each target data set, where the target data set includes the metric information type and at least one numerical data.

[0038] On the basis of the above steps, in the process of constructing the target data set, first, each digital data is bound and stored with the index information type of the target index log respectively matched with it. For example, the digital data "75.3" is bound with the index information type of CPU usage rate in the target index log and stored as CPU usage rate: 75.3. In this way, each original index data respectively matched with each target index log is obtained. Although these original index data contain useful information, there may be some problems. For example, there may be null values in the data (such as the CPU usage rate record at a certain time point is null), outliers (values significantly deviating from the normal range), etc. Further, in order to ensure the quality of the data, the system will perform data cleaning operations on each original index data based on preset data cleaning rules. The data cleaning rules are a set of rules preset for dealing with problems in the original data. For example, null value data is deleted, and the data recorded as null is directly removed from the data set; for outliers, statistical methods can be used for identification and processing, which may be to replace the outliers with reasonable estimated values or directly delete them. This embodiment does not limit this. After the data cleaning operation, the obtained is each cleaned index data. Finally, the system will aggregate and store the cleaned index data with the same index information type, so as to obtain each target data set. For example, all the cleaned index data of CPU usage rate are aggregated together to form a target data set of CPU usage rate. This target data set not only contains the index information type of CPU usage rate, but also contains a series of digital data, such as "75.3", "80.1", "79.8", etc. In this way, through a series of operations, the system successfully obtains at least one target data set in the index log library based on the data acquisition rules set by the user, providing high-quality data support for subsequent data analysis, system performance diagnosis and other work.

[0039] S120. Sort each target data set according to the preset data sorting rules to obtain the original tensor matrices respectively matched with each target data set.

[0040] Among them, sorting each target data set according to the preset data sorting rules to obtain the original tensor matrices respectively matched with each target data set includes: searching for the target index logs respectively matched with each digital data, and obtaining the log attributes of the target index logs as the data attributes of the digital data matched with the target index logs; extracting the required attribute information matched with the digital data from the data attributes according to the preset data sorting rules; sorting each digital data in the target data set according to the data sorting rules and the required attribute information of each digital data to obtain the original tensor matrix matched with the target data set.

[0041] In a specific implementation manner of this embodiment, when organizing data, it is first necessary to search for target metric logs that respectively match each digital data. For example, in a target dataset, there is digital data "75.3", and its corresponding target metric log may be a log recording the CPU usage rate data at a certain moment, that is, the metric log that generates the above digital data "75.3" in step S120. After finding this target metric log, obtain its log attributes, and these log attributes serve as the data attributes of the digital data that matches this target metric log. Next, according to the preset data organization rules, extract the required attribute information that matches the digital data from these data attributes, where the preset data organization rules are a series of specifications and requirements formulated according to specific business requirements and analysis purposes. For example, if the focus of the analysis is the resource usage of the system at different times, then the data organization rules may stipulate that the two attribute information of timestamp and metric information type be extracted as the required attribute information. Taking the target dataset related to CPU usage rate as an example, extract the timestamp and metric information type from the log attributes of the target metric log corresponding to each digital data. For the digital data "75.3", extract the timestamp "2025-03-10 10:00:00" and the metric information type CPU usage rate. Finally, based on the data organization rules and the required attribute information of each digital data, perform a sorting operation on each digital data in the target dataset, and then obtain the original tensor matrix that matches the target dataset. The way of the sorting operation is various, depending on the data organization rules and actual requirements. For example, if the data organization rules require sorting the digital data in chronological order, then in the target dataset of CPU usage rate, according to the timestamp corresponding to each digital data, the digital data will be arranged in chronological order. Assuming that there are multiple digital data of CPU usage rate and their corresponding timestamps in this target dataset, they will be arranged in order from earliest to latest to form an ordered sequence, and this sequence is a simple form of the original tensor matrix. In practical applications, the original tensor matrix may be arranged more complexly according to the dimensions of the data and analysis requirements. This embodiment does not limit the formation rules of the original tensor matrix, and it can be set and modified by relevant personnel according to actual needs. For example, group according to the metric information type, and then arrange the digital data in chronological order within each group to form a multi-dimensional tensor structure, so as to be better applied to subsequent operations such as the pooling layer processing of the convolutional neural network, providing strong support for log analysis and system performance diagnosis.

[0042] S130. Obtain the data sampling requirements of the user, and use the pre-set convolutional neural network to process each original tensor matrix based on the data sampling requirements to obtain each sampling data matrix that respectively matches each original tensor matrix.

[0043] S140: Aggregate each sampling data matrix to obtain a sampling data result that matches the data sampling requirement.

[0044] Optionally, after aggregating the sampling data matrices to obtain sampling data results that match the data sampling requirements, the method further includes: visually displaying the sampling data results based on visualization rules set by the user to obtain visualization icons; wherein the visualization rules include: icon type, color scheme, and indicator information type display method.

[0045] The aggregation operation is an operation that combines and aggregates multiple data units with similar structures or associations during data processing to form a more comprehensive and generalized data result. In the above-mentioned index log data processing scenario, the aggregation operation is performed on multiple sampled data matrices.

[0046] In this embodiment, the method of the aggregation operation can be row-by-row aggregation, column-by-column aggregation, and aggregation based on specific indicators, etc., which can be set by relevant personnel according to the actual implementation scenario, and this embodiment does not limit this. Specifically, row aggregation is to merge the corresponding rows of each sampling data matrix. For example, suppose there are two sampling data matrices A and B, the first row of matrix A is [75.3, 60.7], representing the values ​​of CPU usage and "memory usage" at a certain moment, and the first row of matrix B is [80.1, 62.4], which also represents the value of the same indicator at another moment. When aggregating by row, these two rows of data can be merged into a new row vector [75.3, 60.7, 80.1, 62.4], and so on, all rows of the sampling data matrix are subjected to such a merging operation to form a larger sampling data result matrix. Aggregating by column is to merge the columns of each matrix, such as summarizing the column data representing the CPU usage in all matrices together to form a new column vector to highlight the changes of a certain indicator at different times or under different conditions. Aggregation based on specific indicators is more flexible. It can filter and summarize data of specific indicators according to data sampling requirements. For example, if the data sampling requirements focus on the disk IO indicators of the system under high load, then the data of the disk IO indicators can be extracted from each sampling data matrix and aggregated according to certain rules (such as time sequence, load level, etc.) to obtain a sampling data result specifically for the disk IO indicator.

[0047] Optionally, after the aggregation operation of each sampled data matrix is completed and the sampled data result that matches the data sampling requirement is obtained, in order to more intuitively display the data information and facilitate user understanding and analysis, the sampled data result can also be visually displayed based on the visualization rules set by the user, and then a visualization chart is obtained. The visualization rules are a series of specifications formulated by the user according to their own understanding and display requirements of the data, and mainly include aspects such as chart type, color scheme, and display method of index information type. By visually displaying the sampled data result based on the visualization rules set by the user, the generated visualization chart can present complex data to the user in an intuitive and easy-to-understand manner, helping the user analyze the data more efficiently, discover problems, and thus better meet various needs of the user in data sampling, providing strong support for system performance optimization, business decision-making, etc.

[0048] In the technical solution of the embodiment of the present invention, at least one target data set is obtained from the index log library based on the data acquisition rule set by the user, and then each target data set is sorted according to the preset data sorting rule to obtain an original tensor matrix respectively matching each target data set. Then, the data sampling requirement of the user is obtained, and each original tensor matrix is processed based on the data sampling requirement using a pre-set convolutional neural network to obtain each sampled data matrix respectively matching each original tensor matrix. Finally, an aggregation operation is performed on each sampled data matrix to obtain a sampled data result that matches the data sampling requirement, realizing data sampling of the index log, reducing the data dimension and data redundancy of the index log, and improving the efficiency of index data sampling.

[0049] Embodiment Two

[0050] Figure 2 It is a flowchart of a method for data sampling of index logs provided in Embodiment Two of the present invention. This embodiment is refined based on the above embodiment. Specifically, in this embodiment, the method of using a pre-set convolutional neural network to process each original tensor matrix based on the data sampling requirement to obtain each sampled data matrix respectively matching each original tensor matrix is refined.

[0051] As Figure 2 shown, the method includes:

[0052] S210. Obtain at least one target data set from the index log library based on the data acquisition rule set by the user.

[0053] S220. Sort each target data set according to the preset data sorting rule to obtain an original tensor matrix respectively matching each target data set.

[0054] S230. Obtain the data sampling requirements of the user, and respectively obtain each pooling rule in the data sampling requirements as the target pooling rule.

[0055] Among them, the data sampling requirements include: at least one pooling rule, a pooling window size respectively matched with each pooling rule, and a pooling window step length respectively matched with each pooling rule. The pooling rule includes a pooling type and an index information type matched with the pooling type.

[0056] Furthermore, the pooling type may include: max pooling and / or average pooling. Max pooling means selecting the maximum value within the data window as the output, and this method can effectively extract local prominent features. For example, when processing the index data of CPU usage, if the values within a data window are [75.3, 76.5, 74.8], after the max pooling operation, the output result is 76.5, which can highlight the peak situation of CPU usage during this time period. While average pooling calculates the average value of all the values within the data window as the output, which reflects the overall trend of the data within the window. For example, for the above data window [75.3, 76.5, 74.8], the result of average pooling is (75.3 + 76.5 + 74.8) ÷ 3 ≈ 75.53, enabling the user to understand the average level of CPU usage during this time period.

[0057] Further, the pooling window size determines the data range selected each time the pooling operation is performed. For example, if the pooling window size is set to 3, it means that 3 consecutive data will be selected for pooling calculation each time. Taking the CPU usage data sequence [75.3, 76.5, 80.1, 79.8, 82.4] as an example, when the pooling window size is 3, the first pooling operation will be performed on the 3 data of [75.3, 76.5, 80.1]. The pooling window stride determines the interval at which the pooling operation moves on the data sequence. If the stride is 2, then in the above example, the first pooling operation is for [75.3, 76.5, 80.1], and the second will be for [79.8, 82.4] (skipping one data before [79.8]) for the pooling operation. Appropriate settings of the pooling window size and stride can effectively reduce the data volume while retaining the key features of the data, improving the efficiency of data analysis. It should be noted that the specific content of the pooling rule, the pooling window size, and the pooling window stride can all be set by relevant personnel according to the actual implementation scenario, and this embodiment does not limit them. By accurately obtaining the user's data sampling requirements and separately obtaining each pooling rule from them as the target pooling rule, including elements such as the pooling type, index information type, pooling window size, and pooling window stride, it lays a solid foundation for subsequent data processing and analysis work, helping to extract valuable information from a large amount of index-based log data more efficiently and accurately to meet the user's data analysis needs.

[0058] S240. Search for the original tensor matrix that matches the pooling rule in each original tensor matrix according to the index information type in the target pooling rule as the target tensor matrix.

[0059] Exemplarily, in an actual large-scale data processing scenario, there may be multiple original tensor matrices, and each matrix may contain different index information. Taking a complex cloud computing system as an example, after the index-based log data generated by it is processed, it may form original tensor matrices that respectively record different aspects such as system resource usage and network traffic conditions. At this time, it is particularly important to search according to the index information type in the target pooling rule. Suppose the index information type in the target pooling rule is "CPU usage rate", the system will traverse all the original tensor matrices and check the index information in each matrix. If a certain original tensor matrix contains a data column related to "CPU usage rate", then this matrix is determined as the target tensor matrix that matches the pooling rule. This search operation ensures that the subsequent pooling process is performed on specific index data, avoiding ineffective processing of irrelevant data and improving the pertinence and efficiency of data processing.

[0060] S250. Obtain the target window size and target window stride that match the target pooling rule in the data sampling requirements, and configure the pre-set convolutional neural network according to the target pooling rule, target window size, and target window stride.

[0061] Among them, the convolutional neural network (CNN for short) is a powerful deep learning model with significant advantages in data processing and feature extraction. The pre-set convolutional neural network is a network structure that has been initially built and parameter-initialized. Configuring it according to the obtained target pooling rule, target window size, and target window stride means passing these parameters to the pooling layer in the convolutional neural network. The pooling layer is an important part of the CNN, responsible for dimensionality reduction of data, reducing the amount of data while retaining key features. Through configuration, the pooling layer can process the input data according to the rules set by the user. For example, if the target pooling rule is max pooling, the target window size is 3, and the target window stride is 2, then in the pooling layer, a group of 3 data will be used, and each time the interval of 2 data is moved, the maximum value in each group will be selected from the input data as the output, thereby realizing the pooling operation on the data. This configuration enables the convolutional neural network to better adapt to specific data sampling requirements and improve the accuracy and efficiency of data processing.

[0062] S260. Perform pooling processing on the original tensor matrix through the configured convolutional neural network to obtain a sampled data matrix that matches the original tensor matrix.

[0063] Specifically, based on the above steps, when the original tensor matrix is passed as input to the convolutional neural network, the pooling layer in the network will perform pooling operations on it according to the configured target pooling rule, target window size, and target window stride. Taking a simple original tensor matrix [75.3, 76.5, 80.1, 79.8, 82.4, 81.5, 78.9, 80.2] as an example, if max pooling is used, the target window size is 2, and the target window stride is 2, the pooling layer will first process the data group [75.3, 76.5], select the maximum value 76.5 as the output; then move the window and process [80.1, 79.8], output the maximum value 80.1, and so on. After such processing, the original tensor matrix is converted into a new sampled data matrix [76.5, 80.1, 82.4, 80.2]. This sampled data matrix is obtained through pooling operations on the basis of the original tensor matrix. It retains the key features of the original data while reducing the data volume, making it more convenient for subsequent data analysis and processing. For example, when performing system performance monitoring and analysis, the sampled data matrix can be more efficiently used for tasks such as anomaly detection and trend analysis to help users quickly discover potential problems and patterns in the system. Through this series of operations, from finding the target tensor matrix according to the target pooling rule, configuring the convolutional neural network, to obtaining the sampled data matrix through pooling processing, a complete data processing flow is formed, which can effectively meet the user's data sampling requirements and provide strong support for subsequent data analysis and decision-making.

[0064] S270. Aggregate each sampled data matrix to obtain a sampled data result that matches the data sampling requirements.

[0065] In the technical solution of the embodiment of the present invention, at least one target data set is obtained from the index log library based on the data acquisition rule set by the user. Then, according to the preset data sorting rule, each target data set is sorted to obtain an original tensor matrix respectively matching each target data set. Then, the data sampling requirement of the user is obtained, and each pooling rule is obtained as a target pooling rule in the data sampling requirement. And according to the index information type in the target pooling rule, an original tensor matrix matching the pooling rule is found in each original tensor matrix as a target tensor matrix. Then, a target window size and a target window step length matching the target pooling rule are obtained in the data sampling requirement. And the pre-set convolutional neural network is configured according to the target pooling rule, the target window size and the target window step length. Then, the original tensor matrix is subjected to pooling processing through the configured convolutional neural network to obtain a sampling data matrix matching the original tensor matrix. Finally, an aggregation operation is performed on each sampling data matrix to obtain a sampling data result matching the data sampling requirement, realizing the data sampling of the index log, reducing the data dimension and data redundancy of the index log, and improving the efficiency of index data sampling.

[0066] Embodiment III

[0067] Figure 3 FIG. 7 is a structural schematic diagram of a data sampling device for index logs provided in Embodiment III of the present invention.

[0068] As Figure 3 shown, the device includes:

[0069] A data set acquisition module 310, configured to obtain at least one target data set from the index log library based on a data acquisition rule set by a user;

[0070] A data sorting module 320, configured to sort each target data set according to a preset data sorting rule to obtain an original tensor matrix respectively matching each target data set;

[0071] A convolutional neural network module 330, configured to obtain a data sampling requirement of a user, and use a pre-set convolutional neural network to process each original tensor matrix based on the data sampling requirement to obtain each sampling data matrix respectively matching each original tensor matrix;

[0072] A result generation module 340, configured to perform an aggregation operation on each sampling data matrix to obtain a sampling data result matching the data sampling requirement.

[0073] The technical solution of the embodiment of the present invention obtains at least one target data set in the index log library based on the data acquisition rule set by the user, then sorts out each target data set according to the preset data sorting rule to obtain an original tensor matrix respectively matched with each target data set, then obtains the data sampling requirement of the user, and uses the preset convolutional neural network to process each original tensor matrix based on the data sampling requirement to obtain each sampling data matrix respectively matched with each original tensor matrix, and finally performs an aggregation operation on each sampling data matrix to obtain a sampling data result matched with the data sampling requirement, realizing data sampling of the index log, reducing the data dimension and data redundancy of the index log, and improving the efficiency of index data sampling.

[0074] Based on the above embodiment, the data set acquisition module 310 includes:

[0075] The information type acquisition unit is used to obtain the index information type of each index log according to the log attribute of each index log in the index log library;

[0076] The target setting unit is used to set the index log as a target index log if it is determined that the index information type of the index log matches the data acquisition rule;

[0077] The data extraction unit is used to extract the digital data in each target index log and obtain each target data set based on each digital data and the target index log respectively matched with each digital data.

[0078] Based on the above embodiment, the data extraction unit includes:

[0079] The information binding unit is used to bind and store each digital data and the index information type of the target index log respectively matched with each digital data to obtain each original index data respectively matched with each target index log;

[0080] The data cleaning unit is used to perform a data cleaning operation on each original index data based on the preset data cleaning rule to obtain each cleaned index data;

[0081] The aggregation storage unit is used to aggregate and store the cleaned index data with the same index information type to obtain each target data set, wherein the target data set includes the index information type and at least one digital data.

[0082] Based on the above embodiment, the data sorting module 320 includes:

[0083] The log search unit is used to search for the target index log respectively matched with each digital data and obtain the log attribute of the target index log as the data attribute of the digital data matched with the target index log;

[0084] A requirement attribute acquisition unit, configured to extract requirement attribute information matching the digital data from the data attributes according to a preset data collation rule;

[0085] A data sorting unit, configured to perform a sorting operation on each digital data in the target data set according to the data collation rule and the requirement attribute information of each digital data, so as to obtain an original tensor matrix matching the target data set.

[0086] Based on the above embodiments, the convolutional neural network module 330 includes:

[0087] A pooling rule acquisition unit, configured to respectively obtain each pooling rule in the data sampling requirement as a target pooling rule;

[0088] A matrix search unit, configured to search for an original tensor matrix matching the pooling rule in each original tensor matrix as a target tensor matrix according to the index information type in the target pooling rule;

[0089] A convolution configuration unit, configured to obtain a target window size and a target window step length matching the target pooling rule in the data sampling requirement, and configure a preset convolutional neural network according to the target pooling rule, the target window size, and the target window step length;

[0090] A pooling processing unit, configured to perform pooling processing on the original tensor matrix through the configured convolutional neural network to obtain a sampled data matrix matching the original tensor matrix.

[0091] In fact, the result generation module 340 is further configured to: after aggregating each sampled data matrix to obtain a sampled data result matching the data sampling requirement, perform visual display on the sampled data result based on a visualization rule set by a user to obtain a visualization icon; wherein, the visualization rule includes: icon type, color scheme, and display mode of index information type.

[0092] A data sampling device for metric logs provided by an embodiment of the present invention can execute a data sampling method for metric logs provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0093] Embodiment 4

[0094] Figure 4FIG. 0 shows a schematic structural diagram of an electronic device 10 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0095] As Figure 4 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0096] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0097] The processor 11 can be various general-purpose and / or special-purpose processing components having processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as a data sampling method for a metric log.

[0098] Correspondingly, the method includes:

[0099] Obtain at least one target data set from the metric log library based on the data acquisition rules set by the user;

[0100] Arrange each target data set according to the preset data arrangement rules to obtain an original tensor matrix respectively matching each target data set;

[0101] Obtain the user's data sampling requirements, and use a pre-set convolutional neural network to process each original tensor matrix based on the data sampling requirements to obtain each sampling data matrix respectively matching each original tensor matrix;

[0102] Perform an aggregation operation on each sampling data matrix to obtain a sampling data result matching the data sampling requirements.

[0103] In some embodiments, a data sampling method for metric logs can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by the processor 11, one or more steps of the data sampling method for metric logs described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute a data sampling method for metric logs in any other suitable manner (e.g., by means of firmware).

[0104] The various embodiments of the systems and technologies described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0105] A computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0106] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0107] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0108] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0109] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0110] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

Claims

1. A data sampling method for index logs, characterized in that, Including: Obtain at least one target data set in the metric log library based on the data acquisition rule set by the user; Sort out each target data set according to the preset data sorting rule to obtain an original tensor matrix respectively matching each target data set; Obtain the user's data sampling requirement, and use a pre-set convolutional neural network to process each original tensor matrix based on the data sampling requirement to obtain each sampled data matrix respectively matching each original tensor matrix; Perform an aggregation operation on each sampled data matrix to obtain a sampled data result matching the data sampling requirement.

2. The method according to claim 1, wherein Obtain at least one target data set in the metric log library based on the data acquisition rule set by the user, including: Obtain the metric information type of each metric log according to the log attribute of each metric log in the metric log library; If it is determined that the metric information type of the metric log matches the data acquisition rule, set the metric log as the target metric log; Extract the numerical data in each target metric log, and obtain each target data set based on each numerical data and the target metric log respectively matching each numerical data.

3. The method according to claim 2, wherein Obtain each target data set based on each numerical data and the target metric log respectively matching each numerical data, including: Bind and store each numerical data and the metric information type of the target metric log respectively matching each numerical data to obtain each original metric data respectively matching each target metric log; Perform a data cleaning operation on each original metric data according to the preset data cleaning rule to obtain each cleaned metric data; Aggregate and store the cleaned metric data with the same metric information type to obtain each target data set, where the target data set includes the metric information type and at least one numerical data.

4. The method according to claim 1, wherein Sort out each target data set according to the preset data sorting rule to obtain an original tensor matrix respectively matching each target data set, including: Search for the target metric log respectively matching each numerical data, and obtain the log attribute of the target metric log as the data attribute of the numerical data matching the target metric log; Extract the required attribute information matching the numerical data from the data attribute according to the preset data sorting rule; Sort the numerical data in the target data set according to the data sorting rule and the required attribute information of each numerical data to obtain an original tensor matrix matching the target data set.

5. The method according to claim 1, wherein The data sampling requirement includes: at least one pooling rule, a pooling window size respectively matching each pooling rule, and a pooling window step length respectively matching each pooling rule, where the pooling rule includes a pooling type and a metric information type matching the pooling type.

6. The method according to any one of claims 1-4, characterized in that Use a pre-set convolutional neural network to process each original tensor matrix based on the data sampling requirement to obtain each sampled data matrix respectively matching each original tensor matrix, including: Obtain each pooling rule in the data sampling requirement as the target pooling rule respectively; Search for the original tensor matrix matching the pooling rule in each original tensor matrix according to the metric information type in the target pooling rule as the target tensor matrix; Obtain a target window size and a target window stride that match the target pooling rule in the data sampling requirement, and configure a pre-set convolutional neural network according to the target pooling rule, the target window size, and the target window stride; Perform pooling processing on the original tensor matrix through the configured convolutional neural network to obtain a sampled data matrix that matches the original tensor matrix.

7. The method according to claim 1, characterized in that, After aggregating each sampled data matrix to obtain a sampled data result that matches the data sampling requirement, it further includes: Visually display the sampled data result based on a visualization rule set by the user to obtain a visualization icon; Among them, the visualization rule includes: icon type, color scheme, and display method of metric information type.

8. A data sampling device for index logs, characterized in that, It includes: A dataset acquisition module, configured to acquire at least one target dataset in a metric log library based on a data acquisition rule set by the user; A data arrangement module, configured to arrange each target dataset according to a preset data arrangement rule to obtain an original tensor matrix that matches each target dataset respectively; A convolutional neural network module, configured to obtain the data sampling requirement of the user, and use a pre-set convolutional neural network to process each original tensor matrix based on the data sampling requirement to obtain each sampled data matrix that matches each original tensor matrix respectively; A result generation module, configured to aggregate each sampled data matrix to obtain a sampled data result that matches the data sampling requirement.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute a data sampling method for a metric log according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the processor to implement a data sampling method for a metric log according to any one of claims 1-7 when executed.