A method for managing resources of a commercial data management platform and a commercial data management platform
By obtaining log and indicator data in the commercial data management platform, feature extraction and fusion is performed, and using the decision tree detection model, the problem of fault detection under the microservice architecture is solved, and the operation status of the commercial data management platform is achieved is achieved, which improves the accuracy of abnormal detection and the reliability of the system.
Patent Information
- Application Number
- CN202510577825.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-05-07
AI Technical Summary
Existing commercial data management platforms face the problem of fault detection under the microservice architecture, especially in high concurrency and complex business scenarios. Traditional single data source detection methods cannot provide a comprehensive view of the system's operating status, resulting in low abnormal detection accuracy and difficulty in dealing with dynamic challenges.
By obtaining log data and key operation indicators, feature extraction and fusion are performed, and the decision tree detection model after hyperparameter optimization is used to detect the operation status of the commercial data management platform. Combining log features and indicator characteristics, a resource management method for the commercial data management platform is constructed.
It realizes rapid and comprehensive detection of the operating status of the commercial data management platform, improves the accuracy of abnormal detection and the reliability of the system, and can promptly detect and report any abnormalities, ensuring the availability and reliability of the system.
Smart Images

Figure CN120105214B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of computer management, and particularly to a method for managing resources of a commercial data management platform and a commercial data management platform. Background Art
[0002] Currently, data management platforms have been widely applied in related industries such as finance, power, and the Internet. The data management platform can separate data from business, support rapid iteration, and provide efficient services. In existing commercial data management platforms, a monolithic architecture is usually adopted, that is, the entire application is developed, deployed, and maintained as a single unit. Each functional module shares the same code library and database, and is usually developed using a unified technology stack. This architecture mode is simple and direct, and is suitable for small applications. However, with the continuous expansion of business scale and the increasing complexity, it often faces challenges such as limited scalability, cumbersome deployment processes, and high coupling degree between systems. The traditional monolithic architecture is unable to cope with distributed and high-concurrency scenarios. Therefore, a microservices architecture with high flexibility and scalability and that can better meet the requirements in large-scale, high-concurrency, and complex business scenarios has emerged.
[0003] Microservices split a traditional monolithic architecture application into multiple sub-modules by setting the logical boundaries of business functions, achieving the purpose of splitting complex monolithic architecture applications. Each sub-module can be regarded as an individual sub-service according to the corresponding business function. Different sub-services communicate through communication protocols and complete mutual calls to implement user requests. Each sub-service can be regarded as an independent monolithic application, and different technical frameworks can be used internally to implement the corresponding functions. Fundamentally speaking, the microservices architecture manages these multiple sub-services and uses a unified interface standard to achieve interaction between them to meet more complex user requirements.
[0004] The microservices architecture, through its highly decoupled and modular characteristics, while reducing the workload of software developers, brings unprecedented challenges to software operation and maintenance. Facing the operation and maintenance challenges brought by the microservices architecture, how to achieve timely detection of faults has become an urgent problem to be solved. Summary of the Invention
[0005] In order to solve the technical problems existing in the prior art, embodiments of this application provide a method for managing resources of a commercial data management platform and a commercial data management platform, which determine the overall running state of the platform by obtaining multiple types of running data of the current system and based on multiple features corresponding to the data.
[0006] To achieve the above object, the technical solutions adopted in the embodiments of this application are as follows:
[0007] In a first aspect, a method for managing resources of a business data management platform is provided. The method includes: extracting features from the obtained log sequence and key operation metrics that match the container, to obtain log features for each log and metric features for each key operation metric; fusing the log features and the metric features to obtain a feature vector associated with the operation of the container; and performing state detection on the feature vector through a detection model in a converged state after hyperparameter optimization, to determine the current operation state of the container.
[0008] In some specific implementation manners, the log features include semantic features, sequence features, and weight features; obtaining the log features for each log includes: respectively extracting the semantic feature, sequence feature, and weight feature of the log, and merging the semantic feature and the sequence feature to obtain the log feature, where the weight feature is used to indicate the weight corresponding to the log feature.
[0009] In some specific implementation manners, fusing the log features and the metric features includes: screening the log features and the metric features based on the sorting of the weight features, and performing fusion based on the weight features through a fully connected layer.
[0010] In some specific implementation manners, screening the log features and the metric features based on the sorting of the weight features includes: removing the log features that do not meet the sorting requirements and the metric features at the corresponding time nodes; and performing fusion based on the weight features through a fully connected layer includes: updating the weights of the log features and the corresponding metric features based on the weight features, and performing fusion on the updated log features and metric features through the fully connected layer.
[0011] In some specific implementation manners, the method further includes: parsing the obtained log that matches the container, including: removing the content in the log that matches a preset regular expression, and grouping the logs by the log message length to form multiple log groups, where each log group corresponds to a different log message length; performing secondary grouping on each log group based on the log prefix to form multiple different log subgroups, and calculating the similarity between the message logs in the current log subgroup and the log template in the current group; when the calculated similarity is higher than a predetermined threshold, determining that the log message is classified into the current log group, otherwise creating a new log group for this log message; and adding the ID of the corresponding log message to the ID list corresponding to the log group and updating the corresponding log template to form a final log message sequence; the log message sequence is represented as a word list containing multiple log events.
[0012] In some specific implementation manners, extracting the semantic features of the log includes: performing semantic embedding representation on a plurality of the word lists, converting each word in each word list sequence into a fixed-length semantic vector, integrating the corresponding semantic vectors, and converting the word list sequence into a log vector sequence.
[0013] In some specific implementation manners, extracting the sequence features includes: performing sliding window segmentation on the log sequence to obtain a plurality of subsequences, obtaining a plurality of log template subsequences corresponding to the subsequences through the mapping relationship between the log template and the log sequence, and performing soft one-hot encoding on the log template subsequences to obtain corresponding vectorized subsequences and sequence vectors.
[0014] In some specific implementation manners, the detection model is constructed based on a decision tree. The hyperparameter optimization of the detection model is performed by calculating the fitness of a plurality of initial parameter groups, and screening and iterating through the fitness of the plurality of initial parameter groups until the termination condition is met, and screening the parameter with the highest fitness among the plurality of initial parameter groups as the optimal hyperparameter corresponding to the optimization result.
[0015] In a second aspect, a commercial data management platform is provided, including: a data acquisition unit, a data storage unit, and a data processing unit. The data acquisition unit includes: a running data retrieval module for obtaining log data; an information providing module for collecting process identifiers and performing container positioning; an information collection module for obtaining key running metrics and sending the log data and the key running metrics to the data storage unit; the data storage unit transmits the collected data to the data processing unit, and the data processing unit is used for the resource management method of the commercial data management platform described in any one of the above.
[0016] In some specific implementation manners, the data processing unit includes: a feature processing module for respectively extracting features of the log sequence and the key running metrics obtained and matched with the container to obtain log features for each log and metric features for each key running metric; a feature fusion module for fusing the log features and the metric features to obtain a feature vector associated with the container operation; a detection module for performing state detection on the feature vector through a detection model in a converged state after hyperparameter optimization to determine the current running state of the container.
[0017] In the technical solution provided by the embodiments of the present application, by obtaining the log data and key operation index data in the current management platform system, and performing feature processing and screening and fusion on the data to obtain the fusion features reflecting the operation status of the current commercial data management platform, and processing and detecting the fusion features through a detection model after hyperparameter optimization to determine the operation status of the current system, so as to determine whether the system has an abnormality. The method and platform provided by the embodiments of the present application can collect various data reflecting the operation status, and process the data through a detection model to obtain the status of the platform system operation. Compared with the prior art, it has the rapidity and comprehensiveness of processing, and can capture the complete operation status to make the detection more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0019] The methods, systems, and / or programs in the drawings will be further described according to exemplary embodiments. These exemplary embodiments will be described in detail with reference to the drawings. These exemplary embodiments are non-limiting exemplary embodiments, where the example numbers represent similar mechanisms in the various views of the drawings.
[0020] Figure 1 It is a schematic structural diagram of a commercial data management platform provided by the embodiments of the present application.
[0021] Figure 2 It is a schematic flowchart of the resource management method of the commercial data management platform provided by the embodiments of the present application.
[0022] Figure 3 It is a schematic structural diagram of the data processing unit provided by the embodiments of the present application.
[0023] Figure 4 It is a schematic structural diagram of the terminal device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] In order to better understand the above technical solutions, the following will make a detailed description of the technical solutions of the present application through the drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solutions of the present application, rather than limitations on the technical solutions of the present application. Without conflict, the technical features in the embodiments of the present application and the embodiments can be combined with each other.
[0025] In the following detailed description, numerous specific details are set forth by way of examples in order to provide a thorough understanding of the relevant teachings. However, it will be apparent to those skilled in the art that the present application may be practiced without these details. In other instances, well-known methods, procedures, systems, components, and / or circuits have been described at a relatively high level without detail in order to avoid unnecessarily obscuring aspects of the present application.
[0026] Flowcharts are used in the present application to illustrate the execution processes performed by the systems according to embodiments of the present application. It should be clearly understood that the execution processes of the flowcharts may not be executed in sequence. On the contrary, these execution processes may be executed in reverse order or simultaneously. Additionally, at least one other execution process may be added to the flowchart. One or more execution processes may be deleted from the flowchart.
[0027] Before further elaborating on the embodiments of the present invention, the nouns and terms involved in the embodiments of the present invention are described. The nouns and terms involved in the embodiments of the present invention are applicable to the following explanations.
[0028] (1) Responsive to, which is used to represent the conditions or states upon which the performed operations depend. When the dependent conditions or states are met, one or more operations to be performed may be real-time or may have a set delay; without special instructions, there is no limitation on the execution order of the multiple operations to be performed.
[0029] (2) Based on, which is used to represent the conditions or states upon which the performed operations depend. When the dependent conditions or states are met, one or more operations to be performed may be real-time or may have a set delay; without special instructions, there is no limitation on the execution order of the multiple operations to be performed.
[0030] The present application provides a commercial data management platform, which is established based on a microservices architecture. Specifically, it can be a commercial platform, on which multiple microservice terminals are configured. It can be an e-commerce platform or a SaaS commercial software. In this embodiment, it is mainly a commercial platform configured in the payment scenario, used to manage the data generated during the payment service process. For the payment scenario, especially the users and payment services involved in the C-end payment scenario are relatively complex, and a large amount of payment data can be generated in a short time, with high concurrency and data complexity. If the existing system architecture is adopted, challenges such as limited scalability, cumbersome deployment processes, and high coupling between systems will be faced. Therefore, in this embodiment, a microservices architecture is adopted for the platform system construction.
[0031] Among them, the microservices architecture has become one of the mainstream models in modern software development. Compared with traditional monolithic applications, the microservices architecture improves the scalability and flexibility of the system by decomposing complex applications into a set of lightweight and loosely coupled services, enabling each service to be deployed and scaled independently. Since microservices usually consist of multiple service components, these components may be deployed in different physical locations and communicate and exchange data with each other through the network. Each microservice may run independently in different operating environments and may be developed and maintained by different teams using different technology stacks. This characteristic of diversification and decentralization greatly improves the flexibility and agility of the system. However, the distributed nature and dynamics of the microservices architecture also bring many challenges, especially in the management of service stability and reliability. In the microservices architecture, the underlying call relationships between services are intricate. Once a problem occurs in a service, it may quickly spread throughout the system, triggering cascading failures and systemic problems, thus causing huge economic losses. Therefore, it is necessary to perform anomaly detection on microservices in this system.
[0032] The goal of microservice anomaly detection is to monitor the running state of microservices in order to detect and report any anomalies in a timely manner, ensuring the availability and reliability of the system. However, most existing detection methods rely only on a single data source, which cannot provide a comprehensive view of the system running state in many cases. In some cases, if the key data source is lost or incomplete, it will directly affect the accuracy of anomaly detection. In addition to the fact that a single data source cannot fully describe the system, the dynamics of the microservices system are also a key challenge for anomaly detection.
[0033] Therefore, the embodiments of this application provide a commercial data management platform that can detect the running state in the system, especially anomalies. Refer to Figure 1 , the commercial data management platform 10 includes a data collection unit 11, a data storage unit 12, and a data processing unit 13. Among them, the data collection unit 11 is used to obtain the running data in the current commercial data management platform; the running data of the commercial data management platform in this embodiment includes background data and front-end running data, where the background data is log data in this embodiment, and the front-end running data is key running metric data. And because the commercial data management platform in this embodiment has a microservices architecture, there are multiple containers in the commercial data management platform, and the sources of the above-obtained data are the background data and front-end data of multiple containers. Therefore, it is necessary to match the above data with the corresponding containers.
[0034] The data storage unit 12 is used to store the real-time running data and store the above running data at preset time intervals to form a data sequence.
[0035] Specifically, in this embodiment, to achieve the above technical objectives, the data acquisition unit in this embodiment includes: an operation data retrieval module 111 for obtaining log data; an information providing module 112 for collecting process identifiers and performing container positioning; and an information collection module 113 for obtaining key operation metrics and sending the log data and the key operation metrics to the data processing unit.
[0036] Specifically, the operation data retrieval module 111 uses eBPF technology to achieve more efficient and accurate log data acquisition. Among them, the log data is underlying data, while the commercial data management platform pays more attention to higher-level information such as microservice information and user-defined configurations. And in this embodiment, to better obtain the running status of the commercial data management platform, the above two types of information need to be effectively combined by setting an information providing module. The information providing module is used to collect process identifiers to accurately locate specific containers, and then obtain the running data of the containers. It should be noted that the operation data retrieval module and the information providing module are independent of each other, and data acquisition is directly performed when the processes in the container change.
[0037] Regarding the above operation data stored in the data storage unit 12, in this embodiment, a data processing unit is set for processing. The data processing unit retrieves the above operation data, and a commercial data management platform resource management method is configured in the data processing unit to implement the processing of the collected data to determine the running status of the current commercial data management platform. Specifically, for this method, reference can be made to Figure 2 , including the following steps:
[0038] Step S21. Feature extraction is respectively performed on the obtained log sequence and key operation metrics matching the container to obtain log features for each log and metric features for each key operation metric.
[0039] In this embodiment, the operation data retrieval module can obtain log data. The log data is arranged based on a time series sequence to form a log sequence, and each sequence stores the log data corresponding to the current time point based on the time point. The key operation metrics are obtained through the information collection module. The key operation metrics are also sequence data arranged based on a time series sequence, and the log data and the key operation metrics are set corresponding to each other through the running time point.
[0040] Among them, the original log data usually contains various mixed information, such as timestamps, log levels, log IDs, and specific log messages, etc. These information each have different formats. In order to effectively extract valid information from the complex original log data, it is necessary to first convert the unstructured log into a more easily processed structured form, and then perform feature extraction on the converted log data.
[0041] Specifically, in this embodiment, the format conversion of the log data is implemented by means of log parsing. Among them, for log parsing, first remove the content in the log that matches the preset regular expression. For example, remove the IP address to simplify the log message. Then, construct a log group for the above log message. In this embodiment, a log group is to divide log messages with the same characteristics into the same group to simplify the parsing degree. Among them, the same points of the characteristics include length and log prefix. For the length, it means that assuming that log messages with the same log template have the same message length, divide the above log messages into different log groups according to the log message length. For the log prefix, it means that assuming that the mark at the start position of the log message is a constant, perform secondary grouping on the log messages in the above divided log groups to form multiple different log subgroups, further refining the search path. According to the above grouping results, calculate the similarity between the message log in the current group and the log template in the current group. If the calculated similarity is higher than the predetermined threshold, then determine that this log message is divided into the current log group; otherwise, create a new log group for this log message. When determining the log group through the above steps, the ID corresponding to the log message will be added to the ID list corresponding to the log group and the corresponding log template will be updated to form the final log message sequence. Among them, the performance of the log message sequence is a word list containing multiple log events, and such a sequence can be represented as a word list sequence 。
[0042] Through the above parsing process, the updated log message is obtained, and then feature extraction is performed on the updated log message. Among them, the log message in this embodiment has multiple features, at least including semantic features and sequence features, and in order to pay more attention to important information during subsequent anomaly detection, the features extracted in this embodiment also include the weight value corresponding to each feature, which is the weight feature in this embodiment.
[0043] Among them, semantic features refer to information such as event content and event type reflected in log data, which is particularly important during detection and can directly affect the final detection result. In this embodiment, the extraction of semantic features is implemented using the SBERT model. The SBERT model converts natural language into a fixed-length vector form by combining the siamese network architecture and the semantic processing ability of the BERT model. Specifically, through semantic embedding representation of the word list, each word in each word list is converted into a fixed-length semantic vector, and the corresponding semantic vectors are integrated to convert the word list sequence into a log vector list.
[0044] Sequence features are used to reflect the logical order of event occurrence. After each original log message is parsed in the above parsing process, a corresponding specific log template and parameter list will be obtained. Each log template has a unique identifier, which is called a template ID. In this embodiment, the extraction of sequence features is achieved by performing sliding window segmentation on the log sequence to obtain multiple subsequences, and obtaining the corresponding log template subsequences of the multiple subsequences through the mapping relationship between the log template and the log sequence. Then, by performing soft one-hot encoding on the log template subsequences for transformation, each template subsequence is converted into a corresponding vectorized representation, and the transformed vectorized representations are combined based on the sequence to obtain the final sequence vector.
[0045] Through the above method, the semantic features and sequence features corresponding to the log messages can be obtained. Among them, the sequence features are used to indicate the logical order of the logs, and the semantic features are used to indicate the entity content of the logs. Therefore, in this embodiment, the main task is to combine the above two features to obtain the final log features.
[0046] Specifically, to combine the above two features, the torch.cat function can be used to concatenate the above two different vectors along a new dimension to generate a log feature vector.
[0047] In this embodiment, log data is used to reflect the running state of the system at the backend. For the running state of the system at the frontend, i.e., the task end, feedback on the key running metrics when the system performs tasks also needs to be obtained. Therefore, in this embodiment, in addition to obtaining log features, metric features also need to be obtained to improve the integrity and accuracy during detection.
[0048] Specifically, there are many types of abnormal patterns for key running metrics. Metric features can be divided into local features and global features. Among them, local features can capture local anomalies or mutations, while global features can provide the overall system behavior and trends. Therefore, for the comprehensiveness of the features, in this embodiment, the metric features should be the fused features after fusing global features and local features, so as to achieve a complete capture of the metric features.
[0049] Among them, in order to achieve the above-mentioned effects, the extraction of indicator features is implemented through a multi-scale time series feature extraction module. Among them, the multi-scale time series feature extraction module first performs preliminary feature extraction on the indicator data sequence by setting up an LSTM model, generates a hidden state vector containing multiple time points, and obtains a hidden state matrix by splicing the hidden state vectors. For the obtained hidden state matrix, in order to realize the extraction of global features and local features, the matrix is subjected to global convolution calculation and local convolution calculation respectively. When performing global convolution calculation, the convolution kernel is used to perform convolution calculation on the matrix row by row, so that each convolution kernel fully covers the data in the hidden state matrix to obtain a global hidden state feature matrix. When performing local convolution calculation, the hidden state matrix is convolved row by row through a one-dimensional convolution kernel, and the local features in the convolution result are extracted through the maximum pooling layer. Finally, the local hidden state feature matrix is obtained through maximum pooling processing. Then, the hidden state vector of LSTM at the final moment is used as the query vector, and the row vectors of the global hidden state feature matrix and the local hidden state feature matrix are used as the key vector and value vector to calculate the attention score and corresponding attention weight of each row. Based on the attention weight, the rows of the global hidden state feature matrix and the local hidden state feature matrix are weighted and merged to obtain the global encoding vector and the local encoding vector. Finally, the global encoding vector and the local encoding vector are fused to generate the final encoding vector, which is the feature vector.
[0050] Step S22: Fusing the log feature with the indicator feature to obtain a feature vector associated with the container operation.
[0051] In this embodiment, log features and indicator features are features on a long time series, and the collected data and corresponding features are relatively large. If all features are fused during feature fusion, the processing cost will be too high. Therefore, in this embodiment, before fusion of features and processing based on the fused features, it is necessary to eliminate the features and retain only the important features, so as to reduce the high computational cost caused by the large number of features in subsequent processing. For this reason, the weight feature is introduced in this embodiment to evaluate the weights of the log features and indicator features corresponding to the same time, and the features that do not meet the requirements are eliminated based on the corresponding weight feature.
[0052] Among them, the weight feature is obtained by calculating the similarity score between the log feature and the indicator feature corresponding to each time point. The calculation is performed through dot product attention, and the weight of the log feature and the indicator feature corresponding to each time point is obtained by normalization through the loss function, that is, the weight feature corresponding to this embodiment.
[0053] When performing feature fusion based on the obtained weight features above, first, log features and metric features that do not meet the weight feature sorting requirements are eliminated, and only log features and metric features that meet the weight value requirements are retained. Then, the filtered log features and metric features are fused based on a fully connected layer. Specifically, when performing the fusion, the above log features and metric features are first updated through the weight features, and the updated log features and metric features are obtained by multiplying the weight values with the log features and metric features in a weighted manner. Among them, for feature fusion of the fully connected layer, the method in the existing technology can be directly used for implementation, and it will not be elaborated in this embodiment.
[0054] Step S23. Perform status detection on the feature vector through the detection model in a converged state after hyperparameter optimization, and determine the current running state of the container.
[0055] Regarding steps S21 - S22, data reflecting the running state of the system can be obtained, and the features corresponding to the data are obtained through feature extraction, and multiple features are screened to serve as the input for subsequent status detection.
[0056] Among them, the detection model in this embodiment is constructed based on a decision tree. Specifically, the decision tree in this embodiment is a decision tree based on gradient boosting. And, in order to reduce the training cost of the detection model and improve the detection accuracy of the detection model, the selection of hyperparameters in the decision tree in this embodiment is optimized through a genetic algorithm.
[0057] Specifically, first, multiple initial populations are generated. Each individual represents a set of hyperparameter configurations of the detection model and is represented by a randomly generated binary string. Calculate the fitness corresponding to each set of hyperparameter configurations. Specifically, RMSE is selected as the fitness function for fitness calculation. According to the fitness of each individual, iterative fitness calculation is performed through crossover, mutation, and generation of a new population in the genetic algorithm until the termination condition is met, and then the parameter with the highest fitness among multiple parameter groups is selected as the optimization result, that is, the optimal hyperparameters in this embodiment.
[0058] The detection model configured with the optimal hyperparameters is the final model in this embodiment. Then, the fused features obtained in step S22 are input into the detection model. In the detection model, first, the input feature values are discretized and mapped to a finite number of integer values, and a histogram with a fixed width is constructed based on the finite number of integer values. Using the discretized feature values as indexes, iteration is performed on the histogram to obtain statistics. Among them, for iteration in this embodiment, the leaf node with the largest classification gain is preferentially selected for splitting, and then this process is looped. Among them, for the detection task in this embodiment, its corresponding mapping relationship is: , where X is the input feature and Y is the set of fault categories. By determining the final statistic, the mapping result with the widest distribution corresponding to the statistic, i.e., the fault category, is taken as the detection result in this embodiment.
[0059] It should be noted that for the detection result in this embodiment, it is the proportional relationship of the fault type in the statistical data. And for this embodiment used to determine the system operation state, the fault type is used to indicate a certain fault existing in the system. Therefore, when making the final state judgment, not only the fault type with the largest distribution needs to be determined, but also it is necessary to determine whether such a fault occurs through its distribution proportion relationship.
[0060] Specifically, obtain the distribution proportion of the corresponding fault and determine its relationship with the preset threshold. When the distribution proportion of the fault exceeds the preset threshold, it indicates that such a fault exists in the current system. When the distribution proportion of the fault does not exceed the preset threshold, it is necessary to retrieve the sequence of key operation indicators corresponding to such a fault. This sequence is a continuous value. By statistically calculating the change values corresponding to the maximum and minimum values of the key operation indicator data in this sequence, determine whether the change degree of the change value exceeds the threshold corresponding to this type of data. When this change value is greater than the threshold, it indicates that such a fault exists in the current system; when the change value is less than the threshold, it indicates that such a fault does not exist in the current system.
[0061] In this embodiment, by obtaining the log data and key operation indicator data in the current commercial data management platform, and performing feature processing and screening and fusion on the data to obtain the fusion features reflecting the operation state of the current commercial data management platform, and processing and detecting the fusion features through a detection model after hyperparameter optimization to determine the operation state of the current commercial data management platform, thereby determining whether the commercial data management platform has an anomaly.
[0062] In this embodiment, refer to Figure 3 , for the data processing unit 13 used to execute the processing process of steps S21 - S23, this unit includes the following modules:
[0063] The feature processing module 131 is used to respectively extract features from the obtained log sequence and key operation indicators matching the container to obtain the log features of each log and the index features of each key operation indicator;
[0064] The feature fusion module 132 is used to fuse the log features and the index features to obtain a feature vector associated with the container operation;
[0065] The detection module 133 is used to perform state detection on the feature vector through a detection model in a converged state after hyperparameter optimization to determine the current operation state of the container.
[0066] Refer to Figure 4 In other embodiments, the above method may also be integrated into the provided terminal device 40. Due to relatively large differences that may occur in the device due to configuration or performance, it may include one or more processors 401 and a memory 402. One or more application programs or data may be stored in the memory 402. Among them, the memory 402 may be short-term storage or persistent storage. The application programs stored in the memory 402 may include one or more modules (not shown in the figure), and each module may include a series of computer-executable instructions in the terminal device. Further, the processor 401 may be set to communicate with the memory 402 to execute a series of computer-executable instructions in the memory 402 on the terminal device. The terminal device may also include one or more power supplies 403, one or more wired / wireless network interfaces 404, one or more input / output interfaces 405, one or more keyboards 406, etc.
[0067] In a specific embodiment, the terminal device includes a memory and one or more programs. One or more of the programs are stored in the memory, and one or more of the programs may include one or more modules, and each module may include a series of computer-executable instructions in the terminal device, and is configured to be executed by one or more processors. The one or more programs include the following computer-executable instructions:
[0068] Extract features from the obtained log sequence and key operation metrics that match the container respectively to obtain log features for each log and metric features for each key operation metric;
[0069] Fuse the log features and the metric features to obtain a feature vector associated with the container operation;
[0070] Perform state detection on the feature vector through a detection model in a converged state after hyperparameter optimization to determine the current operation state of the container.
[0071] The following is a specific introduction to the components of the processor:
[0072] Among them, in this embodiment, the processor is an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0073] Optionally, the processor can execute various functions by running or executing software programs stored in the memory and calling data stored in the memory. For example, execute the Figure 2 method shown above.
[0074] In a specific implementation, as an embodiment, the processor may include one or more microprocessors.
[0075] The memory is used to store the software program for executing the solution of the present application and is controlled by the processor for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated here.
[0076] Optionally, the memory may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but not limited thereto. The memory may be integrated with the processor or exist independently and be coupled to the processing unit through the interface circuit of the processor. The embodiments of the present application do not make specific limitations on this.
[0077] It should be noted that the structure of the processor shown in this embodiment does not constitute a limitation on the device. The actual device may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0078] In addition, the technical effects of the processor can refer to the technical effects of the method described in the above method embodiments, which will not be elaborated here.
[0079] It should be understood that the processor in the embodiments of the present application may be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0080] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0081] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0082] In the present application, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or a similar expression means any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0083] It should be understood that in various embodiments of the present application, the order numbers of the above processes do not indicate the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0084] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0085] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0086] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0087] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0088] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0089] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0090] As described above, it is only the specific implementation manner of the present application. However, the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims described above.
Claims
1. A method for managing resources of a commercial data management platform, characterized in that, The method includes: Performing feature extraction on the obtained log sequence and key running metrics that match the container respectively to obtain log features for each log and metric features for each key running metric; the metric features are fused features after fusing global features and local features; Fusing the log features and the metric features to obtain a feature vector associated with the container operation; Performing state detection on the feature vector through a detection model in a converged state after hyperparameter optimization to determine the current running state of the container; the detection model is constructed based on a decision tree. In the detection model, the input feature values are first discretized and mapped to a finite number of integer values. A histogram with a fixed width is constructed based on the finite number of integer values. Iteration is performed on the histogram using the discretized feature values as indices, and statistics are obtained. The mapping result with the widest distribution of the statistics, that is, the fault category, is used as the detection result. Specifically, it includes: obtaining the distribution proportion corresponding to the fault and determining the relationship with a preset threshold. When the distribution proportion of the fault does not exceed the preset threshold, the key running metric sequence corresponding to this type of fault is retrieved. By statistically calculating the change values corresponding to the maximum and minimum values of the key running metric data in the sequence, it is determined whether the degree of change of the change value exceeds the threshold corresponding to this type of data. When the change value is greater than the threshold, it indicates that there is such a fault in the current system.
2. The resource management method of the business data management platform according to claim 1, wherein The log features include semantic features, sequence features, and weight features; Obtaining log features for each log includes: respectively extracting the semantic features, sequence features, and weight features of the log, and merging the semantic features and the sequence features to obtain the log features. The weight features are used to indicate the weights corresponding to the log features.
3. The resource management method of the business data management platform according to claim 2, wherein, Fusing the log features and the metric features includes: screening the log features and the metric features based on the sorting of the weight features, and performing fusion based on the weight features through a fully connected layer.
4. The method for managing resources of a commercial data management platform according to claim 3, wherein Screening the log features and the metric features based on the sorting of the weight features includes: removing the log features that do not meet the sorting requirements and the metric features at the corresponding time nodes; and performing fusion based on the weight features through a fully connected layer includes: updating the weights of the log features and the corresponding metric features based on the weight features, and fusing the updated log features and metric features through the fully connected layer.
5. The resource management method of the business data management platform according to claim 2, characterized in that, The method further includes: parsing the obtained logs matching the container, including: removing the content in the logs matching a preset regular expression, and grouping the logs by the log message length to form multiple log groups, each of the log groups corresponding to a different log message length; secondarily grouping each log group based on the log prefix to form multiple different log subgroups, and calculating the similarity between the message logs in the current log subgroup and the log template in the current group; when the calculated similarity is higher than a predetermined threshold, determining that the log message is classified into the current log group, otherwise creating a new log group for this log message; and adding the ID of the corresponding log message to the ID list corresponding to the log group and updating the corresponding log template to form a final log message sequence; the log message sequence is represented as a word list containing multiple log events.
6. The method for managing resources of a commercial data management platform according to claim 5, characterized in that Extracting the semantic features of the logs, including: performing semantic embedding representation on multiple of the word lists, converting each word in each word list sequence into a fixed-length semantic vector, and integrating the corresponding semantic vectors to convert the word list sequence into a log vector sequence.
7. The resource management method of the business data management platform according to claim 5, characterized in that Extracting the sequence features, including: performing sliding window segmentation on the log sequence to obtain multiple subsequences, obtaining multiple log template subsequences corresponding to the multiple subsequences through the mapping relationship between the log template and the log sequence, and performing soft one-hot encoding on the log template subsequences to obtain corresponding vectorized subsequences and sequence vectors.
8. The method for managing resources of a commercial data management platform according to claim 1, wherein, The hyperparameter optimization of the detection model is performed by calculating the fitness of multiple initial parameter groups, and screening and iterating through the fitness of the multiple initial parameter groups until the termination condition is met, and screening the parameter with the highest fitness among the multiple initial parameter groups as the optimal hyperparameter corresponding to the optimization result.
9. A commercial data management platform, characterized in that, Including: A data acquisition unit, a data storage unit, and a data processing unit, the data acquisition unit includes: A running data retrieval module for obtaining log data; an information providing module for collecting process identifiers and performing container positioning; an information collection module for obtaining key running metrics and sending the log data and the key running metrics to the data storage unit; The data storage unit transmits the collected data to the data processing unit, and the data processing unit is used to execute the commercial data management platform resource management method according to any one of claims 1-8.
10. The commercial data management platform according to claim 9, characterized in that, The data processing unit includes: A feature processing module for respectively extracting features of the obtained log sequence matching the container and the key running metrics to obtain log features for each log and metric features for each key running metric; A feature fusion module for fusing the log features and the metric features to obtain a feature vector associated with the container operation; A detection module for performing state detection on the feature vector through a detection model in a convergent state after hyperparameter optimization to determine the current running state of the container.
Citation Information
Patent Citations
Software detection method and device, chip and computer readable storage medium
CN117992416A
Abnormity detection and processing method and system based on deep learning
CN119293490A