Distributed operation and maintenance data management method based on multi-modal log aggregation framework
Through the distributed operation and maintenance data governance method of the multimodal log aggregation framework, dynamic adjustment of collection strategies and real-time monitoring of log file changes solve the problems of log data loss and load imbalance under the centralized architecture, and realize efficient and reliable log data collection and storage.
Patent Information
- Application Number
- CN202510916165.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-10-24
AI Technical Summary
In large-scale distributed systems, the centralized log processing architecture has difficulty balancing real-time performance and system load due to its fixed-period single-point collection mechanism, leading to log data loss and system load imbalance.
A distributed operation and maintenance data governance method based on a multimodal log aggregation framework is adopted. By dynamically adjusting the collection strategy, combining event collection, periodic collection and compensation collection, using multi-dimensional indicators to calculate the node load status, and monitoring log file changes in real time, data integrity and efficient storage are achieved.
It improves the integrity and reliability of log data collection, reduces system overhead, ensures data accuracy and efficient storage, and optimizes resource utilization.
Smart Images

Figure CN120832290A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of electronic digital data processing, and in particular to a distributed operation and maintenance data governance method based on a multimodal log aggregation framework. Background Art
[0002] As cloud computing and distributed systems continue to expand, the amount of log data required for system operations and maintenance is growing exponentially. This log data includes not only traditional text logs but also multimodal data such as binary files, images, and videos, resulting in enormous volumes. Efficient log collection, transmission, and storage are crucial for ensuring stable system operation, rapid fault diagnosis, and intelligent operations and maintenance decision-making.
[0003] In related technologies, log collection systems employ a centralized log processing architecture, where log files generated by each node are collected and processed uniformly by a processing system. The system periodically polls log files within a fixed time window (typically 30 seconds). The collected data is parsed using a pre-set regular expression template and then transmitted to a central server via a persistent TCP connection. This data is ultimately stored in a traditional relational database and subjected to anomaly detection using threshold rules.
[0004] However, in a large-scale distributed environment, the centralized log processing system has difficulty balancing real-time performance and system load due to its fixed-period single-point collection mechanism, resulting in serious data loss problems when the log volume increases dramatically. Summary of the Invention
[0005] This application provides a distributed operation and maintenance data governance method based on a multimodal log aggregation framework, which is used to improve the integrity of log data collection in large-scale distributed systems.
[0006] In a first aspect, the application provides a distributed operation and maintenance data management method based on a multi-modal log aggregation framework, applied to a data processing system. The method comprises: obtaining system resource occupancy and log write frequency of a target node, and calculating a node load state indicator; determining node collection parameters including a collection period and a collection trigger condition according to the node load state indicator; when the log file change amount of the target node is greater than the collection trigger condition, performing event collection to obtain an event collection data stream; when the log file change amount is less than or equal to the collection trigger condition, performing polling collection according to the collection period to obtain a periodic collection data stream; calculating a log coverage range based on the event collection data stream and the periodic collection data stream, generating a data compensation instruction and performing incremental collection to obtain a compensation data stream; merging the event collection data stream, the periodic collection data stream and the compensation data stream into complete log data, and performing type differentiation on the complete log data to generate a multi-modal data packet; and distributing the multi-modal data packet to a corresponding storage engine to obtain multi-modal log aggregation data; the storage engine includes a text index engine, a time series data engine and an object storage engine.
[0007] In the above embodiment, the data processing system dynamically adjusts the collection strategy based on the system resource occupancy and log write frequency of the target node, adopts event collection when the log change amount is large, adopts periodic polling when the change amount is small, and performs data compensation by calculating the log coverage range; at the same time, the collected data is distributed to different storage engines according to the type, realizing complete collection and efficient storage of log data, and improving the completeness and reliability of log data collection in large-scale distributed systems.
[0008] In some embodiments in combination with the first aspect, the step of obtaining the system resource occupancy and the log write frequency of the target node and calculating the node load state indicator specifically comprises: collecting the processor occupancy, the memory usage and the disk read-write time of the target node, and monitoring the log file size change amount and the write operation count of the target node; calculating the system resource indicator according to the processor occupancy, the memory usage and the disk read-write time; calculating the log write indicator according to the log file size change amount and the write operation count; and performing normalization calculation on the system resource indicator and the log write indicator to obtain the node load state indicator.
[0009] In the above embodiment, the data processing system calculates the node load state through multi-dimensional indicators, including system resource indicators such as processor occupancy, memory usage and disk read-write time, and log write indicators such as log file change amount and write operation count, and obtains accurate node load state indicators through normalization, providing a reliable basis for subsequent formulation of collection strategies.
[0010] In some embodiments of the first aspect, in the step of determining the node collection parameters including the collection period and the collection trigger condition according to the node load state indicator, the step specifically includes: determining a load level of the target node according to the node load state indicator, and determining a period reference value according to the load level; obtaining a success ratio of a historical collection success number to a total collection number of the target node, and adjusting the period reference value according to the success ratio to obtain the collection period; and determining the collection trigger condition according to a log writing frequency of the target node.
[0011] In the above embodiments, the data processing system determines the load level and the period reference value according to the node load state indicator, dynamically adjusts the collection period in combination with the historical collection success rate, and sets the collection trigger condition based on the log writing frequency, so as to realize the adaptive adjustment of the collection parameters and improve the efficiency and reliability of log collection.
[0012] In some embodiments of the first aspect, before the step of performing event collection to obtain an event collection data stream, the method further includes: constructing a file system monitoring channel of the log file, creating a data collection buffer based on a file change notification mechanism of the corresponding file system monitoring channel to obtain a data temporary storage area; recording current reading position information of the log file in the data temporary storage area to obtain file reading position records; detecting a file descriptor state of the log file according to the file reading position records to obtain file state information; and updating the file reading position records based on the file state information to obtain the latest file reading position identifier.
[0013] In the above embodiments, the data processing system constructs the file system monitoring channel, creates the data collection buffer, and records and updates the file reading position in real time, so as to ensure that the change position of the log file can be accurately located during event collection, avoid data duplication collection or omission, and improve the accuracy of data collection.
[0014] In some embodiments of the first aspect, the steps of constructing the file system monitoring channel of the log file, creating the data collection buffer based on the file change notification mechanism of the corresponding file system monitoring channel to obtain the data temporary storage area specifically include: registering an event listener of the file system to obtain an event notification handle; determining a monitoring type corresponding to the event notification handle to obtain a file operation filter; determining an event response strategy according to a trigger threshold of the file operation filter, and starting a monitoring thread of the event listener to obtain a monitoring running state; initializing the file change notification mechanism according to the monitoring running state, and creating the data collection buffer based on the file change notification mechanism to obtain the data temporary storage area.
[0015] In the above embodiment, the data processing system establishes a complete file change monitoring mechanism by registering a file system event listener, setting a file operation filter, and formulating an event response strategy, thereby realizing real-time and accurate monitoring of changes in the log file and laying a foundation for efficient event collection.
[0016] In combination with some embodiments of the first aspect, in some embodiments, after the step of distributing the multi-modal data packet to the corresponding storage engine to obtain the multi-modal log aggregation data, the method further comprises: constructing a data index structure based on the multi-modal log aggregation data to obtain a retrieval access interface; determining a life cycle rule of the multi-modal log aggregation data according to the retrieval access interface to obtain a storage management strategy; performing hierarchical storage on the multi-modal log aggregation data based on the storage management strategy to obtain a hierarchical storage architecture; and updating the storage state of the multi-modal log aggregation data based on the hierarchical storage architecture.
[0017] In the above embodiment, the data processing system constructs an index structure based on the multi-modal log aggregation data, formulates a life cycle rule, and performs hierarchical storage, so that the storage and access of data are more efficient, and the traceability and consistency of data are ensured by updating the storage state.
[0018] In combination with some embodiments of the first aspect, in some embodiments, after the step of updating the storage state of the multi-modal log aggregation data based on the hierarchical storage architecture, the method further comprises: obtaining data access frequencies of each storage layer in the hierarchical storage architecture to obtain a data heat distribution; calculating a storage level migration threshold according to the data heat distribution to obtain a data migration strategy; identifying a data set to be migrated based on the data migration strategy to obtain a migration task list; performing a data migration operation in the migration task list to obtain a migration execution result, and updating storage location information of the multi-modal log aggregation data according to the migration execution result.
[0019] In the above embodiment, the data processing system realizes intelligent hierarchical storage of data by monitoring data access frequencies, calculating a storage level migration threshold, and performing a data migration operation, thereby optimizing storage resource utilization and improving data access efficiency.
[0020] In the second aspect, the embodiments of the present application provide a data processing system, which comprises one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code comprising computer instructions, and the one or more processors invoke the computer instructions to enable the data processing system to perform the method described in the first aspect and any possible implementation manner of the first aspect.
[0021] In a third aspect, the embodiments of the present application provide a computer program product comprising instructions which, when executed on a data processing system, cause the data processing system to carry out the method according to the first aspect and any possible implementation of the first aspect.
[0022] In a fourth aspect, the embodiments of the present application provide a computer-readable storage medium comprising instructions which, when executed on a data processing system, cause the data processing system to carry out the method according to the first aspect and any possible implementation of the first aspect.
[0023] It can be understood that the data processing system provided by the second aspect, the computer program product provided by the third aspect and the computer storage medium provided by the fourth aspect are all used to execute the method provided by the embodiments of the present application. Therefore, the beneficial effects that can be achieved thereby can refer to the beneficial effects in the corresponding method, which will not be described here again.
[0024] The one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. Since the adaptive collection strategy based on the node load state and the log writing feature is adopted, combined with the multi-modal data collection mechanism of event collection, periodic collection and compensation collection, and the distributed storage engine for different types of data, the collection behavior can be dynamically adjusted according to the actual load of the system, the system overhead is reduced while ensuring data integrity, the problems of data loss and system load imbalance caused by fixed periodic collection in the prior art are effectively solved, and efficient and reliable collection and storage of log data in a large-scale distributed system are realized.
[0025] 2. Since the event-driven mechanism based on the file system listening channel is adopted, combined with the precise positioning strategy of the data collection buffer and the file reading position record, and the real-time updating mechanism of the file descriptor state detection, the change of the log file can be timely perceived and precisely positioned, accurate collection and temporary storage of data are realized, the problems of log file change detection lag and inaccurate position in the prior art are effectively solved, and the real-time and accuracy of event collection are realized.
[0026] 3. Since the index structure and retrieval interface based on multi-modal log aggregation data are adopted, combined with the data life cycle management and hierarchical storage architecture, and the real-time updating mechanism of the storage state, efficient retrieval and intelligent storage management of data can be realized, the problems of low storage efficiency and complex management of massive heterogeneous log data in the prior art are effectively solved, and efficient storage, fast retrieval and intelligent management of log data are realized. BRIEF DESCRIPTION OF DRAWINGS
[0027] Figure 1is a schematic diagram of a multi-modal log aggregation framework evolution in the embodiments of the present application; Figure 2 is a process schematic diagram of a distributed operation and maintenance data governance method based on a multi-modal log aggregation framework in the embodiments of the present application; Figure 3 is another process schematic diagram of a distributed operation and maintenance data governance method based on a multi-modal log aggregation framework in the embodiments of the present application; Figure 4 is an entity device structure schematic diagram of a data processing system in the embodiments of the present application. DETAILED DESCRIPTION
[0028] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to be limiting to the present application. As used in the specification of the present application, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "and / or" used in the present application, means any or all possible combinations of one or more of the listed items.
[0029] Hereinafter, the terms "first" and "second" are only for the purpose of description, and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more of the features, and in the description of the embodiments of the present application, the meaning of "multiple" is two or more, unless otherwise specified.
[0030] In order to facilitate understanding, the application scenarios of the embodiments of the present application are introduced as follows.
[0031] A large e-commerce platform faces the challenge of collecting massive log data during Double Eleven. The platform contains thousands of service nodes, each of which generates hundreds of business logs, system logs and access logs per second. Due to the huge difference in load of each node and the violent fluctuation of log writing frequency, the traditional fixed period collection method is difficult to balance the real-time collection and system overhead. Especially during the order peak period, the CPU usage of some core nodes exceeds 90%, if high-frequency log collection continues, it may cause business processing delay; while reducing the collection frequency may cause the loss of important log information. The platform needs a log collection scheme that can be intelligently adjusted according to the node state.
[0032] Please refer to Figure 1 , which is a schematic diagram of a multi-modal log aggregation framework evolution in the embodiments of the present application. In Figure 1 , the evolution process of log processing architecture is shown by comparison: ( Figure 1The prior art adopts a centralized single log processing system, and all text files need to be collected and analyzed through a unified processing system. This architecture will cause processing bottlenecks when the system scale expands. The distributed multi-modal log aggregation framework proposed in the present application Figure 1 The lower half) adopts a decentralized design, and multiple server nodes work cooperatively. Each node can independently process multiple types of data including text, binary files, pictures and videos. The system dynamically adjusts the collection strategy according to the load state of each node, triggers the event collection mechanism when the data change volume is large, and executes periodic polling when the change volume is small. Data compensation is performed by calculating the log coverage range. Finally, the collected data is distributed to different storage engines according to the type. This distributed architecture not only improves the processing capacity and scalability of the system, but also enhances the integrity and reliability of log data collection through the cooperative processing of multi-modal data.
[0033] This architecture change essentially optimizes the mathematical model of the log processing system: from a single linear processing model to a parallel distributed processing model. For example, when processing data volume N, the processing time complexity of the traditional architecture is O(N), while the distributed architecture can reduce the time complexity to O(N / M) through parallel processing of M nodes, while the dynamic load balancing algorithm ensures that the resource utilization of each node remains at the optimal level.
[0034] In related technologies, a fixed-period batch collection method can be used to collect log data using uniform collection parameters for all nodes. At the same time, a global collection trigger threshold is set to trigger additional collection operations when the log file size exceeds the threshold. This method achieves basic log collection functions, but cannot dynamically adjust the collection strategy according to the actual load of the nodes. The following describes the scenario of using the distributed operation and maintenance data governance method based on the multi-modal log aggregation framework in related technologies.
[0035] A financial institution uses a traditional centralized log collection system, which deploys collection agents on each business node and periodically transfers log files to a central server. The system uses a uniform collection period (30 seconds) and trigger threshold (10 MB) for all nodes. During peak trading periods, the core transaction nodes frequently trigger collection operations, causing the node CPU usage to soar to 95%, affecting transaction processing speed. Some low-load monitoring nodes also perform collection operations at fixed intervals even if the log volume is small, resulting in resource waste. At the same time, due to the lack of data integrity verification mechanism, log loss often occurs, affecting problem positioning and auditing work.
[0036] The distributed operation and maintenance data management method based on the multi-modal log aggregation framework in the embodiment of the present application realizes intelligent adjustment of the collection strategy by monitoring the system resource occupancy rate and the log write frequency of the node in real time, dynamically calculating the node load state index, and automatically adjusting the collection parameters based on the index. Not only the performance impact of the collection operation on the high-load node is avoided, but also the collection real-time performance of the low-load node is improved. The following introduces the scene using the distributed operation and maintenance data management method based on the multi-modal log aggregation framework in the present application.
[0037] A cloud service provider implemented a load-aware adaptive log collection scheme. The system dynamically calculates the load state of the node by monitoring the CPU, memory and disk I / O indicators of the node in real time, combined with the log write frequency. For the database node with high load, the system automatically adjusts the collection period to 60 seconds and increases the trigger threshold to 20 MB; for the log node with low load, the collection period of 5 seconds and the trigger threshold of 1 MB are used. At the same time, by establishing an event listening mechanism, the collection is triggered immediately when a large amount of log write is detected. The system also implements a data compensation mechanism to ensure that the lost log data can be automatically supplemented after the node failure is recovered.
[0038] As can be seen, the distributed operation and maintenance data management method based on the multi-modal log aggregation framework in the embodiment of the present application realizes intelligent adjustment of the collection strategy by monitoring the system resource occupancy rate and the log write frequency of the node in real time, dynamically calculating the node load state index, and automatically adjusting the collection parameters based on the index. Not only the performance impact of the collection operation on the high-load node is avoided, but also the collection real-time performance of the low-load node is improved.
[0039] For ease of understanding, the method provided by the present embodiment is described in the following flow. Please refer to Figure 2 , a flowchart of the distributed operation and maintenance data management method based on the multi-modal log aggregation framework in the embodiment of the present application.
[0040] S201, the system resource occupancy rate and the log write frequency of the target node are obtained, and the node load state index is calculated.
[0041] The target node represents a server node in the distributed system that needs to perform log collection; the system resource occupancy rate is a comprehensive index of the system resource usage of the node, such as CPU usage, memory usage, and disk I / O load; the log write frequency represents the amount of log data generated by the target node per unit time, which is usually measured by the number of log lines or bytes written per second; the node load state index is a numerical index reflecting the current workload level of the node, with a value range of 0 to 1.
[0042] When starting a log collection task, the data processing system needs to first evaluate the load situation of the target node to ensure that the collection process does not affect the normal business operation of the node. Specifically, the data processing system obtains the CPU usage, memory usage, disk I / O and other indicators of the target node in real time through a system call interface, and calculates the log write frequency by monitoring the growth rate of the log file. Then, the data processing system standardizes these raw indicators and calculates the weighted sum according to the preset weights, and finally obtains a node load state indicator between 0 and 1. The closer the indicator is to 1, the higher the node load.
[0043] In some embodiments, the calculation of the node load state indicator can be implemented in various ways: optionally, the data processing system can use a sliding time window-based average value calculation method to collect multiple sets of system resource usage data within a fixed size time window (such as 5 minutes), calculate the average value of each indicator, and then perform weighted summation; optionally, the data processing system can also use an exponential moving average-based calculation method to calculate the weighted average value of each indicator by giving higher weights to recent data, and then perform normalization to obtain the final load state indicator. It can be understood that other mathematical models can also be used to implement quantitative calculation of the load state, which is not limited here. In order to make the calculation result more accurate, the historical load trend and business peak period of the node also need to be considered.
[0044] In actual application, due to the differences in hardware configuration and business characteristics of each node in the distributed system, directly using a unified threshold to judge the node load may produce deviation. For this purpose, the data processing system can use an adaptive threshold adjustment mechanism: first, establish a baseline load model of the node to record the normal load range of the node at different time periods; then calculate the mean and standard deviation of the load indicator based on historical data to dynamically adjust the judgment threshold; finally, set differentiated load evaluation standards according to business rules, so as to more accurately reflect the actual load situation of different nodes.
[0045] S202, according to the node load state indicator, determine the node collection parameters including the collection period and the collection trigger condition.
[0046] Among them, the collection period represents the time interval at which the system performs regular log polling; the collection trigger condition refers to the log change threshold that triggers event collection; the node collection parameters include a set of configuration items such as collection period, collection trigger condition, and collection data upper limit. The node load state indicator is used to quantitatively represent the current workload level of the node, and the value range is 0 to 1.
[0047] The data processing system needs to dynamically adjust the log collection strategy based on the load status of the node to balance the real-time performance and system overhead. Specifically, the data processing system first compares the node load status indicators with the preset load level threshold to determine the current load level of the node (such as low load, medium load, and high load). Then, according to different load levels, the system sets the corresponding collection period reference value and dynamically adjusts the reference value in combination with the historical collection success rate. At the same time, the system also needs to set reasonable collection trigger conditions according to the log write frequency characteristics of the node to ensure that the event collection can be triggered in time in the case of a sudden increase in logs.
[0048] In some embodiments, the dynamic adjustment of the collection parameters can be achieved in various ways: optionally, the data processing system can use a feedback control-based parameter adjustment method to dynamically adjust the collection period and trigger threshold by continuously monitoring the execution effect of the collection task (such as collection success rate, data loss rate, etc.); optionally, the system can also use a machine learning-based method to automatically optimize the collection parameter configuration by analyzing historical collection data to build a prediction model. It can be understood that other algorithms can also be used to realize the adaptive adjustment of the collection parameters, which are not limited here. In order to improve the accuracy of parameter adjustment, the periodic variation characteristics of business load also need to be considered.
[0049] In practice, due to the possible sharp fluctuations in node load, the adjustment of the collection parameters may be delayed or oscillated. For this purpose, the data processing system can use a smooth adjustment mechanism: first, set the minimum interval time for parameter adjustment to avoid too frequent adjustment; second, introduce a parameter adjustment step limit to ensure that the adjustment amplitude is within a reasonable range; and finally, set the effective value range of the parameter to prevent the parameter value from deviating from the reasonable range. Through these measures, the adjustment of the collection parameters can be more stable and reliable.
[0050] S203、When the log file change amount of the target node is greater than the collection trigger condition, perform event collection to obtain an event collection data stream.
[0051] Among them, the log file change amount represents the number of bytes of the log file increased in unit time; event collection represents a real-time collection operation triggered based on the file system event; event collection data stream refers to a sequence of real-time log data obtained through event collection. The collection trigger condition is a system-preset log change amount threshold.
[0052] The data processing system monitors the change of the log file in real time through a file system listening mechanism, and starts collection immediately when a sudden increase in log writing is detected. Specifically, the system first establishes a file system event listening channel to obtain log file writing events in real time; then calculates the file change amount in a unit of time and compares it with a preset trigger threshold; when the change amount exceeds the threshold, the system immediately starts a collection thread to read the newly added data from the last collection position and encapsulates the collected data into an event collection data stream.
[0053] In some embodiments, event collection can be implemented in various ways: optionally, the data processing system can use the file system listening interface (such as inotify) provided by the operating system to establish an event notification mechanism, realize real-time monitoring of file changes and data collection; alternatively, the system can also implement event-triggered collection by maintaining file descriptor states and periodically checking file size changes. It can be understood that other technical means can also be used to implement event-based data collection, which is not limited here. The data collection buffer size, read timeout time and other parameters also need to be reasonably configured.
[0054] In a high-concurrency scenario, frequent event triggering can cause excessive consumption of system resources. To this end, the data processing system can use an event aggregation mechanism: first, set a minimum event interval, and merge and process multiple events within this interval; second, establish an event queue to cache and batch process multiple events within a short period of time; and finally, implement a back pressure mechanism to appropriately reduce the event collection frequency when the processing capacity is insufficient. In this way, data real-time can be ensured while avoiding system overload.
[0055] S204, when the log file change amount is less than or equal to the collection trigger condition, performing polling collection according to a collection period to obtain a periodic collection data stream.
[0056] Among them, polling collection means a collection method of scanning and reading the log file at fixed time intervals; the collection period refers to the time interval between two polling collections; the periodic collection data stream refers to the log data sequence obtained by polling. The file change amount refers to the size change of the log file within the detection period.
[0057] When the log writing frequency is relatively stable, the data processing system uses a periodic polling method for data collection to reduce system overhead. Specifically, the system checks the log file state at a preset collection period, obtains the current file size and compares it with the last collection position to calculate the amount of new data; if the file has been updated, the new content is read and encapsulated into a periodic collection data stream; at the same time, the file reading position record is updated to prepare for the next collection. The system also needs to dynamically adjust the collection period according to the collection result to ensure the balance between collection efficiency and system load.
[0058] In some embodiments, the polling collection can be implemented in various ways: optionally, the data processing system can use a timing task scheduling framework to perform the collection task according to a configured cron expression, supporting flexible collection strategy configuration; optionally, the system can also use a thread pool management mechanism to achieve periodic collection by controlling the sleep and wake-up of the collection thread. It can be understood that other scheduling methods can also be used to achieve periodic data collection, which is not limited here. Priority management and resource quota control of the collection task also need to be considered.
[0059] In practical applications, fixed-period polling collection may not be able to adapt well to changes in log write rate. In this regard, the data processing system can implement an adaptive collection period adjustment mechanism: first, analyze historical data to calculate the time distribution characteristics of log writing; then dynamically adjust the collection period according to the trend of the writing frequency; finally, set upper and lower limit constraints for the collection period to prevent the collection interval from being too long or too short. This way, the polling collection can better adapt to changes in business load.
[0060] S205, calculate the log coverage range based on the event collection data stream and the periodic collection data stream, generate data compensation instructions and perform incremental collection to obtain a compensation data stream.
[0061] Among them, the log coverage range represents the completeness of the collected data in the time dimension; the data compensation instruction is a collection task description used to fill in the data gap; the incremental collection means supplementing collection for a specific time period of data gap; the compensation data stream is a sequence of supplementary data obtained by incremental collection.
[0062] The data processing system needs to ensure the integrity of the collected data by calculating the data coverage and supplementing the missing data. Specifically, the system first analyzes the data timestamps of event collection and periodic collection to identify possible data gap intervals; then generates corresponding compensation collection tasks according to the gap intervals, including time range, priority, etc.; finally, execute the incremental collection task, and encapsulate the collected supplementary data as a compensation data stream. The system also needs to maintain the version information of the data to ensure that the compensation data can be correctly merged into the complete data set.
[0063] In some embodiments, data compensation can be implemented in various ways: optionally, the data processing system can use a sliding time window method to detect data gaps by comparing the continuity of adjacent time windows; optionally, the system can also maintain data checksums to identify areas that need to be compensated by comparing the integrity information of data from different collection sources. It can be understood that other algorithms can also be used to implement data integrity checking and compensation, which is not limited here. The processing strategy for data duplication and conflict also needs to be considered.
[0064] In a large-scale distributed environment, frequent data compensation can cause the system to be overloaded. To this end, the data processing system can implement an intelligent compensation scheduling mechanism: first, prioritize compensation tasks, and process data gaps with high importance first; second, implement batch processing of compensation tasks, and execute multiple small compensation tasks together; and finally, set a limit on the concurrency of compensation execution to avoid excessive pressure on the system. This can ensure data integrity while maintaining stable system operation.
[0065] S206, merging the event collection data stream, the periodic collection data stream, and the compensation data stream into complete log data, and performing type differentiation on the complete log data to generate multi-modal data packets.
[0066] Among them, the complete log data represents the complete data set after merging and deduplication processing; the multi-modal data packet refers to a set of data packets classified according to data types and characteristics; type differentiation refers to the process of classifying data according to data format, content characteristics, etc.; data stream merging refers to the process of integrating data from different sources in chronological order.
[0067] The data processing system needs to integrate and classify data streams from multiple sources to form structured data packets. Specifically, the system first aligns and sorts the timestamps of the three data streams, and performs deduplication processing by comparing the unique identifiers of the data; then, according to the preset data pattern rules, the data content is parsed and feature extracted, and the type attribute of the data is identified; finally, data with the same type characteristics are grouped and packaged, and corresponding metadata tags are added, generating multi-modal data packets. The system also needs to ensure the integrity and consistency of the data packets to ensure that subsequent processing can correctly identify and use the data.
[0068] In some embodiments, data merging and classification can be achieved in various ways: optionally, the data processing system can use a rule-based engine method to identify and process different types of data by configuring flexible classification rules; alternatively, the system can also use a machine learning model to automatically classify data by learning the features of the data content. It can be understood that other technical means can also be used to achieve intelligent classification and processing of data, which are not limited here. Data format conversion and encoding consistency processing also need to be considered.
[0069] In the case of large data volume and complex types, the data merging and classification process may encounter performance bottlenecks. To this end, the data processing system can implement a parallel processing mechanism: first, shard the data according to time segments to achieve parallel merging of data; second, use a multi-level classification strategy to first perform coarse-grained classification and then refine the processing; finally, use a caching mechanism to save commonly used classification results to improve processing efficiency. This can improve the throughput and response speed of data processing.
[0070] S207, distribute the multi-modal data packet to the corresponding storage engine to obtain multi-modal log aggregation data.
[0071] Among them, the storage engine represents a dedicated storage system for different types of data, including a text index engine, a time series data engine, and an object storage engine; data distribution refers to the process of selecting the appropriate storage engine according to data characteristics; multi-modal log aggregation data refers to a complete data set stored in different storage engines.
[0072] The data processing system needs to select the most suitable storage method according to the data characteristics to achieve efficient data management. Specifically, the system first analyzes the storage requirements of the multi-modal data packet, including access mode, query characteristics, storage period, etc.; then according to the characteristics of different storage engines, it formulates a data distribution strategy and routes the data packet to the corresponding storage engine; finally, it establishes an index structure in each storage engine to optimize data access performance. The system also needs to maintain the location mapping relationship of the data to support cross-storage engine data retrieval and correlation analysis.
[0073] In some embodiments, data distribution and storage can be achieved in various ways: optionally, the data processing system can use a strategy-based storage routing to determine the storage location and method of data by configuring rules; optionally, the system can also implement dynamic load balancing to adaptively adjust the data distribution strategy according to the load status of the storage engine. It can be understood that other methods can also be used to achieve intelligent storage management of data, which is not limited here. It is also necessary to consider the life cycle management and storage cost optimization of data.
[0074] When multiple storage engines work together, data consistency and performance balancing problems may occur. To this end, the data processing system can implement a hierarchical storage architecture: first, establish a data temperature grading mechanism to determine the storage level of data according to access frequency; second, implement data automatic migration to transfer cold data to a lower-cost storage layer; finally, provide a unified access interface to mask the differences between the underlying storage. This way, you can optimize the use efficiency of storage resources while ensuring data availability.
[0075] In the above embodiment, real-time perception of log file changes is achieved through an event listening mechanism. In actual application, different listening strategies can also be configured according to business characteristics, such as enabling more fine-grained listening for key business logs and using a larger trigger threshold for ordinary logs to further optimize system resource usage. The scenario of this embodiment is supplemented as follows.
[0076] Based on the application of the present scheme, an intelligent manufacturing enterprise further optimizes the storage management strategy. The system allocates data to different storage levels according to the importance and access characteristics of log data. Real-time alarm logs are stored in the high-performance SSD layer, ordinary operation logs are stored in the HDD layer, and historical logs are automatically archived to object storage. Through machine learning algorithm analysis of data access patterns, the system can predict the hot and cold change trend of data and perform data migration in advance. For example, when it is predicted that a certain type of device will enter the maintenance period, the related historical logs are migrated to the fast storage layer in advance to facilitate maintenance personnel to query. This intelligent storage management not only optimizes the storage cost, but also improves the data access efficiency.
[0077] After combining the above scenarios, the following further more specific flow description of the method provided by the present embodiment is given. Please refer to Figure 3 , another flowchart of the distributed operation and maintenance data governance method based on the multi-modal log aggregation framework in the embodiments of the present application.
[0078] S301, acquire the system resource occupation rate and log write frequency of the target node, and calculate the node load state index.
[0079] Referring to step S201, the data processing system will collect and analyze the node resource usage and log write condition.
[0080] In some embodiments, the data processing system will collect and monitor node resource indicators and log write conditions, and calculate the load state, that is, the data processing system will collect the processor occupation rate, memory usage rate and disk read-write time of the target node, and monitor the log file size change amount and write operation count of the target node; calculate the system resource index according to the processor occupation rate, memory usage rate and disk read-write time; calculate the log write index according to the log file size change amount and write operation count; normalize the system resource index and the log write index to obtain the node load state index.
[0081] Among them, the processor occupation rate represents the percentage value of CPU usage; the memory usage rate refers to the proportion of memory usage to total memory; the disk read-write time represents the average response time of I / O operation; the log file size change amount refers to the byte number of log file growth per unit time; the write operation count represents the number of log write per unit time; the system resource index refers to the comprehensive index reflecting the overall resource usage of the system; the log write index is used to represent the quantitative value of log data generation rate; the normalization calculation refers to the mathematical processing process of converting different dimension indexes to a unified interval; the node load state index represents the standardized value of the current workload of the target node.
[0082] The data processing system needs to comprehensively evaluate the running state of the node when starting resource monitoring. Specifically, the data processing system first collects the usage of CPU, memory and disk I / O through the operating system interface every fixed time interval (such as 1 second), and records the growth rate and writing frequency of the log file; then the CPU, memory and disk indicators are weighted and summed according to the preset weight (such as 0.4, 0.3, 0.3) to obtain the system resource indicator; the size change and writing count of the log file are also weighted and calculated to obtain the log writing indicator; finally, the two indicators are mapped to the [0, 1] interval using the Min-Max normalization method, and the final load state indicator is calculated according to the importance ratio of the resource indicator and the writing indicator (such as 6:4).
[0083] In some embodiments, node load calculation can be achieved in various ways: optionally, the data processing system can use a sliding window average method, collect raw indicators every second within a 5-minute time window, calculate the average value after removing outliers, then map each indicator to the (0, 1) interval using a sigmoid function, and finally obtain the load indicator by weighted summation; optionally, the data processing system can also use an exponential moving average method, apply a decreasing weight sequence to the last 30 sampling data, calculate the weighted average of each indicator, and then normalize it by a logarithmic function. It can be understood that other mathematical models can also be used to calculate the load indicator, which is not limited here. It is also necessary to consider the sampling frequency, outlier processing and weight dynamic adjustment mechanism when making supplementary explanations.
[0084] During the load indicator calculation process, there may be a problem that the load mutation caused by too large sampling interval cannot be perceived in time. For this purpose, the data processing system can implement an adaptive sampling mechanism: first, establish a load change rate monitoring model, when a rapid load change is detected (such as a change rate exceeding 50% / s), automatically increase the sampling frequency (such as from 1 time / second to 4 times / second); at the same time, use Kalman filtering algorithm to filter the high-frequency sampling data to avoid false load fluctuations; finally, use exponential smoothing method to predict the trend of the filtered data and give an early warning of possible load peaks. In this way, a balance between resource overhead and monitoring accuracy can be achieved.
[0085] S302, according to the node load state indicator, determine the node collection parameters including the collection period and the collection trigger condition.
[0086] Referring to step S202, the data processing system will dynamically adjust the log collection strategy based on the node state.
[0087] In some embodiments, the data processing system dynamically adjusts the collection strategy parameters according to the node load, that is, the data processing system determines the load level of the target node according to the node load state indicator, and determines the cycle reference value according to the load level; obtains the success ratio of the historical collection success times and the total collection times of the target node, and adjusts the cycle reference value according to the success ratio to obtain the collection cycle; and determines the collection trigger condition according to the log write frequency of the target node.
[0088] Among them, the load level represents the classification result of the node load state, which is usually divided into three levels of low, medium and high; the cycle reference value is the standard reference value of the collection cycle under different load levels; the historical collection success times represent the number of collection tasks successfully completed in the past period of time; the total collection times are the total number of collection tasks executed; the success ratio represents the success rate of the collection task; the collection cycle is the time interval between two collection operations; and the collection trigger condition is the log change threshold that triggers real-time collection.
[0089] When determining the collection parameters, the data processing system needs to comprehensively consider the node load and the historical collection effect. Specifically, the data processing system first compares the node load state indicator with the preset level threshold (such as 0.3 and 0.7), and divides the load state into three levels of low load (0-0.3), medium load (0.3-0.7) and high load (0.7-1.0); then sets the corresponding cycle reference value according to different load levels (such as low load 5 seconds, medium load 10 seconds, and high load 20 seconds); then statistics the collection task execution in the last 24 hours, calculates the success ratio and corrects the cycle reference value according to the preset adjustment coefficient, such as extending the collection cycle by 1.5 times when the success ratio is less than 80%; and finally analyzes the recent log write frequency distribution characteristics of the target node, and sets an appropriate trigger threshold to ensure that the collection can be triggered in time when the log suddenly increases.
[0090] In some embodiments, the determination of the collection parameters can be realized in various ways: optionally, the data processing system can adopt an adaptive adjustment method based on historical data, predict the collection success probability under different cycles by establishing a regression model of collection success rate and collection cycle, select the minimum cycle value that can guarantee 90% success rate, and set a fluctuation range according to the load level; optionally, the data processing system can also adopt a parameter optimization method based on time series analysis, analyze the periodicity of log writing by using ARIMA model, predict the future writing peak, and dynamically adjust the collection trigger condition. It can be understood that other algorithms can also be used to realize the dynamic optimization of parameters, which are not limited here. It is also necessary to consider the smoothness and stability control mechanism of parameter adjustment when making supplementary explanations.
[0091] In the scenario of business load mutation, the fixed parameter adjustment strategy may not respond quickly. To this end, the data processing system can implement a predictive parameter adjustment mechanism: first, a load change prediction model is constructed, and a time series analysis method is used to predict the load trend in the next 30 minutes; then, based on the prediction result, the collection parameters are adjusted in advance, such as increasing the collection period in advance when the load is rising; at the same time, a parameter adjustment buffer interval is set to avoid frequent fluctuations in parameters due to prediction errors; finally, the accuracy of the prediction model is continuously optimized by comparing the predicted value with the actual value. In this way, the load change can be responded to in advance, and the stable execution of the collection task can be ensured.
[0092] S303, a file system monitoring channel of the log file is constructed, a data collection buffer is created based on a file change notification mechanism of the corresponding file system monitoring channel, and a data temporary storage area is obtained.
[0093] Among them, the file system monitoring channel represents a system-level event notification mechanism for monitoring file changes; the file change notification mechanism refers to the file operation event notification service provided by the operating system; the data collection buffer refers to a memory area for temporarily storing collected data; and the data temporary storage area represents a management unit containing file location information and data cache.
[0094] The data processing system needs to establish an efficient file monitoring mechanism to ensure that it can capture changes in log files in a timely manner. Specifically, the system first registers the file system event listener, configures the file operation types and event filtering rules that need to be monitored; then initializes the monitoring thread pool, establishes the event notification queue and processing pipeline; finally, a data buffer associated with the monitoring channel is created for temporarily storing the collected data. The system also needs to maintain the state information of the monitoring channel to ensure that the monitoring service can be restored in time in the event of an exception.
[0095] In some embodiments, file monitoring can be implemented in various ways: optionally, the data processing system can use the inotify service of the operating system to implement real-time file change notification by registering a file event listener; alternatively, the system can also use a polling-based file monitoring method to identify changes by periodically checking the file state. It can be understood that other technologies can also be used to implement file change monitoring, which is not limited here. The fault tolerance and recovery mechanism of the monitoring channel also needs to be considered.
[0096] In a high-concurrency scenario, the file monitoring channel may have event accumulation and processing delay. To this end, the data processing system can implement a multi-level buffering mechanism: first, set up an event preprocessing queue to filter and aggregate the original events; second, implement dynamic expansion of the buffer, automatically adjust the buffer size according to the data volume; finally, establish an overflow handling mechanism to trigger emergency data landing when the buffer is about to overflow. In this way, the stability of the system under sudden high load can be improved.
[0097] In some embodiments, the data processing system can establish a file monitoring mechanism and create a data buffer, i.e., the data processing system can register an event listener of a file system to obtain an event notification handle; determine a listening type corresponding to the event notification handle to obtain a file operation filter; determine an event response strategy according to a trigger threshold of the file operation filter, and start a listening thread of the event listener to obtain a listening running state; initialize a file change notification mechanism according to the listening running state, and create a data collection buffer based on the file change notification mechanism to obtain a data temporary storage area.
[0098] In some embodiments, the data processing system can establish a file monitoring mechanism and create a data buffer, i.e., the data processing system can register an event listener of a file system to obtain an event notification handle; determine a listening type corresponding to the event notification handle to obtain a file operation filter; determine an event response strategy according to a trigger threshold of the file operation filter, and start a listening thread of the event listener to obtain a listening running state; initialize a file change notification mechanism according to the listening running state, and create a data collection buffer based on the file change notification mechanism to obtain a data temporary storage area.
[0099] When starting the file monitoring service, the data processing system needs to establish a complete event processing link. Specifically, the data processing system first registers a file system event listener with the operating system to obtain an event notification handle for subsequent operations; then configures the working parameters of the listener, including the file operation types (such as write, append, truncate, etc.) to be monitored and the corresponding filtering rules; then sets the event response strategy according to the system resource status and business requirements, such as event cache size, processing timeout, etc.; then starts a special listening thread to initialize the event queue and processing pipeline; finally, establish a file change notification mechanism, including configuring the delivery method of event notification, creating a data buffer, etc., to form a complete data temporary storage area.
[0100] In some embodiments, the file monitoring service can be implemented in various ways: optionally, the data processing system can use a file system hook-based implementation method, inject a listening code at a key operation point of the file system, set event filtering conditions and trigger thresholds, configure event cache strategies, and finally start a listening thread pool to process the captured events; optionally, the data processing system can also use a polling-based implementation method, by setting a polling interval, establishing a file state cache, configuring difference comparison rules, and starting a polling thread to detect file changes. It can be understood that other technologies can also be used to implement file change monitoring, which is not limited here. It is also necessary to consider the event loss retry mechanism and abnormal recovery strategy when making the supplementary explanation.
[0101] In high-concurrency write scenarios, event processing may be delayed and accumulated. To this end, the data processing system can implement an adaptive event processing mechanism: first, an event priority model is established, and processing priorities are assigned to events according to file importance and change frequency; then, dynamic thread pool management is implemented, and the number of processing threads is automatically adjusted according to the length of the event queue; at the same time, a batch processing strategy is adopted, and multiple similar events in a short period of time are combined for processing; finally, an overflow processing mechanism is set, which triggers an emergency processing procedure when the event queue approaches the upper limit of the capacity. In this way, file change events can be processed in a timely manner under various load conditions.
[0102] S304, record the current reading position information of the log file in the data temporary storage area to obtain a file reading position record.
[0103] Among them, the reading position information represents the offset and timestamp metadata of the file; the file reading position record refers to the data structure used to track the file reading progress; the data temporary storage area is a memory space for storing file position information and cached data.
[0104] The data processing system needs to accurately record the reading position of the file to ensure the continuity and integrity of data collection. Specifically, the system first obtains the current offset and last modification time of the file; then encapsulates this information into a position record object, including file descriptor, byte offset, timestamp, etc.; finally, the position record is stored in the temporary storage area and indexed to support fast lookup. The system also needs to periodically persist the position record to prevent progress loss due to system exceptions.
[0105] In some embodiments, position record management can be implemented in various ways: optionally, the data processing system can use a memory mapping-based approach to directly map the position information to the file system; alternatively, the system can use a distributed cache to store the position record in a distributed memory. It can be understood that other ways can also be used to manage file position information, which are not limited here. Version control and concurrent access of the position record also need to be considered.
[0106] In the event of an abnormal system restart, the position record may not be accurate. To this end, the data processing system can implement a position recovery mechanism: first, establish a checkpoint mechanism for the position record, and save reliable position information regularly; second, implement position verification logic to verify the accuracy of the position by file signature; finally, provide a rollback mechanism that can roll back to the nearest reliable position when an exception is detected. In this way, the system can still accurately track the file reading position in abnormal situations.
[0107] S305, detect the file descriptor state of the log file according to the file reading position record to obtain file state information.
[0108] wherein, the file descriptor status represents the system-level status of the file such as open, close, rename, etc.; the file status information refers to the information set containing attributes such as file size, access permission, modification time, etc.; the file read position record is used to locate the current read progress of the file.
[0109] The data processing system needs to monitor the file status in real time to adapt to the dynamic changes of the file. Specifically, the system first acquires the file descriptor according to the file read position record, checks whether the file is accessible; then acquires the detailed attribute information of the file, including file size, modification time, access permission, etc.; finally, integrates and compares these status information with the historical record to identify whether the file has undergone renaming, rotation, etc. The system also needs to maintain the change history of the file status to support status backtracking in abnormal cases.
[0110] In some embodiments, file status detection can be implemented in various ways: optionally, the data processing system can use system call interface to directly acquire file status; optionally, the system can also track file status changes through file system event listening. It can be understood that other methods can also be used to detect and maintain file status, which are not limited here. The impact of file permission changes and system security policies also needs to be considered.
[0111] When the log file is rotated or deleted, it may cause the status detection to fail. For this purpose, the data processing system can implement a file tracking mechanism: first, establish a file inode mapping table to track the real location of the file through inode; second, implement file alias resolution to handle file renaming and soft / hard linking; finally, provide file recovery strategy to try to reposition the file when it is inaccessible. This can ensure the accuracy of status detection in various file operation scenarios.
[0112] S306, updating the file read position record based on the file status information to obtain the latest file read position identifier.
[0113] wherein, the file read position identifier represents the latest file read progress information; the position record update refers to the process of adjusting the read position according to the file status change; the latest position identifier contains the updated offset and timestamp information.
[0114] The data processing system needs to update the read position information in a timely manner according to the changes of the file status. Specifically, the system first analyzes the change content in the file status information to determine whether the read position needs to be adjusted; then calculates the new read offset according to the file size change and updates the related timestamp information; finally, the updated position information is stored persistently to ensure that the collection task can continue to execute from the correct position. The system also needs to handle the position update of special cases such as file truncation, content rollback, etc.
[0115] In some embodiments, the position update can be implemented in various ways: optionally, the data processing system can use atomic operations to ensure consistency of the position update; optionally, the system can also use a write-ahead logging mechanism to record the operation sequence of the position update. It can be understood that other techniques can also be used to ensure the reliability of the position update, which are not limited here. It is also necessary to consider the processing of concurrent updates and version conflicts.
[0116] In the high-frequency file change scenario, frequent position updates can affect system performance. In this regard, the data processing system can implement a batch update mechanism: first, set a buffer for position updates to temporarily store multiple updates in a short period of time; second, implement an incremental update strategy to only record the position information that has changed; and finally, use an asynchronous persistence method to decouple the position update operation from the collection process. This can improve system performance while ensuring the accuracy of the position.
[0117] S307、In the case that the log file change amount of the target node is greater than the collection trigger condition, event collection is performed to obtain an event collection data stream.
[0118] Referring to step S203, the data processing system will trigger event collection when the log change amount exceeds the threshold.
[0119] S308、In the case that the log file change amount is less than or equal to the collection trigger condition, polling collection is performed according to the collection period to obtain a periodic collection data stream.
[0120] Referring to step S204, the data processing system will perform periodic collection when the log change amount is small.
[0121] S309、Based on the event collection data stream and the periodic collection data stream, the log coverage range is calculated, the data compensation instruction is generated and the incremental collection is performed to obtain a compensation data stream.
[0122] Referring to step S205, the data processing system will calculate the data coverage and supplement the missing data.
[0123] S310、The event collection data stream, the periodic collection data stream and the compensation data stream are merged into complete log data, and the complete log data is type-distinguished to generate a multi-modal data packet.
[0124] Referring to step S206, the data processing system will integrate multi-source collection data and perform modal classification.
[0125] S311、The multi-modal data packet is distributed to the corresponding storage engine to obtain multi-modal log aggregation data.
[0126] Referring to step S207, the data processing system will distribute the data to the appropriate storage engine for persistence.
[0127] S312, constructing a data index structure based on the multi-modal log aggregated data, to obtain a retrieval access interface.
[0128] Wherein, the data index structure represents an organization form for accelerating data retrieval; the retrieval access interface refers to a program interface that provides a unified data access method; the multi-modal log aggregated data contains a collection of log information of different types and formats.
[0129] The data processing system needs to establish an efficient indexing mechanism for different types of data. Specifically, the system first analyzes the access mode and query characteristics of the data to determine the dimension and granularity of the index; then selects appropriate index structures according to different data types, such as using inverted indexes for text data and timeline indexes for time series data; finally, a unified access interface is constructed to support multi-dimensional data retrieval and aggregation analysis. The system also needs to maintain index updates and optimization to ensure that retrieval performance remains stable as data volume grows.
[0130] In some embodiments, index construction can be achieved in various ways: optionally, the data processing system can use a hierarchical index structure to keep hot data indexes in memory; optionally, the system can also use spatial indexing techniques to support multi-dimensional data fast positioning. It can be understood that other indexing techniques can also be used to optimize data access efficiency, which are not limited here. The storage overhead and maintenance cost of the index also need to be considered.
[0131] In the scenario of frequent updates of index structure, performance jitter may occur. To this end, the data processing system can implement an index update optimization mechanism: first, use incremental index update strategy to only rebuild the index of the changed data; second, implement index sharding mechanism to distribute large indexes to multiple nodes; finally, provide index preheating function to pre-load commonly used indexes when the system load is low. In this way, stable retrieval performance can be provided.
[0132] S313, determining the life cycle rules of the multi-modal log aggregated data according to the retrieval access interface, to obtain a storage management strategy.
[0133] Wherein, the life cycle rules represent the management strategy of data from creation to archiving or deletion; the storage management strategy refers to a set of rules that control data storage location and retention time; the retrieval access interface provides data access and analysis capabilities.
[0134] A data processing system needs to develop a reasonable storage strategy according to the importance of data and access characteristics. Specifically, the system first analyzes the access frequency and business value of the data, divides the data into different importance levels; then according to the storage cost and performance requirements, it develops a corresponding storage period and location strategy for each level; finally, it realizes the automatic archiving and cleaning mechanism to ensure the effective use of storage resources. The system also needs to provide a policy adjustment interface to support dynamic modification of management rules according to business needs.
[0135] In some embodiments, storage management can be achieved in various ways: optionally, the data processing system can use a time-based hierarchical storage strategy, gradually migrating data to low-cost storage as it ages; optionally, the system can also use a dynamic storage strategy based on access frequency to keep hot data in high-performance storage. It can be understood that other methods can also be used to optimize storage resource usage, which are not limited here. Data compliance and security requirements also need to be considered.
[0136] During the execution of the storage strategy, resource allocation may not be balanced. For this, the data processing system can implement an intelligent storage scheduling mechanism: first, establish a storage resource usage model to predict the storage needs of various types of data; second, implement storage quota management to prevent a single business from occupying too many resources; finally, provide a storage balancing mechanism to dynamically adjust data distribution among multiple storage nodes. This can achieve the rational use of storage resources.
[0137] S314, hierarchical storage of the multi-modal log aggregation data based on the storage management strategy is performed to obtain a hierarchical storage architecture.
[0138] Among them, the hierarchical storage architecture represents a hierarchical storage structure that allocates data to different storage media according to access characteristics; hierarchical storage refers to the process of selecting the appropriate storage level according to data characteristics; the storage management strategy defines the rules for data migration between different storage levels.
[0139] A data processing system needs to establish a multi-level storage architecture to optimize storage cost and access performance. Specifically, the system first divides the storage resources into multiple levels, such as memory layer, SSD layer, disk layer, and archive layer, etc.; then according to the access frequency, importance, and other characteristics of the data, it determines the initial storage level of the data; finally, it establishes data migration channels between levels to support automatic migration of data between different levels according to changes in access patterns. The system also needs to monitor the resource usage of each storage level to ensure the balance of storage capacity and performance.
[0140] In some embodiments, hierarchical storage can be implemented in various ways: optionally, a data processing system can use a time-based tiering strategy to store the latest data in a high-speed tier; optionally, the system can also use a value-based tiering approach to select a storage tier according to the business importance of the data. It can be understood that other tiering strategies can also be used to optimize storage effects, which are not limited here. The performance impact during data migration also needs to be considered.
[0141] In the scenario of rapid growth of data volume, unreasonable allocation of storage tiers may occur. To this end, the data processing system can implement a dynamic tier optimization mechanism: first, a data access pattern analysis model is established to predict the hot and cold change trend of the data; second, the elastic expansion and contraction of the storage tier is realized, and the resource ratio of each tier is dynamically adjusted according to the load; finally, the load balancing function between tiers is provided to avoid resource bottlenecks in a certain tier. In this way, the efficient operation of the storage architecture can be maintained.
[0142] S315, updating the storage state of the multi-modal log aggregation data based on the hierarchical storage architecture.
[0143] Among them, the storage state represents the location, access permission and other attribute information of the data in the storage system; the storage state update refers to the process of updating the metadata according to the data migration; the hierarchical storage architecture defines the storage tiers and migration rules of the data.
[0144] The data processing system needs to maintain accurate data storage state information to support efficient data access. Specifically, the system first monitors the migration of data between tiers, records the current location and state of the data block; then updates the access path and permission information of the data to ensure that the application can correctly access the migrated data; finally, the data index is updated synchronously to enable the retrieval request to quickly locate the latest location of the data. The system also needs to ensure the atomicity of state update to avoid data access errors.
[0145] In some embodiments, state update can be implemented in various ways: optionally, a data processing system can use a distributed metadata service to centrally manage the storage state information of the data; optionally, the system can also use a local cache mechanism to speed up the state query of frequently accessed data. It can be understood that other ways can also be used to manage the data state, which are not limited here. The disaster recovery mechanism of state information also needs to be considered.
[0146] In the process of large-scale data migration, state update delay or inconsistency may occur. To this end, the data processing system can implement a state synchronization mechanism: first, a version control strategy is adopted to track the change history of data state; second, incremental state synchronization is implemented to update only the changed state information; and finally, a state repair function is provided to automatically correct the state when inconsistency is detected. In this way, the accuracy and reliability of data state information can be ensured.
[0147] In some embodiments, the data processing system implements data tiered storage and hotness migration management, that is, the data processing system obtains the data access frequency of each storage layer in the tiered storage architecture to obtain the data hotness distribution; calculates the storage level migration threshold according to the data hotness distribution to obtain the data migration strategy; identifies the data set to be migrated based on the data migration strategy to obtain a migration task list; executes the data migration operation in the migration task list to obtain a migration execution result, and updates the storage location information of the multi-modal log aggregation data according to the migration execution result.
[0148] Among them, the data access frequency represents the number of times of reading or writing data; the data hotness distribution refers to the statistical distribution of different data in the access frequency; the storage level migration threshold represents the condition value that triggers the migration of data between different storage layers; the data migration strategy is used to define the rules for transferring data between storage levels; the migration task list refers to the set of data migration operations to be executed; the migration execution result represents the completion status of the data migration operation; and the storage location information refers to the specific location identifier of the data in the tiered storage architecture.
[0149] When managing tiered storage, the data processing system needs to dynamically optimize data distribution. Specifically, the data processing system first collects access statistics of data in each storage layer, including read / write times, access time distribution, etc., to build a data hotness profile; then calculates the optimal migration threshold based on the preset storage level characteristics (such as performance, cost, etc.), for example, when the access frequency of data in the high-performance layer is less than 1 time per hour, trigger migration to the low-speed layer; then scan the data of each storage layer, identify the data set that meets the migration condition, and generate a migration task list containing source location, target location, priority, etc.; finally, execute data migration according to the task priority, and update the location index information of the data after migration to ensure correct routing of subsequent access.
[0150] In some embodiments, data migration management can be achieved in various ways: optionally, the data processing system can adopt a migration method based on heat prediction, predict the access trend in the next 24 hours by establishing a time series model of data access, identify the inflection point of access heat change, plan the data migration path in advance, and finally execute the migration task according to the principle of minimum impact; optionally, the data processing system can also adopt a migration method based on cost optimization, calculate the cost-effective index of different storage schemes by constructing a comprehensive evaluation model of storage cost and access delay, and select the migration scheme with the optimal total cost. It can be understood that other algorithms can also be used to achieve data migration optimization, which is not limited here. It is also necessary to consider the data consistency guarantee and failure recovery mechanism during the migration process.
[0151] In the process of large-scale data migration, the problem of unbalanced load of storage hierarchy may occur. For this purpose, the data processing system can implement a load balancing migration scheduling mechanism: first, establish a load monitoring model of the storage hierarchy, and collect indicators such as IOPS and bandwidth utilization of each layer in real time; then calculate the load balancing factor, and trigger load balancing when the load of a certain level exceeds 150% of the average value; then optimize the scheduling order of the migration task through genetic algorithm, so that the load of each storage layer tends to be balanced; finally, implement a migration flow control mechanism to control the amount of migration data in a unit of time, and avoid affecting normal business access. In this way, the stable operation and efficient utilization of the storage system can be ensured.
[0152] In the embodiments of the present application, due to the adoption of the adaptive collection mechanism based on the node load state, the real-time monitoring mechanism based on the file system event, and the multi-level intelligent storage architecture, the dynamic adjustment of log collection parameters, the real-time response of data collection, and the intelligent allocation of storage resources can be realized, effectively solving the problems of system resource waste, data collection delay, and low storage efficiency caused by fixed collection strategies in traditional log collection systems, and further realizing the efficiency, reliability, and economy of log collection. Specifically, intelligent adjustment of collection parameters is achieved through load sensing, avoiding performance impact on business; the real-time and completeness of data collection are ensured through event listening; and the use efficiency of storage resources is optimized through layered storage, while providing flexible data access capabilities. The overall scheme not only improves the performance and reliability of the log collection system, but also reduces the system maintenance cost and storage overhead.
[0153] The data processing system in the embodiments of the present application will be described from the perspective of hardware processing. Please refer to Figure 4 , which is a schematic diagram of an entity device structure of the data processing system in the embodiments of the present application.
[0154] It should be noted that Figure 4The illustrated structure of the data processing system is only one example and should not impose any limitations on the functions and the range of use of the embodiments of the present application.
[0155] As shown in FIG. 4, the data processing system includes a CPU 401 which can perform various appropriate actions and processes in accordance with a program stored in a ROM 402 or a program loaded from a storage section 408 into a RAM 403, such as the method described in the above embodiments. In the RAM 403, various programs and data required for the operation of the system are also stored. The CPU 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An I / O interface 405 is also connected to the bus 404. Figure 4
[0156] The following components are connected to the I / O interface 405: an input section 406 including an audio input device, a push button switch, and the like; an output section 407 including a Liquid Crystal Display (LCD), an audio output device, an indicator lamp, and the like; the storage section 408 including a hard disk, and the like; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card, a modem, and the like. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as necessary. A removable media 411 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is attached to the drive 410 as necessary, so that a computer program read therefrom is installed into the storage section 408 as necessary.
[0157] In particular, according to embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program comprising computer programs for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication section 409, and / or installed from the removable media 411. When the computer program is executed by the CPU 401, various functions defined in the present application are performed.
[0158] The flowcharts and the block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved.
[0159] Specifically, the data processing system of the embodiment includes a processor and a memory, and the memory stores a computer program. When the computer program is executed by the processor, the distributed operation and maintenance data governance method based on the multi-modal log aggregation framework is realized.
[0160] As another aspect, the application further provides a computer-readable storage medium. The storage medium can be included in the data processing system described in the above embodiments, or can exist independently without being assembled into the data processing system. The storage medium carries one or more computer programs. When the one or more computer programs are executed by a processor of the data processing system, the data processing system realizes the distributed operation and maintenance data governance method based on the multi-modal log aggregation framework provided in the above embodiments.
[0161] The above-described embodiments are only used to illustrate the technical solutions of the present application, but not limit the present application; even though the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
[0162] In the above embodiments, according to the context, the term "when" can be interpreted as meaning "if" or "after" or "in response to determining" or "in response to detecting". Similarly, according to the context, the phrase "upon determining" or "if detecting (the stated condition or event)" can be interpreted as meaning "if determining" or "in response to determining" or "upon detecting (the stated condition or event)" or "in response to detecting (the stated condition or event)".
Claims
1. A distributed operation and maintenance data governance method based on a multi-modal log aggregation framework, characterized in that, Applied to a data processing system, the method comprises: Obtaining the system resource occupation rate and log write frequency of a target node, and calculating a node load state index; According to the node load state index, determining the node collection parameters including the collection period and the collection trigger condition; When the log file change amount of the target node is greater than the collection trigger condition, performing event collection to obtain an event collection data stream; When the log file change amount is less than or equal to the collection trigger condition, performing polling collection according to the collection period to obtain a periodic collection data stream; Based on the event collection data stream and the periodic collection data stream, calculating the log coverage range, generating a data compensation instruction and performing incremental collection to obtain a compensation data stream; Merging the event collection data stream, the periodic collection data stream and the compensation data stream into complete log data, and performing type differentiation on the complete log data to generate a multi-modal data package; Distributing the multi-modal data package to the corresponding storage engine to obtain multi-modal log aggregation data; the storage engine includes a text index engine, a time series data engine and an object storage engine.
2. The method of claim 1, wherein, The step of obtaining the system resource occupation rate and log write frequency of the target node, and calculating the node load state index, specifically comprises: Collecting the processor occupation rate, memory usage and disk read-write time of the target node, and monitoring the log file size change amount and write operation count of the target node; According to the processor occupation rate, memory usage and disk read-write time, calculating the system resource index; According to the log file size change amount and write operation count, calculating the log write index; Normalizing the system resource index and the log write index to obtain the node load state index.
3. The method of claim 1, wherein, The step of determining the node collection parameters including the collection period and the collection trigger condition according to the node load state index, specifically comprises: According to the node load state index, determining the load level of the target node, and according to the load level, determining the periodic reference value; Obtaining the success ratio of the historical collection success times and the total collection times of the target node, and adjusting the periodic reference value according to the success ratio to obtain the collection period; According to the log write frequency of the target node, determining the collection trigger condition.
4. The method of claim 1, wherein, Before the step of performing event collection to obtain an event collection data stream, the method further comprises: Building a file system listening channel of the log file, creating a data collection buffer based on the file change notification mechanism corresponding to the file system listening channel to obtain a data temporary storage area; Recording the current reading position information of the log file in the data temporary storage area to obtain file reading position records; According to the file reading position records, detecting the file descriptor state of the log file to obtain file state information; Based on the file state information, updating the file reading position records to obtain the latest file reading position identifier.
5. The method of claim 4, wherein, The step of constructing the file system monitoring channel of the log file, creating a data collection buffer based on a file change notification mechanism corresponding to the file system monitoring channel, and obtaining a data temporary storage area specifically includes: Registering an event listener of a file system to obtain an event notification handle; Determining a monitoring type corresponding to the event notification handle to obtain a file operation filter; Determining an event response strategy according to a trigger threshold of the file operation filter, and starting a monitoring thread of the event listener to obtain a monitoring running state; Initializing a file change notification mechanism according to the monitoring running state, and creating a data collection buffer based on the file change notification mechanism to obtain a data temporary storage area.
6. The method of claim 1, wherein, After the step of distributing the multi-modal data packet to the corresponding storage engine to obtain multi-modal log aggregation data, the method further includes: Constructing a data index structure based on the multi-modal log aggregation data to obtain a retrieval access interface; Determining a life cycle rule of the multi-modal log aggregation data according to the retrieval access interface to obtain a storage management strategy; Performing hierarchical storage on the multi-modal log aggregation data based on the storage management strategy to obtain a hierarchical storage architecture; Updating a storage state of the multi-modal log aggregation data based on the hierarchical storage architecture.
7. The method of claim 6, wherein, After the step of updating the storage state of the multi-modal log aggregation data based on the hierarchical storage architecture, the method further includes: Obtaining a data access frequency of each storage layer in the hierarchical storage architecture to obtain a data heat distribution; Calculating a storage level migration threshold according to the data heat distribution to obtain a data migration strategy; Identifying a data set to be migrated based on the data migration strategy to obtain a migration task list; Performing a data migration operation in the migration task list to obtain a migration execution result, and updating storage location information of the multi-modal log aggregation data according to the migration execution result.
8. A data processing system, characterized by The data processing system includes one or more processors and a memory; the memory is coupled with the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors invoke the computer instructions to make the data processing system execute the method in any one of claims 1-7.
9. A computer-readable storage medium comprising instructions, characterized in that, When the instructions run on the data processing system, the data processing system executes the method in any one of claims 1-7.
10. A computer program product, characterised in that, When the computer program product runs on the data processing system, the data processing system executes the method in any one of claims 1-7.
Citation Information
Cited By
Data transmission method and electronic equipment
CN121560248A
Data transmission methods and electronic devices
CN121560248B