Mass real-time data scene-oriented knowledge base construction method
Through the knowledge base construction method for massive real-time data scenarios, real-time data access, observable tools, machine learning strategies and load balancing schedulers are used to solve the problems of performance bottlenecks and low data processing efficiency in traditional methods, and efficient and stable massive real-time data processing and query are achieved.
Patent Information
- Application Number
- CN202510036135.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-06-06
AI Technical Summary
When traditional knowledge base construction methods face massive real-time data environments, performance bottlenecks and low data processing efficiency make it difficult to achieve the accuracy of historical data comparison and low latency of query.
The knowledge base construction method for massive real-time data scenarios is adopted, including real-time massive data access module, observable tool real-time acquisition technology, machine learning cleaning strategy generation module and load balancing scheduler. Through cleaning, classification balancing and storage optimization, data processing efficiency and system stability are improved.
It realizes efficient cleaning and processing of massive real-time data, reduces the pressure on various components, improves the stability and reliability of the knowledge base construction system, and ensures the accuracy of historical data comparison and low latency of query.
Smart Images

Figure CN120106188A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network communication and data transmission, and relates to a knowledge base construction method for massive real-time data scenarios. Background Art
[0002] With the rapid development of Internet technology, the widespread application and continuous innovation of 5G communication technology, the amount of global data has exploded, and the generation, transmission, storage and processing of data are facing unprecedented challenges. In particular, in the face of the surge in real-time Internet traffic, how to quickly process PB-level network traffic and how to store this data efficiently and securely have become issues that companies urgently need to solve. In addition, when building a knowledge base, it is crucial to ensure the accuracy of historical data comparison and low latency of queries. However, when dealing with massive real-time data environments, traditional knowledge base construction methods are often limited by performance bottlenecks and data processing efficiency, and will face a series of problems. Summary of the invention
[0003] In view of the problems existing in the prior art, the purpose of the present invention is to provide a knowledge base construction method for massive real-time data scenarios to solve the challenges and difficulties in the process of knowledge base construction.
[0004] The present invention mainly includes the following three technical points:
[0005] The technical solution adopted by the present invention is as follows:
[0006] A method for constructing a knowledge base for massive real-time data scenarios, the steps of which include:
[0007] 1) The real-time massive data access module collects the real-time performance index data of each component in the server cluster and sends it to the message queue interceptor; the components include the database and its node server;
[0008] 2) The message queue interceptor cleans and filters the real-time performance indicator data of each component to be processed according to the current cleaning interception strategy, obtains the valid data of each component and adds it to the message queue;
[0009] 3) using observable tools to collect the status of valid data in the message queue and input it into the machine learning strategy module, generating a cleaning interception strategy and sending the message queue interceptor;
[0010] 4) The knowledge base construction plug-in performs content analysis and extraction on the valid data in the message queue, and then pushes the extracted data to the load balancing strategy module;
[0011] 5) The load balancing strategy module classifies and balances the data generated by the knowledge base construction plug-in, and stores the classification and balancing results in the database.
[0012] Furthermore, the method for training the machine learning strategy module is:
[0013] 21) Annotate the real-time performance indicator data of each component in the server cluster; the annotated task types include: CPU type task, memory type, and disk type;
[0014] 22) Obtaining the load status information of the node server under different task types, as well as the response and waiting time under the corresponding load status, and normalizing the obtained information to construct a data set for training the machine learning strategy module;
[0015] 23) Using the data set to train the selected neural network model to obtain a machine learning strategy module, the training goal is to optimize the response and waiting time of each node server in the server cluster.
[0016] Furthermore, the load balancing strategy module classifies and balances the data generated by the knowledge base construction plug-in in the following way: first, the geographical scope of the data collected by the data access module is divided into grids, and the data generated by the knowledge base construction plug-in is split according to time and geographical location to obtain multiple data subsets; the geographical location and time of each data subset are concatenated and the hash value is taken as the hash value of the corresponding data subset, each storage node is numbered, and a modulo operation is performed according to the hash value of the data subset and the number of nodes to determine the storage node corresponding to each data subset, and each data subset is distributed to the corresponding database storage node for storage.
[0017] Furthermore, the real-time performance indicator data includes: the operating status of the database and the CPU usage, memory usage, disk usage, network delay, bandwidth, and packet loss rate of the node server.
[0018] The main contents of the knowledge base construction method of the present invention for massive real-time data scenarios include:
[0019] Real-time massive data access module: cleans the received data, analyzes valid data, and builds the required knowledge base;
[0020] Real-time collection technology of observable tools: It can collect the status of each component in real time, including but not limited to (Flink, Kafka, database and other software and the node servers where these software are deployed), and link with the real-time massive data access module to clean the data, reduce the pressure of each component, and improve the stability and reliability of the knowledge base construction system;
[0021] Machine learning cleaning strategy generation module: collects real-time performance indicator data of each node in the server cluster; inputs the performance indicator data into a pre-trained machine learning strategy module; and cleans the data in the message queue based on the output results of the machine learning strategy module.
[0022] Load balancing scheduler: Classify and balance the result data generated by the knowledge base construction plug-in, and store the classification and balancing results in the database.
[0023] Furthermore, the massive data access module mainly uses RDMA technology to improve the availability of Kafka, and enables a custom Kafka interceptor to intercept a batch of data.
[0024] Furthermore, observable tools include but are not limited to Prometheus, Grafana, and one or more of your own observable tools.
[0025] Furthermore, the observability tool observation targets include but are not limited to the running status of software such as Flink, Kafka, and database, as well as the CPU, memory, disk usage, network latency, bandwidth, packet loss rate and other parameters of the node servers where these software are deployed, the blocking status, backlog, back pressure, processing pressure and other parameters of the components, the total number of data items in the data status, the number of log items that meet the set characteristics, etc.
[0026] Furthermore, after the machine learning strategy module observable tool collects information from each component, it uses the machine learning strategy module to analyze data quality issues, thereby obtaining a cleaning interception strategy. For example, when the input data in the message queue suddenly increases, the observable tool detects this phenomenon, analyzes the real-time data, and concludes that the real-time data contains invalid attack data. It then extracts the features of the invalid data, loads them into the message queue interceptor, and intercepts the data, thereby reducing the pressure on the message queue.
[0027] Furthermore, the load balancing scheduler classifies the result data generated by the knowledge base construction plug-in and distributes the result data to the storage nodes, thereby achieving the purpose of load balancing.
[0028] Furthermore, the algorithm of the machine learning strategy module includes the following steps:
[0029] (1) Label the different task types in the real-time performance indicator data of each node; label the task type of each indicator data: CPU type task, memory type, disk type. For example, when the memory of a node reaches a certain threshold, the data is cleaned and suspected attack data is intercepted; when the disk storage reaches a certain threshold, stop writing data to the node; if all indicators do not reach the threshold, but the value obtained after normalization reaches the threshold, the node will also be controlled.
[0030] (2) Obtain the load status information of the computing node server under different task types, as well as the response and waiting time under these load conditions, and construct a data set for training the prediction model by normalizing various parameters (such as the server's CPU, memory, disk usage, network latency, bandwidth, packet loss rate, and other parameters, and the component's blocking status, backlog, back pressure, processing pressure, and other parameters);
[0031] (3) The selected neural network model is trained using the data set. The model output is the task allocation decision. The reward function is designed to minimize the response time and task waiting time of the server cluster. The training process simulates the system and collects data. After multiple iterations, the model gradually learns to optimize task allocation under different load conditions to achieve dual optimization of system response time and task waiting time. Let S = {s 1 ,s 2 ,...,s n} is n nodes in the server cluster. For each node s i , we define the following variables: x i =(x 1i ,x 2i ,...,x ki ): A vector of k performance indicators (such as CPU usage, memory usage, etc.).
[0032] Furthermore, dynamically adjusting the model training task's weights for different resource loads includes the following steps:
[0033] (1) Obtain the current task type T and the load of each node in the node set S, and use the model, where f is the trained machine learning strategy module and is the model parameter;
[0034] (2) Use the node corresponding to y to forward the task and perform calculations;
[0035] (3) For newly added task types, the incremental learning module is called to accumulate data, followed by incremental training. At this time, the forwarding nodes for this type of task are calculated according to the default weights.
[0036] Furthermore, the machine learning algorithm also includes the step of regularly retraining the machine learning strategy module.
[0037] (1) Load the structural parameters and weight files of the most recent training to initialize the model;
[0038] (2) Use the accumulated data to train the model;
[0039] (3) Resave the model structure and weight files.
[0040] Furthermore, the partition calculation method adopts spatiotemporal coding partition calculation, and the specific steps are as follows:
[0041] (1) Based on the spatiotemporal attributes of the data collected by the real-time massive data access module, a one-dimensional spatiotemporal coding sliding window partition calculation is constructed;
[0042] (2) Each node calculates the required indicator content and temporarily stores it in the cache;
[0043] (3) Pre-aggregate partition data, using time and geographic location as dimensions, calculate min, max, sum, and count, and record them in the metadata of the temporary data block;
[0044] (4) Aggregate the data blocks according to multiple time levels such as hours, days, and weeks, and multiple geographical levels such as geographical location codes of different digits to obtain the overall data situation.
[0045] Further, the one-dimensional space-time coding sliding window partition calculation process:
[0046] (1) The geographical scope of the data collected by the data access module is divided into grids, each grid representing a small geographical area. The geographical scope can be divided into multiple grids using coordinate information such as longitude and latitude. In this way, the statistical analysis problem of the geographical scope at the provincial and national levels can be transformed into the problem of statistical analysis of each grid separately.
[0047] (2) A dynamic adjustment method is used to select the grid to ensure that the window covers data within a specific geographical range. That is, the starting point and end point of the sliding window correspond to the upper left corner and lower right corner of the grid in the window respectively.
[0048] (3) When matching the currently collected data with the historical data in the sliding window, a spatial index structure is used to support both quadtree and R-tree modes to accelerate the data matching process. By building a spatial index structure, the grid data that falls within the window can be quickly located and filtered, reducing unnecessary calculations and traversal operations.
[0049] (4) As the sliding window slides, new data enters the window and old data leaves the window. In order to keep the data updated and maintained in real time, an incremental calculation method is used. Each time the window slides, only the newly added data and the removed data are statistically calculated, thereby reducing the amount of calculation and delay.
[0050] (5) The data split by region and time is hashed according to the geographical and time-split subsets, and then the database storage node corresponding to each data subset is determined based on the hash value, and the data is distributed to the corresponding database storage node for storage. The geographical location information and time of each subset are concatenated and the hash value is taken as the hash value of the corresponding subset. Each storage node is numbered, and a modulo operation is performed based on the hash value and the number of nodes to determine the storage node corresponding to each subset.
[0051] The advantages of the present invention are as follows:
[0052] 1) High-speed matching technology based on massive data cleaning cache: This technology uses RDMA technology to implement Kafka's high-speed cache and optimizes the three most densely populated data paths in the network (record production, record replication, and record consumption). The Kafka interceptor is combined with observable tools to achieve anomaly identification and efficient cleaning of massive data.
[0053] 2) Load balancing scheduler algorithm based on machine learning: Perform machine learning on some performance indicators to obtain load balancing and data partitioning strategies.
[0054] 3) Real-time data collection technology for observable tools: Use plug-ins to collect the running status of tools such as Kafka and the running status of the server, including but not limited to CPU, memory, disk usage, network latency, bandwidth, packet loss rate and other parameters, and record data status parameters. And link with the massive data cleaning cache technology to filter and clean abnormal data, effectively reduce the pressure on Kafka components, and improve the stability and reliability of knowledge base construction.
[0055] 4) Machine learning cleaning strategy generation module: Through machine learning, the data in the message queue is analyzed, a cleaning strategy is generated, and it is linked with the message queue to reduce the pressure on the Kafka component. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 Build flow charts for knowledge bases of massive real-time data scenarios.
[0057] Figure 2 This is a flowchart of the machine learning algorithm.
[0058] Figure 3 Schematic diagram of space-time coding partitioning. DETAILED DESCRIPTION
[0059] The present invention is further described in detail below in conjunction with the accompanying drawings. The examples given are only used to explain the present invention but not to limit the scope of the present invention.
[0060] Figure 1 A flow chart of the knowledge base construction technology for massive real-time data scenarios provided by the present invention is given, and the flow chart is as follows:
[0061] s1: High-speed cache matching technology for cleaning massive data: optimize the message queue to achieve high-speed cache of the message queue. Use the message queue interceptor to clean and filter abnormal data in massive data, thereby reducing the pressure on the message queue.
[0062] s2: The machine learning strategy module is used to generate cleaning interception strategies: Figure 2 As shown, a cleaning strategy generation algorithm based on machine learning is characterized in that it includes the following steps:
[0063] (1) Collect real-time performance indicator data of each node in the server cluster;
[0064] (2) inputting the performance indicator data into a pre-trained machine learning strategy module;
[0065] (3) The machine learning strategy module generates a cleaning interception strategy based on the input performance indicator data and sends it to the message queue interceptor.
[0066] s201 performance indicator data includes CPU usage, memory usage, network throughput, disk I / O, etc.
[0067] s202: the method of obtaining a machine learning strategy module by pre-training according to the performance indicator data comprises the following steps:
[0068] (1) Label different task types in the input data;
[0069] (2) Obtain the load status information of the computing node server under different task types, as well as the response and waiting time under these load conditions, and construct a data set for training the prediction model by normalizing various parameters;
[0070] (3) Use the selected model to train using the dataset, with the goal of optimizing system response and task waiting time. Let S = {s 1 ,s 2 ,...,s n} is n nodes in the server cluster. For each node s i , we define the following variables: x i =(x 1i ,x 2i ,...,x ki): A vector of k performance indicators (such as CPU usage, memory usage, etc.); the selected model is the decision tree model CART (Classification and Regression Trees); the task data is input into the model after standardization, and the information gain is used as the loss function for training. Through multiple trainings, the loss is reduced to a certain extent, and a suitable prediction model is obtained. Specific training process:
[0071] Initialization: Initialize the root node of the decision tree.
[0072] Feature selection: Select the best segmentation features and segmentation points so that the segmented sub-nodes have the largest information gain.
[0073] Recursive segmentation: recursively perform feature selection and segmentation on each child node until the stopping condition is met (the depth of the tree reaches the maximum value).
[0074] Pruning: Prune the generated decision tree to prevent overfitting.
[0075] Finally, the decision trees of each task are integrated and optimized according to the input task type.
[0076] s203 dynamically adjusts the task's load weights for different resources, including the following steps:
[0077] (1) Obtain the current task type T and the load S of each node, and use the model y = f(S, T, θ), where f is the trained machine learning strategy module and θ is the model parameter;
[0078] (2) Use the node corresponding to y to forward the task and perform calculations;
[0079] (3) For newly added task types, the incremental learning module is called to accumulate data, followed by incremental training. At this time, the forwarding nodes for this type of task are calculated according to the default weights.
[0080] s204 is a step of periodically retraining the machine learning strategy module.
[0081] (1) Load the structural parameters and weight files of the most recent training to initialize the model;
[0082] (2) Use the accumulated data to train the model;
[0083] (3) Resave the model structure and weight files.
[0084] s3: Real-time data collection technology for observable tools: Use plug-ins to collect the running status of tools such as Kafka and the running status of the server, including but not limited to CPU, memory, disk usage, network latency, bandwidth, packet loss rate and other parameters, and record data status parameters. Observable tools include but are not limited to Prometheus, Grafana (monitoring instrument system), and one or more of the self-built observation tools. Linked with the massive data cleaning cache technology, filter and clean abnormal data, effectively reduce the pressure on Kafka components, and improve the stability and reliability of knowledge base construction.
[0085] s4: Knowledge base construction plug-in: Implements the database construction method in the form of a plug-in. Implements according to different business logics and decouples from the entire system. Accesses valid data in the message queue, performs content analysis and extraction, and then pushes it to the load balancing strategy module.
[0086] s5: Load balancing strategy: The load balancing strategy includes partitioning the data generated by the knowledge base construction plug-in. The partitioning strategy mainly includes the following two methods:
[0087] Spatiotemporal coding partitioning computation
[0088] (1) Based on the spatiotemporal attributes of the collected data, a one-dimensional spatiotemporal coding sliding window partition calculation is constructed;
[0089] (2) Each node calculates the required indicator content and temporarily stores it in the cache;
[0090] (3) Pre-aggregate partition data, using time and geographic location as dimensions, calculate min, max, sum, and count, and record them in the metadata of the temporary data block;
[0091] (4) Aggregate the data blocks according to multiple time levels such as hours, days, and weeks, and multiple geographical levels such as geographical location codes of different digits to obtain the overall data situation.
[0092] like Figure 3 As shown, the one-dimensional space-time coding sliding window partition calculation process steps include:
[0093] (1) Divide the geographic scope into grids, each grid representing a small geographic area. The geographic scope can be divided into multiple grids using coordinate information such as longitude and latitude. In this way, the statistical analysis problem of the geographic scope at the provincial and national levels can be transformed into the problem of statistical analysis of each grid separately.
[0094] (2) A dynamic adjustment method is used to select the grid to ensure that the window covers data within a specific geographical range. The starting point and end point of the sliding window correspond to the upper left corner and lower right corner of the grid in the window respectively.
[0095] (3) When matching data within the sliding window, a spatial index structure is used to support both quadtree and R-tree modes to accelerate the data matching process. By building a spatial index structure, the grid data that falls within the window can be quickly located and filtered, reducing unnecessary calculations and traversal operations.
[0096] (4) As the sliding window slides, new data enters the window and old data leaves the window. In order to keep the data updated and maintained in real time, an incremental calculation method is used. Each time the window slides, only the newly added data and the removed data are statistically calculated, thereby reducing the amount of calculation and delay.
[0097] (5) Divide the data into multiple subsets and assign them to multiple database storage nodes for storage.
[0098] Although the specific embodiments of the present invention are disclosed for the purpose of illustration, the purpose is to help understand the content of the present invention and implement it accordingly, those skilled in the art will understand that various substitutions, changes and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the best embodiment, and the scope of the present invention is subject to the scope defined in the claims.
Claims
1. A method for constructing a knowledge base for massive real-time data scenarios, comprising the following steps: 1) The real-time massive data access module collects the real-time performance indicator data of each component in the server cluster and sends it to the message queue interceptor; The components include a database and a node server where it is located; 2) The message queue interceptor cleans and filters the real-time performance indicator data of each component to be processed according to the current cleaning interception strategy, obtains the valid data of each component and adds it to the message queue; 3) using observable tools to collect the status of valid data in the message queue and input it into the machine learning strategy module, generating a cleaning interception strategy and sending the message queue interceptor; 4) The knowledge base construction plug-in performs content analysis and extraction on the valid data in the message queue, and then pushes the extracted data to the load balancing strategy module; 5) The load balancing strategy module classifies and balances the data generated by the knowledge base construction plug-in, and stores the classification and balancing results in the database.
2. The method according to claim 1, characterized in that The method for training the machine learning strategy module is: 21) Annotate the real-time performance indicator data of each component in the server cluster; the annotated task types include: CPU type task, memory type, and disk type; 22) Obtaining the load status information of the node server under different task types, as well as the response and waiting time under the corresponding load status, and normalizing the obtained information to construct a data set for training the machine learning strategy module; 23) Using the data set to train the selected neural network model to obtain a machine learning strategy module, the training goal is to optimize the response and waiting time of each node server in the server cluster.
3. The method according to claim 1, characterized in that The method for the load balancing strategy module to classify and balance the data generated by the knowledge base construction plug-in is as follows: first, the geographical range of the data collected by the data access module is divided into grids, and the data generated by the knowledge base construction plug-in is split according to time and geographical location to obtain multiple data subsets; the geographical location and time of each data subset are spliced together and the hash value is taken as the hash value of the corresponding data subset, each storage node is numbered, and a modulo operation is performed according to the hash value of the data subset and the number of nodes to determine the storage node corresponding to each data subset, and the data subsets are distributed to the corresponding database storage nodes for storage.
4. The method according to claim 1, characterized in that: The real-time performance indicator data includes: the operating status of the database and the CPU usage, memory usage, disk usage, network delay, bandwidth, and packet loss rate of the node server.