Real-time streaming data classification labeling and hierarchical storage management system based on deep learning
Through the real-time streaming data classification, annotation and hierarchical storage management system based on deep learning, the problem of difficulty in processing and analyzing real-time streaming data in a timely manner is solved, efficient data classification, annotation and storage is achieved, and corporate decision-making and production efficiency is improved.
Patent Information
- Application Number
- CN202510223233.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-06
AI Technical Summary
The existing real-time stream data classification annotation and hierarchical storage management system is difficult to process and analyze real-time stream data in a timely and efficient manner, affecting enterprise decision-making and production efficiency, and the classification and storage are inaccurate, making it inconvenient to find.
The real-time stream data classification, annotation and hierarchical storage management system based on deep learning is adopted, including data acquisition module, deep learning classification module, labeling module, hierarchical storage management module and user interaction interface. The real-time stream data is classified and annotated through the deep learning model, and stored in layers according to the data attributes.
It realizes timely classification and labeling and efficient storage of real-time streaming data, improves the accuracy of data analysis and decision-making efficiency, reduces losses caused by information lag, and improves the utilization rate of system resources.
Smart Images

Figure CN120105262A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data classification and management systems, and specifically to a real-time streaming data classification, labeling and hierarchical storage management system based on deep learning. Background Art
[0002] In today's digital age, with the rapid development of the Internet of Things, social media, financial transactions, industrial monitoring and other fields, real-time streaming data has shown explosive growth. This data is continuously generated in a high-speed manner and contains a large amount of valuable information, such as user behavior patterns, equipment operating status, market trends, etc.
[0003] Most of the existing real-time streaming data classification, labeling and hierarchical storage management systems find it difficult to process and analyze real-time streaming data in a timely and effective manner, which affects the effective decision-making of enterprises and organizations, and affects production efficiency and competitiveness. At the same time, some systems find it difficult to accurately classify and store data during classification and storage, and are not convenient for searching based on category and other information.
[0004] To this end, we proposed a real-time streaming data classification and annotation and hierarchical storage management system based on deep learning to solve the above problems. Summary of the invention
[0005] The purpose of the present invention is to provide a real-time streaming data classification, labeling and hierarchical storage management system based on deep learning to solve the problems raised in the above background technology.
[0006] To achieve the above-mentioned object, the present invention provides the following technical solutions: a real-time streaming data classification, annotation and hierarchical storage management system based on deep learning, comprising a data acquisition module, a data acquisition module, a deep learning classification module, an annotation module, a hierarchical storage management module and a user interaction interface; The data acquisition module is used to collect real-time streaming data from different data sources; The data preprocessing submodule preprocesses the collected stream data to improve the data quality; The deep learning classification module classifies the collected real-time streaming data using a pre-trained deep learning model; The labeling module labels the real-time stream data according to the classification result; The hierarchical storage management module hierarchically stores the annotated real-time streaming data in different storage media or storage areas according to the category, importance or other relevant attributes of the data; The user interaction interface is used to display classification results, annotation information and provide a user operation interface for users to view and manage data.
[0007] Preferably, a plurality of data source access submodules are provided for connecting different types of data sources, such as sensor networks, network log servers, social media platforms, etc.
[0008] Preferably, the data preprocessing submodule is used to perform preprocessing operations such as cleaning, denoising, and normalization on the collected original real-time stream data to improve the accuracy of subsequent classification and labeling.
[0009] Preferably, the deep learning model used by the deep learning classification module includes but is not limited to the following types: Convolutional neural networks are suitable for classifying data with spatial structure features such as images and videos; Recurrent neural networks and their variants, long short-term memory networks, and gated recurrent units, are suitable for classifying data with temporal structure characteristics, such as time series data and natural language text; Transformer architecture, suitable for processing large-scale text data and complex sequence relationship modeling; Model training function, which continuously updates and optimizes the parameters of deep learning models based on new annotated data to improve classification performance.
[0010] Preferably, the marking module includes: According to the classification results output by the deep learning classification module, add corresponding category labels to the real-time streaming data; For some complex or difficult to accurately classify data, manual review can also be combined with professionals to correct and confirm the preliminary labeling results.
[0011] Preferably, the hierarchical storage management module performs data hierarchical storage according to the following principles: Store critical data that is of high importance and frequently accessed in high-speed storage devices, such as solid-state drives (SSDs), to ensure fast reading, writing, and retrieval of data; For general data, it is stored in a large-capacity ordinary hard disk array; For data with a long history and low access probability, it can be archived and stored using low-cost storage media such as tape libraries; It has data migration function, which can dynamically adjust the data storage location according to factors such as data access frequency and storage time, so as to ensure the optimal balance between data storage efficiency and cost-effectiveness.
[0012] Preferably, the entire system meets the following performance requirements: Real-time: Real-time streaming data can be processed in a timely manner, and the total time delay from data collection to completion of classification, labeling and storage does not exceed the preset threshold, such as [X] milliseconds; Accuracy: The classification accuracy of the deep learning classification model should reach [X]% or above; Scalability: The system can easily add new data sources, expand storage capacity, and upgrade deep learning models to adapt to growing business needs.
[0013] The present invention provides a real-time streaming data classification, annotation and hierarchical storage management system based on deep learning, which has the following beneficial effects: Through the coordinated use of the data acquisition module and the data collection module, the purpose that can be achieved is to classify, label and analyze the data in a timely manner, analyze the market trends in real time, and users can make decisions quickly and effectively based on the timely information provided by the system, thereby reducing losses caused by information lags. At the same time, faults and hidden dangers can be discovered in a timely manner, maintenance plans can be arranged in advance, downtime can be reduced, production efficiency can be improved, and a faster response to market changes can be achieved, providing high-quality products and services, thereby increasing market competitiveness.
[0014] Through the coordinated use of the labeling module, the hierarchical storage management module and the deep learning classification module, the purpose of accurate and effective information classification and storage can be achieved. The data can be stored in storage devices at different levels according to factors such as the importance of the data and the frequency of use. A clear index structure can be established for the data, and the storage location of the data can be quickly located according to the classification label, shortening the data retrieval time, so that the required information can be quickly retrieved. At the same time, it is convenient to classify and store infrequently used data and frequently used data, reducing the situation where all data occupies high-performance storage resources and improving the utilization of system resources. BRIEF DESCRIPTION OF THE DRAWINGS In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0016] Figure 1 It is a schematic diagram of the module flow of the present invention.
[0017] In the figure: 1. Data acquisition module; 2. Data preprocessing module; 3. Deep learning classification module; 4. Labeling module; 5. Hierarchical storage management module; 6. User interaction interface. DETAILED DESCRIPTION The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0019] This real-time streaming data classification, annotation and hierarchical storage management system based on deep learning adopts a distributed architecture, including data acquisition module 1, deep learning classification module 3, annotation module 4, hierarchical storage management module 5 and user interaction interface 6. Multiple functional modules work together, and each module is deployed on different server nodes. Data transmission and interaction are carried out through network communication protocols. The specific architecture is as follows: Figure 1 As shown, The data acquisition module 1 is used to collect real-time streaming data from different data sources. The data acquisition module 1 is provided with multiple data source access submodules for connecting different types of data sources, such as sensor networks, network log servers, social media platforms, etc.; Specifically, for different types of data sources, corresponding data source access sub-modules are developed. For example, for sensor network data sources, by installing adaptive sensor drivers and communication protocol stacks, connections with various types of sensors are achieved, and data uploaded by the sensors are collected in real time; for network log server data sources, web crawler technology and log parsing tools are used to regularly capture log files on the server, parse the log contents, and extract useful information; for social media platform data sources, the open interfaces provided by the platform are called to obtain text, pictures, videos and other content posted by users on the platform.
[0020] The data preprocessing submodule preprocesses the collected stream data to improve the data quality. The data preprocessing submodule is used to perform preprocessing operations such as cleaning, denoising, and normalization on the collected original real-time stream data; Specifically, during the data collection process, the data preprocessing submodule is started to clean the raw data, for example, to remove outliers in sensor data, which may be caused by sensor failure or interference; for network log data, irrelevant information such as advertising links and static resource request records is filtered out, and only valid information related to the business is retained; Perform denoising on the cleaned data, use filtering algorithms to remove noise interference in the data, and improve the quality of the data; Normalization processing unifies data of different magnitudes into the same scale range for subsequent processing by deep learning models. For example, the temperature data collected by the sensor is normalized to the interval [0, 1]. The specific interval can be selected based on actual needs.
[0021] The deep learning classification module 3 uses a pre-trained deep learning model to classify the collected real-time streaming data; Specifically, the deep learning model used by the deep learning classification module 3 includes but is not limited to the following types, and a suitable deep learning model is selected according to the characteristics of the data: Convolutional neural networks are suitable for classifying data with spatial structure characteristics such as images and videos; Recurrent neural networks and their variants, long short-term memory networks, and gated recurrent units, are suitable for classifying data with temporal structure characteristics, such as time series data and natural language text; Transformer architecture, suitable for processing large-scale text data and complex sequence relationship modeling; Model training function, which continuously updates and optimizes the parameters of deep learning models based on new annotated data.
[0022] Collect a large amount of labeled historical data as a training data set, train the selected deep learning model, and deploy the trained deep learning model to the production environment. When the data acquisition module 1 collects new real-time streaming data and preprocesses it, the data is input into the deployed deep learning model for classification. The model outputs the category prediction results of the data and sends the results to the labeling module 4 for subsequent processing.
[0023] The labeling module 4 labels the real-time stream data according to the classification results. The labeling module 4 includes: According to the classification results output by the deep learning classification module 3, add corresponding category labels to the real-time streaming data; For some complex or difficult to accurately classify data, manual review can also be used.
[0024] Specifically, the labeling module 4 receives the classification result from the deep learning classification module 3, and adds a corresponding category label to each real-time stream data according to a predefined category label system; For some complex or difficult to accurately classify data, a manual review mechanism is set up, and the automatic annotation results and the original data content can be viewed on the user interaction interface 6. If the annotation is found to be incorrect or ambiguous, the annotation result can be corrected. The corrected annotation result will be fed back to the deep learning classification module 3 as new training data for further optimization of the model; Formulate hierarchical storage rules based on factors such as data category, importance, and access frequency. For example, key data related to core business can be divided into high-importance data layers; routine data generated by daily business can be divided into general-importance data layers; historical archived data can be divided into low-importance data layers; at the same time, considering the access frequency of data, frequently accessed data can be stored in high-speed storage devices, while data with a low access probability can be stored in large-capacity but relatively low-speed storage media; In order to ensure data storage efficiency and cost-effectiveness, the hierarchical storage management module 5 also has a data migration function, which regularly performs statistical analysis on data access. For data that has not been accessed for a long time, it is automatically migrated from the high-speed storage device to the large-capacity storage medium; and for data with an increased access frequency recently, it is migrated from the large-capacity storage medium back to the high-speed storage device.
[0025] The user interaction interface 6 is used to display classification results, annotation information and provide a user operation interface for users to view and manage data; Specifically, a visual display area is set on the user interaction interface 6 to intuitively display the classification statistics of real-time streaming data, the distribution of various types of data and other information in the form of charts and graphs.
[0026] Working principle: The data acquisition module 1 collects actual stream data from different data sources. The data preprocessing submodule preprocesses the collected data first, cleans, denoises, and normalizes the original real-time stream data. Then the deep learning classification module 3 uses the pre-trained deep learning model to classify the collected real-time stream data. The labeling module 4 labels the real-time stream data according to the classification results. Finally, the user interaction interface 6 displays the classification results, labeling information, and provides a user operation interface for users to view and manage data.
[0027] The above is the entire working principle of the present invention.
[0028] Finally, a few points should be explained: First, in the description of the present application, it should be noted that, unless otherwise specified and limited, the terms "installed", "connected" and "connected" should be understood in a broad sense, and can be mechanical or electrical connections, or internal connectivity between two components, or direct connections. "Up", "down", "left" and "right" are only used to indicate relative position relationships. When the absolute position of the described object changes, the relative position relationship may change; secondly, in the drawings of the embodiments disclosed in the present invention, only the structures involved in the embodiments disclosed in the present invention are involved. Other structures can refer to the general design. In the absence of conflict, the same embodiment and different embodiments of the present invention can be combined with each other; finally, the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A real-time streaming data classification, labeling and hierarchical storage management system based on deep learning, characterized by: It includes a data collection module (1), a data acquisition module (2), a deep learning classification module (3), a labeling module (4), a hierarchical storage management module (5) and a user interaction interface (6); The data acquisition module (1) is used to collect real-time streaming data from different data sources; The data preprocessing submodule preprocesses the collected stream data to improve the data quality; The deep learning classification module (3) uses a pre-trained deep learning model to classify the collected real-time stream data; The labeling module (4) labels the real-time stream data according to the classification result; The hierarchical storage management module (5) stores the annotated real-time stream data in different storage media or storage areas in a hierarchical manner according to the category, importance or other relevant attributes of the data; The user interaction interface (6) is used to display classification results, annotation information and provide a user operation interface for users to view and manage data.
2. The real-time streaming data classification, labeling and hierarchical storage management system based on deep learning according to claim 1 is characterized by: The data acquisition module (1) is provided with a plurality of data source access submodules for connecting to different types of data sources, such as sensor networks, network log servers, social media platforms, etc.
3. The real-time streaming data classification, labeling and hierarchical storage management system based on deep learning according to claim 1, characterized in that: The data preprocessing submodule is used to perform preprocessing operations such as cleaning, denoising, and normalization on the collected original real-time streaming data.
4. The real-time streaming data classification, labeling and hierarchical storage management system based on deep learning according to claim 1, characterized in that: The deep learning model used by the deep learning classification module (3) includes but is not limited to the following types: Convolutional neural networks are suitable for classifying data with spatial structure characteristics such as images and videos; Recurrent neural networks and their variants, long short-term memory networks, and gated recurrent units, are suitable for classifying data with temporal structure characteristics, such as time series data and natural language text; Transformer architecture, suitable for processing large-scale text data and complex sequence relationship modeling; Model training function, which continuously updates and optimizes the parameters of deep learning models based on new annotated data.
5. The real-time streaming data classification, labeling and hierarchical storage management system based on deep learning according to claim 1, characterized in that: The marking module (4) comprises: According to the classification results output by the deep learning classification module (3), add corresponding category labels to the real-time streaming data; For some complex or difficult to accurately classify data, manual review can also be used.
6. The real-time streaming data classification, labeling and hierarchical storage management system based on deep learning according to claim 1, characterized in that: The hierarchical storage management module (5) performs data hierarchical storage according to the following principles: Store critical data that is of high importance and frequently accessed in high-speed storage devices, such as solid-state drives (SSDs); For general data, it is stored in a large-capacity ordinary hard disk array; For data with a long history and low access probability, it can be archived and stored using low-cost storage media such as tape libraries; It has data migration function and can dynamically adjust the data storage location according to factors such as data access frequency and storage time.
7. The real-time streaming data classification, labeling and hierarchical storage management system based on deep learning according to any one of claims 1 to 6, characterized in that: The entire system meets the following performance requirements: Real-time: Real-time streaming data can be processed in a timely manner, and the total time delay from data collection to completion of classification, labeling and storage does not exceed the preset threshold, such as [X] milliseconds; Accuracy: The classification accuracy of the deep learning classification model should reach [X]% or above; Scalability: The system can easily add new data sources, expand storage capacity, and upgrade deep learning models.