Intermediate data management method and system suitable for machine learning algorithm integration system
By storing and managing intermediate data in machine learning algorithm integration system, the problem of poor intermediate data management in the prior art is solved, and more efficient data processing and system scalability are achieved.
Patent Information
- Application Number
- CN202510281558.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
The existing machine learning platform is not effective enough for intermediate data management, resulting in inaccurate data analysis results, affecting the quality of decision-making, and possibly causing waste of resources.
It provides an intermediate data management method suitable for machine learning algorithm integration system, by acquiring raw data, selecting data processing models and preprocessing algorithms, collecting and storing intermediate data, and storing and managing it with the original data and data processing model.
It improves the flexibility of data management, improves data processing efficiency and system scalability, and ensures the integrity and consistency of intermediate data.
Smart Images

Figure CN120216491A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data management, and particularly relates to an intermediate data management method and system applicable to a machine learning algorithm integration system. Background Art
[0002] In the current application practice of machine learning technology, the architecture and operation mode of existing platforms require users to have high professional programming capabilities and in-depth understanding of specific fields, which limits the popularity of the technology and the breadth of cross-disciplinary applications. At the same time, existing platforms lack effective management of intermediate data generated during the machine learning process, such as preprocessed data, intermediate products of feature engineering, and model output results, and are difficult to effectively utilize. And failures or inefficiencies in intermediate data management may not only lead to inaccurate data analysis results, affecting decision-making quality, but also cause resource waste.
[0003] Therefore, designing an efficient and reliable intermediate data management method is of crucial significance for ensuring the stability and prediction accuracy of machine learning systems, promptly responding to and handling data-related risks, and improving efficiency. Summary of the Invention
[0004] In view of the deficiencies in the related art, the present invention provides an intermediate data management method applicable to a machine learning algorithm integration system to store and manage intermediate data, optimize the data processing flow, and improve management efficiency.
[0005] The present invention provides an intermediate data management method applicable to a machine learning algorithm integration system, including the following steps:
[0006] Obtain original data;
[0007] Select a data processing model and a preprocessing algorithm, set the selected data processing model and / or preprocessing algorithm, and use the set data processing model and preprocessing algorithm to process the original data;
[0008] Collect and store intermediate data that appears during the processing of the original data;
[0009] Store and manage the processed data, intermediate data, data processing model, and preprocessing algorithm as a data packet.
[0010] This technical solution stores intermediate data together with the original data, the data processing model, and the preprocessing algorithm used during the processing of the original data, facilitating the management of intermediate data, improving the flexibility of data management, enhancing data processing efficiency and system scalability, and ensuring the integrity and consistency of intermediate data.
[0011] In some of these embodiments, the intermediate data in the data packet is indexed and associated through an intermediate data association tree, and the intermediate data association tree is a tree-shaped data structure.
[0012] In some of these embodiments, the root node of the intermediate data association tree points to the original data, the internal nodes point to the output data of the preprocessing algorithm or the data processing model, and the leaf nodes are the output data of the data processing model.
[0013] In some of these embodiments, each node of the intermediate data association tree represents a data entity, and the data entity is input data, model information and parameters, intermediate features, output results. The data entity node attributes include but are not limited to data type, data size, data source, timestamp.
[0014] In some of these embodiments, the similar intermediate data in different data packets is associated with each other through a prefix tree, and the prefix tree is a tree-shaped data structure.
[0015] In some of these embodiments, there is one and only one prefix tree, and the prefix tree is updated and adjusted as the data packets, data processing models, and preprocessing algorithms are added, deleted, and / or changed.
[0016] In some of these embodiments, the root node of the prefix tree points to the root directory of the data storage jointly formed by multiple data packets, the internal nodes point to the data processing model and the preprocessing algorithm, and the leaf nodes point to the output data after the data processed by the parent node of the node.
[0017] In some of these embodiments, the data processing model is a well-trained model, and the data processing model includes but is not limited to image recognition, speech recognition, and natural language processing models, and there is more than one data processing model.
[0018] In some of these embodiments, the preprocessing algorithm includes but is not limited to interpolation processing, feature selection, and dimensionless processing.
[0019] In addition, the present invention also provides an intermediate data management system, including:
[0020] A logic layer that pre-stores data processing models and preprocessing algorithms;
[0021] A data input layer for obtaining original data and setting the data processing models and preprocessing algorithms in the logic layer;
[0022] A data management layer for processing the original data using the data processing models and preprocessing algorithms and collecting the intermediate data generated during the original data processing process;
[0023] Data storage layer. The data storage layer is used to store data packets, and the data packets at least include original data, a data processing model and a preprocessing algorithm for processing the original data, and intermediate data generated during the processing of the original data.
[0024] Based on the above technical solution, in the embodiment of the present invention, the intermediate data management method applicable to the machine learning algorithm integration system stores the intermediate data together with the original data, the data processing model and the preprocessing algorithm used during the processing of the original data, so as to facilitate the management of the intermediate data, improve the flexibility of data management, enhance the data processing efficiency and the scalability of the system, and also ensure the integrity and consistency of the intermediate data. Brief Description of the Drawings
[0025] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0026] Figure 1 It is a flowchart of an embodiment of the intermediate data management method applicable to the machine learning algorithm integration system of the present invention;
[0027] Figure 2 It is the intermediate data association tree structure in an embodiment of the intermediate data management method applicable to the machine learning algorithm integration system of the present invention;
[0028] Figure 3 It is the prefix tree structure in an embodiment of the intermediate data management method applicable to the machine learning algorithm integration system of the present invention;
[0029] Figure 4 It is the system architecture of an embodiment of the intermediate data management method applicable to the machine learning algorithm integration system of the present invention;
[0030] Figure 5 It is the data entity design of the intermediate data association tree in an embodiment of the intermediate data management method applicable to the machine learning algorithm integration system of the present invention. Detailed Embodiments
[0031] Next, the technical solutions in the embodiments will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present invention.
[0032] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "center", "lateral", "longitudinal", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings. These are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention.
[0033] The terms "first", "second", "third" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", "third" may explicitly or implicitly include one or more of such features.
[0034] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "installed", "connected", "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0035] As Figure 1 shown, in a schematic embodiment of the intermediate data management method applicable to the machine learning algorithm integration system of the present invention, the intermediate data management method applicable to the machine learning algorithm integration system includes the following steps:
[0036] Obtain the original data;
[0037] Select a data processing model and a preprocessing algorithm, set the selected data processing model and / or preprocessing algorithm, and use the set data processing model and preprocessing algorithm to process the original data;
[0038] Collect and store the intermediate data that appears during the processing of the original data;
[0039] Store and manage the processed data, intermediate data, data processing model, and preprocessing algorithm as a data packet.
[0040] By storing the intermediate data together with the original data, the data processing model, and the preprocessing algorithm used during the processing of the original data, the above intermediate data management method facilitates the management of the intermediate data, improves the flexibility of data management, enhances the data processing efficiency and the scalability of the system, and also ensures the integrity and consistency of the intermediate data.
[0041] The present invention also provides an intermediate data management system, which includes a logic layer, a data input layer, a data processing layer, and a data storage layer.
[0042] The logic layer pre-stores data processing models and preprocessing algorithms. The data processing models are models perfected through machine learning training. The data processing models include, but are not limited to, image recognition, speech recognition, and natural language processing models, and there may be more than one data processing model. During the processing of raw data, the data processing models and preprocessing algorithms used are not fixed and may change during data processing; the preprocessing algorithms include, but are not limited to, interpolation processing, feature selection, and dimensionless processing.
[0043] It should be noted that the logic layer includes multiple data processing models and multiple preprocessing algorithms, so that the intermediate data management system can process different raw data.
[0044] It should also be noted that if the data processing models stored in the logic layer cannot meet the data processing requirements, new data processing models can be trained through machine learning based on the raw data. This belongs to the prior art in this field and will not be elaborated here.
[0045] The data input layer is used to obtain raw data and set the data processing models and preprocessing algorithms in the logic layer.
[0046] In some embodiments, the data input layer includes a system operation interface. The user uploads a data file through the system operation interface, so that the data input layer obtains the data in the data file as raw data. It should be noted that the data file can be uploaded locally by the user or select the intermediate data of the intermediate data management system, that is, the intermediate data stored in the intermediate data management system can also be used as raw data for data processing.
[0047] The file types supported by the data input layer include, but are not limited to, common data formats such as CSV, MP3, and JPG.
[0048] In some other embodiments, the data input layer can also be referred to as the user interface layer. The user interface layer is also used to implement the interaction between the user and the system, so as to set the preprocessing algorithms and data processing models selected from the logic layer.
[0049] The data processing layer includes a data management engine. The data management engine not only integrates the functions of data generation, processing, and storage, but also provides a unified interface to manage the intermediate data in the machine learning algorithm integration system through a highly layered design. The functions of the data management engine include but are not limited to: automatically processing the original data according to the task processing flow designed by the user, using the selected preprocessing algorithms and data processing models; during the task execution, dynamically generating, collecting, and organizing the intermediate data to ensure the integrity and consistency of the data, and storing the intermediate data, original data, preprocessing algorithms, and data processing models used in the data processing process in a data packet; maintaining the data packet and providing operations for saving and deleting the data packet, allowing the user to store the data permanently or process it temporarily according to needs.
[0050] The data storage layer is used to store data packets. The data packet at least includes the original data, the data processing model and preprocessing algorithms for processing the original data, and the intermediate data generated during the processing of the original data.
[0051] As Figure 2 shown, an intermediate data association tree with a tree-shaped data structure is used to effectively index and associate the intermediate data in the data packet to support efficient data indexing and analysis.
[0052] The root node of the intermediate data association tree points to the original data, the internal nodes point to the output data of the preprocessing algorithms or data processing models, and the leaf nodes are the output data of the data processing models. It should be noted that the intermediate data association tree has a unique identification string that can uniquely point to a specific intermediate data association tree, and each node in the tree represents a data entity; the data entity can be input data, model information and parameters, intermediate features, output results, etc., and its node attributes include but are not limited to data type, data size, data source, timestamp, etc.
[0053] In some embodiments, each data packet can be represented by an intermediate data association tree. The root node of the intermediate data association tree points to the data file uploaded or selected by the user, the internal nodes point to the selected algorithms, models, or the output data of the algorithms and models, and the leaf nodes are the output data of the algorithms and models. The system can perform a depth-first traversal or breadth-first traversal on the intermediate data association tree to obtain the data packet.
[0054] Usually, each time the data processing layer processes the original data, at least one data packet will be generated. Therefore, there are usually multiple data packets stored in the data storage layer, and the data packets are searched and managed through the unique identification string in the database.
[0055] As Figure 3As shown, the same type of intermediate data in different data packets is associated with each other through a prefix tree, where the prefix tree is a tree-shaped data structure. The root node of the prefix tree points to the root directory of data storage, the internal nodes point to data processing models and preprocessing algorithms, and the leaf nodes point to the output data after the parent node of the node processes the data.
[0056] It should be noted that there is one and only one prefix tree, and the prefix tree is automatically updated and adjusted as the data packets, data processing models, and preprocessing algorithms are added, deleted, and / or changed.
[0057] In this embodiment, as Figure 4 shown, the intermediate data management system includes a user interface layer, a data processing layer, a logic layer, and a data storage layer; the user interface layer is responsible for the interaction between the user and the system; the data management engine of the data processing layer is responsible for passing data to the logic layer, processing the data generated by the logic layer, and returning to the user interface layer after obtaining or storing data from the data storage layer; the logic layer includes various preprocessing algorithms and machine learning models, processes business logic and sends data to the data management engine; the data storage layer is used to store intermediate data.
[0058] In some embodiments, the data is carried by a general data model, and the general data model is also stored in the intermediate data management system to carry the acquired raw data and intermediate data. The raw data and intermediate data need to be normalized before being stored in the general data model.
[0059] In this embodiment, a general data model is designed in the form of key-value pairs to accommodate structured, semi-structured, and unstructured data. Its class diagram structure is as Figure 5 shown, and the general data model includes the following components:
[0060] 1. Data model architecture definition, which includes: architecture version identifier ($schema), following the JSON Schema Draft 07 standard.
[0061] 2. Data model title (title), clearly indicating that the name of the model is "General Data Model".
[0062] 3. Data model description (description), detailing the purpose of the model, that is, a general model designed to store different types of data.
[0063] 4. Data model type (type), defined as object, indicating that the model is a collection of key-value pairs.
[0064] 5. Data attribute definition (properties), including:
[0065] The data unique identifier (dataId), as the unique identifier of the data item, is of the string type.
[0066] The data type (dataType) is used to identify the category of the data content, with values of "structured", "semi-structured", and "unstructured", defined in the form of an enumeration (enum).
[0067] The data content (dataContent) allows the model to adopt different structural representations according to different data types.
[0068] The metadata is used to store auxiliary description information of the data and is defined as an object type.
[0069] 6. The required attribute list (required) specifies the set of attributes that must be included in the data model, including the data / algorithm / model unique identifier, data type, data content, and metadata.
[0070] 7. The reusable model definitions (definitions) are used to define reusable components in the model, including:
[0071] The structured data model includes the table name (tableName) and row data (rows), where the row data is defined as an array of objects.
[0072] The semi-structured data model defines the JSON content (jsonContent) for storing data in JSON format.
[0073] The unstructured data model includes text data (textData) and binary data (binaryData), where the binary data is stored in string format and indicates the binary format (format: "binary").
[0074] 8. The metadata attribute definitions are further refined as:
[0075] The creator records the identity of the data creator and is of the string type.
[0076] The creation date records the time point when the data was created and is of the datetime string type.
[0077] Data source (dataSource), which describes the source information of the data and is of string type.
[0078] Data unique identifier (dataId), which serves as the unique identifier for the data item and is of string type.
[0079] Data destination (datatarget), which describes the destination of the data and is of string type.
[0080] Through the above definitions, the general data model realizes the unified representation and organization of different types of data and is applicable to various data storage and processing scenarios.
[0081] The following takes the wheel fault detection scenario as an example to introduce the intermediate data management method in detail. The data processing engine is the mysql database, and the mysql database also serves as the data storage layer. The intermediate data management method in the wheel fault detection scenario includes the following steps:
[0082] Step 1: The user uploads the wheel data file to be processed or selects the intermediate data managed by the system. The system will automatically identify the data file and organize the data in the data file into raw data. It should be noted that if the data does not exist in the cloud, the uploaded wheel data file will be normalized and temporarily stored in the general data model.
[0083] Step 2: The user visually selects the preprocessing algorithm and data processing model to be used on the system interface. The information of the selected preprocessing algorithm and data processing model will be automatically collected by the system, normalized, and temporarily stored in the general data model.
[0084] Step 3: After the user completes the modeling and clicks to save the model settings, the system will save the data generated in Steps 1 and 2 and store it in the mysql database.
[0085] Step 4: When the user clicks to start running, the system processes the wheel data file uploaded or selected by the user through the algorithm or model set by the user. The system will automatically capture the intermediate data generated by each algorithm or model, normalize it, and temporarily store it in the general data model.
[0086] Step 5: After the task is completed, when the user clicks to save the data packet, the data generated in Step 4 will be stored in the mysql database, and the system will analyze the data stored in the database and construct the data packet for this task. At the same time, the system will automatically update the prefix tree information for managing the same type of output data between multiple data packets.
[0087] In this embodiment, the intermediate data management system constructs the overall application through the Vue3 + Django framework. The combination of the Vue3 + Django framework provides a responsive and extensible user interface, allowing users to dynamically adjust the data processing flow.
[0088] In this embodiment, the system encapsulates the preprocessing algorithm and the wheel fault detection model into independent layers. Both the preprocessing algorithm and the wheel fault detection layer define standard input and output, as well as standard output, for data transmission between different algorithms and models and the creation of data packets.
[0089] In this embodiment, each algorithm and model has a unique identifier. When the algorithm and model are encapsulated into the processing engine, the system assigns a string with a unique identifier to them.
[0090] In this embodiment, the visualization for selecting the preprocessing algorithm or model technology uses the Vue3 plugin library
[0091] is implemented by @v3e / vue3-draggable-resizable.
[0092] In this embodiment, the algorithms and models integrated by the system have been improved and trained. The involved model parameters have been built into the data management engine, so they can be directly used without additional training.
[0093] In this embodiment, the system uses the pytorch_lightning framework to replace the pytorch framework commonly used in deep learning, retaining the flexibility of pytorch while providing better layering capabilities.
[0094] In this embodiment, the intermediate data managed by the system in step one is obtained by the data management engine traversing the prefix tree. By performing different conditional queries on the prefix tree, the intermediate data generated by different algorithms or models integrated by the system can be obtained, or the data files uploaded by users can also be used.
[0095] In this embodiment, the Python treelib library is used to build and manage the prefix tree. The prefix tree is automatically updated as the data packets, algorithms, and models are added and removed.
[0096] In this embodiment, step two for automatically collecting algorithm and model information is completed by the front-end component. It is implemented using JavaScript technology. By listening to the user's operation behaviors on the interface, especially the click events on the dropdown menu selection and the submit button, the real-time capture of the preprocessing algorithms and models selected by the user is achieved. This layer adopts an event listening and processing mechanism to accurately extract the user's selection and encapsulate it into a data object in JSON format. Subsequently, these data are asynchronously sent to the server side via AJAX technology for further processing and storage.
[0097] In this embodiment, step four automatically captures the intermediate data generated by each algorithm or model and uses the decorator of Python as a hook function to implement. After each algorithm or model is called by the system, its output result is normalized and stored in the general data model.
[0098] In this embodiment, step five for analyzing the database data is automatically completed by the data management engine. The data management engine automatically captures the metadata of all data models stored in the mysql database for this task, and uses the SQL lineage analysis tool SQLLineage developed in Python to perform data lineage analysis on it and generate corresponding data packets. The analysis process mainly analyzes the source and destination of the data in the metadata to construct the lineage relationship, and stores the analyzed lineage relationship in the directed graph (lineage relationship graph) created by the DiGraph class of the networkx library, where the nodes represent data entities and the edges represent the data flow direction.
[0099] In this embodiment, each data packet has a consistent data format in the database, and the data format is: data packet unique identifier string + data packet tree unique identifier string + event processing time.
[0100] In this embodiment, the data packet consists of the intermediate data association tree and its auxiliary information (such as task creation time, etc.), and it has a unique identifier. The unique identifier is automatically generated as a unique string by the system after the user finishes the operation in step four. The intermediate data association tree is generated according to the lineage relationship graph by calling the recursive algorithm and the treelib library of Python. The entire tree structure is stored in the database as a json object, and the system can search for this intermediate data association tree in the database according to the unique identifier.
[0101] In this embodiment, the data in the data packet includes the original input of the wheel data, the preprocessed data, the intermediate data generated by the running system, the selected algorithm and model information (including model parameters, data types, etc.), and the event processing time.
[0102] In this embodiment, the user can save or delete the data packet. While operating on the data association unit, the system will automatically analyze the metadata of each part of it and update the prefix tree information.
[0103] The above intermediate data management method and system applicable to the machine learning algorithm integration system not only improve the data processing efficiency and system scalability by providing a modular data management engine, but also ensure the integrity and consistency of the intermediate data. The user-friendly visual configuration interface and metadata maintenance mechanism improve the usability of the system and the flexibility of data management.
[0104] Finally, it should be noted that the embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0105] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that: modifications can still be made to the specific implementation manners of the present invention, or equivalent replacements can be made to some technical features; without departing from the spirit of the technical solutions of the present invention, they should all be covered within the scope of the technical solutions claimed by the present invention.
Claims
1. An intermediate data management method suitable for a machine learning algorithm integration system, characterized in that: The following steps are involved: Get the original data; Selecting a data processing model and a preprocessing algorithm, setting the selected data processing model and / or preprocessing algorithm, and using the set data processing model and preprocessing algorithm to process the original data; Collect and store intermediate data generated during the processing of raw data; The processed data, intermediate data, data processing models, and preprocessing algorithms are stored and managed as a data package.
2. The intermediate data management method applicable to the machine learning algorithm integration system according to claim 1, characterized in that: The intermediate data in the data packet is indexed and associated through an intermediate data association tree, and the intermediate data association tree is a tree data structure.
3. The intermediate data management method applicable to the machine learning algorithm integration system according to claim 2, characterized in that: The root node of the intermediate data association tree points to the original data, the internal nodes point to the output data of the preprocessing algorithm or the data processing model, and the leaf nodes are the output data of the data processing model.
4. The intermediate data management method applicable to the machine learning algorithm integration system according to claim 2, characterized in that: Each node of the intermediate data association tree represents a data entity, and the data entity is input data, model information and parameters, intermediate features, and output results. The data entity node attributes include but are not limited to data type, data size, data source, and timestamp.
5. The intermediate data management method applicable to the machine learning algorithm integration system according to claim 4, characterized in that: The same type of intermediate data in different data packets are associated with each other through a prefix tree, and the prefix tree is a tree-shaped data structure.
6. The intermediate data management method applicable to the machine learning algorithm integration system according to claim 5, characterized in that: There is only one prefix tree, and the prefix tree is updated and adjusted as the data packets, the data processing model, and the preprocessing algorithm are added or deleted and / or changed.
7. The intermediate data management method applicable to the machine learning algorithm integration system according to claim 5, characterized in that: The root node of the prefix tree points to the root directory of the data storage formed by the plurality of data packets, the internal nodes point to the data processing model and the preprocessing algorithm, and the leaf nodes point to the output data after the parent node of the node processes the data.
8. The intermediate data management method applicable to the machine learning algorithm integration system according to claim 1, characterized in that: The data processing model is a fully trained model, and the data processing model includes but is not limited to image recognition, speech recognition and natural language processing models, and the data processing model includes but is not limited to one.
9. The intermediate data management method applicable to the machine learning algorithm integration system according to claim 1, characterized in that: Preprocessing algorithms include but are not limited to interpolation processing, feature selection, and dimensionless processing.
10. An intermediate data management system, characterized in that: include: A logic layer, wherein the logic layer pre-stores a data processing model and a pre-processing algorithm; A data input layer, used to obtain raw data and set the data processing model and the preprocessing algorithm in the logic layer; The data management layer is used to process the raw data using the data processing model and preprocessing algorithm and to collect the intermediate data generated during the raw data processing; The data storage layer is used to store data packets, which include at least original data, a data processing model and a preprocessing algorithm for processing the original data, and intermediate data generated during the processing of the original data.