Data set version management method and device based on computing power of intelligent computing center
By deploying database and object storage services in the intelligent computing center and creating data set version management tables, the problem of the intelligent computing center lacking data set version management is solved, data set version control and rollback are realized, and the efficiency and flexibility of large-scale model training is improved.
Patent Information
- Application Number
- CN202510058821.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
The intelligent computing center lacks dataset version management capabilities, which makes it difficult to achieve when it is necessary to roll back to the historical dataset version, reproduce the model weights of previous versions, or compare changes in different versions of datasets.
By deploying databases, object storage services and dataset operation environments in the Intelligent Computing Center, creating dataset tables and version management tables, recording the version number of the dataset, and performing version management when the dataset is updated, the version control and rollback of the dataset is realized.
It improves the effective utilization rate of computing power resources in the intelligent computing center, avoids repeated calculations and waste of computing power resources, enhances the flexibility and efficiency of data processing, supports flexible training and fine-tuning of large models, and improves training efficiency and accuracy.
Smart Images

Figure CN119988351A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the technical field of computing power infrastructure, and in particular, to a data set version management method and device based on the computing power of an intelligent computing center. Background Art
[0002] With the development of artificial intelligence technology and computing power technology, the concept of intelligent computing center has emerged. "Intelligent computing center" refers to the use of large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, mainly for artificial intelligence applications (such as artificial intelligence deep learning model development, model training and model reasoning and other scenarios) to provide the required computing power, data and algorithms. Intelligent computing center covers facilities, hardware, software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.
[0003] In today's artificial intelligence field, large language models (LLM) and multimodal large models are increasingly becoming core forces. These models not only perform well in general tasks such as natural language processing and image understanding, but also show great potential in applications in specific fields. As the infrastructure supporting the training and reasoning of these large models, the Intelligent Computing Center provides powerful computing power and storage resources, making it possible to process large-scale data sets and fine-tune models.
[0004] Although the current large language models (LLM) and multimodal large models are performing better and better in general fields, there is still significant room for improvement in specific tasks in vertical fields. In order to enhance the capabilities of large models in these vertical fields, relevant text training sets or image-text pair training sets are usually prepared, and supervised fine-tuning (SFT) is performed on the large models.
[0005] The preparation of the dataset and the fine-tuning of the large model is a dynamic process that involves constantly adding, updating, or deleting content in the dataset. At different time points, the evaluation indicators of the dataset used for fine-tuning may also be different. As the number of experiments increases, the large model obtained by fine-tuning the current dataset may not be optimal, so sometimes it is necessary to roll back to a previous version of the dataset for re-fine-tuning. In addition, it may be necessary to reproduce the large model weight file obtained by fine-tuning a previous version of the dataset, or compare the changes in a sample in different versions of the dataset.
[0006] At present, the computing power of intelligent computing centers can usually only access all data contents in the current state of the data set. After adding, updating or deleting sample data, the new data set formed also only contains data in the current state. This makes it difficult to roll back to a data set in a certain historical state, reproduce a large model obtained by fine-tuning a previous version, or compare sample changes in different versions of data sets, facing the problem of lack of data set version management capabilities. Summary of the invention
[0007] The embodiments of the present invention provide a method and device for data set version management based on the computing power of an intelligent computing center to solve the problem that the current intelligent computing center lacks data set version management capabilities.
[0008] In order to solve the above-mentioned technical problems, the present invention is achieved as follows:
[0009] In a first aspect, an embodiment of the present application provides a data set version management method based on the computing power of an intelligent computing center, the method comprising:
[0010] Step S1: deploying a database, an object storage service, and an operating environment for processing data sets in an intelligent computing center, and connecting the database, the object storage service, and the operating environment;
[0011] Step S2: creating a dataset table in the database, wherein the dataset table includes the following fields: a dataset name field, and a file field for indicating the path of the parquet files constituting the dataset;
[0012] Step S3: Obtain a training data set for training a large multimodal model and / or a large language model;
[0013] Step S4: dividing the training data set into multiple blocks, and saving the multiple blocks as multiple parquet files;
[0014] Step S5: storing the multiple parquet files in the object storage service, saving the paths of the multiple parquet files to the file field, saving the name of the training dataset to the name field, creating a dataset version management table based on the dataset table, and recording the current version number of the training dataset in the dataset version management table;
[0015] Step S6: modifying the data in the training data set to obtain a modified training data set, re-dividing the modified training set into multiple blocks, saving the multiple blocks as multiple new parquet files, storing the multiple new parquet files in the object storage service, and determining the paths where the multiple new parquet files are located;
[0016] Step S7: using the paths of the multiple new parquet files, updating the file field in the data set table;
[0017] Step S8: Execute a submission operation to add a new record in the dataset version management table, wherein the new record is a new version number of the modified training dataset.
[0018] Optionally, after step S8, the method further includes:
[0019] Step S9: Execute a database version rollback operation, roll back to a specified version number of the training data set, and call out multiple parquet files stored in the object storage service corresponding to the specified version number.
[0020] Optionally, after step S8, the method further includes:
[0021] Step S10: Check multiple parquet files stored in the object storage service that correspond to all version numbers of the training data set.
[0022] Optionally, after step S9 or step S10, the method further includes:
[0023] Step S11: obtaining the data content of the training data set corresponding to the specified version number, or determining a target version number among all the version numbers, and obtaining the data content of the training data set corresponding to the target version number;
[0024] Step S12: Based on the data content and the computing resources of the intelligent computing center, the multimodal large model and / or the large language model is trained.
[0025] Optionally, step S6 includes:
[0026] Step S61: receiving a user instruction for the data set table, wherein the instruction is used to select a file field;
[0027] Step S62: selecting file fields according to the instruction, and merging the selected file fields into a file field list;
[0028] Step S63: Read the parquet file indicated by the file field list based on the Pandas tool to form a DataFrame file;
[0029] Step S64: modify the data in the DataFrame file, re-save the modified DataFrame file as multiple new parquet files, and determine the paths where the multiple new parquet files are located.
[0030] In a second aspect, an embodiment of the present application provides a data set version management device based on the computing power of an intelligent computing center, the device comprising:
[0031] A deployment module, used to deploy a database, an object storage service, and an operating environment for processing data sets in an intelligent computing center, and connect the database, the object storage service, and the operating environment;
[0032] An execution module, configured to create a data set table in the database, wherein the data set table includes the following fields: a data set name field, and a file field for indicating a path where a parquet file constituting the data set is located;
[0033] Obtaining a training dataset for training a large multimodal model and / or a large language model;
[0034] Divide the training data set into multiple blocks, and save the multiple blocks into multiple parquet files;
[0035] Store the multiple parquet files in the object storage service, save the paths of the multiple parquet files to the file field, save the name of the training dataset to the name field, create a dataset version management table based on the dataset table, and record the current version number of the training dataset in the dataset version management table;
[0036] Modify the data in the training data set to obtain a modified training data set, re-divide the modified training set into multiple blocks, save the multiple blocks as multiple new parquet files, store the multiple new parquet files in the object storage service, and determine the paths where the multiple new parquet files are located;
[0037] Using the paths of the multiple new parquet files, update the file field in the dataset table;
[0038] A submission operation is performed to add a new record in the dataset version management table, wherein the new record is a new version number of the modified training dataset.
[0039] Optionally, the execution module is further used to, after executing a commit operation to add a new record in the dataset version management table, wherein the new record is a new version number of the modified training dataset, perform a database version rollback operation to roll back to a specified version number of the training dataset, and call out multiple parquet files corresponding to the specified version number and stored in the object storage service.
[0040] In a third aspect, an embodiment of the present invention provides an electronic device, comprising: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of a method for data set version management based on the computing power of an intelligent computing center as described in the first aspect.
[0041] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by the processor, the steps of a data set version management method based on the computing power of an intelligent computing center as described in the first aspect are implemented.
[0042] In a fifth aspect, an embodiment of the present invention provides a computer program product, comprising computer instructions, which, when executed by the processor, implement the steps of a data set version management method based on the computing power of an intelligent computing center as described in the first aspect.
[0043] The method provided in the embodiment of the present application can improve the effective utilization rate of the computing resources of the intelligent computing center by increasing the dataset version management capability of the intelligent computing center, avoid repeated calculations and waste of computing resources, and the combination of computing resources and dataset version management capabilities can improve the flexibility and efficiency of data processing, so that in the scenario of using the computing resources of the intelligent computing center to train or fine-tune large models, the required datasets can be flexibly called, effectively supporting the training of large models and improving the efficiency and accuracy of large model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0045] Figure 1A flowchart of a data set version management method based on the computing power of an intelligent computing center provided by an embodiment of the present invention;
[0046] Figure 2 A flowchart of a data set version management method based on the computing power of an intelligent computing center provided in an embodiment of the present application;
[0047] Figure 3 A structural block diagram of a data set version management method based on the computing power of an intelligent computing center provided in an embodiment of the present application;
[0048] Figure 4 Schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0049] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0050] The technical terms involved in the present invention are briefly described below.
[0051] The "computing power" mentioned in the present invention is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to execute certain computing requirements. It is the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity. It mainly provides services to the society through computing power infrastructure.
[0052] The "computing power" (CP) described in the present invention is the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, super computing power and intelligent computing power. The commonly used unit of measurement is the number of floating point operations performed per second (FLOPS: Floating Point Operations Per Second, 1EFLOPS = 10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1EFLOPS is about the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream notebooks. The calculation formula is: CP = CP general + CP intelligent + CP super.
[0053] The "Network Power" (NP) described in the present invention is a manifestation of the data transmission capability of computing facilities, including comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, etc. The carrying capacity involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities. In the embodiment of the present invention, the carrying capacity uses the memory bandwidth.
[0054] The "Storage Power" (SP) described in the present invention is the comprehensive ability of a data center in terms of data storage capacity, performance, safety and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and server built-in storage devices. The commonly used unit of measurement for storage capacity is exabyte (EB, 1EB = 2^60bytes), and the commonly used unit of measurement for performance is the number of reads and writes per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB). The disaster recovery ratio is an important manifestation of safety and reliability.
[0055] The "computing power infrastructure" described in the present invention is a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity. It can realize centralized computing, storage, transmission and application of information, and presents characteristics such as diversity and ubiquity, intelligence and agility, security and reliability, and green and low-carbon.
[0056] The “computing power” mentioned in the present invention includes general computing power, intelligent computing power and super computing power.
[0057] The “general computing power” mentioned in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0058] The "intelligent computing power" described in the present invention is a computing platform for large-scale deployment of special chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit) for various innovative applications of artificial intelligence, such as natural language processing and machine vision.
[0059] The "super computing power" mentioned in the present invention is mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, and gene analysis.
[0060] The "intelligent computing center" described in the present invention refers to a facility that provides the required computing power, data and algorithms for artificial intelligence applications (such as artificial intelligence deep learning model development, model training and model reasoning scenarios) by using large-scale heterogeneous computing power resources, including general computing power (CPU: Central Processing Unit) and intelligent computing power (GPU: Graphics Processing Unit, FPGA: Field Programmable Gate Array, ASIC: Application Specific Integrated Circuit, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.
[0061] The "computing resources" mentioned in the present invention refer to the technologies and facilities with information calculation, transmission, storage and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPU (Central Processing Unit), GPU (Graphics Processing Unit), network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, as well as supporting and guarantee resources such as wind, fire, water and electricity.
[0062] The "big model" described in the present invention includes: a big language model and a multimodal big model.
[0063] The “large language model” mentioned in the present invention refers to a large language model (LLM), which is a language model with a large parameter scale. It is designed to understand and generate human language. It is trained with a large amount of text data and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.
[0064] The "Multimodal Large Models" mentioned in the present invention refer to models that combine multimodal information such as text, images, videos, audio, etc. for training, including but not limited to multimodal large language models.
[0065] The "parquet file" described in the present invention refers to: a columnar storage file, which is mainly used for big data processing and analysis, and stores text data, or image metadata (image path) and corresponding text data in an image-text pair.
[0066] The "object storage service" described in the present invention may be s3, which is easy to expand and can store and retrieve any amount of data. In the present invention, s3 will mainly store parquet and image data.
[0067] The "original training set data" or "training data set" described in the present invention is the initial data set, which can be in csv format, jsonl format, or data managed in webui mode. In the final data persistence stage, they will all be stored in parquet format. It should be noted that text, data, etc. are all stored in parquet files, and the pictures in the picture-text pair are stored in the object storage service. The path of the picture will be saved in the parquet file together with the text.
[0068] The “database” mentioned in the present invention, such as PostgreSQL, can be used to manage data set version information.
[0069] The "database interaction tool" described in the present invention, such as SQLAlchemy, is a Python SQL toolkit and object-relational mapping (ORM) library, which provides an efficient and flexible way to interact with relational databases.
[0070] The "database version management plug-in" described in the present invention, such as SQLAlchemy-Continuum, can facilitate version management of tables using this plug-in.
[0071] The “data processing tool” described in the present invention, such as pandas, is a python data processing tool, which is mainly used here to read the original sample data and store it in the parquet file format.
[0072] The "data frame Dataframe" mentioned in the present invention comes from pandas, is a data table stored in the memory, and is a connection variable connecting the original training set data and the data finally saved in S3.
[0073] The "data file replacement" described in the present invention uses file-level processing because the amount of fine-tuning data is not very large overall. The parquet files corresponding to different versions of data sets are completely replaced, which is different from the row-level MOR (merge on read) and CoW (copy on write) replacements of professional data tools. In the present invention, exemplarily, when the data file replacement method is used, the file value of the first version is [file_0_0.parquet, file_0_1.parquet, ..., file_0_10.parquet], and the file value of the second version is [file_1_0.parquet, file_1_1.parquet, ..., file_1_10.parquet].
[0074] The "database commit" described in the present invention uses SQLAlchemy to perform data modification in a python script. These operations will not be saved to the database until the commit operation is performed.
[0075] Figure 1 A data set version management method based on the computing power of an intelligent computing center provided in accordance with an embodiment of the present application is shown. Figure 1 As shown, the method includes:
[0076] Step S1: deploy a database, an object storage service, and an operating environment for processing data sets in the intelligent computing center, and connect the database, the object storage service, and the operating environment;
[0077] Step S2: Create a data set table in the database;
[0078] The dataset table includes the following fields: a dataset name field, and a file field indicating the path of the parquet files that make up the dataset;
[0079] Step S3: Obtain a training data set for training a large multimodal model and / or a large language model;
[0080] Step S4: Divide the training data set into multiple blocks, and save the multiple blocks as multiple parquet files;
[0081] Step S5: Store multiple parquet files in the object storage service, save the paths of the multiple parquet files to the file field, save the name of the training dataset to the name field, create a dataset version management table based on the dataset table, and record the current version number of the training dataset in the dataset version management table;
[0082] Step S6: modify the data in the training data set to obtain a modified training data set, re-divide the modified training set into multiple blocks, save the multiple blocks as multiple new parquet files, store the multiple new parquet files in the object storage service, and determine the paths where the multiple new parquet files are located;
[0083] Step S7: Use the paths of multiple new parquet files to update the file field in the dataset table;
[0084] Step S8: Execute a submission operation to add a new record in the dataset version management table, where the new record is the new version number of the modified training dataset.
[0085] In step S1, in the intelligent computing center, you first need to deploy a database, object storage service, and an operating environment for processing data sets. The database is used to store structured data, and the object storage service can be s3, which is easy to expand and can store and retrieve any amount of data. Figure 1 In the method shown, s3 will mainly store parquet and image data. The runtime environment is the infrastructure for data processing and model training. For example, the database can be PostgreSQL, and the runtime environment is installed with Python, SQLAlchemy, SQLAlchemy-Continuum, etc. Through this deployment, efficient data storage and management can be ensured, while providing the necessary computing resources and environment for subsequent data processing and model training.
[0086] In step S2, a dataset table needs to be created in the database. The table contains two main fields: the dataset name field and the file field. The name field is used to identify different datasets, while the file field stores the path information of the parquet file associated with the dataset.
[0087] In step S3, a training data set needs to be obtained. This step involves obtaining training data sets for training multimodal large models and / or large language models from various sources (such as data warehouses, data lakes, etc.). These data sets may contain data in various forms such as text, images, and audio. The training data set can also be called the original training set data, which is the initial data set. It can be in csv format, jsonl format, or data managed in webui mode. In the final data persistence stage, they will all be stored in parquet format. It should be noted that text, data, etc. are all stored in parquet files, and the images in the image-text pair are stored in the object storage service. The path of the image will be saved in the parquet file together with the text. The obtained training data set is the basis for subsequent model training, ensuring that the model can learn rich features and patterns, thereby improving its performance.
[0088] In step S4, the acquired training data set can be divided into multiple blocks, usually according to the size of the data set and the processing power. Each block is saved as a parquet file, which is an efficient columnar storage format suitable for large-scale data processing. By dividing the data set into blocks, the efficiency of data reading and processing can be improved, especially when processing large-scale data, memory and computing resources can be effectively utilized.
[0089] In step S5, multiple parquet files can be uploaded to the object storage service, and the file field can be updated in the dataset table to record the paths of these parquet files. At the same time, a dataset version management table is created to record the current version number. This process implements persistent storage and version management of datasets, allowing users to easily track changes and versions of datasets, ensuring data traceability and standardized management. It should be noted that the creation of a dataset version management table can be achieved with the help of database interaction tools, such as SQLAlchemy's database version management plug-in, such as SQLAlchemy-Continuum. Using this plug-in, it is convenient to manage the version of the table.
[0090] In step S6, necessary modifications (such as cleaning, enhancement, etc.) may be made to the original training data set to generate a new training data set. The new data set is then divided into blocks and saved as new parquet files, uploaded to the object storage service, and the paths of the multiple new parquet files are determined.
[0091] It should be noted that, in a possible implementation, step S6 includes: step S61: receiving user instructions for a data set table, wherein the instructions are used to select file fields; step S62: selecting file fields according to the instructions, and merging the selected file fields into a file field list; step S63: reading the parquet file indicated by the file field list based on the Pandas tool to form a DataFrame file; step S64: modifying the data in the DataFrame file, resaving the modified DataFrame file as multiple new parquet files, and determining the paths of the multiple new parquet files.
[0092] In step S7, the file field in the dataset table can be updated with the new parquet file path to reflect the latest file storage information. This ensures that the information in the dataset table is always up to date, and users can easily obtain the latest training dataset file location to avoid using outdated data.
[0093] In step S8, a commit operation can be performed to add a new record in the dataset version management table to record the new version number of the modified training dataset. This helps track the evolution history of the dataset. Through version management, users can go back to the previous version at any time, ensuring the flexibility and security of data use, and also providing a basis for the continuous improvement of the model.
[0094] In summary, by increasing the dataset version management capability of the intelligent computing center, the effective utilization rate of the computing resources of the intelligent computing center can be improved, repeated calculations and waste of computing resources can be avoided, and the combination of computing resources and dataset version management capabilities can improve the flexibility and efficiency of data processing, so that in the scenarios where the computing resources of the intelligent computing center are used for large-scale model training or fine-tuning, the required datasets can be flexibly called, effectively supporting the training of large models and improving the efficiency and accuracy of large-scale model training.
[0095] In a possible implementation, after step S8, the method further includes: step S9: executing a database version rollback operation, rolling back to a specified version number of the training data set, and calling out multiple parquet files stored in the object storage service corresponding to the specified version number.
[0096] It should be noted that users can roll back the training dataset to a specified version number, which means that users can select a previous version and restore it to the state of that version. This operation usually involves the following steps: find the specified version number, that is, the system queries the dataset version management table to find the record corresponding to the specified version number; obtain the path of the parquet file, that is, extract the corresponding parquet file path from the record and determine the file to be called out; perform a rollback operation on the database to restore the state of the current dataset to the state of the specified version. Therefore, if a problem is found in the new version of the data, the user can quickly roll back to the previous stable version, and during model training and experiments, the user can try different dataset versions and easily roll back to any version for retraining or analysis.
[0097] In a possible implementation, after step S8, the method further includes: step S10: checking multiple parquet files stored in the object storage service corresponding to all version numbers of the training data set.
[0098] It should be noted that in this step, the Intelligent Computing Center provides a function that allows users to view all version numbers of the training dataset and their corresponding parquet files. This allows users to clearly understand the changes in each version, facilitating data auditing and compliance checks; when necessary, users can select the most appropriate dataset for model training or data analysis based on the version information.
[0099] In one possible implementation, Figure 2 As shown, after step S9 or step S10, the method further includes:
[0100] Step S11: obtaining the data content of the training data set corresponding to the specified version number, or, among all version numbers, determining the target version number, and obtaining the data content of the training data set corresponding to the target version number;
[0101] Step S12: Based on the data content and the computing resources of the intelligent computing center, train the multimodal large model and / or the large language model.
[0102] In summary, the addition of dataset version management capabilities in the intelligent computing center has significantly improved the flexibility and controllability of dataset management. Users can not only easily roll back to previous versions to ensure data stability and reliability, but also fully understand the version history of the dataset. This provides a more flexible and reliable dataset foundation for large model training.
[0103] The method shown in the embodiment of the present invention mainly utilizes the computing power in intelligent computing, combined with the parquet file format, object storage service, database and SQLAlchemy-Continuum, etc. to realize the management of data set versions, so as to roll back to a specific version of data or compare the changes of a sample of different versions of data sets.
[0104] The following uses the image-text pairs required for fine-tuning a large multimodal model as an example to illustrate the method shown in the embodiment of the present application, and it should be noted that in the training of the large multimodal model, the original training set data includes image-text pairs and images. In the training of the large language model, there are only no images.
[0105] The specific steps are as follows:
[0106] 1. Build the dataset table and dataset version management table, which is divided into the following three steps: deploy the database PostgreSQL; prepare the operating environment, that is, install Python, SQLAlchemy, SQLAlchemy-Continuum, etc.; create the dataset table and dataset version management table.
[0107] And creating dataset tables and dataset version management tables includes: using SQLAlchemy-Continuum to initialize version control; using SQLAlchemy to create the base class Base; defining the dataset model Datasets. This model inherits from the base class, including the auto-increment id, dataset name name, and all parquet file paths file; using SQLAlchemy to create a database engine. Connect to PostgreSQL; using SQLAlchemy to create dataset tables and dataset version management tables. Because the dataset model inherits from Base, using Base.metadata.create_all(engine) can automatically create dataset tables and dataset version management tables in the database.
[0108] 2. Prepare the data set according to the vertical field tasks. Upload the pictures to s3, and the picture information in the picture-text pair is the s3 path of the corresponding picture.
[0109] 3. Use pandas to read the above dataset and convert it into a dataframe.
[0110] 4. Block storage of parquets corresponding to the current dataset version. According to the block requirements, for example, each parquet can store up to 100k samples. If the dataset is large, block storage is required and persisted as multiple parquets.
[0111] 5. Save the file path of the above parquets to file and the dataset name to name. Because the default version number of SQLAlchemy-Continuum starts from 1, the version number is 1.
[0112] 6. Update data and version number. It is divided into the following five steps: first use SQLAlchemy to read the id, name and file of the data set; then use pandas to read and merge the specific parquets corresponding to the file to form a dataframe; modify a sample in the dataframe. For example, further modify the text content of the image and text of a sample; save the modified dataframe as a new version of parquets by replacing the data file, and replace the path file of these new versions of parquets with the value of the previous file; execute the database commit. In this way, the table based on SQLAlchemy-Continuum will automatically add a new version to the table.
[0113] 7. Dataset rollback can be divided into the following three aspects: Use SQLAlchemy to read the dataset and initialize it as the variable dataset; Use dataset.revert_to(1) to roll back to the specified version, where 1 refers to the initial version; Use dataset.versions.all() to view all versions.
[0114] Therefore, by using SQLAlchemy-Continuum to manage the dataset table and using the table field file to track the parquets corresponding to different versions of the dataset, the version management of the dataset can be achieved.
[0115] Figure 3 A data set version management device based on the computing power of an intelligent computing center according to an embodiment of the present application is shown. Figure 3 As shown, the device comprises:
[0116] A deployment module 301 is used to deploy a database, an object storage service, and an operating environment for processing data sets in an intelligent computing center, and connect the database, the object storage service, and the operating environment;
[0117] An execution module 302 is used to create a data set table in a database, wherein the data set table includes the following fields: a name field of the data set, and a file field for indicating the path of the parquet files constituting the data set;
[0118] Obtaining a training dataset for training a large multimodal model and / or a large language model;
[0119] Divide the training data set into multiple blocks and save the multiple blocks as multiple parquet files;
[0120] Store multiple parquet files in the object storage service, save the paths of the multiple parquet files to the file field, save the name of the training dataset to the name field, create a dataset version management table based on the dataset table, and record the current version number of the training dataset in the dataset version management table;
[0121] Modify the data in the training data set to obtain a modified training data set, divide the modified training set into multiple blocks again, save the multiple blocks as multiple new parquet files, store the multiple new parquet files in the object storage service, and determine the paths of the multiple new parquet files;
[0122] Update the file field in the dataset table with the paths of multiple new parquet files;
[0123] Execute the submit operation to add a new record in the dataset version management table. The new record is the new version number of the modified training dataset.
[0124] In one possible implementation, the execution module 302 is also used to perform a database version rollback operation to roll back to the specified version number of the training dataset after performing a commit operation to add a new record in the dataset version management table, where the new record is the new version number of the modified training dataset, and call out multiple parquet files stored in the object storage service corresponding to the specified version number.
[0125] In summary, by increasing the dataset version management capability of the intelligent computing center, the effective utilization rate of the computing resources of the intelligent computing center can be improved, repeated calculations and waste of computing resources can be avoided, and the combination of computing resources and dataset version management capabilities can improve the flexibility and efficiency of data processing, so that in the scenarios where the computing resources of the intelligent computing center are used for large-scale model training or fine-tuning, the required datasets can be flexibly called, effectively supporting the training of large models and improving the efficiency and accuracy of large-scale model training.
[0126] Please refer to Figure 4 An embodiment of the present invention further provides an electronic device 40, including a processor 401, a memory 402, and a computer program stored in the memory 402 and executable on the processor 401. When the computer program is executed by the processor 401, each process of the above-mentioned embodiment of the data set version management method based on the computing power of the intelligent computing center is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0127] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, each process of the above-mentioned data set version management method embodiment based on the computing power of an intelligent computing center is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0128] An embodiment of the present invention also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, the various processes of the above-mentioned data set version management method embodiment based on the computing power of the intelligent computing center are implemented, and the same technical effect can be achieved. To avoid repetition, they will not be repeated here.
[0129] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0130] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0131] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation modes, which are merely illustrative rather than restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are within the protection of the present invention.
Claims
1. A data set version management method based on the computing power of an intelligent computing center, characterized in that: The method comprises: Step S1: deploying a database, an object storage service, and an operating environment for processing data sets in an intelligent computing center, and connecting the database, the object storage service, and the operating environment; Step S2: creating a dataset table in the database, wherein the dataset table includes the following fields: a dataset name field, and a file field for indicating the path of the parquet files constituting the dataset; Step S3: Obtain a training data set for training a large multimodal model and / or a large language model; Step S4: dividing the training data set into multiple blocks, and saving the multiple blocks as multiple parquet files; Step S5: storing the multiple parquet files in the object storage service, saving the paths of the multiple parquet files to the file field, saving the name of the training dataset to the name field, creating a dataset version management table based on the dataset table, and recording the current version number of the training dataset in the dataset version management table; Step S6: modifying the data in the training data set to obtain a modified training data set, re-dividing the modified training set into multiple blocks, saving the multiple blocks as multiple new parquet files, storing the multiple new parquet files in the object storage service, and determining the paths where the multiple new parquet files are located; Step S7: using the paths of the multiple new parquet files, updating the file field in the data set table; Step S8: Execute a submission operation to add a new record in the dataset version management table, where the new record is the new version number of the modified training dataset.
2. The method according to claim 1, characterized in that After step S8, the method further includes: Step S9: Execute a database version rollback operation, roll back to a specified version number of the training data set, and call out multiple parquet files stored in the object storage service corresponding to the specified version number.
3. The method according to claim 1, characterized in that After step S8, the method further includes: Step S10: Check multiple parquet files stored in the object storage service that correspond to all version numbers of the training data set.
4. The method according to claim 2 or 3, characterized in that: After step S9 or step S10, the method further includes: Step S11: obtaining the data content of the training data set corresponding to the specified version number, or determining a target version number among all the version numbers, and obtaining the data content of the training data set corresponding to the target version number; Step S12: Based on the data content and the computing resources of the intelligent computing center, the multimodal large model and / or the large language model is trained.
5. The method according to claim 1, characterized in that The step S6 comprises: Step S61: receiving a user instruction for the data set table, wherein the instruction is used to select a file field; Step S62: selecting file fields according to the instruction, and merging the selected file fields into a file field list; Step S63: Read the parquet file indicated by the file field list based on the Pandas tool to form a DataFrame file; Step S64: modify the data in the DataFrame file, re-save the modified DataFrame file as multiple new parquet files, and determine the paths where the multiple new parquet files are located.
6. A data set version management device based on the computing power of an intelligent computing center, characterized in that: The device comprises: A deployment module, used to deploy a database, an object storage service, and an operating environment for processing data sets in an intelligent computing center, and connect the database, the object storage service, and the operating environment; An execution module, configured to create a data set table in the database, wherein the data set table includes the following fields: a data set name field, and a file field for indicating a path where a parquet file constituting the data set is located; Obtaining a training dataset for training a large multimodal model and / or a large language model; Divide the training data set into multiple blocks, and save the multiple blocks into multiple parquet files; Store the multiple parquet files in the object storage service, save the paths of the multiple parquet files to the file field, save the name of the training dataset to the name field, create a dataset version management table based on the dataset table, and record the current version number of the training dataset in the dataset version management table; Modify the data in the training data set to obtain a modified training data set, re-divide the modified training set into multiple blocks, save the multiple blocks as multiple new parquet files, store the multiple new parquet files in the object storage service, and determine the paths where the multiple new parquet files are located; Using the paths of the multiple new parquet files, update the file field in the dataset table; A submission operation is performed to add a new record in the dataset version management table, wherein the new record is a new version number of the modified training dataset.
7. The device according to claim 6, characterized in that The execution module is further used to, after executing a commit operation to add a new record in the dataset version management table, wherein the new record is a new version number of the modified training dataset, perform a database version rollback operation to roll back to a specified version number of the training dataset, and call out multiple parquet files corresponding to the specified version number and stored in the object storage service.
8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the data set version management method based on the computing power of an intelligent computing center as described in any one of claims 1 to 5.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the data set version management method based on the computing power of the intelligent computing center are implemented as described in any one of claims 1 to 5.
10. A computer program product, characterized in that It includes computer instructions, which, when executed by the processor, implement the steps of the data set version management method based on the computing power of the intelligent computing center as described in any one of claims 1-5.
Citation Information
Patent Citations
Data interaction method and device based on training system and object storage system
CN115858473A
Training data set version management method and system
CN118964335A
Generating datasets for scenario-based training and testing of machine learning systems
US20230394327A1