Data management method and device
By automatically generating and running target scripts to process sample data of the big model, and extracting and converting data based on user interaction, the problem of inefficient sample data management is solved, and efficient and accurate data management and processing is achieved.
Patent Information
- Application Number
- CN202412000183.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-12-30
AI Technical Summary
The sample data management efficiency of large models is low, resulting in extended data processing time, increased error rate and inconsistency in data.
By obtaining the training sample data uploaded by the user and its data type and format, the target script is automatically generated and run for data processing, and the target sample data is extracted from the data according to user interaction and converted into model training data.
It improves the efficiency and accuracy of data management, reduces labor costs, realizes the standardization and accuracy of data records, and ensures efficient management and processing of large-scale data sets.
Smart Images

Figure CN119917573A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the computer field, and more specifically, to a data management method and device. Background Art
[0002] With the widespread application of big models, companies from all walks of life have put forward the need to train their own big models, and these models will be updated according to specific needs or over time. Since the pre-training of big models requires a large amount of text data, companies usually obtain data sets from open source websites for training each time the model is updated. However, due to differences in the content and ratio of different data sets, the trained models will also differ in capabilities, and the processing methods of raw data also have their own advantages. Therefore, there is an urgent need for a management system specifically for pre-training data sets to more efficiently support model pre-training tasks. The common model pre-training process first obtains open source data sets or self-collected data. The raw data formats are diverse, such as json, txt, etc. It is necessary to extract the required text from the raw data by writing scripts and save them in new files. The processing plan is manually recorded for reference when the next big model is updated.
[0003] However, raw data usually comes from different sources, so its format and structure vary, and may include JSON, TXT and other types. Each format is processed differently, and usually requires a large number of custom scripts to be written for processing. However, manually rewriting these processing scripts every time the newly acquired data is processed is prone to omissions or improper processing, resulting in some important information in the data being lost or incorrectly cleaned, which ultimately affects the performance of the model. In addition, if the data set is large, the workload of manual processing is very huge, which further increases the risk of errors and may also lead to a significant increase in data processing time.
[0004] In addition, the diversity of raw data is a key factor in ensuring that large model pre-training tasks can cover a wide range of scenarios and complex situations. Each time a large model is pre-trained, the relevant data information will be recorded and retained manually for reference when the large model is updated later. However, as the data set is constantly updated and expanded over time, the workload of manual recording also increases. The updated data set may contain new categories, language variants, contextual information, or different data formats, which places higher demands on the accuracy of manual recording. As the amount of data increases and the complexity increases, the probability of errors in the manual recording process will also increase significantly. For example, the recorder may cause incorrect annotations or omissions due to fatigue, cognitive bias, or inaccurate understanding, which will affect the quality and effect of the pre-training task. In addition, there may be standardization issues in the manual recording process. Due to differences in understanding and recording methods among different personnel, this will lead to inconsistent data, which will affect the training effect of subsequent models.
[0005] Therefore, there is a technical problem in the related art that the sample data management efficiency of large models is low. Summary of the invention
[0006] The embodiments of the present application provide a data management method and device to at least solve the technical problem of low efficiency in managing sample data of large models existing in the related art.
[0007] According to one embodiment of the present application, a data management method is provided, including: obtaining training sample data uploaded by a user and the data type and data format of the training sample data; in response to a first interactive operation, generating and running a target script, processing the training sample data through the target script, wherein the target script is determined by a target key-value pair and a script template specified by the first interactive operation, and the script template is a script template corresponding to the data format; in response to a second interactive operation, extracting target sample data from the training sample data according to a target ratio, and converting the target sample data into model training data, wherein the target data type and the target ratio of the target sample data are specified by the second interactive operation, the data type includes the target data type, and the model training data represents a training set used to train a target large model.
[0008] In an exemplary embodiment, after obtaining the training sample data uploaded by the user and the data type and data format of the training sample data, the method further includes: storing the training sample data in a target node; generating a storage path in response to storing the training sample data in the target node, wherein the storage path is used to indicate the storage address of the training sample data on the target node; and storing the data identifier, the data type, the data format and the storage path corresponding to the training sample data in a database.
[0009] In an exemplary embodiment, storing the training sample data in a target node includes: obtaining the data volume of the training sample data; searching for the target node on a target server based on the data volume, wherein the target node represents an idle node whose storage space satisfies the data volume; and in response to finding the target node, storing the training sample data in the target node.
[0010] In an exemplary embodiment, in response to a first interactive operation, generating and running a target script, and processing the training sample data through the target script, includes: in response to the first interactive operation, determining the script template based on the data format; in response to the first interactive operation, determining the target key-value pair according to the key fields in the training sample data; using the target key-value pair to replace the target field in the script template to generate the target script; and running the target script to process the training sample data through the target script.
[0011] In an exemplary embodiment, different data formats correspond to different target scripts, and running the target script to process the training sample data through the target script includes at least one of the following: when a sampling operation needs to be performed on the training sample data, running the target script to perform a sampling operation on the training sample data through the target script; when a cleaning operation needs to be performed on the training sample data, running the target script to perform a cleaning operation on the training sample data through the target script; when a splicing operation needs to be performed on the training sample data, running the target script to perform a splicing operation on the training sample data through the target script.
[0012] In an exemplary embodiment, in response to the second interactive operation, target sample data is extracted from the training sample data according to a target ratio, and the target sample data is converted into model training data, including: in response to the second interactive operation, determining a data set identifier and the target ratio; extracting target sample data from the training sample data according to the data set identifier and the target ratio; and calling a format conversion module to perform format conversion on the target sample data to obtain the model training data.
[0013] In an exemplary embodiment, the method further includes: after obtaining the training sample data uploaded by the user and the data type and data format of the training sample data, obtaining the data volume of the training sample data; searching for the target node on the target server based on the data volume, wherein the target node represents an idle node whose storage space satisfies the data volume; in response to finding the target node, storing the training sample data in the target node; in response to storing the training sample data in the target node, generating a storage path, wherein the storage path is used to indicate the storage address of the training sample data on the target node; storing the data identifier, the data type, the data format and the storage path corresponding to the training sample data in a database; in response to the first interactive operation, determining the script template based on the data format, and generating a storage path according to the training sample data. The method comprises the steps of: determining the target key-value pair based on the key field in the data; using the target key-value pair to replace the target field in the script template, automatically generating the target script, and creating a data processing task; running the target script to execute the data processing task, and reporting the processing result of the data processing task to the database in real time for task status marking; in response to the second interactive operation, displaying the data identifier, the data type, the data format and the storage path corresponding to the training sample data with the task status marked as completed in the database; determining the data set identifier and the target ratio, wherein the data set identifier indicates an identifier of a data identifier that satisfies a preset identification condition; extracting target sample data from the training sample data according to the data set identifier and the target ratio; calling a format conversion module to perform format conversion on the target sample data to obtain the model training data.
[0014] According to another embodiment of the present application, a data management device is provided, including: an acquisition module, used to acquire training sample data uploaded by a user and the data type and data format of the training sample data; a first execution module, used to generate and run a target script in response to a first interactive operation, and process the training sample data through the target script, wherein the target script is determined by a target key-value pair and a script template specified by the first interactive operation, and the script template is a script template corresponding to the data format; a second execution module, used to extract target sample data from the training sample data according to a target ratio in response to a second interactive operation, and convert the target sample data into model training data, wherein the target data type and the target ratio of the target sample data are specified by the second interactive operation, the data type includes the target data type, and the model training data represents the training set used to train the target large model.
[0015] According to another embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.
[0016] According to another embodiment of the present application, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0017] Through this application, structured data management and automated data processing technology are adopted. The system first obtains the training sample data uploaded by the user and the related data type and format information. In response to the user's first interactive operation, the system automatically generates and runs the target script, which is determined by the target key-value pair specified by the user and the preset script template. The template corresponds to the data format, ensuring the standardization and automation of data processing. Further, in response to the user's second interactive operation, the system extracts the target sample data from the training sample data according to the target ratio specified by the user, and converts it into model training data, ensuring the flexibility and accuracy of data matching. Such a process not only improves the efficiency and accuracy of data management, but also maximizes the utilization of computing resources, reduces labor costs, and realizes the standardization and accuracy of data records, achieving the purpose of efficient management and processing of large-scale data sets, thereby achieving the technical effect of improving data processing efficiency and accuracy, and solving the technical problem of low management efficiency of sample data of large models existing in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a hardware structure block diagram of a server device of a data management method according to an embodiment of the present application;
[0019] Figure 2 is a flow chart of a data management method according to an embodiment of the present application;
[0020] Figure 3 is a schematic diagram of a hardware environment of a data management method according to an embodiment of the present application;
[0021] Figure 4 is a logical environment diagram of a data management method according to an embodiment of the present application;
[0022] Figure 5 is a schematic diagram of a data uploading process of a data management method according to an embodiment of the present application;
[0023] Figure 6 is a schematic diagram of a data processing flow of a data management method according to an embodiment of the present application;
[0024] Figure 7 is a schematic diagram of a data extraction process of a data management method according to an embodiment of the present application;
[0025] Figure 8 It is a structural block diagram of a data management device according to an embodiment of the present application. DETAILED DESCRIPTION
[0026] The embodiments of the present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0028] The method embodiments provided in the embodiments of the present application can be executed in a server device or a similar computing device. Taking running on a server device as an example, Figure 1 1 is a hardware structure block diagram of a server device of a data management method according to an embodiment of the present application. Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the server device may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above server device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.
[0029] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the data management method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the server device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0030] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the server device. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0031] In this embodiment, a data management method is provided. Figure 2 is a flow chart of a data management method according to an embodiment of the present application, such as Figure 2 As shown, the process includes the following steps:
[0032] Step S202, obtaining the training sample data uploaded by the user and the data type and data format of the training sample data;
[0033] Step S204, in response to the first interactive operation, generating and running a target script, and processing the training sample data through the target script, wherein the target script is determined by a target key-value pair and a script template specified by the first interactive operation, and the script template is a script template corresponding to the data format;
[0034] Step S206, in response to the second interactive operation, extract target sample data from the training sample data according to the target ratio, and convert the target sample data into model training data, wherein the target data type and target ratio of the target sample data are specified by the second interactive operation, the data type includes the target data type, and the model training data represents the training set used to train the target large model.
[0035] The execution subject of the above steps may be a server, a terminal, etc., but is not limited thereto.
[0036] The execution order of step S202 and step S204 is interchangeable, that is, step S204 may be executed first, and then step S202.
[0037] Optionally, in this embodiment, some nouns involved in this application are first explained:
[0038] TB: Terabyte, trillion bytes.
[0039] json: is an open standard file format and data exchange format.
[0040] txt: A text format.
[0041] key / value: data key-value pair, key is the value used to identify unique data, and value is the data associated with the key.
[0042] bin: In this application, it specifically refers to the input file format required for large model pre-training.
[0043] Optionally, in the embodiment of the present application, the above-mentioned "training sample data" may include but is not limited to the original data sets uploaded by users for training large models. These data sets usually contain a large amount of text information and are used to train and optimize machine learning models. These data sets may come from different fields, such as news articles, social media posts, academic papers, etc. They can be structured data, such as records in a database, or unstructured data, such as text files. For example, the training sample data can be a series of files in JSON format, which contain user comments and rating information, or it can be a file in TXT format, which contains a large amount of book text content.
[0044] Optionally, in the embodiment of the present application, the above-mentioned "data type and data format" may include but is not limited to the organization and storage format of the data, which is essential for determining how to correctly process and parse the data. Data types can be classified according to Wikipedia's standards, including art and culture, geography, humanities and social sciences, nature and natural sciences, engineering and applied sciences, etc. The data format refers to the physical structure of the data, such as JSON, TXT, etc. For example, if the data type is "Engineering and Applied Science", then the data format may be TXT, representing the text content of a series of technical documents; if it is "Humanities and Social Sciences", it may be JSON, containing survey results and statistical data from social science research.
[0045] Optionally, in an embodiment of the present application, the above-mentioned "target script" may include but is not limited to a series of predefined code templates, which are used to process data in a specific format and can be customized according to the key-value pairs specified by the user. The role of the target script is to extract, clean and transform data from the raw data to meet the needs of subsequent model training. For example, if the data format is JSON, the target script may contain code for parsing JSON objects and extracting specific fields; if it is in TXT format, the script may contain text cleaning and word segmentation logic. The generation and operation of the target script are based on the key-value pairs specified by the user through the first interactive operation. These key-value pairs tell the system what specific information needs to be extracted from the data.
[0046] Optionally, in an embodiment of the present application, the above-mentioned "model training data" may include but is not limited to data sets that are processed and converted for training the target large model. These data sets are carefully selected and proportioned to ensure that the model can learn diverse and balanced information. The model training data not only includes the target data type, but also reflects the diversity and complexity of the data. For example, if the target large model is a language model, the model training data may include text data in different languages to ensure that the model can understand and generate content in multiple languages; if it is an image recognition model, the model training data may include images of different categories and styles to improve the generalization ability of the model. Through the second interactive operation, the user can specify the required data type and proportion, and the system will extract and convert a data set suitable for model training from the training sample data accordingly.
[0047] It should be noted that the training sample data uploaded by users can come from a variety of different sources, including but not limited to open source data sets, internal enterprise data accumulation, web content captured by web crawlers, and social media data collected by users themselves. The types and formats of these data are also diverse, for example, they can be text data, image data, audio data, or structured data such as CSV files. This application does not limit this.
[0048] In addition, when responding to the first interactive operation, the target script generated and run can adopt different processing logic according to different data formats. For example, for text data, the script may include preprocessing steps such as word segmentation and stop word removal; for image data, it may include operations such as image resizing and normalization; for audio data, it may involve noise reduction, feature extraction and other processing. The designation of the target key-value pairs may also vary depending on the content of the data, such as keywords and themes may be focused on in text data, and objects and scenes may be focused on in image data. This application is not limited to this.
[0049] On the other hand, in the above-mentioned response to the second interactive operation, the process of extracting target sample data from the training sample data can be adjusted according to different model training requirements. For example, for training a multilingual translation model, it may be necessary to extract text data according to the proportion of different languages; for training an image recognition model, it may be necessary to extract image data according to the proportion of different categories. The process of converting target sample data into model training data can also be customized according to the specific needs of the model. For example, for deep learning models, it may be necessary to convert data into a specific tensor format; for traditional machine learning models, it may be necessary to convert data into feature vectors. This application is not limited to this.
[0050] For example, assume that you are building a data preprocessing system for training a large natural language processing (NLP) model, which is capable of processing large-scale text datasets and converting them into a format suitable for model training.
[0051] S1: The user uploads training sample data. The user uploads a 1TB text dataset through the system interface. These datasets contain files in multiple formats, such as JSON and TXT. The user specifies the data type as "Humanities and Social Sciences" and the data format as JSON and TXT.
[0052] S2: Generate a target script in response to the first interactive operation. The user performs the first interactive operation through the system interface and specifies the target key-value pair. For example, "title" and "content" are specified as key-value pairs in the JSON data, indicating that the title and content need to be extracted from the JSON file. The system automatically generates two target scripts based on the key-value pairs and data formats specified by the user: one is a script for processing JSON format data, and the other is a script for processing TXT format data. These script templates are pre-defined and can handle common data cleaning and format conversion tasks.
[0053] S3: Run the target script to process the training sample data. The system starts to execute the two target scripts. For data in JSON format, the script parses each JSON object, extracts the "title" and "content" fields, and saves this information to a new TXT file. For data in TXT format, the script performs text cleaning operations, such as removing special characters and extra spaces. The processed data is stored in the specified path of the server.
[0054] S4: Extracting target sample data in response to the second interactive operation. The user performs the second interactive operation and specifies to extract target sample data from the processed data in a ratio of 60% news articles and 40% social media posts. The system extracts 60% of the data from the news article dataset and 40% of the data from the social media post dataset according to the ratio specified by the user.
[0055] S5: Convert the target sample data into model training data. The system merges the extracted target sample data into a result file and converts the data into the format required for model training according to the requirements of NLP model training. For example, if the model requires the input data to be a sequence of fixed length, the system will perform word segmentation and padding operations on the text to ensure that the length of each sample is consistent.
[0056] Through the above process, the system realizes efficient management and processing of large-scale text data sets. First, through automated script generation and running, the workload of manual script writing and debugging is greatly reduced, and the speed and accuracy of data processing are improved. Secondly, through automated data extraction and format conversion, the system can quickly respond to user needs and generate data sets suitable for model training, which not only improves the efficiency of data processing, but also ensures the quality and consistency of the data set. Finally, through structured data management and optimization of distributed computing resources, the system can process large-scale data sets while maximizing resource utilization and reducing costs. These technical effects jointly improve the efficiency and effectiveness of large model pre-training, laying a solid foundation for building high-performance NLP models.
[0057] Through the embodiment of the present application, structured data management and automated data processing technology are adopted. The system first obtains the training sample data uploaded by the user and the related data type and format information. In response to the user's first interactive operation, the system automatically generates and runs the target script. These scripts are determined by the target key-value pairs specified by the user and the preset script template. The template corresponds to the data format, ensuring the standardization and automation of data processing. Further, in response to the user's second interactive operation, the system extracts the target sample data from the training sample data according to the target ratio specified by the user, and converts it into model training data, ensuring the flexibility and accuracy of data matching. Such a process not only improves the efficiency and accuracy of data management, but also maximizes the utilization of computing resources, reduces labor costs, and realizes the standardization and accuracy of data records, achieving the purpose of efficient management and processing of large-scale data sets, thereby achieving the technical effect of improving data processing efficiency and accuracy, and solving the technical problem of low management efficiency of sample data of large models existing in related technologies.
[0058] In an exemplary embodiment, after obtaining the training sample data uploaded by the user and the data type and data format of the training sample data, the method further includes: storing the training sample data in a target node; generating a storage path in response to storing the training sample data in the target node, wherein the storage path is used to indicate the storage address of the training sample data on the target node; and storing the data identifier, data type, data format and storage path corresponding to the training sample data in a database.
[0059] Optionally, in the embodiment of the present application, the above-mentioned "target node" may include but is not limited to one or more servers, cloud storage service nodes, nodes in a distributed file system, or any computing resource that can store data. These target nodes are responsible for receiving and storing uploaded training sample data. For example, the target node can be a server cluster within Inspur Electronic Information Industry Co., Ltd., or an S3 storage bucket in Amazon AWS cloud service, or a storage node in Google Cloud Platform.
[0060] Optionally, in an embodiment of the present application, the above-mentioned data identifier, data type, data format and storage path may constitute a sample label, which may include but is not limited to a set of metadata used to describe and identify the content and attributes of the training sample data. The data identifier in the sample label is part of the sample label and is used to uniquely identify each data set or data file. For example, the sample label may include information such as the name, source, collection date, version number of the data set, etc., and the data identifier may be a unique serial number or hash value for quickly retrieving and referencing a specific data set in a database.
[0061] Optionally, in the embodiment of the present application, the "storage path" may include, but is not limited to, a file system path, a URL, a record location in a database, or any identifier that can uniquely determine the data storage location. The storage path is used to indicate the storage address of the training sample data on the target node, so that the system can accurately access and retrieve the data. For example, the storage path can be an absolute path on a server, such as " / server / data / dataset1", or a relative path, such as "data / dataset1", or a URL in a cloud storage service, such as "s3: / / bucket-name / dataset1".
[0062] Optionally, in the embodiments of the present application, "database" may include but is not limited to a relational database, a non-relational database, a distributed database, or any system that can store and manage data. The database is used to store information such as data identifiers, data types, data formats, and storage paths corresponding to the training sample data. This information helps the system manage and retrieve data, as well as perform subsequent data preprocessing and model training tasks. For example, the database can be MySQL, MongoDB, or a Hadoop distributed file system, which can store structured or unstructured data and provide query and data management functions.
[0063] It should be noted that the process of storing the training sample data in the target node may involve a variety of different storage media and environments, which is not limited in this application. These storage media may include, but are not limited to, local hard disks, network attached storage (NAS), cloud storage services (such as Amazon S3, Google Cloud Storage), and distributed file systems (such as Hadoop HDFS). Each storage medium has its specific access speed, capacity, cost, and reliability characteristics, and is suitable for different application scenarios.
[0064] In addition, the storage path generated in response to the training sample data being stored in the target node can have a variety of different formats and structures, which are not limited in this application. The storage path can be a simple file system path, such as " / data / dataset1", or a complex path containing multiple subdirectories, such as " / data / 2024 / Q1 / dataset1". In addition, the storage path can also be a bucket and object identifier in a cloud storage service, such as "s3: / / my-bucket / dataset1", or a table and record ID in a database, such as "table:dataset1,record:12345".
[0065] On the other hand, the operation of storing the data identifier, data type, data format and storage path corresponding to the training sample data in the database can adopt a variety of database technologies and models, which are not limited in this application. These databases can be relational databases such as MySQL and PostgreSQL, which organize data through tables and relational models; they can also be non-relational databases such as MongoDB and Cassandra, which store data through documents or key-value pairs; they can also be time series databases such as InfluxDB, which are specifically used to process time series data. Each database technology has its own specific query language, optimization strategy and applicable scenarios. The most suitable database solution can be selected according to the characteristics and access patterns of the data.
[0066] In an exemplary embodiment, storing training sample data in a target node includes: obtaining the data volume of the training sample data; searching for a target node on a target server based on the data volume, wherein the target node represents an idle node whose storage space satisfies the data volume; and in response to finding the target node, storing the training sample data in the target node.
[0067] Optionally, in an embodiment of the present application, the above-mentioned "data volume" may include but is not limited to the total size of the training sample data, which can be measured in units such as bytes, kilobytes (KB), megabytes (MB), gigabytes (GB) or terabytes (TB). The size of the data volume directly affects the needs of data processing and storage, so it is necessary to accurately obtain it to determine the appropriate storage node. For example, if a data set contains a large number of high-definition images or video files, its data volume may reach several GB or even TB level; while a text data set may only be a few MB to a few hundred MB in size. Understanding the data volume helps the system evaluate the storage resources that need to be allocated and the expected processing time.
[0068] Optionally, in the embodiments of the present application, the "target server" may include but is not limited to a single physical server, a virtual server, a server cluster or a cloud service platform. The target server is a computing resource used by the system to store and process training sample data. For example, the target server can be a physical server configured with a high-performance CPU and a large amount of memory, which is specially used to process large-scale data sets; or it can be a cloud service platform, such as AWS or Azure, which provides elastic computing and storage resources and can dynamically adjust resource allocation according to the amount of data.
[0069] Optionally, in an embodiment of the present application, a "target node" may include, but is not limited to, a physical storage device with a specific storage capacity, a virtual storage space, or a cloud storage node. The target node is the location where the training sample data is actually stored, and there must be enough storage space to meet the needs of the data volume. For example, if the data volume is 500GB, the system may find a storage node with 1TB of free space to store the data; if the data volume is larger, such as reaching several TB, the system may need to find multiple nodes or a node with sufficient capacity to store the data. The selection of the target node will take into account factors such as storage space, read and write speed, data security, and cost-effectiveness.
[0070] It should be noted that the above step of obtaining the data volume of the training sample data may involve a variety of data measurement methods, which is not limited in this application. The data volume can be obtained by automatic statistics of the file size, or it can be estimated by the number of data record entries, or it can be calculated by the logical storage unit occupied by the data. For example, for text data, the data volume may be calculated based on the number of characters or words; for image data, it may be estimated based on the total number of pixels or file size; for time series data, it may be measured based on the number of time points. These measurement methods can be flexibly selected according to the characteristics of the data and the usage scenarios.
[0071] In addition, the above process of searching for the target node on the target server based on the amount of data may involve a variety of storage resource allocation strategies, which is not limited in this application. When searching for the target node, the system may consider multiple dimensions such as the node's storage capacity, read and write speeds, network latency, and energy consumption. For example, the system may give priority to nodes with larger storage capacity and faster read and write speeds to improve data processing efficiency; at the same time, it may also consider the node's energy consumption and cooling capacity to achieve the goal of green energy saving. In addition, the system may also select storage nodes of different performance levels based on the access frequency and importance of the data to optimize the overall storage cost and performance.
[0072] On the other hand, the above-mentioned operation of storing the training sample data to the target node in response to finding the target node may involve a variety of data transmission and storage technologies, which are not limited in this application. Data transmission can be synchronous or asynchronous, and can be directly transmitted through the network or through physical media such as hard disk copying. Storage technologies may include block storage, file storage, or object storage. For example, for data that requires frequent random access, block storage or file storage may be selected; and for large-scale unstructured data, object storage may be selected. In addition, data deduplication, compression and other technologies can be used to optimize the use of storage space during data storage. The selection of these technologies can be determined based on the characteristics of the data and business needs to achieve efficient data storage and management.
[0073] In an exemplary embodiment, in response to a first interactive operation, a target script is generated and run, and training sample data is processed through the target script, including: in response to the first interactive operation, determining a script template based on a data format; in response to the first interactive operation, determining a target key-value pair according to key fields in the training sample data; using the target key-value pair to replace the target field in the script template to generate a target script; and running the target script to process the training sample data through the target script.
[0074] In the embodiments of the present application, the above-mentioned "script template" may include but is not limited to a set of predefined code frameworks, which contain the basic logic and structure required to process specific data formats. Script templates are designed for different types of data formats, and they can include codes for common operations such as data parsing, data cleaning, and data conversion. For example, for data in JSON format, the script template may contain code for parsing JSON objects and arrays; for data in CSV format, the script template may contain code for reading and writing CSV files. Script templates can vary depending on the data format to meet different data processing requirements.
[0075] Optionally, in an embodiment of the present application, "target key-value pairs" may include, but are not limited to, key information pairs extracted from the training sample data, which are used to specify data fields that require special attention in the data processing script. The target key-value pairs are specified by the user according to data processing requirements. They can be column names and values in the database, keys and values in a JSON object, or any data identifiers and corresponding data that need to be extracted from the original data. For example, when processing data from an e-commerce platform, the target key-value pairs may include key fields such as "product ID" and "price", "user rating", etc.; when processing medical and health data, the target key-value pairs may include key information such as "patient ID" and "diagnosis results", "drug name", etc.
[0076] Optionally, in an embodiment of the present application, the "target field" may include, but is not limited to, variables or identifiers used as placeholders in the script template, and these placeholders will be replaced by target key-value pairs during the script generation process. The target field is pre-defined in the script template and is used to indicate the specific data content that the script needs to process when it is executed. For example, in a script template that processes text data, the target field may be "text content" or "author name"; in a script template that processes image data, the target field may be "image path" or "image label". Replacing the target field is a key step in generating the target script, which ensures that the script can perform the expected data processing operations for a specific data set.
[0077] Optionally, in an embodiment of the present application, "running the target script" may include but is not limited to the process of executing a generated script to process the training sample data. This process involves applying the logic defined in the script to the actual data to implement operations such as data cleaning, conversion, and extraction. For example, if the target script is to extract specific keywords from text data, then running the target script will involve scanning the text, identifying the keywords, and saving them to a new data structure; if the target script is to convert image data into the format required for model training, then running the target script will involve reading the image file, applying the image processing algorithm, and saving the processed image to a specified storage location. Running the target script is a core step in the data processing process, which directly affects the results of data preprocessing and the effect of subsequent model training.
[0078] It should be noted that, in response to the first interactive operation, the process of determining the script template based on the data format can select the corresponding script template according to the characteristics and processing requirements of different data formats, and this application does not limit this. The data format may include structured data such as CSV, JSON, XML, semi-structured data such as log files, or unstructured data such as text and images. For example, for data in CSV format, the script template may contain specific reading and parsing logic; for image data, the script template may contain image recognition and processing algorithms. In addition, the selection of the script template may also be based on the source, purpose and expected processing results of the data. For example, for social media data, a specific text analysis template may be required.
[0079] Furthermore, the above-mentioned process of determining the target key-value pairs according to the key fields in the training sample data in response to the first interactive operation can be determined according to the data content and business requirements, and this application does not limit this. Key fields may include, but are not limited to, the primary key, foreign key, index field of the data, or key attributes in the business logic. For example, when processing customer data, key fields may include "customer ID", "purchase history", and "credit score"; when processing sensor data, key fields may include "timestamp", "sensor ID", and "measurement value". The determination of the target key-value pairs can also take into account the privacy and security requirements of the data to ensure that sensitive information is properly handled.
[0080] In addition, the process of generating the target script by replacing the target field in the script template with the target key-value pair can be customized according to different data processing logics and business rules, which is not limited in this application. The replacement of the target field may involve operations such as data extraction, conversion, and aggregation. For example, when processing financial data, you may need to replace the "amount" and "date" fields to generate financial reports; when processing medical data, you may need to replace the "patient name" and "diagnosis result" fields to generate medical record summaries. The replacement of the target field can also be adjusted according to the timeliness and frequency of change of the data to adapt to the dynamically changing data environment.
[0081] On the other hand, the above-mentioned process of running the target script and processing the training sample data through the target script can be executed according to different data processing goals and performance requirements, which is not limited in this application. The operation of the target script may involve batch processing or real-time processing, and it may be necessary to consider the parallelization and optimization of data processing. For example, when processing large-scale text data, it may be necessary to use a distributed computing framework to speed up the processing process; when processing real-time transaction data, it may be necessary to ensure low latency and high throughput of data processing. The operation of the target script can also be adjusted according to the storage location and access mode of the data, such as executing in the cloud, local or edge computing environment to optimize performance and cost.
[0082] In an exemplary embodiment, different data formats correspond to different target scripts, and running the target script to process the training sample data through the target script includes at least one of the following: when a sampling operation needs to be performed on the training sample data, running the target script to perform a sampling operation on the training sample data through the target script; when a cleaning operation needs to be performed on the training sample data, running the target script to perform a cleaning operation on the training sample data through the target script; when a splicing operation needs to be performed on the training sample data, running the target script to perform a splicing operation on the training sample data through the target script.
[0083] Optionally, in the embodiment of the present application, the above-mentioned "sampling operation" may include but is not limited to the process of selecting a portion of data from the training sample data set as a sample, and this process may be based on different sampling techniques such as random sampling, stratified sampling, and systematic sampling. The purpose of the sampling operation is to reduce the size of the data set while maintaining the representativeness of the data set as much as possible, which is particularly important for training large-scale machine learning models. For example, when processing a large-scale image data set, the sampling operation may include randomly selecting a certain number of images as training samples; when processing financial transaction data, the sampling operation may include selecting transaction records within a specific time period as samples.
[0084] Optionally, in the embodiment of the present application, the above-mentioned "cleaning operation" may include but is not limited to preprocessing the training sample data to eliminate noise, fill missing values, remove outliers, etc., to improve the data quality. The cleaning operation is a key step in data preprocessing, which helps to improve the effect of subsequent model training. For example, when processing text data, the cleaning operation may include removing useless punctuation, stop word filtering, stemming, etc.; when processing sensor data, the cleaning operation may include filtering, outlier detection, and filling missing sensor readings.
[0085] Optionally, in an embodiment of the present application, the above-mentioned "stitching operation" may include but is not limited to the process of merging multiple training sample data sets into a large data set, which may be based on simple data merging, complex feature alignment, or data fusion based on specific business logic. The purpose of the stitching operation is to integrate data from different sources to enrich the features and information of the data set. For example, when processing data from multiple sensors, the stitching operation may include aligning and merging time series data from different sensors; when processing text data from different sources, the stitching operation may include merging the contents of different documents into a large text file for subsequent text mining and analysis.
[0086] It should be noted that, when it is necessary to perform sampling operations on the training sample data, the process of running the target script can be adjusted according to different sampling strategies and data characteristics, and this application does not limit this. The sampling operation may involve random sampling, stratified sampling, or sampling based on specific conditions. For example, when processing large-scale data sets, random sampling may be used to reduce the amount of data while maintaining data diversity; when it is necessary to maintain data balance, stratified sampling may be used to ensure the proportion of each category or group in the sample. In addition, the sampling operation can also be performed according to the time series characteristics, spatial distribution characteristics, or business logic characteristics of the data. For example, in time series analysis, sampling may need to be performed at time intervals.
[0087] In addition, when it is necessary to perform cleaning operations on the training sample data, the process of running the target script can be customized according to different data quality issues and business needs, and this application does not limit this. Cleaning operations may include removing duplicate records, correcting erroneous data, filling missing values, or smoothing noise. For example, when processing text data, cleaning may include removing irrelevant characters, standardizing text formats, or eliminating sensitive information; when processing numerical data, it may be necessary to identify and process outliers, fill missing values, or perform data standardization. Cleaning operations can also be performed based on the source, format, and intended use of the data. For example, when processing medical data, special attention may need to be paid to protecting patient privacy.
[0088] On the other hand, in the above-mentioned case where it is necessary to perform a splicing operation on the training sample data, the process of running the target script can be designed according to different data integration requirements and data structures, and this application does not limit this. The splicing operation may involve merging data sets from different sources, aligning data in different formats, or fusing data of different granularities. For example, when integrating data from different databases, the splicing operation may need to deal with the alignment of different data models and data structures; when processing multimodal data, it may be necessary to fuse text, image, and sound data together. The splicing operation can also be optimized based on the time sensitivity, spatial correlation, or business logic of the data. For example, when integrating data from different geographical locations, it may be necessary to consider the proximity and correlation of geographical locations.
[0089] In an exemplary embodiment, in response to a second interactive operation, target sample data is extracted from the training sample data according to a target ratio, and the target sample data is converted into model training data, including: in response to the second interactive operation, determining a data set identifier and a target ratio; extracting target sample data from the training sample data according to the data set identifier and the target ratio; calling a format conversion module to perform format conversion on the target sample data to obtain model training data.
[0090] Optionally, in an embodiment of the present application, the above-mentioned "dataset identification" may include but is not limited to a set of information used to uniquely identify a training sample data set, which may be the name, number, label, or other features of the data set that can distinguish different data sets. The purpose of the data set identification is to quickly locate and reference a specific data set in a complex data management environment. For example, in a data warehouse containing multiple sources and types, the data set identification can be a naming convention consisting of a date and a source channel, such as "2024_06_medical_imaging", which means medical imaging data collected in June 2024; it can also be an internally generated unique identifier, such as a UUID, which is used to index and retrieve data sets in a database.
[0091] Optionally, in an embodiment of the present application, the above-mentioned "target ratio" may include but is not limited to a sample extraction ratio specified according to specific requirements or algorithms. This ratio may be a fixed value or the result of a dynamic calculation. The target ratio is used to determine the number of samples extracted from the original data set to meet the needs of model training or other data processing tasks. For example, when performing model training, the target ratio may be the ratio of the training set determined based on the size of the validation set and the test set, such as 70% of the data for training, 15% for validation, and 15% for testing; when processing unbalanced data, the target ratio may be the oversampling or undersampling ratio of samples of a specific category, such as increasing the proportion of minority class samples to be equal to that of the majority class.
[0092] Optionally, in an embodiment of the present application, the above-mentioned "format conversion module" may include, but is not limited to, a set of functions for converting target sample data into a format suitable for model training. The format conversion module can process different types of data, such as text, images, audio, or video, and convert them into an input form that the model can understand. For example, for text data, the format conversion module may include steps such as word segmentation, part-of-speech tagging, and word embedding construction; for image data, it may include operations such as resizing, normalization, and color space conversion; for audio data, it may include sampling rate conversion, spectrogram extraction, and other processing. The design of the format conversion module is intended to improve the efficiency and effectiveness of model training and ensure that data is input into the model in the best manner.
[0093] It should be noted that the process of determining the data set identifier and the target ratio in response to the second interactive operation can be adjusted according to a variety of different user needs and business scenarios, and this application does not limit this. The data set identifier can be defined based on multiple dimensions such as data source, content theme or timestamp. For example, it can be a news data set of a specific news website, a collection of posts on social media within a specific time period, or a sensor data set collected in a specific region. The target ratio can be determined based on the size of the data set, the needs of model training, or the needs of data balance. For example, it can be to extract samples according to category balance, extract samples according to the proportion of data distribution, or extract samples according to the traditional proportions of 70%, 15%, and 15% of the model training set, validation set, and test set.
[0094] In addition, the above-mentioned operation of extracting target sample data from the training sample data according to the data set identifier and the target ratio can be customized according to the different characteristics and processing objectives of the data, and this application does not limit this. The extraction operation may involve random selection of data, stratified selection, or selection based on specific rules. For example, it can be stratified sampling based on the statistical characteristics of the data, time series sampling based on the time series characteristics of the data, or key sampling based on the importance of the data. In addition, the extraction operation can also be screened according to the quality and completeness of the data, for example, excluding data records with a large number of missing values, selecting data samples with rich information, or selecting according to the novelty and diversity of the data.
[0095] On the other hand, the process of calling the format conversion module to convert the target sample data into a format and obtaining the model training data can be adjusted according to the specific requirements of the model training and the original format of the data, and this application does not limit this. Format conversion may involve standardization, normalization or encoding conversion of data. For example, text data can be converted into word vectors, image data can be converted into pixel value arrays, and audio data can be converted into spectrograms. In addition, format conversion can also be customized according to the input requirements of the model. For example, for deep learning models, data may need to be converted into a specific tensor format; for traditional machine learning models, data may need to be converted into feature vectors. The format conversion module is designed to ensure that data can be input in a form that is most suitable for model training, thereby improving the efficiency and effectiveness of model training.
[0096] In an exemplary embodiment, the method further includes: after obtaining the training sample data uploaded by the user and the data type and data format of the training sample data, obtaining the data volume of the training sample data; searching for a target node on the target server based on the data volume, wherein the target node represents an idle node whose storage space satisfies the data volume; in response to finding the target node, storing the training sample data in the target node; in response to storing the training sample data in the target node, generating a storage path, wherein the storage path is used to indicate the storage address of the training sample data on the target node; storing the data identifier, data type, data format and storage path corresponding to the training sample data in a database; in response to the first interactive operation, determining a script template based on the data format, and generating a storage path according to the training sample data. The key fields in the data determine the target key-value pairs; the target key-value pairs are used to replace the target fields in the script template, the target script is automatically generated, and a data processing task is created; the target script is run to execute the data processing task, and the processing results of the data processing task are reported to the database in real time for task status marking; in response to the second interactive operation, the data identifier, data type, data format and storage path corresponding to the training sample data with the task status marked as completed in the database are displayed; the data set identifier and target ratio are determined, wherein the data set identifier indicates that the data identifier satisfies the preset identification condition; the target sample data is extracted from the training sample data according to the data set identifier and the target ratio; the format conversion module is called to perform format conversion on the target sample data to obtain model training data.
[0097] For example, assume that a large-scale e-commerce review dataset is being processed, which will be used to train a sentiment analysis model to identify sentiment tendencies in customer reviews.
[0098] S1: Users uploaded a dataset containing 10 million reviews, which are stored in JSON format. Each review contains a user ID, review text, and rating.
[0099] S2: The system obtains the data volume of the data set, which is 500 GB in size.
[0100] S3: Based on the amount of data, the system searches for a target node with sufficient free storage space on the target server. There are 5 nodes in the server cluster, and 3 of them have free space greater than 500GB.
[0101] S4: The system selects one of the idle nodes and stores the training sample data in the node. For example, the storage path is ` / server / datasets / ecommerce comments / `.
[0102] S5: The system generates a storage path and stores it in the database together with the data identifier, data type (text), and data format (JSON), recording detailed information of the data set.
[0103] S6: In response to the first interactive operation, the system determines a script template based on the JSON data format, and determines a target key-value pair according to key fields (user ID, comment text, and rating).
[0104] S7: The system replaces the target field in the script template with the target key-value pair, automatically generates the target script, and creates data processing tasks, such as cleaning data and extracting features.
[0105] S 8: The system runs the target script to execute the data processing task, and reports the processing result to the database in real time to mark the task status, for example, marking it as "processing completed".
[0106] S9: In response to the second interactive operation, the system displays the data identifier, data type, data format and storage path corresponding to the training sample data with the task status marked as completed in the database.
[0107] S10: The system determines the dataset identifier and target ratio. For example, reviews with a rating of 4 stars or more (including 5 stars) are selected as training data, and the target ratio is 80%.
[0108] S 11: The system extracts target sample data from the training sample data according to the data set identifier and the target ratio, that is, extracts 80% of the high-rated reviews.
[0109] S12: The system calls the format conversion module to convert the target sample data into a format, for example, converting comments in JSON format into the CSV format required for model training to obtain model training data.
[0110] Through the embodiments of the present application, by utilizing automated storage space management and data storage, the system can quickly locate and store large-scale data sets, ensuring the efficiency of the first step of data processing, namely data storage. Secondly, through the generation and operation of automated scripts, the system can perform customized data processing for specific formats and key fields, improving the flexibility and accuracy of data processing. Thirdly, through real-time task status updates and data extraction, the system can ensure that only high-quality data is used for model training, thereby improving the effectiveness and efficiency of model training. Finally, through the format conversion module, the system can convert data into a format suitable for model training, further optimizing the input data for model training, and providing a solid data foundation for building a high-performance sentiment analysis model.
[0111] The following is a further explanation of this application with reference to specific examples:
[0112] This application designs a large model pre-training data processing system to automatically record the data matching process, ensure the integrity and consistency of the data, and automatically identify and generate corresponding processing scripts according to different data formats (such as json, txt, etc.) to reduce manual intervention and improve processing efficiency.
[0113] The core principle of this application is structured data management, using databases to store raw data information; in addition, it can also optimize distributed computing resources, write programs to mount information of all current servers, and facilitate task distribution; on the other hand, this application can realize automated data processing, and reduce the workload of code development by pre-setting processing scripts and manually replacing key / value by users. Through multi-level table structure management and distributed computing design, combined with automated script generation and real-time task keep-alive mechanism, the efficiency of data processing and the maintainability of the system are guaranteed.
[0114] Figure 3 is a schematic diagram of the hardware environment of the data management method according to an embodiment of the present application, such as Figure 3 As shown, the various modules of this application are introduced as follows:
[0115] Data upload and download module: responsible for uploading the original data and associating the data tags specified by the user, and synchronously recording the upload path and tag information into the database. This design facilitates the rapid identification and management of uploaded data and improves data access efficiency.
[0116] Database module: Data storage table design. The newly added data storage table includes fields such as data set name, data source, data organization format, data type and data storage path. The data organization format field records the format of the original data (such as json, txt, etc.), and the data type field is classified according to encyclopedia standards, including [arts and culture, geography, humanities and social sciences, nature and natural sciences, engineering technology and applied sciences]. This design helps to quickly locate and organize data. Data task table: used to manage the life cycle of tasks, with fields including task information, task creation time, task end time, task update time and task status. The task information field records the task allocation nodes and their specific requirements in detail; the task status field indicates the current stage of the task, and the status values include [initialization, running, completion, abnormality]. Data cache table: records the decompressed high-frequency data, and the fields are set to the data set name, data type, original data compression package path and decompressed file path. This table improves data access speed and avoids repeated decompression.
[0117] Data processing module: Users can manually specify key / value pairs on the interface. The system automatically replaces the placeholders in the preset text processing code templates based on these specified key / value pairs, and generates processing scripts such as cleaning, sampling, and splicing for the current data set. This function simplifies the code writing work for data processing. The preset text processing code templates are common multi-process data processing codes without special algorithms. The number of code templates corresponds to the number of types of original data formats. For example, for the json format, there is a code template for processing json data, and for the txt format, there is a code template for processing txt data.
[0118] Data conversion module: The management system collects the status information of all nodes (the task status information indicates that the conversion is completed) to decide whether to convert the data into bin format. Before converting to bin, the user will be asked to enter the ratio of different types of data, and then all eligible data will be integrated and converted into the process data format. This operation is only executed when all tasks are completed and the status is normal, to ensure the integrity and accuracy of data processing.
[0119] Figure 4 is a logical environment diagram of a data management method according to an embodiment of the present application, such as Figure 4 As shown, including but not limited to the following logical implementations:
[0120] S 1: The data upload and download module is responsible for processing the original data set uploaded by the user. For example, if a user uploads a compressed package containing 1 TB of text data, the module will receive the data and store it on the specified path of the server. This step involves the reception and preliminary storage of data, laying the foundation for subsequent data processing.
[0121] S2: The data decompression module decompresses the uploaded data compression package. For example, if the uploaded data is a compressed package in ZIP format, the module will perform the decompression operation and restore the data to its original format, such as TXT or JSON file, to facilitate subsequent cleaning and processing.
[0122] S3: The data cleaning module cleans the decompressed data. For example, for text data, the module may remove irrelevant characters, correct misspellings, unify text formats, etc. to improve data quality and ensure data consistency and accuracy.
[0123] S4: The data processing module automatically replaces the placeholders in the preset text processing code template according to the key / value pairs specified by the user on the interface, and generates processing scripts such as cleaning, sampling, and splicing for the current data set. For example, if the user specifies that the "title" and "content" fields in the JSON data need to be extracted, the system will generate the corresponding script to extract this information from the JSON file.
[0124] S5: The data sampling module extracts samples from the cleaned data according to the sampling strategy specified by the user. For example, the user may need to randomly extract 10% of the samples from the data set for model training, and the module will perform the sampling operation according to this requirement.
[0125] S6: The data concatenation module concatenates the sampled data according to certain rules to form a data set suitable for model training. For example, if model training requires a continuous text sequence, this module will concatenate multiple text fragments into a long sequence.
[0126] S7: After all data processing tasks are completed, the bin conversion module converts the processed data into the bin format required for model training. For example, if the model training requires a fixed-size input, this module will convert the text data into a fixed-length binary format for easy model reading and processing.
[0127] S8: The database module records the data storage path, data type, data format, task status and other information throughout the process. For example, the database will record the name, source, storage location of each data set, as well as the creation time, end time and status of the data processing task, to facilitate data management and tracking.
[0128] S9: The nodes in the server layer are responsible for executing the tasks of each module in the above application layer. For example, node 1 may be responsible for the tasks of the data upload and download module, node 2 may be responsible for the tasks of the data decompression module, and so on. Each node performs corresponding tasks according to its computing resources and load conditions to achieve distributed processing and resource optimization.
[0129] Figure 5 is a schematic diagram of a data uploading process of a data management method according to an embodiment of the present application, such as Figure 5As shown in the figure, when uploading data, the user specifies the name of the current data set and the type of the data set. The management system automatically searches for available storage space based on the size of the data set. If the storage space is insufficient, the task is directly exited. If storage space is found, task information is generated and recorded in the task table of the database. According to the status of the current child node, the task is sent to any idle child node to complete the data upload task. Finally, the data storage path and the user-specified data set name, type and other information are recorded in the data storage table of the database.
[0130] Figure 6 is a schematic diagram of a data processing flow of a data management method according to an embodiment of the present application, such as Figure 6 As shown, data processing is performed on the uploaded data. Taking json format data as an example, the user specifies the required key / value, and then replaces the user input string into the template script for json data processing, and passes the path of the uploaded data to the script. After executing the script, the required data text can be obtained.
[0131] Figure 7 is a schematic diagram of a data extraction process of a data management method according to an embodiment of the present application, such as Figure 7 As shown in the figure, the processed data is matched and binned. The user selects the data set to be used and specifies the proportion of each type of data. The system extracts the corresponding amount of data from each data set according to the data selected by the user and the matching information, merges them into a result file, and then calls the binning module to process the data into the input file format required for large model pre-training.
[0132] Through the embodiments of the present application, the efficiency and accuracy of data management can be improved. Through structured data storage (such as data storage tables, task tables, and cache tables), the system can efficiently manage a large amount of data and its related meta-information. This approach ensures clear and standardized data management, facilitates rapid query and retrieval, and avoids data redundancy and confusion.
[0133] In addition, resource utilization can be maximized. Through the reasonable optimization of distributed data processing and computing resources (such as reusing GPU servers for efficient computing), the system avoids the waste of CPU resources and realizes efficient utilization of computing resources. Tasks are counted and allocated according to the original compressed package, which avoids idle computing nodes and improves processing efficiency.
[0134] On the other hand, it can reduce labor costs and improve work efficiency. Due to the high degree of automation of the data processing process (such as automated script generation, task status monitoring, and automated scheduling), the system can greatly reduce human intervention and reduce dependence on human resources. This not only improves work efficiency, but also reduces labor costs.
[0135] On the other hand, it can achieve standardization and accuracy of data records, and transfer the recording of information such as data ratio from manual management to system standardized processing, effectively reducing the probability of human errors and ensuring the accuracy and consistency of records.
[0136] That is, this application involves the management and use of large model pre-training data processing, including:
[0137] Efficient management of large-scale data sets: This application achieves efficient management of large-scale data and its related metadata through structured data storage, ensures that the data processing process is orderly and standardized, and optimizes data retrieval and access efficiency.
[0138] Fast data matching and automatic recording: This application supports fast data matching and automatically records each matching result into the database, avoiding omissions and errors that may occur during manual recording and ensuring the accuracy and traceability of the matching process.
[0139] Automated data format processing: For different data formats, this application classifies the data formats in advance. For data in formats such as JSON, the system quickly generates processing scripts by specifying key / value pairs, avoiding the time and workload required for manual script writing.
[0140] Maximize CPU resource utilization: This application can use only CPU resources and can be deployed in a GPU server cluster. When the GPU server is under high load, the system can make full use of idle CPU resources (for example, for decompressing the original data compression package), avoiding the waste of CPU resources.
[0141] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.
[0142] In the present embodiment, a data management device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0143] Figure 8 is a structural block diagram of a data management device according to an embodiment of the present application, such as Figure 8 As shown, the device comprises:
[0144] An acquisition module 802 is used to acquire the training sample data uploaded by the user and the data type and data format of the training sample data;
[0145] A first execution module 804 is used to generate and run a target script in response to the first interactive operation, and process the training sample data through the target script, wherein the target script is determined by a target key-value pair specified by the first interactive operation and a script template, and the script template is a script template corresponding to the data format;
[0146] The second execution module 806 is used to extract target sample data from the training sample data according to the target ratio in response to the second interactive operation, and convert the target sample data into model training data, wherein the target data type and target ratio of the target sample data are specified by the second interactive operation, the data type includes the target data type, and the model training data represents the training set used to train the target large model.
[0147] In an exemplary embodiment, the above-mentioned device is also used to: after obtaining the training sample data uploaded by the user and the data type and data format of the training sample data, store the training sample data in the target node; in response to the training sample data being stored in the target node, generate a storage path, wherein the storage path is used to indicate the storage address of the training sample data on the target node; store the data identifier, data type, data format and storage path corresponding to the training sample data in the database.
[0148] In an exemplary embodiment, the above-mentioned device is used to store the training sample data to the target node in the following manner: obtain the data volume of the training sample data; search for the target node on the target server based on the data volume, wherein the target node represents an idle node whose storage space satisfies the data volume; in response to finding the target node, store the training sample data to the target node.
[0149] In an exemplary embodiment, the above-mentioned device is used to generate and run a target script in response to a first interactive operation in the following manner, and process training sample data through the target script: in response to the first interactive operation, determine a script template based on the data format; in response to the first interactive operation, determine a target key-value pair according to key fields in the training sample data; use the target key-value pair to replace the target field in the script template to generate a target script; run the target script, and process the training sample data through the target script.
[0150] In an exemplary embodiment, different data formats correspond to different target scripts, and the above-mentioned device is used to run the target script in at least one of the following ways to process the training sample data through the target script: when it is necessary to perform a sampling operation on the training sample data, run the target script, and perform the sampling operation on the training sample data through the target script; when it is necessary to perform a cleaning operation on the training sample data, run the target script, and perform the cleaning operation on the training sample data through the target script; when it is necessary to perform a splicing operation on the training sample data, run the target script, and perform the splicing operation on the training sample data through the target script.
[0151] In an exemplary embodiment, the above-mentioned device is used to extract target sample data from the training sample data according to the target ratio and convert the target sample data into model training data in response to the second interactive operation in the following manner: in response to the second interactive operation, determine the data set identifier and the target ratio; extract the target sample data from the training sample data according to the data set identifier and the target ratio; call the format conversion module to perform format conversion on the target sample data to obtain model training data.
[0152] In an exemplary embodiment, the above-mentioned device is also used to: after obtaining the training sample data uploaded by the user and the data type and data format of the training sample data, obtain the data volume of the training sample data; search for the target node on the target server based on the data volume, wherein the target node represents an idle node whose storage space satisfies the data volume; in response to finding the target node, store the training sample data in the target node; in response to storing the training sample data in the target node, generate a storage path, wherein the storage path is used to indicate the storage address of the training sample data on the target node; store the data identifier, data type, data format and storage path corresponding to the training sample data in the database; in response to the first interactive operation, determine the script template based on the data format, and generate a storage path according to the training sample data. The key fields in the data determine the target key-value pairs; the target key-value pairs are used to replace the target fields in the script template, the target script is automatically generated, and a data processing task is created; the target script is run to execute the data processing task, and the processing results of the data processing task are reported to the database in real time for task status marking; in response to the second interactive operation, the data identifier, data type, data format and storage path corresponding to the training sample data with the task status marked as completed in the database are displayed; the data set identifier and target ratio are determined, wherein the data set identifier indicates that the data identifier satisfies the preset identification condition; the target sample data is extracted from the training sample data according to the data set identifier and the target ratio; the format conversion module is called to perform format conversion on the target sample data to obtain model training data.
[0153] It should be noted that the above modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.
[0154] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.
[0155] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0156] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0157] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0158] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail herein.
[0159] Obviously, those skilled in the art should understand that the above modules or steps of the present application can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.
[0160] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the principles of the present application shall be included in the protection scope of the present application.
Claims
1. A data management method, characterized in that: include: Obtaining training sample data uploaded by a user and the data type and data format of the training sample data; In response to the first interactive operation, a target script is generated and executed, and the training sample data is processed by the target script, wherein the target script is determined by a target key-value pair and a script template specified by the first interactive operation, and the script template is a script template corresponding to the data format; In response to the second interactive operation, target sample data is extracted from the training sample data according to a target ratio, and the target sample data is converted into model training data, wherein the target data type and the target ratio of the target sample data are specified by the second interactive operation, the data type includes the target data type, and the model training data represents the training set used to train the target large model.
2. The method according to claim 1, characterized in that After obtaining the training sample data uploaded by the user and the data type and data format of the training sample data, the method further includes: Storing the training sample data in a target node; In response to storing the training sample data in the target node, generating a storage path, wherein the storage path is used to indicate a storage address of the training sample data on the target node; The data identifier, the data type, the data format and the storage path corresponding to the training sample data are stored in a database.
3. The method according to claim 2, characterized in that The storing the training sample data to the target node includes: Obtaining the data volume of the training sample data; Searching for the target node on the target server based on the data volume, wherein the target node represents an idle node whose storage space satisfies the data volume; In response to finding the target node, the training sample data is stored in the target node.
4. The method according to claim 1, characterized in that: The step of generating and running a target script in response to the first interactive operation, and processing the training sample data by the target script, includes: In response to the first interactive operation, determining the script template based on the data format; In response to the first interactive operation, determining the target key-value pair according to the key field in the training sample data; Use the target key-value pair to replace the target field in the script template to generate the target script; Run the target script to process the training sample data through the target script.
5. The method according to claim 4, characterized in that Different data formats correspond to different target scripts, and running the target script to process the training sample data through the target script includes at least one of the following: In the case where a sampling operation needs to be performed on the training sample data, running the target script, and performing a sampling operation on the training sample data through the target script; In the case where a cleaning operation needs to be performed on the training sample data, running the target script, and performing a cleaning operation on the training sample data through the target script; When it is necessary to perform a splicing operation on the training sample data, the target script is run, and the splicing operation is performed on the training sample data through the target script.
6. The method according to claim 1, characterized in that In response to the second interactive operation, extracting target sample data from the training sample data according to a target ratio, and converting the target sample data into model training data, comprises: In response to the second interaction operation, determining a data set identifier and the target ratio; Extracting target sample data from the training sample data according to the data set identifier and the target ratio; The format conversion module is called to perform format conversion on the target sample data to obtain the model training data.
7. The method according to claim 1, characterized in that The method further comprises: After obtaining the training sample data uploaded by the user and the data type and data format of the training sample data, obtaining the data volume of the training sample data; Searching for the target node on the target server based on the data volume, wherein the target node represents an idle node whose storage space satisfies the data volume; In response to finding the target node, storing the training sample data in the target node; In response to storing the training sample data in the target node, generating a storage path, wherein the storage path is used to indicate a storage address of the training sample data on the target node; Storing the data identifier, the data type, the data format and the storage path corresponding to the training sample data in a database; In response to the first interactive operation, determining the script template based on the data format, and determining the target key-value pair according to the key field in the training sample data; Using the target key-value pair to replace the target field in the script template, automatically generating the target script, and creating a data processing task; Running the target script to execute the data processing task, and reporting the processing result of the data processing task to the database in real time for task status marking; In response to the second interactive operation, displaying the data identifier, the data type, the data format and the storage path corresponding to the training sample data in the database whose task status is marked as completed; Determine a data set identifier and the target ratio, wherein the data set identifier indicates an identifier of a data identifier that satisfies a preset identifier condition; Extracting target sample data from the training sample data according to the data set identifier and the target ratio; The format conversion module is called to perform format conversion on the target sample data to obtain the model training data.
8. A data management device, characterized in that: include: An acquisition module, used to acquire the training sample data uploaded by the user and the data type and data format of the training sample data; A first execution module, configured to generate and run a target script in response to a first interactive operation, and process the training sample data through the target script, wherein the target script is determined by a target key-value pair and a script template specified by the first interactive operation, and the script template is a script template corresponding to the data format; A second execution module is used to extract target sample data from the training sample data according to a target ratio in response to a second interactive operation, and convert the target sample data into model training data, wherein the target data type and the target ratio of the target sample data are specified by the second interactive operation, the data type includes the target data type, and the model training data represents the training set used to train the target large model.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method described in any one of claims 1 to 7 when executed by a processor.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Model training method, model prediction method, and model control system
CN112799850A
Generation method of training data set and training method and device of neural network
CN114120064A
Training sample generation method and device of bill recognition model, server and medium
CN118486036A
Training method of script generation model, script generation method and related device
CN118503093A
Industrial large model optimization method, apparatus and device, and storage medium
CN118708695A