Artificial intelligence model-based training data set construction method and system

Through the data acquisition and processing technology of the cloud platform, high-quality and diverse artificial intelligence training data sets are generated, which solves data quality and security problems and improves the training efficiency and performance of the model.

CN120277412APending Publication Date: 2025-07-08BEIJING ZHONGKE JINCAI TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510432439.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, the data quality of the training data set is difficult to guarantee, the diversity is insufficient, and data privacy and security are not protected, which affects the training efficiency and performance of artificial intelligence models.

Method used

Data requirements are obtained through the cloud platform based on the training tasks of artificial intelligence models, data acquisition requirements algorithm, data positioning technology and value resource matching algorithm are collected, preprocessed, classified annotated and divided, training data sets are generated, and data security is protected through quantum encryption storage technology, and data is monitored and updated in real time.

Benefits of technology

Ensure data quality and diversity, protect data privacy, improve the training efficiency and performance of artificial intelligence models, and support dynamic updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277412A_ABST
    Figure CN120277412A_ABST
Patent Text Reader

Abstract

The invention discloses a training data set construction method and system based on an artificial intelligence model, and relates to the technical field of artificial intelligence. The method comprises the following steps: analyzing a training task to obtain construction elements, obtaining a data acquisition demand required by training through a data acquisition demand generation algorithm, and obtaining predicted acquisition data, the method comprises the following steps: selecting data required by training through a value resource matching algorithm, collecting to generate an artificial intelligence model training initial data set, classifying and labeling to generate an artificial intelligence model training classification labeling data set, dividing the data set into a training set, a test set and a verification set, and encrypting and storing through a quantum encryption storage technology. And monitoring data in real time and dynamically updating the training data set. According to the invention, data can be acquired according to acquisition requirements to ensure data quality and data diversity; in addition, data values and occupied resources can be analyzed during data collection, it is ensured that the collected data is high in value and small in occupied resources, and sustainable development of the field of artificial intelligence is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a method and system for constructing a training data set based on an artificial intelligence model. Background Art

[0002] With the rapid development of artificial intelligence technology, machine learning models have achieved remarkable results in the fields of computer vision, natural language processing, speech recognition, recommendation systems, etc. However, the performance of the model highly depends on the quality, scale, and diversity of the training data. Currently, the data quality in the constructed training data set is difficult to guarantee, the data diversity is insufficient, and data privacy and security cannot be protected. In view of the above problems, the present invention proposes a method and system for constructing a training data set based on an artificial intelligence model, aiming to improve data quality and diversity through intelligent and automated methods, while protecting data privacy and security, and supporting the dynamic update of the training data set; the method and system will significantly improve the training efficiency and performance of the artificial intelligence model and have broad application prospects. Summary of the Invention

[0003] The present invention provides a method for constructing a training data set based on an artificial intelligence model, including:

[0004] Step S1, the cloud platform obtains the construction elements of the artificial intelligence model according to the training task of the artificial intelligence model and obtains the data collection requirements required for training the artificial intelligence model through a data collection requirement algorithm;

[0005] Step S2, the cloud platform obtains the collection data sources according to the data collection requirements required for training the artificial intelligence model and obtains the expected collected data through data location technology, collects the data required for training based on a value resource matching algorithm, and generates an initial data set for training the artificial intelligence model;

[0006] Step S3, the cloud platform classifies according to the initial data set for training the artificial intelligence model through a classification condition technology, performs annotation and division through a semi-automatic annotation technology, and generates a training set, a test set, and a validation set for the classified and annotated data for training the artificial intelligence model;

[0007] Step S4, the cloud platform generates a training data set for the artificial intelligence model according to the training set, the test set, and the validation set for the classified and annotated data for training the artificial intelligence model, encrypts it through a quantum encryption storage technology, and stores it in a quantum storage device, and dynamically updates the training data set for the artificial intelligence model by monitoring the data sources of the collected data in real time to obtain data update results.

[0008] A method for constructing a training dataset based on an artificial intelligence model as described above, wherein the cloud platform obtains the construction elements of the artificial intelligence model according to the training task of the artificial intelligence model and obtains the data collection requirements for training the artificial intelligence model through a data collection requirements algorithm, including the following sub-steps:

[0009] Step S11: The cloud platform obtains the construction elements of the artificial intelligence model based on the training task of the artificial intelligence model;

[0010] Step S12: The cloud platform obtains the data collection requirements for training the artificial intelligence model through a data collection requirements generation algorithm according to the construction elements of the artificial intelligence model.

[0011] A method for constructing a training dataset based on an artificial intelligence model as described above, wherein the cloud platform obtains the data collection sources according to the data collection requirements for training the artificial intelligence model and obtains the expected collected data through a data positioning technology, and collects the data required for training based on a value resource matching algorithm and generates an initial training dataset for the artificial intelligence model, including the following sub-steps:

[0012] Step S21: The cloud platform obtains the data collection sources according to the data collection requirements for training the artificial intelligence model, and locates the data through the data positioning technology according to the data collection sources to obtain the expected collected data;

[0013] Step S22: The cloud platform calculates the value resource matching degree through a value resource matching algorithm according to the data collection requirements for training the artificial intelligence model and the expected collected data, and collects the data required for training based on the value resource matching degree;

[0014] Step S23: The cloud platform preprocesses the data required for training through a preprocessing technology to generate an initial training dataset for the artificial intelligence model.

[0015] A method for constructing a training dataset based on an artificial intelligence model as described above, wherein the cloud platform classifies the initial training dataset of the artificial intelligence model through a classification condition technology, performs annotation and division through a semi-automatic annotation technology, and generates a training set, a test set, and a validation set of the classified and annotated data for training the artificial intelligence model, including the following sub-steps:

[0016] Step S31: The cloud platform classifies the initial training dataset of the artificial intelligence through a classification condition technology to generate a classified dataset for training the artificial intelligence model;

[0017] Step S32: The cloud platform performs annotation on the classified dataset for training the artificial intelligence model through a semi-automatic annotation technology to generate a classified and annotated dataset for training the artificial intelligence model;

[0018] Step S33: The cloud platform divides the classified and labeled dataset for AI training to generate a training set, a test set, and a validation set of the classified and labeled data for AI model training.

[0019] A method for constructing a training dataset based on an AI model as described above, wherein the cloud platform generates an AI model training dataset according to the training set, test set, and validation set of the classified and labeled data for AI model training, encrypts it through quantum encryption storage technology, and stores it in a quantum storage device. The data source for real-time monitoring and acquisition of data obtains a data update result and dynamically updates the AI model training dataset, including the following sub-steps:

[0020] Step S41: The cloud platform generates an AI model training dataset according to the training set, test set, and validation set of the classified and labeled data for AI model training, encrypts the AI model training dataset through quantum encryption storage technology, and stores it in a quantum storage device;

[0021] Step S42: The cloud platform real-time monitors the data source for data acquisition, judges the data update status, obtains a data update result according to the data update status, and dynamically updates the AI model training dataset according to the data update result.

[0022] The present invention also provides a system for constructing a training dataset based on an AI model, including:

[0023] A module for generating data acquisition requirements for training, which obtains the construction elements of the AI model according to the training task of the AI model and obtains the data acquisition requirements for AI model training through a data acquisition requirements algorithm;

[0024] A module for generating an initial dataset for AI model training, which obtains the data acquisition source according to the data acquisition requirements for AI model training, obtains the expected data to be acquired through data positioning technology, and acquires the data required for training based on a value resource matching algorithm and generates an initial dataset for AI model training;

[0025] A module for data classification, labeling, and division, which classifies the initial dataset for AI model training through classification condition technology, performs labeling through semi-automatic labeling technology, and divides it to generate a training set, a test set, and a validation set of the classified and labeled data for AI model training;

[0026] Data Encryption Storage Update Module: Based on the training of the classification-annotated data training set, test set, and validation set of the artificial intelligence model, generate the training data set of the artificial intelligence model, encrypt it through quantum encryption storage technology, and store it in the quantum storage device. Real-time monitor the data source of the collected data to obtain the data update result and dynamically update the training data set of the artificial intelligence model.

[0027] A training data set construction system based on the artificial intelligence model as described above, wherein the training data acquisition requirement generation module specifically includes:

[0028] Construction Element Generation Sub-module: Based on the training task of the artificial intelligence model, obtain the construction elements of the artificial intelligence model;

[0029] Training Data Acquisition Requirement Generation Sub-module: According to the construction elements of the artificial intelligence model, obtain the training data acquisition requirements of the artificial intelligence model through the data acquisition requirement generation algorithm.

[0030] A training data set construction system based on the artificial intelligence model as described above, wherein the initial training data set generation module of the artificial intelligence model specifically includes:

[0031] Expected Data Acquisition Sub-module: According to the training data acquisition requirements of the artificial intelligence model, obtain the data acquisition source, and locate the data through the data location technology according to the data acquisition source to obtain the expected data to be collected;

[0032] Training Data Acquisition Sub-module: According to the training data acquisition requirements of the artificial intelligence model and the expected data to be collected, calculate the value resource matching degree through the value resource matching algorithm, and collect the training data based on the value resource matching degree;

[0033] Initial Training Data Set Generation Sub-module of the Artificial Intelligence Model: Preprocess the training data through the preprocessing technology to generate the initial training data set of the artificial intelligence model.

[0034] A training data set construction system based on the artificial intelligence model as described above, wherein the data classification and annotation division module specifically includes:

[0035] Data Classification Sub-module: Classify the initial training data set of the artificial intelligence through the classification condition technology to generate the training classification data set of the artificial intelligence model;

[0036] Data Annotation Sub-module: Annotate the training classification data set of the artificial intelligence model through the semi-automatic annotation technology to generate the training classification and annotated data set of the artificial intelligence model;

[0037] The data division sub-module divides the classified and labeled data set for artificial intelligence training to generate a training set of classified and labeled data for artificial intelligence model training, a test set of classified and labeled data for artificial intelligence model training, and a validation set of classified and labeled data for artificial intelligence model training.

[0038] A training data set construction system based on an artificial intelligence model as described above, wherein the data encryption storage and update module specifically includes:

[0039] The data encryption storage sub-module generates an artificial intelligence model training data set based on the training set of classified and labeled data for artificial intelligence model training, the test set of classified and labeled data for artificial intelligence model training, and the validation set of classified and labeled data for artificial intelligence model training, encrypts the artificial intelligence model training data set through quantum encryption storage technology, and stores it in a quantum storage device;

[0040] The data update sub-module monitors the data source of the collected data in real time, judges the data update status, obtains the data update result according to the data update status, and dynamically updates the artificial intelligence model training data set according to the data update result.

[0041] The beneficial effects achieved by the present invention are as follows: The present invention can obtain data requirements according to the model training task before training data collection, collect data according to the clear data collection requirements for training, and ensure data quality and data diversity; The present invention can also analyze the data value and the occupied resources when collecting data to ensure that the collected data has high value and small occupied resources, which is conducive to the sustainable development of the artificial intelligence field. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings.

[0043] Figure 1 It is a flowchart of a method for constructing a training data set based on an artificial intelligence model provided in Embodiment 1 of the present application;

[0044] Figure 2 It is a schematic diagram of a training data set construction system based on an artificial intelligence model provided in Embodiment 2 of the present application. DETAILED DESCRIPTION OF THE INVENTION

[0045] The following clearly and completely describes the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0046] Embodiment 1

[0047] As Figure 1 shown, Embodiment 1 of the present application provides a method for constructing a training data set based on an artificial intelligence model. The method includes the following steps:

[0048] Step S1: The cloud platform obtains the construction elements of the artificial intelligence model according to the training task of the artificial intelligence model and obtains the data collection requirements required for training the artificial intelligence model through the data collection requirement algorithm;

[0049] Further, the cloud platform obtains the construction elements of the artificial intelligence model according to the training task of the artificial intelligence model and obtains the data collection requirements required for training the artificial intelligence model through the data collection requirement algorithm, including the following sub-steps:

[0050] Step S11: The cloud platform obtains the construction elements of the artificial intelligence model based on the training task of the artificial intelligence model;

[0051] Specifically, the construction elements of the artificial intelligence model include the model application scenario, the model application field, the model type, the model complexity, and the model scale; among them, the model application scenario includes, but is not limited to, medical scenarios, educational scenarios, traffic scenarios, shopping scenarios, and smart home scenarios; the model application field includes, but is not limited to, the field of image recognition, the field of natural language processing, the field of language, and the field of recommendation systems; the model type includes, but is not limited to, supervised learning models, reinforcement learning models, neural network models, graph neural models, classification models, generative models, regression models, and sequence models.

[0052] Step S12: The cloud platform obtains the data collection requirements required for training the artificial intelligence model through the data collection requirement generation algorithm according to the construction elements of the artificial intelligence model;

[0053] Specifically, the data collection requirements required for training the artificial intelligence model include the types of data required, the types of data required, and the amount of data required.

[0054] The specific implementation method of the data collection requirement generation algorithm is as follows:

[0055] According to the model application scenario, the model application field, and the model type in the artificial intelligence model construction elements, through the required quantity type generation formula Generate a set of required data types in the data collection requirements for training an artificial intelligence model, where SZL zlj is the set of required data types, zl(cj, ly, mx) is the required data type generation function, and zl i (cj, ly, mx) is the i-th required data type generated according to the application scenario, application field, and model type. n is the total number of generated required data types, and the value range of i is [1, n].

[0056] Represent the generated set of required data types SZL zlj as SZL zlj ={zl1, zl2, …, zl i , …, zl n}, where zl i is the i-th required data type in the set of required data types SZL zlj .

[0057] Generate a set of required data types for the data collection requirements for training an artificial intelligence model through the required data type generation formula according to the model application scenario, model application field, and model type in the artificial intelligence model construction elements , where SLX lxj is the set of required data types, lx(cj, ly, mx) is the required data type generation function, and zl j (cj, ly, mx) is the j-th required data type generated according to the application scenario, application field, and model type. m is the total number of generated required data types, and the value range of j is [1, m].

[0058] Represent the generated set of required data types SLX zlj as SLX zlj ={lx1, lx2, …, lx j , …, lx m}, where lx j is the j-th required data type in the set of required data types SLX zlj .

[0059] According to the set of required data types SZL zlj ={zl1, zl2, …, zl i , …, zl n}, the set of required data types SLX zlj ={lx1, lx2, …, lx j , …, lx m}, and the model complexity and model scale in the artificial intelligence model construction elements through the formula Calculate the required data volume; where XQL is the required data volume, ω is the total influence degree of the required data types on the required data volume, α i is the importance weight of the i-th required data type, |zl i | is the sample quantity of the i-th required data type, n is the total number of required data types, the value range of i is [1, n], ξ is the total influence degree of the required data types on the required data volume, β j is the importance weight of the j-th required data type, |lx j | is the sample quantity of the j-th required data type, m is the total number of required data types, the value range of j is [1, m], ψ is the total influence degree of the model complexity on the required data volume, fz is the model complexity, μ is the total influence degree of the model scale on the required data volume, gm is the model scale.

[0060] Step S2: The cloud platform obtains the data collection source according to the data collection requirements for training the artificial intelligence model and obtains the expected collected data through the data location technology, and collects the data required for training based on the value resource matching algorithm and generates the initial data set for training the artificial intelligence model;

[0061] Furthermore, the cloud platform obtains the data collection source according to the data collection requirements for training the artificial intelligence model and obtains the expected collected data through the data location technology, and collecting the data required for training based on the value resource matching algorithm and generating the initial data set for training the artificial intelligence model includes the following sub-steps:

[0062] Step S21: The cloud platform obtains the data collection source according to the data collection requirements for training the artificial intelligence model, and locates the data through the data location technology according to the data collection source to obtain the expected collected data;

[0063] Specifically, the data collection source includes but is not limited to business databases, log files, sensors, social media, user information, news websites, intelligent devices; the data location technology includes but is not limited to IP location technology, Bluetooth location technology, ultra-fast band location technology, file path location technology, distributed hash table location technology, geospatial location technology; select the appropriate data location technology according to the data collection source to locate the data.

[0064] Step S22: The cloud platform calculates the value resource matching degree according to the data collection requirements for training the artificial intelligence model and the expected collected data, and collects the data required for training based on the value resource matching degree;

[0065] Specifically, the value-resource matching algorithm is used to determine whether the magnitude of data value is proportional to the resources it occupies. When the data value is higher and the occupied resources match its value, the data value is proportional to the occupied resources, and the two match; when the data value is low and the occupied resources do not match its value, the data value is inversely proportional to the occupied resources, and the two do not match.

[0066] The value-resource matching algorithm calculates the value-resource matching degree of the expected collected data based on the data collection requirements and the expected collected data for the training of the artificial intelligence model through the formula where JZP is the value-resource matching degree, max(lgkz ij ) is the maximum available resource estimation parameter for the combined data of the i-th required data type and the j-th required data type in the expected collected data, κ ij is the value importance weight of the combined data of the i-th required data type and the j-th required data type in the expected collected data, is the value of the i-th required data type in the expected collected data, zl i is the i-th required data type in the expected collected data, lnα is the value parameter of the required data type, is the value of the j-th required data type in the expected collected data, lx j is the j-th required data type in the expected collected data, lnβ is the value parameter of the required data type, zzy ij is the resource amount occupied by the combined data of the i-th required data type and the j-th required data type in the expected collected data, n is the total number of data types in the expected collected data, the value range of i is [1, n], m is the total number of data types in the expected collected data, and the value range of j is [1, m].

[0067] Set a preset value-resource matching degree according to the training task of the artificial intelligence model. When the value-resource matching degree of the expected collected data is greater than or equal to the preset value-resource matching degree, it indicates that the expected collected data meets the data collection requirements, and then data collection is performed. When the value-resource matching degree of the expected collected data is less than the preset value-resource matching degree, it indicates that the expected collected data does not meet the data collection requirements, and then data collection is not performed.

[0068] When the total number of data that meets the collection requirements and is collected in the expected collected data is greater than or equal to the required data volume in the data collection requirements for the training of the artificial intelligence model, the data collection is completed and the training required data is generated.

[0069] Step S23: The cloud platform preprocesses the training required data through preprocessing technology to generate an initial data set for the training of the artificial intelligence model;

[0070] Specifically, the data preprocessing technology includes data cleaning, data transformation, data integration, data reduction, and data augmentation. By using the data preprocessing technology, the quality and generalization ability of the data required for training are improved, and an initial data set for artificial intelligence model training is generated based on the data required for training after preprocessing.

[0071] Step S3: The cloud platform classifies the initial data set for artificial intelligence model training through the classification condition technology, labels it through the semi-automatic annotation technology, and divides it to generate a training set, a test set, and a validation set of the classified and labeled data for artificial intelligence model training;

[0072] Furthermore, the cloud platform classifies the initial data set for artificial intelligence model training through the classification condition technology, labels it through the semi-automatic annotation technology, and divides it to generate a training set, a test set, and a validation set of the classified and labeled data for artificial intelligence model training, including the following sub-steps:

[0073] Step S31: The cloud platform classifies the initial artificial intelligence training data set through the classification condition technology to generate a classified data set for artificial intelligence model training;

[0074] Specifically, the initial data set for artificial intelligence model training is divided into T training data units, which is denoted as XLJ csj ={xl1, xl2, …, xl t , …, xl T}), where XLJ csj is the initial data set for artificial intelligence model training, xl t is the t-th training data unit in the initial data set for artificial intelligence model training, T is the total number of training data units, and the value range of t is [1, T].

[0075] Specifically, the specific implementation method of the classification condition technology is as follows:

[0076] Obtain the model training objective according to the training task of the artificial intelligence model, generate the training data classification standard according to the model training objective, and construct a classification condition set FLJ = {FL1, FL2, …, FL j , …, FL J}, where FLJ is the classification condition set, FL j is the j-th classification condition in the classification condition set FLJ, J is the total number of classification conditions, and the value range of j is [1, J].

[0077] Label each classification condition FL j (xl) with a classification condition number. For example, the classification condition number of the first classification condition FL1(xl) is 1, and the j-th classification condition FL jThe classification condition number of (xl) is j, and the value range of the classification condition number is [1, J].

[0078] Construct an empty set of classification data for artificial intelligence model training according to the classification condition number and the classification condition set FLJ, and represent it as XLJ fls ={flj1, flj2, …, flj j , …, flj J}, where XLJ fls is an empty set of classification data for artificial intelligence model training, and flj j is a classification subset of the classification condition with the classification condition number j, and J is the total number of classification subsets.

[0079] According to the initial data set XLJ of artificial intelligence model training csj and the classification condition set FLJ, calculate the classification condition number of the training data unit xl that meets through the formula t , where flh is the set of classification condition numbers of the training data unit, j is the classification condition number, ψ(FL(xl)) is the classification indication function, and FL j (xl t ) is a judgment function for determining whether the t-th training data unit xl t meets the j-th classification condition FL j . When the judgment function FL j (xl t ) holds, ψ(FL j (xl t )) = 1. When the judgment function FL j (xl t ) does not hold, ψ(FL j (xl t )) = 0. T is the total number of training data units, the value range of t is [1, T], J is the total number of classification conditions, and the value range of j is [1, J].

[0080] Store each training data unit in the initial data set XLJ of artificial intelligence model training csj into the corresponding classification subset of the empty set XLJ of classification data for artificial intelligence model training to generate a classification data set for artificial intelligence model training. fls

[0081] Step S32: The cloud platform performs annotation on the classification data set for artificial intelligence model training through semi-automatic annotation technology to generate a classification annotation data set for artificial intelligence model training;

[0082] ​There is a method of annotating the training classification dataset of an artificial intelligence model through semi-automatic annotation technology. First, the data is preliminarily annotated by an annotation machine and annotation suggestions are given. Then, humans review and correct according to the annotation results and suggestions given by the machine, and an artificial intelligence model training classification annotation dataset is generated based on the completed artificial intelligence model training classification dataset.

[0083] Step S33: The cloud platform divides the artificial intelligence training classification annotation dataset to generate an artificial intelligence model training classification annotation data training set, an artificial intelligence model training classification annotation data test set, and an artificial intelligence model training classification annotation data validation set;

[0084] Specifically, obtain the data division ratio according to the training task of the artificial intelligence model, and divide the artificial intelligence model training classification annotation dataset according to the data division ratio. When dividing, make the data ratio in the artificial intelligence model training classification annotation data training set, the artificial intelligence model training classification annotation data test set, and the artificial intelligence model training classification annotation data validation set the same as the data ratio in the artificial intelligence model training classification annotation dataset; for example, if the positive sample data in the artificial intelligence model training classification annotation dataset accounts for 60% and the negative sample data accounts for 40%, the ratio of positive sample data and negative sample data in the divided artificial intelligence model training classification annotation data training set, the artificial intelligence model training classification annotation data test set, and the artificial intelligence model training classification annotation data validation set also needs to be 60% and 40%.

[0085] Step S4: The cloud platform generates an artificial intelligence model training dataset based on the artificial intelligence model training classification annotation data training set, test set, and validation set, encrypts it through quantum encryption storage technology and stores it in a quantum storage device, and dynamically updates the artificial intelligence model training dataset by obtaining the data update result by real-time monitoring the data source of the collected data;

[0086] Furthermore, the cloud platform generates an artificial intelligence model training dataset based on the artificial intelligence model training classification annotation data training set, test set, and validation set, encrypts it through quantum encryption storage technology and stores it in a quantum storage device, and dynamically updates the artificial intelligence model training dataset by obtaining the data update result by real-time monitoring the data source of the collected data, including the following sub-steps:

[0087] Step S41: The cloud platform generates an artificial intelligence model training dataset based on the artificial intelligence model training classification annotation data training set, the artificial intelligence model training classification annotation data test set, and the artificial intelligence model training classification annotation data validation set, encrypts the artificial intelligence model training dataset through quantum encryption storage technology and stores it in a quantum storage device;

[0088] Specifically, construct an empty set of training data for the artificial intelligence model and store the training set of classified and labeled data for the artificial intelligence model, the test set of classified and labeled data for the artificial intelligence model, and the validation set of classified and labeled data for the artificial intelligence model into this empty set to generate a training data set for the artificial intelligence model.

[0089] Encrypt the training data set of the artificial intelligence model through quantum encryption technology and store the encrypted training data set of the artificial intelligence model into a quantum storage device through quantum storage technology.

[0090] Step S42: The cloud platform monitors the data source of the collected data in real time, judges the data update status, obtains the data update result according to the data update status, and dynamically updates the training data set of the artificial intelligence model according to the data update result;

[0091] Specifically, monitor the data source of the collected data in real time, judge whether the collected data is updated. If the data is not updated, the update status is "not updated". If the data is updated, the update status is "updated". According to the update status and the updated data content collected by the data source, dynamically update the training data set of the artificial intelligence model according to the updated content of the data and the training task of the artificial intelligence model, ensuring that the data in the training data set of the artificial intelligence model is timely, accurate and real-time, so that the function of the artificial intelligence model trained according to the training data set of the artificial intelligence model is more adaptable and generalized.

[0092] Embodiment 2

[0093] As Figure 2 shown, Embodiment 2 of the present application provides a system for constructing a training data set based on an artificial intelligence model, including:

[0094] A training data acquisition requirement generation module 21, which obtains the construction elements of the artificial intelligence model according to the training task of the artificial intelligence model and obtains the training data acquisition requirements of the artificial intelligence model through a data acquisition requirement algorithm;

[0095] Further, the training data acquisition requirement generation module 21 includes the following sub-steps:

[0096] A construction element generation sub-module, which obtains the construction elements of the artificial intelligence model based on the training task of the artificial intelligence model;

[0097] Specifically, the construction elements of the artificial intelligence model include the model application scenario, the model application field, the model type, the model complexity, and the model scale. Among them, the model application scenarios include, but are not limited to, medical scenarios, educational scenarios, traffic scenarios, shopping scenarios, and smart home scenarios. The model application fields include, but are not limited to, the field of image recognition, the field of natural language processing, the language field, and the recommendation system field. The model types include, but are not limited to, supervised learning models, reinforcement learning models, neural network models, graph neural models, classification models, generative models, regression models, and sequence models.

[0098] The training data collection requirement generation sub-module obtains the training data collection requirements for the artificial intelligence model through the data collection requirement generation algorithm according to the construction elements of the artificial intelligence model.

[0099] Specifically, the training data collection requirements for the artificial intelligence model include the required data types, the required data formats, and the required data volumes.

[0100] The specific implementation method of the data collection requirement generation algorithm is as follows:

[0101] According to the model application scenario, the model application field, and the model type in the artificial intelligence model construction elements, through the required quantity type generation formula Generate the set of required data types in the training data collection requirements for the artificial intelligence model, where SZL zlj Is the set of required data types, zl(cj,ly,mx) is the required data type generation function, zl i (cj,ly,mx) is the i-th required data type generated according to the application scenario, application field, and model type, n is the total number of generated required data types, and the value range of i is [1,n].

[0102] Express the generated set of required data types SZL zlj As SZL zlj ={zl1,zl2,…,zl i ,…,zl n}, where zl i Is the i-th required data type in the set of required data types SZL zlj .

[0103] According to the model application scenario, the model application field, and the model type in the artificial intelligence model construction elements, through the required data format generation formula Generate the set of required data formats in the training data collection requirements for the artificial intelligence model, where SLX lxj Is the set of required data formats, lx(cj,ly,mx) is the required data format generation function, zl j(cj, ly, mx) is the j-th required data type generated according to the application scenario, application field, and model type, where m is the total number of generated required data types, and the value range of j is [1, m].

[0104] The generated set of required data types SLX zlj is represented as SLX zlj = {lx1, lx2, …, lx j , …, lx m}, where lx j is the j-th required data type in the set of required data types SLX zlj .

[0105] Based on the set of required data types SZL zlj = {zl1, zl2, …, zl i , …, zl n}, the set of required data types SLX zlj = {lx1, lx2, …, lx j , …, lx m}, and the model complexity and model scale in the construction elements of the artificial intelligence model, the required data volume is calculated through the formula ; where XQL is the required data volume, ω is the total influence degree of the required data types on the required data volume, α i is the importance weight of the i-th required data type, |zl i | is the sample quantity of the i-th required data type, n is the total number of required data types, the value range of i is [1, n], ξ is the total influence degree of the required data types on the required data volume, β j is the importance weight of the j-th required data type, |lx j | is the sample quantity of the j-th required data type, m is the total number of required data types, the value range of j is [1, m], ψ is the total influence degree of the model complexity on the required data volume, fz is the model complexity, μ is the total influence degree of the model scale on the required data volume, and gm is the model scale.

[0106] The initial dataset generation module 22 for artificial intelligence model training obtains the data acquisition data source according to the data acquisition requirements for training the artificial intelligence model and obtains the expected acquisition data through data positioning technology, and acquires the training required data based on the value resource matching algorithm and generates the initial dataset for artificial intelligence model training;

[0107] Furthermore, the initial dataset generation module 22 for artificial intelligence model training includes the following sub-steps:

[0108] The expected acquisition data obtaining sub-module obtains the acquisition data source according to the data acquisition requirements for the training of the artificial intelligence model, and locates the data through the data location technology according to the acquisition data source to obtain the expected acquisition data;

[0109] Specifically, the acquisition data source includes but is not limited to business databases, log files, sensors, social media, user information, news websites, intelligent devices; the data location technology includes but is not limited to IP location technology, Bluetooth location technology, ultra-fast band location technology, file path location technology, distributed hash table location technology, and geospatial-based location technology; the corresponding data location technology is selected according to the acquisition data source to locate the data.

[0110] The training required data acquisition sub-module calculates the value resource matching degree through the value resource matching algorithm according to the data acquisition requirements for the training of the artificial intelligence model and the expected acquisition data, and acquires the training required data based on the value resource matching degree;

[0111] Specifically, the value resource matching algorithm is used to judge whether the size of the data value is proportional to the resources it occupies. When the data value is higher and the occupied resources conform to its value, the data value is proportional to the occupied resources, and the two match; when the data value is low and the occupied resources do not conform to its value, the value of the data is inversely proportional to the occupied resources, and the two do not match.

[0112] The value resource matching algorithm calculates the value resource matching degree of the expected acquisition data through the formula according to the data acquisition requirements for the training of the artificial intelligence model and the expected acquisition data; where, JZP is the value resource matching degree, max(lgkz ij ) is the maximum available resource estimation parameter of the combined data of the i-th required data type and the j-th required data type in the expected acquisition data, κ ij is the value importance weight of the combined data of the i-th required data type and the j-th required data type in the expected acquisition data, is the value of the i-th required data type in the expected acquisition data, zl i is the i-th required data type in the expected acquisition data, lnα is the value parameter of the required data type, is the value of the j-th required data type in the expected acquisition data, lx j is the j-th required data type in the expected acquisition data, lnβ is the value parameter of the required data type, zzy ij is the resource amount occupied by the combined data of the i-th required data type and the j-th required data type in the expected acquisition data, n is the total number of data types in the expected acquisition data, the value range of i is [1,n], m is the total number of data types in the expected acquisition data, and the value range of j is [1,m].

[0113] Set a preset value resource matching degree according to the training task of the artificial intelligence model. When the value resource matching degree of the expected collected data is greater than or equal to the preset value resource matching degree, it indicates that the expected collected data meets the data collection requirements, and then data collection is performed. When the value resource matching degree of the expected collected data is less than the preset value resource matching degree, it indicates that the expected collected data does not meet the data collection requirements, and then data collection is not performed.

[0114] When the total number of data that meets the collection requirements and is collected in the expected collected data is greater than or equal to the required data volume in the required data collection requirements for training the artificial intelligence model, the data collection is completed and the data required for training is generated.

[0115] The initial data set generation sub-module for artificial intelligence model training preprocesses the data required for training through preprocessing techniques to generate an initial data set for artificial intelligence model training;

[0116] Specifically, the data preprocessing techniques include data cleaning, data transformation, data integration, data reduction, and data augmentation. The quality and generalization ability of the data required for training are improved through data preprocessing techniques, and an initial data set for artificial intelligence model training is generated according to the data required for training after preprocessing is completed.

[0117] The data classification annotation and division module 23 classifies according to the initial data set for artificial intelligence model training through classification condition techniques, annotates and divides through semi-automatic annotation techniques, and generates a training set, a test set, and a validation set for artificial intelligence model training classification annotation data;

[0118] Furthermore, the data classification annotation and division module 23 includes the following sub-steps:

[0119] The data classification sub-module classifies the initial data set for artificial intelligence training through classification condition techniques to generate a classification data set for artificial intelligence model training;

[0120] Specifically, the initial data set for artificial intelligence model training is divided into T training data units, which are represented as XLJ csj ={xl1, xl2, …, xl t , …, xl T}, where XLJ csj is the initial data set for artificial intelligence model training, xl t is the t-th training data unit in the initial data set for artificial intelligence model training, T is the total number of training data units, and the value range of t is [1, T].

[0121] Specifically, the specific implementation method of the classification condition technique is as follows:

[0122] Obtain the model training objective according to the training task of the artificial intelligence model, generate the training data classification standard according to the model training objective, and construct the classification condition set FLJ = {FL1, FL2, …, FL j , …, FL J}, where FLJ is the classification condition set, FL j is the j-th classification condition in the classification condition set FLJ, J is the total number of classification conditions, and the value range of j is [1, J].

[0123] Number the classification conditions FL j (xl) according to the classification condition number. For example, the classification condition number of the first classification condition FL1(xl) is 1, and the classification condition number of the j-th classification condition FL j (xl) is j, and the value range of the classification condition number is [1, J].

[0124] Construct an empty set of artificial intelligence model training classification data according to the classification condition number and the classification condition set FLJ, and represent it as XLJ fls = {flj1, flj2, …, flj j , …, flj J}, where XLJ fls is the empty set of artificial intelligence model training classification data, flj j is the classification subset of the classification condition with the classification condition number j, and J is the total number of classification subsets.

[0125] Calculate the classification condition number of the classification condition that the training data unit xl csj meets according to the initial data set XLJ of the artificial intelligence model training and the classification condition set FLJ through the formula t , where flh is the set of classification condition numbers of the training data unit, j is the classification condition number, ψ(FL(xl)) is the classification indicator function, and FL j (xl t ) is the judgment function for determining whether the t-th training data unit xl t satisfies the j-th classification condition FL j . When the judgment function FL j (xl t ) holds, ψ(FL j (xl t )) = 1. When the judgment function FL j (xl t ) does not hold, ψ(FL j (xl t )) = 0. T is the total number of training data units, the value range of t is [1, T], J is the total number of classification conditions, and the value range of j is [1, J].

[0126] Store each training data unit in the initial artificial intelligence model training data set XLJ into the corresponding classification subset of the empty artificial intelligence model training classification data set XLJ according to the classification condition number set of the training data unit, and generate the artificial intelligence model training classification data set. csj Store each training data unit in XLJ of the initial data set for training the artificial intelligence model into the corresponding classification subset of the empty set XLJ of classification data for training the artificial intelligence model, and generate the classification data set for training the artificial intelligence model. fls Generate an artificial intelligence model training classification data set.

[0127] A data annotation sub-module that annotates the artificial intelligence model training classification data set through semi-automatic annotation technology to generate an artificial intelligence model training classification annotation data set;

[0128] Specifically, when annotating the artificial intelligence model training classification data set through semi-automatic annotation technology, first conduct preliminary annotation of the data by an annotation machine and give annotation suggestions, and then manually review and correct according to the annotation results and suggestions given by the machine, and generate an artificial intelligence model training classification annotation data set based on the completed artificial intelligence model training classification data set.

[0129] A data division sub-module that divides the artificial intelligence training classification annotation data set to generate an artificial intelligence model training classification annotation data training set, an artificial intelligence model training classification annotation data test set, and an artificial intelligence model training classification annotation data validation set;

[0130] Specifically, obtain the data division ratio according to the training task of the artificial intelligence model, divide the artificial intelligence model training classification annotation data set according to the data division ratio, and ensure that the data ratios in the artificial intelligence model training classification annotation data training set, the artificial intelligence model training classification annotation data test set, and the artificial intelligence model training classification annotation data validation set are the same as those in the artificial intelligence model training classification annotation data set when dividing; for example, if the positive sample data in the artificial intelligence model training classification annotation data set accounts for 60% and the negative sample data accounts for 40%, the ratios of positive sample data and negative sample data in the divided artificial intelligence model training classification annotation data training set, the artificial intelligence model training classification annotation data test set, and the artificial intelligence model training classification annotation data validation set also need to be maintained at 60% and 40%.

[0131] A data encryption storage and update module 24 that generates an artificial intelligence model training data set based on the artificial intelligence model training classification annotation data training set, test set, and validation set, encrypts it through quantum encryption storage technology and stores it in a quantum storage device, and dynamically updates the artificial intelligence model training data set by monitoring the data source of the collected data in real time to obtain the data update result;

[0132] Furthermore, the data encryption storage and update module 24 includes the following sub-steps:

[0133] The data encryption storage sub-module generates an AI model training data set based on the training set of classified and labeled data for AI model training, the test set of classified and labeled data for AI model training, and the validation set of classified and labeled data for AI model training, encrypts the AI model training data set through quantum encryption storage technology, and stores it in a quantum storage device.

[0134] Specifically, construct an empty set of AI model training data, and store the training set of classified and labeled data for AI model training, the test set of classified and labeled data for AI model training, and the validation set of classified and labeled data for AI model training into this empty set to generate an AI model training data set.

[0135] Encrypt the AI model training data set through quantum encryption technology, and store the encrypted AI model training data set in a quantum storage device through quantum storage technology.

[0136] The data update sub-module monitors the data source of the collected data in real time, judges the data update status, obtains the data update result according to the data update status, and dynamically updates the AI model training data set according to the data update result.

[0137] Specifically, monitor the data source of the collected data in real time, and judge whether the collected data is updated. If the data is not updated, the update status is "not updated"; if the data is updated, the update status is "updated". According to the update status and the updated data content collected from the data source, dynamically update the AI model training data set according to the updated content of the data and the training task of the AI model, ensuring that the data in the AI model training data set is timely, accurate, and real-time, so that the function of the AI model trained based on the AI model training data set is more adaptable and generalized.

[0138] The specific implementation manners described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific implementation manners of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for constructing a training data set based on an artificial intelligence model, characterized in that Including: Step S1: The cloud platform obtains the construction elements of the artificial intelligence model according to the training task of the artificial intelligence model and obtains the data collection requirements for training the artificial intelligence model through the data collection requirement algorithm; Step S2: The cloud platform obtains the data collection data source according to the data collection requirements for training the artificial intelligence model and obtains the expected collected data through the data positioning technology, collects the data required for training based on the value resource matching algorithm, and generates the initial data set for training the artificial intelligence model; Step S3: The cloud platform classifies the initial data set for training the artificial intelligence model through the classification condition technology, performs annotation and division through the semi-automatic annotation technology, and generates the training set, test set, and validation set of the classified and annotated data for training the artificial intelligence model; Step S4: The cloud platform generates the data set for training the artificial intelligence model according to the training set, test set, and validation set of the classified and annotated data for training the artificial intelligence model, encrypts it through the quantum encryption storage technology, and stores it in the quantum storage device, and monitors the data source of the collected data in real time to obtain the data update result and dynamically update the data set for training the artificial intelligence model.

2. The method for constructing a training data set based on an artificial intelligence model according to claim 1, wherein, The cloud platform obtains the construction elements of the artificial intelligence model according to the training task of the artificial intelligence model and obtains the data collection requirements for training the artificial intelligence model through the data collection requirement algorithm, including the following sub-steps: Step S11: The cloud platform obtains the construction elements of the artificial intelligence model based on the training task of the artificial intelligence model; Step S12: The cloud platform obtains the data collection requirements for training the artificial intelligence model through the data collection requirement generation algorithm according to the construction elements of the artificial intelligence model.

3. The method for constructing a training data set based on an artificial intelligence model according to claim 1, characterized in that The cloud platform obtains the data collection data source according to the data collection requirements for training the artificial intelligence model and obtains the expected collected data through the data positioning technology, collects the data required for training based on the value resource matching algorithm, and generates the initial data set for training the artificial intelligence model, including the following sub-steps: Step S21: The cloud platform obtains the data collection data source according to the data collection requirements for training the artificial intelligence model, and locates the data through the data positioning technology according to the data collection data source to obtain the expected collected data; Step S22: The cloud platform calculates the value resource matching degree through the value resource matching algorithm according to the data collection requirements for training the artificial intelligence model and the expected collected data, and collects the data required for training based on the value resource matching degree; Step S23: The cloud platform preprocesses the data required for training through the preprocessing technology to generate the initial data set for training the artificial intelligence model.

4. A method for constructing a training data set based on an artificial intelligence model according to claim 1, characterized in that, The cloud platform classifies the initial data set for training the artificial intelligence model through the classification condition technology, performs annotation and division through the semi-automatic annotation technology, and generates the training set, test set, and validation set of the classified and annotated data for training the artificial intelligence model, including the following sub-steps: Step S31: The cloud platform classifies the initial data set for artificial intelligence training through the classification condition technology to generate the classification data set for training the artificial intelligence model; Step S32: The cloud platform performs annotation on the classification data set for training the artificial intelligence model through the semi-automatic annotation technology to generate the classification and annotated data set for training the artificial intelligence model; Step S33: The cloud platform divides the classified and labeled dataset for AI training to generate a training set, a test set, and a validation set of the classified and labeled data for AI model training.

5. The method for constructing a training data set based on an artificial intelligence model according to claim 1, wherein The cloud platform generates a training dataset for the AI model based on the training set, test set, and validation set of the classified and labeled data for AI model training, encrypts it using quantum encryption storage technology, and stores it in a quantum storage device. The data source for real-time monitoring of the collected data obtains the data update result and dynamically updates the training dataset for the AI model, including the following sub-steps: Step S41: The cloud platform generates a training dataset for the AI model based on the training set, test set, and validation set of the classified and labeled data for AI model training, encrypts the training dataset for the AI model using quantum encryption storage technology, and stores it in a quantum storage device. Step S42: The cloud platform monitors the data source of the collected data in real time, determines the data update status, obtains the data update result based on the data update status, and dynamically updates the training dataset for the AI model according to the data update result.

6. A training data set construction system based on an artificial intelligence model, characterized in that, Including: A training data collection requirement generation module that obtains the construction elements of the AI model based on the training task of the AI model and obtains the training data collection requirements of the AI model through a data collection requirement algorithm. An initial dataset generation module for AI model training that obtains the data collection source based on the training data collection requirements of the AI model, obtains the estimated collected data through data localization technology, and collects the training required data based on the value resource matching algorithm and generates an initial dataset for AI model training. A data classification, labeling, and partitioning module that classifies the initial dataset for AI model training through classification condition technology, labels it through semi-automatic labeling technology, and partitions it to generate a training set, a test set, and a validation set of the classified and labeled data for AI model training. A data encryption, storage, and update module that generates a training dataset for the AI model based on the training set, test set, and validation set of the classified and labeled data for AI model training, encrypts it using quantum encryption storage technology, and stores it in a quantum storage device. The data source for real-time monitoring of the collected data obtains the data update result and dynamically updates the training dataset for the AI model.

7. The training data set construction system based on an artificial intelligence model according to claim 6, characterized in that, The training data collection requirement generation module specifically includes: A construction element generation sub-module that obtains the construction elements of the AI model based on the training task of the AI model. A training data collection requirement generation sub-module that obtains the training data collection requirements of the AI model through a data collection requirement generation algorithm based on the construction elements of the AI model.

8. The training data set construction system based on an artificial intelligence model according to claim 6, wherein The initial dataset generation module for AI model training specifically includes: An estimated collected data acquisition sub-module that obtains the data collection source based on the training data collection requirements of the AI model, locates the data through data localization technology according to the data collection source, and obtains the estimated collected data. The data acquisition sub-module for training, according to the data acquisition requirements for the training of the artificial intelligence model and the expected data to be acquired, calculates the value resource matching degree through the value resource matching algorithm, and acquires the data required for training based on the value resource matching degree; The initial data set generation sub-module for the training of the artificial intelligence model preprocesses the data required for training through preprocessing techniques to generate the initial data set for the training of the artificial intelligence model.

9. The training data set construction system based on an artificial intelligence model according to claim 6, characterized in that The data classification, annotation and division module specifically includes: The data classification sub-module classifies the initial data set for the training of the artificial intelligence through classification condition techniques to generate the classified data set for the training of the artificial intelligence model; The data annotation sub-module annotates the classified data set for the training of the artificial intelligence model through semi-automatic annotation techniques to generate the classified and annotated data set for the training of the artificial intelligence model; The data division sub-module divides the classified and annotated data set for the training of the artificial intelligence to generate the training set of the classified and annotated data for the training of the artificial intelligence model, the test set of the classified and annotated data for the training of the artificial intelligence model, and the validation set of the classified and annotated data for the training of the artificial intelligence model.

10. The training dataset construction system based on an artificial intelligence model according to claim 6, wherein The data encryption storage and update module specifically includes: The data encryption storage sub-module generates the data set for the training of the artificial intelligence model according to the training set of the classified and annotated data for the training of the artificial intelligence model, the test set of the classified and annotated data for the training of the artificial intelligence model, and the validation set of the classified and annotated data for the training of the artificial intelligence model, encrypts the data set for the training of the artificial intelligence model through quantum encryption storage technology and stores it in the quantum storage device; The data update sub-module monitors the data source of the acquired data in real time, judges the data update status, obtains the data update result according to the data update status, and dynamically updates the data set for the training of the artificial intelligence model according to the data update result.