A method and system for constructing an artificial intelligence model training data set

By creating a balanced extracted model cluster in artificial intelligence model training and performing collaborative training, the problem of building high-quality training data sets is solved, and the accuracy and robustness of the model are improved.

CN119830044BActive Publication Date: 2025-06-10BEIJING ZHONGKE JINCAI TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510310225.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-10
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

The existing technology is difficult to efficiently and accurately construct the AI ​​model training dataset, resulting in limited model performance.

Method used

By creating a balanced extraction model cluster for the distributed storage architecture and collaboratively training it, the distributed balanced storage data is extracted, sorted into the original data set, and sampled and data augmented to form a high-quality training data set.

Benefits of technology

It realizes the construction of high-quality training data sets, so that the model can learn the correct features and rules, and improves the accuracy and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119830044B_ABST
    Figure CN119830044B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for constructing an artificial intelligence model training data set, which relates to the technical field of artificial intelligence model training. The method includes: presetting multiple data sources for the construction system and storing the data obtained from different data sources in a distributed manner; creating an equilibrium extraction model cluster for the distributed storage architecture and performing collaborative training on the created equilibrium extraction model cluster; using the trained model cluster to extract the evenly distributed stored data and organizing it into an original data set; sampling the data in the original data set to obtain high-quality training samples; organizing the obtained high-quality training samples into a training data set and then performing persistence for use when training an artificial intelligence model. The training data set constructed by this method can fully reflect the data distribution in the real world, enabling the model to learn correct features and rules, thereby improving the accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence model training, and particularly to a method and system for constructing an artificial intelligence model training data set. Background Art

[0002] With the rapid development of artificial intelligence technology, artificial intelligence models are increasingly widely used in various fields. From image recognition, speech recognition to natural language processing, artificial intelligence is gradually changing people's living and working styles. However, the performance of artificial intelligence models largely depends on the quality and quantity of their training data sets. Therefore, how to efficiently and accurately construct an artificial intelligence model training data set has become a hot and difficult point in current research. Summary of the Invention

[0003] The present invention provides a method for constructing an artificial intelligence model training data set, including:

[0004] Step1: Preset multiple data sources for the construction system, and distribute and store the data obtained from different data sources;

[0005] Step2: Create an equilibrium extraction model cluster for the distributed storage architecture, and perform collaborative training on the created equilibrium extraction model cluster;

[0006] Step3: Use the trained model cluster to extract the evenly distributed stored data and organize it into an original data set;

[0007] Step4: Sample the data in the original data set to obtain high-quality training samples;

[0008] Step5: Organize the obtained high-quality training samples into a training data set and perform persistence, to be used when training an artificial intelligence model.

[0009] For the method for constructing an artificial intelligence model training data set as described above, where an equilibrium extraction model cluster is created for the distributed storage architecture and collaborative training is performed on the created equilibrium extraction model cluster, it is specifically divided into the following sub-steps:

[0010] The central server distributes the initialized equilibrium extraction model to each storage node locally;

[0011] Each storage node uses its own local data to train its local equilibrium extraction model, and after the training is completed, uploads the extracted data samples and the local model parameters to the central server;

[0012] The central server globally validates the distribution characteristics of all data samples. If the validation passes, the training stops. If not, it fuses all local model parameters to generate a new set of model parameters and distributes them to each storage node.

[0013] Each storage node uses the received new model parameters to update the local equilibrium extraction model and conducts a new round of training.

[0014] A method for constructing an artificial intelligence model training dataset as described above, in which data in the original dataset is sampled to obtain high-quality training samples, which is specifically divided into the following sub-steps:

[0015] Perform clustering analysis on the data in the original dataset to obtain multiple data clusters;

[0016] Conduct continuous sampling according to the distance between the data in each data cluster and the cluster center;

[0017] Increment the sampling quantity during continuous sampling to obtain more valuable training samples.

[0018] A method for constructing an artificial intelligence model training dataset as described above, in which the obtained high-quality training samples are sorted into a training dataset and then persisted, which is specifically divided into the following sub-steps:

[0019] Standardize the data format of the training samples and store them as a training dataset;

[0020] Perform data augmentation on the training samples to further enrich the training dataset;

[0021] Transmit the processed training dataset to the cloud for persistence.

[0022] The present invention also provides an artificial intelligence model training dataset construction system, including: a data acquisition module, an extraction model cluster creation module, a data extraction module, a data sampling module, and a data persistence module;

[0023] The data acquisition module is used to preset multiple data sources for the system and distribute and store the data obtained from different data sources;

[0024] The extraction model cluster creation module is used to create an equilibrium extraction model cluster for the distributed storage architecture and conduct collaborative training on the created equilibrium extraction model cluster;

[0025] The data extraction module is used to extract the evenly distributed storage data by using the trained model cluster and organize it into an original dataset;

[0026] The data sampling module is used to sample the data in the original dataset to obtain high-quality training samples;

[0027] A data persistence module, which is used to organize the obtained high-quality training samples into a training data set and perform persistence, and is retrieved when the artificial intelligence model is trained.

[0028] The beneficial effects achieved by the present invention are as follows: it can fully reflect the data distribution of the real world, enable the model to learn correct features and rules, thereby improving the accuracy of the model; by including diverse data, the training set can help the model maintain stable performance when facing different inputs and enhance the robustness of the model. Description of the Drawings

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0030] Figure 1 It is a flowchart of a method for constructing an artificial intelligence model training data set provided in Embodiment 1 of the present application;

[0031] Figure 2 It is a schematic diagram of a system for constructing an artificial intelligence model training data set provided in Embodiment 2 of the present application. Detailed Embodiments

[0032] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0033] Embodiment 1

[0034] As Figure 1 shown, Embodiment 1 of the present application provides a method for constructing an artificial intelligence model training data set, including:

[0035] Step S10: Preset multiple data sources for the construction system and store the data obtained from different data sources in a distributed manner;

[0036] The preset data source can be a local database of a company, institution, or individual, or it can be various modalities of data publicly available on the network. The characteristics of this multi-source heterogeneous data can make the constructed training set more diverse and more realistic. The extracted multi-source heterogeneous data must be uniformly quantified before storage to ensure the consistency of the training data. On the other hand, to improve the read and write performance of the construction system, the obtained data will be stored distributively, and a part of the data processing service is also provided locally at the storage nodes to relieve the pressure on the central server and improve the response rate.

[0037] Step S20: Create a balanced extraction model cluster for the distributed storage architecture and perform collaborative training on the created balanced extraction model cluster.

[0038] The balance of data distribution is crucial for an efficient and accurate training data set. However, the distribution characteristics of data from different data sources are different, and the amount of data obtained from multiple data sources is also very large. It is difficult to quickly and accurately extract data samples in the expected distribution state. Therefore, the concept of using a cluster is proposed here to overcome the above difficulties. Specifically:

[0039] Step S21: The central server distributes the initialized balanced extraction model to each local storage node.

[0040] The balanced extraction model is used to extract data samples in each storage node that meet the expected distribution state. Its mathematical expression is: , where Y is the output of the balanced extraction model, S is the set composed of the data samples already extracted by the balanced extraction model. In the initial state, the S set has only one item of data, that is, any item of data selected from the storage data set X set. Subsequently, the S set will be updated every time the balanced extraction model outputs. is any item of data remaining after removing all elements in the S set from the X set. is any item of data in the S set. is the size of the S set (i.e., the current number of extracted samples). is any item of data in the expected data distribution set (describing the set composed of data samples in the user's expected distribution state, for a certain data source) P set. is the size of the P set (the number of samples in the P set). is for contribution parameter of. represents for contribution parameter of, h is the bandwidth parameter, used to control the smoothness of the model.

[0041] Initializing the model means setting the bandwidth parameter and each contribution parameter in the model to a default value. The default value set in this embodiment is 0.5.

[0042] Step S22: Each storage node trains its local equilibrium extraction model using its respective local data, and after the training is completed, uploads the extracted data samples and the local model parameters to the central server.

[0043] During the training process, the chi-square test method is continuously used to verify whether the data distribution state in the extracted sample set S is consistent with the data distribution state in the expected data distribution set P, and the parameters in the model are iteratively adjusted according to the verification results until the verification passes, at which point the training stops, and the last extracted sample set S and the model parameters at this time are uploaded to the central server.

[0044] Step S23: The central server globally verifies the distribution characteristics of all data samples. If the verification passes, the training stops; if not, a new set of model parameters is generated by fusing all local model parameters and sent to each storage node.

[0045] The global verification also uses the chi-square test method to verify whether the data distribution state of all data sample sets uploaded by the storage nodes is consistent with the data distribution state in the expected global data distribution set C (a set composed of global data samples describing the user's expected distribution state, applicable to all data sources). If they are consistent, the verification passes, and a stop training instruction is sent to each storage node; otherwise, the verification fails, and the formula: is used to calculate the fused value of each model parameter in turn, and the fused new model parameters are sent to each storage node, where is the fused model parameter, represents the data distribution distance between the data sample set uploaded by the j-th storage node and the global data distribution set (obtained using the KL divergence or Wasserstein distance algorithm), j takes values from 1 to m, and m is the total number of storage nodes, represents the data distribution distance between the data sample set uploaded by the k-th storage node and the global data distribution set, k takes values from 1 to m, represents the local model parameter uploaded by the j-th storage node.

[0046] Step S24: Each storage node updates its local equilibrium extraction model using the received new model parameters and conducts a new round of training.

[0047] After each storage node updates its local equilibrium extraction model using the new model parameters, it re-trains the updated model using its local data. After the training is completed, it uploads the sample set and model parameters to the central server for verification, and repeats this process until it receives a training stop instruction.

[0048] Step S30: Use the trained model cluster to extract evenly distributed stored data and organize it into an original data set;

[0049] The trained model cluster can extract data sets in the user-expected distribution state from each storage node. After the central server receives these data sets, it organizes them into a large set, named the original data set, and subsequent training data sets are generated from the original data set.

[0050] Step S40: Sample the data in the original data set to obtain high-quality training samples;

[0051] Sampling the original data set can reduce data noise and improve the value of samples. A good sampling method can directly affect the quality of the final training samples. The sampling method provided in this application is specifically divided into the following sub-steps:

[0052] Step S41: Perform clustering analysis on the data in the original data set to obtain multiple data clusters;

[0053] Clustering analysis can quickly divide the original data set into multiple data clusters with high internal similarity. Sampling based on these data clusters can ensure the balance and diversity of sampling.

[0054] Step S42: Perform continuous sampling according to the distance between the data in each data cluster and the clustering center;

[0055] Continuous sampling means starting from the clustering center, collecting a data sample every L distance. The initial number of collected samples is 200 cases. The initial sampling distance L is calculated by the formula: where is the influence coefficient of data density on sampling frequency, is the influence coefficient of data distribution on sampling frequency, and represent the t-th and (t + 1)-th data from the clustering center outward. T is the total amount of data in the current data cluster, represents and the distance between is the mean of the current data cluster, represents and the distance between

[0056] Step S43: Increase the sampling quantity during continuous sampling to obtain more valuable training samples;

[0057] The farther away from the clustering center, the lower the similarity of the data, and the higher the training value for the model. Therefore, the farther out, the larger the sample size to be collected, so as to obtain more valuable training samples. The increasing coefficient set in this embodiment is 0.05, that is, starting from the completion of the first sample collection, every other L distance, the number of data collected should be 5% more until the sampling is completed. When all calculation results are decimals during the sampling process, they are rounded down.

[0058] Step S50: Organize the obtained high-quality training samples into a training data set and perform persistence for use when training the artificial intelligence model;

[0059] The persistent data set can be quickly loaded during model training, and at the same time, the consistency and repeatability of the data are ensured. The organization of the training samples specifically includes the following sub-steps:

[0060] Step S51: Standardize the data format of the training samples and store them as a training data set;

[0061] Organize the training samples into a table form (such as CSV, Parquet), ensure that each column has a clear field name. For unstructured data, such as images, texts, audios, etc., ensure the file naming specification, and store the metadata (such as annotation information) in a separate table or JSON file. After organization, store all training samples in the training data set.

[0062] Step S52: Perform data augmentation on the training samples to further enrich the training data set;

[0063] There are many existing technologies for data augmentation, which can be selected as needed and will not be elaborated here.

[0064] Step S53: Transmit the processed training data set to the cloud for persistence;

[0065] Use cloud storage technologies such as AWS S3 and Google Cloud Storage to perform persistence on the processed training data set. Its advantages such as high availability, distributed access, and privacy protection can meet all the needs of storing large-scale training data.

[0066] Embodiment 2

[0067] As Figure 2 shown, Embodiment 2 of the present application provides an artificial intelligence model training data set construction system, including: a data acquisition module 21, a model cluster extraction and creation module 22, a data extraction module 23, a data sampling module 24, and a data persistence module 25;

[0068] The data acquisition module 21 is used to preset multiple data sources for the system and store the data obtained from different data sources in a distributed manner;

[0069] The preset data sources can be the local databases of companies, institutions or individuals, or various modalities of data publicly available on the network. The characteristics of this multi-source heterogeneous data can make the constructed training set more diverse and more real. The multi-source heterogeneous data extracted must be uniformly quantified before storage to ensure the consistency of the training data. On the other hand, in order to improve the read and write performance of the constructed system, the obtained data will be stored in a distributed manner, and the storage nodes also provide a part of data processing services locally to relieve the pressure on the central server and improve the response rate.

[0070] The extraction model cluster creation module 22 is used to create a balanced extraction model cluster for the distributed storage architecture and perform collaborative training on the created balanced extraction model cluster; specifically, it includes the following sub-modules:

[0071] 1. The balanced extraction model distribution sub-module, deployed on the central server, is used to distribute the initialized balanced extraction model to each local storage node;

[0072] The balanced extraction model is used to extract data samples that meet the expected distribution state in each storage node, and its mathematical expression is: , where Y is the output of the balanced extraction model, S is the set composed of the data samples already extracted by the balanced extraction model. The S set has only one item in the initial state, that is, any item of data selected from the storage data set X set. Subsequently, the S set will be updated each time the balanced extraction model outputs. is any item of data remaining in the X set after removing all elements in the S set, is any item of data in the S set, is the size of the S set (that is, the current number of extracted samples), is any item of data in the expected data distribution set (describing the set composed of data samples in the user's expected distribution state, for a certain data source) P set, is the size of the P set (the number of samples in the P set), is for contribution parameter of, represents for contribution parameter of, h is the bandwidth parameter, used to control the smoothness of the model.

[0073] Initializing the model means setting both the bandwidth parameter and each contribution parameter in the model to a default value. The default value set in this embodiment is 0.5.

[0074] 2. The balanced extraction model training sub-module is deployed locally on each storage node and is used to train the local balanced extraction model using local data, and upload the extracted data samples and local model parameters to the central server after training is completed;

[0075] During the training process, the chi-square test method is continuously used to verify whether the data distribution state in the extracted sample set S is consistent with the data distribution state in the expected data distribution set P, and various parameters in the model are iteratively adjusted according to the verification results until the verification passes, at which point training stops, and the last extracted sample set S and the model parameters at this time are uploaded to the central server.

[0076] 3. The global verification sub-module is deployed on the central server and is used to globally verify the distribution characteristics of all data samples. If the verification passes, training stops; if not, a new set of model parameters is generated by fusing all local model parameters and sent to each storage node;

[0077] The global verification also uses the chi-square test method to verify whether the data distribution state in all data sample sets uploaded by the storage nodes is consistent with the data distribution state in the expected global data distribution set C (a set composed of global data samples describing the user's expected distribution state, applicable to all data sources). If they are consistent, the verification passes, and a stop training instruction is sent to each storage node; otherwise, the verification fails, and the formula: is used to calculate the fused values of each model parameter in turn, and the fused new model parameters are sent to each storage node, where is the fused model parameter, represents the data distribution distance between the data sample set uploaded by the j-th storage node and the global data distribution set (obtained using the KL divergence or Wasserstein distance algorithm), j takes values from 1 to m, and m is the total number of storage nodes, represents the data distribution distance between the data sample set uploaded by the k-th storage node and the global data distribution set, k takes values from 1 to m, represents the local model parameters uploaded by the j-th storage node.

[0078] 4. The balanced extraction model update sub-module is deployed locally on each storage node and is used to update the local balanced extraction model using the received new model parameters and conduct a new round of training;

[0079] After each storage node updates the local balanced extraction model using the new model parameters, it re-uses the local data to conduct a new round of training on the updated model. After training is completed, the sample set and model parameters are uploaded to the central server for verification, and this process is repeated until a training stop instruction is received.

[0080] The data extraction module 23 is used to extract evenly distributed stored data by using the trained model cluster and organize it into an original data set.

[0081] The trained model cluster can extract data sets in the user-expected distribution state from each storage node. After the central server receives these data sets, it organizes them into a large set, named the original data set, and the subsequent training data sets are generated from the original data set.

[0082] The data sampling module 24 is used to sample the data in the original data set to obtain high-quality training samples. It specifically includes the following sub-modules:

[0083] 1. The clustering analysis sub-module is used to perform clustering analysis on the data in the original data set to obtain multiple data clusters.

[0084] Clustering analysis can quickly divide the original data set into multiple data clusters with high internal similarity. Sampling based on these data clusters can ensure the balance and diversity of sampling.

[0085] 2. The continuous sampling sub-module is used to perform continuous sampling according to the distance between the data in each data cluster and the clustering center, and increase the sampling quantity during continuous sampling to obtain more valuable training samples.

[0086] Continuous sampling means starting from the clustering center, collecting a data sample every L distance. The initial collection quantity is 200 cases, and the initial collection distance L is calculated by the formula: where is the influence coefficient of data density on the sampling frequency, is the influence coefficient of data distribution on the sampling frequency, and represent the t-th and (t + 1)-th data from the clustering center outward, T is the total amount of data in the current data cluster, represents and the distance between them, is the mean value of the current data cluster, represents and the distance between them.

[0087] During continuous sampling, increase the number of samples to obtain more valuable training samples. The farther away from the clustering center, the lower the similarity of the data and the higher the training value for the model. Therefore, the farther out, the larger the sample size should be to obtain more valuable training samples. The increment coefficient set in this embodiment is 0.05. That is to say, starting from the completion of the first sample collection, every other L distance, the number of collected data should be 5% more until the sampling is completed. When there are decimals in all calculation results during the sampling process, round down.

[0088] The data persistence module 25 is used to organize the obtained high-quality training samples into a training data set and perform persistence for use when training the artificial intelligence model; it specifically includes the following sub-modules:

[0089] 1. The data format standardization sub-module is used to standardize the data format of the training samples and store them as a training data set;

[0090] Organize the training samples into a table form (such as CSV, Parquet), ensure that each column has a clear field name. For unstructured data such as images, texts, and audios, ensure the file naming convention and store the metadata (such as annotation information) in a separate table or JSON file. After finishing the organization, store all the training samples into the training data set.

[0091] 2. The data augmentation sub-module is used to perform data augmentation on the training samples to further enrich the training data set;

[0092] There are many existing technologies for data augmentation. Just select according to needs and will not be elaborated here.

[0093] 3. The data storage sub-module is used to transfer the processed training data set to the cloud for persistence;

[0094] Use cloud storage technologies such as AWS S3 and Google Cloud Storage to perform persistence on the processed training data set. Its advantages such as high availability, distributed access, and privacy protection can meet all the requirements for storing large-scale training data.

[0095] Corresponding to the above embodiment, an embodiment of the present invention provides a computer storage medium, including: at least one memory and at least one processor;

[0096] The memory is used to store one or more program instructions;

[0097] The processor is used to run one or more program instructions to execute a method for constructing an artificial intelligence model training data set.

[0098] Corresponding to the above embodiments, an embodiment of the present invention provides a computer-readable storage medium. The computer storage medium contains one or more program instructions, and the one or more program instructions are used to be executed by a processor to implement a method for constructing an artificial intelligence model training data set.

[0099] An embodiment disclosed by the present invention provides a computer-readable storage medium. Computer program instructions are stored in the computer-readable storage medium. When the computer program instructions run on a computer, the computer is enabled to execute the above-mentioned method for constructing an artificial intelligence model training data set.

[0100] In an embodiment of the present invention, the processor may be an integrated circuit chip with the ability to process signals. The processor may be a general-purpose processor, a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0101] It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention may be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The processor reads the information in the storage medium and combines its hardware to complete the steps of the above method.

[0102] The storage medium may be a memory, for example, it may be a volatile memory or a non-volatile memory, or may include both a volatile memory and a non-volatile memory.

[0103] Among them, the non-volatile memory may be a read-only memory (ROM for short), a programmable read-only memory (PROM for short), an erasable programmable read-only memory (EPROM for short), an electrically erasable programmable read-only memory (EEPROM for short), or a flash memory.

[0104] The volatile memory may be a Random Access Memory (RAM) which serves as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM).

[0105] The storage media described in the embodiments of the present invention are intended to include, but are not limited to, these and any other suitable types of memory.

[0106] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the present invention can be implemented by a combination of hardware and software. When applying software, the corresponding functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transfer of a computer program from one place to another. The storage media can be any available medium accessible by a general or special purpose computer.

[0107] The specific embodiments described above further elaborate on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for constructing an artificial intelligence model training data set, characterized in that: include: Step 1: Preset multiple data sources for building the system, and store the data obtained from different data sources in a distributed manner; Step 2: Create a balanced extraction model cluster for the distributed storage architecture and perform collaborative training on the created balanced extraction model cluster. The specific steps are as follows: The central server sends the initialized balanced extraction model to each storage node locally; The balanced extraction model is used to extract data samples that meet the expected distribution state in each storage node. Its mathematical expression is: , where Y is the output of the balanced extraction model, S is the set of data samples extracted by the balanced extraction model. In the initial state, the S set has only one data item, that is, any data item selected from the storage data set X. The subsequent balanced extraction model will update the S set once each time it outputs. It is any data remaining after removing all elements from set S from set X. is any data in the set S, is the size of the S set, that is, the number of samples currently extracted, It is any data in the expected data distribution set P. The expected data distribution set P is used to describe the set of data samples in the user's expected distribution state for a certain data source. is the size of the P set, that is, the number of samples in the P set, yes right The contribution parameter, express right The contribution parameter, h is the bandwidth parameter, which is used to control the smoothness of the model; Each storage node uses its own local data to train its local balanced extraction model, and uploads the extracted data samples and local model parameters to the central server after training is completed; The central server performs global verification on the distribution characteristics of all data samples. If the verification passes, training is stopped. If not, all local model parameters are integrated to generate a new set of model parameters and sent to each storage node. If all data sample sets uploaded by the storage node are consistent with the expected global data distribution set data distribution status, the verification passes and a stop training instruction is sent to each storage node. Otherwise, the verification fails and the formula is used: , calculate the fused values ​​of each model parameter in turn, and send the fused new model parameters to each storage node, where are the model parameters after fusion, It represents the data distribution distance between the data sample set uploaded by the jth storage node and the global data distribution set. j ranges from 1 to m, where m is the total number of storage nodes. It represents the data distribution distance between the data sample set uploaded by the k-th storage node and the global data distribution set. k ranges from 1 to m. Represents the local model parameters uploaded by the jth storage node; Each storage node uses the received new model parameters to update the local balanced extraction model and conduct a new round of training; Step 3: Use the trained model cluster to extract the evenly distributed storage data and organize it into the original data set; Step 4: Sample the data in the original data set to obtain high-quality training samples; Step 5: Organize the obtained high-quality training samples into a training data set and persist it for use during artificial intelligence model training.

2. The method for constructing an artificial intelligence model training data set according to claim 1, characterized in that: Sampling the data in the original data set to obtain high-quality training samples is divided into the following sub-steps: Perform cluster analysis on the data in the original data set to obtain multiple data clusters; Continuous sampling is performed based on the distance between the data in each data cluster and the cluster center; The number of samples is increased during continuous sampling to obtain more valuable training samples.

3. The method for constructing an artificial intelligence model training data set according to claim 1, characterized in that: The obtained high-quality training samples are organized into training data sets and persisted, which is specifically divided into the following sub-steps: Standardize the data format of training samples and store them as training datasets; Perform data enhancement on training samples to further enrich the training data set; The processed training dataset is transferred to the cloud for persistence.

4. An artificial intelligence model training data set construction system, used to execute an artificial intelligence model training data set construction method according to any one of claims 1 to 3, characterized in that: include: Data acquisition module, extraction model cluster creation module, data extraction module, data sampling module, data persistence module; The data acquisition module is used to preset multiple data sources for the system and store the data acquired from different data sources in a distributed manner; The extraction model cluster creation module is used to create a balanced extraction model cluster for the distributed storage architecture and to perform collaborative training on the created balanced extraction model cluster; The data extraction module is used to extract evenly distributed storage data using the trained model cluster and organize it into the original data set; The data sampling module is used to sample the data in the original data set to obtain high-quality training samples; The data persistence module is used to organize the obtained high-quality training samples into training data sets and persist them for use during artificial intelligence model training.

5. The artificial intelligence model training data set construction system according to claim 4, characterized in that: The extraction model cluster creation module specifically includes: The balanced extraction model delivery submodule is deployed on the central server and is used to deliver the initialized balanced extraction model to each storage node locally; The balanced extraction model training submodule is deployed locally on each storage node and is used to train its local balanced extraction model using local data. After the training is completed, the extracted data samples and local model parameters are uploaded to the central server. The global verification submodule is deployed on the central server and is used to globally verify the distribution characteristics of all data samples. If the verification passes, the training is stopped. If not, all local model parameters are integrated to generate a new set of model parameters, which are then sent to each storage node. The balanced extraction model update submodule is deployed locally on each storage node and is used to update the local balanced extraction model using the received new model parameters and perform a new round of training.

6. The artificial intelligence model training data set construction system according to claim 4, characterized in that: The data sampling module specifically includes: The cluster analysis submodule is used to perform cluster analysis on the data in the original data set to obtain multiple data clusters; The continuous sampling submodule is used to perform continuous sampling according to the distance between the data in each data cluster and the cluster center, and to increase the number of samples during continuous sampling to obtain more valuable training samples.

7. The artificial intelligence model training data set construction system according to claim 4, characterized in that: The data persistence module specifically includes: The data format standardization submodule is used to standardize the data format of the training samples and store them as training data sets; The data enhancement submodule is used to enhance the training samples and further enrich the training data set; The data storage submodule is used to transfer the processed training data set to the cloud for persistence.

Citation Information

Patent Citations

  • Distributed model training method for deep learning in edge computing and related device

    CN115114982A

  • Imbalanced data Internet of Things intrusion detection method and device

    CN117118689A