Model training method and system, electronic equipment and storage medium

By sampling, sorting and quality weight determination of training data of multiple sub-data sets, data ratios are optimized to improve the quality of model training data, the problem of insufficient model accuracy in the prior art is solved and higher model accuracy is achieved.

CN119961667APending Publication Date: 2025-05-09GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311473351.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-07
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art is not accurate enough in model training, resulting in the model accuracy still need to be improved.

Method used

By obtaining the training data of multiple sub-data sets, sampling and sorting, the quality weight of each sub-data set is determined, and the data ratio is determined based on the weight, thereby collecting training data for model training.

Benefits of technology

Improve the quality of model training data, and thus improve the accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961667A_ABST
    Figure CN119961667A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a model training method and system, an electronic device and a storage medium, and the model training method comprises the steps: obtaining a training data set which comprises a plurality of sub-data sets; sampling the training data in each sub-data set to obtain a sampling data set, and sorting the training data in the sampling data set according to the quality; according to the sorting position of the training data sampled from each sub-data set in the sampling data set, the quality weight of the corresponding sub-data set is determined, and the quality weight represents the overall training data quality of the single sub-data set; and according to a data ratio determined according to the quality weight, training data is collected from each sub-data set so as to carry out model training, and the data ratio represents the proportion of the training data collected from each sub-data set in the corresponding sub-data set. And the model precision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a model training method, system, electronic device and storage medium. Background Art

[0002] In the field of artificial intelligence, before a model is officially put into use, it is often necessary to train the model based on training data. Generally, the more training data there is and the higher the quality of the training data, the higher the accuracy of the trained model. The quality of training data mainly refers to whether there are errors in the training data, such as whether there are missing values ​​or invalid values ​​in the training data. Training data with fewer errors can have higher quality, and training data with more errors can have lower quality.

[0003] At present, in some technologies, the selection of training data is not accurate enough, resulting in the accuracy of the trained model needs to be improved. Summary of the invention

[0004] In view of this, an embodiment of the present invention provides a model training method, a model training system, an electronic device and a computer-readable storage medium, which can improve the accuracy of model training.

[0005] In one aspect, the present invention provides a model training method, the method comprising:

[0006] Acquire a training data set, where the training data set includes multiple sub-data sets;

[0007] Sampling the training data in each of the sub-data sets to obtain a sampled data set, and sorting the training data in the sampled data set according to quality;

[0008] Determining a quality weight of a corresponding sub-dataset according to a sorting position of training data sampled from each of the sub-datasets in the sampled data set, wherein the quality weight represents the overall training data quality of a single sub-dataset;

[0009] According to the data ratio determined according to the quality weight, training data is collected from each of the sub-datasets to perform model training, and the data ratio represents the proportion of the training data collected from each of the sub-datasets in the corresponding sub-dataset.

[0010] Another aspect of the present invention further provides a model training system, wherein the file includes a plurality of file pages, and the system includes:

[0011] A data acquisition module, used to acquire a training data set, wherein the training data set includes a plurality of sub-data sets;

[0012] A sampling module, used for sampling the training data in each of the sub-data sets to obtain a sampled data set, and sorting the training data in the sampled data set according to quality;

[0013] A weight calculation module, used to determine the quality weight of the corresponding sub-dataset according to the sorting position of the training data sampled from each sub-dataset in the sampled data set, wherein the quality weight represents the training data quality of a single sub-dataset;

[0014] The model training module is used to collect training data from each of the sub-datasets according to the data ratio determined according to the quality weight to perform model training, wherein the data ratio represents the proportion of the training data collected from each of the sub-datasets in the corresponding sub-dataset.

[0015] Another aspect of the present invention provides an electronic device, which includes a processor and a memory, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, the method described above is implemented.

[0016] Another aspect of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the method as described above is implemented.

[0017] In the technical solutions of some embodiments of the present application, the training data sampled from each sub-dataset is sorted according to the quality, and the quality weight of the corresponding sub-dataset is determined based on the sorting position of the training data sampled from each sub-dataset in the sampled data set. This method can accurately quantify the quality weight of the sub-dataset, and then the data ratio determined based on the quality weight can be more accurate. In this way, during model training, the training data collected from each sub-dataset can have a higher quality, so that the trained model can have a higher accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The features and advantages of the present invention will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the present invention in any way. In the accompanying drawings:

[0019] Figure 1 A schematic diagram of a flow chart of a model training method provided by an embodiment of the present application is shown;

[0020] Figure 2 A module schematic diagram of a model training system provided by an embodiment of the present application is shown;

[0021] Figure 3 A schematic diagram of an electronic device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0023] In some model training scenarios, such as training a large-scale language model applied to a specific business field, in order to improve the accuracy of model training, different categories of training data are usually obtained to train the model. Among them, the categories of training data can be divided according to the channel source, content type, content format, etc. of the training data. For example, taking the channel source as an example, the training data obtained through the website and the training data obtained through the system documents can be two categories of training data. For another example, taking the content type as an example, the legal-related training data and the education-related training data can be two categories of training data.

[0024] Since different categories of training data may have different qualities, for example, training data obtained through the website may have lower quality, while training data obtained through system documents may have higher quality, therefore, training data can be screened by quality. Simply put, during the model training process, the sampling amount of high-quality training data is increased and the sampling amount of low-quality training data is reduced. In this way, the model is trained with more high-quality training data to achieve the purpose of improving model accuracy. For example, assuming that the training data of category A includes 2000 data and the training data of category B includes 500 data, the training data of category A has lower quality and the training data of category B has higher quality, then during model training, only 1000 training data can be used out of the 2000 training data of category A, and at the same time, the training data of category B can be reused, that is, category B also has 1000 training data participating in model training. In this way, on the one hand, it ensures that there is sufficient training data for model training, and on the other hand, it improves the quality of training data, thereby improving model accuracy.

[0025] At present, in some technologies, technicians manually judge the quality of each category of training data, and then give a rough usage amount of each category of training data in the model training process. This method is not accurate enough, resulting in the accuracy of the trained model still needs to be improved.

[0026] In view of this, the present application provides a model training method that can improve the model accuracy. The model training method can be applied to electronic devices. Electronic devices include but are not limited to tablet computers, laptops, desktop computers, servers, etc. Figure 1 , which is a flow chart of a model training method provided in one embodiment of the present application. Figure 1 In , the model training method includes the following steps:

[0027] Step S11, obtaining a training data set, where the training data set includes multiple sub-data sets.

[0028] The training data set refers to a set of training data that can be used for model training. In this embodiment, the model is a large-scale language model applied to a specific business field, and the training data in the training data set can also be called text corpus.

[0029] After the training data in the training data set is divided into categories, the training data in each category can constitute a sub-data set. For example, after the training data in the training data set is divided into categories according to the channel source of the training data, the training data from the website can constitute one sub-data set, and the training data from the system document can constitute another sub-data set.

[0030] There may be a data ratio between sub-datasets. The data ratio is used to characterize the relative amount of data between different sub-datasets. Specifically, the data ratio may be the ratio of the storage space occupied by each sub-dataset. For example, sub-dataset A occupies 4G storage space, sub-dataset B occupies 2G storage space, sub-dataset C occupies 5G storage space, and sub-dataset D occupies 1G storage space. Then the data ratio between sub-datasets A, B, C, and D is 4:2:5:1.

[0031] Step S12: sampling the training data in each sub-data set to obtain a sampled data set, and sorting the training data in the sampled data set according to quality.

[0032] Specifically, the so-called sampling means sampling part or all of the training data from each sub-data set. The present application does not limit the amount of training data sampled from each sub-data set. For example, the training data in each sub-data set can be sampled according to the data ratio between the sub-data sets to obtain a sampled data set. For example, assuming that the data ratio between sub-data sets A, B, C, and D is 4:2:5:1, then after sampling 4 training data in sub-data set A, 2 training data can be sampled in sub-data set B, 5 training data can be sampled in sub-data set C, and 1 training data can be sampled in sub-data set D. That is, in the final sampled data set, the ratio of training data sampled from sub-data sets A, B, C, and D is 4:2:5:1. Alternatively, the same amount of training data can be sampled from each sub-data set. For example, 100 training data are sampled from each sub-data set. In a specific embodiment of the present application, a better sampling method is introduced, which will not be repeated here.

[0033] After obtaining the sample data set, each training data in the sample data set can be input into the trained scoring model, and the scoring model outputs a quality score for each training data, wherein the quality score is used to characterize the quality of each training data. Based on the quality score, the training data in the sample data set can be sorted according to quality.

[0034] In this embodiment, since the training data in the sampled data set is a text corpus, a general large-scale language model can be used as a scoring model. The so-called general large-scale language model refers to a large-scale language model that is not applied to a specific business field. In simple terms, when training a large-scale language model applied to a specific business field, a general large-scale language model can be used to assist in scoring the training data in the sampled data set.

[0035] Specifically, the large-scale language model can score the training data in the sampled data set based on the language scoring criteria. The language scoring criteria include but are not limited to universality, readability and content richness. For each piece of training data in the sampled data set, the general large-scale language model can output three scores in the format of [score1, score2, score3]. Among them, score1, score2, and score3 can represent the universality score, readability score, and content richness score of the training data, respectively. For any training data, after averaging the scores of score1, score2, and score3 of the training data, the average value obtained can be used as the quality score of the training data. The higher the quality score of the training data, the higher the quality. According to the order of quality scores from high to low, the training data in the sampled data set can be sorted in the order of quality from high to low. According to the order of quality scores from low to high, the training data in the sampled data set can be sorted in the order of quality from high to low. Since the quality score can objectively reflect the quality of each training data, sorting the training data in the sampled data set based on the quality score can reduce the sorting error.

[0036] Step S13, determining the quality weight of the corresponding sub-dataset according to the sorting position of the training data sampled from each sub-dataset in the sampled data set, the quality weight representing the overall training data quality of the single sub-dataset.

[0037] Specifically, the sorting position may be a sorting sequence number. Take, for example, sorting the training data in the sampled data set in descending order of quality. It is understandable that if most or all of the training data sampled from a sub-data set are located at the front of the sorting queue, it means that among all the sub-data sets, the training data sampled from the sub-data set has a higher quality; if most or all of the training data sampled from a sub-data set are located at the back of the sorting queue, it means that among all the sub-data sets, the training data sampled from the sub-data set has a lower quality. The quality of the sampled data can, to some extent, represent the quality of the training data in the sub-data set where the sampled data is located. Therefore, the overall training data quality of the corresponding sub-data set can be determined based on the sorting position of the training data sampled from each sub-data set in the sampled data set.

[0038] In this embodiment, the quality weight can be used to quantify the training data quality of a single sub-dataset as a whole. The larger the quality weight of a sub-dataset, the better the training data quality of the sub-dataset as a whole. Specifically, when the training data in the sampled data set are sorted in descending order of quality, the quality weight of each sub-dataset can be determined according to expression (1):

[0039] P_A=sum(M-Rank_Ai) / N_A(1)

[0040] Where M represents the number of training data in the sampled data set, Rank_Ai represents the ranking position of the i-th training data sampled from the A-th sub-dataset in the sampled data set, N_A represents the number of training data in the A-th sub-dataset, and P_A represents the quality weight of the A-th sub-dataset.

[0041] It can be understood that in expression (1), if the training data sampled from a sub-dataset is in a forward ranking position, then Rank_Ai will be relatively small, M-Rank_Ai will be relatively large, and P_A will be relatively large; conversely, if the training data sampled from a sub-dataset is in a backward ranking position, then Rank_Ai will be relatively large, M-Rank_Ai will be relatively small, and P_A will be relatively small. Therefore, the value of P_A can reflect the ranking position of the training data sampled from the sub-dataset, and thus can reflect the quality of the training data.

[0042] When the training data in the sampled data set are sorted in order from low to high quality, the quality weight of each sub-data set can be determined according to expression (2):

[0043] P_A=sum(Rank_Ai) / N_A(2)

[0044] The relevant principle of expression (2) is basically similar to that of expression (1), and will not be repeated here.

[0045] For any sub-dataset, expressions (1) and (2) comprehensively consider the ranking positions of the training data sampled from the sub-dataset in the sampled data set, and the quantification of the quality weight is relatively accurate.

[0046] Step S14, according to the data ratio determined according to the quality weight, training data is collected from each sub-dataset to perform model training, and the data ratio represents the proportion of the training data collected from each sub-dataset in the corresponding sub-dataset.

[0047] Specifically, each sub-dataset may have a corresponding data ratio. For any sub-dataset, the larger the data ratio of the sub-dataset, the larger the proportion of the training data collected from the sub-dataset in the sub-dataset. For example, the data ratio of sub-dataset A is 0.8, and the data ratio of sub-dataset B is 0.4, which means that 80% of the training data in sub-dataset A can be used for model training, and 40% of the training data in sub-dataset B can be used for model training.

[0048] It can be understood that if the quality weight of a sub-dataset is greater than the quality weights of other sub-datasets, it means that the training data quality of the sub-dataset is higher than the training data quality of other sub-datasets, therefore, more training data can be collected from the sub-dataset for model training, that is, the sub-dataset can have a higher data ratio. Conversely, if the quality weight of a sub-dataset is smaller than the quality weights of other sub-datasets, it means that the training data quality of the sub-dataset is lower than the training data quality of other sub-datasets, therefore, less training data can be collected from the sub-dataset for model training, that is, the sub-dataset can have a smaller data ratio.

[0049] In this embodiment, the data ratio may be determined based on the following method:

[0050] The quality weight of each sub-data set is normalized to a preset data ratio range to obtain the data ratio corresponding to each sub-data set.

[0051] Specifically, the minimum value of the data matching range can correspond to the minimum quality weight of the sub-data set, and the maximum value of the data matching range can correspond to the maximum quality weight of the sub-data set. Based on this correspondence, each quality weight other than the minimum quality weight and the maximum quality weight can be respectively corresponded to one of the values ​​in the data matching range. This process can be called normalization.

[0052] For ease of understanding, take the data ratio range of 0.5 to 2 as an example, assuming there are three quality weights, 0.7, 0.6, and 0.8, and correspond the minimum quality weight 0.6 to the minimum value 0.5 of the data ratio range, and the maximum quality weight 0.8 to the maximum value 2 of the data ratio range. 0.7 is between 0.6 and 0.8, so 0.7 should correspond to the number in the middle of the data ratio range of 0.5 to 2, that is, 1.25. In this way, normalization is completed.

[0053] It should be noted that if the data ratio of a sub-dataset is greater than 1, it means that during the model training process, part or all of the training data of the sub-dataset needs to be repeatedly collected. For example, the number of training data in a sub-dataset is 500, and the data ratio of the sub-dataset is 1.5, which means that the 500 training data need to be partially and repeatedly collected, that is, 750 training data need to be collected from the sub-dataset for model training.

[0054] The preset data ratio range can be set according to actual needs. For example, in some model training scenarios, it is necessary to collect 10% of the training data from the sub-dataset with the smallest quality weight, and collect 120% of the training data from the sub-dataset with the largest quality weight. In this case, the data ratio range can be set to 0.1-1.2. From this example, it can be seen that by adjusting the data ratio range, the generation logic of the data ratio can be adjusted. This solution can obtain data ratios according to different logics and has better adaptability.

[0055] It is understandable that, in addition to determining the data ratio based on the data ratio range, there are many other ways to achieve the conversion between quality weight and data ratio, such as establishing a conversion equation between quality weight and data ratio, etc. This application does not limit the conversion method between quality weight and data ratio.

[0056] After obtaining the data ratio, training data can be collected from the corresponding sub-datasets according to the data ratios corresponding to each sub-dataset to perform model training.

[0057] In summary, in the technical solutions of some embodiments of the present application, the training data sampled from each sub-dataset is sorted according to the quality, and the quality weight of the corresponding sub-dataset is determined based on the sorting position of the training data sampled from each sub-dataset in the sampled data set. This method can accurately quantify the quality weight of the sub-dataset, and then, the data ratio determined based on the quality weight can be more accurate. In this way, during model training, the training data collected from each sub-dataset can have a higher quality, so that the trained model can have a higher accuracy.

[0058] The solution of this application is further described below.

[0059] Through the above description, it can be understood that if the sampled training data can accurately reflect the overall training data quality of each sub-dataset, then in the sampled data set, the sorting position of the training data should be able to accurately reflect the training data quality differences between different sub-datasets, so that the quality weight and data ratio obtained based on the sorting position of the training data should also be relatively accurate, thereby improving the model accuracy. However, if the quality of the sampled training data cannot accurately reflect the overall training data quality of each sub-dataset, then in the sampled data set, the sorting position of the training data should be inaccurate, which will affect the accuracy of the quality weight and data ratio, and thus affect the model accuracy. For example, when the overall training data quality in a sub-dataset is relatively good, this sub-dataset may also have a small amount of low-quality training data. If during the sampling process, the training data sampled from this sub-dataset is exactly this part of low-quality training data, then based on the sampled training data, it may be judged that the overall training data quality of the sub-dataset is low, and then the quality weight and data ratio determined for the sub-dataset will be relatively small. In this way, during model training, the training data collected from this sub-dataset is relatively small, which ultimately affects the model accuracy.

[0060] In view of this, it is necessary to fully sample each sub-dataset. However, if sampling is performed according to the sampling method described in step S12, there may be a problem of insufficient sampling of some sub-datasets. For example, when sampling is performed according to the data ratio between the sub-datasets, the amount of training data sampled from the sub-dataset with less training data will be relatively small, resulting in insufficient sampling of the sub-dataset with less training data. When sampling the same training data from each sub-dataset, the sub-dataset with less training data may be sampled more fully, but the sub-dataset with more training data may be sampled insufficiently.

[0061] In summary, in some embodiments, the present application proposes a better data sampling method. Specifically, the above-mentioned sampling of the training data in each sub-data set to obtain the sampled data set may include:

[0062] The training data in each sub-dataset is sampled according to different sampling ratios to obtain multiple sampled data sets, where the sampling ratio represents the ratio of the number of training data sampled from each sub-dataset;

[0063] Accordingly, for any sub-dataset, determining the quality weight of the sub-dataset may include:

[0064] Determine the quality weight of each sub-dataset based on the ranking position of the training data sampled from the sub-dataset in each sampled data set;

[0065] The quality weights of the sub-dataset determined based on different sampling data sets are averaged to obtain the quality weight of the sub-dataset.

[0066] In simple terms, multiple samplings are performed according to different sampling ratios to obtain multiple sampling data sets, and then the quality weights of the sub-data sets are determined based on each sampling data set, and finally the quality weights based on different sub-data sets are averaged to obtain the final quality weight of the sub-data set. For example, suppose there are sub-data sets A, B, C, and D, and multiple different sampling ratios are 1:2:4:3 and 3:3:1:2. Then the first sampling data set can be obtained by sampling according to the first sampling ratio 1:2:4:3, that is, for every 1 training data sampled from sub-data set A, 2 training data are sampled from sub-data set B, 4 training data are sampled from sub-data set C, and 3 training data are collected from sub-data set D, so as to obtain the first sampling data set. Then, the second sampling data set can be obtained by continuing to sample according to the second sampling ratio 3:3:1:2. Furthermore, based on the sorting position of the training data sampled from each sub-data set in the first sampling data set, the first quality weight of the corresponding sub-data set can be obtained; based on the sorting position of the training data sampled from each sub-data set in the second sampling data set, the second quality weight of the corresponding sub-data set can be obtained. For any sub-dataset, the first quality weight and the second quality weight of the sub-dataset are averaged, and the obtained average value can be used as the final quality weight of the sub-dataset.

[0067] In the above embodiment, the training data in each sub-data set is sampled multiple times according to different sampling ratios, so that the sampling of the sub-data set can be more sufficient, and thus the quality weight finally obtained can be more accurate.

[0068] Furthermore, in some embodiments, the training data in each sub-data set is sampled according to different sampling ratios to obtain multiple sample data sets, which may include:

[0069] For any sub-dataset, a first quantity of training data is sampled from the sub-dataset, and a second quantity of training data is sampled from all other sub-datasets other than the sub-dataset, to obtain a sampling data set with the sub-dataset as the main sampling set, wherein the first quantity is greater than or equal to the second quantity.

[0070] In simple terms, the number of sampling data sets and sub-data sets is the same, and the sampling data sets and sub-data sets correspond one to one. Each sampling data set uses the corresponding sub-data set as the main sampling set, that is, each sampling data set samples more training data from the corresponding sub-data set. Specifically, in this embodiment, assuming that the amount of training data in the sampling data set is M, the first data amount and the second data amount can be 50%*M respectively. That is, 50%*M of training data is sampled from the sub-data set corresponding to the sampling data set, and then a total of 50%*M of training data is sampled from all other sub-data sets. Among them, when sampling training data in other sub-data sets, sampling can be performed according to the data ratio of other sub-data sets, that is, according to the data ratio of other sub-data sets, a total of 50%*M training data is sampled.

[0071] For ease of understanding, the following example is used to illustrate. Assuming that there are sub-datasets A, B, C, and D, and the data ratio of sub-datasets A, B, C, and D is 5:3:4:1, then sampling can be used to obtain sampled data sets A, B, C, and D, where:

[0072] For the sampled dataset A, 50%*M of training data can be sampled from sub-dataset A and 50%*M of training data can be sampled from sub-dataset B. The training data is sampled from sub-dataset C. The training data is sampled from sub-dataset C. training data;

[0073] For the sampled dataset B, 50%*M of the training data can be sampled from sub-dataset B and 50%*M of the training data can be sampled from sub-dataset A. The training data is sampled from sub-dataset C. The training data is sampled from the sub-dataset D training data;

[0074] And so on.

[0075] In the above embodiment, each sampling data set uses the corresponding sub-data set as the main sampling set, so that each sub-data set can be fully sampled, the sampling accuracy is improved, and the sorting error is reduced.

[0076] Furthermore, in the above embodiment, since data is sampled from the sub-datasets according to a certain ratio, relevant constraints are required to ensure that each sub-dataset has sufficient training data for sampling. Specifically, the sub-datasets include a minimum sub-dataset, which is a sub-dataset that includes the least training data; the training data in each sub-dataset can be sampled according to one or more of the following constraints to obtain a sampled data set:

[0077] 1) The first number is smaller than the number of training data included in the smallest sub-data set, so that the first number of training data can be sampled from each sub-data set.

[0078] 2) The product of the second number and the minimum ratio is greater than 1, and the minimum ratio is the ratio of the number of training data included in the smallest sub-data set to the total number of training data included in all sub-data sets. In this way, it can be ensured that at least one training data is sampled from each sub-data set.

[0079] At this point, the description of all technical solutions of this application is completed.

[0080] Corresponding to the above model training method, this application also provides a model training system. Figure 2 , which is a module diagram of a model training system provided in one embodiment of the present application. Figure 2 In the model training system, the model training system includes:

[0081] A data acquisition module is used to acquire a training data set, where the training data set includes multiple sub-data sets;

[0082] A sampling module is used to sample the training data in each sub-dataset to obtain a sampled data set, and sort the training data in the sampled data set according to quality;

[0083] A weight calculation module, used to determine the quality weight of the corresponding sub-dataset according to the sorting position of the training data sampled from each sub-dataset in the sampled data set, and the quality weight represents the training data quality of a single sub-dataset;

[0084] The model training module is used to collect training data from each sub-dataset according to the data ratio determined according to the quality weight to perform model training. The data ratio represents the proportion of the training data collected from each sub-dataset in the corresponding sub-dataset.

[0085] See also Figure 3 , is a schematic diagram of an electronic device provided by an embodiment of the present application. The electronic device includes a processor and a memory, the memory is used to store a computer program, and when the computer program is executed by the processor, the above method is implemented.

[0086] The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.

[0087] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as program instructions / modules corresponding to the method in the embodiment of the present invention. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions and modules stored in the memory, that is, implementing the method in the above method embodiment.

[0088] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the processor, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0089] One embodiment of the present application further provides a computer-readable storage medium, which is used to store a computer program. When the computer program is executed by a processor, the above method is implemented.

[0090] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A model training method, characterized in that: The method comprises: Acquire a training data set, where the training data set includes multiple sub-data sets; Sampling the training data in each of the sub-data sets to obtain a sampled data set, and sorting the training data in the sampled data set according to quality; Determining a quality weight of a corresponding sub-dataset according to a sorting position of training data sampled from each of the sub-datasets in the sampled data set, wherein the quality weight represents the overall training data quality of a single sub-dataset; According to the data ratio determined according to the quality weight, training data is collected from each of the sub-datasets to perform model training, and the data ratio represents the proportion of the training data collected from each of the sub-datasets in the corresponding sub-dataset.

2. The method according to claim 1, characterized in that The step of sampling the training data in each of the sub-data sets to obtain a sampled data set includes: Sampling the training data in each of the sub-data sets according to different sampling ratios to obtain a plurality of the sampled data sets, wherein the sampling ratio represents the ratio of the number of training data sampled from each of the sub-data sets; For any of the sub-datasets, determining the quality weight of the sub-dataset includes: Determining the quality weight of the sub-datasets respectively based on the sorting positions of the training data sampled from the sub-datasets in each of the sampled data sets; The quality weights of the sub-dataset determined based on different sampling data sets are averaged to obtain the quality weight of the sub-dataset.

3. The method according to claim 2, characterized in that The training data in each of the sub-data sets are sampled according to different sampling ratios to obtain a plurality of the sampled data sets, including: For any of the sub-datasets, a first number of training data is sampled from the sub-dataset, and a second number of training data is sampled from all other sub-datasets other than the sub-dataset, to obtain a sampling data set with the sub-dataset as the main sampling set, wherein the first number is greater than or equal to the second number.

4. The method according to claim 3, characterized in that The sub-dataset includes a minimum sub-dataset, and the minimum sub-dataset is a sub-dataset including the least training data; The training data in each of the sub-data sets are sampled according to one or more of the following constraints to obtain the sampled data sets: The first number is smaller than the number of training data included in the minimum sub-data set; The product of the second number and the minimum ratio is greater than 1, and the minimum ratio is the ratio of the number of training data included in the smallest sub-data set to the total number of training data included in all sub-data sets.

5. The method according to any one of claims 1 to 4, characterized in that: When the training data in the sample data set are sorted in descending order of quality, the quality weight of each sub-data set is determined according to the following expression: sum(M-Rank_Ai) / N_A When the training data in the sample data set are sorted in order of quality from low to high, the quality weight of each sub-data set is determined according to the following expression: sum(Rank_Ai) / N_A Wherein, M represents the number of training data in the sampled data set, Rank_Ai represents the ranking position of the i-th training data sampled from the A-th sub-data set in the sampled data set, and N_A represents the number of training data in the A-th sub-data set.

6. The method according to claim 1, characterized in that The data ratio is determined based on the following method: The quality weight of each of the sub-data sets is normalized to a preset data ratio range to obtain the data ratio corresponding to each of the sub-data sets.

7. The method according to claim 1, characterized in that The step of sorting the training data in the sample data set according to quality includes: Input each training data in the sample data set into a trained scoring model, and the scoring model outputs a quality score for each training data, wherein the quality score is used to characterize the quality of each training data; Based on the quality score, the training data in the sample data set are sorted according to quality.

8. A model training system, characterized in that: The system comprises: A data acquisition module, used to acquire a training data set, wherein the training data set includes a plurality of sub-data sets; A sampling module, used for sampling the training data in each of the sub-data sets to obtain a sampled data set, and sorting the training data in the sampled data set according to quality; A weight calculation module, used to determine the quality weight of the corresponding sub-dataset according to the sorting position of the training data sampled from each sub-dataset in the sampled data set, wherein the quality weight represents the training data quality of a single sub-dataset; The model training module is used to collect training data from each of the sub-datasets according to the data ratio determined according to the quality weight to perform model training, wherein the data ratio represents the proportion of the training data collected from each of the sub-datasets in the corresponding sub-dataset.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: The electronic device comprises a processor and a memory, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.