Model training method and device, equipment and storage medium
By evaluating importance and text compression of the initial sample sequence, combined with feature fusion processing, the low model accuracy problem caused by too long or too short text is solved, the accuracy and training efficiency of the model are improved, and the robustness and generalization ability of the model are enhanced.
Patent Information
- Application Number
- CN202510328607.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-11
AI Technical Summary
When training models, the existing technology faces the problem of low model accuracy due to too long text or too short text, especially in the application fields of big data financial risk control or credit prediction. The existing methods are difficult to effectively compress text without increasing training difficulty and cost.
By evaluating the importance of the initial sample sequence, text compression is performed using position coding and preset importance thresholds, combining multiple compression and feature fusion, important features are gradually screened out and model parameters are adjusted to improve model accuracy.
During the model training process, important information is retained, sequence length is shortened, model accuracy is improved, computing resource waste is reduced, and model robustness and generalization ability are enhanced.
Smart Images

Figure CN120296414A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of artificial intelligence technology, and in particular, to a model training method, device, equipment and storage medium. Background Art
[0002] In the current field of big data financial risk control or credit prediction applications, as the amount of data grows, the time window also expands, enabling rich historical information to be provided when training a model. However, when the provided historical information is too long, the computational complexity during model training is large, and it is difficult to capture effective information, resulting in a poor accuracy of the trained model. In addition, when training a model with short historical information, it is easy to lack key information and miss some information, leading to a poor accuracy of the trained model. Therefore, there is an urgent need for a method to solve the problem of low accuracy of the trained model caused by too long or too short text. Summary of the Invention
[0003] The purpose of the present invention is to provide at least a model training method, device, equipment and storage medium, which can at least solve the problem of low accuracy of the trained model caused by too long or too short text, and can at least improve the accuracy of the trained model.
[0004] To solve the above technical problems, at least one embodiment of the present application provides a model training method, including: obtaining an initial sample, a sample label of the initial sample, an initial sample sequence, and a sample text length; determining the number of compression times based on the sample text length, and obtaining an initial execution count, where the initial execution count is 0; inputting the initial sample sequence into an initial model, and repeatedly performing importance evaluation on the information corresponding to each position in the initial sample sequence in the initial model to obtain a first importance sequence corresponding to the initial sample sequence; determining a first compressed text based on the first importance sequence, the initial sample, and a preset importance threshold; using the first compressed text as a new initial text, obtaining the initial sample sequence of the initial text, and using the value obtained by adding one to the initial execution count of the first compressed text as an intermediate execution count; returning to execute inputting the initial sample sequence into the initial model until the intermediate execution count is equal to the number of compression times, thereby obtaining a plurality of first text sequences, and using the first text sequence obtained most recently as a target sequence; performing feature fusion processing on the plurality of first text sequences to obtain a fusion sequence; the initial model performing prediction processing based on the target sequence and the fusion sequence to obtain a predicted value; adjusting the model parameters of the initial model based on the predicted value and the sample label to obtain a trained intermediate model.
[0005] At least one embodiment of the present application further provides a model training device, including: an acquisition module, configured to acquire an initial sample, a sample label of the initial sample, an initial sample sequence, and a sample text length; a determination module, configured to determine the number of compression times based on the sample text length, and acquire an initial execution count, where the initial execution count is 0; an input module, configured to input the initial sample sequence into an initial model, and repeatedly perform importance evaluation on information corresponding to each position in the initial sample sequence in the initial model to obtain a first importance sequence corresponding to the initial sample sequence; the determination module is further configured to determine a first compressed text based on the first importance sequence, the initial sample, and a preset importance threshold; the acquisition module is further configured to use the first compressed text as a new initial text, acquire the initial sample sequence of the initial text, and use the value obtained by incrementing the initial execution count by one for the first compressed text as an intermediate execution count; the input module is further configured to return to execute inputting the initial sample sequence into the initial model until the intermediate execution count is equal to the number of compression times, thereby obtaining a plurality of first text sequences, and using the latest obtained first text sequence as a target sequence; a fusion module, configured to perform feature fusion processing on the plurality of first text sequences to obtain a fusion sequence; a prediction module, configured to perform prediction processing on the target sequence and the fusion sequence by the initial model to obtain a predicted value; an adjustment module, configured to adjust model parameters of the initial model based on the predicted value and the sample label to obtain a trained intermediate model.
[0006] At least one embodiment of the present application further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned model training method.
[0007] At least one embodiment of the present application further provides a computer-readable storage medium, storing a computer program, and the computer program realizes the above-mentioned model training method when being executed by a processor.
[0008] The model training method provided by the embodiments of the present application includes: obtaining an initial sample, a sample label of the initial sample, an initial sample sequence, and a sample text length; determining the number of compression times based on the sample text length, and obtaining an initial execution number, where the initial execution number is 0; inputting the initial sample sequence into an initial model, and repeatedly performing importance evaluation on the information corresponding to each position in the initial sample sequence in the initial model to obtain a first importance sequence corresponding to the initial sample sequence; determining a first compressed text based on the first importance sequence, the initial sample, and a preset importance threshold; using the first compressed text as a new initial text, obtaining the initial sample sequence of the initial text, and using the value obtained by adding one to the initial execution number of the first compressed text as an intermediate execution number; returning to execute inputting the initial sample sequence into the initial model until the intermediate execution number is equal to the number of compression times, thereby obtaining a plurality of first text sequences, and using the first text sequence obtained most recently as a target sequence; performing feature fusion processing on the plurality of first text sequences to obtain a fusion sequence; the initial model performing prediction processing based on the target sequence and the fusion sequence to obtain a predicted value; adjusting the model parameters of the initial model based on the predicted value and the sample label to obtain a trained intermediate model. In model training, the long text sequence is compressed multiple times, retaining more important information and gradually shortening the length of the sequence to be processed by the model, making the accuracy of the trained intermediate model relatively high, and further making the accuracy of the target model obtained based on this model training method higher. And during the model training process, the model learns how to screen out important features step by step and how to compress the long text sequence step by step in each compression process, avoiding the difficulty of obtaining key information caused by the long text sequence, reducing the risk of poor accuracy of the trained model, and reducing the waste of computing resources for training the model due to the large amount of redundant information in the long text sequence.
[0009] In some alternative embodiments, adjusting the model parameters of the initial model based on the predicted value and the sample label to obtain a trained intermediate model includes: determining a loss value based on the predicted value and the sample label; adjusting the model parameters of the initial model based on the loss value to obtain a trained intermediate model; when the loss value is not greater than a preset loss threshold, using the intermediate model as the target model put into use, where the target model is used to predict the importance or risk of the input information to obtain a predicted value of the importance degree or risk degree of the input information; when the loss value is greater than the preset loss threshold, using the intermediate model as a new initial model, and returning to execute the steps of obtaining the sample label, the initial sample sequence, and the sample text length of the initial sample until the loss value is not greater than the preset loss threshold, and using the intermediate model as the target model put into use. By adjusting the initial model parameters through the calculated loss value, the error of the initial model is gradually reduced to obtain a target model with an error within a preset range that can be put into use. The accuracy of the target model is improved.
[0010] In some alternative embodiments, determining the first compressed text based on the first importance sequence, the initial sample, and a preset importance threshold includes: determining whether the information at each position in the initial sample needs to be retained based on the preset importance threshold and the first importance sequence; deleting the information in the initial sample that does not need to be retained to obtain the first compressed text. By deleting the information corresponding to the positions in the first importance sequence that are lower than the preset importance threshold, irrelevant and redundant information is deleted, reducing the length of the initial sample while retaining useful information, making the accuracy of the target model trained based on the compressed text higher and the training efficiency also higher.
[0011] In some alternative embodiments, performing an importance assessment on the information corresponding to each position in the initial sample sequence to obtain a first importance sequence corresponding to the initial sample sequence includes: adding a position encoding to the information corresponding to each position in the initial sample sequence to obtain an encoded sample sequence; performing an importance assessment on the information corresponding to each position encoding in the encoded sample sequence to obtain an importance assessment corresponding to each position encoding, and further obtaining a first importance sequence corresponding to the initial sample sequence; where the first importance sequence includes the importance assessment values corresponding to each position encoding. Performing position encoding on the initial sample sequence, compressing it, and still performing position encoding on the compressed initial sample sequence makes it so that although the position information is lost after compression in the initial sample sequence, new position information is supplemented through re - performing position encoding. After compression, position encoding is still performed, ensuring that the initial model can maintain sensitivity to position information even when the sequence length changes.
[0012] In some alternative embodiments, after obtaining the initial sample sequence of the initial text, the method further includes: processing the initial sample sequence by using an average pooling method to obtain the global feature of the initial sample sequence; performing a superposition process on each piece of information in the initial sample sequence and the global feature to obtain a superposed sample sequence; and using the superposed sample sequence as a new initial sample sequence. By using the average pooling method to obtain the global feature that represents the entire initial sample sequence and superposing this global feature with each piece of information in the initial sample sequence, the expression ability of the information at each time step is enhanced.
[0013] In some alternative embodiments, the method further includes: obtaining long text information; converting the long text information into a long text sequence recognizable by the target model; and inputting the long text sequence into the target model to obtain a predicted value corresponding to the long text information, where the predicted value is used to evaluate the importance degree of the long text information or evaluate the risk degree of the long text information. The target model can perform better compression processing on the input long text sequence, extract key information for prediction, and improve the efficiency and accuracy of prediction.
[0014] In some alternative embodiments, the method further includes: obtaining an initial training sample set, deleting the samples with missing information or partially missing information in the initial training sample set to obtain an intermediate training sample set; deleting the samples with labels not conforming to the preset labels in the intermediate training sample set to obtain a target training sample set; and the obtaining of the initial sample and the sample label of the initial sample includes: obtaining the initial sample and the corresponding sample label from the target training sample set. By deleting the samples with incomplete information and selecting the samples meeting the training requirements, the accuracy of the trained target model is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings, and these exemplary illustrations do not constitute a limitation on the embodiments.
[0016] Figure 1 is a schematic flowchart of a model training method provided by an embodiment of the present application;
[0017] Figure 2 is a schematic structural diagram of encoding processing of text provided by another embodiment of the present application;
[0018] Figure 3 is a schematic diagram of a model training device provided by another embodiment of the present application;
[0019] Figure 4 is a schematic structural diagram of an electronic device provided by another embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will elaborate on each embodiment of this application with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of this application, many technical details are provided to help readers better understand this application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in this application can still be implemented. The division of the following embodiments is for convenience of description and should not constitute any limitation on the specific implementation of this application. The various embodiments can be combined and cross-referenced with each other on the premise of not being contradictory.
[0021] To facilitate the understanding of the embodiments of this application, the relevant content of the background technology will be introduced here first.
[0022] In the current field of big data financial risk control or credit prediction applications, as the amount of data grows, the time window also expands, enabling rich historical information to be provided when training a model. However, when the provided historical information is too long, the computational cost during model training is large, and it is difficult to capture effective information, resulting in a relatively poor accuracy of the trained model. In addition, when training a model with relatively short historical information, it is easy to lack key information and miss some information, also leading to a relatively poor accuracy of the trained model. Therefore, there is an urgent need for a method to solve the problem of relatively poor accuracy of the trained model caused by overly long or overly short text. Specifically, in the prior art, sampling methods can be adopted to reduce the length of the input text, but this method is likely to exclude important information, resulting in a relatively low accuracy of the trained model. In addition, a convolutional module can be used to compress the text. Although this method can reduce information loss to a certain extent, some important information will still be lost, and this method lacks flexibility. Or a module with a special attention mechanism can be used to process the entire input text sequence. Although this method maintains the integrity of the input text sequence, it increases the complexity and computational cost of the model. Finally, features can be extracted with the help of a pre-trained encoder, and then important features can be further selected through a feature selection model. However, this method will make the processing flow more complex and increase the difficulty of training the model. Therefore, in the prior art, there is a lack of a solution for effectively compressing long text to train a model with better accuracy without increasing the difficulty and cost of training the model.
[0023] To solve the above technical problem of relatively poor accuracy of the trained model caused by overly long or overly short text, the present invention proposes a model training method. The following will specifically describe the implementation details of the model training method of this embodiment. The following content is only provided for convenience of understanding and is not necessary for implementing this solution.
[0024] It should be noted that the acquisition or use of the data in the embodiments of the present application requires the consent of the user. Relevant data can only be obtained after the user authorizes and permits it, and the acquisition or use of the data complies with the provisions of relevant laws and regulations.
[0025] Embodiment 1:
[0026] The model training method of this embodiment can be applied to an electronic device with communication, computing, and data storage capabilities. Its specific process can be as Figure 1 shown and includes:
[0027] Step 101, obtain an initial sample, a sample label of the initial sample, an initial sample sequence, and a sample text length.
[0028] Specifically, the initial sample is a sample required for training an initial model to obtain a target model. Among them, the initial sample contains various information required to train the target model. For example, if a risk prediction model used in a big data financial risk control scenario needs to be trained, the information included in the initial sample can be user identity information, user account flow information, and the asset situation under the user's name; if a credit assessment model used in a credit prediction scenario needs to be trained, the information included in the initial sample can be user identity information, user non-performing assets, user dishonest event information, and user honest event information, etc. This is only an example for illustration and does not limit the information included in the initial sample. The information included in the initial sample can be set according to actual needs.
[0029] Specifically, the sample label is a label corresponding one-to-one to the initial sample.
[0030] Specifically, the initial sample sequence is to convert the initial sample into information recognizable by the initial model.
[0031] In some examples, the method further includes: obtaining an initial training sample set, deleting samples with missing or partially missing information in the initial training sample set to obtain an intermediate training sample set; deleting samples with labels that do not conform to the preset labels in the intermediate training sample set to obtain a target training sample set; the obtaining of the initial sample and the sample label of the initial sample includes: obtaining the initial sample and the sample label corresponding to the initial sample from the target training sample set.
[0032] Specifically, the initial training sample set includes multiple samples and labels corresponding to the samples, including samples with missing information, incomplete information, and complete information.
[0033] Specifically, the preset labels can be positive sample labels and negative sample labels. For example, in the scenario of big data risk control, if a target model for predicting the risk level of users needs to be trained, the positive sample label can be 1, and the negative sample label can be 0. The positive sample label is used to indicate that the user is a risk-free user, and the negative sample label is used to indicate that the user is a risky user.
[0034] Specifically, the target training sample set is a data subset selected from the initial training sample set for iterative training of the word model, and the target training sample set includes multiple samples.
[0035] Step 102, determine the compression times based on the length of the sample text, and obtain the initial execution times, where the initial execution times are 0.
[0036] In some examples, the determining the compression times based on the length of the sample text includes: obtaining multiple corresponding relationships between preset ranges and times, and taking the time corresponding to the preset range that the length of the sample text conforms to as the compression times.
[0037] Step 103, input the initial sample sequence into the initial model, and repeatedly perform importance evaluation on the information corresponding to each position in the initial sample sequence in the initial model to obtain a first importance sequence corresponding to the initial sample sequence.
[0038] Specifically, when training the initial model for the first time, a random model parameter is set for the initial model using a random initialization algorithm.
[0039] In some examples, the performing importance evaluation on the information corresponding to each position in the initial sample sequence to obtain a first importance sequence corresponding to the initial sample sequence includes: adding position encoding to the information corresponding to each position in the initial sample sequence to obtain an encoded sample sequence; performing importance evaluation on the information corresponding to each position encoding in the encoded sample sequence to obtain the importance evaluation corresponding to each position encoding, and further obtaining a first importance sequence corresponding to the initial sample sequence; where the first importance sequence includes the importance evaluation values corresponding to each position encoding.
[0040] Exemplarily, when the initial model is a deep learning model based on the self-attention mechanism, since the attention mechanism itself does not concern the specific positions of the information in the initial sample sequence, after the initial sample sequence is compressed in this solution, the position information will be damaged because the information at some positions is deleted. To ensure that the initial model can maintain sensitivity to position information even when the sequence length changes, it is necessary to perform position encoding on the initial sample sequence input to the initial model, that is, each time a new initial sample sequence is input to the initial model, it is necessary to perform position encoding on the initial sample sequence to supplement the lost position information due to the compression of the initial sample sequence.
[0041] Step 104, determine a first compressed text based on the first importance sequence, the initial sample, and a preset importance threshold.
[0042] Specifically, the first importance sequence includes the importance values of the information at each position in the initial sample sequence.
[0043] In some examples, determining the first compressed text based on the first importance sequence, the initial sample, and a preset importance threshold includes: determining whether the information at each position in the initial sample needs to be retained based on the preset importance threshold and the first importance sequence; deleting the information in the initial sample that does not need to be retained to obtain the first compressed text.
[0044] In some examples, the prediction module can be used to evaluate the importance of the information corresponding to each position in the initial sample sequence to determine the importance degree of the information corresponding to each position. Exemplarily, the prediction module can be Gumbel-Softmax, which can predict the retention probability of the information corresponding to each position. A retention probability close to 1 indicates that the information corresponding to this position has a high importance degree and needs to be retained. If the retention probability is close to 0, it indicates that the information corresponding to this position has a low importance degree and the information corresponding to this position needs to be deleted. For the relevant content of Gumbel-Softmax, reference can be made to the prior art.
[0045] In some examples, a convolutional layer with an interval of 2 can be used to determine the first compressed text based on the first importance sequence, the initial sample, and a preset importance threshold. Among them, the size of the convolutional kernel of this convolutional layer can be set to 5.
[0046] Optionally, to improve the stability of the training model and accelerate convergence, after obtaining the first compressed text, layer regularization technology can be used to normalize the first compressed text to obtain a new first compressed text.
[0047] In some examples, an attention masking matrix can be constructed based on the first importance sequence. For positions determined to be unimportant, positions with a probability of 0 can be retained, but the attention connections between positions with a probability of 0 and other positions need to be cut off, that is, the attention values are set to 0.
[0048] Therefore, it can be known that during the process of training the initial model, by dynamically learning which information in the initial sample sequence is important and which is unimportant, the initial model can be optimized through backpropagation, and the differentiability during the training process is ensured.
[0049] Step 105: Use the first compressed text as the new initial text, obtain the initial sample sequence of the initial text, and use the value obtained by adding one to the initial execution count as the intermediate execution count for the first compressed text.
[0050] In some examples, the initial text can be encoded by an encoder to obtain an initial sample sequence.
[0051] In some examples, at each compression stage of compressing the initial sample (initial sample sequence), an encoder composed of encoding layers with different numbers of layers can be selected to encode the initial text. For example, as Figure 2 shown, when obtaining the initial text after the first compression, when the text length of the initial text is greater than the first preset threshold, the encoder corresponding to layer k1 can be composed of two encoding layers and encode the initial text to obtain the first text sequence; when obtaining the initial text after the second compression, when the text length of the initial text is greater than the second preset threshold, the encoder corresponding to layer k2 can be composed of six or seven or eight encoding layers and encode the initial text. Among them, each encoding layer in the encoder can be of the Transformer or BERT architecture. By using an encoder composed of multiple encoding layers to encode the initial text to obtain a new first text sequence, each encoding layer can learn different levels of feature abstractions. The bottom encoding layer can capture local features, while the high-level encoding layer can capture more complex patterns and relationships. For the relevant content of the Transformer or BERT architecture, reference can be made to the prior art.
[0052] In some examples, after obtaining the initial sample sequence of the initial text, the method further includes: processing the initial sample sequence using the average pooling method to obtain the global feature of the initial sample sequence; performing a superposition process on each piece of information in the initial sample sequence and the global feature to obtain a superposed sample sequence; and using the superposed sample sequence as the new initial sample sequence.
[0053] Specifically, the global features obtained through average pooling are superimposed with the information (original features) of each time step in the initial sample sequence, enhancing the expressive ability of the information of each time step and facilitating the training of a target model with better accuracy.
[0054] Step 106, return and execute inputting the initial sample sequence into the initial model until the intermediate execution times is equal to the compression times, thereby obtaining multiple first text sequences, and using the most recently obtained first text sequence as the target sequence.
[0055] Step 107, perform feature fusion processing on the multiple first text sequences to obtain a fusion sequence.
[0056] In some examples, since the multiple first text sequences are sequences obtained after compression processing with different numbers of times, the lengths of the first text sequences in the multiple first text sequences are not the same. After the compression processing of the first text sequences, some information will be lost. Therefore, the shortest first text sequence contains the least amount of information. By performing fusion processing on the multiple first text sequences, a more comprehensive and non-redundant fusion sequence can be obtained.
[0057] Specifically, the multiple first text sequences can be compressed to the same length as the shortest first text sequence through a convolutional layer, and then multiplied by preset weight information to obtain the final fusion sequence, where the preset weight information can be the importance values corresponding to the information at each position obtained from the aforementioned first importance sequence.
[0058] Step 108, the initial model performs prediction processing based on the target sequence and the fusion sequence to obtain a predicted value.
[0059] Step 109, adjust the model parameters of the initial model based on the predicted value and the sample label to obtain a trained intermediate model.
[0060] In some examples, adjusting the model parameters of the initial model based on the predicted value and the sample label to obtain a trained intermediate model includes: determining a loss value based on the predicted value and the sample label; adjusting the model parameters of the initial model based on the loss value to obtain a trained intermediate model; if the loss value is not greater than a preset loss threshold, using the intermediate model as the target model put into use, where the target model is used to predict the importance or risk of the input information to obtain a predicted value of the importance or risk level of the input information; if the loss value is greater than the preset loss threshold, using the intermediate model as a new initial model, and returning to execute the steps of obtaining the sample label of the initial sample, the initial sample sequence, and the sample text length until the loss value is not greater than the preset loss threshold, and using the intermediate model as the target model put into use.
[0061] Optionally, when adjusting the model parameters of the initial model based on the loss value, the principle of gradient descent can be followed to adjust the model parameters in the initial model to minimize the loss value.
[0062] In some examples, the method further includes: obtaining long text information; converting the long text information into a long text sequence recognizable by the target model; inputting the long text sequence into the target model to obtain a predicted value corresponding to the long text information, where the predicted value is used to evaluate the importance or risk level of the long text information.
[0063] In summary, the present application obtains an initial sample, a sample label of the initial sample, an initial sample sequence, and a sample text length; determines the number of compression times based on the sample text length, and obtains an initial execution count, where the initial execution count is 0; inputs the initial sample sequence into an initial model, and repeatedly performs importance evaluation on the information corresponding to each position in the initial sample sequence in the initial model to obtain a first importance sequence corresponding to the initial sample sequence; determines a first compressed text based on the first importance sequence, the initial sample, and a preset importance threshold; uses the first compressed text as a new initial text, obtains the initial sample sequence of the initial text, and uses the value obtained by adding one to the initial execution count of the first compressed text as an intermediate execution count; returns to execute inputting the initial sample sequence into the initial model until the intermediate execution count is equal to the number of compression times, thereby obtaining a plurality of first text sequences, and using the first text sequence obtained most recently as a target sequence; performs feature fusion processing on the plurality of first text sequences to obtain a fusion sequence; the initial model performs prediction processing based on the target sequence and the fusion sequence to obtain a predicted value; adjusts the model parameters of the initial model based on the predicted value and the sample label to obtain a trained intermediate model. During model training, the long text sequence is compressed multiple times, retaining more important information and gradually shortening the length of the sequence to be processed by the model, making the accuracy of the trained intermediate model relatively high, and further making the accuracy of the target model obtained based on this model training method higher. Moreover, during the model training process, the model learns how to step by step screen out important features and how to step by step compress the long text sequence, avoiding the difficulty of obtaining key information caused by the long text sequence, reducing the risk of poor accuracy of the trained model, and reducing the waste of computing resources for training the model due to a large amount of redundant information in the long text sequence.
[0064] Specifically, after obtaining the initial samples, an encoder composed of multiple encoding layers and average pooling are used to enhance the sequence features, and the sparsification and encoding can be repeated to complete model prediction and optimization. This method gradually compresses the length of the initial samples during the compression process, reducing the model's demand for computing resources. At the same time, by fusing multiple first text sequences obtained after multiple compression processes, the problem of information loss caused during the compression process is avoided, thereby improving the efficiency of training the initial model and the robustness of the model. It is possible to improve the model training efficiency and performance while avoiding information loss. In addition, by selecting a part of the data from a large number of initial training sample sets as the intermediate training sample set and further extracting the target training sample set from it, the representativeness of the samples for training the initial model is ensured. Selecting specific positive and negative samples as the initial samples to train the initial model helps the model to be exposed to diverse data at an early stage, thereby improving the generalization ability of the initial model and enhancing flexibility. Further selecting the target training sample set for single model iteration training from the intermediate training sample set makes the training process more flexible, and the training data can be dynamically adjusted according to the performance of the initial model during the training process, so as to better adapt to the training requirements of the initial model. By dynamically evaluating the importance of each part of the information in the initial sample sequence and using the Gumbel-Softmax technique to assign a retention probability to each position, the intelligent compression of the initial sample sequence (initial sample) is realized. This not only retains the key information but also improves the ability of the initial model to process long text sequences. Using layer regularization techniques after compression processing helps to stabilize the training process and accelerate convergence. By repeatedly performing the process of sparsification and compression, unimportant information is gradually removed and the initial sample sequence is refined. This method helps to improve the initial model's understanding ability of the input information and ultimately improve the prediction accuracy, efficiency, robustness, and strong generalization ability of the target model, while also having flexibility and scalability.
[0065] Embodiment 2:
[0066] Another embodiment of the present application relates to a model training device. The implementation details of the model training device in this embodiment will be specifically described below. The following content is only the implementation details provided for convenient understanding and is not necessary for implementing the solution. The schematic diagram of the model training device in this embodiment can be as Figure 3 shown, including an acquisition module 301, a determination module 302, an input module 303, a fusion module 304, a prediction module 305, and an adjustment module 306.
[0067] The acquisition module 301 is used to acquire the initial samples, the sample labels of the initial samples, the initial sample sequence, and the sample text length;
[0068] A determination module 302, configured to determine the number of compression times based on the length of the sample text and obtain an initial execution count, where the initial execution count is 0;
[0069] An input module 303, configured to input the initial sample sequence into an initial model, and repeatedly perform importance evaluation on the information corresponding to each position in the initial sample sequence in the initial model to obtain a first importance sequence corresponding to the initial sample sequence;
[0070] The determination module 302 is further configured to determine a first compressed text based on the first importance sequence, the initial sample, and a preset importance threshold;
[0071] The acquisition module 301 is further configured to use the first compressed text as a new initial text, obtain the initial sample sequence of the initial text, and use the value obtained by adding one to the initial execution count of the first compressed text as an intermediate execution count;
[0072] The input module 303 is further configured to return to execute inputting the initial sample sequence into the initial model until the intermediate execution count is equal to the number of compression times, thereby obtaining a plurality of first text sequences, and using the first text sequence obtained most recently as a target sequence;
[0073] A fusion module 304, configured to perform feature fusion processing on the plurality of first text sequences to obtain a fusion sequence;
[0074] A prediction module 305, configured to perform prediction processing on the target sequence and the fusion sequence by the initial model to obtain a predicted value;
[0075] An adjustment module 306, configured to adjust the model parameters of the initial model based on the predicted value and the sample label to obtain a trained intermediate model.
[0076] In one example, when the device is used to adjust the model parameters of the initial model based on the predicted value and the sample label to obtain a trained intermediate model, it is specifically used for: determining a loss value based on the predicted value and the sample label; adjusting the model parameters of the initial model based on the loss value to obtain a trained intermediate model; if the loss value is not greater than a preset loss threshold, using the intermediate model as the target model put into use, where the target model is used to predict the importance or risk of input information to obtain a predicted value of the importance or risk level of the input information; if the loss value is greater than the preset loss threshold, using the intermediate model as a new initial model, and returning to execute the operations of obtaining the sample label, the initial sample sequence, and the sample text length of the initial sample until the loss value is not greater than the preset loss threshold, and using the intermediate model as the target model put into use.
[0077] In one example, when the device is used to determine a first compressed text based on the first importance sequence, the initial sample, and a preset importance threshold, it is specifically used for: determining whether the information at each position in the initial sample needs to be retained based on the preset importance threshold and the first importance sequence; deleting the information in the initial sample that does not need to be retained to obtain the first compressed text.
[0078] Optionally, when the device is used to evaluate the importance of the information corresponding to each position in the initial sample sequence to obtain a first importance sequence corresponding to the initial sample sequence, it is specifically used for: adding position encoding to the information corresponding to each position in the initial sample sequence to obtain an encoded sample sequence; evaluating the importance of the information corresponding to each position encoding in the encoded sample sequence to obtain the importance evaluation corresponding to each position encoding, and further obtaining a first importance sequence corresponding to the initial sample sequence; where the first importance sequence includes the importance evaluation values corresponding to each position encoding.
[0079] Optionally, after obtaining the initial sample sequence of the initial text, the device is further used for: processing the initial sample sequence by using an average pooling method to obtain the global feature of the initial sample sequence; performing a superposition process on each piece of information in the initial sample sequence and the global feature to obtain a superposed sample sequence; using the superposed sample sequence as a new initial sample sequence.
[0080] Optionally, the device is further used for: obtaining long text information; converting the long text information into a long text sequence recognizable by the target model; inputting the long text sequence into the target model to obtain a predicted value corresponding to the long text information, where the predicted value is used to evaluate the importance degree of the long text information or evaluate the risk degree of the long text information.
[0081] Optionally, the device is further configured to: obtain an initial training sample set, delete samples with missing information or partially missing information in the initial training sample set to obtain an intermediate training sample set; delete samples in the intermediate training sample set whose labels do not conform to the preset labels to obtain a target training sample set; when the device is used to obtain the initial samples and the sample labels of the initial samples, it is specifically configured to: obtain initial samples from the target training sample set, and the sample labels corresponding to the initial samples.
[0082] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or implemented as a combination of multiple physical units. In addition, in order to highlight the innovative part of this application, units not closely related to solving the technical problems proposed in this application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.
[0083] Embodiment Three:
[0084] Another embodiment of the present application relates to an electronic device, as Figure 4 shown, including: at least one processor 901; and a memory 902 communicatively connected to the at least one processor 901; wherein, the memory 902 stores instructions executable by the at least one processor 901, and the instructions are executed by the at least one processor 901 to enable the at least one processor 901 to execute the model training methods in the above embodiments.
[0085] Wherein, the memory and the processor are connected by a bus. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be an element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted on the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.
[0086] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store data used by the processor when performing operations.
[0087] Embodiment Four:
[0088] Another embodiment of the present application relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the method embodiments described above are implemented.
[0089] That is, those skilled in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program. This program is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0090] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application. In actual applications, various changes can be made in form and details without departing from the spirit and scope of the present application.
Claims
1. A model training method, characterized in that, Including: Obtain an initial sample, the sample label of the initial sample, the initial sample sequence, and the sample text length; Determine the number of compression times based on the sample text length, and obtain the initial execution times, where the initial execution times are 0; Input the initial sample sequence into the initial model, and repeatedly perform importance evaluation on the information corresponding to each position in the initial sample sequence in the initial model to obtain a first importance sequence corresponding to the initial sample sequence; Determine a first compressed text based on the first importance sequence, the initial sample, and a preset importance threshold; Use the first compressed text as the new initial text, obtain the initial sample sequence of the initial text, and use the value obtained by adding one to the initial execution times of the first compressed text as the intermediate execution times; Return to execute inputting the initial sample sequence into the initial model until the intermediate execution times are equal to the compression times, thereby obtaining multiple first text sequences, and use the first text sequence obtained most recently as the target sequence; Perform feature fusion processing on the multiple first text sequences to obtain a fusion sequence; The initial model performs prediction processing based on the target sequence and the fusion sequence to obtain a predicted value; Adjust the model parameters of the initial model based on the predicted value and the sample label to obtain a trained intermediate model.
2. The model training method according to claim 1, wherein The adjusting the model parameters of the initial model based on the predicted value and the sample label to obtain a trained intermediate model includes: Determine a loss value based on the predicted value and the sample label; Adjust the model parameters of the initial model based on the loss value to obtain a trained intermediate model; If the loss value is not greater than a preset loss threshold, use the intermediate model as the target model put into use, where the target model is used to perform importance or risk prediction on input information to obtain a predicted value of the importance degree or risk degree of the input information; If the loss value is greater than the preset loss threshold, use the intermediate model as the new initial model, and return to execute obtaining the sample label, the initial sample sequence, and the sample text length of the initial sample until the loss value is not greater than the preset loss threshold, and use the intermediate model as the target model put into use.
3. The model training method according to claim 1, wherein The determining the first compressed text based on the first importance sequence, the initial sample, and the preset importance threshold includes: Determine whether the information at each position in the initial sample needs to be retained based on the preset importance threshold and the first importance sequence; Delete the information in the initial sample that does not need to be retained to obtain the first compressed text.
4. The model training method according to claim 1, wherein The performing importance evaluation on the information corresponding to each position in the initial sample sequence to obtain a first importance sequence corresponding to the initial sample sequence includes: Add position encoding to the information corresponding to each position in the initial sample sequence to obtain an encoded sample sequence; Perform importance evaluation on the information corresponding to each position encoding in the encoded sample sequence to obtain the importance evaluation corresponding to each position encoding, thereby obtaining a first importance sequence corresponding to the initial sample sequence; Among them, the first importance sequence includes importance evaluation values corresponding to respective position encodings.
5. The model training method according to claim 1, wherein After obtaining the initial sample sequence of the initial text, the method further includes: Processing the initial sample sequence by using an average pooling method to obtain the global feature of the initial sample sequence; Performing superposition processing on each piece of information in the initial sample sequence and the global feature to obtain a superposed sample sequence; Using the superposed sample sequence as a new initial sample sequence.
6. The model training method according to claim 2, wherein The method further includes: Obtaining long text information; Converting the long text information into a long text sequence recognizable by the target model; Inputting the long text sequence into the target model to obtain a prediction value corresponding to the long text information, where the prediction value is used to evaluate the importance degree of the long text information or evaluate the risk degree of the long text information.
7. The model training method according to any one of claims 1 to 5, characterized in that The method further includes: Obtaining an initial training sample set, deleting samples with missing information or partially missing information in the initial training sample set to obtain an intermediate training sample set; Deleting samples with labels not conforming to a preset label in the intermediate training sample set to obtain a target training sample set; The obtaining of the initial sample and the sample label of the initial sample includes: Obtaining an initial sample from the target training sample set and the sample label corresponding to the initial sample.
8. A model training device, characterized in that, Includes: An obtaining module, configured to obtain an initial sample, a sample label of the initial sample, an initial sample sequence, and a sample text length; A determining module, configured to determine the number of compression times based on the sample text length and obtain an initial execution count, where the initial execution count is 0; An input module, configured to input the initial sample sequence into an initial model, and repeatedly perform importance evaluation on information corresponding to each position in the initial sample sequence in the initial model to obtain a first importance sequence corresponding to the initial sample sequence; The determining module is further configured to determine a first compressed text based on the first importance sequence, the initial sample, and a preset importance threshold; The obtaining module is further configured to use the first compressed text as a new initial text, obtain the initial sample sequence of the initial text, and use the value obtained by adding one to the initial execution count of the first compressed text as an intermediate execution count; The input module is further configured to return to execute inputting the initial sample sequence into the initial model until the intermediate execution count is equal to the compression times, thereby obtaining a plurality of first text sequences, and using the first text sequence obtained most recently as a target sequence; A fusion module, configured to perform feature fusion processing on the plurality of first text sequences to obtain a fusion sequence; A prediction module, configured to perform prediction processing on the target sequence and the fusion sequence by the initial model to obtain a prediction value; An adjustment module, configured to adjust model parameters of the initial model based on the prediction value and the sample label to obtain a trained intermediate model.
9. An electronic device, characterized in that, Includes: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the model training method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the model training method according to any one of claims 1 to 7.