Model training method and device, electronic equipment and storage medium
By screening data on pre-trained language models and constructing target training data, the shortcomings of traditional models in metaphor recognition and generation are solved, and the model's metaphorical expression ability and complex reasoning ability are improved.
Patent Information
- Application Number
- CN202510235391.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-10
AI Technical Summary
Traditional pre-trained language models do not perform well in metaphor recognition and generation, especially in complex details recognition and contextual reasoning, while having limited ability to recognize or generate metaphors in longer texts.
Through a model training method, the pre-trained data is processed and screened using a large language model, the target training data is constructed, and the pre-trained model is trained based on the data to optimize the model's metaphorical recognition and generation ability.
It improves the ability of the target model in metaphor recognition and generation, makes up for the shortcomings of traditional models in complex reasoning and scene adaptability, and provides a stronger basis for metaphorical expression.
Smart Images

Figure CN120124751A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the field of computer application technologies, and more particularly, relates to a model training method and apparatus, an electronic device, and a storage medium. Background Art
[0002] Currently, pre-trained language models such as BERT have achieved remarkable success in text understanding and generation. These models can deeply capture semantics and complex language structures through extensive pre-training on large-scale data. However, they often perform poorly in metaphor extraction and generation. Metaphor understanding requires not only accurate semantic and context understanding but also a detailed analysis of figurative and extended meanings, which poses higher requirements for the model's interpretation ability. Traditional pre-trained language models such as BERT or GPT-2 perform well in processing general text features but lack in identifying complex details and performing context reasoning, resulting in their insufficient performance in metaphor-related tasks. In addition, these models usually have limitations in context length, which restricts their ability to identify or generate metaphors in longer texts. Summary of the Invention
[0003] An object of the present disclosure is to provide a model training method and apparatus, an electronic device, and a storage medium to improve the metaphor recognition or metaphor generation ability of a target model.
[0004] In a first aspect of an embodiment of the present disclosure, a model training method is provided, including: Inputting pre-training data into a first model for processing to obtain an output probability value, and screening the pre-training data according to a comparison result between the output probability value and a first threshold to obtain first data, where the first model is a large language model; Integrating second data and first prompt data to obtain target training data, where the second data is data obtained by preprocessing the first data, and the first prompt data is a preset prompt; Training a pre-trained model according to the target training data to obtain a target model; the target model is used to identify metaphors or generate metaphors.
[0005] Optionally, the inputting pre-training data into a first model for processing to obtain an output probability value includes: Concatenating a preset discriminative prompt with the pre-training data, and calculating the output probability value according to an output probability value calculation formula in the first model; The output probability value calculation formula is:
[0006] Where is the output probability value, is the activation function, W is the network weight, b is the network bias, pool represents the pooling layer, and Encoder represents the Transformer encoding layer. is the embedding vector. is the preset discriminant prompt word. T i is the pre-training data.
[0007] Optionally, screening the pre-training data according to the comparison result between the output probability value and the first threshold to obtain the first data includes: If the output probability value is greater than the first threshold, summarize the corresponding pre-training data to obtain the first data; If the output probability value is less than or equal to the first threshold, clear the corresponding pre-training data.
[0008] Optionally, the model training method further includes: Process the first data to obtain the first string set vector; Calculate the similarity between each substring vector in the first string set vector and the first word vector and the second word vector respectively; Select the corresponding substrings according to the similarity to obtain the second data.
[0009] Optionally, calculating the similarity between each substring vector in the first string set vector and the preset word vector; selecting the corresponding substrings according to the similarity to obtain the second data includes: Vectorize the preset first text and second text to obtain the corresponding first word vector and second word vector; vectorize each substring in the first string set to obtain the corresponding substring vector of each substring; Calculate the similarity between each substring vector and the first word vector to obtain the first similarity set; calculate the similarity between each substring vector and the second word vector to obtain the second similarity set; select the substrings in the interval of the maximum value in the first similarity set and the maximum value in the second similarity set to obtain the second data.
[0010] Optionally, the model training method further includes: designing prompt words according to the metaphor recognition task and the metaphor generation task respectively to obtain the preset prompt words; the preset prompt words include the first target prompt word and the second target prompt word; Concatenate the preset first prompt word, the first question and the first option information and input them into the language model for training to obtain the trained first language model; concatenate the preset second prompt word, the second question and the second option information and input them into the language model for training to obtain the trained second language model; Input the test set into the first language model, and calculate the first accuracy rate of the first language model in the metaphor recognition task and the metaphor generation task according to the accuracy rate calculation formula; input the test set into the second language model, and calculate the second accuracy rate of the second language model in the metaphor recognition task and the metaphor generation task according to the accuracy rate calculation formula; Adjust the first prompt word according to the first accuracy rate to obtain the third prompt word; adjust the second prompt word according to the second accuracy rate to obtain the fourth prompt word; Input the third prompt word into the first language model, and calculate the third accuracy rate of the first language model in the metaphor recognition task and the metaphor generation task; input the fourth prompt word into the second language model, and calculate the fourth accuracy rate of the second language model in the metaphor recognition task and the metaphor generation task; Obtain the first target prompt word according to the third accuracy rate; obtain the second target prompt word according to the fourth accuracy rate.
[0011] Optionally, train the pre-trained model according to the target training data to obtain a target model, including: Based on the target training data, train the pre-trained model to obtain a first model; Use the task instruction dataset to adjust the task to obtain a second model, and the parameters of the second model are different from those of the first model; Use hyperparameter search to optimize the learning rate and the number of training rounds of the second model to obtain a target model.
[0012] In a second aspect of the embodiments of the present disclosure, a model training device is provided, including: A first data processing unit, configured to input pre-training data into a first model for processing to obtain an output probability value, and screen the pre-training data according to a comparison result between the output probability value and a first threshold to obtain first data; the first model is a large language model; A second data processing unit, configured to integrate second data and first prompt word data to obtain target training data, where the second data is data obtained by preprocessing the first data, and the first prompt word data is a preset prompt word; A model training unit, configured to train a pre-trained model according to the target training data to obtain a target model.
[0013] In a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the steps of the above model training method are implemented.
[0014] In a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program, which when executed by a processor implements the steps of the above-described model training method.
[0015] The beneficial effects of the model training method, device, electronic device, and readable storage medium provided by the embodiments of the present disclosure are as follows: By constructing target training data, the problem of insufficient quality of metaphor corpus is solved. A stronger metaphorical expression basis is provided for the pre-trained model, making up for the deficiency of sparse metaphor knowledge in traditional pre-trained models. At the same time, the present disclosure improves the ability of the target model to recognize or generate metaphors, making up for the deficiencies of traditional models in complex reasoning and scenario adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0017] Figure 1 It is a schematic flowchart of the model training method provided by an embodiment of the present disclosure; Figure 2 It is a structural block diagram of the model training device provided by an embodiment of the present disclosure; Figure 3 It is a schematic block diagram of the electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.
[0019] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the following will be described in conjunction with the attached Figures 1 - 3 Illustrated by specific embodiments.
[0020] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of the model training method provided by an embodiment of the present disclosure, and the method includes: S101: Input the pre-training data into the first model for processing to obtain the output probability value. Screen the pre-training data according to the comparison result between the output probability value and the first threshold to obtain the first data. The first model is a large language model.
[0021] In this embodiment, the pre-training data contains noisy data. For example, the format is irregular, there are redundant punctuation marks, or content irrelevant to the target task. In order to improve the data quality and provide support for subsequent training, it is necessary to process the pre-trained data. Therefore, a strategy for cleaning the pre-training data using a large language model is designed.
[0022] Specifically, for the processing of pre-training data, conventional text processing methods are mostly used to remove obvious noisy data in the pre-training data and effectively clean up the messy format and irrelevant information in the pre-training data. For the Chinese metaphor generation / recognition task, whether there is a direct metaphor relationship in the pre-training data is also very important for the quality of the pre-training data. For example, for the sentence "The power of youth can be found everywhere, beautiful mountains and waters, good places, all roads are broad, and there is good wine for friends", although the content is coherent, there is no direct metaphor relationship. This part of the data can be regarded as noisy data for the metaphor generation / recognition task. Therefore, a cleaning strategy based on a large language model as a discriminator (LLMs as Judge) is designed. Use the large language model to judge whether there is a direct metaphor relationship in the text content of the pre-training data, so as to effectively screen the pre-training data.
[0023] In this embodiment, splice the preset discriminative prompt word with the pre-training data, and calculate the output probability value according to the output probability value calculation formula in the first model; The formula for calculating the output probability value is:
[0024] where, is the output probability value, is the activation function, W is the network weight, b is the network bias, pool represents the pooling layer, Encoder represents the Transformer encoding layer, is the embedding vector, is the preset discriminative prompt word, T i is the pre-training data.
[0025] Specifically, splice the preset discriminative prompt word with the training data, convert it into a vector through the embedding layer, and then sequentially pass through the Transformer encoding layer, pooling layer and feed-forward neural network to obtain the output probability value.
[0026] The given discriminative prompt word and pre-training data Concatenate into the input sequence Input = ; Map the concatenated input sequence to the corresponding embedding vectors through the embedding layer to obtain a matrix containing each word embedding ; The matrix containing each word embedding Pass through the Transformer encoding layer to obtain context vectors. Among them, Transformer processes the relationship between each input word and other words based on the self-attention mechanism, generating the context representation of each word; Pool the context vectors through the pooling layer to obtain a fixed-length text ; From Extract the output of the [CLS] token, denoted as ; According to the discrimination task (for example, the discrimination task is whether it contains a metaphor), input into the first model to obtain the output probability value; Judge whether it meets the goal of the discrimination task according to the output probability value; Among them, The pooled representation Will be linearly transformed through a fully connected layer W and a bias b:
[0027] Then the output probability value calculation formula is updated to
[0028] In this embodiment, the pre-training data is screened according to the comparison result between the output probability value and the first threshold to obtain the first data, including: If the output probability value is greater than the first threshold, the corresponding pre-training data is summarized to obtain the first data; If the output probability value is less than or equal to the first threshold, the corresponding pre-training data is cleared.
[0029] Specifically, if the output probability value is greater than the first threshold (0.5), we consider that the text contains metaphor information, that is, the output of the first model is yes, otherwise it is no. Through the output result (YES|NO) of the first model, the corpus data that does not contain metaphor relationships can be effectively filtered out; In this disclosure, the noise data is filtered out through the large language model, while retaining the information related to the metaphor generation and recognition tasks, improving the problem of insufficient metaphor corpus quality, providing a stronger metaphor expression basis for the pre-training model, and making up for the defect of sparse metaphor knowledge in the traditional pre-training model.
[0030] In one example, this disclosure introduces a dynamic threshold adjustment mechanism; Dynamically adjust the first threshold based on the distribution density of metaphorical relationships in the pre-training data; If the proportion of metaphorical sentences in the first data is less than the preset lower limit, then lower the first threshold; If the proportion of metaphorical sentences in the first data is greater than the preset upper limit, then raise the first threshold; F new = F - w 1 ∙ ∙ ∆F 1 where q < q lower ; F new = F + w 2 ∙ ∙ ∆F 2 where q > q upper ; F new = F, where q upper > q > q lower ; where F is the current first threshold, and F new is the adjusted first threshold; q is the proportion of metaphorical sentences in the first data obtained by statistics; q lower is the preset lower limit of the proportion of metaphorical sentences; q upper is the preset upper limit of the proportion of metaphorical sentences; ∆F 1 is the amount by which the threshold is decreased each time when the proportion of metaphorical sentences is below the lower limit, and ∆F 1 = f 1 (D), where D is the distribution density of metaphorical relationships in the pre-training data, and f 1 is a monotonically decreasing function of D, and f 1 (D) ; is a constant, is a coefficient greater than 0; ∆F 2 is the amount by which the threshold is increased each time when the proportion of metaphorical sentences is above the upper limit, and f 2 is a monotonically decreasing function of D, and f 2 (D) ; is a constant, is a coefficient greater than 0; w 1 is the weight coefficient when the proportion of metaphorical sentences is below the lower limit, and 0 < w 1 < 1; w 2 is the weight coefficient when the proportion of metaphorical sentences is above the upper limit, and 0 < w 2 < 1; is a regulation factor used to control the sensitivity of the adjustment.
[0031] The calculation formula of = × ×
[0032] Among them, B is the sample size of the first data (the scale of the current data set); is the reference value of the preset maximum sample size, which is used to normalize the data scale. When the sample size is large, is relatively small, and when the sample size is small, is relatively large; is the actual standard deviation of the proportion of metaphorical sentences in the first data, which is used to measure the data volatility; is the reference value of the average standard deviation, which can be determined according to historical data or experience. When the data volatility is large, is relatively large, and when the data volatility is small, is relatively small; is the evaluation value of the first data (the evaluation index value of the current data quality, such as accuracy, recall, etc., which are quality-related indicators); is the set data quality target value, which reflects the gap between the current data quality and the target quality. When the gap is large, is relatively large to more actively adjust the threshold to improve the data quality. When the gap is small, is relatively small; are the weights of the three factors of sample size, data volatility, and the gap between data quality and the target, respectively, and .
[0033] The dynamic adjustment of the first threshold in the present disclosure takes into account the influence of the distribution density of metaphorical relationships in the pre-training data on the adjustment amount of the first threshold, and more delicate control of the adjustment process through weight coefficients and adjustment factors to achieve more accurate adaptive data cleaning.
[0034] S102: Integrate the second data and the first prompt word data to obtain target training data. The second data is the data obtained by preprocessing the first data, and the first prompt word data is the preset prompt word.
[0035] In this embodiment, processing the first data to obtain the second data includes: processing the first data to obtain a first string set vector; calculating the similarity between each substring vector in the first string set vector and the first word vector and the second word vector respectively; selecting the corresponding substrings according to the similarity to obtain the second data (selecting the substrings in the first string set vector with a similarity greater than the preset similarity).
[0036] Specifically, the preset first text and second text are vectorized to obtain corresponding first word vectors and second word vectors; each substring in the first string set is vectorized to obtain substring vectors corresponding to the respective substrings. Calculate the similarity between each substring vector and the first word vector to obtain a first similarity set; calculate the similarity between each substring vector and the second word vector to obtain a second similarity set; select the substrings within the interval of the maximum value in the first similarity set and the maximum value in the second similarity set to obtain the second data.
[0037] More specifically, the text vectorization method (Term Frequency - Inverse Document Frequency, TF - IDF) is used to process the first data to obtain important features related to metaphors in the first data. Based on the Jieba word segmentation tool, the first data is segmented, and the original sentences in the first data are segmented into a substring set s = { s 1 , s 2 , ..., s n}; the preset tenor and vehicle are vectorized respectively to generate corresponding TF - IDF word vector representations and ; each substring in the sentence is vectorized into ; the semantic relevance between the tenor and vehicle and each substring in the sentence is quantified by calculating the cosine similarity; the argmax function is used to select the most relevant substring index; select and in the maximum interval, that is, select the substrings that are most relevant to the tenor and most relevant to the vehicle among the substrings to ensure the semantic coherence of the data.
[0038] The formula for calculating the similarity between a substring and the tenor is: sim The formula for calculating the similarity between a substring and the vehicle is: sim Use the argmax function to select the most relevant substring index:
[0039]
[0040] where the largest index, is such that The largest index.
[0041] In this embodiment, according to the requirements of the task, the Prompt Engineer method is used to design prompts, obtaining the preset prompts, and integrating the preset words with the second data to obtain the target training data.
[0042] In this embodiment, the (Prompt engineer) prompt engineering is used to design prompts corresponding to the task. By trying different prompt templates, the perception ability of the model for information in this field is activated. The designed prompts need to be able to guide the model to better understand the context of the data and effectively stimulate the model's perception and generation ability of metaphors.
[0043] In this embodiment, the design process of the prompts includes: In the first step, design the loss function between the output of the language model and the true label, that is, the goal of optimizing the training of the language model is the cross-entropy loss:
[0044] Among them, N is the number of samples in the training set, K is the number of options (4 in this scenario), is the sample at the j th option's true label. For the correct option = 1, for the wrong option, = 0. is the probability of the model's prediction of the l th option for the sample j This probability is calculated by the model and represents the probability that option j is selected.
[0045] In the second step, design the first prompt and the second prompt respectively according to the metaphor recognition task and the metaphor generation task; In the metaphor recognition task, the first prompt emphasizes comparison, identifying the ontology, the vehicle, and their commonalities: = "Given a Chinese sentence containing a metaphor about the ontology and the vehicle: \"(specific text)\". Please find the ontology, the vehicle, and the commonalities between the ontology and the vehicle in the sentence." In the metaphor generation task, the second prompt requires the second language model to generate relevant metaphors based on the given ontology, vehicle, and commonalities: ="Given an ontology, a vehicle, and a commonality, please write a short and accurate Chinese sentence containing a metaphor about the ontology and the vehicle based on the commonality between them." In the third step, splice the preset first prompt word, the first question, and the first option information together and input them into the language model for training to obtain the trained first language model; splice the preset second prompt word, the second question, and the second option information together and input them into the language model for training to obtain the trained second language model; ,
[0046] wherein, is the first language model, is the second language model.
[0047] In the fourth step, input the test set into the first language model and calculate the first accuracy rate of the first language model in the metaphor recognition task and the metaphor generation task according to the accuracy rate calculation formula; input the test set into the second language model and calculate the second accuracy rate of the second language model in the metaphor recognition task and the metaphor generation task according to the accuracy rate calculation formula; The accuracy rate calculation formula is:
[0048] wherein, M is the number of samples in the test set, is the prediction result of the model for the m th sample, is the m th true label of the sample; is the indicator function, which returns 1 when the predicted label and the true label output by the model ( and ) are equal, and returns 0 otherwise.
[0049] In the fifth step, adjust the first prompt word according to the first accuracy rate to obtain the third prompt word; adjust the second prompt word according to the second accuracy rate to obtain the fourth prompt word.
[0050] Specifically, perform version iteration and adjustment on the prompt word; starting from the initial version prompt word , for the accuracy rate of each model, that is and , perform manual adjustment to generate subsequent versions of the prompt word and so on.
[0051]
[0052] wherein, For the updated prompt, f is the prompt iterative adjustment mechanism, which improves and generates the updated prompt by comprehensively considering the prompt of the initial version and the accuracy rate to improve and generate the updated prompt , and the accuracy rate is
[0053] Step 6: Input the third prompt into the first language model to calculate the third accuracy rate of the first language model in the metaphor recognition task and the metaphor generation task; input the fourth prompt into the second language model to calculate the fourth accuracy rate of the second language model in the metaphor recognition task and the metaphor generation task; obtain the first target prompt according to the third accuracy rate; obtain the second target prompt according to the fourth accuracy rate.
[0054] Specifically, according to the calculated accuracy rate, select the target prompt of the version with the highest accuracy rate on the test set ;
[0055] where R is the language model, including which is the first language model and which is the second language model
[0056] S103: Train the pre-trained model according to the target training data to obtain the target model; the target model is used to recognize or generate metaphors.
[0057] In this embodiment, based on the target training data, the pre-trained model is trained to obtain the first model; the task is adjusted using the task instruction dataset to obtain the second model; the learning rate and the number of training rounds of the second model are optimized by hyperparameter search to obtain the target model.
[0058] Specifically, based on high-quality data, the model needs to be further continuously trained and supervised fine-tuning (SFT) to adapt to the specific requirements of the metaphor generation task and the metaphor recognition task. Select the basic model with the best performance from the existing pre-trained models (such as Qwen, LlaMa, Yi, etc.) as the starting point. These models have good text generation and semantic representation capabilities, providing a solid foundation for the metaphor generation task.
[0059] Use the target training data to continuously train the pre-trained model. The goal of this stage is to enhance the model's general knowledge understanding ability for the metaphor generation task, including: the semantic connection between the ontology and the vehicle and the language expression characteristics of metaphor generation; the loss function during the training of the pre-trained model is:
[0060] where G is the number of samples in the target training data )For a given input ) and model parameters , on the premise of predicting the next element as probability, is the parameter to be optimized by the model.
[0061] During the training process of the pre-trained model, it predicts subsequent content based on the previous information and minimizes the prediction error. This disclosure uses target training data (which contains a large number of metaphorical sentences and related language data) to help the pre-trained model understand the relationship between the ontology and the vehicle and the language characteristics of metaphors. Therefore, the pre-trained model obtains a more in-depth metaphor generation and metaphor recognition ability, laying a solid general language foundation for subsequent supervised fine-tuning (SFT) of the pre-trained model.
[0062] In this embodiment, a task instruction dataset is used to adjust the task to obtain a second model.
[0063] Specifically, the first model has been baselined on the target training data. This disclosure fine-tunes the first model to obtain a second model, further improving the performance of the model in downstream tasks.
[0064] Select an instruction fine-tuning dataset containing a specific multiple-choice question form , and perform fine-tuning for the two tasks of metaphor generation and metaphor recognition respectively; obtain the fine-tuned second model ( ).
[0065] The objective of training and optimizing the first model is the cross-entropy loss:
[0066] where W is the instruction fine-tuning dataset, E is the number of options, is the true label of the instruction fine-tuning data w for the e-th option. For the correct option = 1, for the wrong option, = 0; is the probability that the model predicts the e-th option for the instruction fine-tuning data w.
[0067] Among them, the instruction fine-tuning dataset. For example, for the metaphor generation task, given the instruction: ="Please select the most appropriate sentence from the following metaphor sentence options based on the ontology, vehicle, and their commonalities. Please answer directly with the corresponding option without any additional response. {question}{options}"; For the metaphor recognition task, the given instruction: ="Given a Chinese sentence containing a metaphor of an ontology and a vehicle: {question}. Please find the ontology, vehicle, and the commonalities between the ontology and the vehicle. {options}.".
[0068] In this embodiment, hyperparameter search is used to optimize the learning rate and number of training epochs of the second model to obtain the target model.
[0069] Specifically, hyperparameter search and optimization are the core links in model training, aiming to find the best training settings by exploring different configurations of training learning rates and epochs, thereby improving model performance. Common search methods include grid search, random search, and Bayesian optimization. These methods accelerate model convergence and improve generalization ability by trying multiple hyperparameter combinations.
[0070] For all possible hyperparameter combinations , the goal of hyperparameter search is to find the optimal hyperparameter combination ;
[0071] Among them, is the accuracy of the selected hyperparameter combination in the downstream multiple-choice task.
[0072] In this embodiment, in the hyperparameter search stage, the present disclosure uses a grid search strategy to search for the optimal hyperparameters and optimizes different combination configurations for the learning rate and number of training epochs.
[0073] Specifically, since the second model already has basic language capabilities in the previous training process, in order to prevent the second model from losing its basic performance, compared with pre-training, we adopt a lower learning rate training strategy (finally selected 2e-5), and adopt a higher number of training epochs (finally selected 10 epochs), so as to ensure that the model has the ability to solve the downstream multiple-choice task.
[0074] Among them, the searched learning rate and number of training epochs are as follows:
[0075] Among them, is the learning rate; is the number of training epochs.
[0076] In the pre-training stage, to accelerate the optimization process, a relatively large learning rate and fewer training epochs are adopted. A relatively large learning rate (such as ) helps to quickly adjust the model parameters, improve the search efficiency, and maintain stability through dynamic scheduling such as cosine annealing. In this stage, 1 training epoch is used to ensure efficient optimization. In the supervised fine-tuning stage, since the second model already has strong semantic understanding ability, more training epochs (such as 10 epochs) are adopted and the learning rate is appropriately reduced (such as ) to finely adjust the weights of the second model to obtain the target model, improving the performance of the model in metaphor generation and recognition tasks; through this configuration, the task adaptation ability of the model is further enhanced.
[0077] In one embodiment, the model training method of the present disclosure further includes testing and validating the target model in a downstream task.
[0078] Specifically, a validation data set is constructed; the ability of the target model in two tasks of metaphor generation and metaphor recognition is evaluated according to the validation data set; the question is used as the input of the target model, the option answered by the target model is used as the output, and the output option is compared with the correct answer to calculate the accuracy rate, and the final accuracy rate is used as the evaluation index.
[0079] In addition, to ensure the consistency between the ability of the target model and the actual application scenario, an artificial evaluation (human evaluation) method is adopted for evaluation, that is, metaphor-related questions are artificially proposed and their answering situations are evaluated, and the performance of the target model in metaphor generation and metaphor recognition tasks is comprehensively evaluated through the above methods.
[0080] It can be concluded from the above that the present disclosure solves the problem of insufficient quality of metaphor corpus by constructing target training data. It provides a stronger metaphor expression basis for the pre-trained model, makes up for the defect of sparse metaphor knowledge in traditional pre-trained models. At the same time, the present disclosure improves the generation quality and adaptation ability of the target model in metaphor tasks, making up for the deficiencies of traditional models in complex reasoning and scenario adaptability.
[0081] Corresponding to the model training method in the above embodiment, Figure 2 is the structural block diagram of a model training device provided by an embodiment of the present disclosure. For the sake of convenience of description, only the parts related to the embodiments of the present disclosure are shown. Referring to Figure 2 , the model training method device 20 includes: a first data processing unit 21, a second data processing unit 22, and a model training unit 23. Among them, the first data processing unit 21 is configured to input pre-training data into the first model for processing to obtain an output probability value, and screen the pre-training data according to the comparison result between the output probability value and the first threshold to obtain first data, where the first model is a large language model; The second data processing unit 22 is configured to integrate the second data and the first prompt data to obtain target training data, where the second data is data obtained by preprocessing the first data, and the first prompt data is a preset prompt; The model training unit 23 is configured to train the pre-trained model according to the target training data to obtain a target model.
[0082] In an embodiment of the present disclosure, the first data processing unit 21 is specifically configured to: Concatenate the preset discriminant prompt with the pre-training data and input it into the first model The first model calculates according to the output probability value calculation formula to obtain the output probability value; In this embodiment, the output probability value calculation formula is
[0083] Among them, is the output probability value, is the activation function, W is the network weight, b is the network bias, pool represents the pooling layer, Encoder represents the Transformer encoding layer, is the embedding vector, is the preset discriminant prompt, T i is the pre-training data.
[0084] In an embodiment of the present disclosure, the first data processing unit 21 is specifically configured to: if the output probability value is greater than the first threshold, summarize the corresponding pre-training data to obtain first data; if the output probability value is less than or equal to the first threshold, clear the corresponding pre-training data.
[0085] In an embodiment of the present disclosure, the second data processing unit 22 is specifically configured to: process the first data to obtain a first string set vector; calculate the similarity between each substring vector in the first string set vector and the first word vector and the second word vector respectively; select the corresponding substrings according to the similarity to obtain the second data.
[0086] In an embodiment of the present disclosure, the second data processing unit 22 is specifically configured to: vectorize the preset first text and second text to obtain the corresponding first word vector and second word vector; vectorize each substring in the first string set to obtain the substring vectors of the corresponding substrings; Calculate the similarity between each substring vector and the first word vector to obtain a first similarity set; calculate the similarity between each substring vector and the second word vector to obtain a second similarity set; select the substrings in the interval of the maximum value in the first similarity set and the maximum value in the second similarity set to obtain second data.
[0087] In an embodiment of the present disclosure, the second data processing unit 22 is specifically configured to: design prompt words according to the metaphor recognition task and the metaphor generation task respectively to obtain preset prompt words; the preset prompt words include a first target prompt word and a second target prompt word; Concatenate the preset first prompt word, the first question and the first option information together, and input them into the language model for training to obtain a trained first language model; concatenate the preset second prompt word, the second question and the second option information together, and input them into the language model for training to obtain a trained second language model; Input the test set into the first language model, and calculate the first accuracy rate of the first language model in the metaphor recognition task and the metaphor generation task according to the accuracy rate calculation formula; input the test set into the second language model, and calculate the second accuracy rate of the second language model in the metaphor recognition task and the metaphor generation task according to the accuracy rate calculation formula; Adjust the first prompt word according to the first accuracy rate to obtain a third prompt word; adjust the second prompt word according to the second accuracy rate to obtain a fourth prompt word; Input the third prompt word into the first language model, and calculate the third accuracy rate of the first language model in the metaphor recognition task and the metaphor generation task; input the fourth prompt word into the second language model, and calculate the fourth accuracy rate of the second language model in the metaphor recognition task and the metaphor generation task; Obtain the first target prompt word according to the third accuracy rate; obtain the second target prompt word according to the fourth accuracy rate.
[0088] In an embodiment of the present disclosure, the model training unit 23 is specifically configured to: train a pre-trained model based on target training data to obtain a first model; adjust the task using a task instruction data set to obtain a second model; optimize the learning rate and the number of training rounds of the second model by hyperparameter search to obtain a target model.
[0089] See Figure 3 , Figure 3 is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 3The electronic device 300 in the present embodiment shown may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store a computer program, and the computer program includes program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module / unit in the above device embodiments, for example Figure 2 the functions of the modules 21 to 22 shown.
[0090] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0091] The input device 302 may include a touchpad, a fingerprint sensor (for collecting the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.
[0092] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.
[0093] In specific implementation, the processors 301, input devices 302, and output devices 303 described in the embodiments of the present disclosure may execute the implementation manners described in the first and second embodiments of the xxxx detection method provided in the embodiments of the present disclosure, and may also execute the implementation manner of the electronic device described in the embodiments of the present disclosure, which will not be elaborated here.
[0094] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the method of the above embodiment are implemented. It can also be completed by instructing related hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0095] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the electronic device. Further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.
[0096] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.
[0097] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, the specific working processes of the above-described electronic devices and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0098] In several embodiments provided in the present application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces or units, and can also be electrical, mechanical or other forms of connection.
[0099] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present disclosure.
[0100] In addition, in each embodiment of the present disclosure, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0101] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present disclosure, and these modifications or substitutions should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A model training method, characterized in that: include: Inputting the pre-training data into the first model for processing to obtain an output probability value, and filtering the pre-training data according to a comparison result between the output probability value and a first threshold to obtain first data, wherein the first model is a large language model; Integrate the second data and the first prompt word data to obtain target training data, wherein the second data is data obtained by preprocessing the first data, and the first prompt word data is a preset prompt word; The pre-trained model is trained according to the target training data to obtain a target model; the target model is used to identify metaphors or generate metaphors.
2. The model training method according to claim 1, characterized in that: The step of inputting the pre-trained data into the first model for processing to obtain an output probability value includes: The preset discriminant prompt words are concatenated with the pre-trained data, and the output probability value is calculated according to the output probability value calculation formula in the first model; The output probability value calculation formula is: in, is the output probability value, is the activation function, W is the network weight, b is the network bias, pool represents the pooling layer, Encoder represents the Transformer encoding layer, is the embedding vector, is the preset discriminant prompt word, T i is the pre-training data.
3. The model training method according to claim 1, characterized in that: The step of screening the pre-training data according to the comparison result between the output probability value and the first threshold value to obtain the first data includes: If the output probability value is greater than the first threshold, the corresponding pre-training data is aggregated to obtain first data; If the output probability value is less than or equal to the first threshold, the corresponding pre-training data is cleared.
4. The model training method according to claim 1, characterized in that: Also includes: Processing the first data to obtain a first character string set vector; Calculate the similarity between each substring vector in the first string set vector and the first word vector and the second word vector respectively; A corresponding substring is selected according to the similarity to obtain second data.
5. The model training method according to claim 4, characterized in that: Calculating the similarity between each substring vector in the first string set vector and the preset word vector; selecting the corresponding substring according to the similarity to obtain the second data, including: Vectorize the preset first text and the second text to obtain corresponding first word vectors and second word vectors; vectorize each substring in the first string set to obtain corresponding substring vectors of each substring; Calculate the similarity between each substring vector and the first word vector to obtain a first similarity set; calculate the similarity between each substring vector and the second word vector to obtain a second similarity set; select a substring between the maximum value in the first similarity set and the maximum value in the second similarity set to obtain second data.
6. The model training method according to claim 1, characterized in that: Also includes: According to the metaphor recognition task and the metaphor generation task, prompt words are designed respectively to obtain preset prompt words; The preset prompt words include a first target prompt word and a second target prompt word; splicing the preset first prompt word, the first question and the first option information together, inputting them into the language model for training, and obtaining a trained first language model; splicing the preset second prompt word, the second question and the second option information together, inputting them into the language model for training, and obtaining a trained second language model; Inputting the test set into the first language model, and calculating the first accuracy of the first language model in the metaphor recognition task and the metaphor generation task according to the accuracy calculation formula; Input the test set into the second language model, and calculate the second accuracy of the second language model in the metaphor recognition task and the metaphor generation task according to the accuracy calculation formula; The first prompt word is adjusted according to the first accuracy rate to obtain a third prompt word; the second prompt word is adjusted according to the second accuracy rate to obtain a fourth prompt word; Input the third prompt word into the first language model, and calculate the third accuracy of the first language model in the metaphor recognition task and the metaphor generation task; Inputting the fourth prompt word into the second language model, and calculating the fourth accuracy of the second language model in the metaphor recognition task and the metaphor generation task; According to the third accuracy rate, the first target prompt word is obtained; A second target prompt word is obtained according to the fourth accuracy rate.
7. The model training method according to claim 1, characterized in that: The pre-trained model is trained according to the target training data to obtain the target model, including: Based on the target training data, the pre-trained model is trained to obtain a first model; Using the task instruction data set to adjust the task to obtain a second model, wherein the parameters of the second model are different from those of the first model; Hyperparameter search is used to optimize the learning rate and number of training rounds of the second model to obtain the target model.
8. A model training device, characterized in that: include: a first data processing unit, configured to input the pre-training data into a first model for processing to obtain an output probability value, and to screen the pre-training data according to a comparison result between the output probability value and a first threshold value to obtain first data, wherein the first model is a large language model; a second data processing unit, configured to integrate second data and first prompt word data to obtain target training data, wherein the second data is data obtained by preprocessing the first data, and the first prompt word data is a preset prompt word; The model training unit is used to train the pre-trained model according to the target training data to obtain the target model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.