Text feature extraction method, model training method and related device
By employing a text feature extraction method involving automatic annotation and multi-round fine-tuning, the problems of low efficiency in manual annotation and insufficient robustness of single training strategies in traditional large-scale model training are solved. This achieves efficient selection of training data and improvement of model performance, enhancing the model's adaptability and accuracy in multi-task scenarios.
Patent Information
- Application Number
- CN202511065898.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-14
AI Technical Summary
In traditional large model training, manual annotation is inefficient and costly, and a single training strategy is not robust enough in multi-task scenarios, making it difficult to achieve good task fusion.
We employ an automatic annotation technique combined with two fine-tuning and reinforcement learning optimization methods for text feature extraction. We use a pre-set text extraction model to annotate the initial dialogue text, construct a preference pair dataset, and use the LORA and DPO algorithms to optimize the base model, thereby improving the model's adaptability and robustness in multi-task scenarios.
It enables efficient selection of training data, reduces manual costs, improves the performance and accuracy of the model in complex and variable environments, and enhances the model's adaptability and robustness.
Smart Images

Figure CN120950666A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of machine learning, specifically relating to a text feature extraction method, a model training method, and related apparatus. Background Technology
[0002] Traditional large-scale model training and development can generally be divided into two main parts: data processing and model training. In the data processing stage, manual annotation is a common method. However, when the data volume is very large, the efficiency of manual annotation decreases significantly, and it is also prone to omissions and mislabeling, leading to inconsistencies in annotation quality. This makes accurate model training on large-scale datasets more difficult and costly. Furthermore, the manual annotation process requires a significant amount of time and human resources, further limiting its application in real-world production environments.
[0003] In model training, traditional methods typically employ a single training strategy for task optimization. However, in real-world applications, when faced with multi-task or multi-objective scenarios, a single training method often fails to address the complex relationships between tasks and the diverse needs of different objectives. Especially in the context of multi-task learning, a single strategy may lead to insufficient model robustness, hindering effective task fusion and consequently impacting the model's performance in practical applications. Therefore, traditional training strategies have significant limitations when handling complex multi-task scenarios. Summary of the Invention
[0004] This application provides a text feature extraction method, a model training method, and related apparatus to achieve automatic labeling of training data, efficient sample screening, and reduced manual costs. At the same time, by combining two fine-tuning processes and reinforcement learning optimization, the overall performance of the model and the adaptability of the training strategy are improved.
[0005] This application provides a training method for a text feature extraction model, including: Get multiple initial dialogue texts; Each initial dialogue text in the plurality of initial dialogue texts is labeled using a preset text extraction model to obtain labeling results, which include label data and summary data for each initial dialogue text. Based on the annotation results, target training sample data is determined from the plurality of initial dialogue texts, wherein the target training sample data includes at least one initial dialogue text and label data and summary data corresponding to the at least one initial dialogue text; The base model performs label data extraction and summary data extraction tasks on the target training dataset, and performs the first supervised fine-tuning of the base model using LORA. The base model after the first supervised fine-tuning is used to perform label data and summary data extraction tasks on the target training dataset, and the base model after the first supervised fine-tuning is then fine-tuned a second time using LORA. Construct a preference pair dataset, which includes multiple preference data pairs, each of which includes a better output text and a worse output text for an initial dialogue text; Based on the preferences, a reinforcement learning algorithm is used to optimize the base model after the second supervised fine-tuning on the dataset to obtain the target text feature extraction model.
[0006] This application also provides a text feature extraction method, including: Get the text to be extracted; The text to be extracted is input into the target text feature extraction model to obtain the keyword data and summary data output by the training sample annotation model. The text feature extraction model is trained based on the above-mentioned text feature extraction model training method.
[0007] This application also provides a training device for a text feature extraction model, comprising: The first acquisition unit is used to acquire multiple initial dialogue texts; The annotation unit is used to annotate each of the multiple initial dialogue texts using a preset text extraction model to obtain annotation results, which include the label data and summary data of each initial dialogue text. A determining unit is configured to determine target training sample data from the plurality of initial dialogue texts based on the annotation results, wherein the target training sample data includes at least one initial dialogue text and label data and summary data corresponding to the at least one initial dialogue text. The first fine-tuning unit is used to perform label data extraction tasks and summary data extraction tasks on the target training dataset through the base model, and to perform the first supervised fine-tuning of the base model. The second fine-tuning unit is used to perform label data and summary data extraction tasks on the target training dataset using the base model after the first supervised fine-tuning, and to perform a second supervised fine-tuning on the base model after the first supervised fine-tuning. A construction unit is used to construct a preference pair dataset, which includes multiple preference data pairs, each of which includes a better output text and a worse output text for an initial dialogue text; The optimization unit is used to optimize the base model after the second supervised fine-tuning on the dataset according to the preferences using a reinforcement learning algorithm, so as to obtain the target text feature extraction model.
[0008] This application also provides a text feature extraction apparatus, including: The second acquisition unit is used to acquire the text to be extracted; The extraction unit is used to input the text to be extracted into the target text feature extraction model to obtain the keyword data and summary data output by the training sample annotation model. The text feature extraction model is trained based on the above-mentioned text feature extraction model training method.
[0009] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a training method or a text feature extraction method as described above for any of the text feature extraction models.
[0010] This application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a training method or a text feature extraction method for any of the text feature extraction models described above.
[0011] This application also provides a computer program product, including a computer program that, when executed by a processor, implements a training method or a text feature extraction method for any of the text feature extraction models described above.
[0012] The text feature extraction method, model training method, and related apparatus provided in this application include a text feature extraction model training method comprising: firstly, acquiring multiple initial dialogue texts; then, labeling each initial dialogue text using a preset text extraction model to obtain labeling results, the labeling results including label data and summary data for each initial dialogue text; then, determining target training sample data from the multiple initial dialogue texts based on the labeling results, the target training sample data including at least one initial dialogue text and corresponding label data and summary data for the at least one initial dialogue text; and finally, performing label data extraction on the target training dataset using a base model. The process involves extracting task and summary data, and then performing a first supervised fine-tuning of the base model using LoRa. The base model, after this first supervised fine-tuning, is then used to perform label and summary data extraction tasks on the target training dataset. A second supervised fine-tuning is performed on the base model using LoRa. Next, a preference pair dataset is constructed, comprising multiple preference data pairs, each including a better output text and a worse output text for an initial dialogue text. Finally, a reinforcement learning algorithm is used to optimize the base model after the second supervised fine-tuning based on the preference pair dataset, resulting in a target text feature extraction model. This application significantly reduces reliance on manual intervention by introducing automatic annotation technology, lowering the complexity and cost of data processing. By automatically annotating the initial dialogue text using a pre-defined text extraction model, the annotation results include not only the text's label data but also summary data. This dual annotation provides richer contextual information for subsequent training tasks, enabling the model to understand and learn text features at a higher level.
[0013] During model training, this application innovatively employs a multi-round fine-tuning strategy. For the first time, supervised fine-tuning optimizes the base model, enabling it to separately complete the label and summary data extraction tasks on the target training dataset. This task-specific approach allows the model to capture features from different tasks more precisely, improving its task adaptability. Fine-tuning the model using LoRa technology not only optimizes its learning ability but also improves training efficiency and shortens the training cycle. Subsequently, further fine-tuning further enhances the model's adaptability to the target data, enabling it to more accurately identify and extract text features, thus improving the quality of the final output.
[0014] Furthermore, this application enhances the model's ability to handle complex text by constructing a preference pair dataset. The preference pair dataset creates multiple preference data pairs containing superior and inferior outputs by comparing different outputs of the initial dialogue text, and then optimizes the base model using reinforcement learning algorithms. During this process, the model can gradually adjust its output strategy based on feedback signals, continuously improving the accuracy and intelligence of text extraction. This reinforcement learning optimization strategy not only improves the model's adaptability but also enhances its robustness in diverse application scenarios, enabling the final text feature extraction model to maintain efficient and accurate performance even in complex and changing environments. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0016] Figure 1 This is one of the flowcharts illustrating a training method for a text feature extraction model provided in this application.
[0017] Figure 2 This is a schematic diagram of the model optimization process provided in this application.
[0018] Figure 3 This is a schematic diagram of a sample initial screening process provided in this application.
[0019] Figure 4 This is a schematic diagram of a sample screening process provided in this application.
[0020] Figure 5 This is a schematic diagram of a sample expansion process provided in this application.
[0021] Figure 6 This is the second flowchart illustrating a training method for a text feature extraction model provided in this application.
[0022] Figure 7 This is a flowchart illustrating a text feature extraction method provided in this application.
[0023] Figure 8 This is a block diagram of the functional units of a training device for a text feature extraction model provided in this application.
[0024] Figure 9 This is a block diagram of the functional units of a text feature extraction device provided in this application.
[0025] Figure 10This application provides a schematic diagram of the structure of an electronic device. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0027] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0028] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0029] Traditional large-scale model training and development typically involves two main parts: data processing and model training. Data processing often uses manual annotation, which is inefficient when dealing with very large datasets, and manual annotation is prone to omissions and errors. Common model training methods employ a single training strategy, which often lacks robustness in multi-task scenarios and struggles to achieve effective task fusion.
[0030] To address the aforementioned problems, embodiments of this application provide a text feature extraction method, a model training method, and related apparatus. The embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0031] Please see Figure 1 , Figure 1 This is one of the flowcharts illustrating a training method for a text feature extraction model provided in this application. The training method for the text feature extraction model includes the following steps.
[0032] S101, Obtain multiple initial dialogue texts.
[0033] The initial dialogue text can be transcribed using Automatic Speech Recognition (ASR), or extracted from a chat interface. This chat interface could be a customer service interface for loan services. For example, it could be the dialogue text between a user and customer service representatives after the user applies for a loan.
[0034] S102, Each initial dialogue text in the plurality of initial dialogue texts is labeled using a preset text extraction model to obtain the labeling results.
[0035] The annotation results include tag data and summary data for each initial dialogue text. Tag data is generated based on keywords or key phrases extracted from the initial dialogue text. These tags help classify, cluster, or analyze the text content, providing a foundation for semantic understanding in subsequent tasks. Summary data is a concise summary of the content of each initial dialogue text. It is usually generated by a model that extracts and summarizes the core information of the text to facilitate subsequent processing and rapid analysis. Summary data can be generated through semantic understanding and context modeling of the dialogue text, typically using natural language processing techniques such as extractive or generative summarization. A summary is not merely a simple condensation of the text; it should accurately capture the essence of the text, expressing the main information and intent of the dialogue. This is particularly important for subsequent applications such as information retrieval, recommendation systems, and sentiment analysis. The pre-set text extraction model can be a model capable of understanding and generating high-quality text content, such as the Qwen series models.
[0036] S103, determine the target training sample data from the plurality of initial dialogue texts based on the annotation results.
[0037] The target training sample data includes at least one initial dialogue text and corresponding label and summary data for that initial dialogue text. In other words, this scheme can filter the initial dialogue text based on the annotation results to obtain training data that can be used for model training.
[0038] S104, the base model performs label data extraction and summary data extraction tasks on the target training dataset, respectively, and performs the first supervised fine-tuning on the base model.
[0039] Specifically, dedicated prompts can be designed for the label data extraction task and the summary data extraction task to clearly distinguish the objectives of the two tasks. Then, LORA is used to fine-tune the base model on the datasets for each task under supervision to ensure that the basic capabilities meet the standards. This base model can be a large language model.
[0040] In practice, the goal of the label data extraction task is to extract keywords, phrases, or key entities from dialogue text, thereby supporting subsequent tasks such as text classification, topic modeling, and information retrieval. To ensure the model can accurately extract label data, dedicated prompts can be designed to guide the model to focus on key content in the text. For example, prompts could be phrases like "Please extract all sentiment-related keywords from the text" or "Please extract labels related to people, time, and place." By training on the labeled data of the target training dataset, the model can learn how to effectively label important information in the text.
[0041] The summary data extraction task focuses on condensing information from each dialogue text to generate concise and accurate summaries. The goal of this task is to extract the core information from longer texts for easier subsequent processing or rapid analysis. For this task, prompts can be designed, such as "Please generate a concise summary based on the following text" or "Please extract the main points and conclusions of the text." When performing the summary data extraction task, the model not only needs to understand the details of the text but also needs to generate fluent and natural summary content with the help of context.
[0042] LORA (Low-Rank Adaptation) is an efficient fine-tuning method that optimizes the model's training process by introducing a low-rank matrix, effectively improving the model's generalization ability on small datasets. During the training of label data extraction and summary data extraction tasks, LORA helps the model optimize the training data without changing a large number of parameters, allowing the model to focus more on learning the details of the target task.
[0043] The capabilities of the pedestal model can be effectively enhanced through the LoRa technique. After the first supervised fine-tuning, the pedestal model will show a significant improvement in performance on the target task, and will be able to extract label data and summary data more accurately, thereby providing high-quality training data and features for subsequent tasks.
[0044] S105, the base model after the first supervised fine-tuning is used to perform the label data and summary data extraction tasks on the target training dataset, and the base model after the first supervised fine-tuning is then subjected to a second supervised fine-tuning.
[0045] During the second fine-tuning, a unified prompt word can be designed, and this can be achieved through [various methods / mechanisms] in the prompt word construction process. <summary>、 <tagging>The task prefix identifier guides the model to output the results of two tasks at once.
[0046] S106, Construct the preference pair dataset.
[0047] The preference pair dataset includes multiple preference data pairs, each of which includes a better output text (chosen) and a worse output text (reject) for an initial dialogue text. The preference pair dataset can be constructed based on manual annotation results, online feedback, etc. For example, the preference pair dataset could be: { "prompt": "Summary of the conversation: User: My order hasn't shipped yet. Customer Service: We will process it for you as soon as possible." "chosen": "
Customer Request
Customer Service Response
Customer Request
Customer Service Response
[0048] S107, Based on the preferences, the base model after the second supervised fine-tuning is optimized using a reinforcement learning algorithm on the dataset to obtain the target text feature extraction model.
[0049] Among them, the reinforcement learning algorithm can be the Direct Preference Optimization (DPO) algorithm.
[0050] For example Figure 2 As shown, firstly, dedicated prompt words, namely prompt_summary and prompt_label, are designed for the summary extraction task and the label extraction task, respectively. Then, based on the corresponding context, the order of the two prompt words is fine-tuned (Supervised Fine-Tuning, SFT) using a Large Language Model (LLM). Next, multi-task mixed data is reorganized and sampled. Based on the mixed samples and the corresponding single prompt words (prompt_label_summary), SFT is performed based on LLM. Finally, based on manually labeled and feedback label and summary data, DPO reinforcement learning is performed on LLM to obtain the fine-tuned model.
[0051] As can be seen, this example can achieve automatic labeling of training data, efficient screening of samples, and reduced manual costs. At the same time, by combining two fine-tuning and reinforcement learning optimizations, it can improve the overall performance of the model and the adaptability of the training strategy.
[0052] In one possible embodiment, the step of annotating each initial dialogue text in the plurality of initial dialogue texts using a preset text extraction model to obtain annotation results includes: obtaining a first prompt word for each group of tag data and a second prompt word for each group of summary data; extracting tags for each initial sample based on the first prompt word using the preset text extraction model to obtain first tag data for each initial sample; extracting summaries for each initial sample based on the second prompt word using the preset text extraction model to obtain first summary data for each initial sample; obtaining annotation results based on the first tag data and the first summary data; and determining target training sample data from the plurality of initial dialogue texts based on the annotation results includes: obtaining a business keyword library and matching rules for the business keyword library; matching the plurality of initial dialogue texts with keywords in the business keyword library according to the matching rules to obtain matching results; and determining target training sample data from the plurality of initial dialogue texts based on the matching results and the annotation results.
[0053] The initial text can be filtered using a business keyword library and a pre-defined text extraction model. The matching rules of this business keyword library can be based on regular expressions to match sample data. Simultaneously, the pre-defined text extraction model extracts labels and summaries for each initial sample, and then determines the sample data that matches the pre-defined labels and summaries as the target training sample data. When using the text extraction model for annotation, closely related or conflicting labels and summaries can be grouped to avoid interference. Then, concise and clear prompts are designed for each group of labels and summaries, and their stability is optimized and verified.
[0054] In practical implementation, when designing prompts, it is necessary to clearly define the task content, such as which tags to extract and what the summary dimensions are. It also needs to include a description of the business significance of each tag and summary dimension, and specify the return format to facilitate subsequent extraction and organization.
[0055] For example Figure 3 As shown, on one side, the sample data (i.e., the initial dialogue text) is filtered using a business keyword library to obtain high-precision sample data. On the other side, the designed prompt words are labeled using a preset text extraction model (e.g., the Qwen3-14B model) to obtain labeled samples with high recall. Then, the samples filtered by the rules and the samples labeled by the model are further filtered to obtain the sample selection results, i.e., the selected training samples.
[0056] As can be seen, in this embodiment, the dual-path collaboration of model annotation and business keyword library rule filtering can simultaneously improve the high recall and high precision of target sample data, and finally output high-quality sample screening results.
[0057] In one possible embodiment, determining target training sample data from the plurality of initial dialogue texts based on the matching result and the annotation result includes: determining reference training sample data from the plurality of initial dialogue texts based on the matching result and the annotation result, wherein the reference training sample data includes at least one initial dialogue text and label data and summary data of the at least one initial dialogue text; obtaining a plurality of third prompt words for each set of label data and a plurality of fourth prompt words for each set of summary data; and performing text extraction on each initial dialogue text in the reference training text using any one of the plurality of third prompt words based on a plurality of different preset text extraction models. Line label extraction yields multiple second label data for each initial dialogue text; based on multiple different text extraction models, each initial dialogue text in the reference training text is summed up using any one of the multiple fourth prompt words to obtain multiple second summary data for each initial dialogue text; the multiple second label data corresponding to each initial dialogue text are compared to see if they are the same, and the multiple second summary data are compared to see if they are the same, to obtain a first comparison result corresponding to the second label data and a second comparison result corresponding to the second summary data; target training sample data is determined from the reference training sample data based on the first comparison result and the second comparison result.
[0058] After obtaining the initial screening sample set reference training samples based on the aforementioned content, closely related or conflicting labels and summary dimensions can be grouped for processing to avoid interference. Concise and clear prompts are designed for each group of labels and summary dimensions, and their stability is optimized and verified. Then, different preset text extraction models and prompts are combined simultaneously to generate the annotation results of the initial text in the reference sample data, and the final bias is determined by voting.
[0059] In the specific implementation, when determining the first comparison result, multiple label data can be directly compared to determine whether the multiple label data are the same. When determining the second comparison result, each second sub-segment data can be fed into an embedding model (such as BERT, Word2Vec, GloVe, FastText, etc.) for encoding processing to obtain a text vector sentence_vec. Then, the cosine similarity of the text vectors is calculated.
[0060] If the cosine similarity is greater than a preset value, such as 0.7, then the two sub-segments are considered to be the same; otherwise, they are considered to be different.
[0061] As can be seen, in this embodiment, the data is refined based on the annotation results of multiple text extraction models for the same initial text to obtain target sample data. This enhances the quality of the model training data, thereby improving the generalization ability and overall performance of subsequent models. Simultaneously, a larger-scale model is used for voting to quickly accumulate a batch of reliable samples, improving label accumulation efficiency.
[0062] In one possible embodiment, determining the target training sample data from the reference training sample data based on the first comparison result and the second comparison result includes: voting on each initial dialogue text in the reference training sample data according to the first comparison result and the second comparison result to obtain a first training sample set, a second training sample set, and a third training sample set. The first training sample set includes initial dialogue texts corresponding to multiple identical second label data and multiple identical second summary data. The second training sample set includes initial dialogue texts corresponding to multiple identical second label data and / or multiple identical second summary data. The third training sample set includes initial dialogue texts corresponding to multiple identical second label data and / or multiple identical second summary data. The initial dialogue text contains multiple different second label data and / or multiple different second summary data; it is determined that the target training sample data includes training data from the first training sample set; the sampling verification results of the label data and summary data corresponding to the initial text data included in the second training sample dataset are obtained, and if the verification results meet preset conditions, it is determined that the target training sample data includes training data from the second training sample set; the relabeled label data and relabeled summary data of each initial text data in the third training sample dataset are obtained to obtain a relabeled training sample set; it is determined that the target training sample data includes training data from the relabeled training sample set.
[0063] This can be achieved through tiered review based on the comparison results. For example, if multiple second-label data and multiple second-subtotal data are completely identical, the corresponding initial text is directly adopted. If there are partial differences, sampling is performed to ensure compliance. If the compliance rate is greater than a preset value (e.g., 90%), the text is adopted; otherwise, the prompt words are optimized and the process is rerun. If the text is completely different, manual re-annotation is performed.
[0064] For example Figure 4 The system uses three different pre-defined text extraction models and different prompt words to annotate interactive samples (i.e., the initial dialogue text) in the interactive text (i.e., the reference sample training set), resulting in multiple annotation results. A voting module then generates three different comparison results. If the results are completely identical, they are directly adopted during the review and added to the incremental sample set. If they are completely different, manual re-annotation is performed during the review, and then the sample is added to the incremental sample set. If they are partially identical, for example, if the voting result shows a 2:1 ratio, random sampling is performed for manual review. If the sampling accuracy is greater than 90%, the sample is added to the incremental sample set; otherwise, the prompt words are optimized, and the pre-annotation process is rerun for further optimization.
[0065] As can be seen, this embodiment utilizes a larger-scale model for voting, rapidly accumulating a batch of reliable samples and improving label accumulation efficiency. Simultaneously, setting different verification methods for different sets can improve sample screening efficiency and accuracy.
[0066] In one possible embodiment, determining the target training sample data from the reference training sample data based on the first comparison result and the second comparison result includes: determining preliminary training sample data from the reference training sample data based on the first comparison result and the second comparison result; acquiring a preset dialogue example and target label data and target summary data of the preset dialogue example; generating a fifth prompt word based on the preset dialogue example; generating at least one supplementary text based on the fifth prompt word using a dialogue generation model; acquiring third label data and third summary data of the at least one supplementary text; determining whether the third label data includes the target label data and whether the third summary data includes the target summary data; if the target label data and the target summary data are included, then generating the target training sample data based on the preliminary training sample data and the at least one supplementary text.
[0067] Among them, for example Figure 5 As shown, a small number of high-quality dialogue examples can be selected from a limited sample library. These examples clearly contain target tags and have a complete structure. Then, prompt words are constructed based on these examples. The prompt word design can include the following requirements: mimicking the style and core intent of the seed dialogue; randomly replacing non-critical points (such as irrelevant details like time and location); and mandating that the generated content must reflect the target tags and summaries. Then, dialogue samples are generated based on a large model (e.g., DeepSeek-R1, DS R1). The generated data is then fed back into the large model from the generated dialogue sample library. The model extracts target summaries / tags based on the constructed prompt words and checks whether the generated conversation data contains the target summaries / tags (for tags, directly compare the tags; for summaries, compare the cosine similarity value (sim) > 0.7 to consider them identical). If they are present, they are given to a human for sampling and verification to ensure the quality of the generated data. If the human sampling verification passes, the generated dialogue data is used as an incremental sample; if it fails, the prompt words are reconstructed for optimization and iteration.
[0068] As can be seen, in this embodiment, the problem of missing specific label and summary data is solved based on dialogue examples. By generating, filtering and validating to expand the high-quality sample library, the diversity and coverage of data can be improved, the adaptability of the model to different scenarios and tasks can be enhanced, and the accuracy and robustness of the model in label and summary generation tasks can be improved, thereby improving the performance and effectiveness of the model in practical applications.
[0069] In one possible embodiment, optimizing the base model after the second supervised fine-tuning using a reinforcement learning algorithm on the dataset according to the preference to obtain a training sample labeled model includes: obtaining a first encoded response of the base model to the better output text and a second encoded response of the output text; obtaining a third encoded response of the reference model to the better output text and a fourth encoded response of the worse output text; obtaining a first relative difference between the first encoded response and the third encoded response and a second relative difference between the second encoded response and the fourth encoded response; constructing a loss function based on the first relative difference and the second relative difference; and optimizing the base model based on the loss function to obtain a target text extraction model.
[0070] One approach is to use the DPO algorithm for reinforcement learning. The DPO algorithm is a preference-pair-based reinforcement learning method consisting of two parts: an optimized policy model and a reference policy model. It does not require external training of the reference policy model; instead, it directly adjusts the model parameters by optimizing the objective function. Its loss function is as follows:
[0071] Where x is the model input, i.e., the preference data pair. It is the encoded response of chosen data. It is the encoded response to the rejected data, and D is the entire preference dataset. The strategy model to be optimized (i.e., the base model) has the following parameters: Updated through training This makes the generated results more consistent with the preferred samples. To reference the strategy model, the parameters are not updated, thus helping to evaluate the model. Is the output of the reference model better than that of the reference model? They should be good. It is the sigmoid function. Convert it to a probability form and calculate the negative logarithmic probability. It is a hyperparameter used to control the intensity of preference. At that time, it will amplify the difference in preferences, causing the model to tend to output preferred samples.
[0072] in,
[0073] Indicates the current optimization model For chosen data The relative difference between the response of the model and the reference model, if give A higher probability results in a positive value, encouraging the model to generate a chosen response.
[0074] in,
[0075] Indicates the current optimization model For rejected data The relative difference between the response of the model and the reference model, if give If the probability is low, the value is negative, which inhibits the model from generating a reject response.
[0076] As can be seen, in this embodiment, the DPO algorithm guides the model to learn preferred data by comparing the scores of the current optimized model and the reference model for the choose and reject responses. This alleviates the "repeating effect" of heavy summarization tasks and enhances the robustness of imbalanced samples.
[0077] Please see Figure 6 ,like Figure 6 As shown, the specific implementation steps of the text feature extraction model training method can be as follows: First, obtain the dialogue text transcribed from ASR, then proceed to the data accumulation step (i.e., obtain target sample training data). This data accumulation step includes initial sample screening (i.e., obtain reference training sample data), followed by fine-tuning (i.e., obtain target training sample data), and also includes small-sample expansion (i.e., obtain supplementary training sample data). Then, during multi-task collaborative training, perform double-prompt word fine-tuning (i.e., the first fine-tuning) and single-prompt word fine-tuning (i.e., the second fine-tuning) based on the target training sample data. Finally, perform reinforcement learning based on the DPO algorithm to obtain the target text feature extraction model, and then deploy the model online.
[0078] Please see Figure 7 This application also provides a text feature extraction method, including the following steps: S701, Get the text to be extracted.
[0079] The text to be extracted can be dialogue text.
[0080] S702, the text to be extracted is input into the target text feature extraction model to obtain the keyword data and summary data output by the training sample annotation model.
[0081] The text feature extraction model is trained based on the training method of the text feature extraction model in the above embodiments.
[0082] As can be seen, in this embodiment, the text to be extracted is first obtained, and then the text to be extracted is input into the target text feature extraction model to obtain the keyword data and summary data output by the training sample annotation model. The text feature extraction model is trained based on the training method of the text feature extraction model in the above embodiment. This can improve the accuracy and efficiency of text feature extraction.
[0083] The following describes a training device for a text feature extraction model provided in this application. The training device for the text feature extraction model described below corresponds to the training method for the text feature extraction model described above.
[0084] Please see Figure 8 The training device 800 for the text feature extraction model includes: a first acquisition unit 801, used to acquire multiple initial dialogue texts; an annotation unit 802, used to annotate each of the multiple initial dialogue texts using a preset text extraction model to obtain annotation results, the annotation results including label data and summary data of each initial dialogue text; a determination unit 803, used to determine target training sample data from the multiple initial dialogue texts based on the annotation results, the target training sample data including at least one initial dialogue text and label data and summary data corresponding to the at least one initial dialogue text; and a first fine-tuning unit 804, used to perform label data extraction on the target training dataset using a base model. The system includes a task and summary data extraction task, and performs a first supervised fine-tuning on the base model; a second fine-tuning unit 805, which performs a label data and summary data extraction task on the target training dataset using the base model after the first supervised fine-tuning, and performs a second supervised fine-tuning on the base model after the first supervised fine-tuning; a construction unit 806, which constructs a preference pair dataset, which includes multiple preference data pairs, each of which includes a better output text and a worse output text for an initial dialogue text; and an optimization unit 807, which optimizes the base model after the second supervised fine-tuning using a reinforcement learning algorithm based on the preference pair dataset to obtain a target text feature extraction model.
[0085] In one possible embodiment, regarding the step of annotating each initial dialogue text in the plurality of initial dialogue texts using a preset text extraction model to obtain annotation results, the annotation unit 802 is specifically used to: acquire a first prompt word for each group of tag data and a second prompt word for each group of summary data; extract tags for each initial sample based on the first prompt word using the preset text extraction model to obtain first tag data for each initial sample; extract summaries for each initial sample based on the second prompt word using the preset text extraction model to obtain first summary data for each initial sample; and obtain annotation results based on the first tag data and the first summary data. Regarding the step of determining target training sample data from the plurality of initial dialogue texts based on the annotation results, the determining unit 803 is specifically used to: acquire a business keyword library and matching rules for the business keyword library; match the plurality of initial dialogue texts with keywords in the business keyword library according to the matching rules to obtain matching results; and determine target training sample data from the plurality of initial dialogue texts based on the matching results and the annotation results.
[0086] In one possible embodiment, regarding the determination of target training sample data from the plurality of initial dialogue texts based on the matching result and the annotation result, the determining unit 803 is specifically configured to: determine reference training sample data from the plurality of initial dialogue texts based on the matching result and the annotation result, the reference training sample data including at least one initial dialogue text and label data and summary data of the at least one initial dialogue text; respectively acquire a plurality of third prompt words for each set of label data and a plurality of fourth prompt words for each set of summary data; and based on a plurality of different preset text extraction models, respectively use any one of the plurality of third prompt words to extract each of the reference training texts. Tag extraction is performed on the initial dialogue text to obtain multiple second tag data for each initial dialogue text. Based on multiple different text extraction models, a summary extraction is performed on each initial dialogue text in the reference training text using any one of the multiple fourth prompt words to obtain multiple second summary data for each initial dialogue text. The multiple second tag data corresponding to each initial dialogue text are compared to see if they are the same, and the multiple second summary data are compared to see if they are the same, to obtain a first comparison result corresponding to the second tag data and a second comparison result corresponding to the second summary data. Target training sample data is determined from the reference training sample data based on the first comparison result and the second comparison result.
[0087] In one possible embodiment, regarding the determination of target training sample data from the reference training sample data based on the first comparison result and the second comparison result, the determining unit 803 is specifically configured to: vote on each initial dialogue text in the reference training sample data based on the first comparison result and the second comparison result to obtain a first training sample set, a second training sample set, and a third training sample set. The first training sample set includes initial dialogue texts corresponding to multiple identical second label data and multiple identical second summary data. The second training sample set includes initial dialogue texts corresponding to multiple identical second label data and / or multiple identical second summary data. The third training sample set... The set includes multiple different second label data and / or multiple different second summary data corresponding to the initial dialogue text; it is determined that the target training sample data includes the training data in the first training sample set; the sampling verification results of the label data and summary data corresponding to the initial text data included in the second training sample dataset are obtained, and if the verification results meet the preset conditions, it is determined that the target training sample data includes the training data in the second training sample set; the re-labeled label data and re-labeled summary data of each initial text data in the third training sample dataset are obtained to obtain the re-labeled training sample set; it is determined that the target training sample data includes the training data in the re-labeled training sample set.
[0088] In one possible embodiment, regarding the determination of target training sample data from the reference training sample data based on the first comparison result and the second comparison result, the determining unit 803 is specifically configured to: determine preliminary training sample data from the reference training sample data based on the first comparison result and the second comparison result; acquire a preset dialogue example and target label data and target summary data of the preset dialogue example; generate a fifth prompt word based on the preset dialogue example; generate at least one supplementary text based on the fifth prompt word using a dialogue generation model; acquire third label data and third summary data of the at least one supplementary text; determine whether the third label data includes the target label data and whether the third summary data includes the target summary data; if the target label data and the target summary data are included, then generate target training sample data based on the preliminary training sample data and the at least one supplementary text.
[0089] In one possible embodiment, in optimizing the base model after the second supervised fine-tuning using a reinforcement learning algorithm on the dataset according to the preferences to obtain a training sample labeled model, the optimization unit 807 is specifically used to: obtain the base model's first encoding response to the better output text and the second encoding response to the output text; obtain the reference model's third encoding response to the better output text and the fourth encoding response to the worse output text; obtain the first relative difference between the first encoding response and the third encoding response and the second relative difference between the second encoding response and the fourth encoding response; construct a loss function based on the first relative difference and the second relative difference; and optimize the base model based on the loss function to obtain a target text extraction model.
[0090] The following describes a text feature extraction device provided in this application. The text feature extraction device described below corresponds to the text feature extraction method described above.
[0091] Please see Figure 9 The text feature extraction device 900 includes: a second acquisition unit 901, used to acquire the text to be extracted; and an extraction unit 902, used to input the text to be extracted into a target text feature extraction model to obtain keyword data and summary data output by the training sample annotation model. The text feature extraction model is trained based on the training method of the text feature extraction model described in the above embodiment.
[0092] It is understood that since the method embodiments and the device embodiments are different presentations of the same technical concept, the content of the method embodiment section in this application should be adapted to the device embodiment section in a synchronous manner, and will not be repeated here.
[0093] Please see Figure 10 , Figure 10 This is a schematic diagram of the structure of the electronic device provided in this application. For example... Figure 10 As shown, the electronic device may include a processor 1010, a communications interface 1020, a memory 1030, and a communication bus 1040. The processor 1010, communications interface 1020, and memory 1030 communicate with each other via the communication bus 1040. The processor 1010 can call logical instructions in the memory 1030 to execute a training method for the text feature extraction model or a text feature extraction method. The training method for this text feature extraction model includes: acquiring multiple initial dialogue texts; annotating each of the multiple initial dialogue texts using a preset text extraction model to obtain annotation results, the annotation results including label data and summary data for each initial dialogue text; determining target training sample data from the multiple initial dialogue texts based on the annotation results, the target training sample data including at least one initial dialogue text and corresponding label data and summary data for the at least one initial dialogue text; performing label data extraction and summary data extraction tasks on the target training dataset using a base model, and performing a first supervised fine-tuning on the base model; performing label data and summary data extraction tasks on the target training dataset using the base model after the first supervised fine-tuning, and performing a second supervised fine-tuning on the base model after the first supervised fine-tuning; constructing a preference pair dataset, the preference pair dataset including multiple preference data pairs, each preference data pair including a better output text and a worse output text for an initial dialogue text; and optimizing the base model after the second supervised fine-tuning using a reinforcement learning algorithm based on the preference pair dataset to obtain the target text feature extraction model.
[0094] The text feature extraction method includes: acquiring the text to be extracted; inputting the text to be extracted into a target text feature extraction model to obtain keyword data and summary data output by the training sample annotation model, wherein the text feature extraction model is trained based on the training method of the text feature extraction model described in the above embodiment.
[0095] Furthermore, the logical instructions in the aforementioned memory 1030 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0096] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a training method or a text feature extraction method for the text feature extraction model provided in the above embodiments. The training method for this text feature extraction model includes: acquiring multiple initial dialogue texts; annotating each of the multiple initial dialogue texts using a preset text extraction model to obtain annotation results, the annotation results including label data and summary data for each initial dialogue text; determining target training sample data from the multiple initial dialogue texts based on the annotation results, the target training sample data including at least one initial dialogue text and corresponding label data and summary data for the at least one initial dialogue text; performing label data extraction and summary data extraction tasks on the target training dataset using a base model, and performing a first supervised fine-tuning on the base model; performing label data and summary data extraction tasks on the target training dataset using the base model after the first supervised fine-tuning, and performing a second supervised fine-tuning on the base model after the first supervised fine-tuning; constructing a preference pair dataset, the preference pair dataset including multiple preference data pairs, each preference data pair including a better output text and a worse output text for an initial dialogue text; and optimizing the base model after the second supervised fine-tuning using a reinforcement learning algorithm based on the preference pair dataset to obtain the target text feature extraction model.
[0097] The text feature extraction method includes: acquiring the text to be extracted; inputting the text to be extracted into a target text feature extraction model to obtain keyword data and summary data output by the training sample annotation model, wherein the text feature extraction model is trained based on the training method of the text feature extraction model described in the above embodiment.
[0098] In another aspect, this application also provides a computer program product, including a computer program that, when executed by a processor, implements a training method or a text feature extraction method for any of the text feature extraction models described above. The training method for this text feature extraction model includes: acquiring multiple initial dialogue texts; annotating each of the multiple initial dialogue texts using a preset text extraction model to obtain annotation results, the annotation results including label data and summary data for each initial dialogue text; determining target training sample data from the multiple initial dialogue texts based on the annotation results, the target training sample data including at least one initial dialogue text and corresponding label data and summary data for the at least one initial dialogue text; performing label data extraction and summary data extraction tasks on the target training dataset using a base model, and performing a first supervised fine-tuning on the base model; performing label data and summary data extraction tasks on the target training dataset using the base model after the first supervised fine-tuning, and performing a second supervised fine-tuning on the base model after the first supervised fine-tuning; constructing a preference pair dataset, the preference pair dataset including multiple preference data pairs, each preference data pair including a better output text and a worse output text for an initial dialogue text; and optimizing the base model after the second supervised fine-tuning using a reinforcement learning algorithm based on the preference pair dataset to obtain the target text feature extraction model.
[0099] The text feature extraction method includes: acquiring the text to be extracted; inputting the text to be extracted into a target text feature extraction model to obtain keyword data and summary data output by the training sample annotation model, wherein the text feature extraction model is trained based on the training method of the text feature extraction model described in the above embodiment.
[0100] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0101] This application also provides a computer storage medium storing a computer program for electronic data interchange, which causes a computer to perform some or all of the steps of any of the methods described in the above method embodiments, wherein the computer includes an electronic device.
[0102] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0103] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0104] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.
[0105] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0106] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0107] If the aforementioned integrated units are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0108] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage device, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0109] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.< / tagging> < / summary>
Claims
1. A training method for a text feature extraction model, characterized in that, include: Get multiple initial dialogue texts; Each initial dialogue text in the plurality of initial dialogue texts is labeled using a preset text extraction model to obtain labeling results, which include label data and summary data for each initial dialogue text. Based on the annotation results, target training sample data is determined from the plurality of initial dialogue texts, wherein the target training sample data includes at least one initial dialogue text and label data and summary data corresponding to the at least one initial dialogue text; The base model is used to perform label data extraction and summary data extraction tasks on the target training dataset, and the base model is then subjected to the first supervised fine-tuning. The base model after the first supervised fine-tuning is used to perform label data and summary data extraction tasks on the target training dataset, and the base model after the first supervised fine-tuning is then subjected to a second supervised fine-tuning. Construct a preference pair dataset, which includes multiple preference data pairs, each of which includes a better output text and a worse output text for an initial dialogue text; Based on the preferences, a reinforcement learning algorithm is used to optimize the base model after the second supervised fine-tuning on the dataset to obtain the target text feature extraction model.
2. The method according to claim 1, characterized in that, The step of annotating each initial dialogue text in the plurality of initial dialogue texts using a preset text extraction model to obtain annotation results includes: Obtain the first prompt word for each group of tag data and the second prompt word for each group of summary data; The preset text extraction model extracts labels from each initial sample based on the first prompt word to obtain the first label data of each initial sample. The preset text extraction model extracts a summary of each initial sample based on the second prompt word to obtain the first summary data of each initial sample. The annotation results are obtained based on the first label data and the first summary data; The step of determining the target training sample data from the plurality of initial dialogue texts based on the annotation results includes: Obtain the business keyword library and its matching rules; According to the matching rules, the multiple initial dialogue texts are matched with keywords in the business keyword library to obtain matching results; Target training sample data is determined from the plurality of initial dialogue texts based on the matching results and the annotation results.
3. The method according to claim 2, characterized in that, The step of determining target training sample data from the plurality of initial dialogue texts based on the matching results and the annotation results includes: Reference training sample data is determined from the plurality of initial dialogue texts based on the matching results and the annotation results. The reference training sample data includes at least one initial dialogue text and the label data and summary data of the at least one initial dialogue text. Obtain multiple third prompt words for each group of tag data and multiple fourth prompt words for each group of summary data; Based on multiple different preset text extraction models, each initial dialogue text in the reference training text is labeled using any one of the multiple third prompt words, resulting in multiple second label data for each initial dialogue text; Based on multiple different text extraction models, each initial dialogue text in the reference training text is extracted by using any one of the multiple fourth prompt words, resulting in multiple second summary data for each initial dialogue text; The first comparison result corresponding to the second tag data and the second comparison result corresponding to the second sub-segment data are obtained by comparing whether the multiple second tag data and the multiple second sub-segment data are the same for each initial dialogue text. The target training sample data is determined from the reference training sample data based on the first comparison result and the second comparison result.
4. The method according to claim 3, characterized in that, The step of determining the target training sample data from the reference training sample data based on the first comparison result and the second comparison result includes: Based on the first comparison result and the second comparison result, each initial dialogue text in the reference training sample data is voted on to obtain a first training sample set, a second training sample set, and a third training sample set. The initial dialogue texts included in the first training sample set have multiple identical second label data and multiple identical second sub-data. The initial dialogue texts included in the second training sample set have multiple identical second label data and / or multiple identical second sub-data. The initial dialogue texts included in the third training sample set have multiple different second label data and / or multiple different second sub-data. The target training sample data is determined to include the training data in the first training sample set; Obtain the sampling verification results of the label data and summary data corresponding to the initial text data included in the second training sample dataset. If the verification results meet the preset conditions, it is determined that the target training sample data includes the training data in the second training sample set. Obtain the relabeled label data and relabeled summary data of each initial text data in the third training sample dataset to obtain the relabeled training sample set; determine that the target training sample data includes the training data in the relabeled training sample set.
5. The method according to claim 3 or 4, characterized in that, The step of determining the target training sample data from the reference training sample data based on the first comparison result and the second comparison result includes: Preparatory training sample data is determined from the reference training sample data based on the first comparison result and the second comparison result; Obtain preset dialogue examples and their target tag data and target summary data; The fifth prompt word is generated based on a preset dialogue example; The dialogue generation model generates at least one supplementary text based on the fifth prompt word; Obtain the third tag data and third summary data of the at least one supplementary text; Determine whether the third tag data includes the target tag data, and whether the third sub-summary data includes the target sub-summary data; If the target label data and the target summary data are included, then target training sample data are generated based on the pre-training sample data and the at least one supplementary text.
6. The method according to claim 1, characterized in that, The step of optimizing the base model after the second supervised fine-tuning using a reinforcement learning algorithm on the dataset according to the preferences to obtain a training sample labeled model includes: Obtain the first encoded response of the base model to the better output text and the second encoded response to the output text; Obtain the third encoded response of the reference model to the better output text and the fourth encoded response to the worse output text; Obtain a first relative difference between the first encoded response and the third encoded response, and a second relative difference between the second encoded response and the fourth encoded response; A loss function is constructed based on the first relative difference and the second relative difference; The base model is optimized based on the loss function to obtain the target text extraction model.
7. A text feature extraction method, characterized in that, include: Get the text to be extracted; The text to be extracted is input into the target text feature extraction model to obtain the keyword data and summary data output by the training sample annotation model. The text feature extraction model is trained based on the training method of the text feature extraction model according to any one of claims 1-6.
8. A training device for a text feature extraction model, characterized in that, include: The first acquisition unit is used to acquire multiple initial dialogue texts; The annotation unit is used to annotate each of the multiple initial dialogue texts using a preset text extraction model to obtain annotation results, which include the label data and summary data of each initial dialogue text. A determining unit is configured to determine target training sample data from the plurality of initial dialogue texts based on the annotation results, wherein the target training sample data includes at least one initial dialogue text and label data and summary data corresponding to the at least one initial dialogue text. The first fine-tuning unit is used to perform label data extraction tasks and summary data extraction tasks on the target training dataset through the base model, and to perform the first supervised fine-tuning of the base model. The second fine-tuning unit is used to perform label data and summary data extraction tasks on the target training dataset using the base model after the first supervised fine-tuning, and to perform a second supervised fine-tuning on the base model after the first supervised fine-tuning. A construction unit is used to construct a preference pair dataset, which includes multiple preference data pairs, each of which includes a better output text and a worse output text for an initial dialogue text; The optimization unit is used to optimize the base model after the second supervised fine-tuning on the dataset according to the preferences using a reinforcement learning algorithm, so as to obtain the target text feature extraction model.
9. A text feature extraction device, characterized in that, include: The second acquisition unit is used to acquire the text to be extracted; The extraction unit is used to input the text to be extracted into the target text feature extraction model to obtain the keyword data and summary data output by the training sample annotation model. The text feature extraction model is trained based on the training method of the text feature extraction model according to any one of claims 1-6.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the training method of the text feature extraction model as described in any one of claims 1 to 6, or the text feature extraction method as described in claim 7.
Citation Information
Cited By
Dialogue text analysis method and device, electronic equipment and computer readable medium
CN122133848A