Method for carrying out question and answer pair scoring based on reward mechanism to realize large model fine tuning
By introducing a reinforcement learning model based on the reward mechanism in the Q&A pair score, combining semantic consistency, keyword coverage and business relevance, the problem of single scoring mechanism in the existing technology is solved, high-quality screening of Q&A pairs and effective fine-tuning of large models is achieved, and the comprehensive performance of the intelligent customer service system is improved.
Patent Information
- Application Number
- CN202510043512.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing Q&A scoring method has a single scoring mechanism when dealing with complex business scenarios, and cannot comprehensively consider semantic consistency, keyword coverage and business relevance, which makes it difficult for the screened Q&A to meet the deep-seated needs of a specific business.
The Q&A pair scoring method based on the reward mechanism is adopted. The reinforcement learning model combines reward functions in three dimensions: semantic consistency, keyword coverage and business correlation, and the Q&A pair is scored, and the dynamic knowledge base update and pre-trained language model fine-tuning is adapted to specific business scenarios.
It realizes the high semantic integrity and business relevance of Q&A pairs, improves the fine-tuning training effect of the big model, and enhances the service capabilities and user experience of the intelligent customer service system.
Smart Images

Figure CN119940467A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and specifically to a method for fine-tuning a large model by scoring question and answer pairs based on a reward mechanism. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, large language models (such as GPT, BERT, etc.) have become important tools in the field of intelligent customer service. In order to improve the adaptability of the model in specific business scenarios, the question-answer pair scoring method has gradually been applied to build high-quality training data sets, thereby achieving fine-tuning of large models. This method significantly improves the service capabilities of the intelligent customer service system by screening key data from large-scale question-answer data and optimizing the model in a targeted manner based on business needs.
[0003] Most existing question-answer pair scoring methods are based on rule matching, statistical learning, or single indicator optimization. They assign quality scores to question-answer pairs by calculating indicators such as the similarity and keyword coverage of question-answer pairs. This method has the advantages of high screening efficiency and simple implementation, and performs well in general scenarios. In addition, some methods combine shallow machine learning algorithms to improve the ability to analyze the matching of questions and answers during the question-answer pair scoring process, which improves the generalization performance of the model after fine-tuning to a certain extent.
[0004] However, existing question-answer pair scoring methods still have limitations when dealing with complex business scenarios. Their scoring mechanism is relatively simple and cannot comprehensively consider multi-dimensional factors such as semantic consistency, keyword coverage, and business relevance, resulting in the screened question-answer pairs being difficult to meet the deep-level needs of specific businesses. In addition, existing methods do not provide sufficient support for the dynamic update of the knowledge base and the specific business adaptability of the model, and cannot effectively meet diverse business scenarios, limiting the application potential of large models in intelligent customer service systems. Summary of the invention
[0005] In view of the shortcomings of the existing technology, the present invention provides a method for fine-tuning a large model by scoring question and answer pairs based on a reward mechanism, which solves the problems in the existing technology that the question and answer scoring mechanism is single and cannot comprehensively consider multi-dimensional factors such as semantic consistency, keyword coverage and business relevance.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method for fine-tuning a large model by scoring question-answer pairs based on a reward mechanism, comprising the following steps: S1. Filter the initial question and answer pairs from the historical chat records of the manual customer service; S2. Score the initial question-answer pairs based on the reward mechanism; S3. Filter high-scoring question-answer pairs and update the business question-answer pair knowledge base; S4. Fine-tune the large model based on the updated business question and answer knowledge base; S5. Use the fine-tuned large model to perform intelligent customer service question-and-answer tasks.
[0007] Preferably, the S1 comprises: Perform data cleaning and preprocessing on historical chat records to remove invalid question and answer pairs; Use the semantic model to calculate the semantic similarity between the question and the answer, and select semantically complete question-answer pairs; Based on the business keyword set, count the number of keywords contained in the question-answer pairs, and filter out question-answer pairs with higher keyword matching degrees; sort the question-answer pairs according to the weighted comprehensive score of semantic similarity and keyword matching degrees, and select the question-answer pairs with the highest score rankings.
[0008] Preferably, S2 includes: Convert the question-answer pair into a semantic feature vector as the state input of the reinforcement learning model; The action in reinforcement learning is defined as the score of the question-answer pair. The reward function includes the following parts: Semantic consistency reward: measures the semantic similarity between the question and the answer; Keyword coverage reward: measures the number of business keywords contained in the question and answer pair; Business relevance rewards: Evaluate the matching degree of question and answer pairs based on specific business scenario rules; Train a reinforcement learning scoring model and optimize the scoring strategy based on the reward function.
[0009] Preferably, said S2 further comprises: The specific form of the reward function r is: r=w 1 ·r consistency +w 2 ·r k eywor d +w 3 ·r re l evance where r consistency represents the semantic consistency reward, r keyword represents the keyword coverage reward, r relevance It represents the business relevance reward, w1, w2, w3 are weight parameters.
[0010] Preferably, S3 includes: Filter question-answer pairs whose reward scores are greater than a preset threshold; Add the selected high-scoring question-answer pairs to the business question-answer pair knowledge base; Check the knowledge base for duplication and redundancy of business questions and answers to ensure data consistency and integrity.
[0011] Preferably, S4 includes: Build a question-answer pair dataset for fine-tuning training. The data source is the updated business question-answer pair knowledge base. Use the cross entropy loss function to optimize the accuracy of the answers generated by the large model; An optimization algorithm is used to update the parameters of the pre-trained language model to make it more suitable for business scenarios.
[0012] Preferably, the S4 further comprises: The pre-trained language model is a Transformer architecture, including but not limited to GPT, BERT or LLaMA models. The pre-trained parameters are retained during fine-tuning and trained on the business data set.
[0013] Preferably, S5 includes: Online text Q&A, which supports real-time responses to user questions via web pages or mobile apps; Voice question and answer, combining speech recognition and speech synthesis technology, realizes intelligent customer service in the form of voice interaction.
[0014] Preferably, the intelligent customer service question-and-answer task also includes automatically generating recommended content related to the user's question, and the recommended content is generated based on a comprehensive analysis of the user's historical records and a business knowledge base.
[0015] Preferably, the method is applicable to a variety of business scenarios, including but not limited to power customer service, financial consulting, after-sales support, and educational Q&A.
[0016] The present invention provides a method for fine-tuning a large model by scoring question-answer pairs based on a reward mechanism. It has the following beneficial effects: 1. The present invention introduces a question-answer pair scoring method based on a reward mechanism, comprehensively considers semantic consistency, keyword coverage and business relevance, ensures that the screened question-answer pairs have high semantic integrity and business relevance, provides high-quality data support for the subsequent construction of the knowledge base, and effectively improves the fine-tuning training effect of the large model.
[0017] 2. The present invention combines the dynamic knowledge base update and the fine-tuning method of the pre-trained language model to enable the large model to have the semantic understanding and generation capabilities of specific business scenarios, and can flexibly adapt to the needs of various scenarios, thereby achieving more accurate and efficient intelligent customer service question and answer services to meet the diverse needs of users.
[0018] 3. By integrating user history analysis, dynamic retrieval of knowledge base and various forms of interaction, the intelligent customer service system of the present invention can provide personalized, real-time services and recommended content, expand the functional depth and application breadth of the system, and improve user experience and service efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 The figure is a flow chart of the method of the present invention. DETAILED DESCRIPTION
[0020] The following will be combined with the drawings in the specification of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0021] Please see attached Figure 1 The embodiment of the present invention provides a method for fine-tuning a large model by scoring question-answer pairs based on a reward mechanism, comprising the following steps: S1. Filter initial question-answer pairs from the historical chat records of manual customer service. The selected question-answer pairs have high relevance and semantic integrity, which can significantly improve the data quality of subsequent model training and ensure the accuracy of the model from the source. S2. Score the initial question-answer pairs based on the reward mechanism. The scoring results can objectively reflect the actual business value of the question-answer pairs, laying a solid foundation for the subsequent screening of high-quality question-answer pairs. S3. Filter high-scoring question-answer pairs and update the business question-answer pair knowledge base. By setting a scoring threshold, ensure that the question-answer pairs entering the knowledge base are of high quality, thereby improving the validity and reliability of the knowledge base. The updated knowledge base covers higher-quality and more relevant business question-answer pair content, supporting targeted fine-tuning of large models. S4. Fine-tune the large model based on the updated business question-answer knowledge base, and use the high-quality data provided by the business knowledge base to perform targeted optimization on the pre-trained model to adapt it to specific scenarios and solve problems in specific fields more effectively. S5. Use the fine-tuned big model to perform intelligent customer service question and answer tasks. The fine-tuned big model can provide more accurate and efficient intelligent customer service through targeted optimization of the knowledge base, improve the efficiency and accuracy of the customer service system, enhance the user experience, and reduce the workload of manual customer service.
[0022] Please see attached Figure 1 In a preferred embodiment of the present invention, S1 includes: Perform data cleaning and preprocessing on historical chat records to remove invalid question-answer pairs. Data cleaning and preprocessing are key steps in data quality management. By removing redundant and invalid data, the subsequent question-answer pair screening process is more efficient and accurate. The cleaned and preprocessed question-answer pair data is more structured and high-quality, which can effectively reduce the noise of subsequent model processing and improve the efficiency and accuracy of the entire method. The semantic model is used to calculate the semantic similarity between the question and the answer, and the semantically complete question-answer pairs are selected. By calculating the semantic similarity, the semantic consistency between the question-answer pairs can be measured, thereby ensuring that the selected question-answer pairs are closely related in content and have no obvious contradictions in semantic expression. Semantic similarity screening makes the question-answer pairs more complete and more consistent, effectively improving the accuracy and relevance of subsequent model training data; Based on the business keyword set, the number of keywords contained in the question-answer pairs is counted, and the question-answer pairs with high keyword matching are screened. Keyword matching screening helps to extract question-answer pairs that are highly relevant to the business scenario, ensuring that the screening results can better cover the core business content. Through keyword matching screening, the question-answer pairs are more focused on the core business content, improving the business adaptability of the knowledge base and subsequent models; Question and answer pairs are sorted according to the weighted comprehensive score of semantic similarity and keyword matching, and the question and answer pairs with the highest scores are selected. The comprehensive score sorting combines the two core indicators of semantic similarity and keyword matching, which not only ensures the semantic integrity of the question and answer pairs, but also strengthens their relevance to business scenarios. The sorted question and answer pair set has both high semantic consistency and high business relevance, providing a high-quality input data set for subsequent reward mechanism scoring and knowledge base updates.
[0023] Please see attached Figure 1 In a preferred embodiment of the present invention, S2 includes: The question-answer pair is converted into a semantic feature vector as the state input of the reinforcement learning model. The semantic feature vector compresses the semantic information of the question and answer into a fixed-dimensional representation, which makes it easier for the reinforcement learning model to capture the semantic relationship and features of the question-answer pair. The feature vector represents the semantic information of the question-answer pair, which effectively improves the reinforcement learning model's ability to evaluate the quality of the question-answer pair and ensures the accuracy and robustness of the scoring process. The action in reinforcement learning is defined as the score of the question-answer pair. The reward function includes the following parts: Semantic consistency reward: measures the semantic similarity between the question and the answer; Keyword coverage reward: measures the number of business keywords contained in the question and answer pair; Business relevance rewards: Evaluate the matching degree of question-answer pairs according to specific business scenario rules, and comprehensively evaluate the quality of question-answer pairs through three dimensions: semantic consistency, keyword coverage, and business relevance, to ensure that the reward function fully covers the different characteristics of question-answer pairs. The multi-dimensional design of the reward function improves the comprehensiveness and business adaptability of question-answer pair scoring, and helps to screen high-quality question-answer pairs. Train the reinforcement learning scoring model and optimize the scoring strategy based on the reward function. Reinforcement learning drives strategy optimization through rewards, so that the scoring strategy can dynamically adjust the evaluation criteria of question and answer pairs to adapt to different data characteristics. After reinforcement learning training, the scoring model has adaptive scoring capabilities and can perform more accurate quality assessment of question and answer pairs, providing strong support for the subsequent screening of high-quality question and answer pairs.
[0024] Please refer to the attached Figure 1 , in a preferred embodiment of the present invention, S2 further comprises; The specific form of the reward function r is: r=w 1 ·r consistency +w 2 ·r k eywor d +w 3 ·r re l evance where r consistency represents the semantic consistency reward, r keyword represents the keyword coverage reward, r relevance represents the business relevance reward, w1, w2, and w3 are weight parameters, and the semantic consistency reward measures the semantic similarity between the question and the answer to evaluate whether the answer is closely related to the question. High similarity means that the question-answer pair has a high logical consistency in content, ensuring that the selected question-answer pairs have high semantic integrity and logical consistency, providing higher quality data for subsequent knowledge base updates and model training.
[0025] Please refer to the attached Figure 1 In a preferred embodiment of the present invention, S3 includes: Filter question-answer pairs whose reward scores are greater than the preset threshold. According to the calculated reward score r, filter the question-answer pairs (Q, A) that meet the following conditions: Among them, tau is a preset threshold, which usually ranges from [0,1] and is dynamically adjusted according to business needs, such as: For scenarios with high quality requirements, a higher threshold (such as tau = 0.8) can be set; For scenarios with higher data coverage, the threshold can be appropriately lowered (e.g., tau = 0.6); The reward score comprehensively reflects the semantic consistency, keyword coverage and business relevance of the question-answer pair. By setting a threshold, the question-answer pairs that meet the quality requirements can be screened out to ensure that the question-answer pairs entering the next stage meet the set quality standards, improve the effectiveness and reliability of the knowledge base content, and reduce the impact of redundant or low-quality data. The selected high-scoring question-answer pairs are added to the business question-answer pair knowledge base. By dynamically adding high-scoring question-answer pairs to the knowledge base, the knowledge base is continuously optimized and expanded, so that its content can gradually cover a wider range of higher-quality business question-answer data, thereby providing more diverse and accurate basic data for subsequent fine-tuning training of large models. Check the duplication and redundancy of the business question and answer knowledge base to ensure data consistency and integrity, perform hash mapping on the question and answer pairs in the knowledge base, and generate a unique identifier: hash(Q,A)=MD5(Q+A) Where Q and A represent questions and answers. Compare the hash values of the newly added question-answer pairs with the existing ones, and remove duplicate data with the same hash value. Redundancy check: Calculate the similarity scores of similar question-answer pairs based on semantic similarity: If sin(Q i ,Q j ) is large, it is considered a redundant question-answer pair, and only the one with a higher score is retained. The duplication and redundancy check can effectively clean up redundant data through hash value comparison and semantic similarity analysis; the consistency check ensures the standardization and uniformity of the data format, eliminates duplicate and redundant data in the knowledge base, ensures the efficiency and storage utilization of the knowledge base, and maintains the integrity and consistency of the data, providing more accurate data support for the training of large models.
[0026] Please see attached Figure 1 In a preferred embodiment of the present invention, S4 includes: Construct a question-answer pair dataset for fine-tuning training. The data source is the updated business question-answer pair knowledge base. The constructed dataset is the basis for fine-tuning training and comes from a high-quality business question-answer pair knowledge base. It ensures the accuracy of the training data in terms of semantics and business relevance, thereby providing a high-quality, well-formatted and diverse training dataset, providing reliable data support for large model fine-tuning training and improving the performance of the fine-tuned model. Use the cross entropy loss function to optimize the accuracy of the answers generated by the large model. The training goal is to minimize the error between the answers generated by the model and the target answers, and to increase the probability of the model generating the correct answer: Using cross entropy loss function Defined as: Where: N is the number of question-answer pairs in the training data; M is the total number of answer words; y ij is the vocabulary distribution of the target answer; the cross entropy loss function guides the model to learn the mapping relationship between question and answer pairs by maximizing the generation probability of the target answer, thereby improving the accuracy of answer generation. The use of the cross entropy loss function can effectively optimize the generation ability of the model, reduce the deviation of the generated answers, and improve the performance of the model in actual business scenarios; Use optimization algorithms to update the parameters of pre-trained language models to make them more suitable for business scenarios. Pre-trained language models based on Transformer architecture (such as GPT, BERT, LLaMA) retain the basic knowledge of pre-training; fine-tune some parameters of the model (such as specific task layers or fully connected layers); The AdamW optimizer is used to update the model parameters, and its optimization formula is: Where: m t and v t are the first-order momentum and the second-order momentum respectively; β1 is the gradient of the loss function, usually taken as 0.9; β2: momentum attenuation coefficient, usually taken as 0.999; The square value of the current gradient, the optimization algorithm iteratively updates the model parameters, gradually minimizes the loss function, and enhances the model's adaptability to specific business question and answer scenarios. The fine-tuned model can better understand the semantic structure of specific business question and answer pairs, generate answers that are more in line with actual business needs, and improve the practicality of the intelligent customer service system.
[0027] Please see attached Figure 1 In a preferred embodiment of the present invention, S4 further comprises: The pre-trained language model is a Transformer architecture, including but not limited to GPT, BERT or LLaMA models. During the fine-tuning process, the pre-trained parameters are retained and trained on the business dataset. The Transformer architecture is centered on the self-attention mechanism, which can efficiently process long sequences of contextual information and is suitable for capturing semantic relationships in question-answering tasks. The use of pre-trained language models can significantly reduce the training cost of specific tasks while ensuring that the model has strong semantic understanding and generation capabilities. Retaining pre-trained parameters can ensure that the model continues to use its common language features in specific tasks. At the same time, fine-tuning task-specific layers can adapt to new scenarios more efficiently. The fine-tuning method that retains pre-trained parameters takes into account both training efficiency and model performance, avoiding the high cost of training the model from scratch while improving the performance of the model in specific tasks.
[0028] Please see attached Figure 1 In a preferred embodiment of the present invention, S5 includes: Online text question and answer system supports real-time responses to user questions via web pages or mobile applications. The online text question and answer system is based on real-time communication protocols and fine-tuned large models. It can efficiently parse natural language questions entered by users and generate answers. The online text question and answer function provides users with fast and convenient intelligent customer service, reduces user waiting time, and improves the immediacy and interactivity of customer experience. Voice Q&A combines speech recognition and speech synthesis technologies to realize intelligent customer service in the form of voice interaction. Users ask questions through telephone or voice input devices, and voice data is transmitted to the intelligent customer service system in the form of audio stream; the system uses speech recognition technology to convert audio into text: Q text =ASR(Q audio ) Among them, Q text For user voice input, Q audio For the text problem after identification; The recognized text question Q text Input the fine-tuned large model for answer generation.
[0029] The large model generates answer text ASR based on the question: A text =Model(Q text ) The voice question and answer function combines speech recognition and speech synthesis technology to convert users' voice questions into text, generate answers through the intelligent customer service system, and then convert them into voice replies, realizing question and answer in the form of voice interaction. The voice question and answer function enables intelligent customer service to support more diverse forms of interaction, which is particularly suitable for telephone customer service or scenarios where text input is not convenient, improving the accessibility of the service and the naturalness of the user experience.
[0030] Please see attached Figure 1 In a preferred embodiment of the present invention, the intelligent customer service question-answering task also includes automatically generating recommended content related to user questions. The recommended content is generated based on a comprehensive analysis of user historical records and a business knowledge base. The automatically generated recommended content dynamically matches the user's current needs with business knowledge through a combination of user question analysis, historical record analysis, and knowledge base retrieval to generate accurate recommended content, thereby improving the initiative and pertinence of the service. This function can provide more accurate and diverse recommended content based on the user's specific questions and personalized needs, enhance the personalization and intelligence of the user experience, and improve the user's satisfaction and stickiness to the system; The comprehensive analysis model combines user behavior data and knowledge base content to capture the dynamic changes in user needs and provide data-driven support for recommended content. Comprehensive analysis can greatly improve the accuracy of recommended content, enhance the intelligent customer service system's ability to understand complex user needs, and further improve the system's service quality and efficiency.
[0031] Please refer to the attached Figure 1 In a preferred embodiment of the present invention, the method is applicable to a variety of business scenarios, including but not limited to power customer service, financial consulting, after-sales support and educational Q&A. The power customer service scenario is highly dependent on the knowledge base and user interaction data. Intelligent customer service improves service efficiency by dynamically updating knowledge and mining user history records, realizes the automation and efficiency of power customer service, reduces the burden of manual customer service, and improves user experience and service response speed. The financial consulting scenario relies on precise rule matching and business knowledge extraction. The intelligent customer service system realizes professional Q&A in the financial field through fine-tuning of the large model, improves the professionalism and user satisfaction of financial consulting services, and reduces users' dependence on traditional manual consulting.
[0032] In order to better understand the present invention, the above contents are described in detail below in conjunction with specific embodiments.
[0033] Example 1: Question-answer pair scoring method without adding a reward mechanism Specific method: Question-answer pairs are scored only based on traditional rule matching and statistical learning methods, and the scoring indicators include semantic similarity and keyword coverage, without combining the reward mechanism of reinforcement learning and business relevance evaluation.
[0034] Implementation effect: The selected question-answer pairs have certain semantic consistency and keyword coverage capabilities. However, in complex business scenarios, the business adaptability of the question-answer pairs is low, and the optimization of the screening results for specific business scenarios is limited.
[0035] Example 2: Question-answer pair scoring method with only a reward mechanism added Specific method: A reward mechanism is introduced based on reinforcement learning to score question-answer pairs. The scoring indicators are semantic consistency, keyword coverage and business relevance, and the knowledge base is not dynamically updated.
[0036] Implementation effect: The quality of the screened question and answer pairs is significantly improved, and can better adapt to complex business scenarios; the accuracy of the fine-tuned large model answers is improved compared to Example 1, but the knowledge base is not adaptable enough in dynamic scenarios.
[0037] Embodiment 3: Adding a reward mechanism and a dynamic knowledge base update method (embodiment of the present invention) Specific method: Based on Example 2, a dynamic knowledge base update mechanism is added to ensure the consistency and integrity of the knowledge base by screening high-scoring question-answer pairs and checking the duplication and redundancy of the knowledge base, and fine-tune the large model in combination with the updated knowledge base.
[0038] Result: The selected question-answer pairs not only have semantic consistency and business relevance, but can also adapt to dynamic business changes in real time. The fine-tuned large model has higher accuracy and answer coverage in specific scenarios, and the overall performance of the system is greatly improved.
[0039] Comparative experiment 1: Question-answer pair screening quality comparison experiment purpose: to verify the advantages of the question-answer pair screening method of the present invention in terms of semantic consistency, keyword coverage and business relevance.
[0040] Experimental setup Experimental group (the present invention): A reward-based approach to question-answer pair scoring is used.
[0041] Scoring metrics include semantic consistency, keyword coverage, and business relevance.
[0042] Dynamically adjust scoring weights to suit different scenarios.
[0043] Control group A (rule matching method): Question and answer pairs are screened based on keyword matching rules.
[0044] It has no semantic analysis capability and relies only on simple rule filtering.
[0045] Control group B (traditional scoring mechanism): Screening is performed based on a fixed weighted scoring formula based on semantic similarity and keyword coverage.
[0046] No reinforcement learning reward mechanism and business relevance evaluation.
[0047] Experimental procedures Data preparation: 500 question-answer pairs were randomly selected from the customer service history records of a certain industry, containing data on various semantic consistency, keyword coverage, and business relevance.
[0048] Question and answer pairs are labeled for semantic consistency (maximum score 1), keyword coverage (maximum score 1), and business relevance (maximum score 10).
[0049] Method application: Control group A used rule matching to screen question-answer pairs.
[0050] Control group B used the traditional scoring formula to screen question-answer pairs.
[0051] The experimental group used the method of the present invention to score and screen high-scoring question-answer pairs.
[0052] Indicator calculation: Count the average values of semantic consistency, keyword coverage, and business relevance of each group of selected question-answer pairs.
[0053] The comparative experimental data are shown in Table 1: Table 1 Question and answer screening quality experiment data table From the data in Table 1, we can get: Control group A (rule matching): The semantic consistency of the screening results is low, and some question-answer pairs are directly excluded due to keyword mismatch; the keyword coverage is high, but the business relevance is poor, and the screened question-answer pairs are mostly generalized content.
[0054] Control group B (traditional scoring): semantic consistency has been improved and keyword coverage is high, but insufficient consideration is given to business relevance in complex scenarios; the screening results perform well in general scenarios, but have limited effects in specific business scenarios.
[0055] Experimental group (the present invention): semantic consistency is significantly improved, and the screened question-answer pairs perform well in keyword coverage and business relevance; dynamic weight adjustment and reinforcement learning scoring mechanism ensure the comprehensiveness and pertinence of question-answer pair screening, and adapt to a variety of complex business scenarios.
[0056] Comparative Experiment 2: Knowledge Base Update Efficiency Comparative Experiment Purpose: To verify the advantages of the present invention in the dynamic update of the knowledge base and to evaluate its performance in terms of repeatability, redundancy and knowledge update delay.
[0057] Experimental setup Experimental group (the present invention): A reward-based mechanism is used to screen question-answer pairs, and dynamic knowledge base updates are achieved through repetition and redundancy checks.
[0058] Adopt automated update mechanism to maintain the consistency and integrity of the knowledge base in real time.
[0059] Control group A (manual update): The knowledge base is updated manually on a regular basis, relying on human judgment to filter question and answer pairs, with no automated support.
[0060] Update frequency is set to once a week.
[0061] Control group B (rule filtering update): New question and answer pairs are added based on rule filtering, and the knowledge base is updated by keyword matching.
[0062] Lack of duplication and redundancy checks, relying only on simple matching rules.
[0063] Experimental procedures Data preparation: A business scenario was simulated to generate a log containing 500 new question and answer records, including duplicate question and answer pairs (15%), redundant question and answer pairs (20%), and high-quality question and answer pairs (65%).
[0064] Method application: Control group A updates the knowledge base through manual screening and performs updates once a week.
[0065] Control group B added new question and answer pairs through rule filtering without duplication and redundancy checks.
[0066] The experimental group adopted the method of the present invention to screen high-quality question-answer pairs according to the reward scores, and dynamically updated the knowledge base through repetitive and redundancy checks.
[0067] Indicator calculation: Repetition rate: the proportion of repeated question and answer pairs in the knowledge base.
[0068] Redundancy rate: the proportion of question and answer pairs in the knowledge base that have no actual business value.
[0069] Update delay: The time interval from adding a new question and answer pair to updating the knowledge base.
[0070] The comparative experimental data are shown in Table 2: Table 2 Knowledge base update efficiency experimental data table From the data in Table 2, we can get: Control group A (manual update): The repetition and redundancy rates in the knowledge base are high, and the efficiency of manual screening is low, resulting in significant update delays; it is impossible to dynamically adapt to new data, and the real-time performance of the knowledge base is poor.
[0071] Control group B (rule filtering): The update efficiency is improved and the update delay is significantly lower than that of control group A. However, due to the lack of duplication and redundancy checks, there is still a high proportion of invalid data in the knowledge base.
[0072] Experimental group (the present invention): the update efficiency is the highest, the repetition rate and redundancy rate in the knowledge base are significantly reduced, and the dynamic update delay is almost negligible; high-scoring question-answer pairs are screened through a reward mechanism, and combined with repetition and redundancy checks to ensure the high quality and dynamic adaptability of the knowledge base content.
[0073] Comparative Experiment 3: Comprehensive Performance Comparison Experiment of Intelligent Customer Service Purpose: To verify the comprehensive performance advantages of the present invention in the intelligent customer service system, and to evaluate user satisfaction, problem solving rate and system response time.
[0074] Experimental setup Experimental group (the present invention): A reward-based question-answer pair scoring method and dynamic knowledge base update are used in combination with a fine-tuned pre-trained language model to process user questions.
[0075] Control group A (no model fine-tuning, only using pre-trained model): Use general pre-trained language models (such as GPT-3) to process user questions without fine-tuning or business adaptation.
[0076] Control group B (traditional fine-tuning method): The language model is fine-tuned using a fixed dataset, but lacks dynamic knowledge base updates and reinforcement learning scoring mechanisms.
[0077] Experimental procedures Data preparation: Simulate the application of intelligent customer service system in power customer service scenarios and design 1,000 user questions, covering common problems and complex business scenarios.
[0078] Method application: Each group of systems processes the same set of questions and records the answers and user feedback.
[0079] User issues include: electricity bill inquiries, power outage notifications, payment method instructions, bill error handling, etc.
[0080] Indicator calculation: User satisfaction: User rating of the system's answers (out of 5 points, based on 100 test users).
[0081] Question solving rate: the proportion of correct answers given by the system to user questions.
[0082] System response time: The average time from when a user submits a question to when an answer is generated.
[0083] The comparative experimental data are shown in Table 3: Table 3 Intelligent customer service comprehensive performance experimental data table From the data in Table 3, we can get: Control group A (no model fine-tuning): User satisfaction is low because the pre-trained model lacks business adaptability, and the answers are highly generalized but not targeted enough; the problem solving rate is the lowest, especially in complex scenarios, but the response time is fast.
[0084] Control group B (traditional fine-tuning): User satisfaction and problem solving rate have improved, but the lack of dynamic knowledge base updates leads to insufficient adaptability of the model after changes in business knowledge; the performance in business scenario matching is average, and there are still misanswers in complex scenarios.
[0085] Experimental group (the present invention): user satisfaction and problem solving rate are significantly higher than those of the control group, thanks to the combination of dynamic knowledge base update and reinforcement learning scoring mechanism; system response time is slightly higher than that of control group A, but within a reasonable range and highly acceptable.
[0086] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for fine-tuning a large model by scoring question-answer pairs based on a reward mechanism, characterized in that: The following steps are involved: S1. Filter the initial question and answer pairs from the historical chat records of the manual customer service; S2. Score the initial question-answer pairs based on the reward mechanism; S3. Filter high-scoring question-answer pairs and update the business question-answer pair knowledge base; S4. Fine-tune the large model based on the updated business question and answer knowledge base; S5. Use the fine-tuned large model to perform intelligent customer service question-and-answer tasks.
2. The method for fine-tuning a large model by scoring question-answer pairs based on a reward mechanism according to claim 1, characterized in that: The S1 includes: Perform data cleaning and preprocessing on historical chat records to remove invalid question and answer pairs; Use the semantic model to calculate the semantic similarity between the question and the answer, and select semantically complete question-answer pairs; Based on the business keyword set, count the number of keywords contained in the question-answer pairs, and filter out question-answer pairs with higher keyword matching degrees; sort the question-answer pairs according to the weighted comprehensive score of semantic similarity and keyword matching degrees, and select the question-answer pairs with the highest score rankings.
3. The method for fine-tuning a large model by scoring question-answer pairs based on a reward mechanism according to claim 1, characterized in that: The S2 includes: Convert the question-answer pair into a semantic feature vector as the state input of the reinforcement learning model; The action in reinforcement learning is defined as the score of the question-answer pair. The reward function includes the following parts: Semantic consistency reward: measures the semantic similarity between the question and the answer; Keyword coverage reward: measures the number of business keywords contained in the question-answer pair; Business relevance rewards: Evaluate the matching degree of question and answer pairs based on specific business scenario rules; Train a reinforcement learning scoring model and optimize the scoring strategy based on the reward function.
4. The method for fine-tuning a large model by scoring question-answer pairs based on a reward mechanism according to claim 1, characterized in that: Said S2 further comprises: The specific form of the reward function r is: r=w 1 ·r consistency +w 2 ·r k eywor d +w 3 ·r re l evance where r consistency represents the semantic consistency reward, r keyword represents the keyword coverage reward, r relevance It represents the business relevance reward, w1, w2, w3 are weight parameters.
5. The method for fine-tuning a large model by scoring question-answer pairs based on a reward mechanism according to claim 1, characterized in that: The S3 includes: Filter question-answer pairs whose reward scores are greater than a preset threshold; Add the selected high-scoring question-answer pairs to the business question-answer pair knowledge base; Check the knowledge base for duplication and redundancy of business questions and answers to ensure data consistency and integrity.
6. The method for fine-tuning a large model by scoring question-answer pairs based on a reward mechanism according to claim 1, characterized in that: The S4 includes: Build a question-answer pair dataset for fine-tuning training. The data source is the updated business question-answer pair knowledge base. Use the cross entropy loss function to optimize the accuracy of the answers generated by the large model; An optimization algorithm is used to update the parameters of the pre-trained language model to make it more suitable for business scenarios.
7. The method for fine-tuning a large model by scoring question-answer pairs based on a reward mechanism according to claim 6, characterized in that: The S4 further comprises: The pre-trained language model is a Transformer architecture, including but not limited to GPT, BERT or LLaMA models. The pre-trained parameters are retained during fine-tuning and trained on the business data set.
8. The method for fine-tuning a large model by scoring question-answer pairs based on a reward mechanism according to claim 1, characterized in that: The S5 includes: Online text Q&A, which supports real-time responses to user questions via web pages or mobile apps; Voice question and answer, combining speech recognition and speech synthesis technology, realizes intelligent customer service in the form of voice interaction.
9. The method for fine-tuning a large model by scoring question-answer pairs based on a reward mechanism according to claim 1, characterized in that: The intelligent customer service question-and-answer task also includes automatically generating recommended content related to user questions, and the recommended content is generated based on a comprehensive analysis of user historical records and a business knowledge base.
10. The method for fine-tuning a large model by scoring question-answer pairs based on a reward mechanism according to claim 1, characterized in that: The method is applicable to a variety of business scenarios, including but not limited to power customer service, financial consulting, after-sales support, and educational Q&A.
Citation Information
Patent Citations
Method for optimizing question and answer models on basis of adversarial network reinforcement learning
CN107423437A
Knowledge base question and answer extraction method and system, mobile terminal and storage medium
CN110825860A
Updating method, device and equipment of intelligent customer service knowledge base and storage medium
CN112148743A
Method for generating double-model corpus
CN118350459A
Method and system for processing task data of intelligent agent based on multi-modal large model
CN118396119A
Cited By
Generation method and device of instruction following data set, server and storage medium
CN120316514A
Expert preference alignment service processing method and device, equipment and medium
CN120407754A
Distribution transformer fault detection method and system for photovoltaic transformer area
CN120429645A
Financial large model training method and business problem processing method
CN120670545A
Intelligent professional content generation method based on AI writing
CN120781844A