Term alignment language generation method and device, equipment and medium
By constructing a term alignment reward function and reinforcement learning training, the problems of misuse of terminology and non-standard expression in large language models in industry scenarios are solved, and the compliance and professionalism of language generation are improved. It is suitable for fields such as banking, finance, insurance and healthcare.
Patent Information
- Application Number
- CN202510830501.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-19
AI Technical Summary
Existing large language models have problems with generalized or incorrect terminology usage and inconsistent language styles in scenarios such as banking, finance, insurance, and healthcare, making it difficult to meet industry standards. Existing methods also lack systematic constraint mechanisms.
A term alignment reward function is constructed, including a term matching score, a language fluency score, an information coverage score, and a sensitive expression penalty score. The language generation model is trained through reinforcement learning to screen out compliant candidate responses and update the policy parameters.
It improves the terminology compliance and speech alignment capabilities in the language generation process, enhances the applicability and compliance of the model in professional scenarios, and reduces deployment and maintenance costs.
Smart Images

Figure CN120671846A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a term alignment language generation method, device, equipment and medium. Background Art
[0002] With the rapid development of Large Language Model (LLM) technology, LLM-based intelligent customer service and automated question-answering systems have gained widespread application in banking, finance, insurance, healthcare, and other scenarios. These large models possess strong language understanding and generation capabilities, effectively supporting intelligent services such as answering customer questions, handling business inquiries, and guiding processes, thereby improving service efficiency and user experience. However, during customer service, customer communication often requires strict adherence to industry standards, with specific requirements for terminology, wording compliance, and language style. For example, standard terminology must be used accurately, non-compliant and non-committal language is prohibited, and sentence structure must conform to prescribed templates. While current mainstream large language models, such as GPT, GLM, and LLaMA, possess general language generation capabilities, they are not specifically tailored for banking, finance, insurance, and healthcare scenarios. This can lead to incorrect terminology and non-compliant wording in their output, increasing business risks.
[0003] Specifically, existing large language models suffer from two main issues: generalized or incorrect terminology; and inconsistent language styles that fail to strictly align with service scripting standards. Currently, the industry largely uses prompt engineering to guide large language models to output compliant content. However, this approach relies heavily on prompt design and lacks a systematic constraint mechanism, making it impossible to achieve real-time control of illegal language. Summary of the Invention
[0004] The present invention provides a term alignment language generation method, apparatus, device and medium to solve the technical problem that the existing large language model has obvious deficiencies in term alignment and expression standardization.
[0005] In a first aspect, a term alignment language generation method is provided, comprising:
[0006] Collect user input and context information in the target domain, combine it with the preset terminology knowledge base, and construct input environment data for language generation;
[0007] Inputting the input environment data into a language generation model to generate a plurality of candidate language responses to obtain a candidate response set;
[0008] Based on the candidate response set, constructing a term alignment reward function; wherein the term alignment reward function includes a term matching score, a language fluency score, an information coverage score, and a sensitive expression penalty score;
[0009] Scoring each candidate language response in the candidate response set according to the term alignment reward function to obtain a scoring result;
[0010] Based on the scoring results, selecting the candidate language response with the highest score as the final output response;
[0011] The final output response is used to update the strategy parameters of the language generation model to obtain a language generation result.
[0012] In a second aspect, a term alignment language generation device is provided, comprising:
[0013] The data collection module is used to collect user input and context information in the target domain, and build input environment data for language generation by combining it with the preset terminology knowledge base;
[0014] a data input module, configured to input the input environment data into a language generation model to generate a plurality of candidate language responses to obtain a candidate response set;
[0015] A data construction module, configured to construct a term alignment reward function based on the candidate response set; wherein the term alignment reward function includes a term matching score, a language fluency score, an information coverage score, and a sensitive expression penalty score;
[0016] a data scoring module, configured to score each candidate language response in the candidate response set according to the term alignment reward function to obtain a scoring result;
[0017] A data screening module, configured to screen the candidate language response with the highest score as the final output response based on the scoring result;
[0018] The data generation module is used to update the strategy parameters of the language generation model using the final output response to obtain a language generation result.
[0019] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned term alignment language generation method when executing the computer program.
[0020] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned term alignment language generation method are implemented.
[0021] In the solution implemented by the above-mentioned term alignment language generation method, device, equipment and medium, user input and context information in the target domain can be collected, and input environment data for language generation can be constructed in combination with a preset term knowledge base; the input environment data is input into the language generation model to generate multiple candidate language responses to obtain a candidate response set; based on the candidate response set, a term alignment reward function is constructed; wherein the term alignment reward function includes a term matching score, a language fluency score, an information coverage score and a sensitive expression penalty score; each candidate language response in the candidate response set is scored according to the term alignment reward function to obtain a scoring result; based on the scoring result, the candidate language response with the highest score is selected as the final output response; the final output response is used to update the strategy parameters of the language generation model to obtain a language generation result. In the present invention, in order to address the obvious deficiencies in the term alignment and expression standardization of existing large language models, a term alignment reward function can be constructed, and the final output response can be scored and selected to update the strategy parameters to obtain a language generation result, thereby improving the compliance of the terminology and the ability to align the speech in the language generation process. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0023] Figure 1 This is a flow chart of a method for generating a term alignment language according to an embodiment of the present invention;
[0024] Figure 2 yes Figure 1 A specific implementation process diagram of step S30 Figure 1 ;
[0025] Figure 3 yes Figure 1 A specific implementation process diagram of step S30 Figure 2 ;
[0026] Figure 4 yes Figure 3 A schematic flow chart of a specific implementation of step S321;
[0027] Figure 5 yes Figure 1 A specific implementation process diagram of step S30 Figure 3 ;
[0028] Figure 6 yes Figure 1 A specific implementation process diagram of step S30 Figure 4 ;
[0029] Figure 7 yes Figure 1 A specific implementation process diagram of step S30 Figure 5 ;
[0030] Figure 8 It is a structural diagram of a term alignment language generation device in one embodiment of the present invention;
[0031] Figure 9 is a structural diagram of a computer device in one embodiment of the present invention;
[0032] Figure 10 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0033] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0034] See also Figure 1 As shown, Figure 1 A flowchart of a method for generating a term alignment language provided in an embodiment of the present invention includes the following steps:
[0035] S10: Collect user input and context information in the target domain, combine it with the preset terminology knowledge base, and construct input environment data for language generation.
[0036] Collect user input information and contextual data in the target business scenario. For example, in a financial scenario, a user may enter "I want to apply for a student credit card", and in a medical scenario, a user may enter "I have been coughing recently". According to different business needs, extract the contextual features related to the current task, including conversation history, intent category, etc., and combine them with the preset terminology knowledge base for structured organization to form a "state representation" in the reinforcement learning environment. The terminology knowledge base contains financial standard terms (such as "limit" and "interest rate") and medical standard terms (such as "duration" and "body temperature range"). This step forms the environmental input data for response generation.
[0037] S20: Inputting the input environment data into a language generation model to generate a plurality of candidate language responses to obtain a candidate response set.
[0038] The constructed input environment data is fed into the language generation model. For example, a policy network structure based on LLaMA or ChatGLM, combined with Top-k sampling or Nucleus sampling strategies, is used to generate multiple candidate language responses, forming a candidate response set. Each candidate response represents a possible output of the language generation model in the current context. For example, in a financial scenario, a response might be "We recommend applying for a student card; the annual fee may be waived," while in a medical scenario, a response might be "We recommend visiting a hospital for a checkup."
[0039] S30: Constructing a term alignment reward function based on the candidate response set; wherein the term alignment reward function includes a term matching score, a language fluency score, an information coverage score, and a sensitive expression penalty score.
[0040] In this step, a term alignment reward function is constructed for reinforcement learning training. This term alignment reward function can comprehensively consider four aspects:
[0041] Term matching score: Based on the terminology knowledge base, this score checks whether the candidate response set accurately uses industry-standard terminology. For example, it checks whether "apply" is used instead of "apply for a card," or whether "use antibiotics" replaces ambiguous expressions such as "take some medicine."
[0042] Language fluency score: assesses the perplexity of a sentence, or uses a syntactic integrity analysis tool to detect whether the subject-verb-object structure is complete;
[0043] Information coverage score: Based on the key slots extracted from the input information (e.g., "Applicant = student"), the candidate response set is compared to see whether they are fully answered;
[0044] Sensitive expression penalty score: Detects ambiguous, illegal, or inappropriate expressions (such as "any card" or "take some medicine and see") in the candidate response set and assigns a negative score.
[0045] The above four reward functions are combined according to weighted rules to form a multi-dimensional evaluation system, which can set different weight strategies for different industry scenarios. A multi-dimensional reward function that can be weighted and combined can be designed as shown in the following formula:
[0046] R total =α·R term +β·R fluency +γ·R coverage +λ·R penalty
[0047] Among them, R term Represents the term matching score, which is scored based on whether the output hits the term knowledge base, with a high weight; Rfluency Indicates the language fluency score, ensuring that the language is natural and in line with human language habits, with a medium weight; R coverage Represents the information coverage score, indicating whether the generated content covers the key points of the user's question, with a medium weight; R penalty Penalty points for vague / prohibited expressions (i.e., sensitive expression penalty points), such as the use of non-standard terms such as "apply for a card" and "get a list", provide negative incentives for prohibited, non-standard, and vague terms, and give them a high weight.
[0048] Of course, different weight combinations can be set for different business scenarios (such as account opening, consultation, and complaint handling), with the sum of the weights being 1, and set according to experience and business needs (such as 0.3, 0.2, 0.2, 0.3). The weights α, β, γ, and λ can be dynamically adjusted through grid search or reinforcement learning. After the reward function is combined with the RL model, it can guide the model to produce "compliant, clear, and natural" responses with terminology as constraints, fluency as a carrier, and integrity as a goal. The RL model is a policy optimization model in reinforcement learning. In this embodiment, it is used to guide the language generation model to generate text with more consistent terminology.
[0049] Combine Figure 2 As shown, the construction of the term alignment reward function specifically includes the following steps:
[0050] S311: Determine a term set that should appear in the candidate language response based on the current task intent as an expected term set for term matching.
[0051] Before the language generation task begins, the set of terms that should appear in the generated content is automatically determined based on the intent label of the current task, and is recorded as the expected term set K_expected. This expected term set can be determined by a task classification template, which is preset by experts and trained based on historical corpus. For example, in financial services, expected terms for tasks such as "card activation" include "limit", "annual fee reduction", "billing date", etc.; in medical services, expected terms for initial diagnosis tasks include "duration", "body temperature", "auscultation", etc. The expected term set can support weight grading. For key terms that must appear (such as "responsibility reminders" in the financial field and "symptom descriptions" in medical scenarios), higher weights are set, and for secondary terms, lower weights are set.
[0052] S312: Perform word segmentation and syntactic analysis on the candidate language response to obtain a set of semantic blocks.
[0053] The candidate responses are analyzed for their linguistic structure. First, a Chinese word segmentation tool is used to segment the sentences. Then, a syntactic analyzer is used to extract syntactic units and semantic components, constructing a set of semantic chunks B = {b1, b2, ..., b}. For example, if the candidate response is "It is recommended that you apply for a student card," the semantic chunks are: "apply," "student card"; if the candidate response is "It is recommended that you test for infection," the semantic chunks are: "test," "infection." This set of semantic chunks serves as the basis for term matching.
[0054] S313: Compare the semantic block set with the terminology knowledge base to determine whether each semantic block is a standard term defined in the terminology knowledge base, and obtain a terminology hit result set.
[0055] Match the semantic block set B with the term knowledge base K one by one to determine the value of each semantic block b i Whether it belongs to the set of standard terms defined in the knowledge base. The term knowledge base K can be a set of terms constructed by business experts in the financial or medical industries, such as K = {t1, t2, ..., t}, which includes recognized standard expressions, such as financial terms such as "product recommendation" and "credit limit increase"; and medical terms such as "electrocardiogram examination." Matching methods can be exact matching, fuzzy matching, or synonym expansion matching, and hits are recorded in the term hit result set.
[0056] S314: Calculate the ratio of the term hit result set to the expected term set to obtain a term matching score.
[0057] Based on the term hit result set and the expected term set, the term matching score is calculated. The specific calculation formula is:
[0058]
[0059] Among them, the numerator is the number of terms in the candidate language response that belong to both the semantic block set B and the terminology knowledge base K, and the denominator is the number of standard terms expected to be hit in this task.
[0060] Combine Figure 3 As shown, the construction of the term alignment reward function further includes the following steps:
[0061] S321: Calculate the perplexity index of the candidate language response using a language model.
[0062] A pre-trained language model (such as GPT-2) or a traditional N-gram-based language model can be used to process the candidate text and calculate its perplexity as a fluency indicator (i.e., perplexity index). The perplexity calculation formula is as follows:
[0063]
[0064] Among them, N represents the number of words in the sentence, P(w i |w1,...w i-1 ) is the conditional probability of generating the i-th word under a given context.
[0065] Perplexity reflects the degree of uncertainty a language model has when predicting the next word. Lower perplexity indicates that the generated text conforms more closely to linguistic patterns and is more natural. This is applicable to professional expression scenarios, such as financial services responses like "Your payment application is being processed" or medical advice like "We recommend visiting a hospital for further examination."
[0066] S322: Alternatively, based on a syntax analysis and language quality detection tool, perform a syntax completeness check and grammatical error identification on the candidate language response to obtain a corresponding language fluency score.
[0067] In some scenarios, if the perplexity of the language model is unavailable, a language quality assessment tool can be used to analyze the syntactic structure and grammatical accuracy of the candidate language response as an alternative indicator (i.e., language fluency score). Specifically, the following tools are available:
[0068] Analyze whether the sentence has a complete subject-verb-object structure. For example, "It is recommended that you submit your application materials" is a valid sentence structure, while "Just submit the materials" lacks a subject.
[0069] Grammatical error identification: Use open source tools such as LanguageTool to identify common grammatical errors, including incorrect verb collocations and improper punctuation.
[0070] In the financial field, standard responses usually require semantic completeness and grammatical standardization, such as "Overdraft consumption is not recommended, please pay attention to the repayment period"; in the medical field, health advice must be accurately conveyed, such as "If symptoms persist for more than three days, please seek medical attention in time."
[0071] S323: Normalize the confusion index or the language fluency score to a preset range to obtain a normalized language fluency score.
[0072] To facilitate subsequent unified evaluation and weighting, the perplexity index or language fluency score is normalized and uniformly mapped to the interval [0, 1]. Specifically, for the perplexity index, a reasonable interval (e.g., 10-100) is set and normalized using a linear mapping method. For the language fluency score, a complete sentence is scored as 1, a partially structured sentence is scored as 0.5, and an incorrect sentence is scored as 0, and all are standardized. For example, if the perplexity of a candidate language response is 30, the normalized score is 1-(30-10) ÷ (100-10) ≈ 0.78.
[0073] Combine Figure 4 As shown, the calculation of the perplexity index of the candidate language response using the language model specifically includes the following steps:
[0074] S3211: Split the candidate language response into word sequences and input them into a pre-trained language model.
[0075] The candidate responses to be evaluated are segmented to form an ordered sequence of words. For example, the candidate response "I want to apply for a student card" is broken down into the word sequence {"I", "want", "apply", "one", "student", "card"}. This word sequence is then fed into a pre-trained language model, such as GPT-2, ChatGLM, or an N-gram-based statistical language model, to prepare for subsequent conditional probability calculations. In a financial scenario, this candidate response might be a standard reply generated by a bank's customer service system. In a medical scenario, a candidate response such as "I recommend going to the hospital for an examination" can also be modeled and processed in the same manner.
[0076] S3212: Based on the pre-trained language model, the conditional probability of each word in the word sequence under the given condition of its preceding word is calculated in sequence.
[0077] Call the pre-trained language model and calculate the probability of occurrence of each word under the previous context conditions based on the input word sequence, which can be formalized as follows:
[0078] P(w2|w1), P(w3|w1,w2),…,P(w _T |w1,w2,…,w T-1 )
[0079] Among them, T represents the i-th word, and the preceding word group {w1,...,w T-1} is the context condition. In the above example, the following probabilities will be calculated:
[0080] P("think"|"I"),
[0081] P("apply" | "I want"),
[0082] P("Card" | "I want to apply for a student card")
[0083] Similarly, in a medical scenario, if the candidate language response is "It is recommended to take antibiotics to treat infection", the language prediction probability such as P("antibiotics" | "It is recommended to take") will be calculated.
[0084] S3213: Calculate the logarithmic mean of the conditional probability and calculate the perplexity value with the negative value as the exponential power; wherein the perplexity value is the perplexity index.
[0085] The perplexity value reflects the model's difficulty in predicting candidate language responses. Lower perplexity indicates more logical sentence structure and more fluent language. Perplexity values typically range from 10 to 100. This value is subsequently normalized to the [0, 1] range through a linear transformation to facilitate integration with other reward metrics.
[0086] For example, in finance, if the candidate sentence "You can log in to your mobile banking account to adjust your credit limit" has a perplexity score of 28 calculated by GPT-2, the system can normalize this score and incorporate it into the term alignment reward function as a fluency score. In a medical question-answering system, if the candidate sentence "Please take the medication prescribed by your doctor on time" has a perplexity score of 15, it indicates that the language structure is standard and natural.
[0087] Combine Figure 5 As shown, the construction of the term alignment reward function further includes the following steps:
[0088] S331: Based on the intent recognition and slot extraction model, key intent slots are extracted from the user input to obtain a set of information elements.
[0089] User input sentences are parsed based on an intent recognition and slot extraction model. The model used can be a pre-trained BERT classifier or a fine-tuned version of it, which is used to identify the key intent in the input sentence and the corresponding slot information. The extracted information element set covers the core parameters involved in the specific business request. For example: in a financial scenario, if the user enters "I want to apply for a student card", the system will extract the slot: Operation = Application, Card Type = Student Card; in a medical scenario, if the user enters "I have been coughing and having a low-grade fever recently", the slot may include: Symptom = Cough, Temperature = Low-grade Fever, Duration = Recently. This information element set serves as the target basis for subsequent comparison of candidate responses.
[0090] S332: Input the candidate language response into a slot extraction module to identify semantic slot elements in the candidate language response.
[0091] The generated candidate language responses are fed into the slot extraction module to identify the semantic slot elements within them. This slot extraction module, which shares the same structure as the user intent extraction module, is used to determine whether the response accurately expresses the key information from the user's original intent. For example, if the response is "It is recommended that you apply for a student card, which will waive the annual fee," the semantic slots identified are: Action = Apply, Card Type = Student Card. If the response is "It is recommended that you take cough medicine," the extracted slots might simply be: Recommendation = Medication, Symptom Description = Cough.
[0092] S333: Compare the information element set with the semantic slot elements to obtain the number of hit slots.
[0093] The extracted user information elements are compared one by one with the identified semantic slot elements, and the number of hits is counted, indicating the system's coverage of the user's concerns. The number of hit slots represents the items in the candidate response that successfully reflect the user's concerns, forming the "hit slot set."
[0094] S334: Calculate the ratio between the number of hit slots and the total number of slots to obtain an information coverage score.
[0095] Information coverage score R coverage The definition is as shown in the following formula:
[0096]
[0097] Among them, "total number of response slots" refers to all slot items extracted by the intent recognition module from the user input as the coverage target; "number of hit slots" refers to the corresponding slot items successfully included in the candidate language response. For example, in a financial scenario, if the user inputs "I want to apply for a student card", the extracted slot operation = application and card type = student, if the candidate language response only mentions "recommendation to apply for a card" without distinguishing the attributes of the student card, then the information coverage score = 1 ÷ 2 = 0.5, indicating insufficient information coverage. In a medical scenario, if the user's symptoms are "cough, fever" but the response only covers "cough", the score will also be reduced accordingly.
[0098] Combine Figure 6 As shown, the construction of the term alignment reward function further includes the following steps:
[0099] S341: Build a sensitive word list.
[0100] A list of sensitive terms is pre-defined by industry experts as the basis for language output risk control. This list can be represented as B = {v1, v2, ..., v}, where v represents ambiguous, unprofessional, or unauthorized expressions that pose a risk of violation. For example, sensitive terms in the financial industry include: "Any card will work," "I'm not sure about this," and "Fill out this form." Sensitive terms in the medical industry include: "Take some medicine first," "It's almost okay," and "This problem is not a big deal." These terms will be included in the banned dictionary if they are unclear, inaccurate, or violate regulatory language standards.
[0101] S342: Input the candidate language response into the sensitive term detection module, and determine whether the candidate language response contains a term in the sensitive word list; if so, assign a corresponding risk penalty weight according to the risk level of each sensitive term.
[0102] The candidate language response is input into the sensitive term detection module. The sensitive term detection module uses text scanning and semantic recognition technology to determine the sensitivity of the expressions in the candidate language response in the following ways:
[0103] String matching: directly check whether it contains entries in the banned dictionary;
[0104] Fuzzy matching: Use Levenshtein distance (edit distance) to identify spelling similarities or colloquial expressions;
[0105] Semantic matching: Use word embedding models (such as Word2Vec) to calculate the semantic similarity between candidate language responses and sensitive terms.
[0106] If a sensitive term v∈B is detected in the candidate language response, the corresponding risk weight P is assigned to the term. risk (v) This weight is determined based on the severity of the term's impact on the business, with higher values indicating greater risk. For example, in financial scenarios, the terms "guaranteed profit" and "guaranteed principal" are high-risk, sensitive terms, and their weights can be set between 0.8 and 1.0. In medical scenarios, "unclear" and "let's see" are low-professional expressions, and their weights can be set between 0.4 and 0.6.
[0107] S343: Accumulate the risk penalty weights to obtain a sensitive expression penalty score.
[0108] Accumulate the risk weights corresponding to all sensitive terms hit as the sensitive expression penalty score R penalty The specific calculation formula is:
[0109]
[0110] Among them, S is the set of sensitive terms actually detected in the candidate language response, P risk (v) is the risk weight corresponding to v. For example, if a candidate language response includes "any card can be used" (weight 0.9) and "take a form and fill it out" (weight 0.6), then R penalty = 0.9 + 0.6 = 1.5. This score will ultimately have a negative impact on the overall score of the candidate language response, used to suppress the output tendency of non-compliant or unprofessional expressions.
[0111] Combine Figure 7 As shown, the construction of the term alignment reward function further includes the following steps:
[0112] S351: Preset the weight configuration of term matching, language fluency, information coverage, and sensitive expression penalty corresponding to each business scenario.
[0113] Based on the specific needs of the target industry scenario, the weight ratios of the term matching score, language fluency score, information coverage score, and sensitive expression penalty score in the reward function are predefined and denoted as α, β, γ, and λ respectively. The corresponding weights are as follows:
[0114] α: term matching score weight, used to emphasize term compliance;
[0115] β: Language fluency weight, used to ensure language naturalness and readability;
[0116] γ: information coverage weight, used to improve content completeness;
[0117] λ: The weight of the sensitive expression penalty term, used to constrain non-compliant or ambiguous expressions.
[0118] Different weight configurations can be adjusted based on scenario specificity:
[0119] In financial scenarios (such as account opening, consulting, and financial advice), terminology compliance and sensitive terminology control are usually prioritized, and the α and λ weights are set to higher values;
[0120] In medical scenarios (such as health consultation and initial diagnosis recommendations), more emphasis is placed on information completeness and language clarity, with γ and β accounting for a larger proportion.
[0121] This weight configuration is embedded into the reward function structure as the initial training parameter to guide subsequent training behavior.
[0122] S352: Perform a weighted combination of the weight configuration and the components of the term alignment reward function, and dynamically adjust the weight configuration based on a preset strategy optimization mechanism to adapt to the actual performance of the results generated by the language generation model at different stages.
[0123] Combine each scoring indicator with the corresponding weight to form the total reward function R of term alignment total , expressed as follows:
[0124] R total =α·R term +β·R fluency +γ·R coverage +λ·R penalty
[0125] Among them, R term represents the term matching score, R fluency represents the language fluency score, R coverage represents the information coverage score, R penalty represents the sensitive expression penalty score, and the four are weighted and synthesized to serve as the strategy optimization target of the reinforcement learning model.
[0126] To further adapt to the dynamic changes in model performance during training, two strategies can be introduced for weight adjustment:
[0127] Grid search mechanism: During the pre-training phase, a grid search is used to traverse multiple weight combination schemes, evaluate their comprehensive performance on the validation set, and select the optimal combination;
[0128] Dynamic reinforcement learning adjustment mechanism: During training, the weight coefficients are adjusted in real time based on feedback indicators such as the accuracy, coverage, and violation rate of the model's current output, allowing the reward function to focus on current performance weaknesses, thereby achieving more efficient strategy convergence.
[0129] For example, when the model generates content in a financial scenario that frequently contains sensitive terms, the value of λ will be automatically increased; in a medical scenario, if the response content frequently omits key information slots, the proportion of γ will be increased accordingly.
[0130] S40: Score each candidate language response in the candidate response set according to the term alignment reward function to obtain a scoring result.
[0131] Based on the term alignment reward function, each candidate language response in each candidate response set is scored one by one to obtain a corresponding scoring result. Responses with high scores are generally characterized by clear semantics, standardized terminology, and comprehensive coverage. Responses with low scores may contain prohibited terms, missing semantics, or incomplete structure.
[0132] S50: Based on the scoring result, the candidate language response with the highest score is selected as the final output response.
[0133] Based on the scoring results, the highest-scoring response is selected from the candidate response set as the final output. If all candidate responses fail to meet the minimum compliance standard, a semantic rewrite mechanism is triggered to regenerate a new round of candidate responses to ensure compliance. For example, in a medical consultation scenario, the system would prioritize outputting "Self-medication is not recommended, please seek medical attention" over "Take some medicine and see."
[0134] S60: Using the final output response, the strategy parameters of the language generation model are updated to obtain a language generation result.
[0135] The final scoring response is used as positive feedback, and reinforcement learning-based policy optimization algorithms (such as PPO) are used to update the policy network parameters of the language generation model. This allows the model to gradually learn strategies for selecting standardized terminology and compliant expressions in different environments. This update process continues over multiple rounds of dialogue and sample iterations, significantly improving the model's terminology alignment capabilities and response reliability in professional scenarios such as finance and healthcare.
[0136] As can be seen, the above technical solution, by introducing a reinforcement learning-based term alignment reward mechanism, systematically addresses the issues of misuse of terminology, non-standard expressions, and difficulty meeting regulatory requirements that exist in existing large language models in industry applications. Compared to traditional methods that rely on prompt engineering, this technical solution, without changing the main model structure and parameters, achieves comprehensive constraints and guidance on terminology use, language fluency, information coverage, and sensitive expressions through the training of a reinforcement learning policy network, thereby improving the compliance and professionalism of language generation results.
[0137] Specifically, this technical solution first constructs a term alignment reward function, taking term accuracy as the primary optimization goal. This allows the language model to be guided by the industry terminology knowledge base during the generation process, effectively avoiding term generalization and misuse. Secondly, based on a multi-dimensional scoring mechanism (such as language fluency, information coverage, and sensitive expression penalties), it achieves an overall improvement in language output quality, enhancing the system's applicability in tasks such as actual customer service conversations and business guidance. In addition, this technical solution supports seamless integration with large open source models (such as GPT, GLM, and LLaMA), without modifying the original model parameters. The training process relies solely on reinforcement learning of the policy network, significantly reducing the cost of model deployment and maintenance difficulty.
[0138] More importantly, the reward function parameters are highly customizable and interpretable. The weights can be flexibly configured according to different business scenarios, and the optimization objectives can be dynamically adjusted based on actual operating data, thereby ensuring that the model achieves the optimal balance between compliance, accuracy and user experience.
[0139] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0140] In one embodiment, a term alignment language generation device is provided, which corresponds to the term alignment language generation method in the above embodiment. Figure 8 As shown, the term alignment language generation device includes: a data acquisition module 101, a data input module 102, a data construction module 103, a data scoring module 104, a data screening module 105, and a data generation module 106. The functional modules are described in detail as follows:
[0141] The data collection module 101 is used to collect user input and context information in the target domain, and build input environment data for language generation by combining it with a preset terminology knowledge base;
[0142] A data input module 102 is configured to input the input environment data into a language generation model to generate a plurality of candidate language responses to obtain a candidate response set;
[0143] A data construction module 103 is configured to construct a term alignment reward function based on the candidate response set; wherein the term alignment reward function includes a term matching score, a language fluency score, an information coverage score, and a sensitive expression penalty score;
[0144] a data scoring module 104, configured to score each candidate language response in the candidate response set according to the term alignment reward function to obtain a scoring result;
[0145] A data screening module 105 is configured to screen the candidate language response with the highest score as the final output response based on the scoring result;
[0146] The data generation module 106 is configured to update the strategy parameters of the language generation model using the final output response to obtain a language generation result.
[0147] In one embodiment, the data construction module 103 is specifically configured to:
[0148] Determining a set of terms that should appear in the candidate language response based on the current task intent as an expected set of terms for term matching;
[0149] Perform word segmentation and syntactic analysis on the candidate language response to obtain a set of semantic blocks;
[0150] Comparing the semantic block set with the terminology knowledge base, determining whether each semantic block is a standard term defined in the terminology knowledge base, and obtaining a terminology hit result set;
[0151] A ratio calculation is performed between the term hit result set and the expected term set to obtain a term matching score.
[0152] In one embodiment, the data construction module 103 is further configured to:
[0153] Calculating a perplexity index of the candidate language response using a language model;
[0154] Alternatively, based on a syntactic analysis and language quality detection tool, the candidate language response is subjected to a syntactic completeness check and grammatical error identification to obtain a corresponding language fluency score;
[0155] The confusion index or the language fluency score is normalized to a preset range to obtain a normalized language fluency score.
[0156] The calculating the perplexity index of the candidate language response by using the language model includes:
[0157] Splitting the candidate language response into word sequences and inputting the word sequences into a pre-trained language model;
[0158] Based on the pre-trained language model, sequentially calculating the conditional probability of each word in the word sequence under the given condition of its preceding word;
[0159] The logarithmic mean of the conditional probability is calculated, and a perplexity value is calculated with a negative value as an exponential power; wherein the perplexity value is the perplexity index.
[0160] In one embodiment, the data construction module 103 is further configured to:
[0161] Based on the intent recognition and slot extraction model, key intent slots are extracted from user input to obtain a set of information elements;
[0162] Inputting the candidate language response into a slot extraction module to identify semantic slot elements in the candidate language response;
[0163] Comparing the information element set with the semantic slot elements to obtain the number of hit slots;
[0164] The ratio between the number of hit slots and the total number of slots is calculated to obtain an information coverage score.
[0165] In one embodiment, the data construction module 103 is further configured to:
[0166] Build a list of sensitive words;
[0167] Input the candidate language response into the sensitive term detection module, and determine whether the candidate language response contains a term in the sensitive word list; if so, assign a corresponding risk penalty weight based on the risk level of each sensitive term;
[0168] The risk penalty weights are accumulated to obtain a sensitive expression penalty score.
[0169] In one embodiment, the data construction module 103 is further configured to:
[0170] Preset weight configurations for term matching, language fluency, information coverage, and sensitive expression penalties for each business scenario;
[0171] The weight configuration is weightedly combined with the components of the term alignment reward function, and the weight configuration is dynamically adjusted based on a preset strategy optimization mechanism to adapt to the actual performance of the results generated by the language generation model at different stages.
[0172] For the specific definition of the term alignment language generation device, please refer to the definition of the term alignment language generation method above, which will not be repeated here. The various modules in the above-mentioned term alignment language generation device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0173] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 9 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of a term alignment language generation method.
[0174] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 10 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of a term alignment language generation method.
[0175] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed, the steps provided in the above embodiment can be implemented.
[0176] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0177] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0178] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0179] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A term alignment language generation method, characterized in that: include: Collect user input and context information in the target domain, combine it with the preset terminology knowledge base, and construct input environment data for language generation; Inputting the input environment data into a language generation model to generate a plurality of candidate language responses to obtain a candidate response set; Based on the candidate response set, constructing a term alignment reward function; wherein the term alignment reward function includes a term matching score, a language fluency score, an information coverage score, and a sensitive expression penalty score; Scoring each candidate language response in the candidate response set according to the term alignment reward function to obtain a scoring result; Based on the scoring results, selecting the candidate language response with the highest score as the final output response; The final output response is used to update the strategy parameters of the language generation model to obtain a language generation result.
2. The term alignment language generation method according to claim 1, characterized in that: The constructing of the term alignment reward function includes: Determining a set of terms that should appear in the candidate language response based on the current task intent as an expected set of terms for term matching; Perform word segmentation and syntactic analysis on the candidate language response to obtain a set of semantic blocks; Comparing the semantic block set with the terminology knowledge base, determining whether each semantic block is a standard term defined in the terminology knowledge base, and obtaining a terminology hit result set; A ratio calculation is performed between the term hit result set and the expected term set to obtain a term matching score.
3. The term alignment language generation method according to claim 1, characterized in that: The constructing of the term alignment reward function further includes: Calculating a perplexity index of the candidate language response using a language model; Alternatively, based on a syntactic analysis and language quality detection tool, the candidate language response is subjected to a syntactic completeness check and grammatical error identification to obtain a corresponding language fluency score; The confusion index or the language fluency score is normalized to a preset range to obtain a normalized language fluency score.
4. The term alignment language generation method according to claim 3, characterized in that: The calculating the perplexity index of the candidate language response by using the language model includes: Splitting the candidate language response into word sequences and inputting the word sequences into a pre-trained language model; Based on the pre-trained language model, sequentially calculating the conditional probability of each word in the word sequence under the given condition of its preceding word; The logarithmic mean of the conditional probability is calculated, and a perplexity value is calculated with a negative value as an exponential power; wherein the perplexity value is the perplexity index.
5. The term alignment language generation method according to claim 1, characterized in that: The constructing of the term alignment reward function further includes: Based on the intent recognition and slot extraction model, key intent slots are extracted from user input to obtain a set of information elements; Inputting the candidate language response into a slot extraction module to identify semantic slot elements in the candidate language response; Comparing the information element set with the semantic slot elements to obtain the number of hit slots; The ratio between the number of hit slots and the total number of slots is calculated to obtain an information coverage score.
6. The term alignment language generation method according to claim 1, characterized in that: The constructing of the term alignment reward function further includes: Build a list of sensitive words; Input the candidate language response into the sensitive term detection module, and determine whether the candidate language response contains a term in the sensitive word list; if so, assign a corresponding risk penalty weight based on the risk level of each sensitive term; The risk penalty weights are accumulated to obtain a sensitive expression penalty score.
7. The term alignment language generation method according to claim 1, characterized in that: The constructing of the term alignment reward function further includes: Preset weight configurations for term matching, language fluency, information coverage, and sensitive expression penalties for each business scenario; The weight configuration is weightedly combined with the components of the term alignment reward function, and the weight configuration is dynamically adjusted based on a preset strategy optimization mechanism to adapt to the actual performance of the results generated by the language generation model at different stages.
8. A term alignment language generation device, characterized in that: include: The data collection module is used to collect user input and context information in the target domain, and build input environment data for language generation by combining it with the preset terminology knowledge base; a data input module, configured to input the input environment data into a language generation model to generate a plurality of candidate language responses to obtain a candidate response set; A data construction module, configured to construct a term alignment reward function based on the candidate response set; wherein the term alignment reward function includes a term matching score, a language fluency score, an information coverage score, and a sensitive expression penalty score; a data scoring module, configured to score each candidate language response in the candidate response set according to the term alignment reward function to obtain a scoring result; A data screening module, configured to screen the candidate language response with the highest score as the final output response based on the scoring result; The data generation module is used to update the strategy parameters of the language generation model using the final output response to obtain a language generation result.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the term alignment language generation method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the term alignment language generation method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Model optimization method and device, electronic equipment, storage medium and program product
CN121503735A
Large model intention understanding fine tuning method and device based on reinforcement learning
CN121787595A