Model alignment method and apparatus

CN122817867APending Publication Date: 2026-09-25ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610969937.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,在这种技术方案中,人工评测的效率较低,并且由于不同评审员标准不统一,导致评测结果具有一定主观性

Benefits of technology

[0019]由以上技术方案可知,本说明书实施例提供的模型对齐方法和设备,一方面,通过目标服务场景对应的多个用户的用户相关数据构建评测样本集,基于评测样本集通过策略模型生成针对用户的一组候选服务输出,能够将目标服务场景的线上用户行为数据反馈至评测集构建;另一方面,基于奖励模型对各个候选服务输出进行多个评测标准维度的评测,确定各个候选服务输出对应的奖励值,能够通过预设的多个评测标准维度高效准确地对大模型针对目标服务场景的服务效果进行客观评测,避免了人工评测效率低以及主观性强的问题;再一方面,基于评测得到的奖励值对策略模型的模型参数进行调整,能够根据评测结果对策略模型进行优化,使得针对策略模型的评测效果与策略模型针对目标服务场景的服务效果保持一致,进而实现通过模型评测驱动服务增长。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817867A_ABST
    Figure CN122817867A_ABST
Patent Text Reader

Abstract

The embodiment of the specification provides a model alignment method and device. The method comprises the following steps: constructing an evaluation sample set through user-related data of a plurality of users corresponding to a target service scenario, generating a group of candidate service outputs for a user through a strategy model based on the evaluation sample set; evaluating each candidate service output in a plurality of evaluation standard dimensions based on a reward model to determine a reward value corresponding to each candidate service output; and adjusting model parameters of the strategy model based on the reward value obtained through the evaluation, so as to optimize the strategy model according to the evaluation result, so that the evaluation effect of the strategy model is consistent with the service effect of the strategy model for the target service scenario, thereby realizing service growth driven by model evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of large model technology, and in particular to a model alignment method and electronic device. Background Technology

[0002] With the development of artificial intelligence technology, the application of LLMs (Large Language Models) is becoming increasingly widespread. How to better apply large models to specific service scenarios has become a focus of attention.

[0003] In related technical solutions, large-scale models are applied to target service scenarios, such as marketing scenarios. A team of human reviewers evaluates the service effectiveness of the large-scale model; for example, the human review team scores the messages generated by the model, and the model parameters are adjusted based on reinforcement learning and feedback. However, in this technical solution, human evaluation is inefficient, and the results are somewhat subjective due to inconsistent standards among different reviewers.

[0004] Therefore, how to efficiently and accurately evaluate the service effect of large models for target service scenarios has become an urgent technical problem to be solved.

[0005] The information in the background section is merely information known only to the inventor and does not imply that such information had entered the public domain before the date of this application, nor does it imply that it can be considered prior art in this disclosure. Summary of the Invention

[0006] This manual provides a model alignment method and electronic device that can efficiently and accurately evaluate the service effect of large models for target service scenarios based on preset standards.

[0007] Firstly, this specification provides a model alignment method applied to a preset language model, the preset language model including a policy model and a reward model, the method comprising: Determine the evaluation sample set corresponding to the target service scenario, wherein the evaluation sample set includes user-related data of multiple users; Based on the evaluation sample set, a set of candidate service outputs for the user is generated through the strategy model; The reward model is used to evaluate each candidate service output across multiple evaluation criteria dimensions to determine the reward value corresponding to each candidate service output; and The model parameters of the strategy model are adjusted based on the reward value.

[0008] In some example embodiments, based on the above scheme, the multiple evaluation criteria dimensions include bottom line, compliance, professionalism and expressiveness, and the reward model includes multiple sub-models, including a bottom line sub-model, a compliance sub-model, a professionalism sub-model and an expressiveness sub-model.

[0009] In some example embodiments, based on the above scheme, the step of evaluating each candidate service output using the reward model across multiple evaluation criteria dimensions to determine the reward value corresponding to each candidate service output includes: The sub-reward value corresponding to the output of each candidate service is determined based on each of the sub-models of the reward model; The reward value corresponding to each candidate service output is obtained by performing a weighted summation operation on each of the sub-reward values.

[0010] In some example embodiments, based on the above scheme, determining the evaluation sample set corresponding to the target service scenario includes: Extract the user-related data of the multiple users from the service logs corresponding to the target service scenario; Based on the user-related data of the multiple users, multiple evaluation sets corresponding to the target service scenario are constructed according to a preset classification standard, wherein the preset classification standard is a classification standard determined based on the service conversion path of the target service scenario.

[0011] In some example embodiments, based on the above scheme, the preset classification standard is a preset funnel classification standard, and the construction of multiple evaluation sets corresponding to the target service scenario based on the user-related data of multiple users according to the preset classification standard includes: The user-related data of the multiple users are classified based on the preset funnel classification criteria; and Based on the classification results, construct the multiple evaluation sets corresponding to the target service scenario.

[0012] In some example embodiments, based on the above scheme, the candidate service output includes candidate opening lines, and the preset funnel classification criteria include opening lines being exposed but not adopted, opening lines being exposed and adopted but the user did not speak, and opening lines being exposed and adopted and the user spoke.

[0013] In some example embodiments, based on the above scheme, adjusting the model parameters of the strategy model based on the reward value includes: The relative advantage value within the group corresponding to each of the candidate service outputs is determined based on the reward value corresponding to each of the candidate service outputs. Based on the reward value corresponding to each candidate service output and the relative advantage value within the group, the alignment model loss of the strategy model is determined; and The model parameters of the policy model are adjusted based on the alignment model loss.

[0014] In some example embodiments, based on the above scheme, before adjusting the model parameters of the strategy model based on the reward value, the method further includes: Based on the reward value corresponding to each of the candidate service outputs, the set of candidate service outputs are sampled through dynamic sampling.

[0015] In some example embodiments, based on the above scheme, the alignment model loss includes a policy change term, which measures the change of the new policy function corresponding to the policy model compared to the old policy function corresponding to the policy model. Determining the alignment model loss based on the reward value corresponding to each candidate service output and the relative advantage value within the group includes: Determine the policy ratio between the new policy function and the old policy function; The strategy change term is determined based on the relative advantage value within the group corresponding to the candidate service output and the strategy ratio.

[0016] In some example embodiments, based on the above scheme, the alignment model loss further includes a pruning term, which is used to limit the change of the policy ratio when the policy ratio is greater than a predetermined range. Determining the policy change term based on the intra-group relative advantage value corresponding to the candidate service output and the policy ratio includes: If the strategy ratio is greater than the predetermined range, then the upper limit of the predetermined range is taken as the target strategy ratio, and the strategy change item is determined based on the relative advantage value within the group and the target strategy ratio; If the strategy ratio is less than the predetermined range, then the lower limit of the predetermined range is taken as the target strategy ratio, and the strategy change term is determined based on the relative advantage value within the group and the target strategy ratio.

[0017] In some example embodiments, based on the above scheme, determining the intra-group relative advantage value corresponding to each candidate service output based on the reward value corresponding to each candidate service output includes: The mean reward and standard deviation of the reward corresponding to the output of each candidate service are determined based on the reward value corresponding to the output of each candidate service. Based on the reward value, the mean reward, and the standard deviation of the reward corresponding to each candidate service output, the relative advantage value within the group corresponding to each candidate service output is determined.

[0018] Secondly, this specification also provides an electronic device, comprising: at least one storage medium storing at least one instruction set for performing model alignment processing; and at least one processor communicatively connected to the at least one storage medium, wherein, when the electronic device is running, the at least one processor reads the at least one instruction set and executes the model alignment method described in the first aspect of this specification according to the instructions of the at least one instruction set.

[0019] As can be seen from the above technical solutions, the model alignment method and device provided in the embodiments of this specification, on the one hand, construct an evaluation sample set through user-related data of multiple users corresponding to the target service scenario, and generate a set of candidate service outputs for users through a strategy model based on the evaluation sample set, which can feed back online user behavior data of the target service scenario to the construction of the evaluation set; on the other hand, evaluate each candidate service output based on a reward model using multiple evaluation standard dimensions, and determine the reward value corresponding to each candidate service output, which can efficiently and accurately evaluate the service effect of the large model for the target service scenario through multiple preset evaluation standard dimensions, avoiding the problems of low efficiency and strong subjectivity of manual evaluation; furthermore, adjust the model parameters of the strategy model based on the reward value obtained from the evaluation, which can optimize the strategy model according to the evaluation results, so that the evaluation effect of the strategy model is consistent with the service effect of the strategy model for the target service scenario, thereby realizing service growth driven by model evaluation.

[0020] Other functionalities of the model alignment methods and apparatus provided in this specification will be partially listed in the following description. The figures and examples described below will be readily apparent to those skilled in the art. The inventive aspects of the model alignment methods and apparatus provided in this specification can be fully understood through practice or use of the methods, apparatus, and combinations described in the detailed examples below. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 A schematic diagram of the implementation environment of a model alignment method provided in an embodiment of this specification is shown; Figure 2 A schematic diagram of the hardware structure of an electronic device 200 provided according to an embodiment of this specification is shown; Figure 3 A flowchart illustrating a model alignment method according to some embodiments of this specification is shown; Figure 4 A schematic diagram of a reinforcement learning process provided according to some embodiments of this specification is shown; Figure 5 A flowchart illustrating a model alignment method provided according to further embodiments of this specification is shown; and Figure 6 A flowchart illustrating a model alignment method provided according to some other embodiments of this specification is shown. Detailed Implementation

[0023] The following description provides specific application scenarios and requirements for this specification, intended to enable those skilled in the art to make and use the contents of this specification. Various partial modifications to the disclosed embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the embodiments shown, but rather to the widest scope consistent with the claims.

[0024] The terminology used herein is for the purpose of describing particular exemplary embodiments only and is not restrictive. For example, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” used herein may also include the plural forms. When used in this specification, the terms “comprising,” “including,” and / or “containing” mean that the associated integers, steps, operations, elements, and / or components are present, but do not exclude the presence of one or more other features, integers, steps, operations, elements, components, and / or groups, or that other features, integers, steps, operations, elements, components, and / or groups may be added to the system / method.

[0025] Considering the following description, these and other features of this specification, as well as the operation and function of the related components of the structure, and the economy of assembly and manufacture of the parts, can be significantly improved. All of these form part of this specification with reference to the accompanying drawings. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.

[0026] The flowcharts used in this specification illustrate operations implemented according to some embodiments of this specification. It should be clearly understood that the operations in the flowcharts may not be implemented in a sequential order. Instead, the operations may be implemented in reverse order or simultaneously. Furthermore, one or more additional operations may be added to the flowcharts. One or more operations may be removed from the flowcharts.

[0027] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0028] Preset Language Model: Includes large or small language models. Preset language models are general natural language processing models trained on massive amounts of data. They have complex reasoning and contextual understanding capabilities and can empower the decision-making and reflection processes of intelligent agents. Examples include Transformer-based neural network models, ChatGPT models, or DeepSeek models.

[0029] LLM (Large Language Model): refers to a generative artificial intelligence model trained on large-scale text corpora based on deep learning technology. It has the ability to understand and generate natural language, such as DeepSeek, LLaMA and other series of models.

[0030] Evaluation set: A structured dataset constructed to evaluate the output quality of large models, containing input context and expected output (or standard score), used to measure the model's performance on a specific task. Evaluation criteria: Define a multi-dimensional evaluation system for the quality of model output. The evaluation criteria for the target service scenario can include dimensions such as compliance, professionalism, expressiveness, and bottom line, which are the core basis of the automated evaluation system.

[0031] Evaluation standard alignment: The process of ensuring that the content generated by a large model is consistent with the preset evaluation standards through algorithmic mechanisms, usually achieved through reward models and reinforcement learning to achieve dynamic optimization.

[0032] Reward Model: An automated scoring system built on evaluation criteria, which uses a large language model to score candidate service outputs in multiple dimensions as a feedback signal in reinforcement learning.

[0033] DAPO (Decoupled Clip and Dynamic sAmpling Policy Optimization): An improved reinforcement learning policy optimization method that removes the KL divergence constraint from the traditional PPO (Proximal Policy Optimization) and introduces a dynamic sampling mechanism, making it suitable for learning high-value samples in long sequence generation tasks.

[0034] To enable LLM models to perform specific tasks in downstream target service scenarios, pre-trained LLM models need to be aligned. In related technical solutions, when applying the LLM model to a target service scenario, such as a marketing scenario, the service performance of the large model is evaluated. Only after passing the evaluation is the large model applied to the target service scenario. However, in this approach, it is difficult to ensure that the evaluation performance of the large model aligns with its performance in the target service scenario.

[0035] Based on the above, embodiments of this specification provide a model alignment method and an electronic device. On the one hand, an evaluation sample set is constructed using user-related data from multiple users corresponding to a target service scenario. Based on the evaluation sample set, a set of candidate service outputs for users is generated through a strategy model, which can feed back online user behavior data of the target service scenario to the construction of the evaluation set. On the other hand, based on a reward model, multiple evaluation criteria dimensions are used to evaluate each candidate service output, and the reward value corresponding to each candidate service output is determined. This allows for efficient and accurate objective evaluation of the service effect of the large model for the target service scenario through multiple preset evaluation criteria dimensions, avoiding the problems of low efficiency and strong subjectivity of manual evaluation. Furthermore, the model parameters of the strategy model are adjusted based on the reward value obtained from the evaluation, which can optimize the strategy model according to the evaluation results, so that the evaluation effect of the strategy model is consistent with the service effect of the strategy model for the target service scenario, thereby achieving service growth driven by model evaluation.

[0036] The technical solutions of the embodiments of this specification will now be described in detail with reference to the accompanying drawings.

[0037] Figure 1 A schematic diagram of the implementation environment of a model alignment method provided in an embodiment of this specification is shown.

[0038] See Figure 1 As shown, the implementation environment 100 may include a terminal 110, a server 130, and a database 140.

[0039] Terminal 110 is connected to server 130 via wireless network or wired network 120. Terminal 110 can be a tablet computer, laptop computer, or desktop computer, but is not limited to these.

[0040] Terminal 110 may store data or instructions for executing the model alignment method described in this specification. Terminal 110 may include hardware devices with data processing capabilities and the necessary programs to drive the hardware devices.

[0041] Server 130 is a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. Server 130 provides background services for applications running on terminal 110.

[0042] Server 130 is equipped with an integrated development platform (IDE). An IDE, also known as an integrated development environment, is an application that provides a program development environment, typically including tools such as a code editor, compiler, debugger, and graphical user interface. Developers can write program code (i.e., program development) on the IDE. The IDE server can be a computing device specifically designed to implement model alignment methods. Server 130 can communicate with both terminal 110 and database 140.

[0043] Furthermore, server 130 may store data or instructions for executing the model alignment method described in this specification. Server 130 may include hardware devices with data processing capabilities and the necessary programs to drive the hardware devices. Of course, server 130 may also be merely a hardware device with data processing capabilities, or simply a program running on the hardware device. In some embodiments, server 130 may also be deployed as a plug-in on terminal 110, in which case server 130 stores data or instructions for executing the model alignment method corresponding to terminal 110 described in this specification.

[0044] Database 140 may store data and / or instructions. In some embodiments, database 140 may store user-related data and evaluation sets, etc. In some embodiments, database 140 may store data and / or instructions executed by server 130 or used to execute the model alignment methods described herein. Terminal 110 and server 130 have access to database 140, and terminal 110 and server 130 may access data or instructions stored in database 140 via a network. In some embodiments, database 140 may be directly connected to terminal 110 and server 130. In some embodiments, database 140 may be part of server 130. In some embodiments, database 140 may include mass storage, removable storage, volatile read-write memory, read-only memory (ROM), or similar content, or any combination thereof. Exemplary mass storage may include non-transitory storage media such as disks, optical discs, and solid-state drives. Exemplary removable storage may include flash drives, floppy disks, optical discs, memory cards, zip disks, magnetic tapes, etc. Typical volatile read-write memory may include random access memory (RAM). Example RAMs may include dynamic RAM (DRAM), dual date rate synchronous dynamic RAM (DDRSDRAM), static RAM (SRAM), thyristor RAM (T-RAM), and zero-capacitance RAM (Z-RAM), etc. Exemplary ROMs may include mask ROM (MROM), programmable ROM (PROM), virtual programmable ROM (PEROM), electronically programmable ROM (EEPROM), optical disc ROM (CD ROM), and digital multifunction disk ROM, etc.

[0045] Those skilled in the art will understand that the number of terminals described above can be more or less. For example, there may be only one terminal, or there may be dozens or hundreds of terminals, or even more, in which case other terminals may also be included in the above implementation environment. This specification does not limit the number of terminals or the type of devices in the embodiments.

[0046] After introducing the implementation environment of the embodiments of this specification, the application scenarios of the embodiments of this specification will be described below in conjunction with the above implementation environment. In the following description, the terminal is also the terminal 110 in the above implementation environment, and the server is also the server 130 in the above implementation environment. The technical solutions provided by the embodiments of this specification can be applied in scenarios of model alignment of large models, such as intelligent customer service models, sales assistance models, and marketing content generation models.

[0047] Taking the application of the technical solution provided in the embodiments of this specification to the model alignment scenario of an intelligent customer service model as an example, the intelligent insurance customer service model includes a strategy model and a reward model. The strategy model is the model to be aligned and optimized, and the reward model is a learnable scoring model, such as a scoring function or a neural network model. Terminal 110 determines the evaluation sample set corresponding to the target service scenario. The evaluation sample set includes user-related data of multiple users. Based on the evaluation sample set, a set of candidate service outputs for users is generated through the strategy model. Based on the reward model, multiple evaluation criteria dimensions are used to evaluate each candidate service output to determine the reward value corresponding to each candidate service output. And the model parameters of the strategy model are adjusted based on the reward value.

[0048] It should be noted that the above description is based on the application of the technical solution provided in the embodiments of this specification to the model alignment scenario of the intelligent customer model. The technical solution provided in the embodiments of this specification can also be applied to other appropriate scenarios, such as the model alignment scenario of wealth management model or health consultation model, etc. The implementation process belongs to the same inventive concept as described above, and will not be repeated here.

[0049] It should be noted that the steps in the model alignment method in the example embodiments of this specification may be partially performed by the client, partially performed by the server, or entirely performed by the server or entirely by the client. This specification does not impose any special limitations on this.

[0050] based on Figure 1 The implementation environment shown below will be combined with... Figures 2-6 This specification provides a detailed description of the model alignment method and electronic device provided in the embodiments. It should be noted that the above-described implementation environment is shown only to facilitate understanding of the spirit and principles of this specification, and the embodiments of this specification are not limited in any way. Rather, the embodiments of this specification can be applied to any applicable scenario.

[0051] Figure 2 This is a schematic diagram of an electronic device 200 provided according to some embodiments of this specification. The electronic device 200 can perform the model alignment method described in this specification. The model alignment method is described in other parts of this specification. The electronic device 200 can be a general-purpose computer or a special-purpose computer. For example, the electronic device 200 can be a server, a personal computer, a portable computer (such as a laptop computer, tablet computer, etc.), or other electronic devices with computing capabilities. Of course, the electronic device can be... Figure 1 The terminal 110 and / or server 130 can also be terminal devices used by multiple developers to develop programs on an integrated development platform.

[0052] The electronic device described in this specification may include one or more of the following components: processor 210, memory 220, input device 230, output device 240, and bus 250. The processor 210, memory 220, input device 230, and output device 240 may be connected to each other via bus 250.

[0053] Processor 210 may include one or more processing cores. Processor 210 connects to various parts of the electronic device using various interfaces and lines, and executes the model alignment method described in this specification by running or executing instructions, programs, code sets, or instruction sets stored in memory 220, and by calling data stored in memory 220. Optionally, processor 210 may be implemented using at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). Processor 210 may integrate one or more of a central processing unit (CPU), graphics processing unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 210 and may be implemented separately using a communication chip.

[0054] The memory 220 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 220 may include non-transitory computer-readable storage medium. The memory 220 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 220 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (e.g., touch function, sound playback function, image playback function, etc.), instructions for implementing the various method embodiments described below, etc. The operating system may be an Android system, including systems deeply developed based on the Android system, an iOS system, including systems deeply developed based on the iOS system, or other systems.

[0055] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to establish data communication between the third-party applications and the operating system. This would allow the operating system to obtain the current scenario information of the third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.

[0056] The input device 230 is used to receive input instructions or data, and includes, but is not limited to, a keyboard, mouse, camera, microphone, or touch device. The output device 240 is used to output instructions or data, and includes, but is not limited to, a display device and a speaker. In one example, the input device 230 and the output device 240 can be combined, and both the input device 230 and the output device 240 can be a touch display screen.

[0057] In addition, those skilled in the art will understand that the structure of the electronic device shown in the above figures does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the electronic device may also include radio frequency circuits, input units, sensors, audio circuits, Wireless Fidelity (WiFi) modules, power supplies, Bluetooth modules, etc., which will not be described in detail here.

[0058] Figure 3 A flowchart of a model alignment method provided according to an embodiment of this specification is shown. As previously described, the electronic device 200 can execute the model alignment method of the embodiment of this specification. Specifically, the processor 210 can read the instruction set stored in its local storage medium and then execute the model alignment method of the embodiment of this specification according to the instructions of the instruction set. Hereinafter, steps S310 to S340 of the model alignment method will be described in detail with reference to the accompanying drawings.

[0059] Reference Figure 3 As shown, in step S310, the evaluation sample set corresponding to the target service scenario is determined. The evaluation sample set includes user-related data of multiple users.

[0060] In the example embodiment, the target service scenario refers to the specific service scenario to be optimized, such as an insurance customer service script generation scenario or a marketing script generation scenario. The evaluation sample set is a set of structured samples used to evaluate and train the strategy model. Each sample represents a typical service interaction context. The evaluation sample set includes user-related data for multiple users within the target service scenario. User-related data includes basic user profiles, user behavior data, and user dialogue data. Basic user profiles include static user attribute information such as age, city, occupation, and social security type. User behavior data includes user operation records, browsing history, and click conversion behavior logs on the service platform, such as page views, click behavior, usage frequency, and search behavior. User dialogue data includes historical dialogue text or structured interaction records between users and customer service, intelligent assistants, or agents.

[0061] Electronic device 200 constructs a corresponding hierarchical evaluation sample set based on the service objectives of the target service scenario. For example, electronic device 200 transforms the original service logs of the target service scenario into multiple evaluation sets graded according to "user behavior quality", with each evaluation set corresponding to a standard rating level such as low / medium / high.

[0062] In some example embodiments, the electronic device 200 extracts user-related data of multiple users from the service logs corresponding to the target service scenario; based on the user-related data of multiple users, it constructs multiple evaluation sets corresponding to the target service scenario according to a preset classification standard. The preset classification standard refers to the classification standard determined based on the service conversion path of the target service scenario. For example, before constructing the evaluation set, a set of objective, operable, and automatically judged rules is predefined according to the core conversion path of the target service scenario, such as exposure-adoption-conversion, etc., to divide the original user-related data into several ordered levels (such as low / medium / high) evaluation subsets. For example, evaluation criteria can be determined based on the service objectives of the target service scenario. Taking the target service scenario as the opening dialogue scenario, the evaluation criteria are used to judge the quality of the opening dialogue. The conversion path of the service scenario includes two key behavioral nodes. The first key behavioral node is whether the opening dialogue can be used directly; the second key behavioral node is whether the opening dialogue can move the user and guide the user to speak. Therefore, according to the above evaluation criteria, the user-related data of multiple users can be divided into three evaluation sets. For example, evaluation set 1: the opening dialogue is exposed but not adopted; 2: the opening dialogue is exposed and adopted but the user does not speak; 3: the opening dialogue is exposed and adopted and the user speaks.

[0063] Furthermore, in the example embodiment, the preset classification standard is a preset funnel classification standard. The funnel is used to represent the multi-stage conversion path of a user from contacting the service to completing the service goal, such as exposure-adoption-conversion. The preset funnel classification standard refers to a predefined multi-level sample classification rule based on key behavioral nodes of the service conversion path in the target service scenario. It is used to divide user interaction data into multiple levels of evaluation sets. The evaluation sets are sorted according to the depth of the funnel. The later the level, the higher the service value and the higher the standard score. Electronic device 200 classifies user-related data of multiple users based on the preset funnel classification standard; and constructs multiple evaluation sets corresponding to the target service scenario based on the classification results. The standard scores corresponding to each evaluation set in the multiple evaluation sets increase sequentially. For example, for the target service scenario, the user conversion funnel is defined to include The funnel consists of an ordered behavior phase: S1→S2→…→Sk, where S1 is the initial exposure and Sk is the final successful conversion. The preset funnel classification standard is a set of Boolean condition functions, which are true when the user behavior meets the completion condition of phase k, and false otherwise.

[0064] For example, candidate service outputs include candidate opening lines. Preset funnel classification criteria include opening lines being exposed but not adopted, opening lines being exposed and adopted but the user did not speak, and opening lines being exposed and adopted and the user spoke. For instance, assuming the target service scenario is insurance customer service opening lines, evaluation set 1: opening lines being exposed but not adopted; 2: opening lines being exposed and adopted but the user did not speak; 3: opening lines being exposed and adopted and the user spoke. Electronic device 200 parses the service logs corresponding to the target scenario service, extracting user basic profiles, user behavior data, and user dialogue data from the parsing results. Based on the preset funnel classification criteria, it maps each user's relevant data samples to three stages: S1, S2, and S3. For example, S1: opening lines being exposed but not adopted; S2: opening lines being exposed and adopted but the user did not speak; S3: opening lines being exposed and adopted and the user spoke.

[0065] In step S320, based on the evaluation sample set, a set of candidate service outputs for the user is generated through the strategy model.

[0066] In the example implementation, the pre-set language model includes a policy model. The policy model is responsible for generating candidate service outputs, such as candidate opening statements, based on the input evaluation samples and prompts. The policy model is typically a pre-trained large language model, such as the DeepSeek model or the ChatGPT model. Large language models can employ the Transformer architecture or its variants (such as GPT, BERT, etc.), which utilizes an attention mechanism to globally model sequential data, efficiently handling long-distance dependencies and thus performing well in natural language tasks. Large language models learn the statistical features and semantic relationships of language through pre-training on large-scale corpora, giving them good generalization capabilities. The core capabilities of large language models include, but are not limited to: understanding contextual semantics, generating coherent and grammatically correct text, performing logical reasoning, and handling multi-task scenarios. Their usage typically includes two modes: direct inference and fine-tuning. In direct inference mode, the user guides the large language model to generate specific outputs by designing prompts. Cue words can be task descriptions or instructions in text form, used to stimulate the semantic understanding and generation capabilities of the large language model. In fine-tuning mode, the large language model is further trained on small-scale datasets in a specific domain to optimize its performance on specific tasks.

[0067] In an example embodiment, the electronic device 200 generates multiple candidate service outputs for each user in the evaluation sample set based on that user's relevant data and target prompts using a strategy model. Each candidate service output can be a candidate opening statement, such as a marketing statement or a recommendation statement. Taking an insurance customer service model as an example with a preset language model, the evaluation sample can include the user-related data of various insurance users. The strategy model of the insurance customer service model generates a marketing statement for the current insurance user based on the user-related data and target prompts.

[0068] For example, for each user in each evaluation set, the electronic device 200 generates multiple candidate service outputs (e.g., 4-16 candidate service outputs) from the policy model through random sampling. These multiple candidate service outputs constitute a group G. Furthermore, the diversity and quality of the candidate service outputs can be controlled by adjusting the sampling parameters. For example, the electronic device 200 often uses techniques such as temperature sampling, Top-k sampling, or Nucleus sampling (Top-p) to avoid generating identical answers. In temperature sampling, the temperature parameter controls the randomness of the generated candidate service outputs; the higher the temperature value, the more random the output; the lower the temperature value, the more concentrated the output. Top-k sampling selects the next term from the k terms with the highest probability, reducing long-tail noise.

[0069] In step S330, the output of each candidate service is evaluated based on multiple evaluation criteria dimensions according to the reward model, and the reward value corresponding to each candidate service output is determined.

[0070] In the example embodiment, the preset language model includes a reward model. The reward model (RM) is a pre-trained model used to evaluate the quality of candidate service outputs generated by the strategy model. The reward model is trained based on labeled sample data and outputs a scalar reward value representing the quality or service effect of the candidate service output. The electronic device 200 inputs each candidate service output into the reward model and evaluates each candidate service output across multiple evaluation criteria dimensions based on the reward model to obtain the reward value corresponding to that candidate service output. The reward value guides the optimization of the strategy model, generating higher-quality candidate service outputs; a higher reward value indicates a higher quality candidate service output.

[0071] In some example implementations, corresponding evaluation criteria dimensions are set for different target service scenarios. For instance, if the target service scenario is an insurance customer service script generation scenario, and the service goal is to increase the first-time speaking rate of dormant users, then the multiple evaluation criteria dimensions for this target service scenario include basic compliance, compliance, professionalism, and expressiveness. Specifically, the basic compliance evaluation criterion dimension, based on legal regulations, checks whether candidate service outputs, such as candidate scripts, contain issues related to pornography, terrorism, or politics. The compliance evaluation criterion dimension, based on industry risk control rule bases, such as the insurance industry risk rule base, checks whether candidate service outputs, such as candidate scripts, contain illegal sensitive words or expressions such as exaggerated promises or inducing policy cancellations. The professionalism evaluation criterion dimension, based on a professional knowledge base, checks whether the professional knowledge contained in candidate service outputs, such as candidate scripts, is correct, such as market interest rates, product terms, product yields, and product coverage. The expressiveness evaluation criterion dimension checks whether the language style of candidate service outputs, such as candidate scripts, conforms to common expression habits, ensuring that the text is fluent and easy to understand.

[0072] In other example embodiments, if the target service scenario is a financial marketing script generation scenario, and the service goal is to prompt users to click "view details" or book an appointment with an advisor, then the multiple evaluation criteria dimensions for this target service scenario include bottom-line compliance, compliance, accuracy of return description, and completeness of risk disclosure.

[0073] It should be noted that although the above evaluation criteria dimensions have been used as examples, those skilled in the art should understand that the evaluation criteria dimensions for the target service scenario can also be other appropriate dimensions, such as guidance or personalization, which are also within the scope of the embodiments in this specification.

[0074] According to the technical solutions in the above example embodiments, setting corresponding evaluation standard dimensions for the target service scenario can transform abstract service goals into a quantifiable, executable, and optimizable multi-dimensional evaluation system, thereby improving the efficiency and stability of subsequent reinforcement learning training.

[0075] Furthermore, the reward model can be pre-trained using evaluation sample sets, such that the reward values ​​corresponding to each evaluation set in multiple evaluation sets increase sequentially. The model structure of the reward model can be a classification model (outputting binary classification labels) or a regression model (outputting continuous reward values). Suppose the target service scenario is insurance customer service opening scripts. The reward model is trained using the following evaluation sets 1 to 3: Evaluation set 1: Opening scripts are exposed but not adopted; Evaluation set 2: Opening scripts are exposed and adopted but the user does not speak; Evaluation set 3: Opening scripts are exposed, adopted, and the user speaks. Therefore, the reward value corresponding to evaluation set 1 is less than the reward value corresponding to evaluation set 2, and the reward value corresponding to evaluation set 2 is less than the reward value corresponding to evaluation set 3.

[0076] The reward model can be structured as a classification model (outputting binary labels) or a regression model (outputting continuous reward values). The reward model generates reward values ​​by integrating multiple reward signals. For example, electronic device 200 inputs each candidate service output into the reward model, which generates the reward value corresponding to that candidate service output by weighted combination of multiple reward sub-models.

[0077] In the example embodiment, the reward model includes multiple sub-models, including a baseline sub-model, a compliance sub-model, a professional sub-model, and an expressive sub-model. The baseline sub-model is a hard constraint sub-model, and the sub-model can be a distribution function with a value of 0 or 1. The electronic device 200 determines the sub-reward value corresponding to each candidate service output based on each sub-model of the reward model; it then performs a weighted summation operation on each sub-reward value to obtain the reward value corresponding to each candidate service output. For example, the reward model can be structured as shown in equation (1): score_overall = N1*score_bottom line + N2*score_compliance + N3*score_professionalism + N4*score_expressiveness (1) Wherein, `score_overall` represents the reward value corresponding to the candidate service output; `score_bottom-line` represents the sub-reward value corresponding to the bottom-line sub-model; `score_compliance` represents the sub-reward value corresponding to the compliance sub-model; `score_professionalism` represents the sub-reward value corresponding to the professionalism sub-model; and `score_expressiveness` represents the sub-reward value corresponding to the expressiveness sub-model. `N1`, `N2`, `N3`, and `N4` represent the hyperparameters of different evaluation criteria dimensions for each sub-model, which need to be adjusted based on empirical values. The hyperparameters corresponding to the bottom-line sub-model are greater than those of the other sub-models, meaning bottom-line performance is a hard constraint. The sub-reward values ​​of the sub-models can be distributed between 0 and 1.

[0078] According to the technical solution in the above example embodiment, on the one hand, the reward model is split into multiple sub-models of evaluation standard dimensions, which evaluate different evaluation standard dimensions respectively, and the final reward value is obtained by weighted summation. This can enhance the maintainability and service adaptation flexibility of the model through dimension decoupling, and achieve a flexible balance of multi-objective optimization. On the other hand, setting the bottom-line sub-model as a hard constraint sub-model can ensure bottom-line security and ensure the security of generated content.

[0079] In step S340, the model parameters of the strategy model are adjusted based on the reward value.

[0080] In an example embodiment, the electronic device 200 updates the parameters of the policy model based on the reward value corresponding to the candidate service output obtained from the reward model, making the policy model more inclined to generate high-reward outputs in the future. For example, the electronic device 200 determines the loss function of the policy model based on the reward value, calculates the gradient of the loss function through backpropagation, and updates the parameters of the policy model using an optimizer such as Adam through gradient. For example, electronic device 200 differentiates the loss function, i.e., the objective function, of the policy model to obtain the model parameters corresponding to the loss function. The gradient is used to update the model parameters of the policy model via gradient descent. For example, the loss function of the policy model in electronic device 200 is used to update the model parameters of the policy model using the Adam optimizer.

[0081] according to Figure 3 The technical solution in the example embodiment, on the one hand, constructs an evaluation sample set through user-related data of multiple users corresponding to the target service scenario, and generates a set of candidate service outputs for users through a strategy model based on the evaluation sample set, which can feed back online user behavior data of the target service scenario to the construction of the evaluation set; on the other hand, it evaluates each candidate service output based on multiple evaluation criteria dimensions based on a reward model, and determines the reward value corresponding to each candidate service output, which can efficiently and accurately evaluate the service effect of the large model for the target service scenario through multiple preset evaluation criteria dimensions, avoiding the problems of low efficiency and strong subjectivity of manual evaluation; furthermore, it adjusts the model parameters of the strategy model based on the reward value obtained from the evaluation, which can optimize the strategy model according to the evaluation results, so that the evaluation effect of the strategy model is consistent with the service effect of the strategy model for the target service scenario, thereby realizing service growth driven by model evaluation.

[0082] Figure 4 A schematic diagram of a reinforcement learning process provided according to some embodiments of this specification is shown.

[0083] Reference Figure 4 As shown, in step S410, the relative advantage value within the group corresponding to each candidate service output is determined based on the reward value corresponding to each candidate service output.

[0084] In some example embodiments, the intra-group relative advantage value represents the standardized reward value of a candidate service output within the group, reflecting its advantage relative to the group average performance. Through group-wise comparison and standardization, the reward values ​​corresponding to each candidate service output are converted into intra-group relative advantage values, thereby providing a stable learning signal for policy optimization. The electronic device 200 determines the intra-group relative advantage value for each candidate service output based on its reward value within the group, using intra-group comparison.

[0085] Furthermore, in the example embodiment, the electronic device 200 determines the mean reward and standard deviation of a set of candidate service outputs based on the reward value corresponding to each candidate service output; and determines the intra-group relative advantage value corresponding to each candidate service output based on the reward value, mean reward, and standard deviation of each candidate service output. This is shown in equation (2) below: (2) Where Ai represents the i-th candidate service output The corresponding intra-group relative advantage value, R The reward value output for the i-th candidate service. and Here, Ai represents the mean and standard deviation of a set of reward values, respectively. A value greater than 0 indicates that the reward of the candidate service output is higher than the group average, while a value less than 0 indicates that the reward of the candidate service output is lower than the group average. The standardized Ai reflects the relative quality of the candidate service output within the group.

[0086] According to the technical solution in the above example embodiment, the reward value is standardized into a relative advantage value Ai by comparing within groups. This relative advantage value guides policy updates, making the model more inclined to generate high-advantage outputs rather than relying on a single reward value from candidate service outputs. The advantage value calculated through within-group reward normalization reflects the superiority or inferiority of a single candidate service output relative to the group average. The relative advantage value can replace traditional value networks, dynamically estimating the baseline through within-group comparisons, reducing variance, and improving training stability.

[0087] Furthermore, in the example embodiment, before adjusting the model parameters of the policy model based on the reward value, the electronic device 200 samples a set of candidate service outputs using dynamic sampling based on the reward value corresponding to each candidate service output. Dynamic sampling requires that the sampled set of responses simultaneously include both correct and incorrect responses to ensure that each batch has a valid gradient signal.

[0088] For example, electronic device 200 determines the statistical difference, such as standard deviation, of the reward values ​​corresponding to the reward values ​​of each candidate service output based on the reward value of each candidate service output. If the statistical difference is greater than a predetermined threshold, the candidate service output is retained. If the statistical difference is less than or equal to the predetermined threshold, the candidate service output is discarded.

[0089] For example, the standard deviation of the group of candidate service outputs is determined based on the reward value corresponding to each candidate service output; or, the intra-group relative advantage value corresponding to each candidate service output is determined based on the reward value corresponding to each candidate service output. If the intra-group relative advantage value corresponding to each candidate service output is 0, then the group of candidate service outputs is discarded.

[0090] According to the technical solution in the above example embodiment, a set of candidate service outputs is sampled by dynamic sampling, and only the output batches with different reward values ​​are processed. This ensures that the model learns on high-value output samples, rather than on low-value samples that it "already knows" or "does not know" at all, thereby improving model training efficiency and enhancing model performance.

[0091] In step S420, the alignment model loss of the strategy model is determined based on the reward value corresponding to each candidate service output and the relative advantage value within the group.

[0092] In the example embodiment, the alignment model loss represents the objective function used to optimize the model parameters of the policy model. The alignment model loss includes a policy change term, which measures the change in the new policy function corresponding to the policy model compared to the old policy function. The new policy function represents the policy function of the policy model after updating its model parameters; the old policy function represents the policy function of the policy model before updating its model parameters. The electronic device 200 determines the policy ratio of the new policy function to the old policy function; and determines the policy change term based on the intra-group relative advantage value corresponding to the candidate service output and the policy ratio of the new policy function to the old policy function.

[0093] According to the technical solution in the above example embodiment, the strategy change item is determined based on the relative advantage value within the group and the strategy ratio, and then the strategy model is updated, so that the direction of strategy change is consistent with the service goal and the dependence on the initial strategy model is reduced.

[0094] In other example embodiments, the alignment model loss also includes a pruning term, which is used to limit the change of the policy ratio when the policy ratio is greater than a predetermined range. If the policy ratio is greater than the predetermined range, the upper limit of the predetermined range is taken as the target policy ratio, and the policy change term is determined based on the relative advantage value within the group and the target policy ratio. If the policy ratio is less than the predetermined range, the lower limit of the predetermined range is taken as the target policy ratio, and the policy change term is determined based on the relative advantage value within the group and the target policy ratio.

[0095] For example, the pruning option is used when the strategy ratio is greater than a predetermined range. The parameter for the change in the ratio of the time-limited strategy The clipping range can be, for example, 0.1-0.2. If the strategy ratio is greater than the predetermined range, then the upper limit of the predetermined range will be used. The target policy ratio is used as the target policy ratio; if the policy ratio is less than a predetermined range, the lower limit of the predetermined range is used. As the target strategy ratio.

[0096] According to the technical solution in the above example embodiment, the change range of the policy ratio between the new policy function and the old policy function is limited by the pruning term, so as to prevent instability caused by excessive policy updates.

[0097] In step S430, the model parameters of the policy model are adjusted based on the alignment model loss.

[0098] In an example embodiment, the electronic device 200 calculates the gradient of the alignment model loss through backpropagation based on the alignment model loss, and updates the parameters of the policy model using an optimizer such as Adam. For example, the electronic device 200 differentiates the loss function of the alignment model loss to obtain the gradient of the model parameters corresponding to the loss function; and updates the model parameters of the policy model using gradient descent.

[0099] according to Figure 4 The technical solution in the example embodiment calculates the relative advantage value within the group based on the reward value of the candidate output, and then constructs the alignment model loss and updates the policy model parameters. This can achieve unsupervised high-quality behavior alignment and significantly improve training stability and convergence efficiency.

[0100] Figure 5 A schematic flowchart of a model alignment method provided according to some other embodiments of this specification is shown.

[0101] Reference Figure 5 As shown, in step S510, real samples of online users are sampled.

[0102] In an example embodiment, electronic device 200 extracts user-related data of sample users from online service logs to construct an online sample set. The user-related data includes basic user profiles, user behavior data, and user dialogue data.

[0103] In step S510, multiple candidate opening statements are generated using a strategy model.

[0104] In the example embodiment, the electronic device 200 inputs user-related data, such as basic user profiles, user behavior data, user dialogue data, and preset prompt words, into a strategy model to generate multiple candidate opening statements. The strategy model can use a large language model, such as DeepSeek, as its base model.

[0105] For example, for each user in each evaluation set, the electronic device 200 samples from the policy model to generate multiple candidate service outputs through random sampling. For example, for the same input sample, the policy model is used to perform multiple sampling and decoding to generate a set of candidate utterances {y_1, y_2, ..., y_N} containing N versions.

[0106] In step S530, the policy model is optimized through reinforcement learning.

[0107] In the example embodiment, the electronic device 200 evaluates candidate opening statements across multiple evaluation criteria dimensions using a reward model, obtaining reward values ​​corresponding to the candidate opening statements. Based on these reward values, the model parameters of the strategy model are adjusted using a group-relative strategy optimization approach. Group-relative strategy optimization directly utilizes N samples within the same sampling group to calculate relative advantage, increasing the generation probability of high-advantage samples (i.e., opening statements that meet the evaluation criteria) while suppressing low-scoring samples. Group-relative strategy optimization employs dynamic sampling: requiring that a sampled set of responses simultaneously contain both correct and incorrect responses to ensure that each batch has an effective gradient signal.

[0108] according to Figure 5The technical solution in the example embodiment, on the one hand, constructs a "layered evaluation set + service funnel mapping" mechanism, which structurally maps service data into multiple evaluation subsets through a funnel, such as exposure without adoption and adoption without approval, so that the evaluation can not only judge "good or bad", but also locate "where it is not good", providing a precise direction for model optimization; on the other hand, it constructs a multi-dimensional automated reward model for target service fields such as insurance, with a four-level evaluation standard system consisting of bottom line, compliance, professionalism and expressiveness, and uses a large model to realize automated scoring, solving the problems of low efficiency and strong subjectivity of traditional manual evaluation; furthermore, it realizes a two-way alignment closed loop of "evaluation standard - service effect", by feeding back online user behavior data to the construction of evaluation set and standard iteration, forming a complete closed loop of "model output → automatic evaluation → reinforcement learning optimization → service improvement → standard update", truly realizing evaluation-driven service growth.

[0109] Figure 6 A flowchart illustrating a model alignment method provided according to some other embodiments of this specification is shown.

[0110] Reference Figure 6 As shown, the target service scenario is a marketing scenario in the insurance field. The service objectives of this scenario are "adoption of opening statements" and "users speaking after adoption." Electronic device 200 generates opening statements for users based on strategy models, such as large language models, for use by insurance sales personnel to improve user engagement. For example, electronic device 200 automatically generates personalized, compliant, and persuasive insurance product recommendation statements based on user profiles and historical interaction information.

[0111] Electronic device 200 classifies user-related data of multiple users based on preset funnel classification criteria; and constructs multiple evaluation sets corresponding to the target service scenario based on the classification results. The preset funnel classification criteria include: opening remarks exposed but not adopted; opening remarks exposed and adopted but the user did not speak; and opening remarks exposed and adopted and the user spoke. For example, assuming the target service scenario is insurance customer service opening remarks, evaluation set 1: opening remarks exposed but not adopted; 2: opening remarks exposed and adopted but the user did not speak; 3: opening remarks exposed and adopted and the user spoke.

[0112] Electronic device 200 inputs each candidate service output into the reward model, evaluates each candidate service output across multiple evaluation criteria dimensions based on the reward model, and obtains the reward value corresponding to that candidate service output. The reward value guides the optimization of the strategy model, generating higher-quality candidate service outputs.

[0113] Based on the service scenario objectives and evaluation criteria, the evaluation results of the three evaluation sets should be: reward value 1 of evaluation set 1 < reward value 2 of evaluation set 2 > reward value 3 of evaluation set 3. If the evaluation results do not meet this expectation through performance comparison, then the model parameters of the policy model need to be optimized, for example, by optimizing the model parameters of the policy model through reinforcement learning.

[0114] This specification, in another aspect, provides a non-transitory storage medium storing at least one set of executable instructions for performing model alignment processing. When these executable instructions are executed by a processor, they instruct the processor to implement the steps of the model alignment method described in this specification. In some possible embodiments, various aspects of this specification can also be implemented as a program product comprising program code. When the program product is run on an electronic device 200, the program code causes the electronic device 200 to perform the steps of the model alignment method described in this specification. The program product for implementing the above method may employ a portable compact disc read-only memory (CD-ROM) containing program code and may run on the electronic device 200. However, the program product of this specification is not limited thereto. In this specification, a readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system. The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. The aforementioned computer-readable storage media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium may also be any readable medium other than a readable storage medium that can send, propagate, or transmit programs for use by or in connection with an instruction execution system, apparatus, or device. Program code contained on a readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof. Program code for performing the operations described herein can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on electronic device 200, partially on electronic device 200, as a standalone software package, partially on electronic device 200 and partially on a remote computing device, or entirely on a remote computing device.

[0115] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than those shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0116] In summary, after reading this detailed disclosure, those skilled in the art will understand that the foregoing detailed disclosure is presented by way of example only and is not restrictive. Although not explicitly stated herein, those skilled in the art will understand that this specification requires various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be made by this specification and are within the spirit and scope of the exemplary embodiments described herein.

[0117] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, "an embodiment," "an embodiment," and / or "some embodiments" mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment of this specification. Therefore, it is to be emphasized and understood that two or more references to "an embodiment" or "an embodiment" or "alternative embodiment" in various parts of this specification do not necessarily refer to the same embodiment. Moreover, specific features, structures, or characteristics may be suitably combined in one or more embodiments of this specification.

[0118] It should be understood that in the foregoing description of the embodiments in this specification, various features are combined in a single embodiment, drawing, or description for the purpose of simplifying the description and to aid in understanding a feature. However, this does not mean that the combination of these features is necessary, and those skilled in the art may readily identify some of the devices as separate embodiments when reading this specification. That is, the embodiments in this specification can also be understood as an integration of multiple secondary embodiments. It is also valid when each secondary embodiment contains fewer than all the features of a single foregoing disclosed embodiment.

[0119] Each patent, patent application, publication of a patent application, and other material such as articles, books, specifications, publications, documents, articles, etc., cited herein may be incorporated by reference, except for any identical content appearing in the related documents that may be inconsistent with or conflict with this document, or any identical document content that may have a limiting effect on the widest scope of the claims. For example, in the event of any inconsistency or conflict between the description, definition, and / or use of terms associated with any included material and those related to this document, the terminology herein shall prevail.

[0120] Finally, it should be understood that the embodiments disclosed herein are illustrative of the principles of the embodiments described in this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can implement the applications described in this specification using alternative configurations based on the embodiments in this specification. Therefore, the embodiments in this specification are not limited to the embodiments precisely described in the applications.

Claims

1. A model alignment method applied to a preset language model, the preset language model including a policy model and a reward model, the method comprising: Determine the evaluation sample set corresponding to the target service scenario, wherein the evaluation sample set includes user-related data of multiple users; Based on the evaluation sample set, a set of candidate service outputs for the user is generated through the strategy model; The reward model is used to evaluate each candidate service output using multiple evaluation criteria dimensions to determine the reward value corresponding to each candidate service output. as well as The model parameters of the strategy model are adjusted based on the reward value.

2. The method according to claim 1, wherein, The multiple evaluation criteria dimensions include bottom line, compliance, professionalism, and expressiveness. The reward model includes multiple sub-models, including a bottom line sub-model, a compliance sub-model, a professionalism sub-model, and an expressiveness sub-model.

3. The method according to claim 2, wherein, The step of evaluating each candidate service output using the reward model across multiple evaluation criteria dimensions to determine the reward value corresponding to each candidate service output includes: The sub-reward value corresponding to the output of each candidate service is determined based on each of the sub-models of the reward model; The reward value corresponding to each candidate service output is obtained by performing a weighted summation operation on each of the sub-reward values.

4. The method according to claim 1, wherein, The evaluation sample set corresponding to the target service scenario includes: Extract the user-related data of the multiple users from the service logs corresponding to the target service scenario; Based on the user-related data of the multiple users, multiple evaluation sets corresponding to the target service scenario are constructed according to a preset classification standard, wherein the preset classification standard is a classification standard determined based on the service conversion path of the target service scenario.

5. The method according to claim 4, wherein, The preset classification standard is a preset funnel classification standard. The construction of multiple evaluation sets corresponding to the target service scenario based on user-related data from multiple users according to the preset classification standard includes: The user-related data of the multiple users are classified based on the preset funnel classification criteria; and Based on the classification results, construct the multiple evaluation sets corresponding to the target service scenario.

6. The method according to claim 5, wherein, The candidate service output includes candidate opening lines, and the preset funnel classification criteria include opening lines being exposed but not adopted, opening lines being exposed and adopted but the user did not speak, and opening lines being exposed and adopted and the user spoke.

7. The method according to claim 1, wherein, The adjustment of the model parameters of the strategy model based on the reward value includes: The relative advantage value within the group corresponding to each of the candidate service outputs is determined based on the reward value corresponding to each of the candidate service outputs. Based on the reward value corresponding to each candidate service output and the relative advantage value within the group, the alignment model loss of the strategy model is determined; and The model parameters of the policy model are adjusted based on the alignment model loss.

8. The method according to claim 7, wherein, Before adjusting the model parameters of the strategy model based on the reward value, the method further includes: Based on the reward value corresponding to each of the candidate service outputs, the set of candidate service outputs are sampled through dynamic sampling.

9. The method according to claim 7, wherein, The alignment model loss includes a policy change term, which measures the change in the new policy function corresponding to the policy model compared to the old policy function corresponding to the policy model. Determining the alignment model loss based on the reward value corresponding to each candidate service output and the relative advantage value within the group includes: Determine the policy ratio between the new policy function and the old policy function; The strategy change term is determined based on the relative advantage value within the group corresponding to the candidate service output and the strategy ratio.

10. The method according to claim 9, wherein, The alignment model loss further includes a pruning term, which is used to limit the change in the policy ratio when the policy ratio is greater than a predetermined range. Determining the policy change term based on the intra-group relative advantage value corresponding to the candidate service output and the policy ratio includes: If the strategy ratio is greater than the predetermined range, then the upper limit of the predetermined range is taken as the target strategy ratio, and the strategy change item is determined based on the relative advantage value within the group and the target strategy ratio; If the strategy ratio is less than the predetermined range, then the lower limit of the predetermined range is taken as the target strategy ratio, and the strategy change term is determined based on the relative advantage value within the group and the target strategy ratio.

11. The method according to claim 7, wherein, The step of determining the intra-group relative advantage value corresponding to each candidate service output based on the reward value corresponding to each candidate service output includes: The mean reward and standard deviation of the reward corresponding to the output of each candidate service are determined based on the reward value corresponding to the output of each candidate service. Based on the reward value, the mean reward, and the standard deviation of the reward corresponding to each candidate service output, the relative advantage value within the group corresponding to each candidate service output is determined.

12. An electronic device, comprising: At least one storage medium storing at least one instruction set for model alignment processing; as well as At least one processor is communicatively connected to the at least one storage medium. When the electronic device is running, the at least one processor reads the at least one instruction set and executes the model alignment method according to any one of claims 1-11 according to the instructions of the at least one instruction set.