Human setting following method and system for role playing agent in decoding stage
By evaluating the importance of character attributes and constructing a reward function during the decoding stage, and dynamically adjusting the generation probability distribution of the large language model, the problem of character profile adaptation and consistency of role-playing agents in multiple scenarios is solved, achieving efficient character performance without additional overhead.
Patent Information
- Application Number
- CN202610023918.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-09
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2046-01-09
AI Technical Summary
Existing technologies are insufficient in adapting and maintaining consistency of character profiles for role-playing agents in dynamic scenarios. In particular, the generated content is prone to deviating from the character settings under conditions of multiple roles and multiple scenarios. Furthermore, existing methods require high computing resources and a large amount of manually labeled data.
During the decoding phase, a reward function is constructed by evaluating the importance scores of the persona attributes in the context to dynamically adjust the generation probability distribution of the large language model, aligning it with the target persona and achieving adaptive and consistent persona maintenance across multiple scenarios.
Without requiring additional computation or storage overhead, it achieves consistency in character performance and realism in generated content across multiple roles and scenarios. It is applicable to various large language model architectures and suitable for scenarios such as social simulation, virtual interaction, educational companionship, and entertainment games.
Smart Images

Figure CN121478960A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to a method and system for character-following intelligent agents in the decoding stage. Background Technology
[0002] With the development of large language models, role-playing agents have been widely used in dialogue systems, social science research, and virtual interaction. By setting specific character profiles, agents can simulate the personality traits, behaviors, and values of characters in dialogues, thereby improving the realism and consistency of interactions. However, current technologies still have significant shortcomings in terms of adapting to and maintaining the consistency of character profiles under dynamic scene conditions; that is, they cannot achieve character-driven following.
[0003] Existing methods are mainly divided into non-parametric methods based on prompts and training-based parametric methods. The former typically relies on prompting engineering techniques, such as adding character attribute descriptions to the input, to guide the model in generating content that conforms to the established settings. However, these methods usually only process prompts at a shallow semantic level, lacking deep modeling of character attributes and context-adaptive capabilities, leading to generated content that easily deviates from the character's established role. In contrast, parametric methods utilize large-scale corpora to enhance the model's role consistency through fine-tuning, but often require high computational resources and a large amount of manually labeled data, especially in complex tasks involving multiple roles and scenarios, where data collection and modeling costs are enormous. Furthermore, existing methods generally lack dynamic adaptability. Psychological research shows that humans activate different personality traits in different situations, but current intelligent agents typically treat character settings as static constraints, unable to adjust attribute weights according to situational changes, thus reducing the realism and context relevance of generated content. Therefore, a dynamically adaptable technology is urgently needed to improve the personality consistency and performance realism of role-playing agents under multi-scenario and multi-role conditions. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides a method and system for character profile following in the decoding phase of a role-playing intelligent agent. This invention enables dynamic attribute weight adjustment directly during the decoding phase without introducing additional computational or storage overhead. It achieves adaptive character profile adaptation and consistency maintenance across multiple scenarios. By estimating the importance of character attributes in real-time during inference and dynamically adjusting the generation distribution, it can adaptively adjust character performance according to different contexts and tasks, thereby generating output that better matches the target character's settings. Furthermore, this invention can be embedded into any language model framework, achieving character profile consistency control without modifying the model structure.
[0005] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for character-following intelligent agents during the decoding phase, comprising: For a large language model, which includes contextual context, a set of persona attributes, and input prompts for user queries, the importance score of each persona attribute in the set of persona attributes is evaluated in the contextual context. The importance score is used to quantify the degree of influence of the corresponding persona attribute on the response generated by the large language model. Based on the importance scores, a reward function is constructed to guide the response of the large language model to align with the target persona. The reward function integrates the weighted contributions of each persona attribute based on its corresponding importance score. During the decoding and inference phase of the large language model, the probability distribution of generated lexical units at each time step is dynamically adjusted using the reward function, so that the generated response conforms to the target persona in the context.
[0006] In one embodiment, evaluating the importance score of each personality attribute in the set of personality attributes in the context specifically includes: Calculate character attributes Importance score : ; in, Indicates input prompt words, Indicates from Remove character attributes The input prompt word after that, The large language model is based on input prompts. The generated response, Represents conditional probability. It means "defined as".
[0007] In one embodiment, constructing a reward function based on the importance score to guide the large language model's response to align with the target persona specifically includes: ; For the initial reward function, Indicates input prompt words, This represents the response generated by the large language model. Represents the personality attribute of the i-th person. Importance score The total number of character attributes. for Corresponding step-by-step character development rewards; The initial reward function is normalized to obtain the final reward function. : .
[0008] In one embodiment, At time step Corresponding step-by-step character design rewards for: ; in, For time step index, For large language models at time steps The generated lexical units, For large language models at time steps The generated response, This indicates that the large language model incorporates persona attributes into the input prompts. Given the distribution of predicted output for the next word, This indicates that the large language model does not include persona attributes in the input prompts. In the case of time step The predicted output distribution of the word units. Indicates from Remove character attributes The input prompt word after that.
[0009] In one embodiment, dynamically adjusting the probability distribution of word units generated by the large language model at each time step using the reward function, so that the generated response conforms to the target persona in the context, specifically includes: Dynamically adjust the lexical units generated at time step t of the large language model probability distribution : ; in, Indicates input prompt words, Indicates at time step The generated word units, Indicates at time step The generated response, Representing the large language model at time steps The predicted output distribution of the word units. Represents the regularization parameter. For the reward function, It is a normalization factor, ensuring It is an efficient probability distribution: ; For the reward function, This indicates that the large language model is at time step The generated lexical units, For large language models at time steps The generated response, This indicates that the large language model incorporates persona attributes into the input prompts. In the case of time step The predicted output distribution of the lexical units.
[0010] Secondly, this invention provides a character-following system for a role-playing intelligent agent during the decoding phase, comprising: Persona Importance Assessment Module: For the large language model, which includes contextual context, persona attribute set, and user query input prompts, assess the importance score of each persona attribute in the contextual context. The importance score is used to quantify the degree of influence of the corresponding persona attribute on the response generated by the large language model. Reward function construction module: Based on the importance score, a reward function is constructed to guide the response of the large language model to align with the target persona. The reward function integrates the weighted contribution of each persona attribute based on its corresponding importance score. Reasoning Module: In the decoding and reasoning stage of the large language model, the probability distribution of generated lexical units in the large language model at each time step is dynamically adjusted using the reward function, so that the generated response conforms to the target persona in the context.
[0011] In one embodiment, evaluating the importance score of each personality attribute in the set of personality attributes in the context specifically includes: Calculate character attributes Importance score : ; in, Indicates input prompt words, Indicates from Remove character attributes The input prompt word after that, The large language model is based on input prompts. The generated response, Represents conditional probability. It means "defined as".
[0012] In one embodiment, constructing a reward function based on the importance score to guide the large language model's response to align with the target persona specifically includes: ; For the initial reward function, Indicates input prompt words, This represents the response generated by the large language model. Represents the personality attribute of the i-th person. Importance score The total number of character attributes. for Corresponding step-by-step character development rewards; The initial reward function is normalized to obtain the final reward function. : .
[0013] In one embodiment, At time step Corresponding step-by-step character design rewards for: ; in, For time step index, For large language models at time steps The generated lexical units, For large language models at time steps The generated response, This indicates that the large language model incorporates persona attributes into the input prompts. Given the distribution of predicted output for the next word, This indicates that the large language model does not include persona attributes in the input prompts. In the case of time step The predicted output distribution of the word units. Indicates from Remove character attributes The input prompt word after that.
[0014] In one embodiment, dynamically adjusting the probability distribution of word units generated by the large language model at each time step using the reward function, so that the generated response conforms to the target persona in the context, specifically includes: Dynamically adjust the lexical units generated at time step t of the large language model probability distribution : ; in, Indicates input prompt words, Indicates at time step The generated word units, Indicates at time step The generated response, Representing the large language model at time steps The predicted output distribution of the word units. Represents the regularization parameter. For the reward function, It is a normalization factor, ensuring It is an efficient probability distribution: ; For the reward function, This indicates that the large language model is at time step The generated lexical units, For large language models at time steps The generated response, This indicates that the large language model incorporates persona attributes into the input prompts. In the case of time step The predicted output distribution of the lexical units.
[0015] The system and method in this invention correspond to each other; the specific technical solutions applicable to the method are also applicable to the system.
[0016] Compared with the prior art, the beneficial technical effects of the present invention are: (1) No additional training or parameter overhead: This method directly estimates the importance of character attributes and dynamically aligns them during the inference stage, without retraining the language model or introducing additional parameterization modules. (2) Dynamic character attribute control: By introducing character attribute importance estimation during the inference stage, the weights of different character attributes can be dynamically adjusted according to the current dialogue context, achieving precise control over the character's performance in different contexts, making the generated text more consistent with the target character's setting. (3) Good versatility and scalability: This method does not rely on a specific model structure or training strategy, can be widely adapted to mainstream large language model architectures, and can be extended to character alignment tasks with multiple roles and multiple scenarios. This method has broad application potential and practical value in scenarios such as social simulation, virtual interaction, educational companionship, and entertainment games. Attached Figure Description
[0017] Figure 1 This is a flowchart of the method in an embodiment of the present invention; Figure 2 This is an overall architecture diagram of an embodiment of the present invention. Detailed Implementation
[0018] A preferred embodiment of the present invention will now be described in detail with reference to the accompanying drawings.
[0019] like Figure 1 As shown, a character setting following method for a role-playing intelligent agent in the decoding stage of this invention includes the following steps: S1, for the large language model including context, persona attribute set and user query input prompt words, evaluate the importance score of each persona attribute in the context, the importance score is used to quantify the degree of influence of the corresponding persona attribute on the response generated by the large language model; S2, Based on the importance score, construct a reward function to guide the response of the large language model to align with the target persona. The reward function integrates the weighted contribution of each persona attribute based on its corresponding importance score. S3, in the decoding and reasoning stage of the large language model, the probability distribution of generated lexical units in the large language model at each time step is dynamically adjusted using the reward function, so that the generated response conforms to the target persona in the context.
[0020] Role-playing intelligent agent technology is generally defined as: an artificial intelligence system used to simulate a specified role, which uses a variety of advanced capabilities of a large language model to simulate human characteristics and perform vivid role-playing performances.
[0021] The process proposed in this invention is as follows: Figure 2 As shown, it includes the following three parts: (1) character importance assessment; (2) character-driven reward function construction; and (3) inference stage alignment. These three parts work together in the inference stage to adjust the output distribution of the large language model in the form of dynamic rewards, thereby achieving real-time alignment of character and context.
[0022] 1. Assessment of the importance of persona.
[0023] According to psychological theory, the influence of persona attributes on behavior is not static, but dynamically changes with the situation. Therefore, this invention introduces a persona importance assessment module, which can assess the importance of each persona attribute in different situations without the need for labeled data.
[0024] Specifically, given complete input prompts Large language model generates output The probability can be expressed as: ; This represents the joint probability of a large language model generating the entire response given complete input prompts. Indicates the context, Indicates inclusion Personal attributes A collection of personality attributes, including style, personality, and MBTI (My Own Professional Personality Test) results, etc. This indicates a user query. Based on information theory principles, a certain persona attribute... For output The impact can be measured using Conditional Mutual Information (CMI): ; in, This indicates the input prompt after removing the character's attributes. Indicates conditional mutual information. This represents conditional entropy. The mutual information represents the persona attributes. For output The greater the information gain, the more significant the impact of the attribute on the output of the large language model.
[0025] Because the output space of a large language model is extremely large, directly calculating the entropy value is computationally expensive. Therefore, this invention uses a single representative output sample to approximate the overall distribution. Assuming there exists a standard output that meets the requirements of a person's profile, the conditional mutual information can be rewritten as: ; ; Thus, the character attributes are obtained. Importance score : ; “ " means "defined as".
[0026] Due to the real standard response in role-playing tasks with multiple roles and multiple scenarios Often unavailable, this makes directly calculating the importance of character attributes challenging. This invention proposes a self-supervised approximation method that utilizes responses generated by a large language model itself under complete prompts. As a proxy, assume the distribution generated by the large language model and If there is a positive correlation, then Can be used as A reliable proxy. This correlation assumption is reasonable because the training objective of large language models is to maximize the probability of generating a realistic response given a cue. This indicates the response generated based on the complete prompt. The importance score obtained Indicates a response based on true standards. The importance score obtained.
[0027] Ultimately, in the absence of real-world labeled responses, this invention provides a theoretically sound and practically reliable method for quantifying the importance of personas, adaptable to the current context: .
[0028] 2. Construction of a persona-driven reward function.
[0029] This section is used to integrate the importance score estimated by the persona importance assessment module into the reward function, so as to adjust the output of the large language model during the inference phase to keep it consistent with the target persona.
[0030] First, define the multi-target alignment problem. Constrained reinforcement learning objectives are achieved by adjusting the behavior of large language models to conform to specific properties through a reward function: .
[0031] in, It is the basic model to be aligned. It is the aligned model distribution. Represents the mathematical expectation. It is a reward function used to quantify a given complete input prompt word. and response preference level, express divergence, This is the regularization parameter. The goal is to achieve output alignment through reward optimization without changing the parameters of a large language model.
[0032] The goal of this invention is to make the response of a large language model match the preset set of persona attributes. Alignment, which is equivalent to maximizing the difference between the unconstrained model policy and the constrained policy. Divergence. Therefore, for each persona attribute It can The term is represented as the prediction of the large language model in the th... Expected log ratio over time steps: .
[0033] Based on this formula, each attribute can be defined. The gradual character development reward is as follows: .
[0034] in, and These represent the large language model with and without persona attributes, respectively. The predicted output distribution for the next lexical term is determined by the time interval. This progressive reward effectively captures the local influence of each character's personality attributes during each lexical term generation process, and can guide the decoding in real time to adjust the generation probability.
[0035] In order to align the character design attributes set Multiple character attributes This paper proposes a multi-objective strategy alignment framework. Based on traditional multi-objective alignment methods, and considering the priority of different persona attributes in the context, a dynamic weighting mechanism is introduced. Specifically, for each persona attribute... Assign a person a score of importance And use these importance scores to construct a weighted reward function. : .
[0036] However, while this weighted approach can achieve multi-objective alignment, it has a key limitation: unconstrained optimization may lead to all... Simultaneously maximizing and blurring the priority hierarchy among objectives hinders the generation of personalized, preference-aware Pareto optimal solutions.
[0037] Therefore, this invention proposes a normalized reward function. To encourage the ranking of desired personas, i.e. And retain the expected priority of the alignment target: ; in, This represents the reward vector. Maximizing the reward... This will prompt individuals to reward themselves. Maintain its corresponding importance score Consistent sorting ensures that the hierarchical structure of character attributes is clearly preserved during the alignment process; This represents the F2 norm.
[0038] 3. Alignment during the inference phase.
[0039] This part can dynamically adjust the output probability distribution of the large language model without fine-tuning, aligning it with the target persona. Specifically, by... Integrating into the decoding process of a large language model, the generation probability at each time step is dynamically adjusted.
[0040] By constructing a function using the Lagrange multiplier method, a closed-form solution for the probability output can be obtained through differentiation. At each time step... The alignment strategy for a large language model can be calculated using the following formula: ; in, It is a normalization factor, ensuring It is an efficient probability distribution: ; In this way, the large language model dynamically adjusts the generation probability at each time step based on the current context and the importance of the persona, thereby generating a response consistent with the target persona.
[0041] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0042] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple steps or stages, which are not necessarily completed at the same time, but may be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but may be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0043] Based on the description of the above method embodiments, the present invention also provides a system. The system may be a system that uses software (applications), modules, components, servers, clients, etc., using the methods described in the embodiments of this specification, combined with necessary implementation hardware. Based on the same innovative concept, the systems in one or more embodiments provided in this disclosure are as described in the following embodiments. Since the implementation schemes and methods for solving the problem are similar, the specific system implementations in the embodiments of this specification can refer to the implementations of the foregoing methods, and repeated details will not be repeated. As used below, the terms "module" or "module group" refer to a combination of software and / or hardware capable of implementing a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, hardware implementations, or a combination of software and hardware, are also possible and contemplated.
[0044] A character-following system for role-playing intelligent agents during the decoding phase includes: Persona Importance Assessment Module: For the large language model, which includes contextual context, persona attribute set, and user query input prompts, assess the importance score of each persona attribute in the contextual context. The importance score is used to quantify the degree of influence of the corresponding persona attribute on the response generated by the large language model. Reward function construction module: Based on the importance score, a reward function is constructed to guide the response of the large language model to align with the target persona. The reward function integrates the weighted contribution of each persona attribute based on its corresponding importance score. Reasoning Module: In the decoding and reasoning stage of the large language model, the probability distribution of generated lexical units in the large language model at each time step is dynamically adjusted using the reward function, so that the generated response conforms to the target persona in the context.
[0045] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0046] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention, and no reference numerals in the claims should be construed as limiting the scope of the claims.
[0047] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for character setting following of a role-playing intelligent agent during the decoding phase, characterized in that, include: For a large language model, which includes contextual context, a set of persona attributes, and input prompts for user queries, the importance score of each persona attribute in the set of persona attributes is evaluated in the contextual context. The importance score is used to quantify the degree of influence of the corresponding persona attribute on the response generated by the large language model. Based on the importance scores, a reward function is constructed to guide the response of the large language model to align with the target persona. The reward function integrates the weighted contributions of each persona attribute based on its corresponding importance score. During the decoding and inference phase of the large language model, the probability distribution of generated lexical units at each time step is dynamically adjusted using the reward function, so that the generated response conforms to the target persona in the context.
2. The method for character setting following of a role-playing intelligent agent in the decoding stage according to claim 1, characterized in that, The evaluation of the importance score of each personality attribute in the set of personality attributes in the context specifically includes: Calculate character attributes Importance score : ; in, Indicates input prompt words, Indicates from Remove character attributes The input prompt word after that, The large language model is based on input prompts. The generated response, Represents conditional probability. It means "defined as".
3. The method for character setting following of a role-playing intelligent agent during the decoding stage according to claim 1, characterized in that, The reward function, constructed based on the importance score, to guide the large language model's response to align with the target persona, specifically includes: ; For the initial reward function, Indicates input prompt words, This represents the response generated by the large language model. Represents the personality attribute of the i-th person. Importance score The total number of character attributes. for Corresponding step-by-step character development rewards; The initial reward function is normalized to obtain the final reward function. : 。 4. The method for character setting following of a role-playing intelligent agent in the decoding stage according to claim 3, characterized in that, At time step Corresponding step-by-step character development rewards for: ; in, For time step index, For large language models at time steps The generated lexical units, For large language models at time steps The generated response, This indicates that the large language model incorporates persona attributes into the input prompts. Given the distribution of predicted output for the next word, This indicates that the large language model does not include persona attributes in the input prompts. In the case of time step The predicted output distribution of the word units. Indicates from Remove character attributes The input prompt word after that.
5. The method for character setting following of a role-playing intelligent agent during the decoding stage according to claim 1, characterized in that, The step of dynamically adjusting the probability distribution of word units generated by the large language model at each time step using the reward function, so that the generated response conforms to the target persona in the context, specifically includes: Dynamically adjust the lexical units generated at time step t of the large language model probability distribution : ; in, Indicates input prompt words, Indicates at time step The generated word units, Indicates at time step The generated response, Representing the large language model at time steps The predicted output distribution of the word units. Represents the regularization parameter. For the reward function, It is a normalization factor, ensuring It is an efficient probability distribution: ; For the reward function, This indicates that the large language model is at time step The generated lexical units, For large language models at time steps The generated response, This indicates that the large language model incorporates persona attributes into the input prompts. In the case of time step The predicted output distribution of the lexical units.
6. A character-following system for a role-playing intelligent agent during the decoding phase, characterized in that, include: Persona Importance Assessment Module: For the large language model, which includes contextual context, persona attribute set, and user query input prompts, assess the importance score of each persona attribute in the contextual context. The importance score is used to quantify the degree of influence of the corresponding persona attribute on the response generated by the large language model. Reward function construction module: Based on the importance score, a reward function is constructed to guide the response of the large language model to align with the target persona. The reward function integrates the weighted contribution of each persona attribute based on its corresponding importance score. Reasoning Module: In the decoding and reasoning stage of the large language model, the probability distribution of generated lexical units in the large language model at each time step is dynamically adjusted using the reward function, so that the generated response conforms to the target persona in the context.
7. A character-following system for a role-playing intelligent agent in the decoding stage according to claim 6, characterized in that, The evaluation of the importance score of each personality attribute in the set of personality attributes in the context specifically includes: Calculate character attributes Importance score : ; in, Indicates input prompt words, Indicates from Remove character attributes The input prompt word after that, The large language model is based on input prompts. The generated response, Represents conditional probability. It means "defined as".
8. A character-following system for a role-playing intelligent agent in the decoding stage according to claim 6, characterized in that, The reward function, constructed based on the importance score, to guide the large language model's response to align with the target persona, specifically includes: ; For the initial reward function, Indicates input prompt words, This represents the response generated by the large language model. Represents the personality attribute of the i-th person. Importance score The total number of character attributes. for Corresponding step-by-step character development rewards; The initial reward function is normalized to obtain the final reward function. : 。 9. A character-following system for a role-playing intelligent agent in the decoding stage according to claim 8, characterized in that, At time step Corresponding step-by-step character development rewards for: ; in, For time step index, For large language models at time steps The generated lexical units, For large language models at time steps The generated response, This indicates that the large language model incorporates persona attributes into the input prompts. Given the distribution of predicted output for the next word, This indicates that the large language model does not include persona attributes in the input prompts. In the case of time step The predicted output distribution of the word units. Indicates from Remove character attributes The input prompt word after that.
10. A character-following system for a role-playing intelligent agent in the decoding stage according to claim 6, characterized in that, The step of dynamically adjusting the probability distribution of word units generated by the large language model at each time step using the reward function, so that the generated response conforms to the target persona in the context, specifically includes: Dynamically adjust the lexical units generated at time step t of the large language model probability distribution : ; in, Indicates input prompt words, Indicates at time step The generated word units, Indicates at time step The generated response, Representing the large language model at time steps The predicted output distribution of the word units. Represents the regularization parameter. For the reward function, It is a normalization factor, ensuring It is an efficient probability distribution: ; For the reward function, This indicates that the large language model is at time step The generated lexical units, For large language models at time steps The generated response, This indicates that the large language model incorporates persona attributes into the input prompts. In the case of time step The predicted output distribution of the lexical units.
Citation Information
Patent Citations
Role playing chat robot implementation method based on human background large language model
CN119646140A
Intelligent agent reinforcement learning method and interaction method
CN119830994A
Large language model security vulnerability automatic detection method and device based on scene nesting
CN119885206A
Consultation shunting dialogue system based on small expert model
CN120296135A
Reward-model based reinforcement learning for performing reasoning tasks
US20240104391A1