Role playing ability optimization method and device of large language model, equipment and medium

By extracting role-playing data from novel text and performing supervised full parameters fine-tuning and iterative optimization, the target large language model is generated, which solves the problem of insufficient role-playing ability of large language models, improves the model's colloquial, stylized expression and emotional expression capabilities, and enhances the interactive experience of virtual companion applications.

CN120409468APending Publication Date: 2025-08-01SHENZHEN YUANZHI INFORMATION TECHNOLOGY DEVELOPMENT CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510342272.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In terms of role-playing ability, large language models have problems such as insufficient colloquial and stylized expression ability, poor emotional performance and plot promotion ability, resulting in the lack of vivid and natural interactive experience of virtual companion applications.

Method used

By extracting role-playing data from novel texts, supervised full parameters fine-tuning and iterative optimization are carried out to generate a target large language model, and improve the model's colloquial, stylized expression and emotional expression capabilities.

Benefits of technology

The role-playing ability of the large language model is improved, so that it can express emotions more accurately and promote the development of the plot, enhancing the interactive experience with users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120409468A_ABST
    Figure CN120409468A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, and particularly discloses a role playing ability optimization method and device of a large language model, equipment and a medium. Performing dialogue recognition and processing on the novel text to obtain role playing data; performing supervised all-parameter fine tuning on the basis of the role playing data to obtain an initial fine tuning model, and reasoning dialogue input data in the data to obtain model reasoning data; and performing iterative optimization based on the role playing data and the model reasoning data to obtain target model parameters, and generating a target large language model. According to the method, the role playing data is extracted from the novel text, so that the role playing data comprises spoken and stylized data in various scenes, emotion can be accurately expressed, story development is promoted, iterative optimization is performed according to the role playing data, and the user experience is improved. Therefore, the model can learn oral and stylized expression ability, emotional expression and plot promotion ability from role playing data, and the role playing ability of the large language model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language processing, and particularly to a method, device, equipment and medium for optimizing the role-playing ability of large language models. Background Art

[0002] With the rapid development of large language models, virtual companion applications developed based on large language models are receiving more and more attention and love. In such applications, users can chat with various characters according to their preferences, develop their virtual relationships, or rebuild relationships in reality. However, the role-playing ability of general large language models is limited. For example, they lack colloquial and stylized expressions, have poor sense of picture and emotion expression ability, and poor plot advancement ability, etc. As a result, existing virtual companion applications generally have a strong "assistant style" and cannot interact vividly and naturally with users. Therefore, how to improve the role-playing ability of large language models has become an urgent problem to be solved. Summary of the Invention

[0003] This application provides a method, device, equipment and medium for optimizing the role-playing ability of large language models to improve the role-playing ability of large language models.

[0004] In a first aspect, this application provides a method for optimizing the role-playing ability of large language models, and the method includes:

[0005] Obtain a novel text, perform dialogue recognition and data processing on the novel text to obtain role-playing data;

[0006] Perform supervised full-parameter fine-tuning based on the role-playing data to obtain an initial fine-tuned model, and perform inference on the dialogue input data in the role-playing data based on the initial fine-tuned model to obtain model inference data;

[0007] Iteratively optimize the model parameters based on the role-playing data and the model inference data to obtain target model parameters, and generate a target large language model based on the target model parameters.

[0008] In a second aspect, this application also provides a device for optimizing the role-playing ability of large language models, and the device includes:

[0009] A role-playing data acquisition module, configured to obtain a novel text, and perform dialogue recognition and data processing on the novel text to obtain role-playing data;

[0010] A model inference data acquisition module, configured to perform supervised full-parameter fine-tuning based on the role-playing data to obtain an initial fine-tuned model, and perform inference on the dialogue input data in the role-playing data based on the initial fine-tuned model to obtain model inference data;

[0011] A target large language model acquisition module, configured to iteratively optimize model parameters based on the role-playing data and the model inference data, obtain target model parameters, and generate a target large language model based on the target model parameters.

[0012] In a third aspect, the present application further provides a computer device, which includes a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and, when executing the computer program, implement the method for optimizing the role-playing ability of the large language model as described above.

[0013] In a fourth aspect, the present application further provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the method for optimizing the role-playing ability of the large language model as described above.

[0014] The present application discloses a method, device, equipment and medium for optimizing the role-playing ability of a large language model. A novel text is obtained, and dialogue recognition and data processing are performed on the novel text to obtain role-playing data; supervised full-parameter fine-tuning is performed based on the role-playing data to obtain an initial fine-tuned model, and the dialogue input data in the role-playing data is inferred based on the initial fine-tuned model to obtain model inference data; based on the role-playing data and the model inference data, model parameters are iteratively optimized to obtain target model parameters, and a target large language model is generated based on the target model parameters. The present application extracts role-playing data from the novel text, making it contain colloquial and stylized data in various scenarios, which can accurately express emotions and promote the development of the plot. Then, iterative optimization is performed according to the role-playing data, enabling the model to learn colloquial and stylized expression abilities, emotion expression and plot advancement abilities from the role-playing data, and improving the role-playing ability of the large language model. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings below are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 It is a schematic flowchart of a method for optimizing the role-playing ability of a large language model provided by the first embodiment of the present application;

[0017] Figure 2It is a schematic flowchart of a method for optimizing the role-playing ability of a large language model provided by the second embodiment of the present application;

[0018] Figure 3 It is a schematic flowchart of optimizing and iterating model parameters of a method for optimizing the role-playing ability of a large language model provided by an embodiment of the present application;

[0019] Figure 4 It is a schematic flowchart of a method for optimizing the role-playing ability of a large language model provided by the third embodiment of the present application;

[0020] Figure 5 It is a schematic block diagram of a device for optimizing the role-playing ability of a large language model provided by an embodiment of the present application;

[0021] Figure 6 It is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0022] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0023] The flowcharts shown in the accompanying drawings are only illustrative, not necessarily including all contents and operations / steps, nor necessarily executed in the described order. For example, some operations / steps can be decomposed, combined, or partially merged, so the actual execution order may change according to the actual situation.

[0024] It should be understood that the terms used in this specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0025] It should also be understood that the term " / and / " used in this specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.

[0026] Embodiments of the present application provide a method, apparatus, device, and medium for optimizing the role-playing ability of a large language model. Among them, the method for optimizing the role-playing ability of the large language model can be applied to a server. By extracting role-playing data from novel texts, it includes colloquial and stylized data in various scenarios, can accurately express emotions, and promote the development of the plot. Furthermore, iterative optimization is performed based on the role-playing data, enabling the model to learn colloquial and stylized expression abilities, emotion expression, and plot advancement abilities from the role-playing data, thereby improving the role-playing ability of the large language model. Among them, the server can be an independent server or a server cluster.

[0027] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0028] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for optimizing the role-playing ability of a large language model provided by an embodiment of the present application.

[0029] As Figure 1 shown, the method for optimizing the role-playing ability of the large language model specifically includes steps S101 to S103.

[0030] S101. Obtain a novel text, perform dialogue recognition and data processing on the novel text, and obtain role-playing data;

[0031] In this embodiment, the role-playing data includes N groups of dialogues (x i , y i ), 0 ≤ i ≤ N, and other annotation information, such as the description of the characters and the background story of each group of dialogues, etc. x i is the i-th dialogue input data in the role-playing data, and y i is the i-th dialogue output data in the role-playing data.

[0032] In one embodiment, the novel text can be obtained by web crawling or other means, or it can be a downloaded electronic text.

[0033] In one embodiment, for all the obtained novel texts, dialogue recognition can be performed. The novel texts can be preprocessed, such as removing irrelevant information (such as redundant punctuation marks and white spaces), segmenting, and clause splitting, etc., for subsequent dialogue extraction; then regular expressions or NLP (Natural Language Processing) tools can be used to identify the dialogue content.

[0034] In one embodiment, the identified dialogue content is subjected to data processing such as filtering and rewriting to obtain high-quality role-playing data.

[0035] S102, performing supervised full-parameter fine-tuning based on the role-playing data to obtain an initial fine-tuning model, and performing inference on dialogue input data in the role-playing data based on the initial fine-tuning model to obtain model inference data;

[0036] In one embodiment, supervised full parameter fine-tuning is a method of further training all parameters of a model through supervised learning based on a pre-trained model to adapt to a specific task.

[0037] Use synthetic high-quality role-playing data to perform supervised full-parameter fine-tuning to obtain the initial fine-tuning model, also known as the previous round model, with parameters θ t Specifically, we choose to use role-playing data to fine-tune the pre-trained model (such as the model based on the Transformer architecture (such as GPT, BERT, Qwen, etc.)) and update the model parameters θ t , and obtain the initial fine-tuned model. The initial fine-tuned model is already able to better understand the emotions and intentions in the character dialogue.

[0038] In one embodiment, the dialogue in the role-playing data is input into the data x i Input into the initial fine-tuning model, and use the initial fine-tuning model to analyze the dialogue input data x i Perform inference to generate output data y obtained by model inference i ', which is the model inference data.

[0039] Combined with the dialogue input data x i , dialogue output data y i And the model inference data y i ', generate triple data (x i ,y i ,y i '), 0≤i≤N.

[0040] S103: Iteratively optimize model parameters based on the role-playing data and the model inference data to obtain target model parameters, and generate a target large language model based on the target model parameters.

[0041] In one embodiment, the model parameters are iteratively optimized using an optimization scheme that combines offline optimization and online optimization.

[0042] Specifically, after a supervised full-parameter fine-tuning model is performed based on high-quality role-playing data, K times (a first preset number of iterations) of offline optimization is performed based on the model inference data.

[0043] When the offline optimization reaches K times, the initial language model deployed online is used to process the user input of some users covered by the model, obtain the output result corresponding to the user input, and perform L times of online optimization based on the user input and the output result.

[0044] When the online optimization reaches L times (the second preset number of iterations), the target model parameters are obtained. Based on the target model parameters, the target large language model finally deployed online is generated.

[0045] The above embodiments provide a method, device, equipment and medium for optimizing the role-playing ability of a large language model, obtaining a novel text, performing dialogue recognition and data processing on the novel text to obtain role-playing data; performing supervised full-parameter fine-tuning based on the role-playing data to obtain an initial fine-tuned model, and performing inference on the dialogue input data in the role-playing data based on the initial fine-tuned model to obtain model inference data; iteratively optimizing the model parameters based on the role-playing data and the model inference data to obtain the target model parameters, and generating a target large language model based on the target model parameters. This application extracts role-playing data from the novel text, making it contain colloquial and stylized data in various scenarios, which can accurately express emotions and promote the development of the plot. Then, iterative optimization is performed according to the role-playing data, enabling the model to learn colloquial and stylized expression abilities, emotion expression and plot promotion abilities from the role-playing data, and improving the role-playing ability of the large language model.

[0046] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a method for optimizing the role-playing ability of a large language model provided by an embodiment of the present application.

[0047] As Figure 2 shown, the method for optimizing the role-playing ability of the large language model specifically includes steps S201 to S204.

[0048] S201. Perform iterative optimization based on a preset parameter calculation formula, the role-playing data, and the model inference data;

[0049] Further, the preset parameter calculation formula is

[0050]

[0051] denotes finding θ that makes the minimum value, and taking the θ when takes the minimum value as the model parameters obtained in this round of iteration; where is an inverse proportional function, γ is a regulation coefficient, Θ is a parameter space; xi is the i-th dialogue input data in the role-playing data, y i is the i-th dialogue output data in the role-playing data, y i ′ is the model inference data corresponding to the i-th dialogue input data, and N is the total number of dialogues in the role-playing data.

[0052] In one embodiment, p θ (y i |x i ) is the probability that the model output in this round is y i when the input is x i ; is the probability that the model output obtained in the previous iteration is y i when the input is x i ; p θ (y i ′|x i ) is the probability that the model output in this round is y i ′ when the input is x i ; p θt (y i ′|x i ) is the probability that the model output obtained in the previous iteration is y i ′ when the input is x i .

[0053] In one embodiment, the role-playing data includes N groups of dialogues (x i , y i ), 0 ≤ i ≤ N, and other annotation information, such as the character description and background story of each group of dialogues, etc. x i is the i-th dialogue input data in the role-playing data, and y i is the i-th dialogue output data in the role-playing data.

[0054] Before the model is put into use online, the model parameters are optimized offline. Specifically, the optimized model parameters θ are obtained by minimizing . When the formula achieves the minimum value, θ is the model parameter of the current iteration round. When the first preset iteration number is reached, the model parameters obtained in the last iteration are used as the initial model parameters.

[0055] In one embodiment, according to the preset parameter calculation formula, the optimization iteration of the model parameter θ can increase the probability that the model obtained in the current round of iteration outputs y i compared with the model obtained in the previous round of iteration, and decrease the probability that the model in this round outputs y iThe probability of '. In the offline case, using synthetic high-quality role-playing data and the data obtained from the model's own reasoning, the model is "self-improved" K times (the first preset number of iterations), which can fully learn the colloquial, stylized expressions, sense of picture, emotional expressions, and plot advancement abilities of high-quality role-playing data, and significantly improve the model in terms of colloquial and stylized expressions, sense of picture, emotional expressions, and plot advancement.

[0056] S202. When reaching the first preset number of iterations, obtain the initial model parameters, and based on the initial model parameters, generate an initial large language model.

[0057] In one embodiment, when the offline optimization number of iterations reaches the first preset number of iterations, obtain the initial model parameters, and according to the initial model parameters, form an initial large language model.

[0058] S203. Process at least one user input based on the initial large language model to obtain an output result corresponding to the user input.

[0059] In one embodiment, after obtaining the initial large language model, deploy the model online and cover some users to perform online optimization on the model parameters. Specifically, the initial large language model performs reasoning based on the user input and outputs the result, and collect the preference results of these users for two sets of output results under the same online input. 0 ≤ i ≤ M. Where is the output result approved by the user when the user input is x i and is the output result not approved by the user when the user input is x i and

[0060] S204. Iteratively optimize the initial model parameters based on a preset parameter optimization formula and the output result to obtain target model parameters.

[0061] Furthermore, the iteratively optimizing the initial model parameters based on a preset parameter optimization formula and the output result to obtain target model parameters includes: obtaining the user's evaluation of the output result, and based on the evaluation, classifying the output result to obtain the output result approved by the user and the output result not approved by the user; iteratively optimizing the model parameters based on the preset parameter optimization formula, the output result approved by the user, and the output result not approved by the user; when reaching the second preset number of iterations, obtain the target model parameters.

[0062] In one embodiment, user evaluation is a key data source for optimizing the model. Users can rate or classify the output results of the model. For example, users can rate the satisfaction of the output results (e.g., from 1 to 5), and users can also mark the output results as "approved" or "disapproved".

[0063] According to the user's evaluation, the output results are divided into "output results approved by users" and "output results not approved by users". The classification criteria are as follows: Approved output results: users give higher ratings (e.g., 4 or 5), Unapproved output results: users give lower ratings (e.g., 1 or 2). Or classify directly according to the "approved" and "disapproved" marks labeled by users.

[0064] In one embodiment, the output results approved and not approved by users are used, combined with a preset parameter optimization formula, to iteratively optimize the model parameters. The optimization goal is to reduce the output not approved by users and increase the output approved by users.

[0065] Furthermore, the preset parameter optimization formula is

[0066]

[0067] denotes finding θ that makes the minimum. When takes the minimum value, θ is used as the model parameter obtained in this round of iteration; where M is the total number of user inputs, is the output result approved by users corresponding to the i-th user input x i , is the output result not approved by users corresponding to the i-th user input x i .

[0068] In one embodiment, is the probability that when the user input is x i , the output of this round of the model is ; is the probability that when the user input is x i , the output of the model obtained in the previous round of iteration is ;[[ID= forty-six]] is the probability that when the user input is x i , the output of this round of the model is ; is the probability that when the user input is x i , the output of the model obtained in the previous round of iteration is .

[0069] By finding The minimum value, when The θ when taking the minimum value is the model parameter of the current iteration round. When the second preset iteration number is reached, the parameter obtained in the last iteration is used as the target model parameter. The target large language model is formed according to the target model parameter.

[0070] In one embodiment, according to the preset parameter optimization formula, the optimization iteration of the model parameter θ is performed, which can increase the probability that the output of this round of the model is compared with the model obtained in the previous round of iteration, and decrease the probability that the output of this round of the model is compared with the previous round of model output. After L times of online self-optimization iteration, the model can learn the user's output preference for the online input distribution and better adapt to the online input distribution in the online scenario.

[0071] Specifically, as Figure 3 shown, after performing supervised full-parameter fine-tuning on the model according to high-quality role-playing data, the initial fine-tuned model is used to infer the dialogue input data in the role-playing data to obtain model inference data. Then, according to the role-playing data and the model inference data, K times (the first preset iteration number) of offline "self-improving" learning (K times of iterative optimization of model parameters) are performed. When the iteration number a reaches K times, the initial large language model is obtained. The initial large language model deployed online is used to infer the user input, and the output result corresponding to the user input is obtained, and the output result is analyzed to obtain user preference data, that is, the output results recognized by the user and the output results not recognized by the user. Online optimization is performed according to the user input and the user preference data corresponding to the user input. When the online iterative optimization reaches the second preset iteration number L times (that is, the iteration number a = K + L), the iteration ends, and the target model parameter is obtained. The final online model (that is, the target large language model) is generated according to the target model parameter.

[0072] In the above embodiment, the iterative optimization of the model parameter is performed by using an optimization scheme combining offline optimization and online optimization. Offline optimization fully learns the colloquial, stylized expressions, sense of picture, emotional expression, and plot advancement capabilities of high-quality role-playing data, making the model significantly improved in aspects such as colloquial and stylized expressions, sense of picture, emotional expression, and plot advancement. Online optimization enables the model to learn the user's output preference for each user input, better adapt to the online scenario, and better show a dialogue style, emotional expression, etc. that meet the user's expectations according to different user task settings, improving the role-playing ability of the model.

[0073] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a method for optimizing the role-playing ability of a large language model provided by an embodiment of the present application.

[0074] AsFigure 4 As shown in the figure, the method for optimizing the role-playing ability of the large language model specifically includes steps S301 to S303.

[0075] S301. Traverse all the novel texts, identify the chapters containing dialogue content, and obtain the dialogue content, the previous information of the dialogue content in the chapter, and the content of the previous chapter of the chapter.

[0076] In one embodiment, traverse all the novel texts and look for local fragments containing large sections of dialogue therein. Specifically, the dialogue content can be identified through regular expressions or NLP tools.

[0077] Suppose a novel contains a total of m chapters, and the i-th to j-th sentences in the n-th (0 < n <= m) chapter contain dialogue (0 <= i < j <= k), where k is the total number of sentences in the n-th chapter. Then the single data selected is (C n-1 , C n,0~i-1 , C n,i~j ), where C n-1 represents the content of the previous chapter of the chapter where the dialogue content is located, C n,0~i-1 represents the previous information of the front part of the dialogue content in this chapter, and C n,i~j represents the dialogue content selected from this chapter.

[0078] S302. Determine whether there is a two-person dialogue with the same scene in the dialogue content.

[0079] In one embodiment, since the large sections of dialogue selected may include conversations of multiple people, even conversations of unimportant characters, or there may be changes in the dialogue scene during the process, the data containing these factors are not required for optimizing the role-playing ability of the model. Therefore, pre-filtering is needed to reduce the data volume.

[0080] Specifically, use a general large language model with a larger size and stronger capabilities (such as XVERSE-MOE-A36B, a multilingual large language model based on a mixture of experts model architecture) to determine whether there is a conversation between two people (i.e., a two-person dialogue) in each data, and the occasion where the conversation occurs remains unchanged.

[0081] S303. When there is a two-person dialogue with the same scene in the dialogue content, rewrite and screen the two-person dialogue based on the previous information and the content of the previous chapter to obtain the role-playing data.

[0082] Further, rewriting and screening the two-person dialogue based on the foregoing information and the content of the previous chapter to obtain the role-playing data, including: identifying two dialogue roles in the two-person dialogue and the past events of each dialogue role based on the foregoing information; obtaining the character descriptions and background stories of the two dialogue roles based on the content of the previous chapter; structurally rewriting the two-person dialogue based on the vocabulary and expressions in the dialogue content to generate a dialogue history; and screening out the data of the character descriptions, the background stories, the dialogue history, and the past events that meet the preset conditions as the role-playing data.

[0083] In one embodiment, referring to the context information C n,0~i-1 Determine the dialogue content C n,i~j The two roles having a conversation in. Then use C n,i~j The original vocabulary and expressions in to structurally rewrite the conversation process between the two roles to generate a "dialogue history". Convert the dialogue in the novel text into a more standardized and clearer dialogue format, so that the rewritten structured dialogue data meets the model input requirements. Then, from the perspectives of accuracy, fluency, etc., filter out the rewritten results that do not meet the requirements. Specifically, NLP techniques and tools can be used to evaluate whether factors such as the accuracy and fluency of the rewritten results meet the requirements, and then judge whether the rewritten dialogue correctly reflects information such as the original intention and emotion of the text, retain the rewritten results that meet the requirements, and delete the rewritten results that do not meet the requirements.

[0084] In one embodiment, based on C n-1 Generate the "character descriptions" and "background stories" of these two roles, and generate the "past events" before the two roles have a conversation based on C n,0~i-1 Specifically, NLP tools can be used to analyze the content of C n-1 Extract key information, which may include the language style, emotional tendency, behavior pattern, personal experience, etc. of the role. Generate character descriptions and background stories based on the extracted key information. Then, from the perspectives of logic and emotion, judge whether the connection between the "character descriptions", the "background stories", and the "past events" and the "dialogue history" is reasonable, and filter out the data that does not meet the requirements.

[0085] Specifically, in terms of logic, NLP techniques and tools can be utilized to judge the logic on preset evaluation dimensions (such as time consistency, behavior consistency, plot coherence, etc.). The events in the dialogue history can be compared with the character timeline in the background story to determine whether time consistency is satisfied; the events in the dialogue history can be compared with the background events and past events to check whether plot coherence is satisfied; according to the character description, the behavior patterns and language styles of the characters can be analyzed to determine whether the dialogue language style in the dialogue history conforms to the character description. In terms of emotion, emotion analysis tools can be used to evaluate the emotional tendencies in the dialogue history, character description, background story, and past events to check whether they are consistent. The data with illogical and / or inconsistent emotions in "character description", "background story", and "past events" compared with the "dialogue history" are deleted, and the remaining data are used as role-playing data.

[0086] In the above embodiments, the novel text can be automatically crawled, and dialogue recognition and data processing can be performed on the novel text to generate high-quality role-playing data, improving the acquisition efficiency of role-playing data. Secondly, the novel text contains a large amount of colloquial and stylized data, which can accurately express emotions and drive the development of the plot. The role-playing data obtained by performing dialogue recognition and data processing on the novel text retains the colloquial and stylized data in various scenarios, can accurately express emotions, and drive the development of the plot. The target large language model obtained by iteratively optimizing using the role-playing data can learn colloquialism, stylization, emotion expression, and plot advancement capabilities from the role-playing data, improving the role-playing ability of the large language model.

[0087] Please refer to Figure 5 , Figure 5 FIG. is a schematic block diagram of an apparatus for optimizing the role-playing ability of a large language model provided by an embodiment of the present application. The apparatus for optimizing the role-playing ability of the large language model is used to execute the method for optimizing the role-playing ability of the large language model described above. Among them, the apparatus for optimizing the role-playing ability of the large language model can be configured in a server.

[0088] As Figure 5 shown, the apparatus 400 for optimizing the role-playing ability of the large language model includes:

[0089] A role-playing data acquisition module 401, configured to acquire a novel text, and perform dialogue recognition and data processing on the novel text to obtain role-playing data;

[0090] A model inference data acquisition module 402, configured to perform supervised full-parameter fine-tuning based on the role-playing data to obtain an initial fine-tuned model, and perform inference on the dialogue input data in the role-playing data based on the initial fine-tuned model to obtain model inference data;

[0091] The target large language model acquisition module 403 is used to iteratively optimize model parameters based on the role-playing data and the model inference data, obtain target model parameters, and generate a target large language model based on the target model parameters.

[0092] Further, the target large language model acquisition module 403 includes:

[0093] An iterative optimization unit for performing iterative optimization based on a preset parameter calculation formula, the role-playing data, and the model inference data;

[0094] An initial model generation unit for obtaining initial model parameters when reaching a first preset number of iterations, and generating an initial large language model based on the initial model parameters;

[0095] An output result acquisition unit for processing at least one user input based on the initial large language model to obtain an output result corresponding to the user input;

[0096] A target model generation unit for iteratively optimizing the initial model parameters based on a preset parameter optimization formula and the output result to obtain target model parameters.

[0097] Further, the preset parameter calculation formula is

[0098]

[0099] It means to find θ that makes the minimum value. When takes the minimum value, θ is used as the model parameter obtained in this round of iteration; where is an inverse proportional function, γ is a regulation coefficient, Θ is a parameter space; x i is the i-th dialogue input data in the role-playing data, y i is the i-th dialogue output data in the role-playing data, y i ′ is the model inference data corresponding to the i-th dialogue input data, and N is the total number of dialogues in the role-playing data.

[0100] Further, the target model generation unit includes:

[0101] An output result classification unit for obtaining the user's evaluation of the output result, and classifying the output result based on the evaluation to obtain the output result recognized by the user and the output result not recognized by the user;

[0102] A parameter iterative optimization unit for iteratively optimizing model parameters based on the preset parameter optimization formula, the output results recognized by the user, and the output results not recognized by the user;

[0103] A target parameter acquisition unit for acquiring the target model parameters when the second preset number of iterations is reached.

[0104] Further, the preset parameter optimization formula is

[0105]

[0106] It means to find θ that makes the minimum value. When takes the minimum value, θ is used as the model parameter obtained in this round of iteration; where M is the total number of user inputs, y choseni is the output result recognized by the user corresponding to the i-th user input x i is the output result not recognized by the user corresponding to the i-th user input x i

[0107] Further, the role-playing data acquisition module 401 includes:

[0108] A dialogue recognition unit for traversing all the novel texts, identifying the chapters containing dialogue content, and obtaining the dialogue content, the previous information of the dialogue content in the chapter, and the content of the previous chapter of the chapter;

[0109] A two-person dialogue judgment unit for judging whether the dialogue content contains two-person dialogues with the same scene;

[0110] A role-playing data acquisition unit for, when the dialogue content contains two-person dialogues with the same scene, rewriting and screening the two-person dialogues based on the previous information and the content of the previous chapter to obtain the role-playing data.

[0111] Further, the role-playing data acquisition unit includes:

[0112] A dialogue role recognition subunit for recognizing the two dialogue roles in the two-person dialogue and the past events of each dialogue role based on the previous information;

[0113] A character description acquisition subunit for obtaining the character descriptions and background stories of the two dialogue roles based on the content of the previous chapter;

[0114] ​​A dialogue history generation subunit, configured to structurally rewrite the two-person dialogue based on the vocabulary and expressions in the dialogue content to generate a dialogue history;

[0115] A data screening subunit, configured to screen out data that meets preset conditions in the character description, the background story, the dialogue history, and the past events as the role-playing data.

[0116] It should be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described device and each module can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0117] The above device can be implemented in the form of a computer program, and the computer program can run on a computer device as shown in Figure 6 shown.

[0118] Please refer to Figure 6 , Figure 6 which is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. The computer device can be a server.

[0119] Referring to Figure 6 , the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a non-volatile storage medium and an internal memory.

[0120] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can be enabled to execute any one of the role-playing ability optimization methods of the large language model.

[0121] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.

[0122] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can be enabled to execute any one of the role-playing ability optimization methods of the large language model.

[0123] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 6 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0124] It should be understood that the processor can be a Central Processing Unit (CPU), and the processor can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0125] Among them, in one embodiment, the processor is used to run a computer program stored in a memory to implement the following steps:

[0126] Obtain a novel text, perform dialogue recognition and data processing on the novel text to obtain role-playing data;

[0127] Perform supervised full-parameter fine-tuning based on the role-playing data to obtain an initial fine-tuned model, and perform inference on the dialogue input data in the role-playing data based on the initial fine-tuned model to obtain model inference data;

[0128] Iteratively optimize the model parameters based on the role-playing data and the model inference data to obtain target model parameters, and generate a target large language model based on the target model parameters.

[0129] In one embodiment, when the processor implements iteratively optimizing the model parameters based on the role-playing data and the model inference data to obtain target model parameters, it is used to implement:

[0130] Perform iterative optimization based on a preset parameter calculation formula, the role-playing data, and the model inference data;

[0131] When the first preset number of iterations is reached, obtain initial model parameters, and generate an initial large language model based on the initial model parameters;

[0132] Process at least one user input based on the initial large language model to obtain an output result corresponding to the user input;

[0133] Iteratively optimize the initial model parameters based on a preset parameter optimization formula and the output result to obtain target model parameters.

[0134] In one embodiment, the preset parameter calculation formula is

[0135]

[0136] denotes finding θ that makes the minimum value. When takes the minimum value, θ is used as the model parameter obtained in this round of iteration; where is an inverse proportional function, γ is a regulation coefficient, Θ is a parameter space; x i is the i-th dialogue input data in the role-playing data, y i is the i-th dialogue output data in the role-playing data, y i is the model inference data corresponding to the i-th dialogue input data, and N is the total number of dialogues in the role-playing data.

[0137] In one embodiment, when the processor iteratively optimizes the initial model parameters based on a preset parameter optimization formula and the output result to obtain target model parameters, it is used to implement:

[0138] Obtain the user's evaluation of the output result, and classify the output result based on the evaluation to obtain the output result recognized by the user and the output result not recognized by the user;

[0139] Iteratively optimize the model parameters based on the preset parameter optimization formula, the output result recognized by the user, and the output result not recognized by the user;

[0140] When reaching the second preset number of iterations, obtain the target model parameters.

[0141] In one embodiment, the preset parameter optimization formula is

[0142]

[0143] denotes finding θ that makes the minimum value. When takes the minimum value, θ is used as the model parameter obtained in this round of iteration; where M is the total number of user inputs, is the output result recognized by the user corresponding to the i-th user input x i and is the output result not recognized by the user corresponding to the i-th user input x i .

[0144] In one embodiment, when the processor implements dialogue recognition and data processing on the novel text to obtain role-playing data, it is used to implement:

[0145] Traverse all the novel texts, identify the chapters containing dialogue content, and obtain the dialogue content, the previous information of the dialogue content in the chapter, and the content of the previous chapter of the chapter;

[0146] Determine whether there is a two-person dialogue with the same scene in the dialogue content;

[0147] When there is a two-person dialogue with the same scene in the dialogue content, rewrite and screen the two-person dialogue based on the previous information and the content of the previous chapter to obtain the role-playing data.

[0148] In one embodiment, when the processor realizes rewriting and screening the two-person dialogue based on the previous information and the content of the previous chapter to obtain the role-playing data, it is used to realize:

[0149] Based on the previous information, identify the two dialogue roles in the two-person dialogue and the past events of each dialogue role;

[0150] Based on the content of the previous chapter, obtain the character descriptions and background stories of the two dialogue roles;

[0151] Based on the vocabulary and expressions in the dialogue content, structurally rewrite the two-person dialogue to generate a dialogue history;

[0152] Screen out the data that meets the preset conditions among the character descriptions, the background stories, the dialogue history, and the past events as the role-playing data.

[0153] In an embodiment of the present application, there is also provided a computer-readable storage medium, which stores a computer program, and the computer program includes program instructions. The processor executes the program instructions to implement any one of the role-playing ability optimization methods of the large language model provided by the embodiments of the present application.

[0154] Among them, the computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiment, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.

[0155] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A method for optimizing the role-playing ability of a large language model, characterized in that, Including: Obtain a novel text, perform dialogue recognition and data processing on the novel text to obtain role-playing data; Perform supervised full-parameter fine-tuning based on the role-playing data to obtain an initial fine-tuned model, and perform inference on the dialogue input data in the role-playing data based on the initial fine-tuned model to obtain model inference data; Iteratively optimize the model parameters based on the role-playing data and the model inference data to obtain target model parameters, and generate a target large language model based on the target model parameters.

2. The method for optimizing the role-playing ability of the large language model according to claim 1, wherein The iteratively optimizing the model parameters based on the role-playing data and the model inference data to obtain target model parameters includes: Perform iterative optimization based on a preset parameter calculation formula, the role-playing data, and the model inference data; When reaching a first preset number of iterations, obtain initial model parameters, and generate an initial large language model based on the initial model parameters; Process at least one user input based on the initial large language model to obtain an output result corresponding to the user input; Iteratively optimize the initial model parameters based on a preset parameter optimization formula and the output result to obtain target model parameters.

3. The method for optimizing the role-playing ability of the large language model according to claim 2, wherein The preset parameter calculation formula is Denote to obtain θ that makes the minimum value. When takes the minimum value, use θ as the model parameter obtained in this round of iteration. Among them, is an inverse proportional function, γ is a regulation coefficient, and Θ is a parameter space; x i is the i-th dialogue input data in the role-playing data, and y i is the i-th dialogue output data in the role-playing data, and y i ' is the model inference data corresponding to the i-th dialogue input data, and N is the total number of dialogues in the role-playing data.

4. The method for optimizing the role-playing ability of a large language model according to claim 2, wherein, The iteratively optimizing the initial model parameters based on a preset parameter optimization formula and the output result to obtain target model parameters includes: Obtain a user's evaluation of the output result, and classify the output result based on the evaluation to obtain an output result recognized by the user and an output result not recognized by the user; Iteratively optimize the model parameters based on the preset parameter optimization formula, the output result recognized by the user, and the output result not recognized by the user; When reaching a second preset number of iterations, obtain the target model parameters.

5. The method for optimizing the role-playing ability of the large language model according to claim 4, characterized in that, The preset parameter optimization formula is Denote to obtain θ that makes the minimum value. When takes the minimum value, θ is used as the model parameter obtained in this round of iteration. Among them, M is the total number of user inputs, y choseni is the output result recognized by the user corresponding to the i-th user input x i , and y rejecti is the output result not recognized by the user corresponding to the i-th user input x i .

6. The method for optimizing the role-playing ability of the large language model according to claim 1, characterized in that The performing dialogue recognition and data processing on the novel text to obtain role-playing data includes: Traverse all of the novel text, identify the chapters containing dialogue content, and obtain the dialogue content, the previous information of the dialogue content in the chapter, and the content of the previous chapter of the chapter; Determine whether the dialogue content contains a two-person dialogue with the same scene; When the dialogue content contains a two-person dialogue with the same scene, rewrite and screen the two-person dialogue based on the previous information and the content of the previous chapter to obtain the role-playing data.

7. The method for optimizing the role-playing ability of the large language model according to claim 6, wherein The rewriting and screening the two-person dialogue based on the previous information and the content of the previous chapter to obtain the role-playing data includes: Based on the previous information, identify the two dialogue roles in the two-person dialogue and the past events of each dialogue role; Based on the content of the previous chapter, obtain the character descriptions and background stories of the two dialogue roles; Structurally rewrite the two-person dialogue based on the vocabulary and expression in the dialogue content to generate a dialogue history; Screen out the data that meets the preset conditions among the character descriptions, the background stories, the dialogue history, and the past events as the role-playing data.

8. An optimization device for the role-playing ability of a large language model, characterized in that, Including: A role-playing data acquisition module, configured to acquire a novel text, perform dialogue recognition and data processing on the novel text, and obtain role-playing data; A model inference data acquisition module, configured to perform supervised full-parameter fine-tuning based on the role-playing data to obtain an initial fine-tuned model, and perform inference on the dialogue input data in the role-playing data based on the initial fine-tuned model to obtain model inference data; A target large language model acquisition module, configured to iteratively optimize model parameters based on the role-playing data and the model inference data to obtain target model parameters, and generate a target large language model based on the target model parameters.

9. A computer device, characterized in that, The computer device includes a memory and a processor; The memory is configured to store a computer program; The processor is configured to execute the computer program and, when executing the computer program, implement the method for optimizing the role-playing ability of the large language model according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the processor is caused to implement the method for optimizing the role-playing ability of the large language model according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Large language model alignment method and system for information extraction task

    CN118427292A

  • Role playing model training method and device

    CN119415677A

  • Method and device for training language model

    CN119415957A

  • Large model preference alignment method improved by using online synchronization strategy

    CN119539082A

  • HVAC system design and operational tool for building infection control

    US20200348038A1