Role dialogue generation method based on contrast learning constraint and related equipment
Patent Information
- Application Number
- CN202611248021.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-18
- Publication Date
- 2026-09-18
AI Technical Summary
[0004]本发明的主要目的在于提供基于对比学习约束的角色对话生成方法和相关设备,旨在解决现有技术中角色交互的生成内容难以保持角色一致性和稳定性的问题
[0009]Beneficial Effects: This invention discloses a method and related equipment for generating role-based dialogues based on contrastive learning constraints. Compared to existing technologies, this invention obtains current user dialogue information and identifies the target interaction role; acquires a role parameter set and constructs a dialogue generation task; encodes the role parameter set into a role reference embedding vector, and determines an adaptive temperature parameter based on the number of conversation rounds and the role complexity coefficient; during decoding, contrastive learning constraints are injected based on the role reference embedding vector and the adaptive temperature parameter to generate candidate responses; multi-dimensional scores for each candidate response are calculated and fused into a comprehensive score; and the target response is determined and output based on the score. By injecting contrastive learning constraints during the decoding stage, candidate responses are made to move closer to the role reference embedding vector, and the constraint strength is dynamically adjusted in conjunction with the adaptive temperature parameter. Simultaneously, the target response output is filtered through multi-dimensional scores, ensuring the role consistency and stability of the generated dialogue content.
Smart Images

Figure CN122777680A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and related equipment for generating role dialogues based on contrastive learning constraints. Background Technology
[0002] With the rapid development of generative large model technology, large-scale language models based on the Transformer architecture have the ability to generate open-domain dialogues. They can generate coherent and natural text responses based on user input and are widely used in scenarios such as intelligent customer service, virtual companionship, education and tutoring, and entertainment interaction.
[0003] However, when users expect the model to engage in dialogue as a specific preset IP character (such as a literary figure, brand ambassador, or game character), existing methods usually rely solely on simple system prompts or character description text for constraint. In long conversations or multi-turn interactions, character tone drift is prone to occur, meaning that the generated content gradually deviates from the initial character setting. It is difficult to maintain the consistency and stability of the generated content in complex interactions such as long conversations or multi-mode switching. Summary of the Invention
[0004] The main objective of this invention is to provide a method and related equipment for generating character dialogues based on contrastive learning constraints, aiming to solve the problem that it is difficult to maintain the consistency and stability of character interaction content in the prior art.
[0005] The technical solution of the present invention is as follows: The first aspect of this invention provides a method for generating role dialogue based on contrastive learning constraints, comprising: Obtain current user conversation information and confirm the target interaction role; Obtain the corresponding role parameter set according to the target interaction role, and construct a dialogue generation task based on the role parameter set and the user dialogue information; The character parameter set is encoded into a corresponding character reference embedding vector, and the adaptive temperature parameter is determined based on the current session round and the character complexity coefficient. The dialogue generation task is decoded, and during the decoding process, contrastive learning constraints are injected based on the role reference embedding vector and the adaptive temperature parameter to generate at least one candidate response. Calculate the scores of each candidate response in a preset dimension, and fuse them to obtain the corresponding comprehensive score; Based on the scores of each preset dimension and the comprehensive score, the target response is directly determined and output from each of the candidate responses, or the target response is adjusted for role adaptation and then output.
[0006] A second aspect of the present invention provides a role dialogue generation apparatus based on contrastive learning constraints, comprising: The data acquisition module is used to acquire current user dialogue information and confirm the target interaction role; The task construction module is used to obtain the corresponding role parameter set according to the target interaction role, and construct a dialogue generation task according to the role parameter set and the user dialogue information; The parameter confirmation module is used to encode the character parameter set into a corresponding character reference embedding vector, and determine the adaptive temperature parameter based on the current session round and the character complexity coefficient. The constraint decoding module is used to decode the dialogue generation task and inject contrastive learning constraints according to the role reference embedding vector and the adaptive temperature parameter during the decoding process to generate at least one candidate response. The response scoring module is used to calculate the scores of each candidate response in a preset dimension and fuse them to obtain the corresponding comprehensive score. The filtering output module is used to directly determine and output the target response from each candidate response based on the scores of each preset dimension and the comprehensive score, or to output the target response after role adaptation adjustment.
[0007] A third aspect of the present invention provides a computer device including at least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to perform the aforementioned role dialogue generation method based on contrastive learning constraints.
[0008] A fourth aspect of the present invention provides a computer-readable storage medium storing computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform the above-described role dialogue generation method based on contrastive learning constraints.
[0009] Beneficial Effects: This invention discloses a method and related equipment for generating role-based dialogues based on contrastive learning constraints. Compared to existing technologies, this invention obtains current user dialogue information and identifies the target interaction role; acquires a role parameter set and constructs a dialogue generation task; encodes the role parameter set into a role reference embedding vector, and determines an adaptive temperature parameter based on the number of conversation rounds and the role complexity coefficient; during decoding, contrastive learning constraints are injected based on the role reference embedding vector and the adaptive temperature parameter to generate candidate responses; multi-dimensional scores for each candidate response are calculated and fused into a comprehensive score; and the target response is determined and output based on the score. By injecting contrastive learning constraints during the decoding stage, candidate responses are made to move closer to the role reference embedding vector, and the constraint strength is dynamically adjusted in conjunction with the adaptive temperature parameter. Simultaneously, the target response output is filtered through multi-dimensional scores, ensuring the role consistency and stability of the generated dialogue content. Attached Figure Description
[0010] To more clearly illustrate the solutions in this invention, the accompanying drawings used in the description of the embodiments of this invention will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0011] Figure 1 A schematic diagram of an application environment for the role dialogue generation method based on contrastive learning constraints provided in an embodiment of the present invention; Figure 2 A flowchart of a role dialogue generation method based on contrastive learning constraints provided in an embodiment of the present invention; Figure 3 A schematic diagram of the functional modules of the role dialogue generation device based on contrastive learning constraints provided in an embodiment of the present invention; Figure 4 A schematic diagram of the hardware structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0012] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention is further described in detail below. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. The embodiments of the invention are described below in conjunction with the accompanying drawings.
[0013] The role dialogue generation method based on contrastive learning constraints provided in this invention can be applied to, for example... Figure 1The illustrated intelligent agent interaction scenario with integrated smart hardware includes a terminal device 101, smart hardware 102, a network 103, and a server 104. The network 103 serves as a medium for providing a communication link between the terminal device 101, smart hardware 102, and server 104. The network 103 can include various connection types, such as wired and / or wireless communication links (e.g., Bluetooth, Wi-Fi, NFC, etc.).
[0014] Users can use terminal device 101 and smart hardware 102 to interact with server 104 via network 103 to receive or send messages, etc. Terminal device 101 may have a client installed that supports virtual scenarios. For example, when the virtual scenario is pet companionship, the client could be a virtual pet application. Users can log in to the client to view, edit, or control the interactions of the bound smart agent in the virtual scenario (this is just an example). Terminal device 101 can be various electronic devices with a display screen and web browsing support, including but not limited to smartphones, tablets, and desktop computers.
[0015] The smart hardware 102 can be used with different toy carriers such as plush toys, desktop ornaments, and toy pendants. Character tags are set inside the toy carriers. The smart hardware 102 can read the tags through near-field communication and perform seamless switching operations such as character recognition, resource switching, and session migration based on the reading results, realizing a convenient interactive process of switching between multiple characters.
[0016] Server 104 can be a server providing various services, such as a backend server supporting the content browsed by users using terminal device 101 and smart hardware 102 (for example only). The backend server can analyze and process received user requests and other data, and feed the processing results back to the user through terminal device 101 and smart hardware 102. Server 104 can be a cloud server, a distributed system server, or a server integrated with blockchain.
[0017] It should be understood that the number of terminal devices 101, smart hardware 102, networks 103, and servers 104 mentioned above is merely illustrative. Depending on the implementation needs, there can be any number of terminal devices 101, smart hardware 102, networks 103, and servers 104. For example, a single user can correspond to one terminal device 101 and one smart hardware 102, and multiple users can achieve interactive linkage between target intelligent agents through the network.
[0018] like Figure 2 As shown, the role dialogue generation method based on contrastive learning constraints provided in this embodiment of the invention specifically includes the following steps: S201. Obtain the current user dialogue information and confirm the target interaction role.
[0019] In this embodiment, user dialogue information refers to multi-source input data when a user interacts with a target interactive role, including user input data, user profile data, conversation history data, and current mode identifiers, etc. By acquiring user dialogue information, the context and user state of the current interaction are clarified, providing foundational data for constructing subsequent dialogue generation tasks. Specifically, user input data refers to the text, voice, or structured commands currently sent by the user to the target interactive role; user profile data refers to pre-stored or dynamically updated user attribute information, including user preferences, interaction habits, historical role selection records, and sentiment tags; conversation history data refers to records of multiple rounds of dialogue that have occurred within the current conversation period, including historical user input and system output; and the current mode identifier is used to identify the current interaction mode type, such as casual conversation mode, task mode, or story mode.
[0020] During role-playing dialogues, user dialogue information is received through a pre-defined communication interface. Upon receipt, the integrity of each field is checked and the format is standardized. If data loss, format errors, or invalid request signatures are found, a request failure message is returned to the user. If the data is complete and valid, metadata such as the request initiation time and user identifier is recorded, the target interaction role is confirmed, and the process proceeds to the subsequent task construction stage.
[0021] For example, a user initiates an interaction request with the virtual character "Ancient Style Poet". The user input data is "The weather is very nice today, can I write a poem?". The user profile data shows that the user prefers the classical literature style, the conversation history data includes the previous three rounds of discussion records about the rules of poetry, and the current mode is identified as casual chat mode. After receiving and verifying that the above data is complete and valid, the target interaction character is confirmed to be "Ancient Style Poet", and the standardized user dialogue information is passed to the subsequent module.
[0022] This embodiment ensures that subsequent dialogue generation tasks can fully integrate the user's current intent, historical interaction context, and interaction mode characteristics by uniformly acquiring and standardizing user input data, user profile data, conversation history data, and current mode identifiers. This provides a complete data foundation for role-based dialogue generation and improves the contextual coherence of dialogue generation and the accuracy of understanding user intent.
[0023] S202. Obtain the corresponding role parameter set according to the target interaction role, and construct a dialogue generation task according to the role parameter set and the user dialogue information.
[0024] In this embodiment, the role parameter set is pre-configured structured constraint data for each target interaction role, used to define the role's identity attributes, language style, acceptable response range, and prohibited response range, etc. Based on the currently confirmed target interaction role, the corresponding role parameter set is retrieved from the database and fused with the fields in the user dialogue information to construct the corresponding dialogue generation task. This dialogue generation task is preferably represented by a structured task object, including input text fields, user profile fields, current mode fields, role constraint fields, history summary fields, and security control fields, etc. For example, the task object is assembled in the order of "role constraint fields—current mode fields—input text fields—user profile fields—history summary fields" to ensure that role constraints and the current mode prioritize constraints on the subsequent generation process.
[0025] Specifically, the acquired character parameter set includes character profile parameters, tone parameters, answerable boundary parameters, and prohibited boundary parameters. Character profile parameters are used to define the character's identity background, social relationships, and behavioral preferences; tone parameters are used to define the character's language style characteristics, including intonation tendencies, common sentence patterns, and range of emotional expression; answerable boundary parameters are used to limit the range of topics the character can respond to and the types of actions the character can perform; prohibited boundary parameters are used to specify the topics the character is prohibited from being involved in, the expressions the character is prohibited from using, and the actions the character is prohibited from performing.
[0026] When constructing a dialogue generation task, each sub-parameter in the role parameter set is mapped to a structured field and embedded into the role constraint field of the task object. At the same time, user input data is filled into the input text field, user profile data is filled into the user profile field, current mode identifier is filled into the current mode field, and conversation history data is extracted and filled into the history summary field. Finally, these are concatenated in order to form a complete dialogue generation task.
[0027] This embodiment constructs a dialogue generation task containing multi-dimensional constraint information by structurally integrating the role parameter set with user dialogue information. This enables the subsequent generation process to constrain role settings, user characteristics, and interaction modes within a unified task framework, effectively avoiding the problem of generated content deviating from role settings and improving the constraint integrity of role-based dialogue generation.
[0028] S203. Encode the character parameter set into a corresponding character reference embedding vector, and determine the adaptive temperature parameter based on the current session round and character complexity coefficient.
[0029] In this embodiment, the character reference embedding vector is a dense vector representation that maps the structured character parameter set to a high-dimensional semantic space, used to provide a computable character semantic reference in the subsequent generation and scoring stages. A pre-trained character encoder encodes the character parameter set. This character encoder is preferably built based on a Transformer architecture, including an embedding layer, a multi-layer self-attention encoding layer, and a pooling output layer. The embedding layer maps each structured field in the character parameter set to an initial embedding vector. The multi-layer self-attention encoding layer captures the semantic relationships between fields through a multi-head attention mechanism and performs a non-linear transformation via a feedforward network. The pooling output layer performs mean pooling on the last hidden state, thereby outputting a fixed-dimensional character reference embedding vector.
[0030] Furthermore, an adaptive temperature parameter is calculated based on the current conversation round number and the role complexity coefficient to dynamically adjust the strength of the contrastive learning constraints. Specifically, as the number of conversation rounds increases and the role becomes more complex, the adaptive temperature parameter decreases accordingly, making the contrastive learning constraints increasingly strict in long conversations, thereby suppressing role tone drift.
[0031] For example, for the "ancient style poet" character, the character parameter set is encoded and the character reference embedding vector is output. The current conversation is in the 5th round, and the preset maximum number of conversation rounds is 20 rounds. An adaptive temperature parameter is calculated based on the current conversation round and the character complexity coefficient. This parameter is lower than the initial value, thereby strengthening the constraint on the character's tone in long conversations and preventing the character's personality from drifting as the number of rounds increases.
[0032] This embodiment encodes the structured role parameter set into role reference embedding vectors in the semantic space, providing a quantifiable role semantic benchmark for subsequent generation and scoring stages. At the same time, it obtains dynamic adaptive temperature parameters based on session depth and role complexity to achieve adaptive control of constraint strength, which ensures the generation flexibility in the early stage of short sessions and suppresses role drift in the later stage of long sessions, thereby ensuring the continuous stability of role consistency.
[0033] S204. Decode the dialogue generation task and inject contrastive learning constraints according to the role reference embedding vector and the adaptive temperature parameter during the decoding process to generate at least one candidate response.
[0034] In this embodiment, the dialogue generation task is input into a pre-trained generative model for decoding. This generative model is constructed using an autoregressive Transformer decoder architecture, including an input embedding layer, a position encoding layer, a multi-layer masked self-attention decoding layer, and an output projection layer. The input embedding layer maps each structured field in the dialogue generation task to an initial input vector. The position encoding layer adds positional information to the input vector. The multi-layer masked self-attention decoding layer generates candidate responses word by word through a causal masking mechanism. Each decoding layer contains a masked self-attention sub-layer, a cross-attention sub-layer, and a feedforward network sub-layer. The output projection layer maps the last hidden state to a probability distribution on the vocabulary.
[0035] During the decoding process, contrastive learning constraints are injected based on the character reference embedding vector and an adaptive temperature parameter. Specifically, at each step of decoding, the semantic correlation between the current candidate word embedding and the character reference embedding vector is calculated, and this correlation is incorporated as a bias term into the decoding score. This guides the generative model to prioritize candidate words that are close to the character's semantic anchor points during word selection, making the generated candidate responses semantically closer to the character reference embedding vector. Simultaneously, the strictness of the contrastive learning constraints is adjusted through the adaptive temperature parameter. The lower the temperature parameter, the stronger the contrastive learning constraint, and the closer the generated candidate response is to the character's setting; conversely, the higher the temperature parameter, the weaker the contrastive learning constraint, and the generated candidate response is allowed a moderate semantic deviation. Based on the above dynamic constraint mechanism, at least one candidate response is generated for subsequent scoring and selection.
[0036] This embodiment injects contrastive learning constraints based on role reference embedding vectors in real time during the decoding process, so that the generative model is guided by role semantic references at every step of autoregressive decoding, thereby ensuring the role consistency of candidate responses from the source. At the same time, the constraint strength is dynamically adjusted by adaptive temperature parameters, taking into account both role stability and semantic diversity, ensuring the generation quality and controllability of candidate responses.
[0037] S205. Calculate the scores of each candidate response in a preset dimension and merge them to obtain the corresponding comprehensive score.
[0038] In this embodiment, the preset dimension scoring is a quantitative evaluation of candidate responses from multiple assessment dimensions to evaluate the degree of matching between each candidate response and role settings, security specifications, and user intent. For example, the preset dimension scoring may include role consistency scoring, security scoring, and relevance scoring. After calculating the scores for each dimension separately through the corresponding scoring model or rule engine, the scores of each dimension are merged into a comprehensive score according to preset weights to characterize the overall quality level of the candidate responses.
[0039] By quantifying and scoring candidate responses from multiple preset dimensions such as role consistency, security, and relevance, and integrating them into a comprehensive score, the selection criteria for candidate responses include a comprehensive evaluation of both single and multi-dimensional aspects. This avoids the one-sidedness of relying solely on role consistency while ignoring security regulations or user intent, and ensures the comprehensiveness and reliability of candidate response selection.
[0040] S206. Based on the scores of each preset dimension and the comprehensive score, directly determine the target response from each candidate response and output it, or adjust the target response for role adaptation and output it.
[0041] In this embodiment, the target response is determined based on the preset dimension scores and comprehensive scores of each candidate response, and output in a corresponding manner, specifically including two methods: direct output and role-adaptive adjustment output. When the comprehensive score and preset dimension scores of the target response meet the output conditions, the target response is directly output; when the target response does not meet the output requirements in a certain preset dimension score, the target response is subjected to corresponding role-adaptive adjustments before output, including partial replacement of segments with inconsistent character tone, blocking or rewriting out-of-bounds or disabled expression segments, or performing secondary generation after re-injecting the conversation context, etc.
[0042] This embodiment makes tiered decisions based on preset dimension scores and comprehensive scores. When the quality of the target response meets the standard, it outputs directly to ensure response efficiency. When a single dimension fails to meet the standard, it performs targeted repair through role adaptation adjustment. This avoids the waste of resources caused by discarding due to single defects and ensures the role consistency and security of the final output response, thereby improving the stability of dialogue generation and user experience.
[0043] In the above embodiments, this invention discloses a role dialogue generation method based on contrastive learning constraints. The method involves: acquiring current user dialogue information and identifying the target interaction role; acquiring a role parameter set and constructing a dialogue generation task; encoding the role parameter set into a role reference embedding vector; determining an adaptive temperature parameter based on the number of conversation rounds and the role complexity coefficient; injecting contrastive learning constraints during decoding to generate candidate responses based on the role reference embedding vector and the adaptive temperature parameter; calculating multi-dimensional scores for each candidate response and fusing them into a comprehensive score; determining the target response based on the score and outputting it. By injecting contrastive learning constraints during the decoding stage, candidate responses are made to move closer to the role reference embedding vector, and the constraint strength is dynamically adjusted in conjunction with the adaptive temperature parameter. Simultaneously, the target response output is filtered through multi-dimensional scores, ensuring the role consistency and stability of the generated dialogue content.
[0044] In one embodiment, determining the adaptive temperature parameter based on the current session round number and role complexity coefficient includes: Get the current session round number and the preset maximum session round number; The number of parameter fields, the number of constraint rules, and the granularity of the emotion range of the target interactive character are determined based on the character parameter set. The role complexity coefficient is determined based on the number of parameter fields, the number of constraint rules, and the granularity of the emotion interval. The adaptive temperature parameter is calculated based on the current session number, the preset maximum session number, and the role complexity coefficient.
[0045] In this embodiment, when determining the adaptive temperature parameter, the current session round number and the preset maximum session round number are first obtained. Then, the number of parameter fields, the number of constraint rules, and the granularity of the emotion interval are extracted from the role parameter set. The role complexity coefficient is determined by these three sub-indicators to quantify the fineness of the constraints on the target interactive role. Specifically, the number of parameter fields is the total number of structured fields configured in the role parameter set; more fields indicate a more refined role setting and higher complexity. The number of constraint rules is the number of explicit constraint rules defined in the role parameter set, including answerable boundary rules, disabled boundary rules, and tone preference rules; more rules indicate stricter role constraints. The granularity of the emotion interval is the degree of subdivision of the range of emotion expression of the role, such as dividing emotion intensity into coarse or fine types; finer granularity indicates more complex emotion control. After normalizing the number of parameter fields, the number of constraint rules, and the granularity of the emotion interval, a weighted sum is obtained according to preset weights to obtain the role complexity coefficient.
[0046] Then, the adaptive temperature parameter is calculated based on the current session round number, the maximum session round number, and the role complexity coefficient. Specifically, the adaptive temperature parameter can be calculated using the following formula. :
[0047] in, The reference temperature parameter is β, and the attenuation coefficient is β. Character complexity coefficient, This represents the current session round number. The maximum number of conversation rounds is set so that characters with higher complexity are subject to stricter contrastive learning constraints in long conversations, thereby suppressing the risk of tone drift for characters with high complexity.
[0048] This embodiment quantifies the role complexity coefficient by introducing the number of parameter fields, the number of constraint rules, and the granularity of the emotion interval. This makes the determination of the adaptive temperature parameter not only dependent on the conversation depth, but also fully considers the fineness of the constraints on the role itself. It realizes differentiated constraint control for roles with different complexities, avoids the problem of insufficient constraints for high-complexity roles or overly strict constraints for low-complexity roles caused by uniform temperature parameters, and ensures the accuracy of adaptive temperature parameter adjustment and role adaptability.
[0049] In one embodiment, step S204 includes: Autoregressive decoding processing is performed based on the aforementioned dialogue generation task; During the decoding process, the first semantic similarity between the current candidate word embedding and the role reference embedding vector is calculated, and the sampling temperature is adjusted according to the adaptive temperature parameter. The first semantic similarity is added to the decoding score as a bias term, and multiple candidate responses are generated by combining different sampling temperatures. The multiple candidate responses include high-constraint candidate responses, balanced candidate responses, and high-creativity candidate responses.
[0050] In this embodiment, the autoregressive decoding process is a decoding method in which the generative model predicts candidate responses word by word from left to right. That is, at each decoding step, the generative model calculates the conditional probability distribution of the next word based on the generated preceding sequence and samples the current word from it. During the decoding process, the contrastive learning constraint module calculates the first semantic similarity between the current candidate word embedding and the role reference embedding vector in real time. This first semantic similarity can be calculated using cosine similarity or dot product similarity to reflect the closeness of the current candidate word to the role's semantic reference point in the semantic space. The first semantic similarity is added as a bias term to the decoding score. Specifically, the original decoding score is added to the semantic similarity bias term according to a preset ratio to obtain a corrected decoding score. Then, word sampling is performed based on the corrected decoding score, thereby guiding the generative model to prioritize words that are close to the role's semantic reference point.
[0051] Furthermore, by adjusting the sampling temperature during the decoding process using adaptive temperature parameters, a higher sampling temperature results in a flatter sampling distribution and more diverse generated results; conversely, a lower sampling temperature results in a sharper sampling distribution and more deterministic generated results. This embodiment adjusts the sampling temperature based on adaptive temperature parameters and generates multiple candidate responses by combining different sampling temperatures. Specifically, Top-K sampling is performed at different sampling temperatures based on the corrected decoding score, generating at least one high-constraint candidate response, one balanced candidate response, and one high-creativity candidate response. The high-constraint candidate response is generated at a low sampling temperature, making the decoding process highly dependent on the corrected decoding score, and the generated results strictly adhere to the character setting. The balanced candidate response is generated at a medium sampling temperature, achieving a balance between character consistency and semantic diversity. The high-creativity candidate response is generated at a high sampling temperature, allowing for greater diversity in vocabulary selection and sentence structure. This enables the generation of multiple candidate responses under different constraint intensities based on the same dialogue generation task, allowing for subsequent selection between character constraints and creative expression.
[0052] This embodiment injects semantic similarity as a bias term into the decoding score in real time during the autoregressive decoding process, so that role constraints take effect from the generation source. At the same time, it generates multiple candidate responses with different constraint strengths by adjusting the sampling temperature in conjunction with adaptive temperature parameters, thereby improving the coverage and screening flexibility of the candidate response set.
[0053] In one embodiment, step S205 includes: Each candidate response is encoded into a candidate embedding vector, and a role consistency score is calculated between each candidate embedding vector and the role reference embedding vector. The corresponding security score is calculated based on the out-of-bounds topic hit rate and the banned topic hit rate of each candidate response. A relevance score is calculated based on the correlation between each candidate response and the user dialogue information; The role consistency score, security score, and relevance score of each candidate response are weighted and summed to obtain the comprehensive score of each candidate response.
[0054] In this embodiment, when scoring each candidate response across multiple preset dimensions and performing a comprehensive score, a role consistency score, a security score, and a relevance score are calculated separately and then weighted and summed to obtain the comprehensive score. Specifically, each candidate response is mapped to a dense vector representation in a high-dimensional semantic space using a response encoder, resulting in a corresponding candidate embedding vector. Based on the candidate embedding vector and the role reference embedding vector, a corresponding role consistency score is calculated to evaluate the degree of deviation of the candidate response from the role's semantic reference point in the overall semantic space.
[0055] The security score is determined by using a rule engine or classification model to detect out-of-bounds topics and prohibited topics in candidate responses. The out-of-bounds topic hit rate is the proportion of candidate responses involving topics outside the answerable boundaries of a role, and the prohibited topic hit rate is the proportion of candidate responses touching prohibited topics of a role. The two together constitute the security score. For example, the security score can be obtained by adding the two together and mapping them according to a preset mapping table. The higher the out-of-bounds topic hit rate or the prohibited topic hit rate, the lower the calculated security score.
[0056] The relevance score is determined based on the correlation between each candidate response and the user's dialogue information. Specifically, it can be determined by calculating the semantic similarity between the candidate response and the user's input data, as well as the contextual coherence between the candidate response and the historical data of the conversation. The semantic similarity reflects the degree to which the candidate response responds to the user's current intent, and the contextual coherence reflects the coherence between the candidate response and the historical dialogue. Similarly, the semantic similarity and contextual coherence can be added together and mapped according to a preset mapping table to obtain the relevance score.
[0057] Finally, the role consistency score, security score, and relevance score of each candidate response are weighted and summed to obtain a comprehensive score. During the weighted summation, the scores of each dimension are first normalized to eliminate differences in units of measurement, and then linearly combined according to preset weight coefficients. These preset weight coefficients can be dynamically adjusted based on the characteristics of the current interaction scenario or the target interaction role. For example, in security-sensitive scenarios, the weight of the security score is increased, and in role-playing scenarios, the weight of the role consistency score is increased, thus adapting to the scoring preferences in different scenarios.
[0058] In one embodiment, calculating the role consistency score between each of the candidate embedding vectors and the role reference embedding vector includes: Calculate the second semantic similarity between each of the candidate embedding vectors and the role reference embedding vector; Based on the adaptive temperature parameters and the pre-stored negative sample role embedding vectors, the contrastive learning loss between each candidate embedding vector and the role reference embedding vector is calculated. The second semantic similarity and the contrastive learning loss are fused to obtain the role consistency score between each candidate embedding vector and the role reference embedding vector.
[0059] In this embodiment, when calculating the role consistency score, the second semantic similarity between each candidate embedding vector and the role reference embedding vector is first calculated to determine the closeness between the candidate embedding vector and the role reference embedding vector in the semantic space. Specifically, the second semantic similarity can be calculated using the following formula:
[0060] in, For candidate embedding vectors, Using the role reference embedding vector, we obtain the second semantic similarity with a value range of [-1, 1]. The closer the value is to 1, the more similar the candidate response is to the role setting.
[0061] Furthermore, based on adaptive temperature parameters and pre-stored negative sample character embedding vectors, the contrastive learning loss between each candidate embedding vector and the character reference embedding vector is calculated. Specifically, the contrastive learning loss can be calculated based on the following loss function:
[0062] Where K represents the number of negative sample roles, such as other roles or common dialogues. Embed the vector for the j-th negative sample role. For adaptive temperature parameters, Let be the semantic similarity between the candidate embedding vector and the embedding vector of the j-th negative sample role.
[0063] This embodiment pre-stores a negative sample role embedding library in the database, containing embedding vectors of multiple non-target roles. When calculating the contrastive learning loss, the candidate embedding vector is used as the query vector, the role reference embedding vector is used as the positive sample key vector, and each vector in the negative sample role embedding library is used as the negative sample key vector. The contrastive learning loss is calculated based on an adaptive temperature parameter. The adaptive temperature parameter directly affects the gradient magnitude of the contrastive learning loss. The lower the temperature, the higher the required discriminative power between positive and negative samples, and the more severe the penalty for deviations from the role's semantics.
[0064] Then, the semantic similarity and contrastive learning loss are fused. The fusion method can be a linear weighted combination, such as... γ1 and γ2 are preset weight coefficients, which enable the role consistency score to simultaneously reflect the absolute proximity of the candidate response to the role semantic reference point and the relative distinguishability in the role space.
[0065] This embodiment introduces contrastive learning loss, which means that when evaluating the role consistency of candidate responses, it is necessary not only to be close to the role reference point in the absolute semantic space, but also to be distinguished from other roles in a relative sense. This strengthens the discrimination boundary of the role semantic space and improves the discriminativeness and reliability of the role consistency score.
[0066] In one embodiment, step S206 includes: The candidate response with the highest overall score is selected as the target response; The overall score of the target response is compared with the first threshold, and the role consistency score, security score, and relevance score of the target response are compared with the second threshold. The target response can be output directly based on the comparison results, or the target response can be adjusted accordingly before being output.
[0067] In this embodiment, the candidate response with the highest overall score is first selected from all candidate responses as the target response. Then, the overall score of the target response is compared with a first threshold, and the role consistency score, security score, and relevance score of the target response are compared with a second threshold. The first threshold is a comprehensive score threshold used to determine whether the overall quality of the candidate response meets the direct output standard, while the second threshold is a single-item score threshold used to determine whether the candidate response meets the qualification standard in each preset dimension. Both the first and second thresholds are determined based on historical interaction data statistics or training with manually labeled samples, and can be dynamically adjusted according to the current interaction scenario.
[0068] Specifically, multiple first and second thresholds for different scene types can be pre-stored, and different threshold configurations can be loaded according to the scene type when candidate responses are output. For example, in the narrative mode scene, both the first and second thresholds are moderately increased to ensure the rigor of the plot progression; in the casual conversation mode scene, the thresholds are moderately decreased to allow for more flexible dialogue styles, and so on.
[0069] When comparing scores, if the overall score is greater than the first threshold, the overall quality of the target response is deemed acceptable. If any of the role consistency score, safety score, or relevance score is greater than the second threshold, that dimension is deemed acceptable. Based on the comparison between the overall score and the first threshold, as well as the comparison between each dimension's score and the second threshold, a decision is made as to whether to directly output the target response or to perform role adaptation adjustments before outputting it. By introducing a dual-threshold judgment mechanism of the first and second thresholds, it is possible to accurately identify defects in individual dimensions while ensuring overall output quality, providing clear triggering conditions for subsequent role adaptation adjustments.
[0070] In one embodiment, the step of directly outputting the target response based on the comparison result, or outputting the target response after performing corresponding role adaptation adjustments, specifically includes: If the overall score is greater than the first threshold, and the role consistency score, security score, and relevance score are all greater than the second threshold, then the target response is directly output. If the overall score is greater than the first threshold and the role consistency score is less than the second threshold, then the role tone segment in the target response is partially replaced and output. If the overall score is greater than the first threshold and the security score is less than the second threshold, then the out-of-bounds segments and / or disabled expression segments in the target response are locally replaced and output. If the overall score is greater than the first threshold and the relevance score is less than the second threshold, then after re-injecting the current conversation mode and historical conversation summary, the dialogue generation task is executed again and a new target response is output.
[0071] In this embodiment, if the overall score of the current target response is greater than the first threshold, and its role consistency score, security score, or relevance score is greater than the second threshold, then the target response is directly output to improve the efficiency of the interaction response. If the overall score of the target response is greater than the first threshold, but any one of the role consistency score, security score, or relevance score is less than the second threshold, then the current target response is adjusted for role adaptation before being output.
[0072] Specifically, corresponding role adaptation adjustments can be performed through tone replacement submodule, safety filtering submodule, and secondary generation submodule. Each submodule is selectively activated based on threshold comparison results. Among them, the tone replacement submodule is used to locally replace the tone-inconsistent segments in the target response when the role consistency score fails to meet the standard. Specifically, it can identify the start and end positions of the deviation segments through a segment localization algorithm, retrieve standard sentence patterns matching the current emotional intensity range from the role tone template library, replace the deviation segments with standard sentence patterns, and ensure that the replaced response remains grammatically and semantically coherent.
[0073] The security filtering submodule is used to partially replace out-of-bounds segments and / or prohibited expressions in the target response when the security score fails to meet the standard. Specifically, it first identifies out-of-bounds segments and / or prohibited expressions in the target response through methods such as pattern matching or classification models. Then, it replaces prohibited expressions with semantically similar but compliant alternatives using synonym replacement or deletion. Finally, it guides out-of-bounds topics back to the role's answerable boundaries through a topic fallback strategy, ensuring the security of the generated content.
[0074] The secondary generation submodule is used to perform secondary generation after re-injecting the current conversation pattern and historical conversation summary when the relevance score is insufficient. Since an insufficient relevance score indicates that the target response has failed to accurately understand the user's intent or is out of context, simple local replacement is insufficient to fundamentally correct semantic biases. Therefore, the current conversation pattern identifier and historical conversation summary are re-injected into the dialogue generation task to strengthen pattern constraints and context constraints. After re-executing the dialogue generation task, a new set of candidate responses is generated, from which a new target response is selected for output. During the secondary generation process, the original target response can be retained as a negative sample reference to avoid the same semantic biases from recurring in the secondary generation results.
[0075] This embodiment performs differentiated role adaptation rewriting for different dimensions of non-compliance, achieving precise repair of target response defects. It avoids regenerating regardless of the defect, ensuring output quality while reducing computational overhead and improving interactive response efficiency and the pertinence of role adaptation adjustments.
[0076] It should be noted that there is no necessary order between the above steps. Those skilled in the art will understand from the description of the embodiments of the present invention that the above steps may have different execution orders in different embodiments, that is, they may be executed in parallel or in turn, etc.
[0077] Further reference Figure 3 As a response to the above Figure 2 The present invention provides an embodiment of a role dialogue generation device based on contrastive learning constraints, which implements the method shown. Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0078] like Figure 3 As shown, the role dialogue generation device 30 based on contrastive learning constraints described in this embodiment includes: Data acquisition module 301 is used to acquire current user dialogue information and confirm the target interaction role; The task construction module 302 is used to obtain the corresponding role parameter set according to the target interaction role, and construct a dialogue generation task according to the role parameter set and the user dialogue information; The parameter confirmation module 303 is used to encode the character parameter set into a corresponding character reference embedding vector, and determine the adaptive temperature parameter according to the current session round and the character complexity coefficient. The constraint decoding module 304 is used to decode the dialogue generation task and inject contrastive learning constraints according to the role reference embedding vector and the adaptive temperature parameter during the decoding process to generate at least one candidate response. The response scoring module 305 is used to calculate the scores of each candidate response in a preset dimension and fuse them to obtain the corresponding comprehensive score. The filtering output module 306 is used to directly determine and output the target response from each candidate response based on the scores of each preset dimension and the comprehensive score, or to output the target response after role adaptation adjustment.
[0079] The module referred to in this invention is a series of computer program instruction segments that can perform specific functions. It is more suitable than a program for describing the execution process of role dialogue generation based on contrastive learning constraints. For specific implementation methods of each module, please refer to the corresponding method embodiments above, which will not be repeated here.
[0080] In one embodiment, the parameter confirmation module 303 includes: The session parameter acquisition unit is used to obtain the current session round number and the preset maximum session round number; The role data confirmation unit is used to confirm the number of parameter fields, the number of constraint rules, and the granularity of the emotion range of the target interactive role based on the role parameter set. The complexity calculation unit is used to determine the role complexity coefficient based on the number of parameter fields, the number of constraint rules, and the granularity of the emotion interval. The temperature parameter calculation unit is used to calculate adaptive temperature parameters based on the current session round number, the preset maximum session round number, and the role complexity coefficient.
[0081] In one embodiment, the constraint decoding module 304 includes: An autoregressive decoding unit is used to perform autoregressive decoding processing based on the dialogue generation task; The constraint calculation unit is used to calculate the first semantic similarity between the current candidate word embedding and the role reference embedding vector during the decoding process, and adjust the sampling temperature according to the adaptive temperature parameter. The candidate response generation unit is used to add the first semantic similarity as a bias term to the decoding score and generate multiple candidate responses by combining different sampling temperatures. The multiple candidate responses include high-constraint candidate responses, balanced candidate responses, and high-creativity candidate responses.
[0082] In one embodiment, the response scoring module 305 includes: A consistency scoring unit is used to encode each of the candidate responses into a candidate embedding vector and calculate a role consistency score between each of the candidate embedding vectors and the role reference embedding vector. A security scoring unit is used to calculate the corresponding security score based on the out-of-bounds topic hit rate and the disabled topic hit rate of each candidate response. A relevance scoring unit is used to calculate a relevance score based on the correlation between each candidate response and the user dialogue information; The weighted summation unit is used to perform a weighted summation of the role consistency score, security score, and relevance score of each candidate response to obtain a comprehensive score for each candidate response.
[0083] In one embodiment, the consistency scoring unit includes: A similarity calculation unit is used to calculate the second semantic similarity between each of the candidate embedding vectors and the role reference embedding vector; Based on the adaptive temperature parameters and the pre-stored negative sample role embedding vectors, the contrastive learning loss between each candidate embedding vector and the role reference embedding vector is calculated. The second semantic similarity and the contrastive learning loss are fused to obtain the role consistency score between each candidate embedding vector and the role reference embedding vector.
[0084] In one embodiment, the filtering output module 306 includes: The target response determination unit is used to determine the candidate response with the highest comprehensive score as the target response; The threshold comparison unit is used to compare the comprehensive score of the target response with a first threshold, and to compare the role consistency score, security score, and relevance score of the target response with a second threshold. The response output unit is used to directly output the target response based on the comparison result, or to perform corresponding role adaptation adjustments on the target response and then output it.
[0085] In one embodiment, the response output unit is specifically used for: If the overall score is greater than the first threshold, and the role consistency score, security score, and relevance score are all greater than the second threshold, then the target response is directly output. If the overall score is greater than the first threshold and the role consistency score is less than the second threshold, then the role tone segment in the target response is partially replaced and output. If the overall score is greater than the first threshold and the security score is less than the second threshold, then the out-of-bounds segments and / or disabled expression segments in the target response are locally replaced and output. If the overall score is greater than the first threshold and the relevance score is less than the second threshold, then after re-injecting the current conversation mode and historical conversation summary, the dialogue generation task is executed again and a new target response is output.
[0086] In the above embodiments, the present invention discloses a role dialogue generation device based on contrastive learning constraints. This device acquires current user dialogue information and identifies the target interaction role; obtains a role parameter set and constructs a dialogue generation task; encodes the role parameter set into a role reference embedding vector, and determines an adaptive temperature parameter based on the number of conversation rounds and the role complexity coefficient; during decoding, it injects contrastive learning constraints based on the role reference embedding vector and the adaptive temperature parameter to generate candidate responses; calculates multi-dimensional scores for each candidate response and merges them into a comprehensive score; and determines and outputs the target response based on the score. By injecting contrastive learning constraints during the decoding stage, candidate responses are made to move closer to the role reference embedding vector, and the constraint strength is dynamically adjusted in conjunction with the adaptive temperature parameter. Simultaneously, the target response output is filtered through multi-dimensional scoring, ensuring the role consistency and stability of the generated dialogue content.
[0087] Specific limitations regarding the character dialogue generation device based on contrastive learning constraints can be found in the limitations of the character dialogue generation method based on contrastive learning constraints described above, and will not be repeated here. Each module in the aforementioned character dialogue generation device based on contrastive learning constraints can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0088] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0089] Another embodiment of the present invention provides a computer device, such as... Figure 4 As shown, computer device 40 includes: One or more processors 401 and memory 402, Figure 4 The following section uses a processor 401 as an example. The processor 401 and the memory 402 can be connected via a bus or other means. Figure 4 Taking the example of a connection between China and Israel via a bus.
[0090] The processor 401 is used to perform various control logics of the computer device 40. It can be any conventional processor, microprocessor, state machine, general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), microcontroller, ARM (Acorn RISC Machine) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components.
[0091] The memory 402, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions corresponding to the role dialogue generation method based on contrastive learning constraints in this embodiment of the invention. The processor 401 executes various functional applications and data processing of the computer device 40 by running the non-volatile software programs, instructions, and units stored in the memory 402, thereby implementing the role dialogue generation method based on contrastive learning constraints in the above method embodiment.
[0092] Another embodiment of the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when executed by one or more processors, perform the steps of the role dialogue generation method based on contrastive learning constraints in any of the above method embodiments.
[0093] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0094] Based on the above description of the embodiments, those skilled in the art will understand that the methods described in the embodiments can be implemented using software plus necessary general-purpose hardware platforms. Of course, they can also be implemented using hardware, but in many cases, the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0095] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The computer program can be stored in a non-volatile, computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The storage medium can be a memory, magnetic disk, floppy disk, flash memory, optical storage, etc.
[0096] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.
[0097] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for generating role dialogue based on contrastive learning constraints, characterized in that, include: Obtain current user conversation information and confirm the target interaction role; Obtain the corresponding role parameter set according to the target interaction role, and construct a dialogue generation task based on the role parameter set and the user dialogue information; The character parameter set is encoded into a corresponding character reference embedding vector, and the adaptive temperature parameter is determined based on the current session round and the character complexity coefficient. The dialogue generation task is decoded, and during the decoding process, contrastive learning constraints are injected based on the role reference embedding vector and the adaptive temperature parameter to generate at least one candidate response. Calculate the scores of each candidate response in a preset dimension, and fuse them to obtain the corresponding comprehensive score; Based on the scores of each preset dimension and the comprehensive score, the target response is directly determined and output from each candidate response, or the target response is adjusted for role adaptation and then output. The step of determining the adaptive temperature parameter based on the current session round number and role complexity coefficient includes: Get the current session round number and the preset maximum session round number; The number of parameter fields, the number of constraint rules, and the granularity of the emotion range of the target interactive character are determined based on the character parameter set. The role complexity coefficient is determined based on the number of parameter fields, the number of constraint rules, and the granularity of the emotion interval. The adaptive temperature parameter is calculated based on the current session number, the preset maximum session number, and the role complexity coefficient.
2. The role dialogue generation method based on contrastive learning constraints according to claim 1, characterized in that, The process of decoding the dialogue generation task, and injecting contrastive learning constraints based on the role reference embedding vector and the adaptive temperature parameter during decoding, generates at least one candidate response, including: Autoregressive decoding processing is performed based on the aforementioned dialogue generation task; During the decoding process, the first semantic similarity between the current candidate word embedding and the role reference embedding vector is calculated, and the sampling temperature is adjusted according to the adaptive temperature parameter. The first semantic similarity is added to the decoding score as a bias term, and multiple candidate responses are generated by combining different sampling temperatures. The multiple candidate responses include high-constraint candidate responses, balanced candidate responses, and high-creativity candidate responses.
3. The role dialogue generation method based on contrastive learning constraints according to claim 1, characterized in that, The calculation of scores for each candidate response across a preset dimension, and the fusion of these scores to obtain a corresponding comprehensive score, includes: Each candidate response is encoded into a candidate embedding vector, and a role consistency score is calculated between each candidate embedding vector and the role reference embedding vector. The corresponding security score is calculated based on the out-of-bounds topic hit rate and the banned topic hit rate of each candidate response. A relevance score is calculated based on the correlation between each candidate response and the user dialogue information; The role consistency score, security score, and relevance score of each candidate response are weighted and summed to obtain the comprehensive score of each candidate response.
4. The role dialogue generation method based on contrastive learning constraints according to claim 3, characterized in that, The calculation of the role consistency score between each of the candidate embedding vectors and the role reference embedding vector includes: Calculate the second semantic similarity between each of the candidate embedding vectors and the role reference embedding vector; Based on the adaptive temperature parameters and the pre-stored negative sample role embedding vectors, the contrastive learning loss between each candidate embedding vector and the role reference embedding vector is calculated. The second semantic similarity and the contrastive learning loss are fused to obtain the role consistency score between each candidate embedding vector and the role reference embedding vector.
5. The role dialogue generation method based on contrastive learning constraints according to claim 3, characterized in that, The step of directly determining and outputting the target response from the candidate responses based on the scores of each preset dimension and the comprehensive score, or outputting the target response after role adaptation adjustment, includes: The candidate response with the highest overall score is selected as the target response; The overall score of the target response is compared with the first threshold, and the role consistency score, security score, and relevance score of the target response are compared with the second threshold. The target response can be output directly based on the comparison results, or the target response can be adjusted accordingly before being output.
6. The role dialogue generation method based on contrastive learning constraints according to claim 5, characterized in that, The step of directly outputting the target response based on the comparison results, or outputting the target response after performing corresponding role adaptation adjustments, specifically includes: If the overall score is greater than the first threshold, and the role consistency score, security score, and relevance score are all greater than the second threshold, then the target response is directly output. If the overall score is greater than the first threshold and the role consistency score is less than the second threshold, then the role tone segment in the target response is partially replaced and output. If the overall score is greater than the first threshold and the security score is less than the second threshold, then the out-of-bounds segments and / or disabled expression segments in the target response are locally replaced and output. If the overall score is greater than the first threshold and the relevance score is less than the second threshold, then after re-injecting the current conversation mode and historical conversation summary, the dialogue generation task is executed again and a new target response is output.
7. A role dialogue generation device based on contrastive learning constraints, characterized in that, include: The data acquisition module is used to acquire current user dialogue information and confirm the target interaction role; The task construction module is used to obtain the corresponding role parameter set according to the target interaction role, and construct a dialogue generation task according to the role parameter set and the user dialogue information; The parameter confirmation module is used to encode the character parameter set into a corresponding character reference embedding vector, and determine the adaptive temperature parameter based on the current session round and the character complexity coefficient. The constraint decoding module is used to decode the dialogue generation task and inject contrastive learning constraints according to the role reference embedding vector and the adaptive temperature parameter during the decoding process to generate at least one candidate response. The response scoring module is used to calculate the scores of each candidate response in a preset dimension and fuse them to obtain the corresponding comprehensive score. The filtering output module is used to directly determine and output the target response from each candidate response based on the scores of each preset dimension and the comprehensive score, or to output the target response after role adaptation adjustment. The parameter confirmation module includes: The session parameter acquisition unit is used to obtain the current session round number and the preset maximum session round number; The role data confirmation unit is used to confirm the number of parameter fields, the number of constraint rules, and the granularity of the emotion range of the target interactive role based on the role parameter set. The complexity calculation unit is used to determine the role complexity coefficient based on the number of parameter fields, the number of constraint rules, and the granularity of the emotion interval. The temperature parameter calculation unit is used to calculate adaptive temperature parameters based on the current session round number, the preset maximum session round number, and the role complexity coefficient.
8. A computer device, characterized in that, Includes at least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the role dialogue generation method based on contrastive learning constraints as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform the role dialogue generation method based on contrastive learning constraints as described in any one of claims 1-6.