Intelligent text dialogue generation method and device based on artificial intelligence

By semantic analysis and business model matching of user input text, and personalized optimization of Transformer model and user emotional information, the fluency and universality of existing text dialogue generation methods in the face of diversified natural language expressions is solved, and more efficient dialogue system performance and user stickiness are achieved.

CN120144715AActive Publication Date: 2025-06-13GUANGDONG POWER GRID CO LTD INFORMATION CENT

Patent Information

Application Number
CN202510257717.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-13
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

Existing text dialogue generation methods are difficult to provide accurate responses when facing diversified and flexible natural language expressions, resulting in poor fluency in the dialogue system and limited coverage of the corpus, which cannot cope with new topics or novel expressions, affecting universality and adaptability.

Method used

Using an intelligent text dialogue generation method based on artificial intelligence, we use text semantic analysis of the user input text, match the preset business model library, use the attention mechanism Transformer model to predict, and personalize the optimization of user emotional information and user portraits to generate the final reply dialogue text.

Benefits of technology

It improves the fluency, versatility and adaptability of the dialogue system, can more accurately capture key information in natural language expression, provide responses that are more in line with user needs, and increases the user's stickiness to the dialogue system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144715A_ABST
    Figure CN120144715A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent text dialogue generation method and device based on artificial intelligence, and the method comprises the steps: carrying out the text semantic analysis of a question dialogue text inputted by a target user, and obtaining the text semantic information, user emotion information and business field information of the question dialogue text; performing matching in a preset business model library based on the business field information to obtain a target business model; inputting the text semantic information into a target business model to obtain an initial reply dialogue text output by the target business model; performing emotion adjustment on the initial reply dialogue text based on the user emotion information to obtain an intermediate reply dialogue text matched with the emotion tendency of the target user; and performing personalized optimization on the intermediate reply dialogue text based on the current dialogue situation and the user portrait of the target user to obtain a target reply dialogue text, and feeding back the target reply dialogue text to the target user. According to the method, the fluency, universality and adaptability of the dialogue system are improved, and the stickiness of a user to the dialogue system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to an intelligent text dialogue generation method and device based on artificial intelligence. Background Art

[0002] With the continuous development of artificial intelligence technology, text dialogue systems have been widely used in many fields, such as intelligent customer service, intelligent chatbots, etc.

[0003] The existing text dialogue generation methods are mainly rule-based text dialogue generation methods and retrieval-based text dialogue generation methods. The rule-based text dialogue generation method matches the input text of the user through set dialogue rules and templates. However, fixed dialogue rules and templates are difficult to cope with diverse and flexible natural language expressions. For example, when the user's expression deviates slightly from the preset rules, an accurate response may not be given, resulting in poor fluency of the dialogue system. The retrieval-based text dialogue generation method retrieves similar questions in a pre-constructed dialogue corpus to generate answers. However, the coverage of the corpus is limited. If a new topic or novel expression not included in the corpus is encountered, a suitable answer may not be provided, seriously affecting the generality and adaptability of the dialogue system. Summary of the Invention

[0004] The present invention provides an intelligent text dialogue generation method and device based on artificial intelligence to improve the fluency, generality and adaptability of the dialogue system and increase the stickiness of users to the dialogue system.

[0005] In a first aspect, the present invention provides an intelligent text dialogue generation method based on artificial intelligence, including:

[0006] Performing text semantic parsing on the problem dialogue text input by the target user to obtain the text semantic information, user emotion information and business domain information of the problem dialogue text;

[0007] Based on the business domain information, performing matching in a preset business model library to obtain a target business model; the target business model is trained based on the sample text semantics and the label results of the corresponding reply dialogue text for a Transformer model based on an attention mechanism;

[0008] Inputting the text semantic information into the target business model to obtain an initial reply dialogue text output by the target business model;

[0009] Performing emotion adjustment on the initial reply dialogue text based on the user emotion information to obtain an intermediate reply dialogue text that matches the emotion tendency of the target user;

[0010] Personalize and optimize the intermediate reply dialogue text based on the current dialogue context and the user profile of the target user to obtain a target reply dialogue text and feedback it to the target user; the user profile is constructed by analyzing the personality characteristics and interest preferences based on the historical information of the target user.

[0011] In a second aspect, the present invention also provides an intelligent text dialogue generation device based on artificial intelligence, which is applied to the intelligent text dialogue generation method based on artificial intelligence as described in the first aspect; the intelligent text dialogue generation device based on artificial intelligence includes:

[0012] A semantic parsing module, configured to perform text semantic parsing on the question dialogue text input by the target user to obtain the text semantic information, user emotion information, and business domain information of the question dialogue text;

[0013] A model matching module, configured to perform matching in a preset business model library based on the business domain information to obtain a target business model; the target business model is trained based on the sample text semantics and the label results of the corresponding reply dialogue text for a Transformer model based on the attention mechanism;

[0014] A model prediction module, configured to input the text semantic information into the target business model to obtain an initial reply dialogue text output by the target business model;

[0015] A text adjustment module, configured to perform emotion adjustment on the initial reply dialogue text based on the user emotion information to obtain an intermediate reply dialogue text that matches the emotion tendency of the target user;

[0016] A text optimization and feedback module, configured to perform personalized optimization on the intermediate reply dialogue text based on the current dialogue context and the user profile of the target user to obtain a target reply dialogue text and feedback it to the target user; the user profile is constructed by analyzing the personality characteristics and interest preferences based on the historical information of the target user.

[0017] In a third aspect, the present invention also provides an electronic device, including: a memory for storing a computer software program; a processor for reading and executing the computer software program to implement the intelligent text dialogue generation method based on artificial intelligence as described in any one of the above.

[0018] In a fourth aspect, the present invention also provides a non-transitory computer-readable storage medium, in which a computer software program is stored, and when the computer software program is executed by a processor, the intelligent text dialogue generation method based on artificial intelligence as described in any one of the above is implemented.

[0019] Fifth aspect, the present invention provides a computer program product, including a computer program, which when executed by a processor implements any of the above artificial intelligence-based intelligent text dialogue generation methods.

[0020] In the intelligent text dialogue generation method based on artificial intelligence provided by the embodiments of the present invention, the target business model trained by the Transformer model based on the attention mechanism is used to predict the text semantic information. Since the multi-head attention mechanism can simultaneously focus on different semantic levels of the text, the target business model can understand the input text more comprehensively and deeply. When facing complex and ever-changing natural language expressions, it can capture key information more accurately. Therefore, it can accurately output the reply dialogue text, improving the fluency of the dialogue system. On the other hand, by matching the business domain information in the user's input text to the corresponding target business model, and personalizing and optimizing the initial reply dialogue text according to the user's emotional information, the current dialogue context, and the user's user profile, the final target reply dialogue text that meets the user's needs is obtained, so as to better meet the dialogue requirements of various different scenarios and topics, no longer limited to retrieving a specific corpus range, improving the flexibility and versatility of the dialogue system. Therefore, the stickiness of users to the dialogue system can be increased. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a schematic flowchart of the intelligent text dialogue generation method based on artificial intelligence provided by the embodiments of the present invention;

[0022] Figure 2 is a schematic structural diagram of the intelligent text dialogue generation device based on artificial intelligence provided by the embodiments of the present invention;

[0023] Figure 3 is an embodiment diagram of the electronic device provided by the embodiments of the present invention;

[0024] Figure 4 is an embodiment diagram of the computer-readable storage medium provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0026] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined. In the description of the present invention, the term "for example" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "for example" in the present invention is not necessarily construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to implement and use the present invention. In the following description, details are set forth for purposes of explanation. It should be understood that those of ordinary skill in the art can recognize that the present invention can be implemented without the use of these specific details. In other instances, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0027] Refer to Figure 1 , Figure 1 FIG. is a schematic flow chart of an intelligent text dialogue generation method based on artificial intelligence provided by the present invention. In the embodiments of the present invention, the execution subject of the intelligent text dialogue generation method based on artificial intelligence is an intelligent dialogue device. Therefore, the intelligent text dialogue generation method based on artificial intelligence includes:

[0028] Step 10: Perform text semantic parsing on the problem dialogue text input by the target user to obtain the text semantic information, user emotion information, and business domain information of the problem dialogue text.

[0029] Among them, the intelligent dialogue device in the embodiments of the present invention provides an intelligent dialogue interface. Therefore, when the user conducts an intelligent dialogue with the intelligent dialogue device, the user needs to input corresponding text on the intelligent dialogue interface. Therefore, the intelligent dialogue device can receive the problem dialogue text input by the target user.

[0030] Further, the intelligent dialogue device performs lexical analysis on the problem dialogue text, splits the problem dialogue text into multiple individual words, and then generates corresponding word vector representations for each word through a preset word vector model, such as the FastText word vector model or the ELMo (Embeddings from Language Models) word vector model where i represents the serial number of the word in the problem dialogue text.

[0031] Further, the intelligent dialogue device constructs a syntax tree through syntactic analysis to determine the syntactic relationships between words. For each node (corresponding to a word or phrase) in the syntax tree, a syntactic weight g is assigned according to its level in the syntax tree and the adjacent node relationships i , where the syntactic weight g i The value range can be set from 0 to 1, and the closer to the root of the syntax tree and the more important syntactic components are associated, the higher the weight

[0032] Further, the intelligent dialogue device identifies the semantic roles (such as agent, patient, etc.) in the question dialogue text through the semantic role labeling method, and assigns a role weight s to each semantic role i , where the role weight s i The value range is from 0 to 1, and it is set according to the role importance

[0033] Further, the intelligent dialogue device constructs the text semantic information S of the question dialogue text according to the word vector representation of each word in the question dialogue text the syntactic weight g i and the role weight s i . The text semantic information S of the question dialogue text can be expressed as:

[0034]

[0035] where n represents the total number of words in the question dialogue text

[0036] Optionally, an emotion dictionary is pre - constructed in the intelligent dialogue device of the embodiment of the present invention. The emotion dictionary contains positive, negative, and neutral emotion words and their corresponding emotion intensity values. For example, the intensity range of positive words is [1,5], the negative words are [-1,-5], and the neutral words are 0

[0037] Further, the intelligent dialogue device performs word matching search on the question dialogue text according to the emotion dictionary to obtain the emotion intensity e corresponding to the emotion words matched by the question dialogue text i . Further, the intelligent dialogue device determines the position information of the emotion words in the question dialogue text, and assigns a corresponding position weight p to the emotion words according to the position information i , for example, the closer the emotion word is to the beginning or the end of the question dialogue text, the higher its position weight, and the value range is from 0 to 1. Further, the intelligent dialogue device determines the text type of the question dialogue text, and assigns a tone weight m to the emotion words in the question dialogue text according to the text type. The text types include declarative sentences, interrogative sentences, and exclamatory sentences. For declarative sentences, the tone weight m = 0.5, for interrogative sentences, the tone weight m = 0.8, and for exclamatory sentences, the tone weight m = 1

[0038] Further, the intelligent dialogue device according to the emotion intensity ei 、Position weight p i and tone weight m to determine the user sentiment information E of the question dialogue text. The user sentiment information E can be expressed as:

[0039]

[0040] where k represents the number of sentiment words matched in the question dialogue text.

[0041] Optionally, a business domain keyword set library is pre - constructed in the intelligent dialogue device of the embodiment of the present invention. Each business domain corresponds to a set of representative keywords and their weights. The weights are set according to the importance and typicality of the keywords in the domain, and the range can be set from 0 to 1.

[0042] Further, the intelligent dialogue device extracts keywords from the question dialogue text through the TF - IDF algorithm to obtain the keyword kw in the question dialogue text j , and determines the weight of the keyword kw j according to the business domain keyword set library Further, the intelligent dialogue device determines the matching degree score M of the keyword kw j and its weight . The matching degree score M of the keyword kw j is calculated as follows: d . The calculation formula of the matching degree score m d is as follows:

[0043]

[0044] where d represents the business domain and l represents the number of keywords related to the business domain extracted.

[0045] Further, the intelligent dialogue device takes the business domain corresponding to the maximum matching degree score m d and converts the domain title text of this business domain into the business domain information T to which the question dialogue text belongs m .

[0046] Step 20: Match in the preset business model library based on the business domain information to obtain the target business model.

[0047] Optionally, a business model library is pre - constructed in the intelligent dialogue device of the embodiment of the present invention. The business model library includes multiple business models. Therefore, the intelligent dialogue device traverses the business domain annotations T corresponding to each business model Model in the preset business model library model , and compares the business domain information T of the question dialogue text m with the business domain annotation T of each business model Model modelCalculate the similarity to obtain the similarity value sim(T m , T model ) for each business model Model. Among them, the similarity value sim(T m , T model ) can be calculated by the cosine similarity algorithm, and its value range is from 0 to 1.

[0048] Furthermore, the intelligent dialogue device traverses the similarity values sim(T m , T model ) of each business model Model in the business model library, and determines the business model corresponding to the maximum similarity value sim(T m , T model ) as the target business model Model target .

[0049] Among them, the target business model in the embodiments of the present invention is obtained by training a Transformer model based on the semantic information of the sample text and the label results of the corresponding reply dialogue text, as specifically described in steps 60 to 90.

[0050] Step 30: Input the text semantic information into the target business model to obtain the initial reply dialogue text output by the target business model.

[0051] Furthermore, the intelligent dialogue device inputs the text semantic information into the target business model. The target business model processes the text semantic information through the attention mechanism and outputs the initial reply dialogue text of the question dialogue text. Therefore, the intelligent dialogue device can obtain the initial reply dialogue text output by the target business model, as specifically described in steps 301 to 303.

[0052] Step 40: Perform emotional adjustment on the initial reply dialogue text based on the user's emotional information to obtain an intermediate reply dialogue text that matches the emotional tendency of the target user.

[0053] Furthermore, the intelligent dialogue device performs emotional adjustment on the initial reply dialogue text through the user's emotional information to obtain an intermediate reply dialogue text that matches the emotional tendency of the target user, as specifically described in steps 401 to 404.

[0054] Step 50: Perform personalized optimization on the intermediate reply dialogue text based on the current dialogue context and the user profile of the target user to obtain the target reply dialogue text and feedback it to the target user.

[0055] Optionally, the intelligent dialogue device pre - constructs a user profile for each user, binds the user profile of each user with the user identifier of each user, and subsequently can directly match the user profile of the target user according to the user identifier of the target user. The construction process of the user profile is as follows: Obtain the historical information of the user, where the historical information includes historical conversation records, historical browsing behaviors, etc. Further, extract the personality traits and interest preferences in the historical information through data mining and analysis techniques, and construct the user profile of the user based on the personality traits and interest preferences. Among them, the personality traits such as the degree of extroversion extro, with a value range of 0 to 1; the degree of decision - making firmness dec, with a value range of 0 to 1. The interest preference can be represented by an interest vector I, and each dimension corresponds to the preference degree in different interest fields, with a value range of 0 to 1.

[0056] Further, the intelligent dialogue device extracts the key elements in the current dialogue context, as well as the personality traits and interest preferences in the user profile of the target user, and performs personalized optimization on the intermediate reply dialogue text according to the key elements in the current dialogue context and the personality traits and interest preferences in the user profile of the target user to obtain the target reply dialogue text. Further, the intelligent dialogue device feeds back the target reply dialogue text as the answer to the question dialogue text to the target user through the intelligent interaction interface, specifically as described in steps 501 to 503.

[0057] In the embodiment of the present invention, the target service model trained by the Transformer model based on the attention mechanism is used to predict the text semantic information. Since the multi - head attention mechanism can simultaneously focus on different semantic levels of the text, the target service model can understand the input text more comprehensively and deeply. When facing complex and variable natural language expressions, it can capture key information more accurately. Therefore, it can accurately output the reply dialogue text, improving the fluency of the dialogue system. On the other hand, by matching the target service model corresponding to the business domain information in the user's input text, and performing personalized optimization on the initial reply dialogue text according to the user's emotional information, the current dialogue context and the user's user profile, the final target reply dialogue text that meets the user's needs is obtained, so as to better meet the dialogue requirements of various different scenarios and topics, no longer limited to retrieving a specific corpus range, improving the flexibility and generality of the dialogue system. Therefore, the stickiness of users to the dialogue system can be increased.

[0058] In one embodiment, the descriptions of steps 301 to 303 are as follows:

[0059] Step 301, pre - process the text semantic information, obtain the position encoding vector corresponding to the text vector at each position in the text semantic information, and perform element - by - element addition on the text vector at each position in the text semantic information and its corresponding position encoding vector to obtain the text semantic information after position encoding.

[0060] Optionally, the target business model in the embodiments of the present invention is trained based on a Transformer model with an attention mechanism. The Transformer model includes an encoder and a decoder.

[0061] For the encoder part, the encoder receives the input text semantic information S, and the dimension of the text semantic information S is d model , and performs feature extraction and encoding conversion on it through multiple encoder layers. Each encoder layer includes a first multi-head attention module and a feed-forward neural network module, and a residual connection and a layer normalization operation module are also added between the two modules.

[0062] For the decoder part, the decoder is composed of multiple decoder layers. Each decoder layer, in addition to including a second multi-head attention module, a feed-forward neural network module, and corresponding residual connection and layer normalization operation modules, also has an additional third multi-head attention module for paying attention to the output of the encoder, and gradually generates a reply dialogue text by combining the features encoded by the input text semantic information.

[0063] Before inputting the text semantic information S into the encoder of the target business model Model target , positional encoding needs to be performed on it because the Transformer model itself does not explicitly model the position information of the input sequence. Therefore, the intelligent dialogue device preprocesses the text semantic information S to obtain the position encoding vector PE i (with dimension d model ) corresponding to the text vector S i (with dimension d model ) at the i-th position in the text semantic information S. Among them, in the preprocessing process, for each dimension j in dimension d model , where j = 0, 1,..., d model , the processing formula is as follows:

[0064]

[0065] Furthermore, the intelligent dialogue device adds each text vector S i in the input text semantic information S element-wise to its corresponding position encoding vector PE i to obtain the input vector after positional encoding The semantic information of the entire text after position encoding

[0066] In step 302, the semantic information of the text after position encoding is input into the target business model. Each encoder layer in the encoder extracts features and performs encoding transformation on the semantic information of the text after position encoding through the first multi-head attention mechanism module, the feed-forward neural network module, and the residual connection and layer normalization operation module, generating the final encoder output.

[0067] Furthermore, the intelligent dialogue device inputs the semantic information of the text after position encoding into the target business model. For each encoder layer in the encoder, the encoder layer extracts features and performs encoding transformation on the semantic information of the text after position encoding through the first multi-head attention mechanism module, the feed-forward neural network module, and the residual connection and layer normalization operation module. After being processed by N encoder layers in sequence, the final encoder output is generated.

[0068] Regarding the calculation process of the encoder layer in the encoder:

[0069] Through the first multi-head attention mechanism (Multi-Head Attention) module, the semantic information of the text after position encoding is linearly transformed to obtain the query vector Q, the key vector K, and the value vector V respectively.

[0070] In one embodiment, W Q , W K , W V are three different learnable weight matrices (with dimensions all being d model *d k , where h is the number of heads of the multi-head attention mechanism). Therefore, the calculation processes of the query vector Q, the key vector K, and the value vector V are as follows:

[0071]

[0072] Therefore, the dimension d of the semantic information of the text after position encoding model is mapped to the dimension d k of the multi-head attention mechanism, and it is split into h heads (with the dimension of each head being d k ), that is, Q = [Q 1 , Q 2 ,..., Q h , K = [K 1 , K 2 ,..., K h , V = [V 1 , V 2 ,..., V h .

[0073] Further, for each head i , i = 1, 2, ..., h, calculate the attention scores

[0074]

[0075] where the softmax function is used to normalize the attention scores so that their sum is 1, which is to scale the dot product result to avoid problems such as vanishing gradients caused by overly large numerical values.

[0076] Further, concatenate the attention outputs of each head to obtain the output M of the first multi - head attention mechanism, and the dimension returns to d model , and then pass it through the linear transformation weight matrix W O (with dimension hd k * d model ) for transformation:

[0077] Further, input the output M of the first multi - head attention mechanism into the Feed Forward Network to obtain the output F of the feed - forward neural network. The feed - forward neural network consists of two linear transformation layers and an activation function. The weight matrix of the first linear transformation is W 1 (with dimension d model * d ff ), the weight matrix of the second linear transformation is W 2 (with dimension d ff * d model ), and the intermediate activation function is a preset non - linear function σ custom . Among them, the non - linear function is a complex function that combines the characteristics of polynomial functions and exponential functions. The output F of the feed - forward neural network is calculated as follows:

[0078] F = σ custom (MW 1 )W 2 .

[0079] Further, perform residual connection and layer normalization processing. Add residual connection and layer normalization operations before and after the multi - head attention mechanism and the feed - forward neural network to help the model train better and avoid problems such as vanishing or exploding gradients. For the output M of the multi - head attention mechanism and the text semantic information of the position - encoded input The output after residual connection and layer normalization is:

[0080] Among them, LayerNorm represents the layer normalization operation, which normalizes each dimension of the vector.

[0081] Similarly, for the output F of the feed-forward neural network and the output M of the multi-head attention mechanism of its input, the output of the final encoder layer is

[0082] After being processed by N encoder layers in sequence, the final encoder output representation H is obtained E 。

[0083] Step 303: Input the final encoder output into the decoder. Each decoder layer in the decoder combines the text semantic information after position encoding and the final encoder output through the second multi-head attention mechanism module, the feed-forward neural network module, the residual connection and the layer normalization operation module, and the third multi-head attention mechanism module to generate the initial reply dialogue text.

[0084] Among them, the first multi-head attention mechanism (the second multi-head attention mechanism module) in the decoder is a masked multi-head attention mechanism (Masked Multi-Head Attention) module, which is used to avoid seeing future information when generating the reply dialogue text. The second multi-head attention mechanism (the third multi-head attention mechanism module) in the decoder is a multi-head attention mechanism module that focuses on the final encoder output of the encoder.

[0085] The calculation process of the masked multi-head attention mechanism module is similar to that of the multi-head attention mechanism in the encoder, but when calculating the attention score , a mask matrix M mask (The upper triangular part elements of the mask matrix M mask are negative infinity, the lower triangular part of the mask matrix M mask is 0, and the diagonal elements of the mask matrix M mask are 0, and the dimension is the same as the attention score matrix) is added, so that when performing softmax normalization, the information weights of future positions become 0, that is:

[0086]

[0087] Subsequent concatenation, linear transformation, and residual connection and layer normalization operations are the same as those of the multi-head attention mechanism in the encoder, and the output of the masked multi-head attention mechanism is obtained

[0088] Furthermore, the second multi-head attention mechanism of the decoder is used to focus on the final encoder output H of the encoder E , and its calculation method is similar to that of the ordinary multi-head attention mechanism, and the output of the masked multi-head attention mechanism As the query vector Q, the output H of the encoder E is respectively used as the key vector K and the value vector V for calculation to obtain the output of the second multi-head attention mechanism and also undergoes residual connection and layer normalization operations

[0089] Furthermore, the output of the second multi-head attention mechanism is input into the feed-forward neural network in the decoder. The calculation process is the same as that of the feed-forward neural network in the encoder. After passing through the activation function and linear transformation, the output is obtained, and then through residual connection and layer normalization operations, the output of the final decoder layer is obtained

[0090] After being processed by Y decoder layers in sequence, the final decoder output representation H is obtained D .

[0091] Furthermore, to generate the initial response dialogue text, the dimension of the final decoder output H D needs to be converted to match the dimension of the vocabulary of the response dialogue text. In one embodiment, the vocabulary size is X v , and through a linear transformation layer (weight matrix W vocab , with dimension d model *X v ), H D is mapped to the vocabulary space to obtain the probability distribution P of each vocabulary v :

[0092] P v = softmax(H D W vocab );

[0093] Furthermore, from the probability distribution P of each vocabulary v , in combination with factors such as the entropy value of the probability distribution and the diversity of the historically generated vocabulary, a sampling method Sample is used to determine the vocabulary, and the vocabulary of the initial response dialogue text is generated one by one until the end-of-generation marker (such as <eos>Up to this point, the specific sampling process can be expressed as follows:

[0094] R initial = concat{Text i = Sample(P iv}}, i = 1, 2,..., L;

[0095] Among them, R initial represents the initial reply dialogue text, concat represents the concatenation operation, Text i represents the i-th generated word, P iv represents the probability distribution of the i-th word, and L represents the length of the initial reply dialogue text.

[0096] In the embodiment of the present invention, the target service model trained by the Transformer model based on the attention mechanism is used to predict the text semantic information. Since the multi-head attention mechanism can simultaneously focus on different semantic levels of the text, the target service model can understand the input text more comprehensively and deeply. When facing complex and ever-changing natural language expressions, it can capture key information more accurately. Therefore, it can accurately output the reply dialogue text, improve the fluency of the dialogue system, and increase the user's stickiness to the dialogue system.

[0097] In one embodiment, the descriptions of steps 401 to 404 are as follows:

[0098] Step 401: Perform lexical analysis on the initial reply dialogue text to obtain multiple words in the initial reply dialogue text, and determine the sentiment intensity of the sentiment words matched by each word in the pre-constructed sentiment dictionary, as well as the importance weight and semantic role weight of the sentiment words in the initial reply dialogue text.

[0099] Optionally, the intelligent dialogue device performs lexical analysis on the initial reply dialogue text R initial and divides the initial reply dialogue text R initial into multiple words r i . For each word r i , check whether there is a matching sentiment word in the sentiment dictionary. If found, record its corresponding sentiment intensity (according to the hierarchical intensity value in the dictionary) and the initial sentiment vector (obtained from the constructed sentiment word vector space). At the same time, determine the grammatical role and importance weight of the sentiment word in the sentence through syntactic analysis (for example, the sentiment word in the subject position has a higher weight than that in the object position, and the weight value with a range of 0 to 1 can be set according to syntactic rules), as well as the semantic role weight of the sentence where it is located (For example, as the agent of the core action, the semantic role weight of the emotional word is relatively high, and the value range is set from 0 to 1).

[0100] Step 402: Normalize the user emotional information to obtain the normalized user emotional information, and determine the emotional adjustment coefficient according to the normalized user emotional information.

[0101] Furthermore, the intelligent dialogue device normalizes the user emotional information E (the value range may vary depending on the calculation method, for example, from -5 to 5 represents the degree of emotional tendency from negative to positive) so that its value range is between 0 and 1, and obtains the normalized user emotional information E norm , where the specific normalization formula is as follows:

[0102]

[0103] Furthermore, the intelligent dialogue device calculates the emotional adjustment coefficient α according to the normalized user emotional information E norm . In the embodiment of the present invention, in order to achieve more refined emotional adjustment, the adjustment coefficient is divided into multiple levels. For example, when E E <0.2, α norm = 0.1 (corresponding to mild emotional adjustment, mainly for situations close to neutral emotions); when 0.2 ≤ E E ≤ 0.5, α norm = 0.3 (moderate emotional adjustment); when E E > 0.5, α norm = 0.5 (high emotional adjustment, where the coefficient level can be adjusted and optimized according to the actual business scenario and the demand for emotional adjustment sensitivity. E

[0104] Step 403: Determine the target emotional tendency according to the positivity and negativity of the user emotional information, and determine the target emotional vector according to the target emotional tendency.

[0105] Furthermore, the intelligent dialogue device determines the target emotional tendency according to the positivity and negativity of the user emotional information E. If the user emotional information E > 0, it means that the user has a positive emotional tendency, then the target emotional vector should be selected from the set of positive emotional word vectors. The specific selection method can be to select the central vector corresponding to the intensity level of the user emotional information E according to the emotional intensity level (for example, for moderately positive emotions, select the average vector of the set of moderately positive emotional word vectors as the target emotional vector ). If the user emotional information E < 0, then select the corresponding target emotional vector from the set of negative emotional word vectors The selection method is similar. If the user emotional information E = 0 (close to neutral emotion), then the target emotional vector remains unchanged (that is ) It mainly fine-tunes the response text through syntax and semantic role weights, rather than a large-scale emotional change.

[0106] Furthermore, considering that the user's emotion may not be absolutely positive or negative, but on a continuous spectrum, the target emotion vector is dynamically adjusted. For example, when the user's emotion information E is between two emotion intensity levels, the target emotion vector is determined by linear interpolation In one embodiment, when the user's emotion information E is between emotion intensity level L 1 and emotion intensity level L 2 , the corresponding target emotion vectors are respectively and Then the target emotion vector can be expressed as:

[0107]

[0108] Therefore, the target emotion vector can more accurately match the degree of the user's emotional tendency.

[0109] Step 404: Perform emotion adjustment on each word according to the emotion adjustment coefficient, the target emotion vector, and the importance weight and semantic role weight of the emotion word of each word, obtain the adjusted word vector, and generate an intermediate response dialogue text that matches the emotional tendency of the target user according to the adjusted word vector.

[0110] Furthermore, for each word r initial in the initial response dialogue text R i , if it is determined to be an emotion word (i.e., exists in the emotion dictionary), according to the above-determined emotion adjustment coefficient α E and the target emotion vector calculate the adjusted word vector The specific calculation formula is:

[0111]

[0112] Among them, multiplying by the syntax weight and semantic role weight in the above formula is to make a differential adjustment according to the importance and semantic role of the word in the sentence when adjusting the word vector, so that the emotion adjustment is more in line with the language logic and semantic expression.

[0113] Furthermore, after completing the adjustment of the word vectors of all emotion words, recombine the adjusted word vectors and convert them into an intermediate response dialogue text R through a language generation model (such as a neural network-based language generator, which can convert a sequence of word vectors into natural language text after being trained with a large amount of corpus) intermediate 。During the generation process, the language generation model takes into account the grammatical structure, semantic coherence, and context information of the sentence to ensure that the adjusted text is natural and fluent in language expression, while accurately reflecting the emotional adjustment effect based on the user's emotional information.

[0114] The embodiments of the present invention more precisely perform emotional adjustment on the initial reply dialogue text based on the user's emotional information, making it more in line with the user's emotional expectations and communication scenarios, improving the interaction quality and user experience of the dialogue system, and increasing the user's stickiness to the dialogue system.

[0115] In one embodiment, the descriptions of steps 501 to 503 are as follows:

[0116] Step 501, extract the key elements in the current dialogue scenario, as well as the personality traits and interest preferences in the user profile.

[0117] Optionally, the intelligent dialogue device extracts the key elements in the current dialogue scenario. Among them, the key elements include the dialogue round circle(t), the dialogue topic relevance score sorce(c), and the dialogue time time. The dialogue round circle(t) is represented by a number indicating which round of the dialogue; the dialogue topic relevance score sorce(c) determines the degree of closeness of the current dialogue topic to the historical dialogue topic based on a preset topic model, with a value ranging from 0 to 1. Among them, the algorithm in the preset topic model is the cosine similarity algorithm. If the current dialogue topic is highly relevant to the historical dialogue topic (sorce(c) = 0.8), then more relevant details discussed before can be cited in the reply; if the relevance is low (sorce(c) = 0.3), then the reply should focus more on independently answering the current dialogue topic; the dialogue time time can be converted into time characteristics related to the business, such as whether it is during the business peak period.

[0118] Furthermore, the intelligent dialogue device extracts the personality traits and interest preferences in the user profile. The personality traits such as the extroversion degree extro, with a value ranging from 0 to 1; the decision-making firmness degree dec, with a value ranging from 0 to 1. The interest preferences can be represented by an interest vector indicating the preference degree of each dimension for different interest fields, with a value ranging from 0 to 1.

[0119] Step 502, construct a context feature vector based on the dialogue round, dialogue topic relevance score, and dialogue time, construct a personalized weight vector based on the extroversion degree, decision-making firmness degree, and interest preference degree, and convert the intermediate reply dialogue text into a corresponding text vector representation.

[0120] Furthermore, the intelligent dialogue device constructs a context feature vector according to the dialogue round, dialogue topic relevance score, and dialogue time It should be noted that in practical applications, the dimension of the context feature vector is not limited to the above-mentioned number of dialogue turns, dialogue topic relevance score, and dialogue time. Further, the intelligent dialogue device respectively determines the component weights of the extroversion degree, decision-making firmness degree, and interest preference according to business requirements, and constructs a personalized weight vector according to the component weights of the extroversion degree, decision-making firmness degree, and interest preference Further, the intelligent dialogue device converts the intermediate reply dialogue text R intermediate into the corresponding text vector representation, similar to the above text vector conversion method

[0121] Step 503, perform fusion based on the context feature vector, personalized weight vector, and text vector representation to obtain the fused text vector representation, and convert the fused text vector representation into the form of natural language text to obtain the target reply dialogue text

[0122] Further, the intelligent dialogue device performs fusion according to the context feature vector, personalized weight vector, and text vector representation to integrate the context and user portrait features into the text to obtain the fused text vector representation Among them, the fused text vector representation can be expressed as:

[0123]

[0124] Among them, · represents the vector dot product operation, represents the vector concatenation operation

[0125] Further, the intelligent dialogue device converts the fused text vector representation into the form of natural language text to obtain the target reply dialogue text

[0126] In the embodiment of the present invention, the current scenario and user portrait features are integrated into the intermediate reply dialogue text for personalized optimization to obtain the target reply dialogue text that finally meets the user's needs, so as to better meet the dialogue requirements of various different scenarios and topics, make the dialogue text more in line with the user's emotional expectations and communication scenarios, improve the interaction quality and user experience of the dialogue system, and increase the user's stickiness to the dialogue system

[0127] In one embodiment, the descriptions of steps 60 to 90 are as follows

[0128] Step 60, input multiple initial sample training data pairs into the data screening model, and obtain the loss function value of each initial sample training data pair output by the data screening model

[0129] Specifically, the intelligent dialogue device obtains a plurality of initial sample training data pairs, where each initial sample training data pair includes a sample dialogue text and its corresponding reply dialogue text.

[0130] Further, the intelligent dialogue device inputs the plurality of initial sample training data pairs into a pre-trained data screening model to obtain the first loss function value of each initial sample training data pair output by the data screening model. In one embodiment, the label sequence probability distribution of the initial sample training data pair output by the data screening model is {p 1 , p 2 ,..., p m}, where m represents the number of initial sample training data pairs, p i represents the label sequence probability of the i-th initial sample training data pair, and the true label sequence of the initial sample training data pair is {l 1 , l 2 ,..., l m}, l i represents the true label corresponding to the reply dialogue text of the i-th initial sample training data pair. The loss function of the data screening model is:

[0131] L i =-log(p i [l i );

[0132] where represents the first loss function value of the i-th initial sample training data pair, and p i [l i represents the probability value corresponding to the true label l i in the predicted probability distribution p i .

[0133] Step 70, traverse the first loss function value of each initial sample training data pair, and determine the target sample training data pair for the initial sample training data pair whose first loss function value is less than the first preset loss threshold.

[0134] Further, the intelligent dialogue device traverses the first loss function value of each initial sample training data pair, eliminates the initial sample training data pairs whose first loss function value is greater than or equal to the first preset loss threshold, and retains the initial sample training data pairs whose first loss function value is less than the first preset loss threshold to obtain the target sample training data pairs of the initial sample training data pairs. The first preset loss threshold is set according to the actual situation, and the first preset loss threshold is, for example, 0.05, 0.08, etc.

[0135] Step 80, extract the text semantics of the sample dialogue text in each target sample training data pair, and label the result label of the reply dialogue text in each target sample training data pair.

[0136] Furthermore, the intelligent dialogue device extracts the text semantics of the sample dialogue text in each target sample training data pair. The specific process is as follows: For the sample dialogue text x in each target sample training data pair i , perform lexical analysis and word vector mapping through the CW2Vec word vector model, and map each word w ij (j represents the j-th word in the sample dialogue text x i ) to a d-dimensional vector Then, through the semantic combination function f sem construct the text semantic representation of the entire sample dialogue text x i Among them, the text semantic representation is:

[0137]

[0138] Among them, n represents the number of words in the sample dialogue text x i , and the semantic combination function fs sem can be a function based on the weighted sum of the syntax tree structure and word vectors. Therefore, the text semantic representation is:

[0139]

[0140] Among them, α ij is the weight calculated according to the depth d ij of each word w i in the syntax tree of the sample dialogue text x ij and the part-of-speech importance p ij of each word w ij . The specific calculation formula is as follows.

[0141]

[0142] Furthermore, the intelligent dialogue device labels the result tags of the reply dialogue text in each target sample training data pair. The specific process is as follows: Convert the reply dialogue text reply i corresponding to the sample dialogue text x in each target sample training data pair i into a label sequence, construct a vocabulary V containing all possible reply words, and map each word w i in the reply dialogue text reply i ′ k to the index k in the vocabulary V to obtain the label sequence {l i ,l 1 ,...,l 2 of the reply dialogue text reply​ m}, where m is the number of reply texts in the reply dialogue i .

[0143] Step 90: Based on the text semantics in each target sample training data pair and the result labels of the reply dialogue texts, train the Transformer model to obtain a target business model.

[0144] Furthermore, the intelligent dialogue device trains the Transformer model based on the text semantics in each target sample training data pair and the result labels of the reply dialogue texts to obtain a target business model. The specific training process is as follows: Set the constraints for the Transformer model training process as: the loss function value of each target sample training data pair is less than or equal to a second preset loss threshold, and the difference between the loss function values of two adjacent second loss function values is less than or equal to a preset difference threshold, where the second preset loss threshold and the preset difference threshold can both be set according to the actual situation.

[0145] Therefore, the intelligent dialogue device inputs the text semantics in each target sample training data pair and the result labels of the reply dialogue texts into the Transformer model. The Transformer model makes predictions based on the text semantics in each target sample training data pair to obtain the prediction results of each target sample training data pair.

[0146] Furthermore, the Transformer model calculates the second loss function value of each target sample training data pair according to the prediction results of each target sample training data pair and the result labels of the reply dialogue texts through the loss function in the Transformer model. Therefore, the intelligent dialogue device can obtain the second loss function value of each target sample training data pair output by the Transformer model. The loss function of the Transformer model is:

[0147]

[0148] where loss i (Transformer) represents the second loss function value of the i-th target sample training data pair, n represents the number of target sample training data pairs, y i represents the result label of the i-th target sample training data pair, represents the sample prediction value of the i-th target sample training data pair.

[0149] Further, the intelligent dialogue device compares the second loss function value of each target sample training data pair with a second preset loss threshold to obtain a size comparison result, and calculates the difference between the loss function values of two adjacent target sample training data pairs.

[0150] Further, if it is determined that the constraint condition is not satisfied according to the second loss function value of each target sample training data pair, or / and, the difference between the loss function values of two adjacent target sample training data pairs, that is, the second loss function value of at least one target sample training data pair is greater than the second preset loss threshold, or / and, the difference between the loss function values of at least one pair of adjacent target sample training data pairs is greater than the preset difference threshold, the intelligent dialogue device updates the parameters in the target optimization function of the Transformer model, where the target optimization function of the Transformer model is:

[0151]

[0152] where θ t+1 represents the parameters of the model at the t+1 step, and θ t represents the parameters of the model at the t step; η t represents the learning rate of the model at the t step, and as the number of steps t increases, the learning rate η t decreases; μ represents the preset momentum parameter, and g t represents the gradient of the model at the t step, η 0 represents the initial learning rate, and β represents the preset decay coefficient.

[0153] Further, the intelligent dialogue device calculates the second loss function value of each target sample training data pair according to the adjusted Transformer model until the second loss function value of each target sample training data pair and the difference between two adjacent second loss function values both satisfy the constraint condition, that is, the loss function value of each target sample training data pair is less than or equal to the second preset loss threshold, and the difference between two adjacent second loss function values is less than or equal to the preset difference threshold, to obtain the target service model.

[0154] In the embodiment of the present invention, the target service model is obtained by training the Transformer model of the attention mechanism. Therefore, the text semantic information can be predicted by the target service model subsequently. Since the multi-head attention mechanism can simultaneously focus on different semantic levels of the text, the target service model can understand the input text more comprehensively and deeply, and can capture key information more accurately when facing complex and changeable natural language expressions. Therefore, it can accurately output the reply dialogue text, improve the fluency of the dialogue system, and increase the user's stickiness to the dialogue system.

[0155] Furthermore, the intelligent text dialogue generation device based on artificial intelligence provided by the present invention will be described below. The intelligent text dialogue generation device based on artificial intelligence described below can be correspondingly referred to the intelligent text dialogue generation method based on artificial intelligence described above.

[0156] Optionally, referring to Figure 2 , Figure 2 is a schematic structural diagram of the intelligent text dialogue generation device based on artificial intelligence provided by the present invention. The intelligent text dialogue generation device based on artificial intelligence includes.

[0157] A semantic parsing module 210, configured to perform text semantic parsing on the question dialogue text input by the target user to obtain the text semantic information, user emotion information, and business domain information of the question dialogue text;

[0158] A model matching module 220, configured to perform matching in a preset business model library based on the business domain information to obtain a target business model; the target business model is obtained by training a Transformer model based on attention mechanism with sample text semantics and the label results of the corresponding reply dialogue text;

[0159] A model prediction module 230, configured to input the text semantic information into the target business model to obtain an initial reply dialogue text output by the target business model;

[0160] A text adjustment module 240, configured to perform emotion adjustment on the initial reply dialogue text based on the user emotion information to obtain an intermediate reply dialogue text that matches the emotion tendency of the target user;

[0161] A text optimization and feedback module 250, configured to perform personalized optimization on the intermediate reply dialogue text based on the current dialogue context and the user profile of the target user to obtain a target reply dialogue text and feedback it to the target user; the user profile is constructed by analyzing the personality characteristics and interest preferences based on the historical information of the target user.

[0162] In the embodiments of the present invention, a target business model trained by a Transformer model based on an attention mechanism is used to predict the text semantic information. Since the multi-head attention mechanism can simultaneously focus on different semantic levels of the text, the target business model can understand the input text more comprehensively and deeply. When facing complex and changeable natural language expressions, it can capture key information more accurately. Therefore, it can accurately output the reply dialogue text, improving the fluency of the dialogue system. On the other hand, by matching the business domain information in the user's input text to obtain the corresponding target business model, and personalizing and optimizing the initial reply dialogue text according to the user's emotional information, the current dialogue context, and the user profile of the user, the final target reply dialogue text that meets the user's needs is obtained, so as to better respond to the dialogue needs of various different scenarios and topics, no longer limited to retrieving a specific corpus range, improving the flexibility and versatility of the dialogue system. Therefore, the stickiness of users to the dialogue system can be increased.

[0163] Please refer to Figure 3 , Figure 3 which is the embodiment diagram of the electronic device provided by the embodiments of the present invention. As Figure 3 shown, the embodiments of the present invention provide an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor 320. When the processor 320 executes the computer program 311, the following steps are implemented:

[0164] Perform text semantic parsing on the question dialogue text input by the target user to obtain the text semantic information, user emotional information, and business domain information of the question dialogue text;

[0165] Match based on the business domain information in a preset business model library to obtain a target business model; the target business model is trained based on a Transformer model based on an attention mechanism using the sample text semantics and the label results of their corresponding reply dialogue texts;

[0166] Input the text semantic information into the target business model to obtain the initial reply dialogue text output by the target business model;

[0167] Perform emotional adjustment on the initial reply dialogue text based on the user emotional information to obtain an intermediate reply dialogue text that matches the emotional tendency of the target user;

[0168] Perform personalized optimization on the intermediate reply dialogue text based on the current dialogue context and the user profile of the target user to obtain the target reply dialogue text and feedback it to the target user; the user profile is constructed by analyzing the personality characteristics and interest preferences based on the historical information of the target user.

[0169] Please refer to Figure 4 , Figure 4 This is an embodiment diagram of the computer-readable storage medium provided by the embodiments of the present invention. As Figure 4 shown, this embodiment provides a computer-readable storage medium 400, on which a computer program 311 is stored. When the computer program 311 is executed by a processor, the following steps are implemented:

[0170] Perform text semantic parsing on the problem dialogue text input by the target user to obtain the text semantic information, user emotion information, and business domain information of the problem dialogue text;

[0171] Match based on the business domain information in a preset business model library to obtain a target business model; the target business model is trained based on the sample text semantics and the label results of the corresponding reply dialogue text for a Transformer model based on the attention mechanism;

[0172] Input the text semantic information into the target business model to obtain an initial reply dialogue text output by the target business model;

[0173] Perform emotion adjustment on the initial reply dialogue text based on the user emotion information to obtain an intermediate reply dialogue text that matches the emotion tendency of the target user;

[0174] Perform personalized optimization on the intermediate reply dialogue text based on the current dialogue context and the user profile of the target user to obtain a target reply dialogue text and feedback it to the target user; the user profile is constructed by analyzing the personality characteristics and interest preferences based on the historical information of the target user.

[0175] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the intelligent text dialogue generation method based on artificial intelligence provided by the above methods. The method includes:

[0176] Perform text semantic parsing on the problem dialogue text input by the target user to obtain the text semantic information, user emotion information, and business domain information of the problem dialogue text;

[0177] Match based on the business domain information in a preset business model library to obtain a target business model; the target business model is trained based on the sample text semantics and the label results of the corresponding reply dialogue text for a Transformer model based on the attention mechanism;

[0178] Input the text semantic information into the target business model to obtain an initial reply dialogue text output by the target business model;

[0179] Perform sentiment adjustment on the initial reply dialogue text based on the user's sentiment information to obtain an intermediate reply dialogue text that matches the sentiment tendency of the target user;

[0180] Perform personalized optimization on the intermediate reply dialogue text based on the current dialogue context and the user profile of the target user to obtain the target reply dialogue text and feedback it to the target user; the user profile is constructed by analyzing the personality characteristics and interest preferences based on the historical information of the target user.

[0181] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0182] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0183] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.< / eos>

Claims

1. An intelligent text dialogue generation method based on artificial intelligence, characterized in that: include: Performing text semantic analysis on the question dialogue text input by the target user to obtain text semantic information, user emotion information and business domain information of the question dialogue text; Based on the business domain information, a match is performed in a preset business model library to obtain a target business model; the target business model is obtained by training a Transformer model based on an attention mechanism based on the label results of the sample text semantics and its corresponding reply dialogue text; Inputting the text semantic information into the target business model to obtain the initial reply dialogue text output by the target business model; Performing emotion adjustment on the initial reply dialogue text based on the user emotion information to obtain an intermediate reply dialogue text that matches the emotional tendency of the target user; Based on the current conversation context and the user portrait of the target user, the intermediate reply conversation text is personalized and optimized to obtain the target reply conversation text and feed it back to the target user; The user portrait is constructed based on analyzing the personality characteristics and interest preferences of the target user based on historical information.

2. The method for generating intelligent text dialogue based on artificial intelligence according to claim 1, characterized in that: The Transformer model includes an encoder and a decoder; the encoder includes a plurality of encoder layers, each of which includes a first multi-head attention mechanism module, a feedforward neural network module, and a residual connection and layer normalization operation module; the decoder includes a plurality of decoder layers, each of which includes a second multi-head attention mechanism module, a feedforward neural network module, a residual connection and layer normalization operation module, and a third multi-head attention mechanism module; The step of inputting the text semantic information into the target business model to obtain the initial reply dialogue text output by the target business model includes: Preprocessing the text semantic information, obtaining a position encoding vector corresponding to a text vector at each position in the text semantic information, and adding the text vector at each position in the text semantic information and its corresponding position encoding vector element by element to obtain the position-encoded text semantic information; Inputting the position-encoded text semantic information into the target business model, each encoder layer in the encoder performs feature extraction and encoding conversion on the position-encoded text semantic information through the first head attention mechanism module, the feedforward neural network module, and the residual connection and layer normalization operation module to generate a final encoder output; The final encoder output is input to the decoder, and each decoder layer in the decoder generates the initial reply dialogue text by combining the position-encoded text semantic information and the final encoder output through the second multi-head attention mechanism module, the feedforward neural network module, the residual connection and layer normalization operation module and the third multi-head attention mechanism module.

3. The method for generating intelligent text dialogue based on artificial intelligence according to claim 2, characterized in that: The encoder layer process is as follows: The position-encoded text semantic information is linearly transformed through the first multi-head attention mechanism module to obtain a query vector, a key vector and a value vector; the position-encoded text semantic information is mapped and concatenated in combination with the query vector, the key vector, the value vector, the position-encoded text semantic information and the dimension of the first multi-head attention mechanism module to obtain a first multi-head attention mechanism output; the first multi-head attention mechanism output is linearly transformed based on the feedforward neural network module to obtain a feedforward neural network output; the position-encoded text semantic information, the first multi-head attention mechanism output and the feedforward neural network output are layer-normalized based on the residual connection and layer normalization operation module to generate the final encoder output; The decoder layer process is as follows: Based on the second multi-head attention mechanism module, the final encoder output of the mask matrix is ​​linearly transformed, mapped and concatenated to obtain the masked multi-head attention mechanism output; based on the third multi-head attention mechanism module, the masked multi-head attention mechanism output is linearly transformed, mapped and concatenated to obtain the second multi-head attention mechanism output; based on the feedforward neural network module and the residual connection and layer normalization operation module, the second multi-head attention mechanism output is linearly transformed and layer normalized to obtain the final decoder output; the final decoder output is mapped to the vocabulary space to obtain the probability distribution of each word, and the initial reply dialogue text is generated according to the probability distribution of each word.

4. The method for generating intelligent text dialogue based on artificial intelligence according to claim 2, characterized in that: The training steps of the target business model include: Inputting a plurality of initial sample training data pairs into a data screening model, and obtaining a loss function value of each of the initial sample training data pairs output by the data screening model; each of the initial sample training data pairs includes a sample dialogue text and its reply dialogue text; Traversing the first loss function value of each of the initial sample training data pairs, and determining the initial sample training data pair whose first loss function value is less than a first preset loss threshold as the target sample training data pair; Extracting the text semantics of the sample dialogue text in each of the target sample training data pairs, and marking the result label of the reply dialogue text in each of the target sample training data pairs; Based on the text semantics and the result label of the reply dialogue text in each of the target sample training data pairs, the Transformer model is trained to obtain the target business model.

5. The method for generating intelligent text dialogue based on artificial intelligence according to claim 4, characterized in that: The constraint conditions for the training process of the Transformer model are: the loss function value of each target sample training data pair is less than or equal to the second preset loss threshold, and the difference between two adjacent second loss function values ​​is less than or equal to the preset difference threshold; the Transformer model is trained based on the text semantics in each target sample training data pair and the result label of the reply dialogue text to obtain the target business model, including: Inputting the text semantics and the result label of the reply dialogue text in each of the target sample training data pairs into the Transformer model, and obtaining the second loss function value of each of the target sample training data pairs output by the Transformer model; If it is determined that the constraint condition is not satisfied according to the second loss function value of each target sample training data pair, or / and the difference between two adjacent second loss function values, then the parameters in the target optimization function of the Transformer model are updated until the second loss function value of each target sample training data pair outputted by the adjusted Transformer model and the difference between two adjacent second loss function values ​​satisfy the constraint condition, thereby obtaining the target business model; Among them, the loss function of the Transformer model is: Among them, loss i (Transformer) represents the second loss function value of the i-th target sample training data pair, n represents the number of target sample training data pairs, t i represents the result label of the i-th target sample training data pair, Represents the sample prediction value of the i-th target sample training data pair; The objective optimization function of the Transformer model is: Among them, θ t+1 represents the parameters of the model at step t+1, θ t represents the parameters of the model at the tth step; η t Represents the learning rate of the model at step t, and as the number of steps t increases, the learning rate η t Decreasing; μ represents the preset momentum parameter, g t represents the gradient of the model at the tth step, η0 represents the initial learning rate, and β represents the preset attenuation coefficient.

6. The method for generating intelligent text dialogue based on artificial intelligence according to claim 1, characterized in that: The step of performing emotion adjustment on the initial reply dialogue text based on the user emotion information to obtain an intermediate reply dialogue text that matches the emotional tendency of the target user includes: Performing lexical analysis on the initial reply dialogue text to obtain a plurality of words in the initial reply dialogue text, and determining the sentiment intensity of the sentiment word matched to each word in a pre-built sentiment dictionary, as well as the importance weight and semantic role weight of the sentiment word in the initial reply dialogue text; Normalizing the user emotion information to obtain normalized user emotion information, and determining an emotion adjustment coefficient according to the normalized user emotion information; Determining a target emotional tendency according to the positivity or negativity of the user emotional information, and determining a target emotional vector according to the target emotional tendency; Perform sentiment adjustment on each word according to the sentiment adjustment coefficient, the target sentiment vector, and the importance weight and semantic role weight of the sentiment word of each word to obtain an adjusted word vector, and generate an intermediate reply dialogue text that matches the sentiment tendency of the target user according to the adjusted word vector.

7. The method for generating intelligent text dialogue based on artificial intelligence according to claim 1, characterized in that: The step of performing personalized optimization on the intermediate reply dialogue text based on the current dialogue context and the user portrait of the target user to obtain the target reply dialogue text includes: Extracting key elements in the current conversation context, as well as personality traits and interest preferences in the user portrait; the key elements include conversation rounds, conversation topic relevance scores, and conversation time; the personality traits include degree of extroversion and degree of decisiveness in decision-making; constructing a context feature vector based on the conversation turn, the conversation topic relevance score and the conversation time, constructing a personalized weight vector based on the degree of extroversion, the degree of decisiveness in decision making and the interest preference, and converting the intermediate reply conversation text into a corresponding text vector representation; Based on the situational feature vector, the personalized weight vector and the text vector representation, a fusion is performed to obtain a fused text vector representation, and the fused text vector representation is converted into a natural language text form to obtain the target reply dialogue text.

8. An intelligent text dialogue generation device based on artificial intelligence, characterized in that: Applicable to the intelligent text dialogue generation method based on artificial intelligence as described in any one of claims 1 to 7; The intelligent text dialogue generation device based on artificial intelligence includes: A semantic analysis module is used to perform text semantic analysis on the question dialogue text input by the target user to obtain text semantic information, user emotion information and business domain information of the question dialogue text; A model matching module is used to match the business domain information in a preset business model library to obtain a target business model; the target business model is obtained by training a Transformer model based on an attention mechanism based on the label results of the sample text semantics and its corresponding reply dialogue text; A model prediction module, used for inputting the text semantic information into the target business model to obtain the initial reply dialogue text output by the target business model; A text adjustment module, used to perform emotion adjustment on the initial reply dialogue text based on the user emotion information to obtain an intermediate reply dialogue text that matches the emotional tendency of the target user; A text optimization feedback module is used to personalize and optimize the intermediate reply dialogue text based on the current dialogue context and the user portrait of the target user, obtain the target reply dialogue text and feed it back to the target user; the user portrait is constructed based on the historical information analysis of the target user's personality characteristics and interest preferences.

9. An electronic device, comprising: Memory for storing computer software programs; A processor, used to read and execute a computer software program, wherein when the computer software program is executed by the processor, the intelligent text dialogue generation method based on artificial intelligence as described in any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer software program stored therein, characterized in that: When the computer software program is executed by a processor, the intelligent text dialogue generation method based on artificial intelligence as claimed in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Generative dialogue system based on multi-round emotion analysis

    CN112163080A

  • Intelligent reply method and device

    CN118607529A

  • AI online education intelligent question and answer information processing method

    CN119149710A

  • Real-time voice stream and text dialogue interaction system based on large language model

    CN119539076A

  • Human-machine dialogue system and method

    US20230395075A1

Cited By

  • Intelligent collaborative management method and platform based on artificial intelligence driving

    CN120374059A

  • Conversation processing method and system based on personality characteristics, intelligent agent and electronic equipment

    CN121413633A