Human-computer dialogue method, device, electronic device and storage medium
The target personality dimensions and attribute information are determined through the pre-constructed personality portrait and related reply statements are generated, which solves the problem that human-computer dialogue systems in the prior art are difficult to effectively store and utilize user multi-dimensional personality attribute information, and improves the interactive experience and system efficiency.
Patent Information
- Application Number
- CN202110413191.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-04-16
AI Technical Summary
In the prior art, it is difficult for human-computer dialogue systems to effectively store and utilize the user's multi-dimensional personality attribute information, resulting in poor interactive experience and occupying a large amount of storage space.
Through the pre-constructed personality portrait, determine the target personality dimensions and attribute information corresponding to the dialogue input, generate relevant reply statements, reduce the need for natural language processing, and reduce the storage space occupation.
The interactive experience of human-computer dialogue interaction is improved. By directly utilizing the information in the pre-constructed personal image, the storage of invalid information is reduced, and the efficiency of the system and user satisfaction are improved.
Smart Images

Figure CN112948565B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of natural language processing technologies, and in particular, to a human-computer dialogue method, apparatus, electronic device, and storage medium. Background Art
[0002] With the development of artificial intelligence technologies, human-computer dialogue systems have gradually become a new generation of interaction modes in the future due to their inherent natural convenience. In the application of human-computer dialogue, interactions are usually carried out through intelligent assistants. People have also gradually developed a need for emotional companionship from intelligent assistants, hoping that intelligent assistants can accompany and understand them for a long time like a person or a friend. In view of this, the industry has gradually posed new challenges to the research on AI technologies of intelligent assistants. Summary of the Invention
[0003] To overcome the problems existing in the related technologies, the present disclosure provides a human-computer dialogue method, apparatus, electronic device, and storage medium.
[0004] According to a first aspect of an embodiment of the present disclosure, there is provided a human-computer dialogue method, including:
[0005] Obtaining a dialogue input;
[0006] Determining a target persona dimension corresponding to the dialogue input in a pre-constructed persona profile, where the persona profile includes multiple persona dimensions and keywords corresponding to each persona dimension;
[0007] Determining target persona attribute information corresponding to the dialogue input according to the semantic relationship between the target keyword corresponding to the target persona dimension and the dialogue input;
[0008] Generating a target reply statement according to the target persona attribute information.
[0009] In some embodiments, the determining target persona attribute information corresponding to the dialogue input according to the semantic relationship between the target keyword corresponding to the target persona dimension and the dialogue input includes:
[0010] Inputting the dialogue input into a trained language processing model, where the language processing model is used to predict, for each target keyword, the persona attribute information corresponding to the target keyword in the dialogue input, and use the persona attribute information as the target persona attribute information corresponding to the dialogue input.
[0011] In some embodiments, the language processing model includes a preprocessing layer, an encoding layer and a classification layer for the persona dimension classification scenario, and an extraction layer for the persona attribute information extraction scenario. The language processing model is trained in the following manner:
[0012] Obtain a plurality of training samples, where the plurality of training samples include classification training samples collected in the scenario of the personal dimension classification and personal attribute information training samples collected in the scenario of the extraction of personal attribute information, and each training sample in the plurality of training samples includes user input text and a corresponding annotation label;
[0013] Input each of the training samples into a preprocessing layer to obtain a character sequence corresponding to the user input text in the training sample;
[0014] When the training sample belongs to the classification training sample, input the character sequence of the training sample into the encoding layer to obtain a semantic vector corresponding to each character, and input the average vector of the semantic vectors of all characters into the classification layer. Based on the classification result output by the classification layer and the annotation label in the training sample, determine the first prediction loss corresponding to the training sample;
[0015] When the training sample belongs to the personal attribute information training sample, input the character sequence of the training sample into the extraction layer, and based on the extraction result output by the extraction layer and the annotation label in the training sample, determine the second prediction loss corresponding to the training sample;
[0016] Adjust the parameters of the language processing model based on the sum of the prediction losses corresponding to the plurality of training samples.
[0017] In some embodiments, the generating the target response statement according to the target personal attribute information includes:
[0018] Determine a response template statement according to the target personal attribute information, where the response template statement is a statement including slots to be filled, and each of the slots to be filled carries a keyword identifier;
[0019] Determine slot information corresponding to each keyword identifier in the slots to be filled in the response template statement according to each keyword identifier;
[0020] Fill the slot information into the slot to be filled in the response template statement corresponding to the slot information according to the semantic information of the keyword identifier and the slot information to generate a target response statement.
[0021] In some embodiments, the determining the response template statement according to the target personal attribute information includes:
[0022] In the case that it is determined that there is no historical personal attribute information corresponding to the target personal dimension in the storage module, use any one of the first type of preset template statements configured in the template library as the response template statement for the dialogue input;
[0023] Determining the slot information corresponding to each of the keyword identifiers of the to-be-filled slots in the reply template statement includes:
[0024] Determining, from the target persona attribute information, the slot information corresponding to each of the keyword identifiers of the to-be-filled slots in the reply template statement.
[0025] In some embodiments, determining the reply template statement according to the target persona attribute information includes:
[0026] When it is determined that there is historical persona attribute information in the storage module that is semantically related to the target persona attribute information but has inconsistent persona attribute information, any one of the second type of preset template statements configured in the template library is used as the reply template statement for the dialogue input.
[0027] In some embodiments, determining the reply template statement according to the target persona attribute information includes:
[0028] When it is determined that there is historical persona attribute information in the storage module that is semantically related to the target persona attribute information, identify the dialogue intention of the dialogue input;
[0029] When there is no attribute information in the target persona attribute information and the historical persona attribute information that satisfies the dialogue intention, any one of the third type of preset template statements configured in the template library is used as the reply template statement for the dialogue input;
[0030] Determining the slot information corresponding to each of the keyword identifiers of the to-be-filled slots in the reply template statement includes:
[0031] For each of the keyword identifiers of the to-be-filled slots in the reply template statement, infer the information corresponding to the keyword identifier according to the target persona attribute information and the historical persona attribute information corresponding to the target persona dimension, and use this information as the slot information corresponding to the keyword identifier.
[0032] In some embodiments, the method further includes:
[0033] When it is determined that there is no historical persona attribute information corresponding to the target persona dimension in the storage module, store the target persona attribute information in the storage module;
[0034] In some embodiments, the method further includes:
[0035] In the case of determining that there is historical persona attribute information in the storage module that is semantically related to the target persona attribute information and inconsistent with the persona attribute information, replace the historical persona attribute information in the storage module that is semantically related to the target persona attribute information and inconsistent with the persona attribute information with the target persona attribute information.
[0036] In some embodiments, determining the target persona dimension corresponding to the conversation input includes:
[0037] Input the conversation input into a trained language processing model, which is used to predict the target persona dimension among all the persona dimensions included in a pre-constructed persona profile to which the conversation input belongs.
[0038] According to a second aspect of the embodiments of the present disclosure, a human-machine dialogue device is provided, including:
[0039] An acquisition module, configured to acquire a conversation input;
[0040] A first determination module, configured to determine, in a pre-constructed persona profile, a target persona dimension corresponding to the conversation input, where the persona profile includes multiple persona dimensions and keywords corresponding to each persona dimension;
[0041] A second determination module, configured to determine target persona attribute information corresponding to the conversation input according to the semantic relationship between the target keyword corresponding to the target persona dimension and the conversation input;
[0042] A generation module, configured to generate a target response statement according to the target persona attribute information.
[0043] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including:
[0044] A processor;
[0045] A memory for storing instructions executable by the processor;
[0046] Wherein, the processor is configured to execute the steps of implementing the human-machine dialogue method provided in the first aspect of the present disclosure. According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of implementing the human-machine dialogue method provided in the first aspect of the present disclosure are realized.
[0047] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0048] Since the persona dimensions are directly recorded in the pre-built persona portrait, the target persona dimensions corresponding to the dialogue input can be obtained without processing the complete natural language recorded, reducing the occupation of storage space by invalid information; the target response statement is determined by using the target persona attribute information in the dialogue input. In this way, the generated target response statement can contain semantic information related to the target persona attribute information, thereby improving the interaction experience of the dialogue interaction.
[0049] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0051] Figure 1 is a schematic diagram of an application scenario of a human-machine dialogue method shown according to an exemplary embodiment of the present disclosure.
[0052] Figure 2 is a flowchart of a human-machine dialogue method shown according to an exemplary embodiment of the present disclosure.
[0053] Figure 3 is a flowchart of another human-machine dialogue method shown according to an exemplary embodiment of the present disclosure.
[0054] Figure 4 is a schematic structural diagram of a human-machine dialogue device shown according to an exemplary embodiment of the present disclosure.
[0055] Figure 5 is a block diagram of an electronic device shown according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0057] Before introducing the human-machine dialogue method provided by the present disclosure, first, the application scenarios involved in each embodiment of the present disclosure are described. The present disclosure can be applied to the process of realizing language interaction through a terminal, and the language interaction refers to the human-machine communication dialogue between a user and the terminal.
[0058] Figure 1FIG. 0 is a schematic diagram of an application scenario of a human-machine dialogue method according to an exemplary embodiment of the present disclosure, where a user has a dialogue with a terminal (through a smart assistant). Among them, the terminal may be a smart phone, a tablet computer, a smart watch, a personal computer, a laptop computer, a smart TV, a PDA (English: Personal Digital Assistant, Chinese: Personal Digital Assistant), or other terminals.
[0059] To improve the human-machine dialogue experience, users usually hope that when having a dialogue with the smart assistant in the terminal, various persona attributes in various persona dimensions of the user and the smart assistant can be stored in the terminal, for example, in the age dimension, the preference dimension, etc. Correspondingly, the persona attribute is the specific attribute value corresponding to each dimension. In this way, the interaction experience between the two parties can be improved through the persona attributes. Considering that in the related art, the persona dimension is usually described in complete natural language, there is a lot of invalid information, and thus a large amount of storage space is occupied.
[0060] In view of this, the present disclosure provides a human-machine dialogue method, device, electronic device, and storage medium. Through this method, based on a pre-constructed persona portrait, a target persona dimension is determined. Since the persona dimension is directly recorded in the pre-constructed persona portrait, the target persona dimension corresponding to the dialogue input can be obtained without processing the recorded complete natural language, reducing the occupation of storage space by invalid information; using the target persona attribute information in the dialogue input to determine the target response sentence. In this way, the generated target response sentence can contain semantic information related to the target persona attribute information, thereby improving the interaction experience of the dialogue interaction.
[0061] Figure 2 FIG. 10 is a flowchart of a human-machine dialogue method according to an exemplary embodiment of the present disclosure, as Figure 2 shown, the human-machine dialogue method is used in a terminal and includes the following steps.
[0062] In step S21, a dialogue input is obtained.
[0063] Exemplarily, the dialogue input may be voice information or text information. In the case where the dialogue input is voice information, the terminal can perform semantic recognition on the voice information, convert it into text information, and then execute the following steps according to the text information.
[0064] It should be noted that before step S21, the human-machine dialogue method further includes: receiving an acquisition instruction, and executing step S21 when the acquisition instruction is received. Exemplarily, the function of triggering the generation of the acquisition instruction can be realized by setting a preset button. Exemplarily, the function of triggering the generation of the acquisition instruction can also be realized by detecting a preset voice in the current environment. The present disclosure does not limit this.
[0065] In step S22, in the pre-constructed persona portrait, determine the target persona dimension corresponding to the dialogue input. The persona portrait includes multiple persona dimensions and keywords corresponding to each persona dimension.
[0066] Exemplarily, the persona dimensions include persona dimensions on the user side and persona dimensions on the intelligent assistant side. The persona dimensions on both the user side and the intelligent assistant side can include dimensions such as preferences, life trajectories, appearance, constellations, age, relatives, skills, etc. For example, taking the dialogue input "I am seven and a half years old this year and you are older than me" as an example, the target persona dimension of this dialogue input can be age. In a possible way, the target persona dimension corresponding to the dialogue input can be determined according to the semantic information of the dialogue input.
[0067] It should be noted that the keyword is a vocabulary that assists the terminal in determining the target persona attribute information corresponding to the target persona dimension in a complete sentence. For example, when the persona dimension is age, the corresponding keywords can be vocabulary related to numbers, years, attribute comparisons, etc.
[0068] In step S23, according to the semantic relationship between the target keyword corresponding to the target persona dimension and the dialogue input, determine the target persona attribute information corresponding to the dialogue input.
[0069] In the present disclosure, the semantic relationship refers to the semantic relationship between the target keyword and the characters in the dialogue input.
[0070] Exemplarily, still taking the above dialogue input "I am seven and a half years old this year and you are older than me" as an example, the corresponding target keywords can be numbers, years, etc. When the target keyword represents a number, the voice information related to the number represented in the dialogue input can be used as the target persona attribute information. Since "you are older than me" can represent a numerical relationship, "you are older than me" can be used as the target persona attribute information; when relying on the target keyword, when the target keyword represents a year, therefore, based on the target keyword "year", "this year" in the dialogue input can be the target persona attribute information corresponding to the dialogue input, and based on the target keyword "number", "seven and a half years old" can be the target persona attribute information corresponding to the dialogue input.
[0071] In step S24, generate a target reply sentence according to the target persona attribute information.
[0072] Through the above technical solution, since the persona dimensions are directly recorded in the pre-constructed persona portrait, the target persona dimension corresponding to the dialogue input can be obtained without processing the recorded complete natural language, reducing the occupation of storage space by invalid information; using the target persona attribute information in the dialogue input to determine the target reply sentence, thus, the generated target reply sentence can contain semantic information related to the target persona attribute information, thereby improving the interaction experience of the dialogue interaction.
[0073] It should be noted that in the constructed character portrait, each character dimension can also include historical character attribute information corresponding to the character dimension. Taking the historical dialogue input of "I am ten years old this year" as an example, "ten years old" and "this year" can both be used as historical character attribute information under the age dimension.
[0074] In a possible manner, the step of determining target character attribute information corresponding to the dialogue input according to the semantic relationship between the target keywords corresponding to the target character dimension and the dialogue input includes: inputting the dialogue into a trained language processing model to obtain the target character attribute information corresponding to the dialogue input.
[0075] It should be noted that the language processing model is used to predict the personality attribute information corresponding to each target keyword in the dialogue input, and use the personality attribute information as the target personality attribute information corresponding to the dialogue input.
[0076] The language processing model can essentially predict, for each target keyword, the position of the personality attribute information corresponding to the target keyword in the dialogue input, which includes the starting position and the ending position, and then intercept the personality attribute information corresponding to the target keyword in the dialogue input based on the starting position and the ending position. For example, still taking the above dialogue input "I am seven and a half years old this year and you are older than me" as an example, the dialogue input includes 10 characters each, and accordingly, the corresponding position of each character is: "I" corresponds to position 1, "today" corresponds to position 2, "year" corresponds to position 3, "seven" corresponds to position 4, "years" corresponds to position 5, "half" corresponds to position 6, "you" corresponds to position 7, "than" corresponds to position 8, "I" corresponds to position 9, and "older" corresponds to position 10. When the target keyword includes numbers, the corresponding predicted starting position is position 4. Combined with the context semantics, the corresponding ending position is position 6. Therefore, the personality attribute information corresponding to the target keyword is between position 4 and position 6 (including the starting position and the ending position). The text information extracted according to the position situation is "seven and a half years old", which means that "seven and a half years old" is the personality attribute information corresponding to the target keyword.
[0077] It should be noted that the language processing model that predicts the target persona attribute information can also predict that the dialogue input belongs to the target persona dimension among all the persona dimensions included in the pre-constructed persona portrait. For example, the language processing model can output the probability of each persona dimension among all the persona dimensions included in the pre-constructed persona portrait to which the dialogue input belongs. In this case, the persona dimension with the highest probability is used as the target persona temperature; in addition, the language processing model can directly output the target persona dimension to which the dialogue input belongs, which is not limited in this embodiment.
[0078] Exemplarily, the language processing model may include a preprocessing layer for preprocessing the input samples; an encoding layer and a classification layer for the character setting dimension classification scenario, and an extraction layer for the character setting attribute information extraction scenario. Among them, the language processing model can be trained in the following ways:
[0079] First, obtain a plurality of training samples.
[0080] It should be noted that the plurality of training samples include classification training samples collected in the character setting dimension classification scenario and character setting attribute information training samples collected in the character setting attribute information extraction scenario. The classification training samples are used for the training of the character setting dimension classification task, and the character setting attribute information training samples are used for the training of the character setting attribute information extraction task. Each training sample in the plurality of training samples includes user input text and a corresponding annotation label. Among them, the annotation label in the classification training sample is a character setting dimension label, and the annotation label of the character setting attribute information training sample is a character setting attribute information label.
[0081] Exemplarily, when the training sample belongs to the classification training sample, the training sample can be (query, class), where query is the user input text and class is the character setting dimension label corresponding to the training sample. When the training sample belongs to the character setting attribute information training sample, the training sample can be (query, {slot 1 ,slot 2 ,slot 3 ,slot 4 ……,slot M}), where query is the user input text, slot 1 ,slot 2 ,slot 3 ,slot 4 ……,slot M is the character setting attribute information label, where the predicted position corresponding to each character setting attribute information label can be span i =[start k ,end n , start k is the starting position, end n is the ending position, i, k, n, and M are natural positive integers, span i is the text segment between the starting position and the ending position, span i is slot i .
[0082] Second, input each training sample into the preprocessing layer to obtain a character sequence corresponding to the user input text in the training sample.
[0083] Exemplarily, the preprocessing layer is used to process the complete text input to the model into a character sequence. Still taking the above-mentioned dialogue input "I am seven and a half years old this year. You are older than me" as an example, the dialogue input can obtain the corresponding character sequence through the preprocessing layer, and this character sequence = {I, this, year, seven, years, half, you, older, than, me}.
[0084] Third, when the training sample belongs to a classification training sample, input the character sequence of the training sample into the encoding layer to obtain a semantic vector corresponding to each character, and input the average vector of the semantic vectors of all characters into the classification layer. Based on the classification result output by the classification layer and the labeled tag in the training sample, determine the first prediction loss corresponding to the training sample.
[0085] Exemplarily, the classification result output by the classification layer: y' = softmax(SW), S is the average vector of the semantic vectors of all characters, W is a model parameter, and Softmax is a function.
[0086] Fourth, when the training sample belongs to a persona attribute information training sample, input the character sequence of the training sample into the extraction layer. Based on the extraction result output by the extraction layer and the labeled tag in the training sample, determine the second prediction loss corresponding to the training sample.
[0087] Exemplarily, the second prediction loss includes the prediction loss at the start position and the prediction loss at the end position. start k = s' k = argmax(softmax(HWs)), end n = e' n = argmax(softmax(HWe)), where argmax is a function, H represents the character sequence of the training sample, and Ws and We are two model parameters.
[0088] Fifth, based on the sum of the prediction losses corresponding to each of the multiple training samples, adjust the parameters of the language processing model.
[0089] Exemplarily, the sum of the prediction losses loss = -y * log(y') - ∑js k * log(s' k ) - ∑je n * log(e' n ), where the first term is the first prediction loss, the second term is the second prediction loss (including the prediction loss at the start position and the prediction loss at the end position), y is the labeled tag corresponding to the classification training sample, s k and e n are respectively the start position and the predicted position corresponding to the labeled tag of the persona attribute information training sample, i is a natural positive integer, and j has M values.
[0090] Through the above technical solution, two related tasks (including the task of classifying any personal setting dimensions and the task of predicting personal setting attribute information) are learned together. The purpose is to make full use of the common knowledge between related tasks and improve the model learning and generalization effects on any single task through the way of shared learning. The two tasks of classifying personal setting dimensions and predicting personal setting attribute information involved in this disclosure have obvious similarities. On the one hand, classifying personal setting dimensions can help the task of predicting personal setting attribute information to be more accurate. On the other hand, knowing certain personal setting attribute information can also deepen the model's understanding of the personal setting dimension category. Therefore, adopting this training method can further improve the classification and extraction effects of the model.
[0091] Figure 3 It is a flowchart of another human-machine dialogue method shown according to an exemplary embodiment of the present disclosure. Refer to Figure 3 , generating the target response statement according to the target personal setting attribute information may include the following steps:
[0092] In step 31, a response template statement is determined according to the target personal setting attribute information, where the response template statement is a statement including unfilled slots, and each unfilled slot carries a keyword identifier.
[0093] It should be noted that the response template statement can be determined by analyzing the existence and consistency of the target personal setting attribute information and the historical personal setting attribute information corresponding to the target personal setting dimension. Among them, existence refers to whether there is information about an attribute related to the target personal setting attribute information in the historical personal setting attribute information. Consistency refers to whether the historical personal setting attribute information related to the target personal setting attribute information is consistent with the target personal setting attribute information.
[0094] Exemplarily, the response template statement can be: "Okay, I remember. You are {time - qualifier}{age - number}", where {time - qualifier} and {age - number} are unfilled slots, time - qualifier is the keyword identifier of {time - qualifier}, which represents a time qualifier, and age - number is the keyword identifier of {age - number}, which represents an age number qualifier. The response template statement can also be: "So it's {age - compare}, I remember", {age - compare} is an unfilled slot, and age - compare is the keyword identifier of {age - compare}, which represents an attribute comparison qualifier.
[0095] In step 32, according to each keyword identifier of the unfilled slots in the response template statement, the slot information corresponding to the keyword identifier is determined.
[0096] For example, if the reply template statement is "Okay, I remember. You are {time - qualifier}{age - number}", therefore, it is necessary to determine the slot information corresponding to time - qualifier and age - number.
[0097] In step 33, according to the semantic information of the keyword identifier and the slot information, fill the slot information into the to - be - filled slot corresponding to the slot information in the reply template statement to generate the target reply statement.
[0098] In the above - mentioned manner, by constructing a reply template statement with to - be - filled slots, it is possible to directly control and select the reply to the dialogue input, solving the problem of uncontrollable replies caused by the generative algorithm in the related technology.
[0099] In some possible implementation manners, the determining the reply template statement according to the target persona attribute information includes:
[0100] In the case that there is no historical persona attribute information corresponding to the target persona dimension in the storage module, any one of the first - type preset template statements configured in the template library is used as the reply template statement for the dialogue input.
[0101] It should be noted that the storage module is used to store the persona attribute information under various persona dimensions that appear in the historical dialogue, including the persona attribute information under different persona dimensions on the user side and the intelligent assistant side.
[0102] For example, the first - type preset template statement can be a declarative statement, which is used to echo the target persona attribute information included in the dialogue input. In addition, the first - type preset template statement can also be a statement used to guide the user to input more persona attribute information by voice.
[0103] In the case that there is no historical persona attribute information corresponding to the target persona dimension in the storage module, Figure 3The steps of determining the slot information corresponding to each keyword identifier of the slot to be filled in the reply template statement may include: determining the slot information corresponding to each keyword identifier from the target persona attribute information according to each keyword identifier of the slot to be filled in the reply template statement. For example, in the case where the reply template statement is a statement in the first type of preset template statement for echoing the persona attribute information included in the dialogue input, still taking the above dialogue input "I am seven and a half years old this year and you are older than me" as an example, the determined target persona attribute information includes "this year", "seven and a half years old", "you are older than me", and the reply template statement is "Okay, I remember. You are {time - qualifier}{age - number} now", then according to the keyword identifiers of the slots to be filled, the "this year" and "seven and a half years old" in the target persona attribute information can be added to each slot to be filled, obtaining the target reply statement "Okay, I remember. You are seven and a half years old this year".
[0104] For example, in the case where the reply template statement is a statement in the first type of preset template statement for guiding the user to input more persona attribute information by voice, taking the first input dialogue input of the user as "My age is 15 years old", correspondingly, the determined target persona attribute information may include "age" and "15 years old", and the determined target reply statement may be: What day is your {birthday - qualifier}?, then according to the keyword identifier "birthday - qualifier" and the target persona attribute information, it can be determined that the slot information corresponding to the {birthday - qualifier} slot may be "birthday", and the generated target reply statement may be "What day is your birthday?", thus, the user can be guided to say the persona attribute information related to their birthday, improving the interestingness of the interaction.
[0105] In some possible implementation manners, the step of determining the reply template statement according to the target persona attribute information may include:
[0106] In the case of determining that there is historical persona attribute information in the storage module that is semantically related to the target persona attribute information and the persona attribute information is inconsistent, any one of the second type of preset template statements configured in the template library is used as the reply template statement for the dialogue input.
[0107] For example, taking the dialogue input "I am seven and a half years old this year and you are older than me" as an example, in the storage module, the persona attribute information of "10 years old" in the user's previous input "I am 10 years old this year" is recorded. Therefore, "seven and a half years old" and "10 years old" are semantically related and inconsistent persona attribute information.
[0108] It should be noted that the second type of preset template statement can be a rhetorical clarification statement to guide the user to input voice to clarify the inconsistent persona attribute information. For example, the statement "You said you were {age-number} before, and now you say you are {age-number}?" in the second type of preset template statement can be used as a reply template statement. Based on the semantic relationship of this reply template statement, the position of each {age-number} is determined. Thus, the target reply statement obtained can be: "You said you were 10 years old before, and now you say you are seven and a half years old?", which can guide the user to input voice to clarify the inconsistent persona attribute information.
[0109] In some possible implementation manners, the step of determining a reply template statement according to the target persona attribute information may include:
[0110] In the case where it is determined that there is historical persona attribute information semantically related to the target persona attribute information in the storage module, identify the dialogue intention of the dialogue input; in the case where there is no attribute information that meets the dialogue intention in the target persona attribute information and the historical persona attribute information, any one of the third type of preset template statements configured in the template library is used as the reply template statement for the dialogue input.
[0111] It should be noted that the third type of preset template statement can be a statement for replying to the dialogue intention of the user's dialogue input.
[0112] Exemplarily, when the dialogue input is "So how old was I the previous year", and the storage module records that the user has previously input "I am 10 years old this year", it can be seen that the dialogue intention of the dialogue input cannot be directly obtained from the historical persona attribute information and the target persona attribute information. In this case, "You are {age-number} {time - qualifier}" in the third type of preset template statement can be used.
[0113] In some possible implementation manners, the step of determining the slot information corresponding to each keyword identifier in the reply template statement according to the keyword identifier may include:
[0114] For each keyword identifier of the slot to be filled in the reply template statement, infer the information corresponding to the keyword identifier based on the target persona attribute information and the historical persona attribute information corresponding to the target persona dimension, and use this information as the slot information corresponding to the keyword identifier.
[0115] Exemplarily, when the dialogue input is "How old was I the previous year?", according to the dialogue intention and the keyword identifier time - qualifier, it can be determined that "the previous year" is the keyword identifier time - qualifier for {time - qualifier}. Further, based on "the previous year", "this year", and "10 years old", the slot information corresponding to {age - number} can be inferred to be 9 years old. Therefore, by filling the determined slot information into "You are {time - qualifier}{age - number} years old", the target reply statement obtained is "You were 9 years old the previous year."
[0116] Considering the current scenario where human - machine dialogue usually only maintains a single - round conversation, in a multi - round conversation scenario, the intelligent assistant cannot interact with the user in the current conversation by combining the information related to the user profile attributes in the historical conversation, and thus cannot achieve the purpose of long - term companionship, understanding, and getting to know the user. Therefore, through the above - mentioned method, by combining the target user profile attribute information and the historical user profile attribute information corresponding to the target user profile dimension to generate the target reply statement for the dialogue input, in the multi - round conversation application scenario, the generated target reply statement can cover the user profile attribute information in the historical conversation, achieving the purpose of long - term companionship, understanding, and getting to know the user, thereby improving the interaction experience of the dialogue interaction.
[0117] In some possible implementation manners, the method further includes: when it is determined that there is no historical user profile attribute information corresponding to the target user profile dimension in the storage module, storing the target user profile attribute information in the storage module;
[0118] In some possible implementation manners, the method further includes: when it is determined that there is historical user profile attribute information in the storage module that is semantically related to the target user profile attribute information but has inconsistent user profile attribute information, replacing the historical user profile attribute information in the storage module that is semantically related to the target user profile attribute information and has inconsistent user profile attribute information with the target user profile attribute information.
[0119] Exemplarily, taking the dialogue input "I am seven and a half years old this year and you are older than me" as an example, if there is no historical user profile attribute information related to the age dimension of the user in the storage module, then "seven and a half years old" and "you are older than me" can be stored in the storage module; if there is historical user profile attribute information related to the age dimension of the user in the storage module and the time is all this year as "10 years old", then the historical user profile attribute information of "10 years old" can be replaced with "seven and a half years old" to achieve the update of the attributes in the storage module.
[0120] Through the above - mentioned method, the recording, addition, and deletion of the all - around user profile attribute information on the user side or the robot side are realized, so as to strengthen the memory and update of the user profile dimensions of different user profiles. Based on the user profile attribute information in the storage module, the intelligent assistant can achieve the purpose of long - term companionship with the user.
[0121] Figure 4 This is a schematic structural diagram of a human - machine dialogue device 40 shown according to an exemplary embodiment of the present disclosure. Referring to Figure 4 , the device 40 includes an acquisition module 41, a first determination module 42, a second determination module 43, and a generation module 44.
[0122] The acquisition module 41 is configured to acquire a dialogue input;
[0123] The first determination module 42 is configured to determine, in a pre - constructed persona portrait, a target persona dimension corresponding to the dialogue input, where the persona portrait includes multiple persona dimensions and keywords corresponding to each persona dimension;
[0124] The second determination module 43 is configured to determine, according to the semantic relationship between the target keyword corresponding to the target persona dimension and the dialogue input, target persona attribute information corresponding to the dialogue input;
[0125] The generation module 44 is configured to generate a target response statement according to the target persona attribute information.
[0126] In some embodiments, the second determination module 43 is specifically configured to input the dialogue input into a trained language processing model, and the language processing model is used to predict, for each target keyword, the persona attribute information corresponding to the target keyword in the dialogue input, and use this persona attribute information as the target persona attribute information corresponding to the dialogue input.
[0127] In some embodiments, the language processing model includes a pre - processing layer, an encoding layer and a classification layer for the persona dimension classification scenario, and an extraction layer for the persona attribute information extraction scenario. The language processing model is trained in the following manner:
[0128] Acquire a plurality of training samples, where the plurality of training samples include classification training samples collected in the persona dimension classification scenario and persona attribute information training samples collected in the persona attribute information extraction scenario. Each training sample in the plurality of training samples includes user input text and a corresponding labeled tag;
[0129] Input each training sample into the pre - processing layer to obtain a character sequence corresponding to the user input text in the training sample;
[0130] When the training sample belongs to the classification training sample, input the character sequence of the training sample into the encoding layer to obtain semantic vectors corresponding to each character, and input the average vector of the semantic vectors of all characters into the classification layer. Based on the classification result output by the classification layer and the labeled tag in the training sample, determine the first prediction loss corresponding to the training sample;
[0131] When the training sample belongs to the persona attribute information training sample, input the character sequence of the training sample into the extraction layer. Based on the extraction result output by the extraction layer and the labeled tag in the training sample, determine the second prediction loss corresponding to the training sample;
[0132] Based on the sum of the prediction losses corresponding to each of the multiple training samples, adjust the parameters of the language processing model.
[0133] In some embodiments, the generation module 44 includes:
[0134] A template generation sub-module, configured to determine a reply template statement, where the reply template statement is a statement including slots to be filled, and each of the slots to be filled carries a keyword identifier;
[0135] A slot information determination sub-module, configured to determine slot information corresponding to each keyword identifier according to each keyword identifier of the slots to be filled in the reply template statement;
[0136] A filling sub-module, configured to fill the slot information into the slot to be filled corresponding to the slot information in the reply template statement according to the semantic information of the keyword identifier and the slot information, and generate a target reply statement.
[0137] In some embodiments, the template generation sub-module includes a first template generation sub-template, configured to, when it is determined that there is no historical persona attribute information corresponding to the target persona dimension in the storage module, use any one of the first type of preset template statements configured in the template library as the reply template statement for the dialogue input;
[0138] The slot information determination sub-module includes a first slot information determination sub-module, configured to determine slot information corresponding to each keyword identifier from the target persona attribute information according to each keyword identifier of the slots to be filled in the reply template statement.
[0139] In some embodiments, the template generation sub-module includes a second template generation sub-template, which is configured to, when it is determined that there is historical persona attribute information in the storage module that is semantically related to the target persona attribute information and has inconsistent persona attribute information, use any one of the second type of preset template statements configured in the template library as the reply template statement for the dialogue input.
[0140] In some embodiments, the template generation sub-module includes a third template generation sub-template, which is configured to, when it is determined that there is historical persona attribute information in the storage module that is semantically related to the target persona attribute information, identify the dialogue intention of the dialogue input;
[0141] When there is no attribute information in the target persona attribute information and the historical persona attribute information that satisfies the dialogue intention, use any one of the third type of preset template statements configured in the template library as the reply template statement for the dialogue input;
[0142] The slot information determination sub-module includes a second slot information determination sub-module, which is configured to, for each keyword identifier of the to-be-filled slots in the reply template statement, infer the information corresponding to the keyword identifier based on the target persona attribute information and the historical persona attribute information corresponding to the target persona dimension, and use the information as the slot information corresponding to the keyword identifier.
[0143] In some embodiments, the human-machine dialogue device 40 further includes:
[0144] A recording module, which is configured to, when it is determined that there is no historical persona attribute information corresponding to the target persona dimension in the storage module, store the target persona attribute information in the storage module;
[0145] In some embodiments, the human-machine dialogue device 40 further includes:
[0146] A replacement module, which is configured to, when it is determined that there is historical persona attribute information in the storage module that is semantically related to the target persona attribute information and has inconsistent persona attribute information, replace the historical persona attribute information in the storage module that is semantically related to the target persona attribute information and has inconsistent persona attribute information with the target persona attribute information.
[0147] In some embodiments, the first determination module 42 is further configured to input the dialogue input into a trained language processing model, and the language processing model is used to predict the target persona dimension among all the persona dimensions included in the pre-constructed persona portrait to which the dialogue input belongs.
[0148] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0149] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, and when the program instructions are executed by a processor, the steps of the human-machine dialogue method provided by the present disclosure are implemented.
[0150] Figure 5 is a block diagram of an electronic device shown according to an exemplary embodiment. For example, the electronic device 500 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0151] Referring to Figure 5 , the electronic device 500 may include one or more of the following components: a processing component 502, a memory 504, a power component 506, a multimedia component 508, an audio component 510, an input / output (I / O) interface 512, a sensor component 514, and a communication component 516.
[0152] The processing component 502 generally controls the overall operation of the electronic device 500, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 502 may include one or more processors 520 to execute instructions to complete all or part of the steps of the above human-machine dialogue method. In addition, the processing component 502 may include one or more modules to facilitate the interaction between the processing component 502 and other components. For example, the processing component 502 may include a multimedia module to facilitate the interaction between the multimedia component 508 and the processing component 502.
[0153] The memory 504 is configured to store various types of data to support the operation of the electronic device 500. Examples of such data include instructions for any application or method operating on the electronic device 500, contact data, phone book data, messages, pictures, videos, etc. The memory 504 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0154] The power component 506 provides power to various components of the electronic device 500. The power component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 500.
[0155] The multimedia component 508 includes a screen that provides an output interface between the electronic device 500 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of a touch or swipe action but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 508 includes a front camera and / or a rear camera. When the electronic device 500 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0156] The audio component 510 is configured to output and / or input audio signals. For example, the audio component 510 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 500 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 504 or transmitted via the communication component 516. In some embodiments, the audio component 510 further includes a speaker for outputting audio signals.
[0157] The I / O interface 512 provides an interface between the processing component 502 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0158] The sensor assembly 514 includes one or more sensors for providing an assessment of the status of the electronic device 500 in various aspects. For example, the sensor assembly 514 can detect the on / off state of the electronic device 500, the relative positioning of components, such as the display and keypad of the electronic device 500. The sensor assembly 514 can also detect a change in the position of the electronic device 500 or a component of the electronic device 500, the presence or absence of user contact with the electronic device 500, the orientation or acceleration / deceleration of the electronic device 500, and a change in the temperature of the electronic device 500. The sensor assembly 514 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 514 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 514 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0159] The communication component 516 is configured to facilitate communication between the electronic device 500 and other devices in a wired or wireless manner. The electronic device 500 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 516 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 516 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0160] In an exemplary embodiment, the electronic device 500 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described human-machine dialogue method.
[0161] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as the memory 504 including instructions, is also provided. The above instructions can be executed by the processor 520 of the electronic device 500 to complete the above-described human-machine dialogue method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0162] In another exemplary embodiment, a computer program product is also provided, which includes a computer program executable by a programmable device. The computer program has a code portion for performing the above-mentioned human-machine dialogue method when executed by the programmable device.
[0163] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0164] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A human-computer dialogue method, characterized in that, the method includes: obtaining dialogue input; in a pre-constructed persona portrait, determining a target persona dimension corresponding to the dialogue input, the persona portrait including multiple persona dimensions and keywords corresponding to each persona dimension, the keywords being used to assist in determining the persona attribute information of the persona dimension; determining target persona attribute information corresponding to the dialogue input according to the semantic relationship between the target keyword corresponding to the target persona dimension and the dialogue input; generating a target response statement according to the target persona attribute information; The determining the target persona attribute information corresponding to the dialogue input according to the semantic relationship between the target keyword corresponding to the target persona dimension and the dialogue input includes: inputting the dialogue input into a trained language processing model, the language processing model being used to predict, for each target keyword, the start position and end position of the persona attribute information corresponding to the target keyword in the dialogue input, intercepting the persona attribute information corresponding to the target keyword in the dialogue input according to the start position and end position, and using the persona attribute information as the target persona attribute information corresponding to the dialogue input.
2. The method according to claim 1, characterized in that, the language processing model includes a preprocessing layer, an encoding layer and a classification layer for the persona dimension classification scenario, and an extraction layer for the persona attribute information extraction scenario, and the language processing model is trained in the following manner: obtaining a plurality of training samples, wherein the plurality of training samples include classification training samples collected in the persona dimension classification scenario and persona attribute information training samples collected in the persona attribute information extraction scenario, and each training sample in the plurality of training samples includes user input text and a corresponding labeled tag; inputting each training sample into the preprocessing layer to obtain a character sequence corresponding to the user input text in the training sample; when the training sample belongs to the classification training sample, inputting the character sequence of the training sample into the encoding layer to obtain a semantic vector corresponding to each character, and inputting the average vector of the semantic vectors of all characters into the classification layer, and determining the first prediction loss corresponding to the training sample based on the classification result output by the classification layer and the labeled tag in the training sample; when the training sample belongs to the persona attribute information training sample, inputting the character sequence of the training sample into the extraction layer, and determining the second prediction loss corresponding to the training sample based on the extraction result output by the extraction layer and the labeled tag in the training sample; adjusting the parameters of the language processing model based on the sum of the prediction losses corresponding to the plurality of training samples.
3. The method according to claim 1, characterized in that, the generating a target response statement according to the target persona attribute information includes: determining a response template statement according to the target persona attribute information, wherein the response template statement is a statement including slots to be filled, and each slot to be filled carries a keyword identifier; Determine the slot information corresponding to each keyword identifier of the to-be-filled slots described in the reply template statement; Fill the slot information into the to-be-filled slots corresponding to the slot information in the reply template statement according to the semantic information of the keyword identifier and the slot information to generate a target reply statement.
4. The method according to claim 3, wherein, the determining the reply template statement according to the target persona attribute information includes: when it is determined that there is no historical persona attribute information corresponding to the target persona dimension in the storage module, taking any one of the first type of preset template statements configured in the template library as the reply template statement for the dialogue input; the determining the slot information corresponding to each keyword identifier of the to-be-filled slots described in the reply template statement includes: determining the slot information corresponding to each keyword identifier from the target persona attribute information according to each keyword identifier of the to-be-filled slots described in the reply template statement.
5. The method according to claim 3, wherein, the determining the reply template statement according to the target persona attribute information includes: when it is determined that there is historical persona attribute information that is semantically related to the target persona attribute information but has inconsistent persona attribute information in the storage module, taking any one of the second type of preset template statements configured in the template library as the reply template statement for the dialogue input.
6. The method according to claim 3, wherein, the determining the reply template statement according to the target persona attribute information includes: when it is determined that there is historical persona attribute information that is semantically related to the target persona attribute information in the storage module, identifying the dialogue intention of the dialogue input; when there is no attribute information that satisfies the dialogue intention in the target persona attribute information and the historical persona attribute information, taking any one of the third type of preset template statements configured in the template library as the reply template statement for the dialogue input; the determining the slot information corresponding to each keyword identifier of the to-be-filled slots described in the reply template statement includes: for each keyword identifier of the to-be-filled slots described in the reply template statement, inferring the information corresponding to the keyword identifier according to the target persona attribute information and the historical persona attribute information corresponding to the target persona dimension, and taking the information as the slot information corresponding to the keyword identifier.
7. The method according to any one of claims 1-6, wherein, the method further includes: when it is determined that there is no historical persona attribute information corresponding to the target persona dimension in the storage module, storing the target persona attribute information in the storage module.
8. The method according to any one of claims 1-6, wherein, the method further includes: In the case where there is historical persona attribute information in the storage module that is semantically related to the target persona attribute information but has inconsistent persona attribute information, replace the historical persona attribute information in the storage module that is semantically related to the target persona attribute information and has inconsistent persona attribute information with the target persona attribute information.
9. The method according to claim 1, wherein, the determining the target persona dimension corresponding to the conversation input includes: inputting the conversation input into a trained language processing model, the language processing model being configured to predict the target persona dimension among all persona dimensions included in a pre-constructed persona portrait to which the conversation input belongs.
10. A human-machine dialogue device, wherein, comprising: an acquisition module configured to acquire a conversation input; a first determination module configured to determine, in a pre-constructed persona portrait, a target persona dimension corresponding to the conversation input, the persona portrait including multiple persona dimensions and keywords corresponding to each persona dimension, the keywords being used to assist in determining the persona attribute information of the persona dimension; a second determination module configured to determine, according to the semantic relationship between the target keyword corresponding to the target persona dimension and the conversation input, the target persona attribute information corresponding to the conversation input; a generation module configured to generate a target reply statement according to the target persona attribute information; the second determination module is used for: inputting the conversation input into a trained language processing model, the language processing model being configured to, for each target keyword, predict the start position and end position of the persona attribute information corresponding to the target keyword in the conversation input, intercept the persona attribute information corresponding to the target keyword in the conversation input according to the start position and end position, and use the persona attribute information as the target persona attribute information corresponding to the conversation input.
11. An electronic device, wherein, comprising: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the human-machine dialogue method according to any one of claims 1-9.
12. A computer-readable storage medium having computer program instructions stored thereon, wherein, when the program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Reply generation method and device
CN110069612A
Question and answer processing method, device and equipment
CN112395398A