Method, device, electronic device and storage medium for human-computer conversation
By extracting users' psychological characteristics and statement vectors from the human-computer dialogue system and using a neural network model to generate responses that conform to users' psychology, the problem of insufficient humanization in existing human-computer dialogue systems is solved, and higher empathy and response satisfaction are achieved.
Patent Information
- Application Number
- CN202210693162.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-17
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2042-06-17
AI Technical Summary
Current human-computer dialogue systems often generate overly formal or rigid dialogue responses, resulting in a low degree of humanization and making it difficult to accurately and genuinely meet the psychological needs of users.
By extracting the user's psychological feature vector and the word embedding vector of the statement, and inputting them into a trained neural network model, candidate responses that integrate the user's psychological features are generated, and finally, the final response that meets the user's psychological needs is determined.
It improves the empathy and humanization of human-computer dialogue, better meeting users' psychological needs and providing practical answers.
Smart Images

Figure CN117033565B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, and more particularly, to a method, apparatus, device and storage medium for human-computer conversation. BACKGROUND
[0002] Human-computer conversation is a core technology for smooth communication between humans and computers, and is one of the ultimate goals of the development of artificial intelligence. In recent years, with the continuous in-depth research of conversation technology, human-computer conversation has gradually entered the public eye, bringing great convenience to people's daily life. From intelligent customer service to personalized chat, from emotional companionship to social robots, human-computer conversation has brought a new intelligent revolution to information acquisition, and has greatly promoted the research and application of natural language processing technology.
[0003] However, current human-computer conversation often generates too official or stereotyped conversation replies, resulting in a low degree of humanization of human-computer conversation. SUMMARY
[0004] The embodiments of the present application provide a method, apparatus, device and storage medium for human-computer conversation, which can reply to the user's question based on the psychological characteristics of the user, help to provide the user with a reply that meets the user's psychological needs, and thus improve the degree of humanization of human-computer conversation.
[0005] In a first aspect, the embodiments of the present application provide a method for human-computer conversation, comprising:
[0006] extracting a psychological characteristic vector of the user according to text content related to the user;
[0007] processing a sentence input by the user to obtain a word embedding vector of the sentence;
[0008] inputting the psychological characteristic vector and the word embedding vector into a trained first neural network model to obtain a first candidate reply to the sentence; wherein the first neural network model is trained according to psychological characteristics of a user sample and human-computer conversation data samples of the user sample;
[0009] determining a final reply to the sentence according to the first candidate reply.
[0010] In a second aspect, the embodiments of the present application provide an apparatus for human-computer conversation, characterized in that it comprises:
[0011] an extraction unit configured to extract a psychological characteristic vector of the user according to text content related to the user;
[0012] an obtaining unit configured to process a sentence input by the user to obtain a word embedding vector of the sentence;
[0013] a first neural network model, configured to input the psychological feature vector and the word embedding vector into the trained first neural network model to obtain a first candidate reply for the sentence; wherein the first neural network model is trained according to psychological features of a user sample and human-computer conversation data samples of the user sample;
[0014] a determination unit, configured to determine a final reply for the sentence according to the first candidate reply.
[0015] In a third aspect, an embodiment of the present application provides an electronic device, comprising:
[0016] a processor, adapted to implement computer instructions; and
[0017] a memory, storing computer instructions, the computer instructions being adapted to be loaded and executed by the processor to implement the method of the first aspect.
[0018] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, which stores computer instructions. When the computer instructions are read and executed by a processor of a computer device, the computer device implements the method of the first aspect.
[0019] In a fifth aspect, an embodiment of the present application provides a computer program product or a computer program, which comprises computer instructions. The computer instructions are stored in a computer readable storage medium. When the computer instructions are read by a processor of a computer device, the processor executes the computer instructions, so that the computer device implements the method of the first aspect.
[0020] Based on the above technical solutions, the psychological feature vector of the user and the word embedding vector of the input sentence are input into the trained first neural network model, so that the first candidate reply which fuses the psychological features of the user can be obtained, and then the final reply for the input sentence of the user can be obtained according to the first candidate reply. Since the first candidate reply fuses the psychological features of the user, the final reply obtained according to the first candidate reply has good empathy ability and can meet the psychological needs of the user, thereby helping to improve the humanization degree of human-computer conversation. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 FIG. 1 is a schematic diagram of a system architecture according to an embodiment of the present application;
[0022] Figure 2 FIG. 2 is a schematic flowchart of a method of human-computer conversation according to an embodiment of the present application;
[0023] Figure 3A schematic diagram of an applicable network architecture provided for an embodiment of the present application;
[0024] Figure 4 A schematic flow chart of another method of human-computer conversation provided for an embodiment of the present application;
[0025] Figure 5 A schematic diagram of another applicable network architecture provided for an embodiment of the present application;
[0026] Figure 6 A schematic flow chart of another method of human-computer conversation provided for an embodiment of the present application;
[0027] Figure 7 A schematic diagram of another applicable network architecture provided for an embodiment of the present application;
[0028] Figure 8 A schematic flow chart of another method of human-computer conversation provided for an embodiment of the present application;
[0029] Figure 9 A schematic diagram of another applicable network architecture provided for an embodiment of the present application;
[0030] Figure 10 A schematic diagram of a Seq2Seq model with attention mechanism provided for an embodiment of the present application;
[0031] Figure 11 A schematic flow chart of another method of human-computer conversation provided for an embodiment of the present application;
[0032] Figure 12 A schematic flow chart of another method of human-computer conversation provided for an embodiment of the present application;
[0033] Figure 13 A schematic block diagram of a device of human-computer conversation provided for an embodiment of the present application;
[0034] Figure 14 A schematic block diagram of an electronic device provided for an embodiment of the present application. DETAILED DESCRIPTION
[0035] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0036] It should be understood that, in the embodiments of the present application, "B corresponding to A" means that B is associated with A. In an implementation, B can be determined according to A. However, it should also be understood that determining B according to A does not mean that B is determined only according to A, but B can also be determined according to A and / or other information.
[0037] In the description of the present application, "at least one" means one or more, and "multiple" means two or more than two. In addition, "and / or" describes the association between the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it. "At least one of the following" or the like means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.
[0038] It should also be understood that the first, second, and the like descriptions appearing in the embodiments of the present application are only for illustrative and distinguishing purposes, and do not have any order or represent a special limitation on the number of devices in the embodiments of the present application, and cannot constitute any limitation on the embodiments of the present application.
[0039] It should also be understood that the specific features, structures or characteristics in the description related to the embodiments include at least one embodiment of the present application. In addition, these specific features, structures or characteristics can be combined in any suitable manner in one or more embodiments.
[0040] In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units does not have to be limited to those clearly listed steps or units, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0041] The embodiments of the present application apply to the field of artificial intelligence technology.
[0042] Among them, artificial intelligence (Artificial Intelligence, AI) is to use digital computers or digital computer controlled machine simulation, extension and expansion of human intelligence, perception of the environment, knowledge acquisition and use of knowledge to obtain the best results of theory, method, technology and application system. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence, and produces a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.
[0043] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software level technology. Artificial intelligence basic technology generally includes, such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology and machine learning / deep learning and other several major directions.
[0044] With the research and progress of artificial intelligence technology, artificial intelligence technology is researched and applied in many fields, such as common smart home, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned vehicles, autonomous vehicles, drones, robots, intelligent medical treatment, intelligent customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0045] The embodiments of the present application can involve natural language processing (Nature Language processing, NLP) technology in artificial intelligence technology. NLP is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can realize effective communication between man and computer using natural language. Natural language processing is a science that integrates linguistics, computer science and mathematics. Therefore, the research in this field will involve natural language, i.e. the language used in daily life, so it has a close relationship with the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question and answer, knowledge graph and other technologies.
[0046] The embodiments of the present application can also relate to machine learning (ML) in artificial intelligence technology. ML is a multi-disciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc. It is a subject that studies how computers simulate or implement human learning behavior to acquire new knowledge or skills, reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent. Its applications are widespread in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and inductive learning.
[0047] In addition, the scheme provided by the embodiments of the present application can also involve model pre-training technology. Pre-training trains a model through a large number of unlabeled or weakly labeled samples to obtain a set of model parameters; the set of parameters is used to initialize the model to achieve a "hot start", and then the parameters are fine-tuned according to the specific task to fit the label data provided by the task on the existing model architecture.
[0048] Figure 1 For a system architecture involved in the embodiments of the present application, as shown in Figure 1 The system architecture can include a user device 101, a data collection device 102, a training device 103, an execution device 104, a database 105, and a content library 106.
[0049] The data collection device 102 is configured to read training data from the content library 106 and store the read training data into the database 105. The training data involved in the embodiments of the present application includes psychological characteristics of user samples and human-computer dialogue data samples of user samples.
[0050] The training device 103 trains a machine learning model based on the training data maintained in the database 105, so that the trained machine learning model can effectively reply to the user's question. The machine learning model obtained by the training device 103 can be applied to different systems or devices.
[0051] In addition, with reference to Figure 1The execution device 104 is configured with an I / O interface 107 to interact with external devices for data exchange. For example, the I / O interface receives the question sentence sent by the user device 101 and the user-related text content. The computing module 109 in the execution device 104 extracts the psychological feature vector of the user according to the user-related text content, obtains the word embedding vector of the question sentence of the user, splices the psychological feature vector and the word embedding vector to obtain an input vector based on psychological features, and inputs the input vector based on psychological features into the trained human-computer dialogue model to output the reply to the question sentence. The human-computer dialogue model can send the corresponding result to the user device 101 through the I / O interface.
[0052] The user device 101 can include a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted terminal, a mobile internet device (MID), or other terminal devices.
[0053] The execution device 104 can be a server.
[0054] The server can be a rack-mounted server, a blade server, a tower server, or a cabinet server, etc. The server can be a standalone test server, or a test server cluster composed of multiple test servers.
[0055] The server can be one or more. When the server is more than one, at least two servers are used to provide different services, and / or at least two servers are used to provide the same service, such as providing the same service in a load balancing manner, and the embodiments of the present application do not limit this.
[0056] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, etc. The server can also be a node of a blockchain.
[0057] In this embodiment, the execution device 104 is connected with the user device 101 through a network. The network can be an Intranet, the Internet, a Global System of Mobile communication (GSM), a Wideband Code Division Multiple Access (WCDMA), a 4G network, a 5G network, Bluetooth, Wi-Fi, a telephony network, or other wireless or wired networks.
[0058] It should be noted that, Figure 1 The positional relationship between the devices, elements, modules, etc. shown in the figure does not constitute any limitation. In some embodiments, the data collection device 102, the user device 101, the training device 103, and the execution device 104 can be the same device. The database 105 can be distributed on one server or multiple servers, and the content library 106 can be distributed on one server or multiple servers.
[0059] Figure 1 The network architecture in the embodiment can be applied to a human-computer dialogue model, and can provide a humanized reply meeting the psychological needs of a user by fusing the psychological features of the user.
[0060] The embodiment of the present application can be applied to the field of intelligent medical treatment of artificial intelligence. Specifically, the embodiment of the present application can be used in a human-computer dialogue system in the field of intelligent medical treatment.
[0061] Medical dialogue, as an important research direction in the field of human-computer dialogue, aims to make people more conveniently obtain medical and health information, and is widely used in the fields of psychological counseling, nursing for the elderly, medical question and answer, and disease diagnosis, etc. For example, in the face of the current sudden epidemic, online medical consultation has become increasingly frequent, which also makes the research of medical dialogue technology more urgent. The related research accelerates the spread of medical information, makes people more conveniently master their own health status, provides an important guarantee for improving the health level of human beings and improving the quality of life of the whole people, and has important theoretical significance and practical value. In medical dialogue, the disease name, symptoms, and body parts involved in the user's question are recognized by recognizing the user's intention, which can assist the human-computer dialogue system to generate effective and professional medical replies and improve the overall performance of the medical dialogue system.
[0062] Meanwhile, mental health is increasingly concerned and valued by the public. How to maintain a healthy psychological state while maintaining physical health is an important guarantee for improving the quality of life. However, the current human-computer dialogue often ignores the psychological state of the user, and it is difficult to accurately and directly provide a reply with empathy ability that meets the psychological state of the user, for example, generating a dialogue reply that is too official or stereotyped, resulting in a low degree of humanization of human-computer dialogue. For example, the current medical dialogue often ignores the psychological state of the user, and it is difficult to accurately and directly provide a reply that meets the psychological state of the user, resulting in the inability to directly provide a reply with empathy ability, and even causing misdiagnosis or missed diagnosis of the user's mental health problems.
[0063] Therefore, the embodiments of the present application provide a method, device, electronic equipment and storage medium for human-computer dialogue, which can reply to the user's question based on the psychological characteristics of the user, help to directly provide a reply that meets the psychological needs of the user, and further improve the humanization degree of human-computer dialogue.
[0064] Specifically, the embodiments of the present application extract the psychological feature vector of the user through the text content related to the user, process the sentence input by the user, obtain the word embedding vector of the sentence, then input the psychological feature vector and the word embedding vector into the trained first neural network model to obtain the first candidate reply to the sentence input by the user. Finally, according to the first candidate reply, the final reply to the sentence input by the user is determined.
[0065] The first neural network model is trained according to the psychological characteristics of the user sample and the human-computer dialogue data sample of the user sample. The human-computer dialogue data can include user question text and reply text corresponding to the question.
[0066] In some embodiments, when the method of human-computer dialogue is applied to a medical dialogue system, the human-computer dialogue data sample used to train the first neural network model can include human-computer dialogue data between doctors and patients, which can be related to medical consultation or psychological consultation, and the present application does not limit this.
[0067] Therefore, the embodiments of the present application can obtain the first candidate reply that fuses the psychological characteristics of the user by inputting the psychological feature vector of the user and the word embedding vector of the input sentence into the trained first neural network model, and then obtaining the final reply to the input sentence of the user according to the first candidate reply. Since the first candidate reply fuses the psychological characteristics of the user, the final reply obtained according to the first candidate reply has good empathy ability and can meet the psychological needs of the user, thereby helping to improve the humanization degree of human-computer dialogue.
[0068] For example, the psychological characteristics can refer to at least one of a user personality characteristic, a depression tendency characteristic, and a suicide tendency characteristic mined based on text content related to the user, and the present application does not limit this.
[0069] As an example, one of the application scenarios of the embodiments of the present application can be a dialogue robot, and the embodiments of the present application can enhance the humanization degree of the dialogue service of the dialogue robot. Specifically, the human-computer dialogue method provided by the embodiments of the present application can be applied to the dialogue robot, and the user's question is replied to individually based on the psychological characteristics of the user, and the psychological burden of the user is timely relieved. For example, the dialogue robot can extract the psychological characteristic vector of the user and obtain the word embedding vector of the user's question according to the historical chat data with the user, input the psychological characteristic vector of the user and the word embedding vector of the user's question into a neural network model to obtain a candidate reply, and finally determine the final reply to the user's question based on the candidate reply. In this scenario, the reply based on the psychological characteristics of the user can timely relieve the psychological burden of the user, help the user to fully and freely communicate in the process of communication, and thus alleviate the psychological health problem of the user, so that the dialogue robot provides more humanized dialogue service.
[0070] As another example, one of the application scenarios of the embodiments of the present application can be a micro medical platform, and the embodiments of the present application can provide one-to-one health answering service of doctors to users on the micro medical platform. Specifically, the human-computer dialogue method provided by the embodiments of the present application can be applied to the micro medical platform, and the question of the user is replied to professionally based on the psychological characteristics of the user, and the psychological health support and answering of the user are realized. For example, when the user feeds back the psychological health status of the user, the micro medical platform can extract the psychological characteristic vector of the user and obtain the word embedding vector of the question of the user, input the psychological characteristic vector of the user and the word embedding vector of the question of the user into a neural network model to obtain a candidate reply, and finally determine the final medical reply to the question of the user based on the candidate reply, realize automatic generation of guidance opinions or solutions of related psychological problems, reduce the labor cost and workload of one-to-one answering of doctors, and improve the solution efficiency of medical consultation problems.
[0071] Hereinafter, the human-computer dialogue method provided by the embodiments of the present application is described in detail in combination with the drawings.
[0072] Figure 2 A schematic flowchart of a human-computer dialogue method 200 provided by the embodiments of the present application. The method 200 can be executed by any electronic device with data processing capability. For example, the electronic device can be implemented as a server or a terminal device, and for example, the electronic device can be implemented as a computing module 109 in the server, and the present application does not limit this. Figure 1
[0073] In some embodiments, a machine learning model can be included (e.g., deployed) in the electronic device, which can be a deep learning model, a neural network model, or other models without limitation. In some embodiments, the machine learning model can be a human-computer dialogue model or a medical dialogue model, etc., which can reply to the user's question based on the user's psychological characteristics.
[0074] Figure 3 is a schematic diagram of a network architecture 300 applicable to embodiments of the present application, which includes a psychological characteristic extraction module 301, a dialogue context acquisition module 302, a first neural network model 303, a second neural network model 304, a selection module 305, a human-computer dialogue data source acquisition module 306, and a retrieval model 307. The network architecture 300 can be referred to as a human-computer dialogue system architecture. In the following, the network architecture will be described in conjunction with a method 200 of human-computer dialogue. Figure 3
[0075] As shown in Figure 2 , the method 200 of human-computer dialogue can include steps 210 to 240.
[0076] 210, according to the user-related text content, extracting the psychological characteristic vector of the user.
[0077] Here, the user-related text content can include at least one of the historical dialogue context of the user's human-computer interaction and the text content associated with the user on the social media network. The text content associated with the user on the social media network can be the text content published by the user on the social media network, or can be the text content read or commented by the user, or other more diverse user-related text content, which is not limited by the present application. Through the diversification of the user-related text content, it can be helpful to more comprehensively extract the long-term psychological characteristics of the user, so as to more accurately extract the psychological characteristic vector of the user.
[0078] In some embodiments, the psychological characteristics can refer to at least one of the user's personality characteristics, depression tendency characteristics, and suicide tendency characteristics mined based on the user-related text content, which is not limited by the present application. Accordingly, the psychological characteristic vector includes at least one of the user's personality characteristic vector, depression tendency characteristic vector, and suicide tendency characteristic vector.
[0079] Specifically, the user's psychological state can be described from two aspects, one is the relatively stable user personality characteristics, and the other is the continuously changing user's mental health state, such as the user's depression tendency characteristics and suicide tendency characteristics, etc. Therefore, through the user's personality characteristics, depression tendency characteristics, or suicide tendency characteristics, the user's mental health state can be more richly and comprehensively represented.
[0080] Referring to Figure 3 The psychological feature vector of the user can be extracted from the user-related text content by the psychological feature extraction module 301.
[0081] In some embodiments, when the above-mentioned psychological feature vector includes a personality feature vector, referring to Figure 4 , the personality feature vector can be specifically obtained through the following processes 2101 to 2103.
[0082] 2101, for each word in the i-th message in the text content, the hidden state of each word is obtained by a first GRU unit. Wherein, i is a positive integer, i is less than or equal to the total number of messages in the text content.
[0083] Referring to Figure 5 , a specific example of the network architecture of the psychological feature extraction module 301 is shown, which can be used to extract the personality feature vector of the user. In Figure 5 , the user-related text content is taken as an example of the user historical dialogue context. For the i-th message in the user historical dialogue context, the word embedding vector of each word can be extracted by the word embedding layer self501. Then, the extracted word embedding vector of each word is input into the first Gated Recurrent Unit (GRU), such as GRU units 502, 503 and 504, to obtain the hidden state of each word.
[0084] For example, for the j-th word in the i-th message, the corresponding hidden state obtained by the GRU unit can be represented as Specifically, the GRU unit can use the following encoding method:
[0085]
[0086] Wherein, r t represents the reset gate of the GRU unit, which is used to capture the short-term dependence relationship in the time series, and can filter out the content associated with the current input from the previous time; z t represents the reset gate of the GRU unit, which is used to capture the long-term dependence relationship in the time series, and can filter out the content to be retained from the current time; h t-1 represents the hidden state of the previous time, represents the memory state of the current time;{W z ,W r ,W} represents the weight matrix; σ represents the activation function, and the activation result is a value between -1 and 1.
[0087] In some embodiments, the word embedding layer 501 can be a BERT (Bidirectional Encoder Representations from Transformers), a long-short term memory (LSTM), a convolutional neural network (CNN), or other models, without limitation. For example, the text vector representation extracted by the word embedding layer 501 can be an embedding vector, without limitation. The word embedding layer can also be referred to as a word representation layer.
[0088] For example, the word embedding layer 501 can convert the one-hot encoding vector of a word into a dense vector of the word based on the user-related text content using a pre-trained language model (such as BERT, LSTM, CNN, etc.).
[0089] It should be noted that the first GRU unit is taken as an example of a unidirectional GRU unit in the above description, but embodiments of the present application are not limited thereto. For example, the first GRU unit can also be a bidirectional GRU unit or other GRU units. Figure 5
[0090] 2102, the hidden state of each word is weighted and summed using the first attention mechanism, and the hidden state of the ith message is obtained through the second GRU unit.
[0091] Continuing to refer to FIG. 5, the first attention mechanism can be used on the hidden state of each word of the ith message obtained by the first GRU unit (such as the GRU units 502, 503, and 504) to obtain the weight of the hidden state of each word, and the hidden state of each word is weighted and summed according to the weight. Figure 5
[0092] For example, the first attention mechanism can perform the following calculation:
[0093]
[0094] wherein, represents the attention weight of the word hidden state, represents the normalized weight score of the word hidden state, s i represents the word sequence of the ith message represents the weighted combination of the corresponding hidden state, l represents the number of words of the ith message, d word represents the word vector used to learn the word-level attention, b word is a bias term, W word a randomly initialized weight matrix representing the words.
[0095] s i The hidden state of the i-th message can be obtained by a second GRU unit, such as the GRU unit 506. For example, the hidden state of the i-th message h i may be represented as follows:
[0096] h i = GRU(s i ) (3)
[0097] It can be understood that for each message in the user-related content text, the hidden state of each message can be determined based on the above manner.
[0098] For example, the second GRU unit can be a unidirectional GRU unit or a bidirectional GRU unit, which is not limited in the present application.
[0099] 2103, the hidden states of each message of the text content are weighted and summed by using the second attention mechanism, and the personality feature vector of the user is obtained by passing through a fully connected layer.
[0100] Continuing to refer to Figure 5 For the hidden state of each message in the user historical dialogue context, a second attention mechanism, such as the message attention mechanism 507, can be applied to obtain the weight of the hidden state of each message, and the hidden state of each message is weighted and summed according to the weight.
[0101] For example, the second attention mechanism can specifically perform the following calculation:
[0102]
[0103] wherein e i represents the attention weight of the message hidden state, β i represents the normalized weight score of the message hidden state, u' represents the weighted combination of the hidden state corresponding to the message {s1, s2, …, s n}, n represents the number of messages of the user-related text content, e message represents the message vector used to learn the message-level attention, b message is a bias term, and W message represents a randomly initialized weight matrix of the message.
[0104] Continuing to refer to Figure 5 After the hidden states of each message are weighted and summed, the weighted and summed result can also be input into a fully connected layer, such as the fully connected layer 508, to obtain the personality feature vector of the user.
[0105] An exemplary calculation formula of the full connection layer is u" = W * u' + b, which is specifically expanded as the following formula (5):
[0106]
[0107] wherein u' is a set of {u'1, u'2,..., u'N}, which is an input of the full connection layer; u" is a set of {u"1, u"2,..., u"N}, which is an output of the full connection layer; n n W is a weight matrix, b is a bias term, b is in a set form. At this time, the output of the full connection layer 508 is the final obtained personality feature vector, which can be expressed as U1.
[0108] In some embodiments, when the above psychological feature vector includes a depression tendency feature vector, referring to Figure 6 , the depression tendency feature vector can be obtained through the following processes 2104 to 2106.
[0109] 2104, extracting depression tendency feature words from the text content.
[0110] In some embodiments, the depression tendency feature words can include at least one of social network feature words, user emotion feature words, user depression field feature words, and user topic selection feature words. Here, the user's social network feature words, user emotion feature words, user depression field feature words, or user topic selection feature words can reflect whether the user is depressed or the degree of depression.
[0111] For example, the user's social network feature words can include the number of fans of the user's social website, the number of microblogs published by the user, etc.; the user's emotion feature words can include emotional words contained in the text published by the user; the user's depression field feature words can include words with depression tendency contained in the text published by the user; and the user's topic selection feature words can include topic words frequently selected by the user to participate in.
[0112] Referring to Figure 7 , a specific example of the network architecture of the psychological feature extraction module 301 is shown, which can be used to extract the depression tendency feature vector of the user. In Figure 7 For example, the user-related text content is taken as the user historical dialogue context. For the user historical dialogue context, the feature words of the depression tendency can be extracted by the feature word extraction unit 701 to obtain related words such as social network features, emotion features, depression field features, and theme features. Then, for the related words such as the social network features, the emotion features, the depression field features, and the theme features, the word embedding vectors of each word can be extracted by the word embedding layer 702.
[0113] 2105, the hidden state of each feature word in the feature words of the depression tendency is obtained by the third GRU unit.
[0114] Continuing to refer to Figure 7 The word embedding vectors of each word extracted by the word embedding layer 702 can be input into the third GRU unit, such as the GRU units 703, 704, 705, and 706, to obtain the hidden state of each word. As shown in Figure 7 The third GRU unit, such as the GRU units 703, 704, 705, and 706, can be a bidirectional GRU unit, which is not limited in the present application.
[0115] For example, when the third GRU unit is a bidirectional GRU unit, the input of the third GRU unit at time t can be represented as x t Correspondingly, the third GRU unit can output two representation vectors, which are as follows:
[0116]
[0117] wherein, represents the forward hidden state at the current time, represents the reverse hidden state at the current time.
[0118] Then, the and can be spliced by the splicing layer 707 to obtain the hidden state at the current time, as follows:
[0119]
[0120] In some embodiments, the word embedding layer 701 can be BERT, LSTM, CNN, or other models. For example, the text vector representation extracted by the word embedding layer 701 can be an embedding vector, which is not limited in the present application.
[0121] It should be noted that, in Figure 7The third GRU unit is taken as an example of the bidirectional GRU unit, but embodiments of the present application are not limited thereto. For example, the third GRU unit can also be a unidirectional GRU unit or other GRU units.
[0122] 2106, the hidden state of each feature word is weighted and summed by using the third attention mechanism, and a depression tendency feature vector of the user is obtained through a fully connected layer.
[0123] Continuing to refer to Figure 7 The third attention mechanism, such as the attention mechanism 708, can be used for the hidden state of each depression tendency feature word, to obtain the weight of the hidden state of each feature word, and the hidden state of each feature word is weighted and summed according to the weight.
[0124] For example, the third attention mechanism can perform the following calculation:
[0125]
[0126] wherein m t represents the attention weight of the hidden state of the depression tendency feature word, a t represents the normalized weight score of the hidden state of the depression tendency feature word, u' i represents the weighted combination of the hidden state of the depression tendency feature word, and n represents the number of feature words.
[0127] Continuing to refer to Figure 7 After the hidden state of the depression tendency feature word is weighted and summed, the weighted and summed result can also be input into a fully connected layer, such as the fully connected layer 709, to obtain the depression tendency feature vector of the user.
[0128] For example, the calculation method of the fully connected layer can refer to the description of the formula (5) above, which will not be repeated here. At this time, the output of the fully connected layer 709 is the final obtained depression tendency feature vector, which can be represented as U2.
[0129] In some embodiments, when the above psychological feature vector includes a suicide tendency feature vector, referring to Figure 8 The suicide tendency feature vector can be obtained through the following processes 2107 to 2110.
[0130] 2107, according to the word frequency of the suicide tendency feature word of each message in the text content, the messages in the text content are sorted.
[0131] For example, messages in a user's relevant text content can be sorted according to the frequency of suicidal tendencies, from highest to lowest frequency. This places messages with a higher risk of suicidal tendencies at the top, potentially increasing the probability of identifying users with suicidal tendencies.
[0132] See Figure 9 This illustrates another specific example of the network architecture of the psychological feature extraction module 301, where the module can be used to extract the user's suicidal tendency feature vector. Figure 9 Taking user-related text content as the user's historical dialogue context as an example, the sorting module 901 can sort the messages in the user's historical dialogue context from high to low according to the word frequency of suicidal tendency feature words. For example, the first message in the sorted user historical dialogue context contains the highest word frequency of suicidal tendency feature words, the second message contains the second highest word frequency of suicidal tendency feature words, and so on.
[0133] 2108. For each word in the i-th message of the text content, the hidden state of each word is obtained through the fourth GRU unit. Here, i is a positive integer, less than or equal to the total number of messages in the text content.
[0134] See also Figure 9 For the i-th message in the sorted user history dialogue context, the word embedding vector of each word can be extracted through the word embedding layer 902. Then, the extracted word embedding vector of each word is input into the fourth GRU unit, such as GRU units 903, 904, and 905, to obtain the hidden state of each word. For example, the hidden state of each word can be represented as follows:
[0135]
[0136] in, This represents the j-th word in the i-th message; This represents the hidden state of the j-th word in the i-th message; GRU units can be used to capture sequence information and long-term dependencies of sentences to extract contextual features.
[0137] For example, the fourth GRU unit can be a unidirectional GRU unit or a bidirectional GRU unit, and this application does not limit it.
[0138] 2109. The hidden state of each word is weighted and summed using the fourth attention mechanism, and the hidden state of the i-th message is obtained through the fifth GRU unit.
[0139] See alsoFigure 9 The fourth attention mechanism, such as the word attention mechanism 906, can be used on the hidden state of each word of the i-th message to obtain the weight of the hidden state of each word, and the hidden state of each word is weighted and summed according to the weight.
[0140] For example, the fourth attention mechanism can perform the following calculation:
[0141]
[0142] The attention weight of the word hidden state, The normalized weight score of the word hidden state, s i The word sequence of the i-th message The weighted combination of the corresponding hidden states, m represents the number of words of the i-th message, d word The word vector used to learn the word-level attention, b is a bias term, W word The word random initialization weight matrix.
[0143] The s of the i-th message i The hidden state of the i-th message can be obtained by the fifth GRU unit, such as the GRU unit 907. For example, the hidden state of the i-th message h i It can be represented as the following formula:
[0144] r i = GRU (s i ) (11)
[0145] It can be understood that for each message in the sorted user-related content text, the hidden state of each message can be determined based on the above method.
[0146] For example, the fifth GRU unit can be a unidirectional GRU unit or a bidirectional GRU unit, which is not limited in the present application.
[0147] 2110, using the fifth attention mechanism to weight and sum the hidden state of each message of the text content, and obtaining the user's suicide tendency feature vector through a fully connected layer.
[0148] Continuing to refer to Figure 9 For the hidden state of each message in the sorted user historical dialogue context, the fifth attention mechanism, such as the message attention mechanism 908, can be applied to obtain the weight of the hidden state of each message, and the hidden state of each message is weighted and summed according to the weight.
[0149] Exemplarily, the fifth attention mechanism can be calculated as follows:
[0150]
[0151] wherein q i denotes the attention weight of the message hidden state, a j denotes the normalized weight score of the message hidden state, u' denotes the weighted combination of the hidden states corresponding to the messages {s1, s2, …, s n m denotes the number of messages of the text content related to the user, v attention denotes the message vector used to learn the message-level attention, b is a bias term, and W denotes a randomly initialized weight matrix of the messages.
[0152] Continuing to refer to Figure 9 After the weighted sum of the hidden states of each message is performed, the result of the weighted sum can be input into a fully connected layer, such as the fully connected layer 509, to obtain the suicide tendency feature vector of the user.
[0153] Exemplarily, the calculation method of the fully connected layer can refer to the description of the above formula (5), which will not be described here again. At this time, the output of the fully connected layer 909 is the finally obtained suicide tendency feature vector, which can be denoted as U3.
[0154] In some embodiments, after the personality feature vector U1, the depression tendency feature vector U2 and the suicide tendency feature vector U3 of the user are obtained, the personality feature vector U1, the depression tendency feature vector U2 and the suicide tendency feature vector U3 can be spliced to obtain a psychological feature vector of the user, denoted as [U1, U2, U3] T .
[0155] 220, processing the sentence input by the user to obtain a word embedding vector of the sentence.
[0156] Exemplarily, referring to Figure 3 The word embedding vector of the sentence input by the user can be obtained by feature extraction on the sentence input by the user through the dialogue context acquisition model 302. As an example, exemplarily, the dialogue context acquisition model 302 can include BERT, LSTM, CNN or other models, which are not limited in the present application.
[0157] 230, inputting the psychological feature vector and the word embedding vector into the trained first neural network model to obtain a first candidate reply to the sentence. The first neural network model is trained according to the psychological features of the user samples and the human-computer dialogue data samples of the user samples.
[0158] With continued reference to Figure 3 The first neural network model 303 can obtain the candidate reply #1 according to the input psychological feature vector and the word embedding vector.
[0159] In some embodiments, the training sample set of the first neural network model includes psychological features of a user sample and human-computer conversation data samples of the user sample. The human-computer conversation data samples include an input sentence of the user sample and a real label corresponding to the input sentence. The real label corresponding to the input sentence can be a reply sentence corresponding to the input sentence in a historical conversation context between the user and a doctor.
[0160] For example, the psychological features of the user sample and the process of obtaining the psychological feature vector of the user sample can refer to the description in the foregoing, which will not be described here again.
[0161] For example, the human-computer conversation data samples of the user sample can include a historical conversation context between the user and a doctor on a medical health question and answer community website. For example, the medical and patient question and answer data on the related website can be obtained by using a crawler technology to obtain the human-computer conversation data samples of the user sample.
[0162] As a specific example, the conversation content containing the word “spirit” or “psychology” in the department can be data screened, and the conversation with the number of conversation turns not less than 2 times is retained, so that the quality of the training data is as good as possible. In addition, the data samples after the final screening processing can be saved in a json format, and the stored data sample fields can include at least one of a conversation number, a doctor department, a disease description, a conversation content and a conversation turn, which are not limited in the present application.
[0163] In some embodiments, the first neural network model can include an encoder-decoder structure, for example, a Seq2Seq model, which is not limited in the present application. For example, the encoder can be a recurrent neural network (RNN) encoder, and the decoder can be an RNN decoder.
[0164] In some embodiments, the psychological feature vector of the user sample and the word embedding vector of the input sentence can be input into the encoder in the first neural network model to encode the psychological feature vector of the user sample and the word embedding vector of the input sentence. Then, the decoder in the first neural network model can decode according to the hidden state obtained by the encoder and the word embedding vector of the target word at the last time to obtain the hidden state at the current time.
[0165] As a possible implementation manner, the psychological feature vector of the user sample and the word embedding vector of the input sentence can be spliced to obtain a psychological feature-based input vector of the user sample, and the psychological feature-based input vector of the user sample is taken as an input part of an encoder in the first neural network model.
[0166] Referring to Figure 10 , an example diagram of a Seq2Seq model with an attention mechanism is shown, which includes an RNN encoder and an RNN decoder. As shown in Figure 10 , the input word embedding vector obtained by splicing the psychological feature vector of the user sample and the word embedding vector of the input sentence can be input into the RNN encoder to generate a hidden state h t through the RNN encoder. The calculation of the hidden state h t is as follows:
[0167] h t =RNN encoder (u i ,h t-1 ) (13)
[0168] wherein h t represents the hidden state of the encoder at time point t, u i represents the encoded information of the psychological feature-based input vector of the user sample, and h t-1 represents the hidden state of the encoder at time point t-1.
[0169] Continuing to refer to Figure 10 , for the first time point in the RNN decoder, the RNN decoder receives the word embedding vector of the target sentence word (i.e., the real label corresponding to the target sentence) and the hidden state of the encoder at the previous time node, and generates a current hidden state.
[0170] For the second and subsequent time points in the RNN decoder, the RNN decoder receives the word embedding vector of the target sentence word and the context vector of the decoder at the previous time node, and generates a current hidden state, and the calculation formula is as follows:
[0171] s t =RNN decoder (y i ,c t-1 ) (14)
[0172] wherein s t represents the hidden state of the decoder at time t, y i represents the word embedding vector of the target sentence word, and c t-1 represents the context vector of the decoder at the previous time node t-1.
[0173] Then, referring back to Figure 10 , the attention mechanism can obtain a score e ij from the hidden state of the RNN encoder and the hidden state of the RNN decoder by the following equation
[0174] e ij = score(s i , h j ) (15)
[0175] where s i represents the hidden state of the decoder and h j represents the hidden state of the encoder.
[0176] For example, the score function can use the dot product to calculate e ij , i.e.
[0177] score(s i , h j ) = s i T · h j
[0178] Then, the attention mechanism can obtain the corresponding context vector c i from the score in equation (15) above, according to the following equation:
[0179]
[0180] where a ij represents the weighted average of each hidden state in the encoder.
[0181] Referring back to Figure 10 , at the next time point, the context vector c i of the encoder is obtained as the input of the RNN decoder, and the next hidden state of the decoder is obtained by applying equation (14) above. The above process can be repeated in turn until the decoding is completed to obtain the hidden state of the decoder at each time point.
[0182] Then, the context vector c i and the hidden state s t of the decoder can be concatenated to obtain the attention hidden state
[0183]
[0184] where c t represents the context vector at time t, and s tdenotes the hidden state of the decoder at time t, and W denotes a weight matrix.
[0185] Finally, the attention hidden state is input to a fully connected layer, and then normalized by an activation function softmax to obtain a prediction probability p, as shown in the following formula:
[0186]
[0187] During training of the first neural network model, parameters of the first neural network model can be updated according to the prediction probability p and a true label of the target sentence, to obtain a trained first neural network model.
[0188] After the first neural network model is trained, the psychological feature vector of the user and the word embedding vector of the user input sentence can be input into the trained first neural network model to obtain a first candidate reply to the user input sentence, which can be denoted as R1. Here, the first candidate reply is generated based on the psychological feature vector of the user, and fuses the psychological state of the user and the change in psychological activity of the user.
[0189] 240, according to the first candidate reply, determining a final reply to the sentence.
[0190] In some embodiments, the first candidate reply can be taken as the final reply to the user input sentence. Continuing to refer to Figure 3 , the selection module 305 can take the candidate reply #1 as the final reply to the user input sentence.
[0191] Therefore, by inputting the psychological feature vector of the user and the word embedding vector of the input sentence into the trained first neural network model, the embodiments of the present application can obtain a first candidate reply that fuses the psychological features of the user, and then obtain a final reply to the user input sentence according to the first candidate reply. Since the first candidate reply fuses the psychological features of the user, the final reply obtained according to the first candidate reply has good empathic ability, and can meet the psychological needs of the user, thereby helping to improve the humanization degree of human-computer dialogue.
[0192] In some optional embodiments, referring to Figure 11 , the following steps 231 and 241 can also be performed to obtain the final reply to the user input sentence.
[0193] 231, input the word embedding vector of the user input sentence into the trained second neural network model to obtain a second candidate reply for the sentence, wherein the second neural network model uses the same temperature parameter as the teacher model and is trained according to the human-computer conversation data sample of the user sample and the soft target output by the teacher model, and the teacher model is trained according to the human-computer conversation data sample of the user sample.
[0194] Continuing to refer to Figure 3 , the second neural network model 304 can obtain a candidate reply #2 according to the input word embedding vector of the user input sentence.
[0195] Specifically, the second neural network model can generate a reply for the user input sentence based on a knowledge distillation method. A neural network model with an encoder-decoder architecture, such as a Seq2Seq model, is a one-way generation model, which is limited by the constraint of one-way generation ability. However, BERT has bidirectional encoding ability and can well solve the problem of being limited by one-way generation ability. However, a pure BERT model is not conducive to performing a generation task, so a knowledge distillation method is proposed.
[0196] The purpose of knowledge distillation is to force the network response of some intermediate layers of a student model (such as a Seq2Seq model) to approximate the response of the corresponding intermediate layers of a teacher model (such as a BERT model), that is, the teacher model transfers feature knowledge to the student model. After pre-training, the teacher model outputs a normalized soft target after softmax. The purpose of using the teacher model is to make the output of the student model after softmax normalization sufficiently close to the output of the teacher model.
[0197] For example, the teacher model can be fine-tuned to obtain a BERT with bidirectional encoding ability, and the student model can be a Seq2Seq model.
[0198] During fine-tuning of BERT, the sequence pair of the source sentence and the target sentence can be connected to obtain (X, Y), where X represents the source sentence and Y represents the target sentence. Then, a part (such as 15%) of the word groups or words Y m , the target probability where y m is a mask word, belongs to Y m , and Y u is a set of unmasked word groups.
[0199] In the process of knowledge distillation, the teacher model is first trained, and then the student model is trained using the trained teacher model. The human-computer conversation data samples used in training the teacher model and the learning model can refer to the description in the foregoing, which will not be described here.
[0200] In the process of training the teacher model, the sequence set and can be used to estimate the probability of the mask word y m . Wherein represents the unmasked word group or word at the time point before the current time t, represents the unmasked word group or word at the time point after the current time t, and belong to Y u , and N represents the number of words in the target sentence. The prediction probability of y m of the teacher model is The prediction result output by the teacher model is the soft target (logits) obtained by performing softmax normalization calculation on the prediction probability.
[0201] In the process of training the student model, the same temperature parameter as the teacher model can be used to make the output of the student model compared with the soft target output by the teacher model, and constantly approach the soft target output by the teacher model. In the process of approaching the output of the learning model and the soft target output by the teacher model, the backpropagation gradient descent algorithm can be used to constantly update the parameters of the student model.
[0202] Taking the teacher model as BERT and the student model as Seq2Seq model as an example, in order to make the word probability provided by the Seq2Seq model match the word probability provided by the BERT model, the Seq2Seq model can be trained according to the following loss function in the process of knowledge distillation:
[0203] l bidi (θ)=-∑ ω∈V [P φ (y t =ω|Y u ,X)·logP θ (y t =ω|y 1:t-1 ,X)] (19)
[0204] Wherein, P φ (y t ) is the soft target estimated by the BERT model, and its learning parameter is φ, which is fixed; V represents the output vocabulary.
[0205] Further, in order to improve the accuracy of the model prediction, the student model, i.e., the Seq2Seq model, can be trained according to the following composite objective function l(0):
[0206] l(0) = a l bidi (0) + (1-a) l xe (0) (20)
[0207] wherein a is a hyperparameter, l xe (0) is the loss function of the original Seq2Seq model. For example, l xe (0) can be as shown in the following formula:
[0208]
[0209] After the second neural network model is trained, the word embedding vector of the user input sentence can be input into the trained second neural network model to obtain a second candidate reply to the user input sentence, which can be denoted as R2, for example.
[0210] 241, according to the first candidate reply and the second candidate reply, determining a final reply to the user input sentence.
[0211] Continuing to refer to Figure 3 , the selection module 305 can determine a final reply to the user input sentence according to the candidate reply #1 and the candidate reply #2. For example, the selection indicator score of the first candidate reply and the second candidate reply can be determined, and the candidate reply with a higher selection indicator score can be selected as the final reply. The selection indicator score can be, for example, a Distinct score, which is not limited in the present application.
[0212] In some embodiments, the final reply to the user input sentence can be determined according to the second candidate reply. Optionally, at this time, the step or process of determining the first candidate reply can not be performed.
[0213] Therefore, the embodiments of the present application can integrate the indications learned by the teacher model into the dialogue generation model (i.e., the student model) for the user input sentence through knowledge distillation, so that the dialogue generation model can better train the data and improve the accuracy of generating reply text for the user input sentence. Therefore, the dialogue generation model trained based on knowledge distillation in the embodiments of the present application has better generalization effect than directly training a single model.
[0214] In some optional embodiments, referring to Figure 12 , the following steps 232, 233, 234 and 242 can also be performed to obtain the final reply to the user input sentence.
[0215] 232, obtaining a human-computer dialogue data set.
[0216] For example, the human-machine conversation dataset can include historical conversation contexts between users and doctors on a medical health question and answer community website. See Figure 3 The medical conversation data source obtaining model 306 can obtain the human-machine conversation dataset.
[0217] 233, according to the user input sentence, obtain the user's psychological keywords.
[0218] Here, the user's psychological keywords, i.e. the word groups or words that can reflect the user's psychological state, such as the keywords of feeling unwell, unhappy, depressed, etc. For example, the text content of the user input sentence can be extracted to obtain the user's psychological keywords.
[0219] 234, based on the frequency of the psychological keywords in the human-machine conversation dataset, determine the similarity of the user input sentence to the candidate replies in the human-machine conversation dataset, and further determine the third candidate reply for the user input sentence in the candidate replies according to the similarity.
[0220] Specifically, a classical retrieval model can be used to retrieve the third candidate reply for the user input sentence from the human-machine conversation dataset according to the user's psychological keywords. See Figure 3 The retrieval model 307 can retrieve the candidate reply #3 from the human-machine conversation dataset according to the psychological keywords in the user input sentence.
[0221] Specifically, the frequency of the psychological keywords in the user's question can be calculated from the human-machine conversation dataset, denoted as TF q The calculation formula can be as follows:
[0222]
[0223] And the frequency of the psychological keywords in the relevant answers can be calculated from the human-machine conversation dataset, denoted as TF a The calculation formula can be as follows:
[0224]
[0225] At the same time, the inverse frequency of the psychological keywords in the human-machine conversation dataset can be determined, denoted as IDF q The calculation formula can be as follows:
[0226]
[0227] And the inverse frequency of the psychological keywords in the human-machine conversation dataset can also be determined, denoted as IDF a, the calculation formula can be as follows:
[0228]
[0229] Then, the product of TF q and IDF q can be calculated as the representation vector of the user question, the product of TF a and IDF a can be calculated as the representation vector of the answer. The representation vector of the user question and the representation vector of the answer can be used as a representation vector group of the user question and the answer to calculate the similarity of the user question and the answer. Wherein, the closer the representation vector of the user question and the representation vector of the answer are, the higher the similarity of the user question and the answer is. For example, the difference between the representation vector of the user question and the representation vector of the answer can be used as the standard of similarity judgment, that is, the smaller the difference between the two is, the higher the similarity between the two is.
[0230] After obtaining the similarity of the user question and the answer, all question-answer pairs in the human-computer dialogue data set that match the user question can be sorted based on the similarity. Here, the answer in the sorted question-answer pair can be used as the alternative reply. For the user question, a representation vector of the question can be calculated, that is, the product of TF q and IDF q of the psychological word of the question, at this time, an alternative reply with the highest similarity to the representation vector of the question can be selected from the already sorted question-answer pairs as the third candidate reply to the user input sentence, for example, it can be recorded as R3, that is, R3={Q,A}={(Q1,A1),(Q2,A2),…,(Q m ,A m )}, wherein (Q m ,A m ) represents the matched question-answer pair.
[0231] 242, according to the first candidate reply, the second candidate reply and the third candidate reply, determining the final reply.
[0232] Continuing to refer to Figure 3 , the selection module 305 can select the candidate reply #1, the candidate reply #2 and the candidate reply #3 as the final reply to the user input sentence. For example, the selection index scores of the first candidate reply, the second candidate reply and the third candidate reply can be determined, and the candidate reply with a higher selection index score can be selected as the final reply. The selection index score, for example, can be a Distinct score, which is not limited in the present application.
[0233] In some embodiments, the final reply to the sentence of the user input can be determined according to the third candidate reply. Optionally, the steps or processes of determining the first candidate reply or the second candidate reply can not be performed at this time.
[0234] In some embodiments, the final reply to the sentence of the user input can be determined according to the first candidate reply and the third candidate reply. Optionally, the steps or processes of determining the second candidate reply can not be performed at this time.
[0235] Therefore, the embodiments of the present application can help improve the overall performance of the dialogue system by fusing the retrieval-based dialogue generation model in the dialogue generation model, which can help improve the accuracy of human-computer dialogue while also helping to improve the humanization degree of human-computer dialogue, thereby helping to improve the reply quality of the human-computer dialogue model.
[0236] In addition, the hardware of the embodiments of the present application can use an electronic computer with a graphics processing unit (GPU), which can include one or more central processing units (CPUs), one or more GPUs, random access memory (RAM) or Gaussian cache memory, and necessary input and output devices, etc.
[0237] The specific embodiments of the present application are described in detail above in combination with the drawings, but the present application is not limited to the specific details in the above embodiments. Within the technical concept of the present application, various simple modifications can be made to the technical solutions of the present application, and these simple modifications all belong to the protection scope of the present application. For example, in the above specific embodiments, various specific technical features described can be combined in any appropriate manner without contradiction. In order to avoid unnecessary repetition, various possible combination manners are not described again in the present application. For example, various different embodiments of the present application can also be combined in any manner, as long as it does not deviate from the idea of the present application, it should also be considered as disclosed in the present application.
[0238] It should also be understood that in various method embodiments of the present application, the size of the sequence number of the above processes does not mean the order of execution, the execution order of the processes should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. It should be understood that these sequence numbers can be interchanged under appropriate circumstances, so that the described embodiments of the present application can be implemented in an order other than those illustrated or described.
[0239] The method embodiments of the present application are described in detail above, and the device embodiments of the present application are described in detail below in combination with Figures 13 to 14 , detailed description of the device embodiments of the present application.
[0240] Figure 13 is a schematic block diagram of the device 1000 for human-computer conversation provided by an embodiment of the present application. As shown in the figure, the device 1000 for human-computer conversation can include an extraction unit 1010, an acquisition unit 1020, a first neural network model 1030, and a determination unit 1040. Figure 13
[0241] The extraction unit 1010 is configured to extract a psychological feature vector of a user according to text content related to the user.
[0242] The acquisition unit 1020 is configured to process a sentence input by the user and acquire a word embedding vector of the sentence.
[0243] The first neural network model 1030 is configured to input the psychological feature vector and the word embedding vector to obtain a first candidate reply to the sentence, wherein the first neural network model is trained according to a psychological feature of a user sample and human-computer conversation data samples of the user sample.
[0244] The determination unit 1040 is configured to determine a final reply to the sentence according to the first candidate reply.
[0245] In some embodiments, the device 1000 further includes a second neural network model configured to input the word embedding vector to obtain a second candidate reply to the sentence, wherein the second neural network model uses the same temperature parameter as a teacher model, is trained according to human-computer conversation data samples of the user sample and soft targets output by the teacher model, and is obtained according to the human-computer conversation data samples of the user sample, wherein the teacher model is trained according to the human-computer conversation data samples of the user sample.
[0246] The determination unit 1040 is specifically configured to:
[0247] determine the final reply according to the first candidate reply and the second candidate reply.
[0248] In some embodiments, the device 1000 further includes a retrieval model configured to:
[0249] acquire a human-computer conversation data set;
[0250] acquire a psychological term according to a sentence input by the user;
[0251] determine a similarity between the sentence and candidate replies in the human-computer conversation data set based on a frequency of the psychological term in the human-computer conversation data set, and determine a third candidate reply to the sentence among the candidate replies according to the similarity.
[0252] The determining unit 1040 is specifically configured to:
[0253] The final reply is determined according to the first candidate reply, the second candidate reply, and the third candidate reply.
[0254] In some embodiments, the psychological feature vector includes at least one of a personality feature vector, a depression tendency feature vector, and a suicide tendency feature vector of the user.
[0255] In some embodiments, when the psychological feature vector includes the personality feature vector, the extracting unit 1010 is specifically configured to:
[0256] For each word in the i-th message in the text content, a hidden state of each word is obtained by a first GRU unit, where i is a positive integer;
[0257] The hidden states of each word are weighted and summed by a first attention mechanism, and a hidden state of the i-th message is obtained by a second GRU unit;
[0258] The hidden states of each message in the text content are weighted and summed by a second attention mechanism, and a personality feature vector of the user is obtained by a fully connected layer.
[0259] In some embodiments, when the psychological feature vector includes the depression tendency feature vector, the extracting unit 1010 is specifically configured to:
[0260] Extracting a feature word of depression tendency from the text content;
[0261] Obtaining a hidden state of each feature word of the feature word of depression tendency by a third GRU unit;
[0262] The hidden states of each feature word are weighted and summed by a third attention mechanism, and a depression tendency feature vector of the user is obtained by a fully connected layer.
[0263] In some embodiments, the feature word of depression tendency includes at least one of:
[0264] Social network feature words, user emotion feature words, user depression field feature words, and user topic selection feature words.
[0265] In some embodiments, when the psychological feature vector includes the suicide tendency feature vector, the extracting unit 1010 is specifically configured to:
[0266] According to the word frequency of the feature word of the suicide tendency of each message in the text content, the messages in the text content are sorted.
[0267] for each word in the i-th message in the text content, obtaining a hidden state of each word by a fourth GRU unit;
[0268] performing weighted summation on the hidden states of each word by a fourth attention mechanism, and obtaining a hidden state of the i-th message by a fifth GRU unit;
[0269] performing weighted summation on the hidden states of each message in the text content by a fifth attention mechanism, and obtaining a suicide tendency feature vector of the user by a fully connected layer.
[0270] In some embodiments, the first neural network model comprises an encoder-decoder structure.
[0271] In some embodiments, the human-computer conversation data sample comprises human-computer conversation data of a doctor and a patient.
[0272] In some embodiments, the user-related text content comprises at least one of a historical conversation context of human-computer interaction of the user and text content associated with the user on a social media network.
[0273] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, no longer described here. Specifically, the device 1000 of the human-computer conversation in this embodiment can correspond to the corresponding subject of the method 200 of performing the embodiments of the present application, and the foregoing and other operations and / or functions of each module in the device 1000 are respectively for realizing the corresponding processes in the method 200 in the above, and for the sake of brevity, no longer described here.
[0274] The device and system of the embodiments of the present application are described above in the functional module perspective. It should be understood that the functional module can be realized by hardware, or by software instructions, or by a combination of hardware and software modules. Specifically, the steps of the method embodiments in the embodiments of the present application can be completed by the integrated logic circuit of hardware and / or software instructions in the processor, and the steps of the method disclosed in the embodiments of the present application can be directly embodied as hardware decoding processor for execution, or executed by a combination of hardware and software modules in the decoding processor. Alternatively, the software module can be located in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, register, etc. The storage medium is located in the memory, and the processor reads the information in the memory, and combines the hardware to complete the steps in the above method embodiments.
[0275] As Figure 14is a schematic block diagram of an electronic device 1100 provided by embodiments of the present application.
[0276] As shown in Figure 14 The electronic device 1100 can include:
[0277] The memory 1110 and the processor 1120, the memory 1110 is used to store computer programs, and the program code is transmitted to the processor 1120. In other words, the processor 1120 can call and run the computer program from the memory 1110 to realize the method in the embodiments of the present application.
[0278] For example, the processor 1120 can be used to execute the steps in the above-mentioned method 200 according to the instructions in the computer program.
[0279] In some embodiments of the present application, the processor 1120 can include but is not limited to:
[0280] General processor, digital signal processor (Digital Signal Processor, DSP), application specific integrated circuit (Application Specific Integrated Circuit, ASIC), field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc.
[0281] In some embodiments of the present application, the memory 1110 includes but is not limited to:
[0282] The non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM) used as an external cache. By way of example, and not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synch link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0283] In some embodiments of the present application, the computer program can be divided into one or more modules, which are stored in the memory 1110 and executed by the processor 1120 to complete the method provided by the present application. The one or more modules can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the electronic device 1100.
[0284] Optionally, as shown in Figure 14 The electronic device 1100 can further include a transceiver 1130, which can be connected to the processor 1120 or the memory 1110. The processor 1120 can control the transceiver 1130 to communicate with other devices, specifically, can send information or data to other devices, or receive information or data sent by other devices. The transceiver 1130 can include a transmitter and a receiver. The transceiver 1130 can further include an antenna, and the number of antennas can be one or more.
[0285] It should be understood that the various components in the electronic device 1100 are connected through a bus system, wherein the bus system includes a data bus, a power supply bus, a control bus, and a state signal bus in addition to the data bus.
[0286] According to an aspect of the present application, a communication apparatus is provided, comprising a processor and a memory for storing a computer program, the processor is configured to invoke and run the computer program stored in the memory, so that the encoder performs the method of the above-mentioned method embodiments.
[0287] According to an aspect of the present application, a computer storage medium is provided, which stores a computer program, the computer program is executed by a computer to enable the computer to perform the method of the above-mentioned method embodiments. Alternatively, the embodiments of the present application also provide a computer program product comprising instructions, the instructions are executed by a computer to perform the method of the above-mentioned method embodiments.
[0288] According to another aspect of the present application, a computer program product or computer program is provided, which comprises computer instructions stored in a computer readable storage medium. The processor of the computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the computer device to perform the method of the above-mentioned method embodiments.
[0289] In other words, when implemented by using software, the embodiments of the present application can be implemented in the form of a computer program product, entirely or partially. The computer program product comprises one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the flow or function of the embodiments of the present application is entirely or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer readable storage medium, or transferred from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (for example, coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example, infrared, wireless, microwave, etc.) manner. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media sets. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, digital video disc (DVD)), or a semiconductor medium (for example, solid state disk (SSD)) and the like.
[0290] It can be understood that in the specific embodiments of the present application, data related to user information and the like can be involved. When the above embodiments of the present application are applied to specific products or technologies, the user's permission or consent needs to be obtained, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of countries and regions.
[0291] Those skilled in the art can understand that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0292] In several embodiments provided in the present application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely schematic, for example, the division of the modules is only a logical function division, and actual implementation can have another division manner, for example, a plurality of modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed ones can be indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.
[0293] The modules described as separate components can or can not be physically separated, and the components displayed as modules can or can not be physical modules, i.e. can be located in one place or can be distributed on a plurality of network units. Part or all of the modules can be selected to achieve the purpose of the embodiments according to actual needs. For example, the functional modules in each embodiment of the present application can be integrated in one processing module, or each module can be physically present separately, or two or more modules can be integrated in one module.
[0294] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method of human-machine dialog, characterized in that, The method comprises: extracting a psychological feature vector of the user according to text content related to the user; the psychological feature vector comprises at least one of a personality feature vector, a depression tendency feature vector and a suicide tendency feature vector of the user; processing a sentence input by the user to obtain a word embedding vector of the sentence; inputting the psychological feature vector and the word embedding vector into a trained first neural network model to obtain a first candidate reply to the sentence; the first neural network model is trained according to a psychological feature of a user sample and human-computer conversation data samples of the user sample; determining a final reply to the sentence according to the first candidate reply; The method further comprises: obtaining a human-computer conversation data set; obtaining a psychological keyword according to the sentence input by the user; 2. The method of claim 1, wherein, determining a similarity between the sentence and a candidate reply in the human-computer conversation data set based on a frequency of the psychological keyword in the human-computer conversation data set, and determining a third candidate reply to the sentence from the candidate replies according to the similarity; The method further comprises: determining the final reply according to the first candidate reply, the second candidate reply and the third candidate reply.
4. The method of any one of claims 1-3, when the psychological feature vector comprises the personality feature vector, the method further comprises:
3. The method of claim 2, wherein, for each word in an ith message in the text content, obtaining a hidden state of each word by a first GRU unit, wherein i is a positive integer; performing weighted summation on the hidden state of each word by a first attention mechanism, and obtaining a hidden state of the ith message by a second GRU unit; performing weighted summation on the hidden state of each message in the text content by a second attention mechanism, and obtaining the personality feature vector of the user by a full connection layer. 5. The method according to any one of claims 1 to 3, characterized in that, In a case that the psychological feature vector comprises the depression tendency feature vector, the extracting the psychological feature vector of the user according to the text content related to the user comprises: extracting a feature word of depression tendency from the text content; obtaining a hidden state of each feature word of the feature word of depression tendency through a third GRU unit; performing weighted summation on the hidden state of each feature word by using a third attention mechanism, and obtaining the depression tendency feature vector of the user through a full connection layer.
6. The method of claim 5, wherein, The feature word of depression tendency comprises at least one of the following: a social network feature word, a user emotion feature word, a user depression field feature word, and a user topic selection feature word.
7. The method according to any one of claims 1 to 3, characterized in that, In a case that the psychological feature vector comprises the suicide tendency feature vector, the extracting the psychological feature vector of the user according to the text content related to the user comprises: sorting messages in the text content according to a word frequency of a feature word of suicide tendency of each message in the text content; obtaining a hidden state of each word in the i-th message in the text content through a fourth GRU unit, wherein i is a positive integer; performing weighted summation on the hidden state of each word by using a fourth attention mechanism, and obtaining a hidden state of the i-th message through a fifth GRU unit; performing weighted summation on the hidden state of each message in the text content by using a fifth attention mechanism, and obtaining the suicide tendency feature vector of the user through a full connection layer.
8. The method according to any one of claims 1 to 3, characterized in that, The first neural network model comprises an encoder-decoder structure.
9. The method according to any one of claims 1 to 3, characterized in that, The human-computer conversation data sample comprises human-computer conversation data of a doctor and a patient.
10. The method according to any one of claims 1 to 3, characterized in that, The text content related to the user comprises at least one of a historical conversation context of human-computer interaction of the user and text content associated with the user on a social media network.
11. An apparatus for human-to-computer dialog, the apparatus comprising: comprise: an extracting unit configured to extract a psychological feature vector of a user according to text content related to the user; the psychological feature vector comprises at least one of a personality feature vector of the user, a depression tendency feature vector, and a suicide tendency feature vector; an obtaining unit configured to process a sentence input by the user, and obtain a word embedding vector of the sentence; a first neural network model configured to input the psychological feature vector and the word embedding vector, and obtain a first candidate reply to the sentence; the first neural network model is trained according to a psychological feature of a user sample and human-computer conversation data of the user sample; a determining unit configured to determine a final reply to the sentence according to the first candidate reply. The extracting unit is specifically configured to: obtain a hidden state of a word or a word in the text content through a gated recurrent unit (GRU); perform weighted summation on the hidden state of the word or the word by using an attention mechanism, and obtain the psychological feature vector of the user through a full connection layer.
12. An electronic device, comprising: The device comprises a processor and a memory, the memory stores instructions, and the processor executes the instructions to perform the method in any one of claims 1-10. The device comprises a processor and a memory, the memory stores instructions, and the processor executes the instructions to perform the method in any one of claims 1-10.
13. A computer storage medium, characterized in that A computer program product for storing a computer program comprising computer program code means adapted to perform the method of any of claims 1-10 when the computer program is run on an electronic device.
14. A computer program product, characterised in that, A computer program comprising computer program code means adapted to perform the method of any of claims 1-10 when the computer program is run on an electronic device.
Citation Information
Patent Citations
Intelligent dialogue generation method and device, computer equipment and computer storage medium
CN110990543A