Dialogue State Determination Method, Device, Computer Equipment and Storage Medium

By weighted fusion processing of word vector sets of multiple dialogue rounds, the dialogue state is determined, and the problem of inaccurate dialogue state recognition in the prior art is solved, and higher accuracy is achieved.

CN111581958BActive Publication Date: 2025-07-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202010462203.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-05-27
Publication Date
2025-07-18
Estimated Expiration
2040-05-27

AI Technical Summary

Technical Problem

In the prior art, the dialogue state determination method relies on a small number of dialogue statements, resulting in inaccurate status recognition.

Method used

By obtaining dialogue sentences for multiple dialogue rounds, using word vector sets and preset dimension identifications for weighted fusion processing, the dialogue state of the target dialogue round is determined, including similarity analysis and keyword extraction of word vector sets.

Benefits of technology

The accuracy of dialogue state is improved, and the accuracy of dialogue state recognition is enhanced by integrating the amount of information from multiple rounds of dialogue statements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111581958B_ABST
    Figure CN111581958B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses a method, device, computer device, and storage medium for determining a dialogue state, belonging to the field of computer technology. The method includes: obtaining dialogue statements of multiple dialogue turns, obtaining a set of word vectors for each dialogue turn, obtaining a first fused feature vector corresponding to each preset dimension identifier according to the dimension vectors of each preset dimension identifier, determining at least one target dimension identifier corresponding to the target dialogue turn, and target keywords corresponding to each target dimension identifier, and determining at least one target dimension identifier and the target keywords corresponding to each target dimension identifier as the dialogue state of the target dialogue turn. By fusing the dialogue statements of the target dialogue turn and the historical dialogue turns, the information volume of the dialogue statements is enriched, and the set of word vectors of multiple dialogue turns is weighted and fused according to each preset dimension identifier, thereby improving the accuracy of the obtained dialogue state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technologies, and particularly to a method, apparatus, computer device, and storage medium for determining a dialogue state. Background Art

[0002] With the development of computer technologies, dialogue systems have been increasingly widely used. Through dialogue statements in a dialogue system, a dialogue state can be identified, such as a hotel name dimension identifier and a hotel name, a hotel address dimension identifier and a hotel address, and corresponding services can be provided to a user subsequently.

[0003] In related technologies, a method for determining a dialogue state is provided. Any dialogue statement in a dialogue turn, including at least one of a question statement and a reply statement, is obtained, and the dialogue state of this dialogue turn is determined by performing state recognition on the dialogue statement in this dialogue turn. Since few dialogue statements are used in the state recognition process, the determined dialogue state is not accurate enough. Summary of the Invention

[0004] Embodiments of the present application provide a method, apparatus, computer device, and storage medium for determining a dialogue state, which can improve the accuracy of the determined dialogue state. The technical solution is as follows:

[0005] On the one hand, a method for determining a dialogue state is provided. The method includes:

[0006] Obtain dialogue statements of multiple dialogue turns, where the dialogue statement of each dialogue turn includes at least one of a question statement or a reply statement, and the multiple dialogue turns include a target dialogue turn and at least one historical dialogue turn before the target dialogue turn;

[0007] According to the dialogue statement of each dialogue turn, obtain a word vector set of each dialogue turn, where the word vector set of the dialogue turn includes word vectors of each word in the dialogue statement of the dialogue turn;

[0008] According to the similarity between the dimension vector of each preset dimension identifier and each word vector in the word vector sets of the multiple dialogue turns, perform weighted fusion processing on the word vector sets of the multiple dialogue turns to obtain a first fusion feature vector corresponding to each preset dimension identifier;

[0009] According to the similarity between the word vector of the preset keyword corresponding to each preset dimension identifier and the corresponding first fusion feature vector, determine at least one target dimension identifier corresponding to the target dialogue turn, and a target keyword corresponding to each target dimension identifier;

[0010] Determine the at least one target dimension identifier and the target keyword corresponding to each target dimension identifier as the dialogue state of the target dialogue turn.

[0011] In a possible implementation manner, for each preset dimension identifier, according to the similarity between the dimension vector of the preset dimension identifier and multiple word vectors in the word vector set of each dialogue turn, perform weighted fusion processing on the multiple word vectors in the word vector set of each dialogue turn to obtain the first feature vector of each dialogue turn, including:

[0012] For each preset dimension identifier and each dialogue turn, respectively determine the first similarity between the dimension vector of the preset dimension identifier and each word vector in the word vector set of the dialogue turn;

[0013] According to the first similarity corresponding to each word vector in the word vector set of the dialogue turn, determine the first weight of each word vector, and the first weight of each word vector is positively correlated with the corresponding first similarity;

[0014] Perform weighted fusion processing on the multiple word vectors according to the first weights of the multiple word vectors in the word vector set to obtain the first feature vector of the dialogue turn.

[0015] In another possible implementation manner, according to the similarity between the dimension vector of the preset dimension identifier and the first feature vectors of the multiple dialogue turns, perform weighted fusion processing on the first feature vectors of the multiple dialogue turns to obtain the second fusion feature vector corresponding to the preset dimension identifier, including:

[0016] Respectively determine the second similarity between the dimension vector of the preset dimension identifier and the first feature vector of each dialogue turn;

[0017] According to the second similarity corresponding to each dialogue turn, determine the second weight of each dialogue turn, and the second weight of each dialogue turn is positively correlated with the corresponding second similarity;

[0018] Perform weighted fusion processing on the first feature vectors of the multiple dialogue turns according to the second weights of the multiple dialogue turns to obtain the second fusion feature vector corresponding to the preset dimension identifier.

[0019] In another possible implementation manner, according to the third similarity between the first feature vector of any dialogue turn and the first feature vectors of the multiple dialogue turns, perform weighted fusion processing on the word vector sets of the multiple dialogue turns, and use the feature vector after the fusion processing as the adjusted first feature vector of any dialogue turn, including:

[0020] Determine the third weight of each conversation turn in the multiple conversation turns according to the third similarity between the first feature vector of any one of the conversation turns and the first feature vectors of the multiple conversation turns. The third weight corresponding to each conversation turn in the multiple conversation turns has a positive correlation with the corresponding third similarity;

[0021] Perform weighted fusion processing on the first feature vectors of the multiple conversation turns according to the third weights corresponding to the multiple conversation turns, and use the feature vector after the fusion processing as the adjusted first feature vector of any one of the conversation turns.

[0022] In another possible implementation manner, the fusing the second fusion feature vector with the first feature vector of the target conversation turn to obtain the first fusion feature vector corresponding to the preset dimension identifier includes:

[0023] Determine the fourth weight of the first feature vector of the target conversation turn and the fifth weight of the second fusion feature vector, and the sum of the fourth weight and the fifth weight is 1;

[0024] Perform weighted fusion processing on the second fusion feature vector and the first feature vector of the target conversation turn according to the fourth weight and the fifth weight to obtain the first fusion feature vector corresponding to the preset dimension identifier.

[0025] On the other hand, a method for determining a conversation state is provided. The method includes:

[0026] Obtain the conversation statements of multiple conversation turns. The conversation statement of each conversation turn includes at least one of a question statement or a reply statement. The multiple conversation turns include a target conversation turn and at least one historical conversation turn before the target conversation turn;

[0027] Call the word vector acquisition sub-model in the conversation state determination model, and obtain the word vector set of each conversation turn according to the conversation statement of each conversation turn. The word vector set of the conversation turn includes the word vectors of each word in the conversation statement of the conversation turn;

[0028] Call the feature vector acquisition sub-model in the conversation state determination model, and perform weighted fusion processing on the word vector sets of the multiple conversation turns according to the similarity between the dimension vector of each preset dimension identifier and each word vector in the word vector sets of the multiple conversation turns, to obtain the first fusion feature vector corresponding to each preset dimension identifier;

[0029] Invoke the dialogue state determination sub-model in the said dialogue state determination model, and determine at least one target dimension identifier corresponding to the target dialogue turn and the target keyword corresponding to each target dimension identifier according to the similarity between the word vectors of the preset keywords corresponding to each preset dimension identifier and the corresponding first fusion feature vectors; determine the at least one target dimension identifier and the target keyword corresponding to each target dimension identifier as the dialogue state of the target dialogue turn.

[0030] In a possible implementation manner, before invoking the second attention layer in the said feature vector acquisition sub-model and performing weighted fusion processing on the first feature vectors of the multiple dialogue turns according to the similarity between the dimension vectors of the preset dimension identifiers and the first feature vectors of the multiple dialogue turns to obtain the second fusion feature vectors corresponding to the preset dimension identifiers, the method further includes:

[0031] Invoke the third attention layer in the said feature vector acquisition sub-model to obtain the position vector of each dialogue turn, where the position vector of the dialogue turn is used to represent the position of the dialogue turn in the multiple dialogue turns; perform fusion processing on the first feature vector of each dialogue turn and the corresponding position vector, and use the feature vector after the fusion processing as the adjusted first feature vector of each dialogue turn.

[0032] On the other hand, a dialogue state determination device is provided, and the device includes:

[0033] A dialogue statement acquisition module, configured to acquire dialogue statements of multiple dialogue turns, where the dialogue statement of each dialogue turn includes at least one of a question statement or a reply statement, and the multiple dialogue turns include a target dialogue turn and at least one historical dialogue turn before the target dialogue turn;

[0034] A first set acquisition module, configured to acquire a word vector set of each dialogue turn according to the dialogue statement of each dialogue turn, where the word vector set of the dialogue turn includes the word vectors of each word in the dialogue statement of the dialogue turn;

[0035] A first fusion processing module, configured to perform weighted fusion processing on the word vector sets of the multiple dialogue turns according to the similarity between the dimension vectors of each preset dimension identifier and each word vector in the word vector sets of the multiple dialogue turns to obtain the first fusion feature vectors corresponding to each preset dimension identifier;

[0036] A first determination module, configured to determine at least one target dimension identifier corresponding to the target conversation turn and a target keyword corresponding to each target dimension identifier according to the similarity between the word vector of the preset keyword corresponding to each preset dimension identifier and the corresponding first fusion feature vector;

[0037] The first determination module is further configured to determine the at least one target dimension identifier and the target keyword corresponding to each target dimension identifier as the conversation state of the target conversation turn.

[0038] In a possible implementation manner, the first fusion processing module includes:

[0039] A first fusion processing unit, configured to perform weighted fusion processing on multiple word vectors in the word vector set of each conversation turn respectively according to the similarity between the dimension vector of the preset dimension identifier and the multiple word vectors in the word vector set of each conversation turn, to obtain a first feature vector of each conversation turn;

[0040] A second fusion processing unit, configured to perform weighted fusion processing on the first feature vectors of the multiple conversation turns according to the similarity between the dimension vector of the preset dimension identifier and the first feature vectors of the multiple conversation turns, to obtain a second fusion feature vector corresponding to the preset dimension identifier;

[0041] A third fusion processing unit, configured to perform fusion processing on the second fusion feature vector and the first feature vector of the target conversation turn to obtain a first fusion feature vector corresponding to the preset dimension identifier.

[0042] In another possible implementation manner, the first fusion processing unit is configured to respectively determine a first similarity between the dimension vector of the preset dimension identifier and each word vector in the word vector set of the conversation turn for each preset dimension identifier and each conversation turn; determine a first weight of each word vector according to the first similarity corresponding to each word vector in the word vector set of the conversation turn, and the first weight of each word vector is positively correlated with the corresponding first similarity; and perform weighted fusion processing on the multiple word vectors according to the first weights of the multiple word vectors in the word vector set to obtain a first feature vector of the conversation turn.

[0043] In another possible implementation, the second fusion processing unit is configured to respectively determine a second similarity between the dimension vector of the preset dimension identifier and the first feature vector of each conversation turn; determine a second weight for each conversation turn according to the second similarity corresponding to each conversation turn, and the second weight of each conversation turn is positively correlated with the corresponding second similarity; and perform weighted fusion processing on the first feature vectors of the multiple conversation turns according to the second weights of the multiple conversation turns to obtain a second fusion feature vector corresponding to the preset dimension identifier.

[0044] In another possible implementation, the apparatus further includes:

[0045] A first adjustment module, configured to perform weighted fusion processing on the first feature vectors of the multiple conversation turns according to a third similarity between the first feature vector of any conversation turn and the first feature vectors of the multiple conversation turns, and use the feature vector after the fusion processing as the adjusted first feature vector of the any conversation turn.

[0046] In another possible implementation, the first adjustment module includes:

[0047] A weight determination unit, configured to determine a third weight for each conversation turn among the multiple conversation turns according to the third similarity between the first feature vector of the any conversation turn and the first feature vectors of the multiple conversation turns, and the third weight corresponding to each conversation turn among the multiple conversation turns is positively correlated with the corresponding third similarity;

[0048] A fourth fusion processing unit, configured to perform weighted fusion processing on the first feature vectors of the multiple conversation turns according to the third weights corresponding to the multiple conversation turns, and use the feature vector after the fusion processing as the adjusted first feature vector of the any conversation turn.

[0049] In another possible implementation, the apparatus further includes:

[0050] A position vector acquisition module, configured to acquire a position vector of each conversation turn, and the position vector of the conversation turn is used to represent the position of the conversation turn in the multiple conversation turns;

[0051] A second fusion processing module, configured to perform fusion processing on the first feature vector of each conversation turn and the corresponding position vector, and use the feature vector after the fusion processing as the adjusted first feature vector of each conversation turn.

[0052] In another possible implementation, the third fusion processing unit is configured to determine a fourth weight of the first feature vector of the target dialogue turn and a fifth weight of the second fusion feature vector, and the sum of the fourth weight and the fifth weight is; according to the fourth weight and the fifth weight, perform weighted fusion processing on the second fusion feature vector and the first feature vector of the target dialogue turn to obtain a first fusion feature vector corresponding to the preset dimension identifier.

[0053] In another possible implementation, the first determination module includes:

[0054] The similarity determination unit is configured to determine, for each preset dimension identifier, a fourth similarity between the word vectors of a plurality of preset keywords corresponding to the preset dimension identifier and the first fusion feature vector;

[0055] The keyword selection unit is configured to select alternative keywords from the plurality of preset keywords, and the fourth similarity corresponding to the alternative keywords is greater than the fourth similarity corresponding to other preset keywords in the plurality of preset keywords;

[0056] The target determination unit is configured to, in response to the alternative keyword not being an empty keyword, determine the preset dimension identifier as the target dimension identifier and determine the alternative keyword as the target keyword.

[0057] On the other hand, a dialogue state determination device is provided, and the device includes:

[0058] The dialogue statement acquisition module is configured to acquire dialogue statements of a plurality of dialogue turns, and the dialogue statement of each dialogue turn includes at least one of a question statement or a reply statement, and the plurality of dialogue turns include a target dialogue turn and at least one historical dialogue turn before the target dialogue turn;

[0059] The first set acquisition module is configured to call a word vector acquisition sub-model in the dialogue state determination model, and acquire a word vector set of each dialogue turn according to the dialogue statement of each dialogue turn;

[0060] The first fusion processing module is configured to call a feature vector acquisition sub-model in the dialogue state determination model, and perform weighted fusion processing on the word vector sets of the plurality of dialogue turns according to the similarity between the dimension vector of each preset dimension identifier and each word vector in the word vector sets of the plurality of dialogue turns to obtain a first fusion feature vector corresponding to each preset dimension identifier;

[0061] A first determination module, configured to call a dialogue state determination sub-model in the dialogue state determination model, and determine at least one target dimension identifier corresponding to the target dialogue turn and a target keyword corresponding to each target dimension identifier according to the similarity between the word vector of the preset keyword corresponding to each preset dimension identifier and the corresponding first fusion feature vector; and determine the at least one target dimension identifier and the target keyword corresponding to each target dimension identifier as the dialogue state of the target dialogue turn.

[0062] In a possible implementation manner, the first fusion processing module includes:

[0063] A first fusion processing unit, configured to call a first attention layer in the feature vector acquisition sub-model, and perform weighted fusion processing on multiple word vectors in the word vector set of each dialogue turn according to the similarity between the dimension vector of the preset dimension identifier and the multiple word vectors in the word vector set of each dialogue turn, to obtain a first feature vector of each dialogue turn;

[0064] A second fusion processing unit, configured to call a second attention layer in the feature vector acquisition sub-model, and perform weighted fusion processing on the first feature vectors of the multiple dialogue turns according to the similarity between the dimension vector of the preset dimension identifier and the first feature vectors of the multiple dialogue turns, to obtain a second fusion feature vector corresponding to the preset dimension identifier;

[0065] A third fusion processing unit, configured to call a feature fusion layer in the feature vector acquisition sub-model, and perform fusion processing on the second fusion feature vector and the first feature vector of the target dialogue turn, to obtain a first fusion feature vector corresponding to the preset dimension identifier.

[0066] In another possible implementation manner, the apparatus further includes:

[0067] A first adjustment module, configured to call a third attention layer in the feature vector acquisition sub-model, and perform weighted fusion processing on the word vector sets of the multiple dialogue turns according to a third similarity between the first feature vector of any dialogue turn and the first feature vectors of the multiple dialogue turns, and use the fused feature vector as the first feature vector after adjustment of the dialogue turn.

[0068] In another possible implementation manner, the apparatus further includes:

[0069] A second adjustment module, configured to call a third attention layer in the feature vector acquisition sub-model to obtain a position vector for each conversation turn, where the position vector of the conversation turn is used to represent the position of the conversation turn among the multiple conversation turns; perform a fusion process on the first feature vector of each conversation turn and the corresponding position vector, and use the feature vector after the fusion process as the adjusted first feature vector of each conversation turn.

[0070] In another possible implementation manner, the apparatus further includes:

[0071] The conversation statement acquisition module is further configured to obtain sample conversation statements of multiple sample conversation turns, and a sample conversation state corresponding to each sample conversation turn. The sample conversation statement of each sample conversation turn includes at least one of a sample user question statement or a sample reply statement. The multiple sample conversation turns include a target sample conversation turn and at least one historical sample conversation turn before the target sample conversation turn;

[0072] The first set acquisition module is further configured to call the word vector acquisition sub-model, and obtain a word vector set for each sample conversation turn according to the sample conversation statement of each sample conversation turn. The word vector set of the sample conversation turn includes word vectors of each word in the sample conversation statement of the sample conversation turn;

[0073] The first fusion processing module is further configured to call the feature vector acquisition sub-model, and perform a weighted fusion process on the word vector sets of the target sample conversation turn and the corresponding historical sample conversation turns according to the similarity between the dimension vectors of each preset dimension identifier and each word vector in the word vector sets of the target sample conversation turn and the corresponding historical sample conversation turns, to obtain a first sample fusion feature vector corresponding to each preset dimension identifier;

[0074] The first determination module is further configured to call the conversation state determination sub-model, and determine at least one predicted dimension identifier corresponding to the target sample conversation turn and a predicted keyword corresponding to each predicted dimension identifier according to the word vectors of multiple preset keywords corresponding to each preset dimension identifier and the corresponding first sample fusion feature vector; determine the at least one predicted dimension identifier and the predicted keyword corresponding to each predicted dimension identifier as the predicted conversation state of the target sample conversation turn;

[0075] A model adjustment module, configured to adjust the word vector acquisition sub-model, the feature vector acquisition sub-model, and the conversation state determination sub-model according to the predicted conversation state of the target sample conversation turn and the corresponding sample conversation state.

[0076] In another possible implementation, the apparatus further includes:

[0077] A second determination module, configured to use any one of the multiple sample conversation turns as the target sample conversation turn, and use the sample conversation turns before the any one sample conversation turn as the historical sample conversation turns of the target sample conversation turn.

[0078] In another possible implementation, the apparatus includes:

[0079] A change probability acquisition module, configured to acquire multiple sample change probabilities for each preset dimension identifier, where the sample change probability represents the difference between the sample conversation states corresponding to every two adjacent sample conversation turns among the multiple sample conversation turns;

[0080] A feature vector processing module, configured to, for each preset dimension identifier, call a state transition model to respectively process the first fusion feature vectors of every two adjacent sample conversation turns corresponding to the preset dimension identifier, to obtain multiple predicted change probabilities corresponding to the preset dimension identifier, and there is a one-to-one correspondence between the multiple sample change probabilities and the multiple predicted change probabilities corresponding to the same preset dimension identifier;

[0081] The model adjustment module includes:

[0082] A model adjustment unit, configured to adjust the word vector acquisition sub-model, the feature vector acquisition sub-model, the conversation state determination sub-model, and the called state transition model according to the obtained multiple predicted change probabilities for each preset dimension identifier, the multiple sample change probabilities for each preset dimension identifier, the predicted conversation state of each sample conversation turn, and the corresponding sample conversation state.

[0083] On the other hand, a computer device is provided, where the computer device includes a processor and a memory, and at least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the conversation state determination method as described in the above aspect.

[0084] On the other hand, a computer-readable storage medium is provided, where at least one instruction is stored in the computer-readable storage medium, and the at least one instruction is loaded and executed by a processor to implement the conversation state determination method as described in the above aspect.

[0085] The beneficial effects brought by the technical solution provided in the embodiments of the present application at least include:

[0086] The method, apparatus, computer device, and storage medium provided by the embodiments of the present application obtain the dialogue statements of multiple dialogue turns, obtain the word vector sets of each dialogue turn, perform weighted fusion processing on the word vector sets of multiple dialogue turns according to the dimension vectors identified by each preset dimension, obtain the first fusion feature vectors corresponding to each preset dimension identifier, determine at least one target dimension identifier corresponding to the target dialogue turn, and the target keywords corresponding to each target dimension identifier, and determine at least one target dimension identifier and the target keywords corresponding to each target dimension identifier as the dialogue state of the target dialogue turn. By fusing the dialogue statements of the target dialogue turn and the historical dialogue turns, the information volume of the dialogue statements is enriched, and the word vector sets of multiple dialogue turns are weighted and fused according to each preset dimension identifier, improving the accuracy of the first fusion feature vectors of each preset dimension identifier, thereby improving the accuracy of the obtained dialogue state. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0088] Figure 1 is a flowchart of a method for determining dialogue state provided by an embodiment of the present application;

[0089] Figure 2 is a flowchart of a method for determining dialogue state provided by an embodiment of the present application;

[0090] Figure 3 is a flowchart of a method for determining dialogue state provided by an embodiment of the present application;

[0091] Figure 4 is a schematic diagram of the weights of a dialogue turn provided by an embodiment of the present application;

[0092] Figure 5 is a method for training a dialogue state determination model provided by an embodiment of the present application;

[0093] Figure 6 is a flowchart of a method for training a dialogue state determination model provided by an embodiment of the present application;

[0094] Figure 7 is a flowchart of a method for training a dialogue state determination model provided by an embodiment of the present application;

[0095] Figure 8 is a schematic structural diagram of a dialogue state determination device provided by an embodiment of the present application;

[0096] Figure 9 It is a schematic structural diagram of a dialogue state determination device provided by an embodiment of the present application;

[0097] Figure 10 It is a schematic structural diagram of a dialogue state determination device provided by an embodiment of the present application;

[0098] Figure 11 It is a schematic structural diagram of a dialogue state determination device provided by an embodiment of the present application;

[0099] Figure 12 It is a schematic structural diagram of a terminal provided by an embodiment of the present application;

[0100] Figure 13 It is a schematic structural diagram of a server provided by an embodiment of the present application. Detailed implementation manners

[0101] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0102] The terms "first", "second", etc. used in the present application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the present application, the first image may be referred to as the second image, and similarly, the second image may be referred to as the first image.

[0103] The terms "at least one", "multiple", "each", "any one" used in the present application, at least one includes one, two, or more than two, multiple includes two or more than two, and each refers to each one of the corresponding multiple, and any one refers to any one of the multiple. For example, multiple elements include 3 elements, and each refers to each of these 3 elements, and any one refers to any one of these 3 elements, which may be the first, the second, or the third.

[0104] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning, and decision-making.

[0105] Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0106] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph, and other technologies.

[0107] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.

[0108] The solution provided by the embodiments of this application, based on the machine learning technology of artificial intelligence, can train a word vector acquisition sub-model, a feature vector acquisition sub-model, and a dialogue state determination sub-model. The training process can also be combined with the invocation of a state transition model. Then, using the trained word vector acquisition sub-model, feature vector acquisition sub-model, and dialogue state determination sub-model, the dialogue state is determined.

[0109] The dialogue state determination method provided by the embodiments of the present application can be used in a computer device, which includes a terminal or a server. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this.

[0110] The method provided by the embodiments of the present application can be used in the scenario of intelligent dialogue.

[0111] For example, in the scenario of intelligent dialogue:

[0112] After the computer device obtains the dialogue statements of multiple dialogue turns, it uses the dialogue state determination method provided by the embodiments of the present application to determine the dialogue state of the target dialogue turn. Subsequently, according to this dialogue state, it generates a reply statement corresponding to the question statement of the user, and provides the reply statement to the user, realizing an intelligent dialogue of human-computer interaction.

[0113] Another example, in the scenario of intelligent translation:

[0114] After the computer device obtains the dialogue statements of multiple dialogue turns, it uses the dialogue state determination method provided by the embodiments of the present application to determine the dialogue state of the target dialogue turn. Subsequently, according to this dialogue state, it determines the translation statement of the question statement sent by the user, and provides the translation statement to the user, realizing an intelligent translation of human-computer interaction.

[0115] Figure 1 It is a flowchart of a dialogue state determination method provided by the embodiments of the present application, which is applied to a computer device, as Figure 1 shown, and the method includes:

[0116] 101. The computer device obtains the dialogue statements of multiple dialogue turns.

[0117] Among them, the dialogue statement of each dialogue turn includes at least one of a question statement or a reply statement. The question statement is a statement input by the user, and the reply statement is a statement replied by the dialogue system to the question statement. In the dialogue statement of each dialogue turn, it may only include the user statement, or may include the question statement and the reply statement, or may only include the reply statement.

[0118] Since during the interaction between the user and the dialogue system, the user inputs one sentence and the dialogue system replies with one sentence, when determining the dialogue turn, the question sentence input by the user and the subsequent reply sentence can be regarded as one dialogue turn, and each dialogue turn includes a question sentence and a reply sentence; or the reply sentence and the subsequent question sentence can be regarded as one dialogue turn, and the first dialogue turn may only include the question sentence. Starting from the second dialogue turn, each dialogue turn includes a reply sentence and a question sentence.

[0119] For example, when the user interacts with the dialogue system, the user inputs a question sentence and the dialogue system replies with a reply sentence, then the question sentence and the reply sentence are regarded as the dialogue sentences of one dialogue turn; or the user inputs the first question sentence, the system replies with the first reply sentence, and when the user inputs the second question sentence, the first question sentence is regarded as one dialogue turn, and the first reply sentence and the second question sentence are regarded as one dialogue turn.

[0120] Multiple dialogue turns include the target dialogue turn and at least one historical dialogue turn before the target dialogue turn. The target dialogue turn can be the last dialogue turn among the multiple dialogue turns. Then, among these multiple dialogue turns, except for the target dialogue turn, the other dialogue turns are the historical dialogue turns of the target dialogue turn. These multiple dialogue turns can be arranged according to the generation order of each dialogue turn. For example, if the multiple dialogue turns obtained are arranged in the generation order as: dialogue turn 1, dialogue turn 2, dialogue turn 3, and dialogue turn 4, then dialogue turn 4 is the target dialogue turn, and dialogue turns 1, 2, and 3 are the historical dialogue turns of dialogue turn 4.

[0121] It should be noted that this computer device is a terminal, and the terminal runs the dialogue system. The terminal obtains the dialogue sentences of multiple dialogue turns through the dialogue system; or this computer device is a server, and the server is the background server of the dialogue system. The terminal interacts through the dialogue system, obtains the dialogue sentences of multiple dialogue turns, and sends the dialogue sentences of these multiple dialogue turns to the server, and the server receives the dialogue sentences of multiple dialogue turns.

[0122] In the embodiments of the present application, the dialogue sentences of the dialogue turn are generated by the interaction between the user and the dialogue system. The obtained dialogue sentences of multiple dialogue turns belong to the same user identifier and are obtained through multiple interactions between the user identifier and the dialogue system. By obtaining the dialogue sentences of multiple dialogue turns corresponding to the same user identifier, the dialogue state of the user identifier in the target dialogue turn can be determined subsequently. Among them, the user logs in to the terminal through the user identifier and interacts with the dialogue system running in the terminal.

[0123] 102. The computer device obtains the word vector set of each conversation turn according to the conversation statements of each conversation turn.

[0124] Among them, the word vector set of a conversation turn includes the word vectors of each word in the conversation statements of the conversation turn, and the word vectors of different words are different.

[0125] For any conversation turn, the conversation statements of this conversation turn include at least one of a question statement or a reply statement. The question statement includes at least one word, and the reply statement includes at least one word. Then the conversation statements of this conversation turn include at least one word. Therefore, the word vectors of each word in this conversation turn are obtained, so as to obtain the word vector set of this conversation turn.

[0126] The word vectors of multiple words in the word vector set can be arranged according to the arrangement order of the multiple words in the conversation statements. For example, the conversation statements of any conversation turn include word 1, word 2, word 3, and word 4. The word vector A of word 1, the word vector B of word 2, the word vector C of word 3, and the word vector D of word 4 are obtained. Then the word vector set of this conversation turn includes word vector A, word vector B, word vector C, and word vector D.

[0127] 103. The computer device performs weighted fusion processing on the word vector sets of multiple conversation turns according to the similarity between the dimension vectors of each preset dimension identifier and each word vector in the word vector sets of multiple conversation turns, and obtains the first fusion feature vector corresponding to each preset dimension identifier.

[0128] In the conversation statements generated in the conversation scenario, there may be a lot of information, and this information can belong to different dimensions, such as the hotel name dimension, the home address dimension, the occupation dimension, and so on. In the embodiments of the present application, the dimension to which the target conversation turn belongs and the information on this dimension can be determined through analysis and processing, so as to determine the conversation state of the target conversation turn.

[0129] Among them, the preset dimension identifier is a pre-set identifier used to indicate the corresponding dimension. The preset dimension identifier can be a hotel name dimension identifier, a hotel address dimension identifier, etc. The dimension vector is used to represent the vector of the corresponding preset dimension identifier. The preset dimension identifiers of different dimensions are different, and the dimension vectors are also different. The computer device can set multiple preset dimension identifiers, and then select the target dimension identifier corresponding to the target conversation turn from the multiple preset dimension identifiers.

[0130] In addition, the above dimension can be called a slot, and the information on the above dimension can also be called slot information.

[0131] The first fusion feature vector is the feature vector after fusing the word vector sets of multiple conversation turns, and is used to represent the fusion feature of the conversation statements of the multiple conversation turns.

[0132] The similarity between the dimension vector of any preset dimension identifier and any word vector is used to represent the possibility that the word corresponding to the word vector belongs to the preset dimension corresponding to the preset dimension identifier. The greater the similarity, the greater the possibility that the word belongs to the preset dimension; the smaller the similarity, the smaller the possibility that the word belongs to the preset dimension.

[0133] For each preset dimension identifier, the higher the similarity between the word vector and the preset dimension identifier, the greater the weight corresponding to the word vector; the lower the similarity between the word vector and the preset dimension identifier, the smaller the weight corresponding to the word vector. Then, according to the similarity between the dimension vector of the preset dimension identifier and each word vector in the word vector set of multiple dialogue turns, weighted fusion processing is performed on the word vector set of multiple dialogue turns, so as to obtain the first fusion feature vector corresponding to the preset dimension identifier.

[0134] Repeat the above steps according to the dimension vectors of multiple preset dimension identifiers, and the first fusion feature vector of each preset dimension identifier can be obtained. Since the dimension vectors of different preset dimension identifiers are different, the similarity between the same word vector and different dimension vectors may also be different. Different first fusion feature vectors corresponding to different preset dimension identifiers may be obtained through the similarity between the dimension vectors of different preset dimension identifiers and each word vector in the word vector set of multiple dialogue turns.

[0135] 104. The computer device determines at least one target dimension identifier corresponding to the target dialogue turn and the target keyword corresponding to each target dimension identifier according to the similarity between the word vector of the preset keyword corresponding to each preset dimension identifier and the corresponding first fusion feature vector.

[0136] Among them, the preset keyword is a keyword belonging to the corresponding preset dimension identifier and can be preset. In the embodiments of the present application, each preset dimension identifier may have one or more preset keywords. For example, the preset keywords corresponding to the hotel name dimension identifier may be: Hotel A, Hotel B, Hotel C, etc.

[0137] Since the first fusion feature vector corresponding to each preset dimension identifier is obtained in the above step 103, at least one target dimension identifier is determined from multiple preset dimension identifiers according to the similarity between the word vector of the preset keyword corresponding to each preset dimension identifier and the corresponding first fusion feature vector, and the target keyword corresponding to each target dimension identifier is determined from the multiple preset keywords corresponding to each target dimension identifier.

[0138] 105. The computer device determines at least one target dimension identifier and the target keyword corresponding to each target dimension identifier as the dialogue state of the target dialogue turn.

[0139] Among them, the dialogue state is used to represent the dialogue topic included in the dialogue statement. By using at least one determined target dimension identifier and the target keyword corresponding to each target dimension identifier as the dialogue state of the target dialogue turn, it indicates that the dialogue statement in the target dialogue turn is about the target dimension, and the information included in the dialogue statement contains the target keyword, so that the intention of the target dialogue turn can be determined based on the dialogue state of the target dialogue turn in the subsequent process, and then a reply statement can be generated for the user.

[0140] The method provided by the embodiments of the present application obtains the dialogue statements of multiple dialogue turns, obtains the word vector set of each dialogue turn, performs weighted fusion processing on the word vector sets of multiple dialogue turns according to the dimension vectors of each preset dimension identifier to obtain the first fusion feature vector corresponding to each preset dimension identifier, determines at least one target dimension identifier corresponding to the target dialogue turn and the target keyword corresponding to each target dimension identifier, and determines at least one target dimension identifier and the target keyword corresponding to each target dimension identifier as the dialogue state of the target dialogue turn. By fusing the dialogue statements of the target dialogue turn and the historical dialogue turns, the amount of information in the dialogue statements is enriched, and the word vector sets of multiple dialogue turns are weighted and fused according to each preset dimension identifier, improving the accuracy of the first fusion feature vector of each preset dimension identifier, thereby improving the accuracy of the obtained dialogue state.

[0141] In a possible implementation manner, performing weighted fusion processing on the word vector sets of multiple dialogue turns according to the similarity between the dimension vector of each preset dimension identifier and each word vector in the word vector sets of multiple dialogue turns to obtain the first fusion feature vector corresponding to each preset dimension identifier includes:

[0142] For each preset dimension identifier, perform weighted fusion processing on multiple word vectors in the word vector set of each dialogue turn respectively according to the similarity between the dimension vector of the preset dimension identifier and multiple word vectors in the word vector set of each dialogue turn to obtain the first feature vector of each dialogue turn;

[0143] Perform weighted fusion processing on the first feature vectors of multiple dialogue turns according to the similarity between the dimension vector of the preset dimension identifier and the first feature vectors of multiple dialogue turns to obtain the second fusion feature vector corresponding to the preset dimension identifier;

[0144] Fuse the second fusion feature vector with the first feature vector of the target dialogue turn to obtain the first fusion feature vector corresponding to the preset dimension identifier.

[0145] In another possible implementation, for each preset dimension identifier, based on the similarity between the dimension vector of the preset dimension identifier and multiple word vectors in the word vector set of each conversation turn, weighted fusion processing is respectively performed on the multiple word vectors in the word vector set of each conversation turn to obtain the first feature vector of each conversation turn, including:

[0146] For each preset dimension identifier and each conversation turn, respectively determine the first similarity between the dimension vector of the preset dimension identifier and each word vector in the word vector set of the conversation turn;

[0147] According to the first similarity corresponding to each word vector in the word vector set of the conversation turn, determine the first weight of each word vector, and the first weight of each word vector is positively correlated with the corresponding first similarity;

[0148] According to the first weights of the multiple word vectors in the word vector set, perform weighted fusion processing on the multiple word vectors to obtain the first feature vector of the conversation turn.

[0149] In another possible implementation, based on the similarity between the dimension vector of the preset dimension identifier and the first feature vectors of multiple conversation turns, perform weighted fusion processing on the first feature vectors of multiple conversation turns to obtain the second fusion feature vector corresponding to the preset dimension identifier, including:

[0150] Respectively determine the second similarity between the dimension vector of the preset dimension identifier and the first feature vector of each conversation turn;

[0151] According to the second similarity corresponding to each conversation turn, determine the second weight of each conversation turn, and the second weight of each conversation turn is positively correlated with the corresponding second similarity;

[0152] According to the second weights of multiple conversation turns, perform weighted fusion processing on the first feature vectors of multiple conversation turns to obtain the second fusion feature vector corresponding to the preset dimension identifier.

[0153] In another possible implementation, before performing weighted fusion processing on the first feature vectors of multiple conversation turns based on the similarity between the dimension vector of the preset dimension identifier and the first feature vectors of multiple conversation turns to obtain the second fusion feature vector corresponding to the preset dimension identifier, the method further includes:

[0154] According to the third similarity between the first feature vector of any conversation turn and the first feature vectors of multiple conversation turns, perform weighted fusion processing on the first feature vectors of multiple conversation turns, and use the feature vector after the fusion processing as the adjusted first feature vector of any conversation turn.

[0155] In another possible implementation, based on the third similarity between the first feature vector of any conversation turn and the first feature vectors of multiple conversation turns, performing weighted fusion processing on the word vector sets of multiple conversation turns, and using the feature vector after the fusion processing as the adjusted first feature vector of any conversation turn, including:

[0156] Based on the third similarity between the first feature vector of any conversation turn and the first feature vectors of multiple conversation turns, determining the third weight of each conversation turn among multiple conversation turns, and the third weight corresponding to each conversation turn among multiple conversation turns has a positive correlation with the corresponding third similarity;

[0157] Based on the third weights corresponding to multiple conversation turns, performing weighted fusion processing on the first feature vectors of multiple conversation turns, and using the feature vector after the fusion processing as the adjusted first feature vector of any conversation turn.

[0158] In another possible implementation, before performing weighted fusion processing on the first feature vectors of multiple conversation turns according to the similarity between the dimension vector of a preset dimension identifier and the first feature vectors of multiple conversation turns to obtain the second fusion feature vector corresponding to the preset dimension identifier, the method further includes:

[0159] Obtaining the position vector of each conversation turn, where the position vector of the conversation turn is used to represent the position of the conversation turn among multiple conversation turns;

[0160] Performing fusion processing on the first feature vector of each conversation turn and the corresponding position vector, and using the feature vector after the fusion processing as the adjusted first feature vector of each conversation turn.

[0161] In another possible implementation, performing fusion processing on the second fusion feature vector and the first feature vector of the target conversation turn to obtain the first fusion feature vector corresponding to the preset dimension identifier, including:

[0162] Determining the fourth weight of the first feature vector of the target conversation turn and the fifth weight of the second fusion feature vector, and the sum of the fourth weight and the fifth weight is 1;

[0163] Based on the fourth weight and the fifth weight, performing weighted fusion processing on the second fusion feature vector and the first feature vector of the target conversation turn to obtain the first fusion feature vector corresponding to the preset dimension identifier.

[0164] In another possible implementation, based on the similarity between the word vector of the preset keyword corresponding to each preset dimension identifier and the corresponding first fusion feature vector, determining at least one target dimension identifier corresponding to the target conversation turn and the target keyword corresponding to each target dimension identifier, including:

[0165] For each preset dimension identifier, determine the fourth similarity between the word vectors of the multiple preset keywords corresponding to the preset dimension identifier and the first fusion feature vector;

[0166] Select alternative keywords from the multiple preset keywords, where the fourth similarity corresponding to the alternative keywords is greater than the fourth similarity corresponding to the other preset keywords among the multiple preset keywords;

[0167] In response to the alternative keyword not being an empty keyword, determine the preset dimension identifier as the target dimension identifier and the alternative keyword as the target keyword.

[0168] In another possible implementation, according to the dialogue statements of each dialogue turn, obtain the word vector set of each dialogue turn, including:

[0169] Invoke the word vector acquisition submodel in the dialogue state determination model, and according to the dialogue statements of each dialogue turn, obtain the word vector set of each dialogue turn;

[0170] According to the similarity between the dimension vector of each preset dimension identifier and each word vector in the word vector sets of multiple dialogue turns, perform weighted fusion processing on the word vector sets of multiple dialogue turns to obtain the first fusion feature vector corresponding to each preset dimension identifier, including:

[0171] Invoke the feature vector acquisition submodel in the dialogue state determination model, and according to the similarity between the dimension vector of each preset dimension identifier and each word vector in the word vector sets of multiple dialogue turns, perform weighted fusion processing on the word vector sets of multiple dialogue turns to obtain the first fusion feature vector corresponding to each preset dimension identifier;

[0172] According to the similarity between the word vector of the preset keyword corresponding to each preset dimension identifier and the corresponding first fusion feature vector, determine at least one target dimension identifier corresponding to the target dialogue turn and the target keyword corresponding to each target dimension identifier; determine at least one target dimension identifier and the target keyword corresponding to each target dimension identifier as the dialogue state of the target dialogue turn, including:

[0173] Invoke the dialogue state determination submodel in the dialogue state determination model, and according to the similarity between the word vector of the preset keyword corresponding to each preset dimension identifier and the corresponding first fusion feature vector, determine at least one target dimension identifier corresponding to the target dialogue turn and the target keyword corresponding to each target dimension identifier; determine at least one target dimension identifier and the target keyword corresponding to each target dimension identifier as the dialogue state of the target dialogue turn.

[0174] In another possible implementation, the feature vector acquisition sub-model in the dialogue state determination model is called, and based on the similarity between the dimensional vector identified by each preset dimension and each word vector in the word vector set of multiple dialogue turns, weighted fusion processing is performed on the word vector set of multiple dialogue turns to obtain the first fusion feature vector corresponding to each preset dimension identifier, including:

[0175] Call the first attention layer in the feature vector acquisition sub-model. For each preset dimension identifier, based on the similarity between the dimensional vector of the preset dimension identifier and multiple word vectors in the word vector set of each dialogue turn, weighted fusion processing is respectively performed on the multiple word vectors in the word vector set of each dialogue turn to obtain the first feature vector of each dialogue turn;

[0176] Call the second attention layer in the feature vector acquisition sub-model. Based on the similarity between the dimensional vector of the preset dimension identifier and the first feature vectors of multiple dialogue turns, weighted fusion processing is performed on the first feature vectors of multiple dialogue turns to obtain the second fusion feature vector corresponding to the preset dimension identifier;

[0177] Call the feature fusion layer in the feature vector acquisition sub-model, and fuse the second fusion feature vector with the first feature vector of the target dialogue turn to obtain the first fusion feature vector corresponding to the preset dimension identifier.

[0178] In another possible implementation, before calling the second attention layer in the feature vector acquisition sub-model to perform weighted fusion processing on the first feature vectors of multiple dialogue turns based on the similarity between the dimensional vector of the preset dimension identifier and the first feature vectors of multiple dialogue turns to obtain the second fusion feature vector corresponding to the preset dimension identifier, the method further includes:

[0179] Call the third attention layer in the feature vector acquisition sub-model. Based on the third similarity between the first feature vector of any dialogue turn and the first feature vectors of multiple dialogue turns, weighted fusion processing is performed on the word vector set of multiple dialogue turns, and the feature vector after the fusion processing is used as the adjusted first feature vector of the dialogue turn.

[0180] In another possible implementation, before calling the second attention layer in the feature vector acquisition sub-model to perform weighted fusion processing on the first feature vectors of multiple dialogue turns based on the similarity between the dimensional vector of the preset dimension identifier and the first feature vectors of multiple dialogue turns to obtain the second fusion feature vector corresponding to the preset dimension identifier, the method further includes:

[0181] Invoke the third attention layer in the feature vector acquisition sub-model to obtain the position vector of each conversation turn. The position vector of the conversation turn is used to represent the position of the conversation turn among multiple conversation turns; fuse the first feature vector of each conversation turn with the corresponding position vector, and use the fused feature vector as the adjusted first feature vector of each conversation turn.

[0182] In another possible implementation, the method further includes:

[0183] Obtain the sample conversation statements of multiple sample conversation turns, and the corresponding sample conversation states of each sample conversation turn. The sample conversation statements of each sample conversation turn include at least one of the sample user question statement or the sample reply statement. The multiple sample conversation turns include the target sample conversation turn and at least one historical sample conversation turn before the target sample conversation turn;

[0184] Invoke the word vector acquisition sub-model, and according to the sample conversation statements of each sample conversation turn, obtain the word vector set of each sample conversation turn. The word vector set of the sample conversation turn includes the word vectors of each word in the sample conversation statement of the sample conversation turn;

[0185] Invoke the feature vector acquisition sub-model, and perform weighted fusion processing on the word vector sets of the target sample conversation turn and the corresponding historical sample conversation turns according to the similarity between the dimension vector of each preset dimension identifier and each word vector in the word vector sets of the target sample conversation turn and the corresponding historical sample conversation turns, to obtain the first sample fusion feature vector corresponding to each preset dimension identifier;

[0186] Invoke the conversation state determination sub-model, and according to the word vectors of multiple preset keywords corresponding to each preset dimension identifier and the corresponding first sample fusion feature vector, determine at least one predicted dimension identifier corresponding to the target sample conversation turn, and the predicted keywords corresponding to each predicted dimension identifier; determine the at least one predicted dimension identifier and the predicted keywords corresponding to each predicted dimension identifier as the predicted conversation state of the target sample conversation turn;

[0187] Adjust the word vector acquisition sub-model, the feature vector acquisition sub-model, and the conversation state determination sub-model according to the predicted conversation state of the target sample conversation turn and the corresponding sample conversation state.

[0188] In another possible implementation, after obtaining the sample conversation statements of multiple sample conversation turns and the corresponding sample conversation states of each sample conversation turn, the method further includes:

[0189] Take any one of the multiple sample conversation turns as the target sample conversation turn, and take the sample conversation turns before any one of the sample conversation turns as the historical sample conversation turns of the target sample conversation turn.

[0190] In another possible implementation manner, before adjusting the word vector acquisition sub-model, the feature vector acquisition sub-model, and the conversation state determination sub-model according to the predicted conversation state of the target sample conversation turn and the corresponding sample conversation state, the method includes:

[0191] Obtain the multiple sample change probabilities of each preset dimension identifier, where the sample change probability represents the difference between the sample conversation states corresponding to every two adjacent sample conversation turns among the multiple sample conversation turns;

[0192] For each preset dimension identifier, call the state transition model to process the first fusion feature vectors of every two adjacent sample conversation turns corresponding to the preset dimension identifier respectively, and obtain the multiple predicted change probabilities corresponding to the preset dimension identifier. The multiple sample change probabilities corresponding to the same preset dimension identifier and the multiple predicted change probabilities correspond one by one;

[0193] Adjusting the word vector acquisition sub-model, the feature vector acquisition sub-model, and the conversation state determination sub-model according to the predicted conversation state of the target sample conversation turn and the corresponding sample conversation state includes:

[0194] Adjust the word vector acquisition sub-model, the feature vector acquisition sub-model, the conversation state determination sub-model, and the called state transition model according to the obtained multiple predicted change probabilities of each preset dimension identifier, the multiple sample change probabilities of each preset dimension identifier, the predicted conversation state of each sample conversation turn, and the corresponding sample conversation state.

[0195] Figure 2 It is a flowchart of a conversation state determination method provided by an embodiment of the present application, which is applied to a computer device, such as Figure 2 As shown, the method includes:

[0196] 201. The computer device obtains the conversation statements of multiple conversation turns.

[0197] Among them, the conversation statement of each conversation turn includes at least one of a question statement or a reply statement, and the multiple conversation turns include a target conversation turn and at least one historical conversation turn before the target conversation turn.

[0198] In a possible implementation, the computer device is a server, which is the background server of the dialogue system, and the terminal runs the dialogue system. Then, step 201 may include: the terminal sends the target user identifier to the server, and the server queries the dialogue statements of multiple dialogue turns corresponding to the target user identifier from the dialogue statement library according to the target user identifier.

[0199] Among them, the target user identifier is the user identifier logged in to the terminal. The dialogue statement library is used to store the dialogue statements of the interaction between the dialogue system and the user. In the dialogue statement library, the dialogue statements are stored corresponding to the user identifiers. Through the target user identifier, the corresponding dialogue statements of multiple dialogue turns can be queried from the dialogue statement library.

[0200] 202. The computer device obtains the word vector set of each dialogue turn according to the dialogue statement of each dialogue turn.

[0201] Among them, the word vector set of the dialogue turn includes the word vectors of each word in the dialogue statement of the dialogue turn.

[0202] 203. For each preset dimension identifier, the computer device performs weighted fusion processing on the multiple word vectors in the word vector set of each dialogue turn according to the similarity between the dimension vector of the preset dimension identifier and the multiple word vectors in the word vector set of each dialogue turn, and obtains the first feature vector of each dialogue turn.

[0203] Among them, the first feature vector is used to represent the feature after the fusion of the multiple word vectors of the corresponding dialogue turn.

[0204] For any preset dimension identifier and any dialogue turn, the word vector set of the dialogue turn includes multiple word vectors. According to the similarity between the dimension vector of the preset dimension identifier and the multiple word vectors in the word vector set of the dialogue turn, the weight of each word vector in the word vector set of the dialogue turn is determined. According to the weight of each word vector, the multiple word vectors in the word vector set of the dialogue turn are subjected to weighted fusion processing, and used as the first feature vector of the dialogue turn corresponding to the preset dimension identifier.

[0205] Since the computer device obtains the word vector sets of multiple dialogue turns, the first feature vector of each dialogue turn corresponding to the preset dimension identifier can be obtained. And since the computer device obtains multiple preset dimension identifiers, the first feature vector of each dialogue turn corresponding to each preset dimension identifier can be obtained. Since the dimension vectors of different preset dimension identifiers are different, the similarity between the dimension vectors of different preset dimension identifiers and the same word vector may be different, and the weight of the same word vector obtained may also be different. Then, for the same dialogue turn, the first feature vectors corresponding to different preset dimension identifiers may also be different.

[0206] For example, if multiple dialogue turns include dialogue turn 1, dialogue turn 2, and dialogue turn 3, and multiple preset dimension identifiers include preset dimension identifier A, preset dimension identifier B, and preset identifier dimension C, then the first feature vectors of dialogue turn 1 corresponding to the obtained preset dimension identifier A, the first feature vectors of dialogue turn 2 corresponding to the preset dimension identifier A, the first feature vectors of dialogue turn 3 corresponding to the preset dimension identifier A, the first feature vectors of dialogue turn 1 corresponding to the preset dimension identifier B, the first feature vectors of dialogue turn 2 corresponding to the preset dimension identifier B, the first feature vectors of dialogue turn 3 corresponding to the preset dimension identifier B, the first feature vectors of dialogue turn 1 corresponding to the preset dimension identifier C, the first feature vectors of dialogue turn 2 corresponding to the preset dimension identifier C, and the first feature vectors of dialogue turn 3 corresponding to the preset dimension identifier C are obtained, a total of 9 first feature vectors, and the 9 first feature vectors may be different.

[0207] In a possible implementation manner, step 203 may include:

[0208] 2031. For each preset dimension identifier and each dialogue turn, respectively determine the first similarity between the dimension vector of the preset dimension identifier and each word vector in the word vector set of the dialogue turn.

[0209] Among them, the first similarity is used to represent the similarity degree between the dimension vector and the word vector. For any word vector and the dimension vector of the preset dimension identifier, the Euclidean distance, cosine distance, etc. can be used to determine the first similarity.

[0210] 2032. According to the first similarity corresponding to each word vector in the word vector set of the dialogue turn, determine the first weight of each word vector.

[0211] Among them, the first weight is used to represent the influence degree of the corresponding word vector on the feature vector of the dialogue turn. The greater the first weight, the greater the influence degree, and the smaller the first weight, the smaller the influence degree. The first weight of each word vector is positively correlated with the corresponding first similarity. The greater the first similarity of the word vector, the greater the first weight of the word vector, and the smaller the first similarity of the word vector, the smaller the first weight of the word vector.

[0212] In a possible implementation manner, step 2032 may include: determining the sum of the first similarities corresponding to multiple word vectors in the word vector set of the dialogue turn as the total similarity, and taking the ratio between the first similarity corresponding to any word vector in the word vector set of the dialogue turn and the total similarity as the first weight of the word vector. Among them, the sum of the first weights corresponding to multiple word vectors in the word vector set of the dialogue turn is 1.

[0213] It should be noted that the embodiments of the present application are described by determining the first weight of each word vector according to the first similarity corresponding to each word vector in the word vector set according to the dialogue turn. In another embodiment, steps 2031-2032 do not need to be executed, and the following steps can be executed to determine the first weight of each word vector: For any preset dimension identifier and any dialogue turn, the product of any word vector in the word vector set of this dialogue turn and the dimension vector of this preset dimension identifier is used as the first weight corresponding to this word vector.

[0214] 2033. According to the first weights of multiple word vectors in the word vector set, perform weighted fusion processing on the multiple word vectors to obtain the first feature vector of the dialogue turn.

[0215] For multiple word vectors in the word vector set of this dialogue turn, perform fusion processing on the product of each word vector and the corresponding first weight to obtain the first feature vector of this dialogue turn.

[0216] In a possible implementation manner, step 2033 may include: According to the first weights of multiple word vectors in the word vector set, perform weighted summation on the multiple word vectors to obtain the first feature vector of the dialogue turn. That is, determine the product of each word vector in the multiple word vectors and the corresponding first weight, and take the sum of the products corresponding to the multiple word vectors as the first feature vector of this dialogue turn.

[0217] In a possible implementation manner, step 2033 may include: According to the first weights of multiple word vectors in the word vector set, perform weighted averaging on the multiple word vectors to obtain the first feature vector of the dialogue turn. That is, determine the product of each word vector in the multiple word vectors and the corresponding first weight, determine the sum of the products corresponding to the multiple word vectors, and determine the ratio between the sum of the products and the number of the multiple word vectors as the first feature vector of this dialogue turn.

[0218] 204. The computer device performs weighted fusion processing on the first feature vectors of multiple dialogue turns according to the similarity between the dimension vector of the preset dimension identifier and the first feature vectors of multiple dialogue turns to obtain the second fusion feature vector corresponding to the preset dimension identifier.

[0219] Among them, the second fusion feature vector is a vector used to represent the features of the dialogue statements of multiple dialogue turns.

[0220] For any preset dimension identifier, according to the similarity between the dimension vector of the preset dimension identifier and the first feature vector of each dialogue turn, determine the weight corresponding to each dialogue turn, and perform weighted fusion processing on the first feature vectors of the multiple dialogue turns according to the weight of each dialogue turn to obtain the second fusion feature vector corresponding to the preset dimension identifier.

[0221] Since the computer device obtains multiple preset dimension identifiers, the second fusion feature vectors corresponding to each preset dimension identifier can be obtained. Since the dimension vectors of different preset dimension identifiers are different, the similarities between the dimension vectors of different preset dimension identifiers and the first feature vector of the same dialogue turn may be different, so the weights of the same dialogue turn obtained may also be different, and the second fusion feature vectors corresponding to different preset dimension identifiers may also be different.

[0222] In a possible implementation manner, step 204 may include:

[0223] 2041. Respectively determine the second similarities between the dimension vectors of the preset dimension identifiers and the first feature vector of each dialogue turn.

[0224] Among them, the second similarity is used to represent the similarity degree between the dimension vector and the first feature vector of the dialogue turn. For the first feature vector of any dialogue turn, according to the dimension vector of the preset dimension identifier and the first feature vector of this dialogue turn, determine the second similarity between the preset dimension and the first feature vector of this dialogue turn. The second similarity can be determined by methods such as Euclidean distance and cosine distance.

[0225] 2042. According to the second similarities corresponding to each dialogue turn, determine the second weights of each dialogue turn.

[0226] Among them, the second weight is used to represent the influence degree of the first feature vector of the corresponding dialogue turn on the first fusion feature vectors of the multiple dialogue turns. The greater the second weight, the greater the influence degree, and the smaller the second weight, the smaller the influence degree. The second weight of each dialogue turn is positively correlated with the corresponding second similarity. The greater the second similarity corresponding to the dialogue turn, the greater the second weight of this dialogue turn, and the smaller the second similarity corresponding to the dialogue turn, the smaller the second weight of this dialogue turn.

[0227] In a possible implementation manner, step 2042 may include: determining the sum of the second similarities corresponding to the multiple dialogue turns as the total similarity, and determining the ratio between the second similarity of any dialogue turn and the total similarity as the second weight of this dialogue turn. Among them, the sum of the second weights of the multiple dialogue turns is 1.

[0228] 2043. According to the second weights of the multiple dialogue turns, perform weighted fusion processing on the first feature vectors of the multiple dialogue turns to obtain the second fusion feature vector corresponding to the preset dimension identifier.

[0229] In a possible implementation, step 2043 may include: performing a weighted sum on the first feature vectors of multiple dialogue turns according to the second weights of the multiple dialogue turns to obtain a second fused feature vector corresponding to a preset dimension identifier. That is, determining the product of the first feature vector of each dialogue turn in the multiple dialogue turns and the corresponding second weight, and taking the sum of the products corresponding to the multiple dialogue turns as the second fused feature vector corresponding to the preset dimension identifier.

[0230] In a possible implementation, step 2043 may include: performing a weighted average on the first feature vectors of multiple dialogue turns according to the second weights of the multiple dialogue turns to obtain a second fused feature vector corresponding to a preset dimension identifier. That is, determining the product of the first feature vector of each dialogue turn in the multiple word vectors and the corresponding second weight, determining the sum of the products corresponding to the multiple dialogue turns, and determining the ratio between the sum of the products and the number of the multiple dialogue turns as the second fused feature vector of the preset dimension identifier.

[0231] It should be noted that in the embodiments of the present application, for any dimension identifier, after determining the first feature vector of each dialogue turn, a weighted fusion process is directly performed on the first feature vector of each dialogue turn according to the dimension vector of the preset dimension identifier to obtain the second fused feature vector of the preset dimension identifier. In another embodiment, after obtaining the first feature vector of each dialogue turn, the first feature vector may also be adjusted, and after adjusting the first feature vector of each dialogue turn, the step of performing a weighted fusion process on the first feature vectors of multiple dialogue turns is executed.

[0232] In a possible implementation, the process of adjusting the first feature vector of each dialogue turn may include the following three methods:

[0233] The first method: performing a weighted fusion process on the first feature vectors of multiple dialogue turns according to the third similarity between the first feature vector of any dialogue turn and the first feature vectors of multiple dialogue turns, and taking the feature vector after the fusion process as the adjusted first feature vector of any dialogue turn.

[0234] The adjusted first feature vector of each dialogue turn is obtained by performing a weighted fusion process on the first feature vectors of the multiple dialogue turns before adjustment. In the above manner, the adjusted first feature vector of each dialogue turn is obtained respectively. By adjusting each dialogue turn, the features of the dialogue sentences of other dialogue turns are incorporated into each dialogue turn, thereby improving the accuracy of the feature vector of each dialogue turn.

[0235] In a possible implementation, the first method may include the following steps 1 - step 2:

[0236] Step 1: Determine the third weight of each conversation turn among the multiple conversation turns according to the third similarity between the first feature vector of any conversation turn and the first feature vectors of the multiple conversation turns.

[0237] Among them, the third weight corresponding to each conversation turn among the multiple conversation turns is positively correlated with the corresponding third similarity. The greater the third similarity, the greater the corresponding third weight; the smaller the third similarity, the smaller the corresponding third weight.

[0238] In a possible implementation manner, step 1 may include: taking the ratio between the third similarity corresponding to any conversation turn among the multiple conversation turns and the sum of the third similarities corresponding to the multiple conversation turns as the third weight corresponding to any conversation turn among the multiple conversation turns.

[0239] Step 2: Perform weighted fusion processing on the first feature vectors of the multiple conversation turns according to the third weights corresponding to the multiple conversation turns, and use the feature vector after the fusion processing as the adjusted first feature vector of any conversation turn.

[0240] In a possible implementation manner, step 2 may include: performing weighted summation on the first feature vectors of the multiple conversation turns according to the third weights corresponding to the multiple conversation turns to obtain the adjusted first feature vector of any conversation turn.

[0241] In a possible implementation manner, step 2 may include: performing weighted averaging on the first feature vectors of the multiple conversation turns according to the third weights corresponding to the multiple conversation turns to obtain the adjusted first feature vector of any conversation turn.

[0242] It should be noted that the first method is only described by adjusting the first feature vector of each conversation turn once. In another embodiment, after adjusting the first feature vector of each conversation turn once, a second adjustment is performed according to the first feature vectors of the multiple conversation turns after the first adjustment to obtain the first feature vectors of the multiple conversation turns after the second adjustment. The above process is repeated. When the number of adjustments reaches the preset number of times, step 204 is executed according to the first feature vectors of the multiple conversation turns adjusted by the preset number of times. Among them, the adjustment methods in each adjustment process are similar and will not be elaborated here.

[0243] The second method: Obtain the position vector of each conversation turn, perform fusion processing on the first feature vector of each conversation turn and the corresponding position vector, and use the feature vector after the fusion processing as the adjusted first feature vector of each conversation turn.

[0244] Among them, the position vector of the dialogue turn is used to represent the position of the dialogue turn among multiple dialogue turns. By fusing the first feature vector of each dialogue turn with the corresponding position vector, the relationship between multiple dialogue turns is enhanced, and the accuracy of the first feature vector is improved.

[0245] The third method: Obtain the position vector of each dialogue turn, perform a fusion process on the first feature vector of each dialogue turn and the corresponding position vector, use the feature vector after the fusion process as the first feature vector after the first adjustment of each dialogue turn, and perform a weighted fusion process on the first feature vectors after the first adjustment of multiple dialogue turns according to the third similarity between the first feature vector after the first adjustment of any dialogue turn and the first feature vectors after the first adjustment of multiple dialogue turns, and use the feature vector after the fusion process as the first feature vector after the second adjustment of any dialogue turn.

[0246] In the process of adjusting the first feature vector of each dialogue turn, first, according to the second method above, fuse the first feature vector of each dialogue turn with the corresponding position vector to obtain the first feature vector after adjustment of each dialogue turn. According to the first feature vector after adjustment of each dialogue turn, and according to the first method above, further adjust the first feature vector after adjustment of each dialogue turn to obtain the first feature vector after further adjustment.

[0247] In addition, after obtaining the first feature vector after the second adjustment, the above process can also be repeatedly executed according to the first method. When the number of adjustments reaches the preset number of times, step 204 above is executed according to the first feature vectors of multiple dialogue turns after being adjusted by the preset number of times.

[0248] 205. The computer device fuses the second fusion feature vector with the first feature vector of the target dialogue turn to obtain the first fusion feature vector corresponding to the preset dimension identifier.

[0249] The embodiment of this application is to obtain the dialogue state of the target dialogue turn. After obtaining the second fusion feature vector after fusing the dialogue statements of multiple dialogue turns, in order to enhance the features of the dialogue statement of the target dialogue turn, the second fusion feature vector is fused with the first feature vector of the target dialogue turn again to obtain the first fusion feature vector corresponding to the preset dimension.

[0250] In a possible implementation manner, step 205 may include:

[0251] 2051. Determine the fourth weight of the first feature vector of the target dialogue turn and the fifth weight of the second fusion feature vector.

[0252] Among them, the sum of the fourth weight and the fifth weight is 1, and the fourth weight and the fifth weight can be set manually or obtained through model training.

[0253] 2052. According to the fourth weight and the fifth weight, perform weighted fusion processing on the second fusion feature vector and the first feature vector of the target dialogue turn to obtain the first fusion feature vector corresponding to the preset dimension identifier.

[0254] When performing weighted fusion processing on the second fusion feature vector and the first feature vector of the target dialogue turn, it is possible to perform weighted summation on the second fusion feature vector and the first feature vector of the target dialogue turn, or perform weighted averaging on the second fusion feature vector and the first feature vector of the target dialogue turn.

[0255] It should be noted that in the embodiments of the present application, it is described that the first feature vector of each dialogue turn is first obtained, and then the first feature vectors of each dialogue turn are subjected to weighted fusion processing to obtain the first fusion feature vector. In another embodiment, steps 203-205 do not need to be executed, and it is possible to perform weighted fusion processing on the word vector sets of multiple dialogue turns according to the similarity between the dimension vectors of each preset dimension identifier and each word vector in the word vector sets of multiple dialogue turns to obtain the first fusion feature vector corresponding to each preset dimension identifier.

[0256] 206. The computer device determines at least one target dimension identifier corresponding to the target dialogue turn and the target keyword corresponding to each target dimension identifier according to the similarity between the word vector of the preset keyword corresponding to each preset dimension identifier and the corresponding first fusion feature vector.

[0257] Among them, the preset keyword is a keyword belonging to the preset dimension and is used to describe the preset dimension, and can be set in advance. In the embodiments of the present application, each preset dimension identifier can have one or more preset keywords. For example, the preset keywords corresponding to the hotel name dimension identifier can be: Hotel A, Hotel B, Hotel C, etc.

[0258] Since the first fusion feature vectors corresponding to multiple preset dimension identifiers are obtained, at least one target dimension identifier is determined from multiple preset dimension identifiers according to the similarity between the word vector of the preset keyword corresponding to each preset dimension identifier and the corresponding first fusion feature vector, and the target keyword corresponding to each target dimension identifier is respectively determined from the multiple preset keywords corresponding to each target dimension identifier.

[0259] In a possible implementation manner, step 206 may include the following steps 2061-2063:

[0260] 2061. For each preset dimension identifier, determine the fourth similarity between the word vectors of multiple preset keywords corresponding to the preset dimension identifier and the first fusion feature vector.

[0261] Among them, the fourth similarity is used to represent the similarity degree between the word vector of the preset keyword and the first fusion feature vector.

[0262] 2062. Select alternative keywords from multiple preset keywords.

[0263] Among them, the fourth similarity corresponding to the alternative keyword is greater than the fourth similarity corresponding to other preset keywords among the multiple preset keywords.

[0264] Optionally, for any preset dimension identifier, determine the similarity between each preset keyword corresponding to the preset dimension identifier and the first fusion feature vector corresponding to the preset dimension identifier, and select the preset keyword with the greatest similarity to the first fusion feature vector from the multiple preset keywords corresponding to the preset dimension identifier as the alternative keyword, so as to obtain multiple preset dimension identifiers and the alternative keywords corresponding to each preset dimension identifier.

[0265] 2063. In response to the alternative keyword not being an empty keyword, determine the preset dimension identifier as the target dimension identifier and the alternative keyword as the target keyword.

[0266] Among them, the empty keyword is used to indicate that the first fusion feature vector corresponding to the preset dimension identifier does not match the preset dimension identifier, that is, the feature of the dialogue statement in the target dialogue turn does not match the preset dimension identifier.

[0267] In the embodiments of the present application, among the multiple preset keywords corresponding to each preset dimension identifier, an empty keyword is included. When the alternative keyword corresponding to any preset dimension identifier is an empty keyword, it means that the preset dimension identifier does not match the feature of the dialogue statement in the target dialogue turn. When the alternative keyword corresponding to any preset dimension identifier is not an empty keyword, it means that the alternative keyword corresponding to the preset dimension identifier matches the feature of the dialogue statement in the target dialogue turn.

[0268] Through the alternative keywords corresponding to the determined multiple preset dimension identifiers, select the preset dimension identifier whose alternative keyword is not an empty keyword from the multiple preset dimension identifiers as the target dimension identifier, and use the alternative keyword corresponding to the target dimension identifier as the target keyword.

[0269] 207. The computer device determines at least one target dimension identifier and the target keyword corresponding to each target dimension identifier as the dialogue state of the target dialogue turn.

[0270] Among them, the dialogue state is used to represent the dialogue topics included in the dialogue statements. For example, the dialogue state of the target dialogue turn includes: the hotel name dimension identifier and hotel name A, and the restaurant name dimension identifier and restaurant name B.

[0271] By using at least one determined target dimension identifier and the target keyword corresponding to each target dimension identifier as the dialogue state of the target dialogue turn, it is indicated that the dialogue statement in this target dialogue turn is about the target dimension, and the information included in the dialogue statement contains the target keyword, so that the intention of the target dialogue turn can be determined based on the dialogue state of the target dialogue turn subsequently, and then a reply statement can be generated for the user.

[0272] It should be noted that the embodiments of this application are described by obtaining the dialogue state of the target dialogue turn. In another embodiment, when the alternative keywords of all preset dimension identifiers are empty keywords, it is determined that the dialogue state of the target dialogue turn has not been obtained. In response to not obtaining the dialogue state of the target dialogue turn, a preset question statement is sent to the user; or, in response to not obtaining the dialogue state of the target dialogue turn, any one of the preset dimension identifiers in the preset dimension identifiers is recommended to the user to prompt the user to input the information of this preset dimension identifier.

[0273] The method provided by the embodiments of this application obtains the dialogue statements of multiple dialogue turns, obtains the word vector set of each dialogue turn, performs weighted fusion processing on the word vector sets of multiple dialogue turns according to the dimension vectors of each preset dimension identifier, obtains the first fusion feature vector corresponding to each preset dimension identifier, determines at least one target dimension identifier corresponding to the target dialogue turn, and the target keyword corresponding to each target dimension identifier, and determines at least one target dimension identifier and the target keyword corresponding to each target dimension identifier as the dialogue state of the target dialogue turn. By fusing the dialogue statements of the target dialogue turn and the historical dialogue turns, the amount of information in the dialogue statements is enriched, and the word vector sets of multiple dialogue turns are weighted and fused according to each preset dimension identifier, improving the accuracy of the first fusion feature vector of each preset dimension identifier, thereby improving the accuracy of the obtained dialogue state.

[0274] Moreover, for any preset dimension identifier, the first feature vector of each dialogue turn is adjusted according to the first feature vectors of multiple dialogue turns, so that the features of the dialogue statements of other dialogue turns are incorporated into the adjusted first feature vector of each dialogue turn, enhancing the relationship between multiple dialogue turns, improving the accuracy of the adjusted first feature vector of each dialogue turn, and thus improving the accuracy of the determined dialogue state.

[0275] Moreover, when adjusting the first feature vector of each dialogue turn, the position vectors of each dialogue turn are fused, enhancing the context connection between multiple dialogue turns, thereby improving the accuracy of the determined dialogue state.

[0276] Based on the Figure 2 embodiment shown, the dialogue state determination model can also process the dialogue statements of multiple dialogue turns to obtain the dialogue state of the target dialogue turn. The specific process is detailed in the following embodiments.

[0277] Figure 3 is a flowchart of a dialogue state determination method provided by an embodiment of the present application, which is applied to a computer device, such as Figure 3 shown, and the method includes:

[0278] 301. The computer device obtains the dialogue statements of multiple dialogue turns.

[0279] This step is similar to step 201 above and will not be elaborated here.

[0280] 302. The computer device calls the word vector acquisition submodel in the dialogue state determination model, and obtains the word vector set of each dialogue turn according to the dialogue statement of each dialogue turn.

[0281] In the embodiment of the present application, the dialogue state determination model can process the dialogue statements of multiple dialogue turns to obtain the dialogue state of the target dialogue turn. The dialogue state determination model includes a word vector acquisition submodel, a feature vector acquisition submodel, and a dialogue state determination submodel. The feature vector acquisition submodel can include a first attention layer, a second attention layer, a third attention layer, and a feature fusion layer.

[0282] Among them, the word vector acquisition submodel is used to obtain the word vector set of each dialogue turn. By inputting the dialogue statement of each dialogue turn into the word vector acquisition submodel, the word vector acquisition submodel processes the dialogue statement of each dialogue turn respectively, thereby obtaining the word vector set of each dialogue turn.

[0283] In addition, since the dialogue statement of each dialogue turn includes at least one of a question statement or a reply statement, when inputting the dialogue statement of each dialogue turn into the feature vector acquisition submodel, the overall statement character can be added at the starting position of the dialogue statement, and a statement separator can be added between the question statement and the reply statement, so that the feature vector acquisition submodel can accurately obtain the word vector set of each dialogue turn.

[0284] For example, if the dialogue statement of any dialogue turn includes question statement 1 and response statement 2, with the overall character of the statement being [CLS] and the statement separator being [SEP], then the statement input to the feature vector acquisition sub-model is "[CLS]Question statement 1[SEP]Response statement 2[SEP]", or the statement input to the feature vector acquisition sub-model is "[CLS]Response statement 2[SEP]Question statement 1[SEP]". When the dialogue statement of this dialogue turn only includes question statement 1, then the statement input to the feature vector acquisition sub-model is "[CLS][SEP]Question statement 1[SEP]", or the statement input to the feature vector acquisition sub-model is "[CLS]Question statement 1[SEP][SEP]".

[0285] In a possible implementation manner, step 302 may include: calling the word vector acquisition sub-model BERT1, and according to the question statement R t and response statement U t included in the dialogue statement of the t-th dialogue turn, obtaining the word vector set h t of the t-th dialogue turn. The question statement R t , response statement U t , and the word vector set h t of the t-th dialogue turn satisfy the following relationship:

[0286] h t = BERT1([R t , U t ),

[0287] where BERT (Bidirectional Encoder Representations from Transformers, bidirectional encoding model) is used to represent the word vector acquisition sub-model.

[0288] 303. The computer device calls the first attention layer in the feature vector acquisition sub-model. For each preset dimension identifier, according to the dimension vector of the preset dimension identifier and the similarity between the dimension vector and multiple word vectors in the word vector set of each dialogue turn, weighted fusion processing is respectively performed on the multiple word vectors in the word vector set of each dialogue turn to obtain the first feature vector of each dialogue turn.

[0289] Among them, the feature vector acquisition sub-model includes a first attention layer, and the first attention layer is used to comprehensively consider multiple word vectors in the word vector set of the dialogue turn according to the dimension vector of the preset dimension identifier to obtain the first feature vector of each dialogue turn.

[0290] Through this first attention layer, the attention mechanism is introduced. By inputting the dimension vectors of each preset dimension and the set of word vectors of each conversation turn into this first attention layer, this first attention layer performs a weighted fusion process on multiple word vectors in the set of word vectors according to the similarity between the dimension vectors identified by the preset dimension and each word vector in the set of word vectors, and obtains the first feature vector of this conversation turn corresponding to the preset dimension identifier, improving the accuracy of the obtained first feature vector.

[0291] In addition, for the method of obtaining the dimension vectors of each preset dimension identifier, the dimension vectors of each dimension identifier can be obtained through other feature vector acquisition models, or the dimension vectors of each preset dimension identifier can be pre-stored in the database. Then, when using the dimension vectors of each preset dimension identifier, they can be directly queried from the database.

[0292] Optionally, call the feature vector acquisition model BERT2. For any preset dimension identifier s, obtain the dimension vector h of this preset dimension identifier s s , the dimension vector h of this preset dimension identifier s s satisfies the following relationship:

[0293] h s = BERT2(s).

[0294] Optionally, call the first attention layer. For any preset dimension identifier s, according to the dimension vector h of this preset dimension identifier s s , and the set of word vectors h of the t-th conversation turn t , obtain the first feature vector of the t-th conversation turn Then the dimension vector h of this preset dimension identifier s s , the set of word vectors h of the t-th conversation turn t and the first feature vector of the t-th conversation turn satisfy the following relationship:

[0295]

[0296] Among them, Multihead() represents the multi-head attention function of the first attention layer.

[0297] 304. The computer device calls the second attention layer in the feature vector acquisition sub-model, and performs a weighted fusion process on the first feature vectors of multiple conversation turns according to the similarity between the dimension vectors of the preset dimension identifier and the first feature vectors of multiple conversation turns, and obtains the second fusion feature vector corresponding to the preset dimension identifier.

[0298] Among them, the feature vector acquisition sub-model includes a second attention layer, which is used to obtain the first fused feature vector after fusing multiple conversation turns by comprehensively considering the first feature vectors of multiple conversation turns according to the dimension vector identified by the preset dimension.

[0299] Through this second attention layer, an attention mechanism is introduced. For any preset dimension identifier, according to the similarity between the dimension vector of the preset dimension identifier and the first feature vector of each conversation turn, the first feature vectors of multiple conversation turns are weighted and fused to obtain the second fused feature vector corresponding to the preset dimension identifier, improving the accuracy of the obtained second fused feature vector.

[0300] In a possible implementation, the second attention layer is called. For any preset dimension identifier s, the dimension vector h of the preset dimension identifier s s , and the set of first feature vectors of multiple conversation turns The second fused feature vector corresponding to the preset dimension identifier s Satisfies the following relationship:

[0301]

[0302] Among them, Multihead() represents the multi-head attention function of the first attention layer, which is used for the set of first feature vectors of multiple conversation turns, and the first feature vector of each conversation turn is used as a dimension of the set of first feature vectors.

[0303] Such as Figure 4 shown, different colors of each word correspond to different weights, and in order from largest to smallest weight are: color 3, color 2, color 1, color 0. Through multi-head attention, conversation turn 1 is determined. When the multi-head attention is different, through the weights of each conversation turn in the multi-head attention, the second attention layer fuses conversation turn 1 and conversation turn 2 according to the similarity between the first feature vector of each conversation turn and the preset dimension identifier to obtain the second fused feature vector.

[0304] It should be noted that in the embodiments of the present application, for any dimension identifier, after determining the first feature vector of each conversation turn, according to the dimension vector of the preset dimension identifier, the first feature vectors of each conversation turn are directly weighted and fused to obtain the second fused feature vector of the preset dimension identifier. In another embodiment, after obtaining the first feature vector of each conversation turn, the first feature vector can also be adjusted, and after adjusting the first feature vector of each conversation turn, the step of weighted fusion processing of the first feature vectors of multiple conversation turns is performed.

[0305] In a possible implementation, for any dimension identifier, the process of adjusting the first feature vector of each dialogue turn may include the following three methods:

[0306] The first method: Invoke the third attention layer in the feature vector acquisition sub-model. According to the third similarity between the first feature vector of any dialogue turn and the first feature vectors of multiple dialogue turns, perform weighted fusion processing on the word vector sets of multiple dialogue turns, and use the feature vector after the fusion processing as the first feature vector after the adjustment of the dialogue turn.

[0307] The adjusted first feature vector of each dialogue turn obtained by invoking the third attention layer is obtained by performing weighted fusion processing on the first feature vectors of these multiple dialogue turns before adjustment. According to the above method, the adjusted first feature vector of each dialogue turn is obtained respectively. By adjusting each dialogue turn, the features of the dialogue statements of other dialogue turns are incorporated into each dialogue turn, thereby improving the accuracy of the feature vector of each dialogue turn.

[0308] It should be noted that the first method only illustrates the adjustment of the first feature vector of each dialogue turn by invoking the third attention layer once. In another embodiment, the third attention layer may include multiple attention units, and each attention unit is used to adjust the first feature vector of each dialogue turn once. The output of the previous attention unit is used as the input of the next attention unit. Through the multiple attention units included in the third attention layer, the first feature vectors of multiple dialogue turns are adjusted multiple times. Subsequently, according to the adjusted first feature vectors of multiple dialogue turns output by the last attention unit, step 304 above is executed. Among them, the adjustment method of each attention unit is similar and will not be elaborated here.

[0309] The second method: Invoke the third attention layer in the feature vector acquisition sub-model to obtain the position vector of each dialogue turn. The position vector of the dialogue turn is used to represent the position of the dialogue turn in multiple dialogue turns; fuse the first feature vector of each dialogue turn with the corresponding position vector, and use the feature vector after the fusion processing as the first feature vector after the adjustment of each dialogue turn.

[0310] The third method: call the third attention layer in the feature vector acquisition submodel, obtain the position vector of each dialogue round, fuse the first feature vector of each dialogue round with the corresponding position vector, use the fused feature vector as the first feature vector after the first adjustment of each dialogue round, perform weighted fusion processing on the first feature vectors of multiple dialogue rounds after the first adjustment according to the third similarity between the first feature vector after the first adjustment of any dialogue round and the first feature vectors after the first adjustment of multiple dialogue rounds, and use the fused feature vector as the first feature vector after the second adjustment of any dialogue round.

[0311] The third attention layer may include two attention units. The first feature vector of each dialogue round and the position vector of each dialogue round are input into the first attention unit. The first attention unit integrates the first feature vector of each dialogue round into the corresponding position vector, and inputs the adjusted first feature vector of each dialogue round into the second attention unit. The second attention unit adjusts the adjusted first feature vector of each dialogue round again and outputs the adjusted first feature vector of each dialogue round.

[0312] In addition, the third attention layer may also include multiple attention units, each of which is used to adjust the first feature vector of each dialogue round once, and the output of the previous attention unit is used as the input of the next attention unit. The first attention unit is used to fuse and adjust the first feature vector of each dialogue round with the corresponding position vector, and the second attention unit is used to adjust the adjusted first feature vector of each dialogue round again. Each attention unit after the second attention unit adjusts the adjusted first feature vectors of multiple dialogue rounds output by the previous attention unit again, thereby obtaining the first feature vector of each dialogue round after multiple adjustments, and then executes the above step 304 according to the adjusted first feature vectors of multiple dialogue rounds output by the last attention unit. Each attention unit may include two norm sublayers, a feedforward neural network sublayer, and a multi-head self-attention sublayer.

[0313] In a possible implementation, the third attention layer may further include X attention units. For any preset dimension identifier s, the first attention unit of the third attention layer performs a fusion adjustment on the first feature vector of each dialogue round and the corresponding position vector to obtain the first first feature vector set m after the first adjustment of multiple dialogue rounds. 0 , the first eigenvector set m 0The second attention unit input to the third attention layer adjusts the first feature vectors after the first adjustment for multiple dialogue turns again. Through the nth attention unit of the third attention layer, the first feature vectors for multiple dialogue turns are adjusted again to obtain the nth set m of the first feature vectors after adjustment for multiple dialogue turns. n Through the Xth attention unit of the third attention layer, the first feature vectors for multiple dialogue turns are adjusted again to obtain the Xth set m of the first feature vectors after adjustment for multiple dialogue turns. X .

[0314] The first set m of the first feature vectors 0 satisfies the following relationship:

[0315]

[0316] where represents the first feature vector of the first dialogue turn corresponding to the preset dimension identifier s, represents the first feature vector of the second dialogue turn corresponding to the preset dimension identifier s, represents the first feature vector of the tth dialogue turn corresponding to the preset dimension identifier s, PE(1) represents the position vector of the first dialogue turn, PE(2) represents the position vector of the second dialogue turn, and PE(t) represents the position vector of the tth dialogue turn.

[0317] The nth set m of the first feature vectors n satisfies the following relationship:

[0318] m n = FFN(Multihead(m n-1 , m n-1 , m n-1 ))

[0319] FFN(x) = max(0, xW1 + b1)W2 + b2

[0320] where n is used to represent the nth attention unit among the X attention units of the third attention layer, n is a positive integer not less than 1, m n-1 represents the (n - 1)th set of the first feature vectors, Multihead( ) represents the multi-head attention function of the third attention layer, FFN( ) represents the transformation function in the third attention layer, max represents taking the maximum value, x is an arbitrary unknown, and W1, b1, W2, and b2 are all adjustment parameters and can be any constants.

[0321] The Xth set m of the first feature vectors X satisfies the following relationship:

[0322]

[0323] Wherein, X represents the total number of attention units included in the third attention layer, and X is a positive integer not less than 1. represents the set of first feature vectors updated for multiple dialogue turns corresponding to the preset dimension identifier s, and the first feature vector of each dialogue turn serves as one dimension of the set of first feature vectors.

[0324] 305. The computer device invokes the feature fusion layer in the feature vector acquisition submodel, fuses the second fusion feature vector with the first feature vector of the target dialogue turn, and obtains the first fusion feature vector corresponding to the preset dimension identifier.

[0325] Wherein, the feature vector acquisition submodel includes a feature fusion layer, and the feature fusion layer is used to fuse the second fusion feature vectors of multiple turns with the first feature vector of the target dialogue turn.

[0326] In a possible implementation manner, step 305 may include: invoking the feature fusion layer to determine the fourth weight g of the first feature vector of the target dialogue turn s,t , and the fifth weight 1 - g s,t of the second fusion feature vector. According to the fourth weight g s,t and the fifth weight 1 - g s,t , perform weighted fusion processing on the second fusion feature vector and the first feature vector of the target dialogue turn to obtain the first fusion feature vector corresponding to the preset dimension identifier s The first fusion feature vector corresponding to the preset dimension identifier s The first feature vector The second fusion feature vector The fourth weight g s,t and the fifth weight 1 - g s,t satisfy the following relationship:

[0327]

[0328]

[0329] Wherein, σ( ) represents the sigmoid (logistic regression) activation function, W g represents the adjustment parameter, which can be any constant. The W g belongs to a matrix of 2d rows and d columns, and d is a positive integer not less than 1. ⊙ represents multiplying vectors dimension by dimension, represents multiplying vectors by dot product.

[0330] It should be noted that this application is described by obtaining the first fusion feature vector through the first attention layer, the second attention layer, and the feature fusion layer. In another embodiment, steps 303-305 do not need to be executed, and the feature vector acquisition sub-model can be called to obtain the first fusion feature vector corresponding to each preset dimension identifier in other ways.

[0331] 306. The computer device calls the dialogue state determination sub-model in the dialogue state determination model, and determines at least one target dimension identifier corresponding to the target dialogue turn and the target keyword corresponding to each target dimension identifier according to the similarity between the word vector of the preset keyword corresponding to each preset dimension identifier and the corresponding first fusion feature vector. The at least one target dimension identifier and the target keyword corresponding to each target dimension identifier are determined as the dialogue state of the target dialogue turn.

[0332] Among them, the dialogue state determination sub-model is used to obtain the dialogue state of the target dialogue turn.

[0333] By inputting the word vector of the preset keyword of each preset dimension and the first fusion feature vector corresponding to each preset dimension into the dialogue state determination sub-model, the dialogue state determination sub-model outputs the dialogue state of the target dialogue turn.

[0334] In addition, for the method of obtaining the word vector of the preset keyword corresponding to each preset dimension identifier, other feature vector acquisition models can be called to obtain the word vector of each preset keyword, or the word vector of the preset keyword corresponding to each preset dimension identifier can be pre-stored in the database, and when using the word vector of the preset keyword corresponding to each preset dimension identifier, it can be directly queried from the database.

[0335] Optionally, call other feature vector acquisition model BERT2. For any preset keyword υ corresponding to any preset dimension identifier s s , obtain the word vector of this preset keyword υ s The word vector of this preset keyword υ The word vector of this preset keyword υ s satisfies the following relationship:

[0336]

[0337] Among them, BERT (Bidirectional Encoder Representations from Transformers, bidirectional encoding word embedding model) is used to represent the word vector acquisition sub-model.

[0338] The method provided by the embodiment of the present application obtains the dialogue statements of multiple dialogue turns, obtains the word vector set of each dialogue turn, performs weighted fusion processing on the word vector sets of multiple dialogue turns according to the dimension vectors identified by each preset dimension, obtains the first fusion feature vector corresponding to each preset dimension identifier, determines at least one target dimension identifier corresponding to the target dialogue turn, and the target keyword corresponding to each target dimension identifier, and determines at least one target dimension identifier and the target keyword corresponding to each target dimension identifier as the dialogue state of the target dialogue turn. By fusing the dialogue statements of the target dialogue turn and the historical dialogue turns, the information volume of the dialogue statements is enriched, and the word vector sets of multiple dialogue turns are weighted and fused according to each preset dimension identifier, improving the accuracy of the first fusion feature vector of each preset dimension identifier, thereby improving the accuracy of the obtained dialogue state.

[0339] Moreover, for any preset dimension identifier, the first feature vectors of each dialogue turn are adjusted respectively according to the first feature vectors of multiple dialogue turns, so that the features of the dialogue statements of other dialogue turns are incorporated into the adjusted first feature vectors of each dialogue turn, enhancing the relationship between multiple dialogue turns, improving the accuracy of the adjusted first feature vector of each dialogue turn, and thus improving the accuracy of the determined dialogue state.

[0340] Moreover, when adjusting the first feature vector of each dialogue turn, the position vectors of each dialogue turn are fused, enhancing the context connection between multiple dialogue turns, and thus improving the accuracy of the determined dialogue state.

[0341] Moreover, the accuracy of the obtained dialogue state is improved by the word vector acquisition sub-model, the feature vector acquisition sub-model, and the dialogue state determination sub-model included in the dialogue state determination model.

[0342] In Figure 3 Based on the embodiment shown, before calling the word vector acquisition sub-model, the feature vector acquisition sub-model, and the dialogue state determination sub-model included in the dialogue state determination model, it is necessary to train the word vector acquisition sub-model, the feature vector acquisition sub-model, and the dialogue state determination sub-model. The training process is detailed in the following embodiments.

[0343] Figure 5 A dialogue state determination model training method provided by an embodiment of the present application is applied to a computer device, as Figure 5 shown, and the method includes:

[0344] 501. The computer device obtains the sample dialogue statements of multiple sample dialogue turns and the sample dialogue state corresponding to each sample dialogue turn.

[0345] Among them, the sample dialogue statements of each sample dialogue turn include at least one of the sample user question statements or sample response statements. The multiple sample dialogue turns include the target sample dialogue turn and at least one historical sample dialogue turn before the target sample dialogue turn. The sample dialogue state corresponding to each sample dialogue turn is used to represent the dialogue topic included in the sample dialogue statement of the sample dialogue turn. The sample dialogue state corresponding to each sample dialogue turn can be obtained by manually annotating the sample dialogue statements of multiple sample dialogue turns.

[0346] 502. The computer device invokes the word vector acquisition sub-model, and according to the sample dialogue statements of each sample dialogue turn, obtains the word vector set of each sample dialogue turn.

[0347] Among them, the word vector acquisition sub-model can be an initialized word vector acquisition sub-model.

[0348] The word vector set of a sample dialogue turn includes the word vectors of each word in the sample dialogue statement of the sample dialogue turn. By inputting the sample dialogue statements of each sample dialogue turn into the word vector acquisition sub-model respectively, the word vector set of each sample dialogue turn can be obtained.

[0349] 503. The computer device invokes the first attention layer. For each preset dimension identifier, according to the similarity between the dimension vector of the preset dimension identifier and each word vector in the word vector sets of the target sample dialogue turn and the corresponding historical sample dialogue turns, performs weighted fusion processing on the multiple word vectors in the word vector sets of the target sample dialogue turn and the corresponding historical sample dialogue turns respectively, to obtain the first feature vector of the target sample dialogue turn and the first feature vector of the corresponding historical sample dialogue turn.

[0350] In the embodiments of the present application, the feature vector acquisition sub-model includes a first attention layer, a second attention layer, and a feature fusion layer. Through the first attention layer, the second attention layer, and the feature fusion layer, the first sample fusion feature vector corresponding to each preset dimension identifier is obtained.

[0351] 504. The computer device invokes the second attention layer. According to the similarity between the dimension vector of the preset dimension identifier and the first feature vectors of the target sample dialogue turn and the corresponding historical sample dialogue turns, performs weighted fusion processing on the first feature vectors of the target sample dialogue turn and the corresponding historical sample dialogue turns, to obtain the second sample fusion feature vector corresponding to the preset dimension identifier.

[0352] 505. The computer device invokes the feature fusion layer, and fuses the second sample fusion feature vector with the first feature vector of the target sample dialogue turn, to obtain the first sample fusion feature vector corresponding to the preset dimension identifier.

[0353] It should be noted that the embodiments of the present application are described by obtaining the first sample fusion feature vector corresponding to each preset dimension identifier through the first attention layer, the second attention layer, and the feature fusion layer. In another embodiment, steps 503-505 do not need to be executed, and the feature vector acquisition submodel can be called to obtain the first sample fusion feature vector corresponding to each preset dimension identifier in other ways.

[0354] 506. The computer device calls the dialogue state determination submodel, and determines at least one predicted dimension identifier corresponding to the target sample dialogue turn and the predicted keyword corresponding to each predicted dimension identifier according to the word vectors of multiple preset keywords corresponding to each preset dimension identifier and the corresponding first sample fusion feature vector, and determines at least one predicted dimension identifier and the predicted keyword corresponding to each predicted dimension identifier as the predicted dialogue state of the target sample dialogue turn.

[0355] 507. The computer device adjusts the word vector acquisition submodel, the first attention layer, the second attention layer, the feature fusion layer, and the dialogue state determination submodel according to the predicted dialogue state of the target sample dialogue turn and the corresponding sample dialogue state.

[0356] The computer device can determine the loss values of the word vector acquisition submodel, the first attention layer, the second attention layer, the feature fusion layer, and the dialogue state determination submodel according to the difference between the predicted dialogue state and the corresponding sample dialogue state, and adjust the word vector acquisition submodel, the first attention layer, the second attention layer, the feature fusion layer, and the dialogue state determination submodel according to the loss values to improve the accuracy of the word vector acquisition submodel, the first attention layer, the second attention layer, the feature fusion layer, and the dialogue state determination submodel.

[0357] It should be noted that the embodiments of the present application are described by training the word vector acquisition submodel, the first attention layer, the second attention layer, the feature fusion layer, and the dialogue state determination submodel. In another embodiment, step 507 does not need to be executed, and the word vector acquisition submodel, the feature vector acquisition submodel, and the dialogue state determination submodel can be adjusted according to the predicted dialogue state of the target sample dialogue turn and the corresponding sample dialogue state.

[0358] It should be noted that the embodiments of the present application take any target sample dialogue turn as an example to describe the training process. In another embodiment, training can also be performed according to multiple target sample dialogue turns.

[0359] In a possible implementation manner, after step 501: Any one of the multiple sample conversation turns is used as the target sample conversation turn, and the sample conversation turns before any one of the sample conversation turns are used as the historical sample conversation turns of the target sample conversation turn.

[0360] Through the obtained multiple sample conversation turns, each sample conversation turn can be used as the target sample conversation turn respectively, so that multiple sets of sample conversation turns can be obtained. Each set of sample conversation turns includes a target sample conversation turn and the corresponding historical sample conversation turns.

[0361] For example, the obtained multiple sample conversation turns include: sample conversation turn 1, sample conversation turn 2, sample conversation turn 3, sample conversation turn 4, sample conversation turn 5. Then, each sample conversation turn is used as the target sample conversation turn respectively, and 5 sets of sample conversation turns can be obtained. The first set of sample conversation turns includes: sample conversation turn 1; the second set of sample conversation turns includes: sample conversation turn 1, sample conversation turn 2; the third set of sample conversation turns includes: sample conversation turn 1, sample conversation turn 2, sample conversation turn 3; the fourth set of sample conversation turns includes: sample conversation turn 1, sample conversation turn 2, sample conversation turn 3, sample conversation turn 4; the fifth set of sample conversation turns includes: sample conversation turn 1, sample conversation turn 2, sample conversation turn 3, sample conversation turn 4, sample conversation turn 5.

[0362] In a possible implementation manner, before step 507, the method further includes the following steps 1 - step 2:

[0363] Step 1: Obtain multiple sample change probabilities for each preset dimension identifier.

[0364] Wherein, the sample change probability represents the difference between the sample conversation states corresponding to every two adjacent sample conversation turns among the multiple sample conversation turns. For the sample conversation states corresponding to any two adjacent sample conversation turns, if the two sample conversation states are the same, the sample change probability corresponding to the any two sample conversation turns is 0; if the two sample conversation states are different, the sample change probability corresponding to the any two sample conversation turns is 1.

[0365] Since multiple sample conversation turns are obtained, there is a sample change probability corresponding to every two adjacent sample conversation turns. For any preset dimension identifier, multiple sample change probabilities can be obtained.

[0366] For example, the multiple obtained sample dialogue turns include: sample dialogue turn 1, sample dialogue turn 2, sample dialogue turn 3, sample dialogue turn 4, and sample dialogue turn 5. Each sample dialogue turn is respectively used as the target sample dialogue turn. According to the sample dialogue states corresponding to each sample dialogue turn, 4 sample change probabilities can be obtained. There is one sample change probability corresponding to sample dialogue turn 1 and sample dialogue turn 2, one sample change probability corresponding to sample dialogue turn 2 and sample dialogue turn 3, one sample change probability corresponding to sample dialogue turn 3 and sample dialogue turn 4, and one sample change probability corresponding to sample dialogue turn 4 and sample dialogue turn 5.

[0367] Step 2: For each preset dimension identifier, call the state transition model to process the first fusion feature vectors of every two adjacent sample dialogue turns corresponding to the preset dimension identifier respectively, and obtain multiple predicted change probabilities corresponding to the preset dimension identifier.

[0368] Among them, the state transition model is used to obtain the predicted change probability according to the first fusion feature vectors of two adjacent sample dialogue turns.

[0369] For the multiple obtained sample dialogue turns, each sample dialogue turn is used as the target sample dialogue turn, and the above steps 503 - 505 are executed, then the first fusion feature vector of each sample dialogue turn can be obtained. For any preset dimension identifier, by calling the state transition model, one predicted change probability can be obtained between every two adjacent sample dialogue turns among the multiple sample dialogue turns. Then, for this preset dimension identifier, multiple preset change probabilities can be obtained, and the number of the preset change probabilities is equal to the number of the multiple sample dialogue turns minus 1. The multiple sample change probabilities corresponding to the same preset dimension identifier and the multiple predicted change probabilities are in one-to-one correspondence.

[0370] Correspondingly, step 507 may include:

[0371] According to the multiple predicted change probabilities of each preset dimension identifier, the multiple sample change probabilities of each preset dimension identifier, the predicted dialogue state of each sample dialogue turn and the corresponding sample dialogue state, adjust the word vector acquisition sub-model, the first attention layer, the second attention layer, the feature fusion layer, the dialogue state determination sub-model and the called state transition model.

[0372] The computer device can determine the loss values of the word vector acquisition sub-model, the first attention layer, the second attention layer, the feature fusion layer, and the dialogue state determination sub-model based on the differences between the predicted dialogue state and the corresponding sample dialogue state, and the differences between the multiple predicted change probabilities of each preset dimension identifier and the multiple sample change probabilities of each preset dimension identifier. Then, the word vector acquisition sub-model, the first attention layer, the second attention layer, the feature fusion layer, and the dialogue state determination sub-model are adjusted according to the loss values to improve the accuracy of the word vector acquisition sub-model, the first attention layer, the second attention layer, the feature fusion layer, and the dialogue state determination sub-model.

[0373] Optionally, for any preset dimension identifier s, according to the first fusion feature vector corresponding to the preset dimension identifier s and the t-th sample dialogue turn and the sample keyword corresponding to the t-th sample dialogue turn of the word vector to determine the first loss value j corresponding to the preset dimension identifier s and the t-th sample dialogue turn, satisfying the following relationship:

[0374]

[0375]

[0376] where U represents the question sentence in the sample dialogue turn, R represents the reply sentence in the sample dialogue turn, U≤t represents the question sentences from the first sample dialogue turn to the t-th sample dialogue turn, R≤t represents the reply sentences from the first sample dialogue turn to the t-th sample dialogue turn, represents any preset keyword in the multiple preset keyword sets Z corresponding to the preset dimension identifier s, exp represents the exponential function with the natural constant e as the base, ||||2 represents the norm, and O s,t represents the feature vector obtained by linearly transforming the first fusion feature vector corresponding to the preset dimension identifier s and the t-th sample dialogue turn , Dropout( ) is used to represent the random selection and discard function, Linear() is used to represent the linear transformation function, and LayerNorm() is used to represent the normalization function.

[0377] Correspondingly, according to multiple preset dimension identifiers and multiple sample dialogue turns, the total first loss value L corresponding to the predicted dialogue state and the corresponding sample dialogue state is obtained dst , satisfying the following relationship:

[0378]

[0379] Among them, Y represents a set of multiple preset dimension identifiers, and G represents the total number of multiple dialogue turns. represents the sample keyword corresponding to the t-th dialogue turn corresponding to any preset dimension identifier s, where s represents any one of the multiple preset dimension identifiers, U represents the question sentence in the sample dialogue turn, R represents the reply sentence in the sample dialogue turn, U ≤ t represents the question sentences from the first sample dialogue turn to the t-th sample dialogue turn, and R ≤ t represents the reply sentences from the first sample dialogue turn to the t-th sample dialogue turn.

[0380] Optionally, for the preset dimension identifier s, the preset change probability between the t-th sample dialogue turn and the (t - 1)-th sample dialogue turn

[0381]

[0382]

[0383] Among them, σ( ) represents the sigmoid (logistic regression) activation function, and W p represents the adjustment parameter matrix, which can be any constant matrix. The adjustment parameter matrix W p is a matrix of d rows and d columns. represents the first fused feature vector after non-linear transformation of the feature vector. represents the first fused feature vector after non-linear transformation of the feature vector. ⊙ represents multiplying vectors dimension by dimension, and W c represents the adjustment parameter vector, which can be any constant vector. The adjustment parameter matrix W c is a column vector of 2d rows. tanh() is used to represent the non-linear transformation function.

[0384] Correspondingly, according to the multiple predicted change probabilities of each preset dimension identifier and the multiple sample change probabilities of each preset dimension identifier, the total second loss value L stp corresponding to the predicted change probability and the sample change probability is obtained, satisfying the following relationship:

[0385]

[0386] Among them, represents the preset change probability between the t-th sample dialogue turn and the (t - 1)-th sample dialogue turn, represents the sample change probability between the t-th sample dialogue turn and the (t - 1)-th sample dialogue turn, Y represents a set of multiple preset dimension identifiers, and G represents the total number of multiple dialogue turns.

[0387] Accordingly, based on the multiple predicted change probabilities of each preset dimension identifier, the multiple sample change probabilities of each preset dimension identifier, the predicted dialogue state of each sample dialogue turn, and the corresponding sample dialogue state, the total loss value L of the determined model joint satisfies the following relationship:

[0388] L joint = L dst + L stp

[0389] wherein, L dst represents the total sum of the first loss values corresponding to the predicted dialogue state and the corresponding sample dialogue state, and L stp represents the total sum of the second loss values corresponding to the predicted change probability and the sample change probability.

[0390] Based on the total loss value L of the determined model joint the word vector acquisition sub-model, the first attention layer, the second attention layer, the feature fusion layer, the dialogue state determination sub-model, and the call state transition model are adjusted.

[0391] It should be noted that in the embodiments of the present application, the feature vector acquisition sub-model includes the first attention layer, the second attention layer, and the feature fusion layer and the dialogue state determination sub-model for illustration. In another embodiment, the feature vector acquisition sub-model can also exist in a separate structure, and then the feature vector acquisition sub-model can be directly trained.

[0392] The method provided by the embodiments of the present application adjusts the word vector acquisition sub-model, the feature vector acquisition sub-model, and the dialogue state determination sub-model by using the sample dialogue statements of multiple sample dialogue turns, enriching the sample dialogue statements for training and improving the accuracy of the trained model. Moreover, during the training process of the word vector acquisition sub-model, the feature vector acquisition sub-model, and the dialogue state determination sub-model, a state transition model is added. Through the joint training of the state transition model, the word vector acquisition sub-model, the feature vector acquisition sub-model, and the dialogue state determination sub-model, the correlation between multiple sample dialogue turns is enhanced, the ability of the model to obtain features from historical sample dialogue turns is improved, and the accuracy of the word vector acquisition sub-model, the feature vector acquisition sub-model, and the dialogue state determination sub-model is improved.

[0393] Table 1 shows the differences between the dialogue state determination model trained by the embodiments of the present application and the models in other existing technologies, as shown in Table 1. From the data in the table, it can be seen that the dialogue filling determination model of the present application has higher joint precision and slot precision than the models in the prior art on different data sets.

[0394] Table 1

[0395]

[0396] Table 2

[0397]

[0398]

[0399] In the dialogue domain dataset 2.1, by comparing the accuracy of the model obtained by jointly training the dialogue state determination model and the state transition model with the model obtained by separately training the dialogue state determination model, it can be seen that the accuracy of the dialogue state determination model obtained by jointly training the dialogue state determination model and the state transition model is high, as shown in Table 2.

[0400] For the dialogue state determination model in this application, the dialogue state determination model includes a word vector acquisition sub-model, a feature vector acquisition sub-model, and a dialogue state determination sub-model. The feature vector acquisition sub-model includes a first attention layer, a feature combination layer, a third attention layer, a second attention layer, and a feature fusion layer. By removing different sub-models or layers, the dialogue state determination model is trained. As shown in Table 3, it can be seen from the data in Table 3 that the accuracy of the dialogue state determination model of this application is the highest.

[0401] Table 3

[0402] Model Dialogue Domain Dataset 2.1 Dialogue State Determination Model 57 Remove Feature Fusion Layer 56.76(-0.24) Remove Second Attention Layer 56.85(-0.15) Remove Third Attention Layer 55.28(-1.72) Remove Feature Fusion Layer, Second Attention Layer, Third Attention Layer 50.28(-6.72)

[0403] In the embodiment of this application, during the model training process, the dialogue statements of historical dialogue turns are used for training. By comparing the training of the model with various training methods, it can be seen that the accuracy of the model obtained by using the dialogue statements of historical dialogue turns for training is high. As shown in Table 4.

[0404] Table 4

[0405] Improvement Type Accuracy Historical Information Reasoning Improvement 64.49% Current Information Reasoning Improvement 34.86% Other Types of Information 0.65%

[0406] Figure 6 It is a flowchart of a method for training a dialogue state determination model provided by an embodiment of this application. The dialogue state determination model includes a word vector acquisition sub-model, a feature vector acquisition sub-model, and a dialogue state determination sub-model. The feature vector acquisition sub-model includes a first attention layer, a feature combination layer, a third attention layer, a second attention layer, and a feature fusion layer. Other feature extraction models are already trained models, which can be used to obtain the dimension vectors of each preset dimension identifier and the word vectors of multiple preset keywords corresponding to each preset dimension identifier.

[0407] During the training process of the model, the state transition model is jointly trained with the dialogue state determination model.

[0408] Dialogue statements of multiple sample dialogue turns are processed through a word vector acquisition sub-model to obtain a set of word vectors for each sample dialogue turn. The set of word vectors for each sample dialogue turn and the dimension vectors of the preset dimension identifiers are input into the first attention layer to obtain the first feature vector for each sample dialogue turn. The first feature vectors of multiple sample dialogue turns are input into the combined feature layer to obtain a combined feature vector for multiple sample dialogue turns, with the first feature vector of each sample dialogue turn serving as one dimension of the combined feature vector. The combined feature vector and the position vector of each sample dialogue turn are input into the third attention layer, which adjusts the first feature vector of each sample dialogue turn to obtain an updated combined feature vector. The updated combined feature vector and the dimension vectors of the preset dimension identifiers are input into the second attention layer, and the second attention layer outputs the second sample fusion feature vector corresponding to the preset dimension identifier. The second sample fusion feature vector and the first feature vector of the t-th sample dialogue turn are input into the feature fusion layer to obtain the first sample fusion feature vector corresponding to the preset dimension identifier.

[0409] The first sample fusion feature vector corresponding to the preset dimension identifier and the word vectors of multiple preset keywords of the preset dimension identifier are input into the dialogue state determination sub-model to obtain the total first loss value. The first sample fusion feature vector corresponding to the preset dimension identifier and the first sample fusion feature vector of the previous sample dialogue turn are input into the state transition model to obtain the total second loss value. The sum of the total first loss value and the total second loss value is used as the total loss value of the state transition model and the dialogue state determination model. Based on this total loss value, the state transition model and the dialogue state determination model are adjusted.

[0410] Figure 7 It is a flowchart of a method for training a dialogue state determination model provided by an embodiment of the present application. During the training process of the dialogue state determination model, a state transition model and a first bidirectional encoding representation model are required. The first bidirectional encoding representation model is a pre-trained model. Through the joint training of the dialogue state determination model and the state transition model, an accurate dialogue state determination model can be obtained. Among them, the first bidirectional encoding representation model is used to obtain the dimension vectors of each preset dimension identifier and the word vectors of preset keywords.

[0411] The dialogue state determination model includes a bidirectional encoding representation layer, a word-level multi-head attention layer, a feature combination layer, a context encoding layer, a sentence-level multi-head attention layer, a gate control layer, and a linear transformation layer. In an embodiment of the present application, the word vector acquisition sub-model is the bidirectional encoding representation layer, the first attention layer is the word-level multi-head attention layer, the third attention layer is the context encoding layer, the second attention layer is the sentence-level multi-head attention layer, the feature fusion layer is the gate control layer, and the dialogue state determination sub-model includes the linear transformation layer.

[0412] Obtain the dialogue statements of multiple sample dialogue turns, and process the dialogue statements of multiple sample dialogue turns through a bidirectional encoding representation layer, a word-level multi-head attention layer, a feature combination layer, a context encoding layer, a sentence-level multi-head attention layer, and a gate control layer. After the gate control layer outputs the first sample fusion feature vector, input the first sample fusion feature vector into a linear transformation layer and a state transition model respectively.

[0413] The linear transformation layer performs a linear transformation on the first sample fusion feature vector to obtain the transformed first sample fusion feature vector, and obtains the total first loss value through the transformed first sample fusion feature vector and the word vector of the preset keyword.

[0414] The state transition model obtains the total second loss value according to the first sample fusion feature vector and the first sample fusion feature vector of the previous sample dialogue turn.

[0415] Take the sum value of the total first loss value and the total second loss value as the total loss value of the state transition model and the dialogue state determination model, and adjust the state transition model and the dialogue state determination model according to the total loss value.

[0416] Figure 8 It is a schematic structural diagram of a dialogue state determination device provided by an embodiment of the present application. As Figure 8 shown, the device includes:

[0417] A dialogue statement acquisition module 801, configured to acquire the dialogue statements of multiple dialogue turns, where the dialogue statements of each dialogue turn include at least one of a question statement or a reply statement, and the multiple dialogue turns include a target dialogue turn and at least one historical dialogue turn before the target dialogue turn;

[0418] A first set acquisition module 802, configured to acquire the word vector set of each dialogue turn according to the dialogue statement of each dialogue turn, and the word vector set of the dialogue turn includes the word vectors of each word in the dialogue statement of the dialogue turn;

[0419] A first fusion processing module 803, configured to perform weighted fusion processing on the word vector sets of multiple dialogue turns according to the similarity between the dimension vector of each preset dimension identifier and each word vector in the word vector sets of multiple dialogue turns, and obtain the first fusion feature vector corresponding to each preset dimension identifier;

[0420] The first determination module 804 is configured to determine at least one target dimension identifier corresponding to the target conversation turn and a target keyword corresponding to each target dimension identifier according to the similarity between the word vector of the preset keyword corresponding to each preset dimension identifier and the corresponding first fusion feature vector;

[0421] The first determination module 804 is further configured to determine at least one target dimension identifier and a target keyword corresponding to each target dimension identifier as the conversation state of the target conversation turn.

[0422] In a possible implementation manner, as Figure 9 shown, the first fusion processing module 803 includes:

[0423] The first fusion processing unit 8301 is configured to, for each preset dimension identifier, perform weighted fusion processing on multiple word vectors in the word vector set of each conversation turn according to the similarity between the dimension vector of the preset dimension identifier and the multiple word vectors in the word vector set of each conversation turn, so as to obtain a first feature vector of each conversation turn;

[0424] The second fusion processing unit 8302 is configured to perform weighted fusion processing on the first feature vectors of multiple conversation turns according to the similarity between the dimension vector of the preset dimension identifier and the first feature vectors of multiple conversation turns, so as to obtain a second fusion feature vector corresponding to the preset dimension identifier;

[0425] The third fusion processing unit 8303 is configured to perform fusion processing on the second fusion feature vector and the first feature vector of the target conversation turn to obtain a first fusion feature vector corresponding to the preset dimension identifier.

[0426] In a possible implementation manner, the first fusion processing unit 8301 is configured to, for each preset dimension identifier and each conversation turn, respectively determine a first similarity between the dimension vector of the preset dimension identifier and each word vector in the word vector set of the conversation turn; determine a first weight of each word vector according to the first similarity corresponding to each word vector in the word vector set of the conversation turn, and the first weight of each word vector is positively correlated with the corresponding first similarity; perform weighted fusion processing on the multiple word vectors according to the first weights of the multiple word vectors in the word vector set to obtain a first feature vector of the conversation turn.

[0427] In a possible implementation, the second fusion processing unit 8302 is configured to respectively determine the second similarity between the dimension vectors of the preset dimension identifiers and the first feature vectors of each conversation turn; determine the second weight of each conversation turn according to the second similarity corresponding to each conversation turn, and the second weight of each conversation turn is positively correlated with the corresponding second similarity; perform weighted fusion processing on the first feature vectors of multiple conversation turns according to the second weights of multiple conversation turns to obtain the second fusion feature vector corresponding to the preset dimension identifier.

[0428] In a possible implementation, as Figure 9 shown, the apparatus further includes:

[0429] The first adjustment module 805 is configured to perform weighted fusion processing on the first feature vectors of multiple conversation turns according to the third similarity between the first feature vector of any conversation turn and the first feature vectors of multiple conversation turns, and use the feature vector after the fusion processing as the adjusted first feature vector of any conversation turn.

[0430] In a possible implementation, as Figure 9 shown, the first adjustment module 805 includes:

[0431] The weight determination unit 8501 is configured to determine the third weight of each conversation turn among multiple conversation turns according to the third similarity between the first feature vector of any conversation turn and the first feature vectors of multiple conversation turns, and the third weight corresponding to each conversation turn among multiple conversation turns is positively correlated with the corresponding third similarity;

[0432] The fourth fusion processing unit 8502 is configured to perform weighted fusion processing on the first feature vectors of multiple conversation turns according to the third weights corresponding to multiple conversation turns, and use the feature vector after the fusion processing as the adjusted first feature vector of any conversation turn.

[0433] In a possible implementation, as Figure 9 shown, the apparatus further includes:

[0434] The position vector acquisition module 806 is configured to acquire the position vector of each conversation turn, and the position vector of the conversation turn is used to represent the position of the conversation turn in multiple conversation turns;

[0435] The second fusion processing module 807 is configured to fuse the first feature vector of each conversation turn with the corresponding position vector, and use the feature vector after the fusion processing as the adjusted first feature vector of each conversation turn.

[0436] In a possible implementation manner, the third fusion processing unit 8303 is configured to determine a fourth weight of the first feature vector of the target dialogue turn and a fifth weight of the second fusion feature vector, where the sum of the fourth weight and the fifth weight is 1; and perform weighted fusion processing on the second fusion feature vector and the first feature vector of the target dialogue turn according to the fourth weight and the fifth weight to obtain a first fusion feature vector corresponding to the preset dimension identifier.

[0437] In a possible implementation manner, as Figure 9 shown, the first determination module 804 includes:

[0438] The similarity determination unit 8401 is configured to determine a fourth similarity between the word vectors of multiple preset keywords corresponding to a preset dimension identifier and the first fusion feature vector for each preset dimension identifier;

[0439] The keyword selection unit 8402 is configured to select alternative keywords from the multiple preset keywords, where the fourth similarity corresponding to the alternative keywords is greater than the fourth similarity corresponding to other preset keywords among the multiple preset keywords;

[0440] The target determination unit 8403 is configured to, in response to the alternative keywords not being empty keywords, determine the preset dimension identifier as the target dimension identifier and determine the alternative keywords as the target keywords.

[0441] Figure 10 FIG. is a schematic structural diagram of a dialogue state determination device provided by an embodiment of the present application. As Figure 10 shown, the device includes:

[0442] The dialogue statement acquisition module 1001 is configured to acquire dialogue statements of multiple dialogue turns, where the dialogue statement of each dialogue turn includes at least one of a question statement or a reply statement, and the multiple dialogue turns include a target dialogue turn and at least one historical dialogue turn before the target dialogue turn;

[0443] The first set acquisition module 1002 is configured to call the word vector acquisition sub-model in the dialogue state determination model and acquire a word vector set for each dialogue turn according to the dialogue statement of each dialogue turn;

[0444] The first fusion processing module 1003 is configured to call the feature vector acquisition sub-model in the dialogue state determination model and perform weighted fusion processing on the word vector sets of multiple dialogue turns according to the similarity between the dimension vector of each preset dimension identifier and each word vector in the word vector sets of multiple dialogue turns to obtain a first fusion feature vector corresponding to each preset dimension identifier;

[0445] The first determination module 1004 is configured to call the dialogue state determination sub-model in the dialogue state determination model, and determine at least one target dimension identifier corresponding to the target dialogue turn and the target keyword corresponding to each target dimension identifier according to the similarity between the word vectors of the preset keywords corresponding to each preset dimension identifier and the corresponding first fusion feature vectors; and determine the at least one target dimension identifier and the target keyword corresponding to each target dimension identifier as the dialogue state of the target dialogue turn.

[0446] In a possible implementation manner, as Figure 11 shown, the first fusion processing module 1003 includes:

[0447] The first fusion processing unit 1031 is configured to call the first attention layer in the feature vector acquisition sub-model, and perform weighted fusion processing on multiple word vectors in the word vector set of each dialogue turn according to the similarity between the dimension vector of the preset dimension identifier and the multiple word vectors in the word vector set of each dialogue turn, respectively, to obtain the first feature vector of each dialogue turn;

[0448] The second fusion processing unit 1032 is configured to call the second attention layer in the feature vector acquisition sub-model, and perform weighted fusion processing on the first feature vectors of multiple dialogue turns according to the similarity between the dimension vector of the preset dimension identifier and the first feature vectors of multiple dialogue turns, to obtain the second fusion feature vector corresponding to the preset dimension identifier;

[0449] The third fusion processing unit 1033 is configured to call the feature fusion layer in the feature vector acquisition sub-model, and perform fusion processing on the second fusion feature vector and the first feature vector of the target dialogue turn to obtain the first fusion feature vector corresponding to the preset dimension identifier.

[0450] In a possible implementation manner, as Figure 11 shown, the apparatus further includes:

[0451] The first adjustment module 1005 is configured to call the third attention layer in the feature vector acquisition sub-model, and perform weighted fusion processing on the word vector sets of multiple dialogue turns according to the third similarity between the first feature vector of any dialogue turn and the first feature vectors of multiple dialogue turns, and use the fused feature vector as the first feature vector of the dialogue turn after adjustment.

[0452] In a possible implementation manner, as Figure 11 shown, the apparatus further includes:

[0453] The second adjustment module 1006 is configured to call the third attention layer in the feature vector acquisition sub-model to obtain the position vector of each conversation turn, where the position vector of the conversation turn is used to represent the position of the conversation turn among multiple conversation turns; fuse the first feature vector of each conversation turn with the corresponding position vector, and use the fused feature vector as the adjusted first feature vector of each conversation turn.

[0454] In a possible implementation manner, as Figure 11 shown, the apparatus further includes:

[0455] The conversation statement acquisition module 1001 is further configured to obtain the sample conversation statements of multiple sample conversation turns, and the sample conversation status corresponding to each sample conversation turn. The sample conversation statement of each sample conversation turn includes at least one of a sample user question statement or a sample reply statement. The multiple sample conversation turns include a target sample conversation turn and at least one historical sample conversation turn before the target sample conversation turn;

[0456] The first set acquisition module 1002 is further configured to call the word vector acquisition sub-model, and obtain the word vector set of each sample conversation turn according to the sample conversation statement of each sample conversation turn. The word vector set of the sample conversation turn includes the word vectors of each word in the sample conversation statement of the sample conversation turn;

[0457] The first fusion processing module 1003 is further configured to call the feature vector acquisition sub-model, and perform weighted fusion processing on the word vector sets of the target sample conversation turn and the corresponding historical sample conversation turns according to the similarity between the dimension vectors of each preset dimension identifier and each word vector in the word vector sets of the target sample conversation turn and the corresponding historical sample conversation turns, to obtain the first sample fusion feature vector corresponding to each preset dimension identifier;

[0458] The first determination module 1004 is further configured to call the conversation state determination sub-model, and determine at least one predicted dimension identifier corresponding to the target sample conversation turn, and the predicted keyword corresponding to each predicted dimension identifier according to the word vectors of multiple preset keywords corresponding to each preset dimension identifier, and the corresponding first sample fusion feature vector; determine at least one predicted dimension identifier, and the predicted keyword corresponding to each predicted dimension identifier, as the predicted conversation state of the target sample conversation turn;

[0459] The model adjustment module 1007 is configured to adjust the word vector acquisition sub-model, the feature vector acquisition sub-model, and the conversation state determination sub-model according to the predicted conversation state of the target sample conversation turn and the corresponding sample conversation state.

[0460] In a possible implementation manner, as Figure 11 shown, the apparatus further includes:

[0461] A second determination module 1008, configured to use any one of the multiple sample conversation turns as a target sample conversation turn, and use the sample conversation turns before any one of the sample conversation turns as the historical sample conversation turns of the target sample conversation turn.

[0462] In a possible implementation, as Figure 11 shown, the apparatus includes:

[0463] A change probability acquisition module 1009, configured to acquire multiple sample change probabilities for each preset dimension identifier, where the sample change probability represents the difference between the sample conversation states corresponding to every two adjacent sample conversation turns among the multiple sample conversation turns;

[0464] A feature vector processing module 1010, configured to, for each preset dimension identifier, call a state transition model to respectively process the first fusion feature vectors of every two adjacent sample conversation turns corresponding to the preset dimension identifier, and obtain multiple predicted change probabilities corresponding to the preset dimension identifier. There is a one-to-one correspondence between the multiple sample change probabilities and the multiple predicted change probabilities corresponding to the same preset dimension identifier;

[0465] A model adjustment module 1007 includes:

[0466] A model adjustment unit 1071, configured to adjust the word vector acquisition sub-model, the feature vector acquisition sub-model, the conversation state determination sub-model, and the called state transition model according to the obtained multiple predicted change probabilities for each preset dimension identifier, the multiple sample change probabilities for each preset dimension identifier, the predicted conversation state of each sample conversation turn, and the corresponding sample conversation state.

[0467] Figure 12 FIG. shows a structural block diagram of an electronic device 1200 provided by an exemplary embodiment of the present application. The electronic device 1200 may be a portable mobile terminal, such as: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer, or a desktop computer. The electronic device 1200 may also be referred to by other names such as a user device, a portable terminal, a laptop terminal, a desktop terminal, etc.

[0468] Generally, the electronic device 1200 includes a processor 1201 and a memory 1202.

[0469] The processor 1201 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1201 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1201 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 1201 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1201 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0470] The memory 1202 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1202 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1202 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 1201 to implement the dialogue state determination method provided in the method embodiments of the present application.

[0471] In some embodiments, the electronic device 1200 may further optionally include: a peripheral device interface 1203 and at least one peripheral device. The processor 1201, the memory 1202, and the peripheral device interface 1203 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1203 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 1204, a display screen 1205, a camera assembly 1206, an audio circuit 1207, a positioning component 1208, and a power supply 1209.

[0472] The peripheral device interface 1203 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1201 and the memory 1202. In some embodiments, the processor 1201, the memory 1202, and the peripheral device interface 1203 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1201, the memory 1202, and the peripheral device interface 1203 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0473] The radio frequency circuit 1204 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1204 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1204 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1204 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 1204 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1204 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.

[0474] The display screen 1205 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1205 is a touch display screen, the display screen 1205 also has the ability to collect touch signals on or above the surface of the display screen 1205. The touch signals can be input as control signals to the processor 1201 for processing. At this time, the display screen 1205 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there may be one display screen 1205, which is provided on the front panel of the electronic device 1200; in other embodiments, there may be at least two display screens 1205, which are respectively provided on different surfaces of the electronic device 1200 or are in a folding design; in other embodiments, the display screen 1205 may be a flexible display screen, which is provided on the curved surface or folding surface of the electronic device 1200. Even, the display screen 1205 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 1205 can be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0475] The camera module 1206 is used to collect images or videos. Optionally, the camera module 1206 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera respectively, so as to realize the function of background blurring by fusing the main camera and the depth camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera module 1206 may also include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. The two-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0476] The audio circuit 1207 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1201 for processing, or input to the radio frequency circuit 1204 to implement voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device 1200. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1201 or the radio frequency circuit 1204 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1207 may further include a headphone jack.

[0477] The positioning component 1208 is used to locate the current geographical location of the electronic device 1200 to implement navigation or LBS (Location Based Service). The positioning component 1208 may be a positioning component based on the US GPS (Global Positioning System), the Chinese Beidou system or the Russian Galileo system.

[0478] The power supply 1209 is used to supply power to each component in the electronic device 1200. The power supply 1209 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 1209 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0479] In some embodiments, the electronic device 1200 further includes one or more sensors 1210. The one or more sensors 1210 include but are not limited to: an acceleration sensor 1211, a gyroscope sensor 1212, a pressure sensor 1213, a fingerprint sensor 1214, an optical sensor 1215 and a proximity sensor 1216.

[0480] The acceleration sensor 1211 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the electronic device 1200. For example, the acceleration sensor 1211 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1201 can control the display screen 1205 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1211. The acceleration sensor 1211 can also be used for collecting game or user's motion data.

[0481] The gyroscope sensor 1212 can detect the body direction and rotation angle of the electronic device 1200. The gyroscope sensor 1212 can cooperate with the acceleration sensor 1211 to collect the 3D actions of the user on the electronic device 1200. Based on the data collected by the gyroscope sensor 1212, the processor 1201 can implement the following functions: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.

[0482] The pressure sensor 1213 can be disposed on the side frame of the electronic device 1200 and / or the lower layer of the display screen 1205. When the pressure sensor 1213 is disposed on the side frame of the electronic device 1200, it can detect the holding signal of the user on the electronic device 1200, and the processor 1201 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 1213. When the pressure sensor 1213 is disposed on the lower layer of the display screen 1205, the processor 1201 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 1205. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0483] The fingerprint sensor 1214 is used to collect the fingerprint of the user. The processor 1201 can identify the user's identity according to the fingerprint collected by the fingerprint sensor 1214, or the fingerprint sensor 1214 can identify the user's identity according to the collected fingerprint. When the identified user identity is a trusted identity, the processor 1201 authorizes the user to perform relevant sensitive operations, and the sensitive operations include unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings, etc. The fingerprint sensor 1214 can be disposed on the front, back, or side of the electronic device 1200. When there is a physical button or a manufacturer's logo on the electronic device 1200, the fingerprint sensor 1214 can be integrated with the physical button or the manufacturer's logo.

[0484] The optical sensor 1215 is used to collect the ambient light intensity. In one embodiment, the processor 1201 can control the display brightness of the display screen 1205 according to the ambient light intensity collected by the optical sensor 1215. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1205 is increased; when the ambient light intensity is low, the display brightness of the display screen 1205 is decreased. In another embodiment, the processor 1201 can also dynamically adjust the shooting parameters of the camera module 1206 according to the ambient light intensity collected by the optical sensor 1215.

[0485] The proximity sensor 1216, also known as the distance sensor, is usually disposed on the front panel of the electronic device 1200. The proximity sensor 1216 is used to collect the distance between the user and the front side of the electronic device 1200. In one embodiment, when the proximity sensor 1216 detects that the distance between the user and the front side of the electronic device 1200 is gradually decreasing, the processor 1201 controls the display screen 1205 to switch from the lit state to the off state; when the proximity sensor 1216 detects that the distance between the user and the front side of the electronic device 1200 is gradually increasing, the processor 1201 controls the display screen 1205 to switch from the off state to the lit state.

[0486] Those skilled in the art can understand that Figure 12 the structure shown in does not constitute a limitation on the electronic device 1200, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements.

[0487] Figure 13 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 1300 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 1301 and one or more memories 1302. Among them, at least one instruction is stored in the memory 1302, and the at least one instruction is loaded and executed by the processor 1301 to implement the methods provided by the above-mentioned various method embodiments. Of course, the server may also have components such as wired or wireless network interfaces, keyboards, and input / output interfaces for input / output. The server may also include other components for implementing the functions of the device, which will not be elaborated here.

[0488] The server 1300 can be used to execute the above-mentioned dialogue state determination method.

[0489] An embodiment of the present application also provides a computer device, which includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the dialogue state determination method of the above embodiment.

[0490] An embodiment of the present application also provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by the processor to implement the dialogue state determination method of the above embodiment.

[0491] An embodiment of the present application also provides a computer program, in which at least one instruction is stored, and the at least one instruction is loaded and executed by the processor to implement the dialogue state determination method of the above embodiment.

[0492] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc.

[0493] The above are only optional embodiments of the embodiments of the present application, and are not intended to limit the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the embodiments of the present application shall be included in the protection scope of the present application.

Claims

1. A method for determining a dialogue state, characterized in that, The method includes: Obtaining dialogue statements of multiple dialogue turns, where the dialogue statements of each dialogue turn include at least one of a question statement or a response statement, and the multiple dialogue turns include a target dialogue turn and at least one historical dialogue turn before the target dialogue turn; Obtaining a set of word vectors for each dialogue turn according to the dialogue statements of each dialogue turn, where the set of word vectors for the dialogue turn includes word vectors of each word in the dialogue statements of the dialogue turn; Performing weighted fusion processing on the sets of word vectors of the multiple dialogue turns according to the similarity between the dimension vector of each preset dimension identifier and each word vector in the sets of word vectors of the multiple dialogue turns, to obtain a first fusion feature vector corresponding to each preset dimension identifier; Determining at least one target dimension identifier corresponding to the target dialogue turn and a target keyword corresponding to each target dimension identifier according to the similarity between the word vector of the preset keyword corresponding to each preset dimension identifier and the corresponding first fusion feature vector; Determining the at least one target dimension identifier and the target keyword corresponding to each target dimension identifier as the dialogue state of the target dialogue turn.

2. The method according to claim 1, characterized in that The performing weighted fusion processing on the sets of word vectors of the multiple dialogue turns according to the similarity between the dimension vector of each preset dimension identifier and each word vector in the sets of word vectors of the multiple dialogue turns, to obtain a first fusion feature vector corresponding to each preset dimension identifier includes: For each preset dimension identifier, performing weighted fusion processing on multiple word vectors in the set of word vectors of each dialogue turn according to the similarity between the dimension vector of the preset dimension identifier and the multiple word vectors in the set of word vectors of each dialogue turn, to obtain a first feature vector of each dialogue turn; Performing weighted fusion processing on the first feature vectors of the multiple dialogue turns according to the similarity between the dimension vector of the preset dimension identifier and the first feature vectors of the multiple dialogue turns, to obtain a second fusion feature vector corresponding to the preset dimension identifier; Performing fusion processing on the second fusion feature vector and the first feature vector of the target dialogue turn to obtain a first fusion feature vector corresponding to the preset dimension identifier.

3. The method according to claim 2, wherein Before the performing weighted fusion processing on the first feature vectors of the multiple dialogue turns according to the similarity between the dimension vector of the preset dimension identifier and the first feature vectors of the multiple dialogue turns, to obtain a second fusion feature vector corresponding to the preset dimension identifier, the method further includes: Performing weighted fusion processing on the first feature vectors of the multiple dialogue turns according to a third similarity between the first feature vector of any dialogue turn and the first feature vectors of the multiple dialogue turns, and using the feature vector after the fusion processing as the adjusted first feature vector of the any dialogue turn.

4. The method according to claim 2, wherein Before performing weighted fusion processing on the first feature vectors of the multiple conversation turns based on the similarity between the dimension vectors identified according to the preset dimensions and the first feature vectors of the multiple conversation turns to obtain the second fusion feature vectors corresponding to the preset dimension identifications, the method further includes: Obtain the position vector of each conversation turn, where the position vector of the conversation turn is used to represent the position of the conversation turn in the multiple conversation turns; Perform fusion processing on the first feature vector of each conversation turn and the corresponding position vector, and use the feature vector after the fusion processing as the adjusted first feature vector of each conversation turn.

5. The method according to claim 1, wherein The determining of at least one target dimension identification corresponding to the target conversation turn and the target keyword corresponding to each target dimension identification according to the similarity between the word vectors of the preset keywords corresponding to each preset dimension identification and the corresponding first fusion feature vectors includes: For each preset dimension identification, determine the fourth similarity between the word vectors of the multiple preset keywords corresponding to the preset dimension identification and the first fusion feature vector; Select alternative keywords from the multiple preset keywords, where the fourth similarity corresponding to the alternative keywords is greater than the fourth similarity corresponding to the other preset keywords in the multiple preset keywords; In response to the alternative keyword not being an empty keyword, determine the preset dimension identification as the target dimension identification and determine the alternative keyword as the target keyword.

6. A method for determining a dialogue state, characterized in that, The method includes: Obtain the conversation statements of multiple conversation turns, where the conversation statement of each conversation turn includes at least one of a question statement or a reply statement, and the multiple conversation turns include a target conversation turn and at least one historical conversation turn before the target conversation turn; Invoke the word vector acquisition sub-model in the conversation state determination model, and according to the conversation statement of each conversation turn, obtain the word vector set of each conversation turn, where the word vector set of the conversation turn includes the word vectors of each word in the conversation statement of the conversation turn; Invoke the feature vector acquisition sub-model in the conversation state determination model, and perform weighted fusion processing on the word vector sets of the multiple conversation turns according to the similarity between the dimension vectors of each preset dimension identification and each word vector in the word vector sets of the multiple conversation turns to obtain the first fusion feature vectors corresponding to each preset dimension identification; Invoke the conversation state determination sub-model in the conversation state determination model, and according to the similarity between the word vectors of the preset keywords corresponding to each preset dimension identification and the corresponding first fusion feature vectors, determine at least one target dimension identification corresponding to the target conversation turn and the target keyword corresponding to each target dimension identification; determine the at least one target dimension identification and the target keyword corresponding to each target dimension identification as the conversation state of the target conversation turn.

7. The method according to claim 6, characterized in that, Calling the feature vector acquisition sub-model in the dialogue state determination model, and performing weighted fusion processing on the word vector sets of the multiple dialogue turns according to the similarity between the dimension vectors identified by each preset dimension and each word vector in the word vector sets of the multiple dialogue turns, to obtain a first fusion feature vector corresponding to each preset dimension identifier, including: Calling the first attention layer in the feature vector acquisition sub-model, for each preset dimension identifier, performing weighted fusion processing on multiple word vectors in the word vector set of each dialogue turn respectively according to the similarity between the dimension vector of the preset dimension identifier and the multiple word vectors in the word vector set of each dialogue turn, to obtain a first feature vector of each dialogue turn; Calling the second attention layer in the feature vector acquisition sub-model, and performing weighted fusion processing on the first feature vectors of the multiple dialogue turns according to the similarity between the dimension vector of the preset dimension identifier and the first feature vectors of the multiple dialogue turns, to obtain a second fusion feature vector corresponding to the preset dimension identifier; Calling the feature fusion layer in the feature vector acquisition sub-model, and performing fusion processing on the second fusion feature vector and the first feature vector of the target dialogue turn, to obtain a first fusion feature vector corresponding to the preset dimension identifier.

8. The method according to claim 7, wherein Before calling the second attention layer in the feature vector acquisition sub-model, and performing weighted fusion processing on the first feature vectors of the multiple dialogue turns according to the similarity between the dimension vector of the preset dimension identifier and the first feature vectors of the multiple dialogue turns, to obtain a second fusion feature vector corresponding to the preset dimension identifier, the method further includes: Calling the third attention layer in the feature vector acquisition sub-model, and performing weighted fusion processing on the word vector sets of the multiple dialogue turns according to a third similarity between the first feature vector of any dialogue turn and the first feature vectors of the multiple dialogue turns, and using the feature vector after the fusion processing as the adjusted first feature vector of the dialogue turn.

9. The method according to claim 6, wherein The method further includes: Obtaining sample dialogue statements of multiple sample dialogue turns, and the sample dialogue state corresponding to each sample dialogue turn, the sample dialogue statements of each sample dialogue turn include at least one of a sample user question statement or a sample response statement, and the multiple sample dialogue turns include a target sample dialogue turn and at least one historical sample dialogue turn before the target sample dialogue turn; Calling the word vector acquisition sub-model, and obtaining the word vector set of each sample dialogue turn according to the sample dialogue statement of each sample dialogue turn, where the word vector set of the sample dialogue turn includes the word vectors of each word in the sample dialogue statement of the sample dialogue turn; Call the feature vector acquisition sub-model, and perform weighted fusion processing on the word vector sets of the target sample dialogue turn and the corresponding historical sample dialogue turns according to the similarity between the dimension vectors identified by each preset dimension and each word vector in the word vector sets, to obtain the first sample fusion feature vector corresponding to each preset dimension identifier; Call the dialogue state determination sub-model, and determine at least one predicted dimension identifier corresponding to the target sample dialogue turn and the predicted keywords corresponding to each predicted dimension identifier according to the word vectors of multiple preset keywords corresponding to each preset dimension identifier and the corresponding first sample fusion feature vector; determine the at least one predicted dimension identifier and the predicted keywords corresponding to each predicted dimension identifier as the predicted dialogue state of the target sample dialogue turn; Adjust the word vector acquisition sub-model, the feature vector acquisition sub-model, and the dialogue state determination sub-model according to the predicted dialogue state of the target sample dialogue turn and the corresponding sample dialogue state.

10. The method according to claim 9, characterized in that, After obtaining the sample dialogue statements of multiple sample dialogue turns and the sample dialogue state corresponding to each sample dialogue turn, the method further includes: Use any one of the multiple sample dialogue turns as the target sample dialogue turn, and use the sample dialogue turns before the any one sample dialogue turn as the historical sample dialogue turns of the target sample dialogue turn.

11. The method according to claim 10, wherein Before adjusting the word vector acquisition sub-model, the feature vector acquisition sub-model, and the dialogue state determination sub-model according to the predicted dialogue state of the target sample dialogue turn and the corresponding sample dialogue state, the method includes: Obtain the multiple sample change probabilities of each preset dimension identifier, where the sample change probability represents the difference between the sample dialogue states corresponding to every two adjacent sample dialogue turns among the multiple sample dialogue turns; For each preset dimension identifier, call the state transition model to process the first fusion feature vectors of every two adjacent sample dialogue turns corresponding to the preset dimension identifier respectively, to obtain multiple predicted change probabilities corresponding to the preset dimension identifier, and the multiple sample change probabilities and the multiple predicted change probabilities corresponding to the same preset dimension identifier correspond one by one; The adjusting the word vector acquisition sub-model, the feature vector acquisition sub-model, and the dialogue state determination sub-model according to the predicted dialogue state of the target sample dialogue turn and the corresponding sample dialogue state includes: Adjust the word vector acquisition sub-model, the feature vector acquisition sub-model, the dialogue state determination sub-model, and the called state transition model according to the obtained multiple predicted change probabilities of each preset dimension identifier, the multiple sample change probabilities of each preset dimension identifier, the predicted dialogue state of each sample dialogue turn and the corresponding sample dialogue state.

12. A dialogue state determination device, characterized in that, The device includes: A dialogue statement acquisition module, configured to acquire dialogue statements of multiple dialogue turns. Each dialogue statement of a dialogue turn includes at least one of a question statement or a reply statement. The multiple dialogue turns include a target dialogue turn and at least one historical dialogue turn before the target dialogue turn; A first set acquisition module, configured to acquire a word vector set of each dialogue turn according to the dialogue statement of each dialogue turn. The word vector set of the dialogue turn includes word vectors of each word in the dialogue statement of the dialogue turn; A first fusion processing module, configured to perform weighted fusion processing on the word vector sets of the multiple dialogue turns according to the similarity between the dimension vector of each preset dimension identifier and each word vector in the word vector sets of the multiple dialogue turns, to obtain a first fusion feature vector corresponding to each preset dimension identifier; A first determination module, configured to determine at least one target dimension identifier corresponding to the target dialogue turn and a target keyword corresponding to each target dimension identifier according to the similarity between the word vector of the preset keyword corresponding to each preset dimension identifier and the corresponding first fusion feature vector; The first determination module is further configured to determine the at least one target dimension identifier and the target keyword corresponding to each target dimension identifier as the dialogue state of the target dialogue turn.

13. A dialogue state determination device, characterized in that, The apparatus includes: A dialogue statement acquisition module, configured to acquire dialogue statements of multiple dialogue turns. Each dialogue statement of a dialogue turn includes at least one of a question statement or a reply statement. The multiple dialogue turns include a target dialogue turn and at least one historical dialogue turn before the target dialogue turn; A first set acquisition module, configured to call a word vector acquisition sub-model in a dialogue state determination model, and acquire a word vector set of each dialogue turn according to the dialogue statement of each dialogue turn. The word vector set of the dialogue turn includes word vectors of each word in the dialogue statement of the dialogue turn; A first fusion processing module, configured to call a feature vector acquisition sub-model in the dialogue state determination model, and perform weighted fusion processing on the word vector sets of the multiple dialogue turns according to the similarity between the dimension vector of each preset dimension identifier and each word vector in the word vector sets of the multiple dialogue turns, to obtain a first fusion feature vector corresponding to each preset dimension identifier; A first determination module, configured to call a dialogue state determination sub-model in the dialogue state determination model, and determine at least one target dimension identifier corresponding to the target dialogue turn and a target keyword corresponding to each target dimension identifier according to the similarity between the word vector of the preset keyword corresponding to each preset dimension identifier and the corresponding first fusion feature vector; and determine the at least one target dimension identifier and the target keyword corresponding to each target dimension identifier as the dialogue state of the target dialogue turn.

14. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one instruction is stored in the memory. The at least one instruction is loaded and executed by the processor to implement the dialogue state determination method according to any one of claims 1 to 5; or to implement the dialogue state determination method according to any one of claims 6 to 11.

15. A computer-readable storage medium, characterized in that, At least one instruction is stored in the computer-readable storage medium. The at least one instruction is loaded and executed by a processor to implement the dialogue state determination method according to any one of claims 1 to 5; or to implement the dialogue state determination method according to any one of claims 6 to 11.