Method, apparatus, and electronic device for intention recognition

By introducing virtual scenes and knowledge graphs into the intention recognition system and combining the intention recognition model, the problem of inaccurate intention recognition in the prior art is solved, and higher recognition accuracy is achieved.

CN114429142BActive Publication Date: 2025-06-20NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210102819.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-27
Publication Date
2025-06-20
Estimated Expiration
2042-01-27

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify the user's true intentions in intention recognition, especially when the same conversation may have different intentions in multiple scenarios.

Method used

By providing a graphical user interface, displaying virtual scenes and virtual characters, responding to the conversation text entered by the player, and obtaining the current scene image. Based on the preset knowledge graph, knowledge subgraphs of dialogue text and scene images are obtained, and intention recognition models are used to predict intentions.

Benefits of technology

It improves the accuracy of intention recognition, can more accurately reflect the player's real intention for virtual scenes, reduces error recognition, and enhances the accuracy of dialogue intention recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114429142B_ABST
    Figure CN114429142B_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus and electronic device for intention recognition, relating to the technical field of information processing. The method includes: in response to receiving the dialogue text input by the player corresponding to the virtual character, acquiring the current scene image corresponding to the virtual scene, and acquiring the respective knowledge sub-graphs corresponding to the dialogue text and the scene image based on a preset knowledge graph, and finally obtaining the predicted intention corresponding to the dialogue text based on the knowledge sub-graph and the intention recognition model. When the technology of the present application performs intention recognition on the dialogue text to which the player belongs, it not only considers the information of the dialogue text itself, but also considers the virtual scene where the player issues the dialogue text, takes the dialogue text and the scene image as the inputs of the intention recognition model together, and the intention determined by the intention recognition model can more accurately reflect the true intention of the player for the virtual scene, effectively avoiding the misrecognition of the player's intention and improving the accuracy of intention recognition for the dialogue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information processing, and in particular, to a method, apparatus, and electronic device for intention recognition. Background Art

[0002] With the development of artificial intelligence, in more and more scenarios, it is necessary to use artificial intelligence technology to recognize the intention of the voice information sent by the user.

[0003] A commonly used recognition method in intention recognition is: starting from the dialogue text information, classifying the dialogue text into intentions, that is, modeling it into a text classification task, and mapping the dialogue to a predefined intention. A more common way is to input the dialogue text into the BERT intention recognition model to obtain the predicted semantics corresponding to the dialogue text.

[0004] However, the above method cannot accurately recognize the true intention of the user. Therefore, how to accurately recognize the true intention behind the information currently input by the user is particularly important. Summary of the Invention

[0005] In view of this, an object of the present invention is to provide a method, apparatus, and electronic device for intention recognition to improve the accuracy of intention recognition.

[0006] In a first aspect, an embodiment of the present invention provides a method for intention recognition. The method provides a graphical user interface through an electronic device, and the content displayed on the graphical user interface includes at least a virtual scene and a virtual character; the method includes: responding to the received dialogue text input by the player corresponding to the virtual character, and obtaining the scene image currently corresponding to the virtual scene; obtaining the knowledge subgraphs corresponding to the dialogue text and the scene image respectively based on a preset knowledge graph; and obtaining the predicted intention corresponding to the dialogue text based on the knowledge subgraph and the intention recognition model.

[0007] Further, the above method further includes: extracting a plurality of sub-images including target objects from the scene image; obtaining the knowledge subgraph corresponding to the scene image based on a preset knowledge graph, including: obtaining the knowledge subgraphs corresponding to the plurality of sub-images respectively based on a preset knowledge graph.

[0008] Further, obtaining the knowledge subgraphs corresponding to the plurality of sub-images respectively based on a preset knowledge graph includes: obtaining the description text corresponding to each of the plurality of sub-images through an image description algorithm; and obtaining the knowledge subgraph corresponding to each sub-image respectively based on the preset knowledge graph and the description text corresponding to each sub-image.

[0009] Further, obtaining the knowledge subgraphs corresponding to the dialogue text and the scenario image based on the preset knowledge graph includes: obtaining the initial knowledge subgraphs corresponding to the dialogue text and the scenario image based on the preset knowledge graph; respectively updating the nodes in the initial knowledge subgraphs based on the preset graph convolutional neural network to obtain the knowledge subgraphs corresponding to the dialogue text and the scenario image.

[0010] Further, the step of obtaining the initial knowledge subgraphs corresponding to the dialogue text and the scenario image based on the preset knowledge graph includes: segmenting the description text corresponding to the dialogue text into multiple first text units, and segmenting the description text corresponding to the scenario image into multiple second text units; determining the first knowledge nodes corresponding to each first text unit and the second knowledge nodes corresponding to each second text unit from the preset knowledge graph; determining the first sub-knowledge nodes whose distances from the first knowledge nodes are less than the preset distance threshold, and the second sub-knowledge nodes whose distances from the second knowledge nodes are less than the preset distance threshold from the preset knowledge graph; determining the subgraph composed of the first knowledge nodes and the first sub-knowledge nodes as the initial knowledge subgraph corresponding to the dialogue text, and determining the subgraph composed of the second knowledge nodes and the second sub-knowledge nodes as the initial knowledge subgraph corresponding to the scenario image.

[0011] Further, obtaining the predicted intention corresponding to the dialogue text based on the knowledge subgraph and the intention recognition model includes: generating a first sequence according to the dialogue text and multiple sub-images; generating a second sequence according to the first type information of the dialogue text and the second type of multiple sub-images; encoding the positions of each object in the first sequence to obtain the encoded representation of each object in the first sequence, and obtaining a third sequence according to the encoded representation; determining a fourth sequence according to the knowledge subgraph; inputting the first sequence, the second sequence, the third sequence, and the fourth sequence into the intention recognition model, and processing through the intention recognition model to obtain the predicted intention corresponding to the dialogue text.

[0012] Further, the step of determining the fourth sequence according to the knowledge subgraph includes: determining the target objects that appear in both the knowledge subgraph corresponding to the dialogue text and the knowledge subgraph corresponding to any sub-image in the first sequence; taking the encoded identifier corresponding to the target object as the knowledge representation of the target object; taking the knowledge subgraphs corresponding to the multiple sub-images as the knowledge representations of the corresponding sub-images; obtaining the fourth sequence according to the knowledge representation.

[0013] Further, a semantic library is also pre-stored in the above electronic device; the step of obtaining the predicted intention corresponding to the dialogue text based on the knowledge subgraph and the intention recognition model includes: determining the confidence of each semantic in the semantic library according to the knowledge subgraph and the intention recognition model; where the confidence is used to represent the probability that the semantic can reflect the true intention of the dialogue text; determining the predicted intention matching the dialogue text according to the confidence of the semantic.

[0014] Further, the above-mentioned intent recognition model includes a semantic representation sub-model and a classifier, and the intent recognition model is trained by the following method: obtaining sample data; wherein, the sample data includes training texts and the true semantics corresponding to the training texts; performing semantic prediction on the training texts through the initial intent recognition model to obtain training predicted semantics; calculating the cross-entropy loss function value of the true semantics and the training predicted semantics; updating the parameters of the semantic representation sub-model according to the calculation result, and / or updating the parameters of the classifier according to the calculation result.

[0015] In a second aspect, an embodiment of the present invention further provides an intent recognition device, which provides a graphical user interface, and the content displayed on the graphical user interface at least includes a virtual scene and a virtual character. The device includes: a scene image acquisition module, configured to acquire the scene image currently corresponding to the virtual scene in response to receiving the dialogue text input by the player corresponding to the virtual character; a knowledge sub-graph determination module, configured to obtain the knowledge sub-graphs corresponding to the dialogue text and the scene image respectively based on a preset knowledge graph; an intent prediction module, configured to obtain the predicted intent corresponding to the dialogue text based on the knowledge sub-graph and the intent recognition model.

[0016] In a third aspect, an embodiment of the present invention further provides an electronic device, including a processor and a memory, where the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the method for intent recognition in the first aspect above.

[0017] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the method for intent recognition in the first aspect above.

[0018] Compared with the prior art, the present invention has the following beneficial effects:

[0019] The method, apparatus, and electronic device for intent recognition provided by the embodiments of the present invention first respond to the received dialogue text input by the player corresponding to the virtual character, obtain the current scene image corresponding to the virtual scene, and obtain the knowledge subgraphs corresponding to the dialogue text and the scene image respectively based on a preset knowledge graph. Finally, based on the knowledge subgraph and the intent recognition model, the predicted intent corresponding to the dialogue text is obtained. When the technology of this application performs intent recognition on the dialogue text belonging to the player, it not only considers the information of the dialogue text itself, but also considers the virtual scene where the player issues the dialogue text. It expands the knowledge graph of the dialogue text and the scene image, and uses the expanded knowledge subgraphs together as the input of the intent recognition model. The predicted intent determined by the intent recognition model can more accurately reflect the true intent of the player for the virtual scene, effectively avoiding the misrecognition of the player's intent and improving the accuracy of intent recognition for the dialogue.

[0020] Other features and advantages of the present disclosure will be described in the following description, or some features and advantages can be inferred from the description or determined without doubt, or can be learned by implementing the above technologies of the present disclosure.

[0021] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0023] Figure 1 It is a schematic structural diagram of an electronic system provided by an embodiment of the present invention;

[0024] Figure 2 It is a flowchart of a method for intent recognition provided by an embodiment of the present invention;

[0025] Figure 3 It is a flowchart of another method for intent recognition provided by an embodiment of the present invention;

[0026] Figure 4 It is a schematic diagram of a virtual scene image and a segmented image obtained after segmentation provided by an embodiment of the present invention;

[0027] Figure 5 It is a flowchart of another method for intent recognition provided by an embodiment of the present invention;

[0028] Figure 6 Schematic diagram of a text information input and its corresponding knowledge subgraph provided by an embodiment of the present invention;

[0029] Figure 7 Flowchart of another method for intent recognition provided by an embodiment of the present invention;

[0030] Figure 8 Schematic diagram of the structure of the BERT model in the prior art;

[0031] Figure 9 Schematic diagram of the structure of an intent recognition model provided by an embodiment of the present invention;

[0032] Figure 10 Schematic diagram of the structure of an intent recognition device provided by an embodiment of the present invention;

[0033] Figure 11 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0035] In many RPG (Role-playing game) games, especially in cultivation-based RPGs, players communicate with NPCs (non-player characters) to express their intentions. However, the same dialogue may present different intentions in different scenarios. Identifying intentions only based on the dialogue without considering the scenario will increase the situation of incorrect recognition results. Based on this, the embodiments of the present invention provide a method, device, and electronic device for intent recognition to improve the accuracy of intent recognition results.

[0036] Refer to Figure 1 the schematic diagram of the structure of the electronic system 100 shown. This electronic system can be used to implement the method and device for intent recognition in the embodiments of the present invention.

[0037] As Figure 1Schematic structural diagram of an electronic system shown. The electronic system 100 includes one or more processing devices 102, one or more storage devices 104, an input device 106, an output device 108, and one or more data acquisition devices 110. These components are interconnected through a bus system 112 and / or other forms of connection mechanisms (not shown). It should be noted that Figure 1 The components and structure of the electronic system 100 shown are exemplary rather than restrictive. According to requirements, the electronic system may also have other components and structures.

[0038] The processing device 102 can be a server, a smart terminal, or a device including a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities. It can process the data of other components in the electronic system 100 and can also control other components in the electronic system 100 to perform an intention recognition function.

[0039] The storage device 104 can include one or more computer program products. The computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage media. The processing device 102 can run the program instructions to implement the client functions in the embodiments of the present invention below (implemented by the processing device) and / or other desired functions. Various application programs and various data can also be stored in the computer-readable storage media, such as various data used and / or generated by the application programs, etc.

[0040] The input device 106 can be a device used by a user to input instructions and can include one or more of a keyboard, a mouse, a microphone, and a touch screen, etc.

[0041] The output device 108 can output various information (such as images or sounds) to the outside (for example, to the user) and can include one or more of a display, a speaker, etc.

[0042] The data acquisition device 110 can acquire the voice signals input by the user and the images corresponding to the graphical user interface presented to the user, and store the voice signals and images in the storage device 104 for use by other components.

[0043] Exemplarily, the devices in the method, apparatus, and electronic device for implementing intent recognition according to the embodiments of the present invention may be integrally provided or dispersedly provided. For example, the processing device 102, the storage device 104, the input device 106, and the output device 108 may be integrally provided, while the data acquisition device 110 is provided at a specified position where voice signals and images can be acquired. When the devices in the above electronic system are integrally provided, the electronic system may be implemented as an intelligent terminal such as a camera, a smart phone, a tablet computer, a computer, a vehicle-mounted terminal, etc.

[0044] Figure 2 FIG. 4 is a flowchart of a method for intent recognition provided by an embodiment of the present invention. The method provides a graphical user interface through an electronic device, and the content displayed on the graphical user interface at least includes a virtual scene and a virtual character. Refer to Figure 2 , and the method includes the following steps:

[0045] S202: In response to receiving the dialogue text input by the player corresponding to the virtual character, obtain the scene image currently corresponding to the virtual scene;

[0046] Among them, the virtual scene is when the user faces the graphical user interface in the electronic device. For example, in the virtual scene of a game, when the player emits a voice signal during a certain game process and faces a certain game screen, then the game screen is the virtual scene, the image corresponding to the game screen is the scene image corresponding to the virtual scene, and the content of the voice signal is the dialogue text input by the player corresponding to the virtual character.

[0047] The process of converting the voice signal into dialogue text may adopt the voice-to-text conversion technology in the prior art, and the present invention does not limit this.

[0048] The scene image is the text information obtained by identifying the image content of the scene image corresponding to the virtual scene and describing the image content with text. For example, for an image containing the sea, its corresponding scene image is a sea. The process of performing image intent recognition on the image will be elaborated in detail below and will not be repeated here.

[0049] S204: Obtain the knowledge subgraphs corresponding to the dialogue text and the scene image respectively based on the preset knowledge graph;

[0050] In order to make the predicted semantics conform to the human thinking mode, the embodiments of the present application use a knowledge graph to expand the semantics of dialogue texts and scene images. The knowledge graph can be the N-Pedia knowledge graph (a Chinese encyclopedia knowledge graph), where each node in the graph represents a knowledge point, and the edge connecting the nodes represents the relationship between the two connected nodes. For example, if the two nodes are "pavilion" and "building", and the classification of the pavilion is building, then in the knowledge graph, the pavilion and the building can be connected, and the connection relationship is classification.

[0051] In the embodiments of the present application, the knowledge graph can be pre-stored in the electronic device. Each word in the dialogue text can be found corresponding to a node in the knowledge graph, and further, the range represented by the node can be expanded, and the nodes related to the node corresponding to the word are also associated with the word to form a knowledge sub-graph of the dialogue text. Similarly, for the scene image, each word in the description text corresponding to the scene image is used to determine the node and the associated nodes in the knowledge graph, and a knowledge sub-graph corresponding to the scene image is formed.

[0052] S206: Obtain the predicted intention corresponding to the dialogue text based on the knowledge sub-graph and the intention recognition model;

[0053] The embodiments of the present application can pre-train the intention recognition model, and perform intention prediction through the intention recognition model and the knowledge sub-graph. Specifically, the intention recognition model can be any form of neural network model. Preferably, it can be a BERT model.

[0054] The method for intention recognition provided by the embodiments of the present invention first responds to the received dialogue text input by the player corresponding to the virtual character, obtains the scene image corresponding to the current virtual scene, and obtains the knowledge sub-graphs corresponding to the dialogue text and the scene image respectively based on the preset knowledge graph. Finally, based on the knowledge sub-graph and the intention recognition model, the predicted intention corresponding to the dialogue text is obtained. When the technology of the present application performs intention recognition on the dialogue text of the player, it not only considers the information of the dialogue text itself, but also considers the virtual scene where the player issues the dialogue text. The knowledge graph of the dialogue text and the scene image is expanded, and the expanded knowledge sub-graphs are jointly used as the input of the intention recognition model. The predicted intention determined by the intention recognition model can more accurately reflect the player's true intention for the virtual scene, effectively avoiding the misrecognition of the player's intention and improving the accuracy of intention recognition for the dialogue.

[0055] Usually, the amount of information contained in a scene image is uneven because the layout of the image is uneven. In order to better utilize the information in the virtual scene image, the embodiments of the present application also provide another method for intention recognition. This method focuses on describing the process of determining the corresponding knowledge sub-graph according to the scene image, such asFigure 3 As shown in the figure, the method may specifically include:

[0056] S302: In response to receiving the dialogue text input by the player corresponding to the virtual character, obtain the scene image currently corresponding to the virtual scene;

[0057] S304: Obtain the knowledge sub-graph corresponding to the dialogue text based on the preset knowledge graph;

[0058] S306: Extract multiple sub-images containing the target object from the scene image;

[0059] Specifically, target detection can be performed on the scene image corresponding to the virtual scene, and the regions containing the target are extracted to obtain multiple sub-images containing the target object. Therefore, in some possible implementation manners, the above embodiments of the present application may further include: extracting multiple sub-images containing the target object from the scene image. Among them, the target object can be objects of multiple categories. For example, sub-images containing various buildings can be extracted with the building as the target object, or sub-images containing different character images can be extracted with the task as the target object.

[0060] For example, the recognized target object can be a person, a car, a building, a tree, etc. The extraction of sub-images can use the Faster-RCNN (Faster Regions with CNN features) algorithm. First, target detection is performed on the scene image, and the regions containing the target are extracted to obtain multiple sub-images: where Faster-RCNN uses a convolutional neural network to extract image features, and uses a Region Proposal Network (RPN) to extract the regions containing the target, and finally uses a pooling layer to obtain the specific classification of the target.

[0061] S308: Obtain the knowledge sub-graphs respectively corresponding to the multiple sub-images based on the preset knowledge graph;

[0062] Specifically, each sub-image obtained after extraction can be described by a piece of text. Therefore, in some examples, the description text corresponding to each sub-image in the multiple sub-images can be obtained through an image description algorithm, and based on the preset knowledge graph and the description text corresponding to each sub-image, the knowledge sub-graph corresponding to each sub-image can be obtained respectively. That is, the sub-image is converted into text information, and the knowledge sub-graph corresponding to the text information is determined in the knowledge graph.

[0063] In some possible implementation manners, the knowledge sub-graphs corresponding to the multiple sub-images can be determined specifically by the following method:

[0064] (1) Obtain the description text corresponding to each sub-image among multiple sub-images through an image description algorithm;

[0065] The process of predicting the text description of an image can be implemented using the DLCT (Dual-Level Collaborative Transformer) image description algorithm.

[0066] (2) Based on a preset knowledge graph and the description text corresponding to each sub-image, obtain the knowledge sub-graph corresponding to each sub-image respectively.

[0067] S310: Based on the knowledge sub-graph and the intent recognition model, obtain the predicted intent corresponding to the dialogue text.

[0068] By segmenting the scene image, it is possible to more precisely capture the effective information contained in the virtual scene image, filter out image noise and background, etc., thereby improving the accuracy of intent recognition. At the same time, it reduces the analysis time for background and noise, etc., and improves the efficiency of intent recognition.

[0069] Figure 4 For the segmented image corresponding to the virtual scene image obtained by the method provided in the above embodiments of the present application (as shown in the topmost image in Figure 4 ), and the image semantic text recognition result corresponding to one of the segmented images is: A knight wearing a cloak is riding a tall and handsome white horse (as shown in the bottom image in Figure 4 ). Figure 4 ).

[0070] Based on the above method, the embodiments of the present application also provide another method for intent recognition. This method focuses on describing the process of determining the knowledge sub-graph from the dialogue text and the description text, as shown in Figure 5 . This method includes:

[0071] S502: In response to receiving the dialogue text input by the player corresponding to the virtual character, obtain the scene image currently corresponding to the virtual scene;

[0072] S504: Based on the preset knowledge graph, obtain the initial knowledge sub-graphs corresponding to the dialogue text and the scene image respectively;

[0073] The knowledge graph contains multiple semantically related words and the connection relationships between the words. Based on this, in some possible implementation manners, the initial knowledge sub-graph can be determined through the following method (L11 - L14):

[0074] L11: Segment the description text corresponding to the dialogue text into multiple first text units;

[0075] L12: Segment the description text corresponding to the scene image into a plurality of second text units;

[0076] L13: Determine a first knowledge node corresponding to each first text unit from a preset knowledge graph;

[0077] L14: Determine a second knowledge node corresponding to each second text unit from a preset knowledge graph;

[0078] L15: Determine, from the preset knowledge graph, a first sub-knowledge node whose distance to the first knowledge node is less than a preset distance threshold;

[0079] L16: Determine, from the preset knowledge graph, a second sub-knowledge node whose distance to the second knowledge node is less than a preset distance threshold;

[0080] L17: Determine the subgraph consisting of the first knowledge node and the first sub-knowledge node as the initial knowledge subgraph corresponding to the dialogue text;

[0081] L18: Determine the subgraph consisting of the second knowledge node and the second sub-knowledge node as the initial knowledge subgraph corresponding to the scene image.

[0082] For example, the dialogue text is "I want to go to the pavilion", and the description text corresponding to a sub-image in the scene image is "fire". Then, first divide the dialogue text into three first text units "I", "I want to go", and "pavilion". Each text unit can determine a corresponding knowledge node in the knowledge graph. Here, it can be determined that the first knowledge node includes: "I, go, pavilion". Set a preset distance threshold, for example, the distance is 1. Then, for the knowledge node "pavilion", the knowledge graph can have the following relationship: (pavilion-isA->building, pavilion-isFor->shade), then the first sub-knowledge node includes: "building, shade". If the distance is set to 2, then, by expanding the relationship range in the knowledge graph, the following relationship can be found: (pavilion-isA->building, pavilion-isFor->shade, building-hasType->residence). Then, the first sub-knowledge node is "building, shade, residence". The determination method of the second knowledge node and the second sub-knowledge node is the same as the determination method of the first knowledge node and the first sub-knowledge node, which will not be repeated here.

[0083] Then, all the information included in the first knowledge node and the first sub-knowledge node is determined as the initial knowledge sub-graph corresponding to the dialogue text. All the information included in the second knowledge node and the second sub-knowledge node is determined as the initial knowledge sub-graph corresponding to the scene image.

[0084] Figure 6It is a schematic diagram of a knowledge subgraph corresponding to the input text information extracted from a knowledge graph. As Figure 6 shown, the text information is a knight in a cloak riding a tall and handsome white horse. The knowledge subgraph extracted from the knowledge graph is as Figure 6 shown on the right, where the content in the parentheses on each connection represents the relationship between two nodes in the knowledge subgraph.

[0085] S506: Based on a preset graph convolutional neural network, update the nodes in the initial knowledge subgraph respectively to obtain the knowledge subgraphs corresponding to the dialogue text and the scene image;

[0086] Specifically, the TransE model that has been trained on CN-Pedia can be used to map each node into a vector representation, so as to determine a knowledge subgraph vector for each node in the knowledge subgraph. After obtaining the knowledge subgraph vectors, the graph convolutional neural network GCN can be further used to update the representation of the nodes. During the update process, the nodes in the knowledge subgraph will continuously aggregate the representations of neighbor nodes to transmit information.

[0087] S508: Based on the knowledge subgraph and the intent recognition model, obtain the predicted intent corresponding to the dialogue text.

[0088] In the above embodiments of the present invention, a knowledge subgraph is constructed for each text unit in the dialogue text and the description text, the "knowledge" association between the text units is obtained, and then the association between different texts is captured, so as to more comprehensively understand the true semantics of the dialogue text input by the current player.

[0089] Through the method provided by the above embodiments of the present application, the knowledge subgraphs corresponding to the dialogue text and the scene image can be obtained. On this basis, the embodiments of the present application also provide another method for intent recognition. This method focuses on describing the specific process of determining the predicted intent according to the knowledge subgraph and the intent recognition model. As Figure 7 shown, this method includes the following steps:

[0090] S702: In response to receiving the dialogue text input by the player corresponding to the virtual character, obtain the scene image corresponding to the current virtual scene;

[0091] S704: Based on a preset knowledge graph, obtain the knowledge subgraphs corresponding to the dialogue text and the scene image respectively;

[0092] S706: Generate a first sequence according to the dialogue text and multiple sub-images;

[0093] S708: Generate a second sequence according to the first type of information of the dialogue text and the second type of multiple sub-images;

[0094] S710: Encode the position of each object in the first sequence to obtain the encoded representation of each object in the first sequence, and obtain the third sequence according to the encoded representation;

[0095] S712: Determine the fourth sequence according to the knowledge subgraph;

[0096] S714: Input the first sequence, the second sequence, the third sequence, and the fourth sequence into the intent recognition model, and process them through the intent recognition model to obtain the predicted intent corresponding to the dialogue text.

[0097] In the above embodiments, there are four input sequences for the intent recognition model, namely the first sequence, the second sequence, the third sequence, and the fourth sequence. The intent recognition model in the embodiments of the present application is a model obtained by improving the traditional BERT model. The structure of the traditional BERT model is as Figure 8 shown, and the structure of the intent recognition model provided in the embodiments of the present application is as Figure 9 shown. Below, in combination with Figure 9 detail the content and determination methods of the first sequence, the second sequence, the third sequence, and the fourth sequence.

[0098] (1) The first sequence

[0099] The first sequence is the text content, including multiple feature vectors, which consists of two parts: (1) the feature vectors corresponding to the dialogue text; (2) the feature vectors of the sub-images corresponding to the scene images.

[0100] For example, if the dialogue text output by the player is "I want to go to the pavilion in front", then each character will correspond to a feature vector. The sub-images corresponding to the current scene image include two, Image 1: a campfire, Image 2: a lake. Then, the first sequence includes: (I, want, to, go, pavilion, campfire, lake water).

[0101] (2) The second sequence

[0102] For each feature vector in the first sequence, its corresponding type constitutes the second sequence, that is, the second sequence includes the dialogue text and the scene image. Continuing with the above example, the second sequence corresponding to (I, want, to, go, pavilion, campfire, lake water) is (dialogue, dialogue, dialogue, dialogue, scene image, scene image).

[0103] (3) The third sequence

[0104] The third sequence is the position information of each feature vector in the first sequence, represented by absolute position encoding. For example, for a conversation text "I want to eliminate him", the conversation text is segmented into four text units: "I", "want to", "eliminate", and "him". Then the position of the text unit "I" in the whole sentence is [0], the position of the text unit "want to" in the whole sentence is [1], the position of the text unit "eliminate" in the whole sentence is [2], and the position of the text unit "him" in the whole sentence is [3]. That is, the position information of each feature vector in the conversation text is formed.

[0105] (4) The fourth sequence

[0106] In some possible implementation manners, the fourth sequence can be determined by the following method (L21 - L24):

[0107] L21: Determine the target objects that appear in both the knowledge sub - graph corresponding to the conversation text and the knowledge sub - graph corresponding to any sub - image in the first sequence;

[0108] For the feature vectors whose type belongs to the conversation text, first determine whether it is a target object, that is, determine that if it exists in

the nodes in the two - hop sub - graph constructed by the conversation text

[0109] For the feature vectors whose type belongs to the scene image, the representation of the whole sub - image is obtained by performing a splicing + pooling operation on the final complete node representations of the two - hop sub - graph constructed by the sub - image as the value of the fourth sequence corresponding to this feature vector.

[0110] L22: Use the encoded representation corresponding to the target object as the knowledge representation of the target object;

[0111] L23: Use the knowledge sub - graphs corresponding to multiple sub - images as the knowledge representations of the corresponding sub - images;

[0112] L24: Obtain the fourth sequence according to the knowledge representation.

[0113] In the above embodiments of the present application, by using the knowledge sub - graph as the input of the intention recognition model, since the knowledge sub - graph not only includes the content of the conversation text and the scene image itself, but also further expands it, it avoids inaccurate recognition due to the inability to find a matching semantics during the intention recognition process, and further improves the accuracy of intention recognition.

[0114] In some possible implementation manners, the intention recognition model in the embodiments of the present application is a pre - trained neural network model, which may include a semantic representation sub - model and a classifier. The training method of the model can be specifically:

[0115] (1) Obtain sample data; wherein, the sample data includes training texts and the true semantics corresponding to the training texts;

[0116] (2) Perform semantic prediction on the training texts through the initial intent recognition model to obtain training predicted semantics;

[0117] (3) Calculate the cross-entropy loss function value of the true semantics and the training predicted semantics;

[0118] (4) Update the parameters of the semantic representation sub-model according to the calculation result, and / or update the parameters of the classifier according to the calculation result.

[0119] It should be noted that in each iteration process, when updating the model parameters, it is possible to only update the parameters of the semantic representation sub-model, or only update the parameters of the classifier, or update the parameters of the semantic representation sub-model and the classifier at the same time. Of course, other update strategies can also be set. The embodiments of the present invention do not limit the update method of the model parameters. The embodiments of the present application adopt the method of separately updating the parameters of the semantic representation sub-model and the classifier, which can make the model better adapt to different virtual scenarios and improve the generalization of the model.

[0120] Based on the intent recognition model obtained through the above training, the confidence levels of each preset semantic information can be obtained, that is, the probability values indicating the true intent that each preset semantic information can express the user input voice information. It can be understood that the richer the preset semantic information, the easier it is to find semantic information with a higher matching degree with the true intent. Therefore, a semantic library corresponding to the virtual scenario can also be pre-stored in the electronic device; multiple semantics for different virtual scenarios are pre-stored in the semantic library. For example, for each virtual scenario, 50 semantics can be preset, or further, more semantics can be preset for each virtual scenario.

[0121] Based on this, the steps of performing intent recognition on the dialogue text in the above embodiments of the present application can specifically be:

[0122] (1) Determine the confidence level of each semantic in the semantic library according to the knowledge sub-graph and the intent recognition model; wherein, the confidence level is used to represent the probability that the semantic can reflect the true intent of the dialogue text;

[0123] (2) Determine the predicted intent that matches the dialogue text according to the confidence level of the semantic.

[0124] For the selection of the predicted intent, the semantic with the highest confidence can be selected as the predicted intent, or a semantic threshold can be set, and the semantic with the highest confidence among those higher than the threshold is used as the predicted intent. For example, if the semantic threshold is set to 0.5, then the semantic with the highest confidence among those with a confidence higher than 0.5 is considered the predicted intent, and the intent recognition is successful this time; otherwise, there is no predicted intent and the intent recognition fails this time.

[0125] The following combines Figure 9 , and introduces a method for intent recognition in a virtual scenario provided by an embodiment of the present invention. This method is illustrated by taking the recognition of the voice signal input by the user during the game process as an example. The method specifically includes:

[0126] Step 1: During the game process, the user inputs a voice signal in the game scenario;

[0127] For example, during the game process in the user's game screen, the user inputs a dialogue through the microphone: I'm so hot. The targets included in the game screen are three: a lake, a bonfire, and a pavilion.

[0128] Step 2: Recognize the text information of the voice signal and the image semantic text corresponding to the current game screen;

[0129] For example, segment the image corresponding to the current game screen to obtain multiple segmented images. The semantic corresponding to one of the segmented images is a clear lake, the semantic corresponding to another segmented image is a blazing bonfire, and the semantic corresponding to another segmented image is a pavilion.

[0130] Step 3: Determine the text information and the image semantic text as the text input;

[0131] Step 4: Determine the category corresponding to each word in the text information and the category corresponding to each word in the image semantic text as the category input;

[0132] Step 5: Determine the position of each word in the text information and the position corresponding to each word in the image semantic text as the position input;

[0133] Step 6: Determine the knowledge information corresponding to the text information and the image semantic text from the knowledge graph;

[0134] Specifically, for the determination of the knowledge information, the method for determining the knowledge information in the above-mentioned embodiment can be referred to. Details are not described herein again.

[0135] Step 7: Input all the above text input, category input, position input, and knowledge information into the K-VisualBERT model to obtain the confidence corresponding to each preset semantic;

[0136] For example, for the game scenario of this embodiment, there are three preset semantics: go to the gazebo to enjoy the cool, go to the lake to drink water, and warm up by the fire. After being predicted by the above K-VisualBERT model, the confidence level of "go to the gazebo to enjoy the cool" is 0.9, the confidence level of "go to the lake to drink water" is 0.85, and the confidence level of "warm up by the fire" is 0.2.

[0137] Step 8: Determine the semantics with the highest confidence level as the result of intent recognition.

[0138] Continuing with the above example, finally, it is determined that the user's intent is "go to the gazebo to enjoy the cool".

[0139] Based on the above method embodiments, an embodiment of the present invention further provides an intent recognition device, which provides a graphical user interface. The content displayed on the graphical user interface includes at least a virtual scene and a virtual character. Refer to Figure 10 As shown, the device includes:

[0140] A scene image acquisition module 1002, configured to respond to the received dialogue text input by the player corresponding to the virtual character, and acquire the scene image currently corresponding to the virtual scene;

[0141] A knowledge sub-graph determination module 1004, configured to acquire the knowledge sub-graphs corresponding to the dialogue text and the scene image respectively based on a preset knowledge graph;

[0142] An intent prediction module 1006, configured to obtain the predicted intent corresponding to the dialogue text based on the knowledge sub-graph and the intent recognition model.

[0143] The above intent recognition device provided by the embodiment of the present invention first responds to the received dialogue text input by the player corresponding to the virtual character, acquires the scene image currently corresponding to the virtual scene, and acquires the knowledge sub-graphs corresponding to the dialogue text and the scene image respectively based on a preset knowledge graph. Finally, based on the knowledge sub-graph and the intent recognition model, the predicted intent corresponding to the dialogue text is obtained. When the technology of this application performs intent recognition on the dialogue text of the player, it not only considers the information of the dialogue text itself, but also considers the virtual scene where the player issues the dialogue text, expands the knowledge graph of the dialogue text and the scene image, and uses the expanded knowledge sub-graphs together as the input of the intent recognition model. The predicted intent determined by the intent recognition model can more accurately reflect the true intent of the player for the virtual scene, effectively avoiding the misrecognition of the player's intent and improving the accuracy of intent recognition for the dialogue.

[0144] The above device further includes: a sub-image extraction module, configured to extract multiple sub-images including the target object from the scene image; the above knowledge sub-graph determination module 1004 is further configured to: acquire the knowledge sub-graphs corresponding to the multiple sub-images respectively based on a preset knowledge graph.

[0145] The above-mentioned knowledge sub-graph determination module 1004 is further configured to: obtain the description text corresponding to each of the multiple sub-images through an image description algorithm; and respectively obtain the knowledge sub-graph corresponding to each sub-image based on a preset knowledge graph and the description text corresponding to each sub-image.

[0146] The above-mentioned knowledge sub-graph determination module 1004 is further configured to: obtain the initial knowledge sub-graphs corresponding to the dialogue text and the scene image respectively based on a preset knowledge graph; and respectively update the nodes in the initial knowledge sub-graphs based on a preset graph convolutional neural network to obtain the knowledge sub-graphs corresponding to the dialogue text and the scene image respectively.

[0147] The process of obtaining the initial knowledge sub-graphs corresponding to the dialogue text and the scene image respectively based on a preset knowledge graph includes: splitting the description text corresponding to the dialogue text into multiple first text units, and splitting the description text corresponding to the scene image into multiple second text units; determining the first knowledge nodes corresponding to each first text unit and the second knowledge nodes corresponding to each second text unit from the preset knowledge graph; determining the first sub-knowledge nodes whose distances from the first knowledge nodes are less than a preset distance threshold, and the second sub-knowledge nodes whose distances from the second knowledge nodes are less than the preset distance threshold from the preset knowledge graph; and determining the sub-graph composed of the first knowledge nodes and the first sub-knowledge nodes as the initial knowledge sub-graph corresponding to the dialogue text, and determining the sub-graph composed of the second knowledge nodes and the second sub-knowledge nodes as the initial knowledge sub-graph corresponding to the scene image.

[0148] The process of obtaining the predicted intention corresponding to the dialogue text based on the knowledge sub-graph and the intention recognition model includes: generating a first sequence according to the dialogue text and the multiple sub-images; generating a second sequence according to the first type information of the dialogue text and the second type of the multiple sub-images; encoding the positions of each object in the first sequence to obtain the encoded representation of each object in the first sequence, and obtaining a third sequence according to the encoded representation; determining a fourth sequence according to the knowledge sub-graph; and inputting the first sequence, the second sequence, the third sequence, and the fourth sequence into the intention recognition model, and processing through the intention recognition model to obtain the predicted intention corresponding to the dialogue text.

[0149] The process of determining the fourth sequence according to the knowledge sub-graph includes: determining the target objects that appear in both the knowledge sub-graph corresponding to the dialogue text and the knowledge sub-graph corresponding to any one of the sub-images in the first sequence; using the encoded identifier corresponding to the target object as the knowledge representation of the target object; using the knowledge sub-graphs corresponding to the multiple sub-images respectively as the knowledge representations of the corresponding sub-images; and obtaining the fourth sequence according to the knowledge representation.

[0150] The above-mentioned electronic device also pre-stores a semantic library; the above-mentioned intention prediction module 1006 is further configured to: determine the confidence of each semantic in the semantic library according to the knowledge sub-graph and the intention recognition model; wherein, the confidence is used to represent the probability that the semantic can reflect the true intention of the dialogue text; determine the predicted intention matching the dialogue text according to the confidence of the semantic.

[0151] The above-mentioned intention recognition model includes a semantic representation sub-model and a classifier, and the intention recognition model is trained by the following method: obtaining sample data; wherein, the sample data includes training texts and the true semantics corresponding to the training texts; performing semantic prediction on the training texts through an initial intention recognition model to obtain training predicted semantics; calculating the cross-entropy loss function value of the true semantics and the training predicted semantics; updating the parameters of the semantic representation sub-model according to the calculation result, and / or updating the parameters of the classifier according to the calculation result.

[0152] The intention recognition device provided by the embodiment of the present invention has the same implementation principle and the same technical effects as those of the foregoing method embodiment. For the sake of brief description, for the parts not mentioned in the embodiment of the above-mentioned device, reference may be made to the corresponding content in the foregoing method embodiment of intention recognition.

[0153] The embodiment of the present invention also provides an electronic device, as Figure 11 shown, which is a schematic structural diagram of the electronic device. Among them, the electronic device includes a processor 1101 and a memory 1102. The memory 1102 stores computer-executable instructions that can be executed by the processor 1101, and the processor 1101 executes the computer-executable instructions to implement the above-mentioned method of intention recognition.

[0154] In Figure 11 the shown embodiment, the electronic device further includes a bus 1103 and a communication interface 1104. Among them, the processor 1101, the communication interface 1104, and the memory 1102 are connected through the bus 1103.

[0155] Among them, the memory 1102 may include high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 1104 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 1103 can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The bus 1103 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 11 only a bidirectional arrow is used in Figure 11 , but it does not mean that there is only one bus or one type of bus.

[0156] The processor 1101 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 1101 or by instructions in the form of software. The above-mentioned processor 1101 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor 1101 reads the information in the memory and combines its hardware to complete the steps of the method for intention recognition in the foregoing embodiments.

[0157] An embodiment of the present invention also provides a computer-readable storage medium storing computer-executable instructions, which, when called and executed by a processor, cause the processor to implement the above-described method for intent recognition. For specific implementation, reference may be made to the foregoing method embodiments and will not be elaborated herein.

[0158] A computer program product of the method, apparatus, and electronic device for intent recognition provided by an embodiment of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the foregoing method embodiments. For specific implementation, reference may be made to the method embodiments and will not be elaborated herein.

[0159] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.

[0160] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, and other various media that can store program code.

[0161] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0162] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for intention recognition, characterized in that, Provide a graphical user interface through an electronic device, and the content displayed on the graphical user interface includes at least a virtual scene and a virtual character; the method includes: In response to receiving the dialogue text input by the player corresponding to the virtual character, obtain the scene image currently corresponding to the virtual scene; Based on a preset knowledge graph, obtain the knowledge sub-graphs corresponding to the dialogue text and the scene image respectively; Based on the knowledge sub-graph and the intent recognition model, obtain the predicted intent corresponding to the dialogue text; The step of obtaining the knowledge sub-graphs corresponding to the dialogue text and the scene image respectively based on a preset knowledge graph includes: Based on a preset knowledge graph, obtain the initial knowledge sub-graphs corresponding to the dialogue text and the scene image respectively; Based on a preset graph convolutional neural network, update the nodes in the initial knowledge sub-graphs respectively to obtain the knowledge sub-graphs corresponding to the dialogue text and the scene image respectively; The step of obtaining the initial knowledge sub-graphs corresponding to the dialogue text and the scene image respectively based on a preset knowledge graph includes: Segment the description text corresponding to the dialogue text into multiple first text units, and segment the description text corresponding to the scene image into multiple second text units; Determine the first knowledge nodes corresponding to each of the first text units and the second knowledge nodes corresponding to each of the second text units from a preset knowledge graph; Determine the first sub-knowledge nodes whose distance from the first knowledge nodes is less than a preset distance threshold, and the second sub-knowledge nodes whose distance from the second knowledge nodes is less than the preset distance threshold from the preset knowledge graph; Determine the sub-graph composed of the first knowledge nodes and the first sub-knowledge nodes as the initial knowledge sub-graph corresponding to the dialogue text, and determine the sub-graph composed of the second knowledge nodes and the second sub-knowledge nodes as the initial knowledge sub-graph corresponding to the scene image.

2. The method according to claim 1, characterized in that, The method further includes: Extract multiple sub-images containing the target object from the scene image; The step of obtaining the knowledge sub-graph corresponding to the scene image based on a preset knowledge graph includes: Based on a preset knowledge graph, obtain the knowledge sub-graphs corresponding to the multiple sub-images respectively.

3. The method according to claim 2, characterized in that, The step of obtaining the knowledge sub-graphs corresponding to the multiple sub-images respectively based on a preset knowledge graph includes: Obtain the description text corresponding to each of the multiple sub-images respectively through an image description algorithm; Based on a preset knowledge graph and the description text corresponding to each sub-image, obtain the knowledge sub-graphs corresponding to each sub-image respectively.

4. The method according to claim 2, characterized in that, Based on the knowledge sub-graph and the intent recognition model, obtaining the predicted intent corresponding to the dialogue text includes: Generate a first sequence according to the dialogue text and the multiple sub-images; Generate a second sequence according to the first type information of the dialogue text and the second type of the multiple sub-images; Encode the positions of each object in the first sequence to obtain the encoded representation of each object in the first sequence, and obtain a third sequence according to the encoded representation; Determine a fourth sequence according to the knowledge sub-graph; Input the first sequence, the second sequence, the third sequence, and the fourth sequence into the intent recognition model, and process them through the intent recognition model to obtain the predicted intent corresponding to the dialogue text.

5. The method according to claim 4, characterized in that, The step of determining the fourth sequence according to the knowledge subgraph includes: Determine the target objects that appear in both the knowledge subgraph corresponding to the dialogue text and the knowledge subgraph corresponding to any of the sub-images in the first sequence; Use the encoding identifier corresponding to the target object as the knowledge representation of the target object; Use the knowledge subgraphs corresponding to the multiple sub-images respectively as the knowledge representations of the corresponding sub-images; Obtain the fourth sequence according to the knowledge representation.

6. The method according to claim 1, characterized in that, A semantic library is also pre-stored in the electronic device; The step of obtaining the predicted intent corresponding to the dialogue text based on the knowledge subgraph and the intent recognition model includes: Determine the confidence of each semantic in the semantic library according to the knowledge subgraph and the intent recognition model; wherein, the confidence is used to represent the probability that the semantic can reflect the true intent of the dialogue text; Determine the predicted intent that matches the dialogue text according to the confidence of the semantic.

7. The method according to any one of claims 1-6, characterized in that The intent recognition model includes a semantic representation sub-model and a classifier, and the intent recognition model is trained by the following method: Obtain sample data; wherein, the sample data includes training texts and the true semantics corresponding to the training texts; Perform semantic prediction on the training text through the initial intent recognition model to obtain the training predicted semantics; Calculate the cross-entropy loss function value of the true semantics and the training predicted semantics; Update the parameters of the semantic representation sub-model according to the calculation result, and / or update the parameters of the classifier according to the calculation result.

8. An intention recognition device, characterized in that The device provides a graphical user interface, and the content displayed on the graphical user interface at least includes a virtual scene and a virtual character. The device includes: A scene image acquisition module, configured to respond to receiving the dialogue text input by the player corresponding to the virtual character, and acquire the scene image corresponding to the virtual scene currently; A knowledge subgraph determination module, configured to obtain the knowledge subgraphs corresponding to the dialogue text and the scene image respectively based on a preset knowledge graph; An intent prediction module, configured to obtain the predicted intent corresponding to the dialogue text based on the knowledge subgraph and the intent recognition model; The knowledge subgraph determination module is further configured to obtain the initial knowledge subgraphs corresponding to the dialogue text and the scene image respectively based on a preset knowledge graph; update the nodes in the initial knowledge subgraphs respectively based on a preset graph convolutional neural network to obtain the knowledge subgraphs corresponding to the dialogue text and the scene image respectively; The knowledge sub-graph determination module is further configured to split the description text corresponding to the conversation text into a plurality of first text units, and split the description text corresponding to the scene image into a plurality of second text units; determine, from a preset knowledge graph, a first knowledge node corresponding to each of the first text units and a second knowledge node corresponding to each of the second text units; determine, from the preset knowledge graph, a first sub-knowledge node whose distance from the first knowledge node is less than a preset distance threshold, and a second sub-knowledge node whose distance from the second knowledge node is less than the preset distance threshold; determine the sub-graph composed of the first knowledge node and the first sub-knowledge node as the initial knowledge sub-graph corresponding to the conversation text, and determine the sub-graph composed of the second knowledge node and the second sub-knowledge node as the initial knowledge sub-graph corresponding to the scene image.

9. An electronic device, characterized in that It includes a processor and a memory, the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Interaction method and device based on intelligent robot

    CN110580516A