An object intention prediction method, device, electronic device, and storage medium

CN116933149BActive Publication Date: 2026-09-11SHENZHEN TCL NEW-TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210365445.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-07
Publication Date
2026-09-11
Estimated Expiration
2042-04-07

AI Technical Summary

Technical Problem

[0003]但是,相关技术中的一般每次人工智能与用户进行交互时,都只依托于本次交互过程中用户提供的信息进行交互,但是用户在交互时,用户可能每次交互时都是出于不同的交互意图,也有可能是延续上一轮交互的交互意图,采取现有的这种交互方式,有些时候会导致人工智能对于用户的交互意图预测不准确,影响用户的交互体验

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116933149B_ABST
    Figure CN116933149B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose an object intention prediction method and device, electronic equipment and storage medium; embodiments of the present application can obtain the current interaction instruction of the target interaction object, map the first interaction instruction to the current interaction semantic vector according to the preset semantic mapping relationship, obtain the historical interaction vector of the target interaction object, the historical interaction vector is obtained by mapping the historical interaction instruction based on the semantic mapping relationship, and the interaction time of the historical interaction instruction is before the current interaction instruction; vector fusion is performed based on the current interaction semantic vector and the historical interaction vector to obtain an intention fusion vector, the intention fusion vector is classified, and the interaction intention of the target interaction object is determined; in the embodiments of the present application, the information provided by the target interaction object in each round of interaction can be combined in the interaction process to determine the actual interaction intention of the target interaction object in the current round of interaction, and the accuracy of the interaction intention prediction of the target interaction object is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and more specifically to an object intent prediction method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the rapid development of science and technology, various types of artificial intelligence are appearing more and more frequently in people's lives. For example, artificial intelligence can imitate humans and interact with humans through voice, behavior, and so on.

[0003] However, in most related technologies, each time artificial intelligence interacts with a user, it relies solely on the information provided by the user during that interaction. But when users interact, they may have different intentions each time, or they may be continuing the intentions of the previous interaction. Using the existing interaction method can sometimes lead to inaccurate prediction of the user's interaction intentions by artificial intelligence, thus affecting the user's interaction experience. Summary of the Invention

[0004] This invention provides a method, apparatus, electronic device, and storage medium for predicting user intent, which can combine information provided by the user during each round of interaction to improve the accuracy of predicting the user's interaction intent.

[0005] This invention provides a method for predicting object intent, comprising:

[0006] Get the current interaction command of the target interaction object;

[0007] According to the preset semantic mapping relationship, the first interaction instruction is mapped to the current interaction semantic vector;

[0008] Obtain the historical interaction vector of the target interactive object. The historical interaction vector is obtained by mapping the historical interaction instructions based on the semantic mapping relationship. The interaction time of the historical interaction instructions is before the current interaction instruction.

[0009] Based on the current interaction semantic vector and the historical interaction vector, vector fusion is performed to obtain the intent fusion vector;

[0010] The intent fusion vector is classified to determine the interaction intent of the target interaction object.

[0011] Accordingly, embodiments of the present invention also provide an object intent prediction device, comprising:

[0012] The instruction acquisition unit is used to acquire the current interaction instruction of the target interaction object;

[0013] The vector mapping unit is used to map the first interaction instruction into the current interaction semantic vector according to a preset semantic mapping relationship;

[0014] A history vector acquisition unit is used to acquire at least one history interaction vector of the target interaction object. The history interaction vector is obtained by mapping the history interaction instructions based on the semantic mapping relationship. The interaction order of the history interaction instructions is before the current interaction instruction.

[0015] A vector fusion unit is used to perform vector fusion on the current interaction semantic vector and the historical interaction vector to obtain an intent fusion vector;

[0016] The intent determination unit is used to classify intents based on the intent fusion vector and determine the interaction intent of the target interaction object.

[0017] Optionally, the object intent prediction device provided in this embodiment of the invention further includes a weight determination unit, used to determine the historical interaction time corresponding to each of the historical interaction instructions;

[0018] Calculate the time interval between each of the historical interaction times and the current interaction time of the current interaction command;

[0019] The computational weights of each of the historical interaction vectors are determined based on the time intervals described.

[0020] Accordingly, the vector fusion unit is used to perform weighted calculation based on the current interaction semantic vector and the historical interaction vector, according to the calculation weight of each of the historical interaction vectors, to obtain the intent fusion vector.

[0021] Optionally, the object intent prediction device provided in this embodiment of the invention further includes a vector generation unit for obtaining at least one historical interaction instruction of the target interactive object;

[0022] Based on the semantic mapping relationship, each of the historical interaction instructions is mapped to a historical interaction semantic vector;

[0023] The historical interaction semantic vectors are classified according to intent to determine the historical interaction intent of each historical interaction semantic vector and the intent matching degree between each historical interaction semantic vector and the corresponding historical interaction intent.

[0024] Generate an intent matching vector based on the intent matching degree described above;

[0025] The historical interaction semantic vectors and the intent matching vectors are fused together to obtain the historical interaction vectors corresponding to each historical interaction instruction.

[0026] Optionally, the vector fusion unit is used to calculate the correlation degree between the current interaction semantic vector and each of the historical interaction vectors to obtain the correlation weight vector corresponding to the current interaction semantic vector;

[0027] The current interaction semantic vector is weighted and calculated with the associated weight vector to obtain the current weighted semantic vector corresponding to the current interaction semantic vector.

[0028] The historical interaction vector is weighted and calculated with the associated weight vector to obtain the historical weighted semantic vector corresponding to the historical interaction vector.

[0029] Vector fusion is performed based on the current weighted semantic vector and the historical weighted semantic vector to obtain the intent fusion vector.

[0030] Optionally, the object intent prediction device provided in this embodiment of the invention further includes an interaction result determination unit, used to determine an interaction database corresponding to the interaction intent based on the interaction intent, wherein the interaction database is used to determine the corresponding interaction result based on the interaction intent;

[0031] The interaction intent is sent to the interaction database, which then determines the corresponding interaction result based on the interaction intent and returns the interaction result.

[0032] Receive the interaction results sent by the interactive database.

[0033] Optionally, the object intent prediction device provided in this embodiment of the invention further includes a loop interaction unit for receiving interaction continuity instructions;

[0034] Based on the interaction persistence instruction, the current interaction instruction is used as the historical interaction instruction;

[0035] Receive the interaction instruction from the target interaction object and use the interaction instruction as the new current interaction instruction;

[0036] The step of mapping the current interaction instruction to a current interaction semantic vector according to the preset semantic mapping relationship is performed until the interaction continuous instruction is no longer received.

[0037] Optionally, the instruction acquisition unit is used to acquire the current voice instruction of the target interactive object;

[0038] Perform speech recognition on the current voice command to obtain the command text corresponding to the current voice command;

[0039] The instruction text is used as the current interaction instruction for the target interactive object.

[0040] Accordingly, embodiments of the present invention also provide an electronic device, including a memory and a processor; the memory stores an application program, and the processor is used to run the application program in the memory to perform the steps in any of the object intent prediction methods provided in embodiments of the present invention.

[0041] Accordingly, embodiments of the present invention also provide a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute steps in any of the object intent prediction methods provided in embodiments of the present invention.

[0042] Furthermore, embodiments of the present invention also provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps in any of the object intent prediction methods provided in embodiments of the present invention.

[0043] The solution of this invention can obtain the current interaction command of the target interaction object, map the first interaction command to the current interaction semantic vector according to a preset semantic mapping relationship, obtain the historical interaction vector of the target interaction object, and obtain the historical interaction vector by mapping the historical interaction command based on the semantic mapping relationship. The interaction time of the historical interaction command is before the current interaction command. Vector fusion is performed based on the current interaction semantic vector and the historical interaction vector to obtain an intent fusion vector. Intent classification is performed on the intent fusion vector to determine the interaction intent of the target interaction object. Since the intent fusion vector is used to determine the interaction intent of the target interaction object in this embodiment of the invention, and the intent fusion vector is obtained based on the current interaction command and the historical interaction command of the target interaction object, the actual interaction intent of the target interaction object in the current round of interaction can be determined by combining the information provided by the target interaction object in each round of interaction, thereby improving the accuracy of predicting the interaction intent of the target interaction object. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 This is a schematic diagram of a scenario for the object intent prediction method provided in an embodiment of the present invention;

[0046] Figure 2 This is a flowchart of the object intent prediction method provided in an embodiment of the present invention;

[0047] Figure 3 This is a schematic diagram of the voice command processing procedure of the object intent prediction method provided in this embodiment of the invention;

[0048] Figure 4 This is a schematic diagram of the skill server application process of the object intent prediction method provided in this embodiment of the invention;

[0049] Figure 5 This is a schematic diagram of the object intent prediction device provided in an embodiment of the present invention;

[0050] Figure 6 This is another structural schematic diagram of the object intent prediction device provided in an embodiment of the present invention;

[0051] Figure 7 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0053] This invention provides an object intent prediction method, apparatus, electronic device, and computer-readable storage medium. Specifically, this invention provides an object intent prediction method suitable for an object intent prediction apparatus, which can be integrated into an electronic device.

[0054] The electronic device can be a terminal or other device, including but not limited to mobile terminals and fixed terminals. For example, mobile terminals include but are not limited to smartphones, smartwatches, tablets, laptops, smart vehicles, etc., while fixed terminals include but are not limited to desktop computers, smart TVs, etc.

[0055] The electronic device can also be a server or other similar device. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, but it is not limited to these.

[0056] The object intent prediction method of this invention can be implemented by a terminal or by both a terminal and a server.

[0057] The following example illustrates the method of predicting the intent of an object, implemented jointly by a terminal and a server.

[0058] like Figure 1 As shown, the object intent prediction system provided in this embodiment of the invention includes a terminal 10 and a server 20, etc.; the terminal 10 and the server 20 are connected through a network, such as through a wired or wireless network, etc., wherein the terminal 10 can exist as a terminal for interacting with the target interactive object.

[0059] Among them, terminal 10 can be a terminal that generates the current interaction instruction in response to the interaction operation of the target interaction object, and is used to send the current interaction instruction of the target interaction object to server 20.

[0060] Server 20 can be used to obtain the current interaction command of the target interaction object, map the current interaction command to a current interaction semantic vector according to a preset semantic mapping relationship, obtain the historical interaction vector of the target interaction object, the historical interaction vector is obtained by mapping the historical interaction command based on the semantic mapping relationship, the interaction time of the historical interaction command is before the current interaction command, perform vector fusion based on the current interaction semantic vector and the historical interaction vector to obtain an intent fusion vector, perform intent classification on the intent fusion vector, and determine the interaction intent of the target interaction object.

[0061] It is understandable that the steps performed by server 20 can also be performed by terminal 10.

[0062] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.

[0063] The embodiments of the present invention will be described from the perspective of an object intent prediction device, which can be integrated into a terminal and / or a server.

[0064] like Figure 2 As shown, the specific process of the object intent prediction method in this embodiment can be as follows:

[0065] 201. Obtain the current interaction command of the target interaction object.

[0066] The target interaction object is the object that interacts with the interactive device, such as the user of the interactive device.

[0067] It is understandable that when the target interactive object interacts with the interactive device, each round of interaction can be treated as an independent interaction process, or multiple rounds of interaction can be treated as a single independent interaction process.

[0068] The current interaction command is the interaction command of the target interaction object in this round of interaction. For example, if the target interaction object has already interacted three times in one interaction, the interaction command in the fourth interaction is the current interaction command.

[0069] Specifically, the current interaction command may include the target interaction object's account information, the target interaction object's interaction content, the current interaction time and / or the current interaction location, etc. For example, the target interaction object's interaction content may be voice, text, images, etc., input by the target interaction object.

[0070] In some embodiments, the content of the target interactive object's interaction can be directly used as the current interaction command, or the current interaction command can be obtained by processing the content of the target interactive object's interaction. Taking the content of the target interactive object's interaction as voice as an example, step 201 may specifically include:

[0071] Get the current voice command of the target interactive object;

[0072] Perform speech recognition on the current voice command to obtain the command text corresponding to the current voice command;

[0073] The instruction text is used as the current interaction instruction for the target interactive object.

[0074] By processing voice commands, audio can be converted into text, saving storage space and improving interaction efficiency.

[0075] For example, such as Figure 3 As shown, IoT (Internet of Things) devices can collect the user's current voice input through the device's corresponding voice acquisition module, and then convert the current voice command in voice form into text command text through the ASR (Automatic Speech Recognition) server.

[0076] 202. Based on the preset semantic mapping relationship, map the current interaction instruction to the current interaction semantic vector.

[0077] In this application embodiment, the method for mapping interactive instructions (including mapping current interactive instructions and mapping historical interactive instructions) involves technologies such as natural language processing and machine learning in the field of artificial intelligence.

[0078] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to have perception, reasoning, and decision-making capabilities. AI technology is a comprehensive discipline involving a wide range of fields, encompassing both hardware and software technologies. AI software technology mainly includes computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and intelligent transportation.

[0079] Natural Language Processing (NLP) is an important area within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close connection with linguistic research. NLP technologies typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.

[0080] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory, among others. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning.

[0081] Specifically, semantic mapping relationships are used to map interactive commands to a predefined vector space. For example, the current interactive command can be mapped to a vector in the vector space. For instance, the semantic mapping relationship can be a predefined functional relationship, or it can be a pre-trained semantic feature extraction model, and so on.

[0082] It is understandable that semantic mapping relationships can include various semantic feature extraction models, such as semantic feature extraction models for text, semantic feature extraction models for images, and so on. Specific semantic feature extraction models can be BERT (Bidirectional Encoder Representations from Transformers) models, neural network models composed of pyramid networks, and so on.

[0083] 203. Obtain the historical interaction vector of the target interactive object. The historical interaction vector is obtained by mapping the historical interaction instructions based on the semantic mapping relationship. The interaction time of the historical interaction instructions is before the current interaction instruction.

[0084] Among them, historical interaction commands are interaction commands generated before the current interaction during a certain interaction process. For example, if the target interaction object has already been interacted with three times during an interaction process, and the fourth interaction is currently underway, then the commands generated in the previous three interactions are the historical interaction commands.

[0085] In some optional embodiments, the historical interaction vector can be a semantic vector obtained by mapping historical interaction instructions through a semantic mapping relationship, or the historical interaction vector can be a vector obtained by further processing the semantic vector obtained by mapping historical interaction instructions through a semantic mapping relationship. For example, before the step "obtaining the historical interaction vector of the target interaction object", the object intent prediction method provided in this embodiment of the invention may further include:

[0086] Obtain at least one historical interaction command from the target interaction object;

[0087] Based on the semantic mapping relationship, each of the historical interaction instructions is mapped to a historical interaction semantic vector;

[0088] The historical interaction semantic vectors are classified according to intent to determine the historical interaction intent of each historical interaction semantic vector and the intent matching degree between each historical interaction semantic vector and the corresponding historical interaction intent.

[0089] Generate an intent matching vector based on the intent matching degree described above;

[0090] The historical interaction semantic vectors and the intent matching vectors are fused together to obtain the historical interaction vectors corresponding to each historical interaction instruction.

[0091] There can be multiple intent matching vectors, each corresponding to a historical interaction semantic vector; or, the intent matching vector can be a single vector used to indicate the intent matching degree corresponding to each historical interaction semantic vector.

[0092] The historical interaction vector can be n vectors obtained by fusing n historical interaction semantic vectors with one or n intent matching vectors, in which case each historical interaction vector corresponds to each historical interaction semantic vector; or, the historical interaction vector can be one vector obtained by fusing n historical interaction semantic vectors with one or n intent matching vectors.

[0093] By fusing historical interaction semantic vectors with intent matching vectors obtained based on intent matching degree, the generated historical interaction vectors can include relevant information about intent matching degree, which is intended to provide more information when predicting the intent of the current interaction command.

[0094] 204. Perform vector fusion based on the current interaction semantic vector and the historical interaction vector to obtain the intent fusion vector.

[0095] Specifically, vector fusion can be achieved through methods such as concatenation, weighted calculation, multiplication, etc. Technical personnel can determine the vector fusion method according to the actual situation.

[0096] In some optional embodiments, the weight of the historical interaction vector during fusion can be determined based on the interaction time corresponding to the historical interaction command. That is, before step 2204, the object intent prediction method provided in this embodiment of the invention may further include:

[0097] Determine the historical interaction time corresponding to each of the aforementioned historical interaction commands.

[0098] Calculate the time interval between each of the historical interaction times and the current interaction time of the current interaction command;

[0099] The computational weights of each of the historical interaction vectors are determined based on the time intervals described.

[0100] Accordingly, the step "performing vector fusion based on the current interaction semantic vector and the historical interaction vector to obtain an intent fusion vector" may specifically include:

[0101] Based on the current interaction semantic vector and the historical interaction vector, a weighted calculation is performed according to the calculation weight of each historical interaction vector to obtain the intent fusion vector.

[0102] For example, the larger the time interval between the historical interaction time and the current interaction time, the smaller the corresponding calculation weight can be, and so on.

[0103] By determining the weights during vector fusion based on time intervals, the interaction information of the target interactive object at different interaction times can have different weights in the final intent fusion vector, thereby improving the accuracy of intent classification.

[0104] In some optional embodiments, a multi-head attention mechanism can be used to process the current interaction semantic vector and the historical interaction vector. The step "performing vector fusion based on the current interaction semantic vector and the historical interaction vector to obtain an intent fusion vector" may specifically include:

[0105] Calculate the correlation degree between the current interaction semantic vector and each of the historical interaction vectors to obtain the correlation weight vector corresponding to the current interaction semantic vector;

[0106] The current interaction semantic vector is weighted and calculated with the associated weight vector to obtain the current weighted semantic vector corresponding to the current interaction semantic vector.

[0107] The historical interaction vector is weighted and calculated with the associated weight vector to obtain the historical weighted semantic vector corresponding to the historical interaction vector.

[0108] Vector fusion is performed based on the current weighted semantic vector and the historical weighted semantic vector to obtain the intent fusion vector.

[0109] A multi-head attention network is used to transform the current interaction semantic vector into a query vector (q), a key vector (k), and a value vector (v). For example, a self-attention network can be used to fuse the current interaction semantic vector with the transformation parameters of the three dimensions to obtain the query vector (q), the key vector (k), and the value vector (v). The query vector (q), the key vector (k), and the value vector (v) are used as the association weight vector of each historical interaction vector.

[0110] 205. Perform intent classification on the intent fusion vector to determine the interaction intent of the target interaction object.

[0111] like Figure 3 As shown, in this embodiment of the invention, intent can be parsed using an NLP server. This embodiment of the invention can provide an object intent prediction system that may include a voice acquisition module, a device control system, a voice generation module, a display module, an ASR (Automatic Speech Recognition) server, an NLP (Natural Language Processing) server, and a skill server.

[0112] The system includes a voice acquisition module for collecting user voice information, a device control system for controlling device information scheduling and device status, a sound generation module for emitting sound, and a display module for displaying processed data, which may include images, videos, and text.

[0113] An ASR server can be used to convert speech into text. An NLP server is a semantic understanding module that can be used to obtain the user's intent based on the speech and text of the target interaction object, and extract key information from the text, such as domains, intents, slots, etc.

[0114] A specific example: The target interactive object's speech text is "Famous attractions in Beijing." The intention of this sentence is to query information about famous attractions in Beijing. An NLP server can convert the target interactive object's query into a formula that a computer can process:

[0115] {"domain":"viewspot","intent":"Scenery","slots":[{"name":"Location","value":"Beijing"},{"name":"Sortby_Hot","value":"famous"}]}.

[0116] It is understandable that the ASR server, NLP server, and skill server can be located on different servers or on the same server.

[0117] In some optional embodiments, the object intent prediction method provided by the embodiments of the present invention may further include:

[0118] Based on the interaction intent, an interaction database corresponding to the interaction intent is determined, wherein the interaction database is used to determine the corresponding interaction result based on the interaction intent;

[0119] The interaction intent is sent to the interaction database, which then determines the corresponding interaction result based on the interaction intent and returns the interaction result.

[0120] Receive the interaction results sent by the interactive database.

[0121] For example, such as Figure 4 As shown, there can be N different interaction databases (skill servers) corresponding to different intents. After the NLP server parses the corresponding intent, the NLP server can send it to the corresponding skill server (interaction database) according to the intent. After receiving the interaction intent, the skill server processes the data and returns the text information and the matching degree of the skill together.

[0122] The target interactive object can provide the current interactive command through voice input. After receiving the voice data, the voice acquisition module sends it to the ASR server to be converted into text information. The NLP server parses the user's intent and then calls the corresponding skill server. After receiving the data, the skill server performs the corresponding business processing and returns the processing result to the IoT device.

[0123] In some optional examples, the object intent prediction method provided in this embodiment of the invention may further include:

[0124] Receive continuous interactive commands;

[0125] Based on the interaction persistence instruction, the current interaction instruction is used as the historical interaction instruction;

[0126] Receive the interaction instruction from the target interaction object and use the interaction instruction as the new current interaction instruction;

[0127] The step of mapping the current interaction instruction to a current interaction semantic vector according to the preset semantic mapping relationship is performed until the interaction continuous instruction is no longer received.

[0128] The interaction persistence command can be automatically generated. For example, if the terminal detects that the target interaction object interacts again within a certain period of time, it can automatically generate the interaction persistence command. Alternatively, the interaction persistence command can be generated based on the target interaction object's operation. For example, if the target interaction object can choose to continue interacting on the terminal, the terminal can generate the interaction persistence command, and so on.

[0129] As can be seen from the above, the embodiments of the present invention can obtain the current interaction command of the target interaction object, map the first interaction command to the current interaction semantic vector according to the preset semantic mapping relationship, obtain the historical interaction vector of the target interaction object, the historical interaction vector is obtained by mapping the historical interaction command based on the semantic mapping relationship, the interaction time of the historical interaction command is before the current interaction command, and perform vector fusion based on the current interaction semantic vector and the historical interaction vector to obtain the intent fusion vector, classify the intent fusion vector, and determine the interaction intent of the target interaction object. Since the intent fusion vector is used to determine the interaction intent of the target interaction object in the embodiments of the present invention, and the intent fusion vector is obtained based on the current interaction command and the historical interaction command of the target interaction object, the actual interaction intent of the target interaction object in the current round of interaction can be determined by combining the information provided by the target interaction object in each round of interaction, thereby improving the accuracy of predicting the interaction intent of the target interaction object.

[0130] To better implement the above methods, the present invention also provides an object intent prediction device.

[0131] refer to Figure 5 The device may include:

[0132] The instruction acquisition unit 501 is used to acquire the current interaction instruction of the target interaction object;

[0133] The vector mapping unit 502 is used to map the first interaction instruction into the current interaction semantic vector according to a preset semantic mapping relationship;

[0134] The history vector acquisition unit 503 is used to acquire at least one history interaction vector of the target interaction object. The history interaction vector is obtained by mapping the history interaction instructions based on the semantic mapping relationship. The interaction order of the history interaction instructions is before the current interaction instruction.

[0135] The vector fusion unit 504 is used to perform vector fusion on the current interaction semantic vector and the historical interaction vector to obtain an intent fusion vector.

[0136] The intent determination unit 505 is used to classify intents based on the intent fusion vector and determine the interaction intent of the target interaction object.

[0137] Optional, such as Figure 6 As shown, the object intent prediction device provided in this embodiment of the invention further includes a weight determination unit 506, used to determine the historical interaction time corresponding to each of the historical interaction instructions;

[0138] Calculate the time interval between each of the historical interaction times and the current interaction time of the current interaction command;

[0139] The computational weights of each of the historical interaction vectors are determined based on the time intervals described.

[0140] Accordingly, the vector fusion unit 504 is used to perform weighted calculation based on the current interaction semantic vector and the historical interaction vector, according to the calculation weight of each of the historical interaction vectors, to obtain the intent fusion vector.

[0141] Optionally, the object intent prediction device provided in this embodiment of the invention further includes a vector generation unit 507, used to obtain at least one historical interaction instruction of the target interactive object;

[0142] Based on the semantic mapping relationship, each of the historical interaction instructions is mapped to a historical interaction semantic vector;

[0143] The historical interaction semantic vectors are classified according to intent to determine the historical interaction intent of each historical interaction semantic vector and the intent matching degree between each historical interaction semantic vector and the corresponding historical interaction intent.

[0144] Generate an intent matching vector based on the intent matching degree described above;

[0145] The historical interaction semantic vectors and the intent matching vectors are fused together to obtain the historical interaction vectors corresponding to each historical interaction instruction.

[0146] Optionally, the vector fusion unit 504 is used to calculate the correlation degree between the current interaction semantic vector and each of the historical interaction vectors to obtain the correlation weight vector corresponding to the current interaction semantic vector.

[0147] The current interaction semantic vector is weighted and calculated with the associated weight vector to obtain the current weighted semantic vector corresponding to the current interaction semantic vector.

[0148] The historical interaction vector is weighted and calculated with the associated weight vector to obtain the historical weighted semantic vector corresponding to the historical interaction vector.

[0149] Vector fusion is performed based on the current weighted semantic vector and the historical weighted semantic vector to obtain the intent fusion vector.

[0150] Optionally, the object intent prediction device provided in this embodiment of the invention further includes an interaction result determination unit 508, which is used to determine the interaction database corresponding to the interaction intent based on the interaction intent, wherein the interaction database is used to determine the corresponding interaction result based on the interaction intent.

[0151] The interaction intent is sent to the interaction database, which then determines the corresponding interaction result based on the interaction intent and returns the interaction result.

[0152] Receive the interaction results sent by the interactive database.

[0153] Optionally, the object intent prediction device provided in this embodiment of the invention further includes a loop interaction unit 509 for receiving interaction continuity instructions;

[0154] Based on the interaction persistence instruction, the current interaction instruction is used as the historical interaction instruction;

[0155] Receive the interaction instruction from the target interaction object and use the interaction instruction as the new current interaction instruction;

[0156] The step of mapping the current interaction instruction to a current interaction semantic vector according to the preset semantic mapping relationship is performed until the interaction continuous instruction is no longer received.

[0157] Optionally, the instruction acquisition unit 501 is used to acquire the current voice instruction of the target interactive object;

[0158] Perform speech recognition on the current voice command to obtain the command text corresponding to the current voice command;

[0159] The instruction text is used as the current interaction instruction for the target interactive object.

[0160] As can be seen from the above, the object intent prediction device can obtain the current interaction command of the target interaction object, map the first interaction command to the current interaction semantic vector according to the preset semantic mapping relationship, obtain the historical interaction vector of the target interaction object, and obtain the historical interaction vector by mapping the historical interaction command based on the semantic mapping relationship. The interaction time of the historical interaction command is before the current interaction command. Vector fusion is performed based on the current interaction semantic vector and the historical interaction vector to obtain the intent fusion vector. The intent fusion vector is then classified to determine the interaction intent of the target interaction object. Since the intent fusion vector is used to determine the interaction intent of the target interaction object in this embodiment of the invention, and the intent fusion vector is obtained based on the current interaction command and the historical interaction command of the target interaction object, the actual interaction intent of the target interaction object in this round of interaction can be determined by combining the information provided by the target interaction object in each round of interaction, thereby improving the accuracy of predicting the interaction intent of the target interaction object.

[0161] Furthermore, embodiments of the present invention also provide an electronic device, which may be a terminal or a server, etc. Figure 7 As shown, it illustrates a structural schematic diagram of the electronic device involved in an embodiment of the present invention, specifically:

[0162] The electronic device may include a radio frequency (RF) circuit 701, a memory 702 including one or more computer-readable storage media, an input unit 703, a display unit 704, a sensor 705, an audio circuit 706, a wireless Fidelity (WiFi) module 707, a processor 708 including one or more processing cores, and a power supply 709, etc. Those skilled in the art will understand that... Figure 7 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0163] RF circuit 701 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and hands it over to one or more processors 708 for processing; additionally, it transmits uplink data to the base station. Typically, RF circuit 701 includes, but is not limited to, an antenna, at least one amplifier, a tuner, one or more oscillators, a Subscriber Identity Module (SIM) card, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc. Furthermore, RF circuit 701 can also communicate wirelessly with networks and other devices. Wireless communication can use any communication standard or protocol, including but not limited to GSM, GPRS, CDMA, WCDMA, LTE, email, and SMS.

[0164] The memory 702 can be used to store software programs and modules. The processor 708 executes various functional applications and data processing by running the software programs and modules stored in the memory 702. The memory 702 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device (such as audio data, telephone directory, etc.). In addition, the memory 702 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 702 may also include a memory controller to provide access to the memory 702 for the processor 708 and the input unit 703.

[0165] The input unit 703 can be used to receive input digital or character information, and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. Specifically, in one embodiment, the input unit 703 may include a touch-sensitive surface and other input devices. The touch-sensitive surface, also known as a touch display or touchpad, can collect user touch operations on or near it (e.g., user operations using fingers, styluses, or any suitable object or accessory on or near the touch-sensitive surface), and drive corresponding connection devices according to a pre-set program. Optionally, the touch-sensitive surface may include a touch detection device and a touch controller. The touch detection device detects the user's touch orientation and the signal generated by the touch operation, transmitting the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to the processor 708, and can receive and execute commands from the processor 708. Furthermore, various types of touch-sensitive surfaces, such as resistive, capacitive, infrared, and surface acoustic wave, can be used. In addition to the touch-sensitive surface, the input unit 703 may also include other input devices. Specifically, other input devices may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.

[0166] Display unit 704 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of electronic devices. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Display unit 704 may include a display panel, optionally configured as a liquid crystal display (LCD), organic light-emitting diode (OLED), or similar form. Furthermore, a touch-sensitive surface may cover the display panel. When the touch-sensitive surface detects a touch operation on or near it, it transmits the information to processor 708 to determine the type of touch event. Subsequently, processor 708 provides corresponding visual output on the display panel according to the type of touch event. Although in Figure 7 In this context, the touch-sensitive surface and the display panel are two separate components for implementing input and output functions. However, in some embodiments, the touch-sensitive surface and the display panel can be integrated to achieve both input and output functions.

[0167] The electronic device may also include at least one sensor 705, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor can adjust the brightness of the display panel according to the ambient light level, and the proximity sensor can turn off the display panel and / or backlight when the electronic device is moved to the ear. As a type of motion sensor, a gravity acceleration sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. Other sensors that may be configured in the electronic device, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.

[0168] Audio circuitry 706, a speaker, and a microphone provide an audio interface between the user and the electronic device. Audio circuitry 706 converts received audio data into electrical signals, transmits them to the speaker, and the speaker converts them into sound signals for output. Conversely, the microphone converts collected sound signals into electrical signals, which are then received by audio circuitry 706, converted back into audio data, and processed by processor 708. The processed data is then transmitted via RF circuitry 701 to, for example, another electronic device, or output to memory 702 for further processing. Audio circuitry 706 may also include an earphone jack to facilitate communication between external headphones and the electronic device.

[0169] WiFi is a short-range wireless transmission technology. Electronic devices using the WiFi module 707 can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 7 WiFi module 707 is shown, but it is understood that it is not a necessary component of the electronic device and can be omitted as needed without changing the nature of the invention.

[0170] The processor 708 is the control center of the electronic device. It connects various parts of the phone via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 702, and by calling data stored in the memory 702, thereby providing overall monitoring of the phone. Optionally, the processor 708 may include one or more processing cores; preferably, the processor 708 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 708.

[0171] The electronic device also includes a power supply 709 (such as a battery) that supplies power to various components. Preferably, the power supply can be logically connected to the processor 708 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 709 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0172] Although not shown, the electronic device may also include a camera, Bluetooth module, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 708 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 702 according to the following instructions, and the processor 708 runs the applications stored in the memory 702 to realize various functions, as follows:

[0173] Get the current interaction command of the target interaction object;

[0174] According to the preset semantic mapping relationship, the first interaction instruction is mapped to the current interaction semantic vector;

[0175] Obtain the historical interaction vector of the target interactive object. The historical interaction vector is obtained by mapping the historical interaction instructions based on the semantic mapping relationship. The interaction time of the historical interaction instructions is before the current interaction instruction.

[0176] Based on the current interaction semantic vector and the historical interaction vector, vector fusion is performed to obtain the intent fusion vector;

[0177] The intent fusion vector is classified to determine the interaction intent of the target interaction object.

[0178] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0179] To this end, embodiments of the present invention provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the object intent prediction methods provided in embodiments of the present invention. For example, the instructions can execute the following steps:

[0180] Get the current interaction command of the target interaction object;

[0181] According to the preset semantic mapping relationship, the first interaction instruction is mapped to the current interaction semantic vector;

[0182] Obtain the historical interaction vector of the target interactive object. The historical interaction vector is obtained by mapping the historical interaction instructions based on the semantic mapping relationship. The interaction time of the historical interaction instructions is before the current interaction instruction.

[0183] Based on the current interaction semantic vector and the historical interaction vector, vector fusion is performed to obtain the intent fusion vector;

[0184] The intent fusion vector is classified to determine the interaction intent of the target interaction object.

[0185] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0186] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0187] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the object intent prediction methods provided in the embodiments of the present invention, the beneficial effects that any of the object intent prediction methods provided in the embodiments of the present invention can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0188] According to one aspect of this application, a computer program product or computer program is also provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the methods provided in the various optional implementations of the above embodiments.

[0189] The foregoing has provided a detailed description of an object intent prediction method, apparatus, electronic device, and storage medium provided by embodiments of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. An object intention prediction method characterized by comprising: include: Get the current interaction command of the target interaction object; Based on a preset semantic mapping relationship, the current interaction instruction is mapped to a current interaction semantic vector; Obtain the historical interaction vector of the target interactive object. The historical interaction vector is obtained by mapping the historical interaction instructions based on the semantic mapping relationship. The interaction time of the historical interaction instructions is before the current interaction instruction. Based on the current interaction semantic vector and the historical interaction vector, vector fusion is performed to obtain the intent fusion vector; The intent fusion vector is classified to determine the interaction intent of the target interaction object; Before obtaining the historical interaction vector of the target interactive object, the method further includes: Obtain at least one historical interaction command from the target interaction object; Based on the semantic mapping relationship, each of the historical interaction instructions is mapped to a historical interaction semantic vector; The historical interaction semantic vectors are classified according to intent to determine the historical interaction intent of each historical interaction semantic vector and the intent matching degree between each historical interaction semantic vector and the corresponding historical interaction intent. Generate an intent matching vector based on the intent matching degree described above; The historical interaction semantic vectors and the intent matching vectors are fused together to obtain the historical interaction vectors corresponding to each historical interaction instruction.

2. The object intention prediction method of claim 1, wherein, Before performing vector fusion based on the current interaction semantic vector and the historical interaction vector to obtain the intent fusion vector, the method further includes: Determine the historical interaction time corresponding to each of the aforementioned historical interaction commands. Calculate the time interval between each of the historical interaction times and the current interaction time of the current interaction command; The computational weights of each of the historical interaction vectors are determined based on the time intervals described. The process of fusing vectors based on the current interaction semantic vector and the historical interaction vector to obtain an intent fusion vector includes: Based on the current interaction semantic vector and the historical interaction vector, a weighted calculation is performed according to the calculation weight of each historical interaction vector to obtain the intent fusion vector.

3. The object intention prediction method of claim 1, wherein, The process of fusing vectors based on the current interaction semantic vector and the historical interaction vector to obtain an intent fusion vector includes: Calculate the correlation degree between the current interaction semantic vector and each of the historical interaction vectors to obtain the correlation weight vector corresponding to the current interaction semantic vector; The current interaction semantic vector is weighted and calculated with the associated weight vector to obtain the current weighted semantic vector corresponding to the current interaction semantic vector. The historical interaction vector is weighted and calculated with the associated weight vector to obtain the historical weighted semantic vector corresponding to the historical interaction vector. Vector fusion is performed based on the current weighted semantic vector and the historical weighted semantic vector to obtain the intent fusion vector.

4. The object intention prediction method of claim 1, wherein, The method further includes: Based on the interaction intent, an interaction database corresponding to the interaction intent is determined, wherein the interaction database is used to determine the corresponding interaction result based on the interaction intent; The interaction intent is sent to the interaction database, which then determines the corresponding interaction result based on the interaction intent and returns the interaction result. Receive the interaction results sent by the interactive database.

5. The object intention prediction method of claim 1, wherein, The method further includes: Receive continuous interactive commands; Based on the interaction persistence instruction, the current interaction instruction is used as the historical interaction instruction; Receive the interaction instruction from the target interaction object and use the interaction instruction as the new current interaction instruction; The step of mapping the current interaction instruction to a current interaction semantic vector according to the preset semantic mapping relationship is performed until the interaction continuous instruction is no longer received.

6. The object intention prediction method according to any one of claims 1-5, characterized by, The step of obtaining the current interaction command of the target interaction object includes: Get the current voice command of the target interactive object; Perform speech recognition on the current voice command to obtain the command text corresponding to the current voice command; The instruction text is used as the current interaction instruction for the target interactive object.

7. An object intention prediction device characterized by comprising: include: The instruction acquisition unit is used to acquire the current interaction instruction of the target interaction object; The vector mapping unit is used to map the current interaction instruction into a current interaction semantic vector according to a preset semantic mapping relationship; A history vector acquisition unit is used to acquire at least one history interaction vector of the target interaction object. The history interaction vector is obtained by mapping the history interaction instructions based on the semantic mapping relationship. The interaction order of the history interaction instructions is before the current interaction instruction. A vector fusion unit is used to perform vector fusion on the current interaction semantic vector and the historical interaction vector to obtain an intent fusion vector; An intent determination unit is used to classify intents based on the intent fusion vector and determine the interaction intent of the target interaction object. A vector generation unit is used to obtain at least one historical interaction instruction of the target interactive object; Based on the semantic mapping relationship, each of the historical interaction instructions is mapped to a historical interaction semantic vector; The historical interaction semantic vectors are classified according to intent to determine the historical interaction intent of each historical interaction semantic vector and the intent matching degree between each historical interaction semantic vector and the corresponding historical interaction intent. Generate an intent matching vector based on the intent matching degree described above; The historical interaction semantic vectors and the intent matching vectors are fused together to obtain the historical interaction vectors corresponding to each historical interaction instruction.

8. An electronic device, characterized in that, It includes a memory and a processor; the memory stores an application program, and the processor is used to run the application program within the memory to perform the steps in the object intent prediction method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the object intent prediction method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the object intent prediction method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • A statement user intention identification method and device

    CN109697282A