Intent information determination methods, apparatus, equipment, storage media and program products

By acquiring audio tags of audio pairs and filtering redundant information in a task-oriented dialogue system, the accuracy of intent information is improved by utilizing intent information to determine the model, thus solving the problem of redundant information interference in multi-turn dialogues.

CN114691920BActive Publication Date: 2025-10-28BEIJING SANKUAI ONLINE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210306852.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-25
Publication Date
2025-10-28
Estimated Expiration
2042-03-25

AI Technical Summary

Technical Problem

Existing task-oriented dialogue systems suffer from poor accuracy in generating intent information due to interference from redundant information in multi-turn dialogues.

Method used

By acquiring the audio tags of multiple audio pairs, filtering out audio pairs whose intent parameters do not meet the preset parameters, and using the intent information determination model to determine the target intent information.

Benefits of technology

It improves the accuracy of intent information, avoids interference from redundant information, and ensures the accurate generation of intent information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114691920B_ABST
    Figure CN114691920B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, storage medium, and program product for determining intent information, belonging to the field of neural network technology. The method acquires multiple first audio pairs, determines an audio tag for each first audio pair, and based on the audio tag of each first audio pair, filters out first audio pairs whose intent parameters do not meet preset parameters from the multiple first audio pairs, obtaining at least one second audio pair. Based on at least one second audio pair, the target intent parameters are determined. Since the audio tag can reflect the intent parameters of the audio pair, audio pairs whose intent parameters do not meet the conditions, i.e., redundant audio pairs, can be filtered out from multiple audio pairs based on the audio tag of each audio pair, obtaining audio pairs that meet the conditions. The target intent information is determined based on the audio pairs that meet the conditions, thereby avoiding interference from redundant information, accurately generating the target intent information, and thus improving the accuracy of the determined target intent information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of neural network technology, and in particular to a method, apparatus, device, storage medium, and program product for determining intent information. Background Technology

[0002] Task-oriented dialogue systems are systems that help users complete specific tasks through dialogue, such as restaurant reservations, weather forecasts, or travel arrangements. These systems can determine the user's intent based on the dialogue and respond accordingly to meet the user's needs. Therefore, determining this intent is a crucial problem that needs to be solved.

[0003] In related technologies, when a system completes a task, the user may not state all their needs at once, and may change their needs during the conversation. Therefore, the system may need to engage in multiple rounds of dialogue with the user. After two rounds of dialogue and before starting a new round, the system needs to determine the user's intent information based on the audio of the previous multiple rounds of dialogue.

[0004] However, the audio of this multi-turn dialogue contains redundant information, which interferes with the generation of intent information, affects the generation effect of intent information, and leads to poor accuracy of the determined intent information. Summary of the Invention

[0005] This application provides a method, apparatus, device, storage medium, and program product for determining intent information, which can improve the accuracy of the determined intent information. The technical solution is as follows:

[0006] Firstly, a method for determining intent information is provided, the method comprising:

[0007] Acquire multiple first audio pairs, where each first audio pair is the audio of a dialogue between a first dialogue object and a second dialogue object;

[0008] Determine an audio tag for each first audio pair, the audio tag being used to represent the intent parameter of the first audio pair;

[0009] Based on the audio tags of each first audio pair, filter out first audio pairs whose intent parameters do not meet the preset parameters from the plurality of first audio pairs to obtain at least one second audio pair;

[0010] Based on the at least one second audio pair, the target intent information of the first dialogue object is determined.

[0011] Secondly, a method for determining intent information is provided, the method comprising:

[0012] Acquire multiple first audio pairs, where each first audio pair is the audio of a dialogue between a first dialogue object and a second dialogue object;

[0013] By using the multiple first audio pairs as input intent information in the model, the target intent information of the first dialogue object is obtained;

[0014] The intent information determination model is used to determine the audio tag of each first audio pair. The audio tag is used to represent the intent parameter of the first audio pair. Based on the audio tag of each first audio pair, first audio pairs whose intent parameters do not meet the preset parameters are filtered out from the plurality of first audio pairs to obtain at least one second audio pair. Based on the at least one second audio pair, the target intent information is determined.

[0015] Thirdly, an intent information determination apparatus is provided, the apparatus comprising:

[0016] The first acquisition module is used to acquire multiple first audio pairs, wherein the first audio pair is the audio of a dialogue between a first dialogue object and a second dialogue object;

[0017] The first determining module is used to determine the audio tag for each first audio pair, wherein the audio tag is used to represent the intent parameter of the first audio pair;

[0018] The filtering module is used to filter out first audio pairs whose intent parameters do not meet preset parameters from the plurality of first audio pairs based on the audio tags of each first audio pair, so as to obtain at least one second audio pair;

[0019] The second determining module is used to determine the target intent information of the first dialogue object based on the at least one second audio pair.

[0020] Fourthly, an intent information determination device is provided, the device comprising:

[0021] The second acquisition module is used to acquire multiple first audio pairs, wherein the first audio pair is the dialogue audio between the first dialogue object and the second dialogue object;

[0022] The input module is used to determine the target intent information of the first dialogue object by inputting the multiple first audio pairs into the intent information determination model;

[0023] The intent information determination model is used to determine the audio tag of each first audio pair. The audio tag is used to represent the intent parameter of the first audio pair. Based on the audio tag of each first audio pair, first audio pairs whose intent parameters do not meet the preset parameters are filtered out from the plurality of first audio pairs to obtain at least one second audio pair. Based on the at least one second audio pair, the target intent information is determined.

[0024] Fifthly, an electronic device is provided, the electronic device including a processor and a memory, the memory storing at least one piece of program code, the at least one piece of program code being loaded and executed by the processor to implement an intent information determination method as described in any of the possible implementations of the first or second aspect above.

[0025] In a sixth aspect, a computer-readable storage medium is provided, wherein at least one piece of program code is stored therein, the at least one piece of program code being loaded and executed by a processor to implement an intent information determination method as described in any of the possible implementations of the first or second aspect above.

[0026] In a seventh aspect, a computer program product is provided, the computer program product storing at least one piece of program code, the at least one piece of program code being loaded and executed by a processor to implement the intent information determination method as described in any of the possible implementations of the first or second aspect above.

[0027] The beneficial effects of the technical solutions provided in this application include at least the following:

[0028] This application provides a method for determining intent information. The method first determines the audio tag of each audio pair. Since the audio tag can reflect the intent parameters of the audio pair, audio pairs whose intent parameters do not meet the conditions, i.e., redundant audio pairs, can be filtered out from multiple audio pairs based on the audio tag of each audio pair, and audio pairs that meet the conditions are obtained. The target intent information is determined based on the audio pairs that meet the conditions, so as to avoid the interference of redundant information, accurately generate the target intent information, and thus improve the accuracy of the determined target intent information. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a schematic diagram of the implementation environment of an intent information determination method provided in an embodiment of this application;

[0031] Figure 2 This is a flowchart of an intent information determination method provided in an embodiment of this application;

[0032] Figure 3 This is a flowchart of an intent information determination method provided in an embodiment of this application;

[0033] Figure 4 This is a schematic diagram illustrating the determination of audio tags provided in an embodiment of this application;

[0034] Figure 5 This is a schematic diagram illustrating the determination of target intent information provided in an embodiment of this application;

[0035] Figure 6 This is a flowchart of an intent information determination method provided in an embodiment of this application;

[0036] Figure 7 This is a schematic diagram of a PR curve provided in an embodiment of this application;

[0037] Figure 8 This is a schematic diagram of the structure of an intent information determination device provided in an embodiment of this application;

[0038] Figure 9 This is a schematic diagram of the structure of an intent information determination device provided in an embodiment of this application;

[0039] Figure 10 This is a structural block diagram of a terminal provided in an embodiment of this application;

[0040] Figure 11 This is a structural block diagram of a server provided in an embodiment of this application. Detailed Implementation

[0041] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0042] The terms "first," "second," "third," and "fourth," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0043] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the audio of conversations involved in this application was obtained with full authorization.

[0044] Figure 1 This is a schematic diagram illustrating the implementation environment of an intent information determination method provided in this application embodiment. See also... Figure 1 The implementation environment includes an electronic device, which can be provided as a terminal 101, or as a combination of a terminal 101 and a server 102. In this embodiment, no specific limitation is made.

[0045] If the electronic device is provided as terminal 101, which is a device that carries the second dialogue object, the first dialogue object and the second dialogue object communicate through the terminal 101. The terminal 101 can record the audio of the dialogue between the first dialogue object and the second dialogue object. The terminal 101 determines the intent information of the first dialogue object based on the audio, determines the statement to reply to the first dialogue object based on the intent information, and sends the statement to the second dialogue object. The second dialogue object replies to the first dialogue object based on the statement, thereby satisfying the needs of the first dialogue object.

[0046] The first dialogue partner can be a user or a device. If the first dialogue partner is a user, the user can directly communicate with the second dialogue partner. If the first dialogue partner is a device, the user communicates with the second dialogue partner through the device. The second dialogue partner is a module with voice function embedded in the terminal, such as a voice assistant. There are no specific limitations on the first and second dialogue partners here.

[0047] If the electronic device provides a terminal 101 and a server 102, then the terminal 101 records the audio of the conversation between the first and second dialogue partners and sends the audio to the server 102. The server 102 determines the intent information of the first dialogue partner based on the audio and then returns the intent information to the terminal 101. The terminal 101 determines the reply statement to the first dialogue partner based on the intent information and sends the statement to the second dialogue partner. Alternatively, the server determines the reply statement to the first dialogue partner based on the intent information, sends the statement to the terminal, the terminal forwards the statement to the second dialogue partner, and the second dialogue partner replies to the first dialogue partner based on the statement, thereby satisfying the needs of the first dialogue partner.

[0048] The method provided in this application can be applied in multiple scenarios, such as in ordering food, booking tickets, querying services, consulting services, and other scenarios where services are provided according to the requirements of the first dialogue partner.

[0049] If this method is applied to a restaurant ordering scenario, the second dialogue partner is the voice assistant in terminal 101. When a user wants to reserve a table, terminal 101 determines the user's intent based on the dialogue, such as the reservation time, number of guests, and location, and then reserves the table for the user, thus fulfilling their dining needs. If this method is applied to a ticketing scenario, the second dialogue partner is a voice-enabled module embedded within an intelligent robot. When a user wants to buy a plane ticket, the intelligent robot determines the user's intent based on the dialogue, such as the departure time, departure city, and destination, and then purchases the ticket for the user, thus fulfilling their ticketing needs.

[0050] Terminal 101 can be at least one of mobile phones, tablets, PCs (Personal Computers), robots, etc. Server 102 can be at least one of a single server, a server cluster consisting of multiple servers, a cloud server, a cloud computing platform, and a virtualization center.

[0051] Figure 2 This is a flowchart of an intent information determination method provided in an embodiment of this application. See also... Figure 2 The method includes:

[0052] Step 201: Obtain multiple first audio pairs, where each first audio pair is the audio of a dialogue between a first dialogue object and a second dialogue object.

[0053] Step 202: Determine the audio tag for each first audio pair. The audio tag is used to represent the intent parameter of the first audio pair.

[0054] Step 203: Based on the audio tag of each first audio pair, filter out the first audio pairs whose intent parameters do not meet the preset parameters from multiple first audio pairs to obtain at least one second audio pair.

[0055] Step 204: Determine the target intent information of the first dialogue object based on at least one second audio pair.

[0056] In one possible implementation, determining the audio tag for each first audio pair includes:

[0057] Determine the slot vector for each slot in the multiple slots and the context vector for the multiple first audio pairs. The context vector is used to represent the context in which the multiple first audio pairs are located.

[0058] For each first audio pair, a first audio vector is determined for the first audio pair. The first audio vector is used to represent the association between the first audio pair and the slot.

[0059] The second audio vector is determined based on the slot vector of each slot, the context vectors of multiple first audio pairs, and the first audio vector of the first audio pair.

[0060] The audio tag of the first audio pair is determined based on the dimension of the second audio vector.

[0061] In another possible implementation, the process of determining the first audio vector of the first audio pair includes:

[0062] The first audio pair is encoded to obtain the third audio vector;

[0063] Based on the third audio vector and the slot vector of each slot, determine the probability distribution of each slot in each vector dimension of the first audio pair;

[0064] The first audio vector is determined based on the probability distribution of each slot in each vector dimension of the first audio pair and the third audio vector.

[0065] In another possible implementation, the process of determining the context vectors of multiple first audio pairs includes:

[0066] Based on the word vector of each word in the first audio pair, determine the probability distribution of each word in the first audio pair;

[0067] Based on the probability distribution of each word in the first audio pair, determine the sentence vector of the first audio pair;

[0068] Based on the sentence vector of the first audio pair, determine the probability distribution of each first audio pair among multiple first audio pairs;

[0069] Based on the probability distribution of each first audio pair and its corresponding sentence vector, the context vectors of multiple first audio pairs are determined.

[0070] In another possible implementation, the process of determining the slot vector includes:

[0071] Determine the domain vector of the slot and the slot vector of the slot to which the slot belongs;

[0072] The slot vector is obtained by determining the sum of the neighborhood vector and the slot position vector.

[0073] In another possible implementation, the audio tag of the first audio pair is determined based on the dimension of the second audio vector, including:

[0074] If the dimension of the second audio vector is the first dimension, the audio label of the first audio pair is determined to be the first label. The first label is used to indicate that the intent parameters of the first audio pair do not meet the preset parameters.

[0075] If the dimension of the second audio vector is the second dimension, the audio label of the first audio pair is determined as the second label. The second label is used to indicate that the intent parameters of the first audio pair meet the preset parameters.

[0076] In another possible implementation, the target intent information of the first dialogue object is determined based on at least one second audio pair, including:

[0077] Determine the target audio vector based on the first audio vector of at least one second audio pair;

[0078] Obtain the slot vector and the first dialogue audio for each of the multiple slots, where the first dialogue audio includes multiple first audio pairs;

[0079] Based on the target audio vector, the slot vector, and the fourth audio vector of the first dialogue audio, the intent information of the slot is determined.

[0080] Based on the intent information of each slot, the target intent information is determined.

[0081] In another possible implementation, the intent information of the slot is determined based on the target audio vector, the slot vector of the slot, and the fourth audio vector of the first dialogue audio, including:

[0082] Based on the target audio vector, the slot vector, and the fourth audio vector of the first dialogue audio, the intent vector of the slot and the predicted intent value are determined.

[0083] Based on the intent vector of the slot, determine the intent type of the slot;

[0084] Based on the intent type and predicted intent value of the slot, the actual intent value of the slot is determined, and the intent information of the slot is obtained.

[0085] This application provides a method for determining intent information. The method first determines the audio tag of each audio pair. Since the audio tag can reflect the intent parameters of the audio pair, audio pairs whose intent parameters do not meet the conditions, i.e., redundant audio pairs, can be filtered out from multiple audio pairs based on the audio tag of each audio pair, and audio pairs that meet the conditions are obtained. The target intent information is determined based on the audio pairs that meet the conditions, so as to avoid the interference of redundant information, accurately generate the target intent information, and thus improve the accuracy of the determined target intent information.

[0086] In this embodiment, the electronic device can directly determine the target intent information based on multiple first audio pairs, or it can determine the target intent information through an intent information determination model; no specific limitation is made in this regard. Here, we will first take the example of the electronic device directly determining the target intent information based on multiple first audio pairs for explanation.

[0087] Figure 3 This is a flowchart of an intent information determination method provided in an embodiment of this application, executed by an electronic device. See also... Figure 3 The method includes:

[0088] Step 301: The electronic device acquires multiple first audio pairs.

[0089] The first audio pair is the audio of a conversation between a first dialogue partner and a second dialogue partner. In this embodiment, the second dialogue partner can provide services to the first dialogue partner and complete tasks according to the needs of the first dialogue partner. Since the first dialogue partner may not state all of its needs at once, and may change its needs during the conversation, the second dialogue partner may need to have multiple rounds of conversation with the first dialogue partner. Therefore, the first audio pair is the audio of each round of conversation between the second dialogue partner and the first dialogue partner, and the beginning audio of the first audio pair is the audio of the second dialogue partner, while the ending audio is the audio of the first dialogue partner.

[0090] The first dialogue partner can be any user, and the second dialogue partner is a voice-enabled module embedded within the electronic device. No specific limitations are placed on the first and second dialogue partners. If the first dialogue partner is a user and the second dialogue partner is a voice assistant in the terminal, then the first audio pair consists of the start audio being the audio of the voice assistant speaking and the end audio being the audio of the user speaking.

[0091] In this step, the electronic device can acquire multiple first audio pairs through the following method: the electronic device acquires first dialogue audio and determines multiple first audio pairs based on the first dialogue audio. This first dialogue audio is the historical dialogue audio between the second dialogue object and the first dialogue object before the start of a new round of dialogue during the completion of the current task.

[0092] In this implementation, the electronic device can divide the multi-turn dialogue in the first dialogue audio into a dialogue history list according to the starting audio being the audio of the second dialogue partner speaking and the ending audio being the audio of the first dialogue partner speaking, and the dialogue history list includes multiple first audio pairs.

[0093] For example, the first dialogue audio H = [R1, U1, R2, U2, ..., R t U t ], the dialogue history list h = [(R1, U1), (R2, U2)..., (R t U t [], where H represents the audio of the first dialogue, R1 represents the audio of the second dialogue partner during the first round of dialogue, U1 represents the audio of the first dialogue partner during the first round of dialogue, and R t U represents the audio of the second dialogue partner during the t-th round of dialogue. t Let (R1, U1), (R2, U2), ..., (R2, U2) represent the audio of the first dialogue partner in the t-th round of conversation. t U t) represents the first audio pair, and the t-th round of dialogue is the last round of dialogue up to the current time.

[0094] Step 302: The electronic device determines the slot vector for each of the multiple slots.

[0095] In this step, for each slot, the electronic device determines the domain vector of the domain to which the slot belongs and the slot position vector of the slot to which the slot belongs, and determines the sum of the domain vector and the slot position vector of the slot to obtain the slot vector of the slot.

[0096] The electronic device can first determine the domain to which the slot belongs, determine the domain code corresponding to that domain, and then determine the domain vector based on the domain code. Correspondingly, the electronic device determines the slot position to which the slot belongs, determines the slot position code corresponding to that slot, and then determines the slot vector based on the slot position code. Therefore, it can be seen that the electronic device uses a combination of domain code and slot position code to construct the slot vector.

[0097] For example, the slot vector can be represented as s j =Embedding(Dp)+Embedding(Sq), where, s j Let represent the slot vector of the j-th slot, where j is the slot number. Embedding(Dp) represents the neighborhood vector, and Embedding(Sq) represents the slot vector, where j is an integer greater than 0.

[0098] Step 303: The electronic device determines the context vectors of multiple first audio pairs.

[0099] In this step, the electronic device can determine the context vector through the following steps (1) to (4):

[0100] (1) The electronic device determines the probability distribution of each word in the first audio pair based on the word vector of each word in the first audio pair.

[0101] The first audio pair includes multiple words. The electronic device can divide the first audio pair into multiple words, and determine the word vector corresponding to each word. This word vector is the encoded word vector. For each word, the electronic device transforms the word vector corresponding to the word to obtain a transformed word vector. Based on the transformed word vector of each word, the probability distribution of each word is determined.

[0102] In this implementation, the electronic device can convert word vectors using a preset conversion function to obtain converted word vectors. Then, a sum of values ​​is determined with a first value as the base and the product of the converted word vector of each word and a second value as the exponent. Finally, a ratio is determined between the sum of values ​​with the first value as the base and the product of the converted word vector of each word and the second value as the exponent, and the sum of values, thereby obtaining the probability distribution of each word.

[0103] For example, electronic devices transform word vectors using the following relation: u im Let tanh() represent the word vector of the m-th word in the i-th first audio pair after transformation, and let tanh() represent the preset transformation function. Let W1 represent the word vector of the m-th word in the i-th first audio pair before transformation, and let W1 and b1 represent the transformation parameters, where i and m are both integers greater than 0. The electronic device determines the probability distribution of each word using the following relationship: α im Let represent the probability of the m-th word in the i-th first audio pair, where n is the total number of words in the i-th first audio pair, the first value is the natural base e, and the second value is u. w .

[0104] (2) The electronic device determines the sentence vector of the first audio pair based on the probability distribution of each word in the first audio pair.

[0105] Electronic devices can determine the sum of the products of the probability of each word and its corresponding word vector to obtain the sentence vector of the first audio pair.

[0106] For example, an electronic device can determine the sentence vector of the first audio pair using the following relation: s i Let represent the sentence vector of the i-th first audio pair.

[0107] (3) The electronic device determines the probability distribution of each first audio pair among multiple first audio pairs based on the sentence vector of the first audio pair.

[0108] The electronic device can transform the sentence vector of the first audio pair to obtain the transformed sentence vector, and determine the probability distribution of each first audio pair based on the transformed sentence vector.

[0109] The electronic device determines the sum of the products of the sentence vector of each first audio pair and the fourth value, with the third value as the base. It then determines the ratio of the sum to the exponent of the product of the sentence vector of each first audio pair and the fourth value, with the third value as the base, to the sum, thus obtaining the probability distribution of each first audio pair.

[0110] For example, the electronic device transforms the sentence vector of the first audio pair using the following relation: r i =tanh(W2s) i +b2), r i Let W2 represent the sentence vector after the transformation of the i-th first audio pair, and let b2 represent the transformation parameters.

[0111] The electronic device determines the probability distribution of each first audio pair using the following relationship: αi Let represent the probability distribution of the i-th first audio pair, t represent the total number of first audio pairs, and the fourth value is u. s The third value can be the same as or different from the first value. Here, we will only take the case where the third value is the same as the first value, which is the natural base e, as an example.

[0112] (4) The electronic device determines the context vector of multiple first audio pairs based on the probability distribution of each first audio pair and its corresponding sentence vector.

[0113] The electronic device determines the sum of the products of the probability of each first audio pair and its corresponding sentence vector, thus obtaining a context vector for multiple first audio pairs. This context vector is used to represent the context in which the multiple first audio pairs are located.

[0114] For example, electronic devices determine the context vector using the following relationship: c represents the context vector.

[0115] In this embodiment, a hierarchical encoding structure is used to determine the context vectors of multiple first audio pairs. First, word-level attention is used to construct a sentence vector for each first audio pair. Then, the sentence vectors are further encoded, and sentence-level attention is used to fuse all the sentence vectors to construct the current required context vector, thereby assisting in determining the audio tag of the first audio pair.

[0116] Step 304: For each first audio pair, the electronic device determines the first audio vector of the first audio pair.

[0117] In this step, the electronic device can determine the first audio vector of the first audio pair through the following steps (1) to (3):

[0118] (1) The electronic device encodes the first audio pair to obtain the third audio vector.

[0119] The electronic device can employ a bidirectional GRU (Gate Recurrent Unit) structure to encode the first audio pair. This process can be as follows: the electronic device can determine the word vector for each word in the first audio pair based on pre-trained GloVe vectors (Global Vectors for Word Representation) and Char vectors. This word vector is obtained by combining the GloVe vectors and Char vectors. Then, this word vector is fed into the bidirectional GRU encoder for encoding, resulting in a third audio vector. This third audio vector is a sequence of word vectors composed of the encoded word vectors of multiple words in the first audio pair.

[0120] For example, an electronic device can represent the word vector encoding the previous word using the following relation: Let m be the word vector of the m-th word in the i-th first audio pair. The electronic device can represent the third audio vector using the following relation: h i Represents the third audio vector. This represents the word vector encoded by the m-th word in the i-th first audio pair.

[0121] It should be noted that the electronic device can also encode the first dialogue audio to determine the fourth audio vector of the first dialogue audio. This process is similar to the process of determining the third audio vector of the first audio pair. It also involves determining the word vector of each word in the first dialogue audio, and then sending the word vector into the bidirectional GRU encoder for encoding to obtain the fourth audio vector.

[0122] For example, the fourth audio vector can be represented as: H represents the fourth audio vector, w q Let q represent the word vector encoded by the q-th word in the first dialogue audio, and z represent the total number of words in the first dialogue audio. Both q and z are integers greater than 0.

[0123] (2) The electronic device determines the probability distribution of each slot in each vector dimension of the first audio pair based on the third audio vector and the slot vector of each slot.

[0124] For each first audio pair, the electronic device determines the sub-vectors of the third audio vector of the first audio pair in each vector dimension. For each slot, it determines the sum of the exponents of the product of the slot vector of the slot and the sub-vectors of the first audio pair in each vector dimension, with the fifth value as the base. It also determines the ratio of the product of the slot vector of the slot and the sub-vectors of the first audio pair in each vector dimension to the sum, with the fifth value as the base, to obtain the probability distribution of the slot in each vector dimension of the first audio pair.

[0125] For example, the electronic device determines the probability distribution of each slot in each vector dimension of each first audio pair using the following relationship: α iv This represents the probability of the i-th first audio pair in the v-th vector dimension. Let s represent the subvector of the i-th first audio pair in the v-th vector dimension, where V represents the total vector dimension of the first audio pair. j Let v represent the slot vector of the j-th slot, where v and V are both integers greater than 0.

[0126] (3) The electronic device determines the first audio vector based on the probability distribution of each slot in each vector dimension of the first audio pair and the third audio vector.

[0127] For each slot and each first audio pair, the electronic device determines the sum of the products of the probability of the slot in each vector dimension of the first audio pair and the third audio vector, thus obtaining the first audio vector associated with the slot for the first audio pair.

[0128] For example, the electronic device determines the first audio vector associated with each slot for the first audio pair using the following relationship: Where, h′ i Represents the i-th first audio pair and slot s j The relevant first audio vector.

[0129] It should be noted that the third audio vector obtained in step (1) of step 304 is independent of the slot, and different slots have different information of interest for the same first audio pair. Therefore, in this embodiment, an attention mechanism is introduced to update the audio vector of the first audio pair with the help of the slot vector, thereby obtaining the first audio vector of each first audio pair associated with each slot.

[0130] Another point to note is that, since it is necessary to determine the first audio vector of the first audio pair relative to the slot based on the slot vector of the slot, in this embodiment of the application, the electronic device first determines the slot vector of each slot, and then determines the first audio vector of each first audio pair relative to each slot. As for the order in which the electronic device determines the context vector of multiple first audio pairs and the slot vector of each slot, it can be set and changed as needed. For example, the electronic device first determines the context vector of multiple first audio pairs and then determines the slot vector of each slot, or the electronic device first determines the slot vector of each slot and then determines the context vector of multiple first audio pairs.

[0131] Step 305: The electronic device determines the second audio vector based on the slot vector of each slot, the context vectors of multiple first audio pairs, and the first audio vector of the first audio pair.

[0132] In this step, for each slot and each first audio pair, the electronic device can concatenate the slot vector of the slot, the context vectors of multiple first audio pairs, and the first audio vector of the first audio pair to obtain the concatenated vector. Then, the concatenated vector is sequentially activated and normalized to obtain the second audio vector of the first audio pair relative to the slot.

[0133] For example, electronic devices determine the second audio vector using the following relationship: l ij =Softmax(relu(W3·[h′) i sj ,c])),l ij Let W3 represent the second audio vector of the i-th first audio pair relative to the j-th slot, Softmax() represents the normalization function, ReLU() represents the activation function, and W3 represents the parameters.

[0134] Step 306: The electronic device determines the audio tag of the first audio pair based on the dimension of the second audio vector.

[0135] For each slot and each first audio pair, if the dimension of the first audio pair relative to the second audio vector of the slot is the first dimension, the electronic device determines that the audio label of the first audio pair relative to the slot is the first label, which is used to indicate that the intent parameters of the first audio pair do not meet the preset parameters.

[0136] If the dimension of the first audio pair relative to the second audio vector of the slot is the second dimension, the electronic device determines that the audio label of the first audio pair relative to the slot is the second label, and the second label is used to indicate that the intent parameters of the first audio pair meet the preset parameters.

[0137] In this embodiment, the first dimension can be 0-dimensional and the second dimension can be 1-dimensional. Binary classification is used to filter out first audio pairs that meet the conditions from multiple first audio pairs. The preset parameters can be set and changed as needed; in this embodiment, no specific limitations are imposed.

[0138] In this embodiment, the electronic device can perform binary classification labeling on the first audio pairs to determine whether each first audio pair is useful relative to each slot in a specific dialogue context. That is, when the intent parameters of the first audio pair do not meet the preset parameters, the electronic device determines that the label of the first audio pair relative to the slot is a useless label, and when the intent parameters of the first audio pair meet the preset parameters, the electronic device determines that the label of the first audio pair relative to the slot is a useful label. This can remove redundant first audio pairs to a certain extent, thereby improving the accuracy of the determined intent information.

[0139] See Figure 4 ,from Figure 4 As can be seen, the electronic device first determines word vectors, then sentence vectors based on the word vectors, and finally context vectors based on the sentence vectors. Based on the context vectors, slot vectors, and the first audio vector, the second audio vector is determined, and the audio tag for the first audio pair is determined based on the dimension of the second audio vector.

[0140] Step 307: Based on the audio tag of each first audio pair, the electronic device filters out first audio pairs whose intent parameters do not meet the preset parameters from multiple first audio pairs to obtain at least one second audio pair.

[0141] The electronic device can filter out first audio pairs whose intent parameters do not meet the preset parameters from multiple first audio pairs, and obtain at least one second audio pair whose intent parameters meet the preset parameters, that is, at least one second audio pair with useful tags is selected from multiple first audio pairs.

[0142] Step 308: The electronic device determines the target audio vector based on the first audio vector of at least one second audio pair.

[0143] For each slot, the electronic device can determine the sum of at least one second audio pair with respect to the first audio vector of that slot, to obtain at least one second audio pair with respect to the target audio vector of that slot.

[0144] In this embodiment of the application, the electronic device can also directly determine the target audio vector based on the audio tag and the first audio vector of each first audio pair. Accordingly, steps 307 and 308 can be replaced as follows: For each slot, the electronic device determines the audio tag value of each first audio pair relative to the slot based on the audio tag of each first audio pair relative to the slot. If the audio tag of the first audio pair relative to the slot is a first tag, that is, a useless tag, then the audio tag value of the first audio pair relative to the slot is determined to be 0. If the audio tag of the first audio pair relative to the slot is a second tag, that is, a useful tag, then the audio tag value of the first audio pair relative to the slot is determined to be 1. Then, the sum of the products of the first audio vector of each first audio pair relative to the slot and its audio tag value is determined to obtain the target audio vector of multiple first audio pairs relative to the slot.

[0145] In this implementation, the electronic device can determine the target audio vector using the following relationship: h select Indicates multiple first audio pairs relative to slot s j The target audio vector, L ij h′ represents the audio tag value of the i-th first audio pair relative to the j-th slot. i This indicates that the i-th first audio pair is relative to slot s. j The first audio vector.

[0146] Step 309: The electronic device acquires the slot vector and the first dialogue audio for each of the multiple slots.

[0147] The electronic device can obtain the slot vector of each slot obtained in step 302 and obtain the first dialogue audio obtained in step 301.

[0148] Step 310: The electronic device determines the intent information of the slot based on the target audio vector, the slot vector of the slot, and the fourth audio vector of the first dialogue audio.

[0149] This step can be achieved through the following steps (1) to (3), including:

[0150] (1) For each slot, the electronic device determines the intent vector and predicted intent value of the slot based on the target audio vector, the slot vector of the slot and the fourth audio vector of the first dialogue audio.

[0151] The process by which an electronic device determines the intent vector of a slot can be as follows: The electronic device inputs the target audio vector and the slot vector of the slot into the decoder. The decoder generates a first latent vector in the first decoding step using an attention mechanism. Based on the first latent vector and the fourth audio vector of the first dialogue audio, a first probability is determined. Based on the first probability and the fourth audio vector of the first dialogue audio, the intent vector of the slot is determined. The first probability can represent the importance of each word in the first dialogue audio. The higher the first probability, the more important the word is. Therefore, the intent vector obtained based on the first probability and the fourth audio vector is the vector obtained by combining the important words in the first dialogue audio.

[0152] For example, an electronic device can input the target audio vector and the slot vector of the slot into the decoder using the following relationship: The input during the first step of decoding by the decoder For slot vector s j h j(k-1) for h select The first hidden vector obtained after the first step of decoding is h. j1 .

[0153] The electronic device determines the product of the first latent vector and the fourth audio vector, and then normalizes this product value to obtain the first probability. For example, the electronic device determines the first probability using the following relationship: This represents the first probability of the j-th slot after the first step of decoding.

[0154] The electronic device determines the intent vector of the slot by multiplying the first probability by the fourth audio vector. For example, the electronic device determines the intent vector of the slot using the following relationship: c j1 This represents the intent vector of the j-th slot after decoding in step 1.

[0155] The process by which an electronic device determines the predicted intent value can be as follows: The decoder generates a second and third latent vector after complete decoding, using an attention mechanism based on the target audio vector and the slot vector. Based on the second latent vector and the fourth audio vector of the first dialogue audio, it determines the second probability of each word in the first dialogue audio. Based on the second latent vector and a pre-defined vocabulary, it determines the third probability of each word in the pre-defined vocabulary. Based on the second probability and the fourth audio vector, it determines the fifth audio vector. Based on the third latent vector, the fourth latent vector obtained from the previous decoding step, and the fifth audio vector, it determines the transition probability. Based on the transition probability, the second probability, and the third probability, it determines the predicted intent value.

[0156] In this implementation, the electronic device can determine the product of the second latent vector and the fourth audio vector, and then normalize the product value to obtain a second probability. This second probability represents the generation probability of each word in the first dialogue audio. For example, the electronic device determines the second probability using the following relationship: H represents the second probability of the j-th slot after the k-th decoding step, where k is the total number of decoding steps in the decoder, H represents the fourth audio vector, and h represents the second probability of the j-th slot after the k-th decoding step. jk Let represent the second hidden vector obtained after decoding in the k-th step, and Softmax() represent the normalization function.

[0157] The electronic device can determine the product of the second latent vector and the vocabulary vector of a preset vocabulary, and then normalize this product value to obtain a third probability, which represents the generation probability of each word in the preset vocabulary. For example, the electronic device determines the third probability using the following relationship: Let E represent the third probability of the j-th slot after the k-th decoding step, and let E represent the word vector of the preset word list.

[0158] The electronic device can determine the product of the second probability and the fourth audio vector to obtain the fifth audio vector. For example, the electronic device determines the fifth audio vector using the following relationship: c jk This represents the fifth audio vector of the j-th slot after the k-th decoding step.

[0159] The electronic device concatenates the third hidden vector, the fourth hidden vector obtained from the previous decoding step, and the fifth audio vector to obtain a concatenated vector. This concatenated vector is then activated to obtain a transition probability. This transition probability indicates whether a word is generated from a preset vocabulary or from the first dialogue audio. For example, the electronic device determines the transition probability using the following relationship: This represents the transition probability of the j-th slot after the k-th decoding step. This represents the third hidden vector of the j-th slot after the k-th decoding step. W4 represents the fourth hidden vector of the j-th slot after the (k-1)-th decoding step, and Sigmoid() represents the activation function.

[0160] The electronic device determines the product of the transition probability and the third probability to obtain the first product value. It then determines the difference between the sixth value and the transition probability, and multiplies this difference with the second probability to obtain the second product value. Finally, it determines the sum of the first and second product values ​​to obtain the predicted intent value, which represents the probability distribution of each word. The sixth value can be set and changed as needed; here, we only use a sixth value of 1 as an example. For instance, the electronic device determines the predicted intent value using the following formula: This indicates the predicted intention value.

[0161] After obtaining the predicted intent value, the electronic device can combine multiple words with higher probabilities based on the probability distribution of each word to obtain the combined word.

[0162] (2) The electronic device determines the intent type of the slot based on the intent vector of the slot.

[0163] In this embodiment of the application, the electronic device can pre-divide five intent types, namely “ptr”, “none”, “dontcare”, “yes” and “no”, and then normalize the intent vector of the slot to obtain the intent type distribution of the slot.

[0164] For example, an electronic device determines the intent type distribution of the slot using the following relationship: G j =Softmax(W5·c j1 ), G j W5 represents the intent type distribution of the j-th slot.

[0165] (3) The electronic device determines the actual intent value of the slot based on the intent type and predicted intent value of the slot, and obtains the intent information of the slot.

[0166] The electronic device uses the intent type with the highest probability in the intent type distribution as the intent type for that slot.

[0167] If the intent type of the slot is "ptr", the electronic device will use the word obtained by combining the predicted intent value as the actual intent value of the slot to obtain the intent information of the slot.

[0168] If the intent type of the slot is any one of "none", "dontcare", "yes", and "no", the electronic device uses this intent type value as the actual intent value for the slot to obtain the intent information of the slot. The intent information of the slot includes: the domain to which the slot belongs, the slot position, and the actual intent value. It can be described by a triple of domain-slot-actual intent value, where the domain is one of a set of preset domains, and the slot position is one of a set of preset slot positions.

[0169] Step 311: The electronic device determines the target intent information based on the intent information of each slot.

[0170] The target intent information is the intent information of the first dialogue object. The electronic device can combine the intent information of multiple slots to obtain the target intent information, which can be represented as: B F ={B1, B2, ..., B f}, B F Indicates target intent information, B1, B2, ... B f This contains intent information for multiple slots.

[0171] See Figure 5 ,from Figure 5 As can be seen, the electronic device determines the intent type and predicted intent value for each slot based on the target audio vector, slot vector, and fourth audio vector determined by the audio tag, and then determines the actual intent value based on the intent type and predicted intent value.

[0172] This application provides a method for determining intent information. The method first determines the audio tag of each audio pair. Since the audio tag can reflect the intent parameters of the audio pair, audio pairs whose intent parameters do not meet the conditions, i.e., redundant audio pairs, can be filtered out from multiple audio pairs based on the audio tag of each audio pair, and audio pairs that meet the conditions are obtained. The target intent information is determined based on the audio pairs that meet the conditions, so as to avoid the interference of redundant information, accurately generate the target intent information, and thus improve the accuracy of the determined target intent information.

[0173] This application uses the example of an electronic device determining target intent information through an intent information determination model to illustrate the following embodiments.

[0174] Figure 6 This is a flowchart of an intent information determination method provided in an embodiment of this application, executed by an electronic device. See also... Figure 6 The method includes:

[0175] Step 601: The electronic device acquires multiple first audio pairs.

[0176] This step is the same as step 301, and will not be repeated here.

[0177] Step 602: The electronic device determines the target intent information by selecting multiple first audio pairs from the input intent information model.

[0178] The electronic device inputs multiple first audio pairs into the intent information determination model. The intent information determination model is used to determine the audio tag of each first audio pair. Based on the audio tag of each first audio pair, the first audio pairs whose intent parameters do not meet the preset parameters are filtered out from the multiple first audio pairs to obtain at least one second audio pair. Based on at least one second audio pair, the target intent information is determined.

[0179] The process by which the electronic device determines the audio tag of each first audio pair through the intent information determination model is the same as the process by which the electronic device performs steps 302-306 above. The process by which the electronic device obtains at least one second audio pair is the same as the process by which the electronic device performs step 307. The process by which the electronic device determines the target intent information based on at least one second audio pair is the same as the process by which the electronic device performs steps 308-311. These details will not be repeated here.

[0180] In this embodiment of the application, the intent information determination model can be trained by the electronic device or by other electronic devices. The electronic device then uses the model to determine the target intent information. There is no specific limitation on this. Here, we only use the example of the electronic device training the intent information determination model for illustration.

[0181] The process of training an electronic device to obtain an intent information determination model can be as follows: the electronic device acquires multiple sample audio pairs and sample intent information, where the sample audio pairs are the dialogue audio between a first sample object and a second sample object. Based on these multiple sample audio pairs and sample intent information, the model is trained to obtain the intent information determination model.

[0182] In this implementation, the electronic device can obtain the real audio tag of each sample audio pair, the real intent value of each sample audio pair relative to each slot, and the real intent type of each slot. Then, it inputs the multiple sample audio pairs into an initial model, predicts the audio tags of the multiple sample audio pairs through the initial model, and determines at least one sample audio pair whose intent parameters satisfy preset parameters based on the audio tags predicted by the initial model. Based on the at least one sample audio pair, it determines the predicted intent value of each sample audio pair relative to each slot and the predicted intent type of each slot.

[0183] The initial model may include: an audio tagging module, an intent generation module, and an intent type model. The loss function for the audio tagging module can be: L h This represents the loss value of the audio tag module, where N represents the total number of slots, and l′ij Indicates the prediction of audio tags, This indicates the actual audio tag.

[0184] The loss function for the intent generation module can be: L p p′ represents the loss value of the module intended to generate. jk Indicates the predicted intention value. This represents the true intention value.

[0185] The loss function for the intent type module can be: L g G′ represents the loss value of the intent type module. j Indicates the type of predicted intent. Indicates the type of true intent.

[0186] During model training, the electronic device jointly trains these three modules together. This means training the model based on the real audio labels, predicted audio labels, real intent values, predicted intent values, real intent types, and predicted intent types, until the preset number of iterations is reached or the total loss value is minimized. This loss value is the sum of the loss values ​​of the three modules, i.e., L = L0. p +L g +L h Minimum, where L is the total loss value.

[0187] When training the model, the electronic device can use the MultiWOZ dataset, which includes 56,668 samples in the training set, 7,374 samples in the validation set, and 7,368 samples in the test set. The ratio of the training set, validation set, and test set is approximately 8:1:1.

[0188] In the embodiments of this application, in the dialogue state tracking task, that is, in the process of determining intent information, the historical dialogue audio is analyzed and utilized in a more granular manner, the inherent structural relationship of the dialogue is used for model training, and the interference of redundant information in the dialogue history on the dialogue state tracking task is reduced by judging the usefulness of the historical dialogue audio, so as to achieve more accurate and stable dialogue state tracking.

[0189] This application provides an intent information determination method. The method inputs multiple first audio pairs into an intent information determination model, which determines the audio tag for each audio pair. Since the audio tag reflects the intent parameters of the audio pair, audio pairs whose intent parameters do not meet the conditions (i.e., redundant audio pairs) can be filtered out from the multiple audio pairs based on their audio tags, resulting in audio pairs that meet the conditions. The target intent information is then determined based on these audio pairs, thereby avoiding interference from redundant information, accurately generating the target intent information, and improving the accuracy of the determined target intent information.

[0190] Next, experimental data will be used to characterize the intent information provided in the embodiments of this application and determine the performance of the model.

[0191] First, the performance of the audio tag module.

[0192] The audio tagging module is primarily used to determine the usefulness of each audio pair for a specific slot in a given context. This can be viewed as a classification subtask; therefore, its performance is measured using classic classification metrics: precision, recall, and F1 score. Referring to Table 1, it can be seen that the precision, recall, and F1 score reached 89.01%, 88.53%, and 88.77% respectively, meeting expectations.

[0193] Table 1 Performance of the audio tag module

[0194] Accuracy (%) Recall rate (%) F1(%) 89.01 88.53 88.77

[0195] See Figure 7 , Figure 7 This is the PR curve, which represents the relationship between precision and recall. The larger the area under the curve, and the closer it is to 1, the better the performance. Figure 7 As can be seen from the data, the area of ​​the curve is 0.94, which is relatively large, indicating that the model performs well.

[0196] Second, the performance of the intent type module.

[0197] In this embodiment of the application, the intent type is expanded to a five-class type. See Table 2. Table 2 compares the performance of the three-class type with that of the five-class type. As can be seen from Table 2, the model does not lose classification performance on the other three types when two new types are added. On the contrary, due to the more detailed classification, the interference of other types during classification is reduced, which optimizes the intent type prediction to a certain extent.

[0198] Table 2 Performance of the Intent Type Module

[0199]

[0200]

[0201] Third, the overall performance of the model

[0202] When analyzing the overall performance of the model, we focus on two key evaluation metrics: joint accuracy and slot prediction accuracy.

[0203] in,

[0204]

[0205] The joint accuracy rate is the percentage of samples where the intent information is completely and correctly predicted, while the slot prediction accuracy rate is the average accuracy rate of intent information prediction per sample. Table 3 shows a performance comparison between the model provided in this application and several other models. As can be seen from Table 3, compared to other models, the model provided in this application has a joint accuracy rate of 50.01% and a slot prediction accuracy rate of 97.18%, demonstrating better experimental results.

[0206] Table 3 Overall performance of each model

[0207] Model Combined accuracy (%) Tank prediction accuracy (%) GLAD 35.57 95.44 Neural Reading 41.10 -- SUMBT 46.65 96.44 TRADE* 47.86 96.85 COMER 48.79 -- The model provided in this application 50.01 97.18

[0208] Joint accuracy is a very stringent evaluation metric. Although different methods have similar slot prediction accuracy, there are significant differences in joint accuracy. This reflects the requirement for the completeness of the predicted intent information in dialogue state tracking. It is hoped that the method can predict the current dialogue state, i.e. the target intent information, as completely and accurately as possible.

[0209] Furthermore, we calculated the performance of the method provided in this application embodiment on the five domains of the dataset. Referring to Table 4, it can be seen from Table 4 that compared with the TRADE model, the performance of this model is improved in all five domains, with the improvement being particularly significant in the taxi domain, where there is an improvement of nearly 20% in joint accuracy. In addition, there are also improvements of 2% to 4% in the other four domains.

[0210] Table 4 Overall performance in different fields

[0211]

[0212] Fourth, model performance under different dialogue lengths.

[0213] The method provided in this application mainly extracts useful information from dialogue audio and minimizes interference from redundant information. The dialogue audio is divided based on the number of dialogue rounds to analyze the model's performance. Referring to Table 5, it can be seen that the model provided in this application exhibits better performance in cases with longer dialogue rounds.

[0214] Table 5. Performance of the model under different dialogue lengths

[0215]

[0216] In summary, the method provided in this application, by determining the usefulness of each audio pair for each slot in a specific context, filters useful information, thereby removing redundant information and improving the accuracy of the determined intent information. Furthermore, in the slot value type determination stage, the use of a five-category determination module effectively improves the performance of this stage.

[0217] Figure 8 This is a schematic diagram of the structure of an intent information determination device provided in an embodiment of this application. See also... Figure 8 The device includes:

[0218] The first acquisition module 801 is used to acquire multiple first audio pairs, where the first audio pair is the dialogue audio between a first dialogue object and a second dialogue object;

[0219] The first determining module 802 is used to determine the audio tag for each first audio pair, where the audio tag is used to represent the intent parameter of the first audio pair;

[0220] The filtering module 803 is used to filter out first audio pairs whose intent parameters do not meet preset parameters from multiple first audio pairs based on the audio tags of each first audio pair, so as to obtain at least one second audio pair;

[0221] The second determining module 804 is used to determine the target intent information of the first dialogue object based on at least one second audio pair.

[0222] In one possible implementation, the first determining module 802 is used to determine the slot vector of each slot in the plurality of slots and the context vector of the plurality of first audio pairs, the context vector being used to represent the context in which the plurality of first audio pairs are located; for each first audio pair, a first audio vector of the first audio pair is determined, the first audio vector being used to represent the association relationship between the first audio pair and the slot; based on the slot vector of each slot, the context vector of the plurality of first audio pairs and the first audio vector of the first audio pairs, a second audio vector is determined; based on the dimension of the second audio vector, the audio tag of the first audio pair is determined.

[0223] In another possible implementation, the first determining module 802 is used to encode the first audio pair to obtain a third audio vector; based on the third audio vector and the slot vector of each slot, determine the probability distribution of each slot in each vector dimension of the first audio pair; and based on the probability distribution of each slot in each vector dimension of the first audio pair and the third audio vector, determine the first audio vector.

[0224] In another possible implementation, the first determining module 802 is used to determine the probability distribution of each word in the first audio pair based on the word vector of each word in the first audio pair; determine the sentence vector of the first audio pair based on the probability distribution of each word in the first audio pair; determine the probability distribution of each first audio pair in multiple first audio pairs based on the sentence vector of the first audio pair; and determine the context vector of multiple first audio pairs based on the probability distribution of each first audio pair and its corresponding sentence vector.

[0225] In another possible implementation, the first determining module 802 is used to determine the domain vector of the domain to which the slot belongs and the slot vector of the slot to which the slot belongs; and to determine the sum of the domain vector and the slot vector of the slot to obtain the slot vector of the slot.

[0226] In another possible implementation, the first determining module 802 is used to determine the audio tag of the first audio pair as a first tag if the dimension of the second audio vector is the first dimension, the first tag being used to indicate that the intent parameters of the first audio pair do not meet the preset parameters; and to determine the audio tag of the first audio pair as a second tag if the dimension of the second audio vector is the second dimension, the second tag being used to indicate that the intent parameters of the first audio pair meet the preset parameters.

[0227] In another possible implementation, the second determining module 804 is used to determine a target audio vector based on a first audio vector of at least one second audio pair; acquire the slot vector of each slot in a plurality of slots and a first dialogue audio, the first dialogue audio including a plurality of first audio pairs; determine the intent information of the slot based on the target audio vector, the slot vector of the slot and a fourth audio vector of the first dialogue audio; and determine the target intent information based on the intent information of each slot.

[0228] In another possible implementation, the second determining module 804 is used to determine the intent vector and predicted intent value of the slot based on the target audio vector, the slot vector of the slot and the fourth audio vector of the first dialogue audio; determine the intent type of the slot based on the intent vector of the slot; and determine the actual intent value of the slot based on the intent type of the slot and the predicted intent value, thereby obtaining the intent information of the slot.

[0229] This application provides an intent information determination device. The device first determines the audio tag of each audio pair. Since the audio tag can reflect the intent parameters of the audio pair, audio pairs whose intent parameters do not meet the conditions, i.e., redundant audio pairs, can be filtered out from multiple audio pairs based on the audio tag of each audio pair, and audio pairs that meet the conditions are obtained. The target intent information is determined based on the audio pairs that meet the conditions, thereby avoiding interference from redundant information, accurately generating the target intent information, and thus improving the accuracy of the determined target intent information.

[0230] Figure 9 This is a schematic diagram of the structure of an intent information determination device provided in an embodiment of this application. See also... Figure 9 The device includes:

[0231] The second acquisition module 901 is used to acquire multiple first audio pairs, where the first audio pair is the dialogue audio between the first dialogue object and the second dialogue object;

[0232] Input module 902 is used to determine the target intent information of the first dialogue object by inputting multiple first audio pairs into the intent information determination model;

[0233] The intent information determination model is used to determine the audio tag for each first audio pair. The audio tag is used to represent the intent parameters of the first audio pair. Based on the audio tag of each first audio pair, the first audio pairs whose intent parameters do not meet the preset parameters are filtered out from multiple first audio pairs to obtain at least one second audio pair. Based on at least one second audio pair, the target intent information is determined.

[0234] In one possible implementation, the device further includes:

[0235] The third acquisition module is used to acquire multiple sample audio pairs and sample intent information. The sample audio pairs are the dialogue audio between the first sample object and the second sample object.

[0236] The training module is used to train the model based on multiple sample audio pairs and sample intent information, and obtain the intent information to determine the model.

[0237] This application provides an intent information determination device. The device inputs multiple first audio pairs into an intent information determination model, and determines the audio tag of each audio pair through the intent information determination model. Since the audio tag can reflect the intent parameters of the audio pair, audio pairs whose intent parameters do not meet the conditions, i.e., redundant audio pairs, can be filtered out from multiple audio pairs based on the audio tag of each audio pair, and audio pairs that meet the conditions are obtained. The target intent information is determined based on the audio pairs that meet the conditions, thereby avoiding interference from redundant information, accurately generating the target intent information, and thus improving the accuracy of the determined target intent information.

[0238] It should be noted that the intent information determination device provided in the above embodiments is only illustrated by the division of the above functional modules when determining intent information. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the electronic device can be divided into different functional modules to complete all or part of the functions described above. In addition, the intent information determination device and the intent information determination method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0239] If the electronic device is provided as a terminal, see [link to relevant documentation]. Figure 10 , Figure 10 This illustration shows a structural block diagram of a terminal 1000 provided in an exemplary embodiment of this application. The terminal 1000 may be a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The terminal 1000 may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.

[0240] Typically, terminal 1000 includes a processor 1001 and a memory 1002.

[0241] Processor 1001 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1001 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1001 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1001 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1001 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0242] The memory 1002 may include one or more computer-readable storage media, which may be non-transitory. The memory 1002 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1002 are used to store at least one program code, which is executed by the processor 1001 to implement the intent information determination method provided in the method embodiments of this application.

[0243] In some embodiments, the terminal 1000 may also optionally include a peripheral device interface 1003 and at least one peripheral device. The processor 1001, memory 1002, and peripheral device interface 1003 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1003 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 1004, a touch display screen 1005, a camera 1006, an audio circuit 1007, a positioning component 1008, and a power supply 1009.

[0244] Peripheral device interface 1003 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1001 and memory 1002. In some embodiments, processor 1001, memory 1002 and peripheral device interface 1003 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1001, memory 1002 and peripheral device interface 1003 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0245] The radio frequency (RF) circuit 1004 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1004 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1004 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1004 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1004 can communicate with other terminals via at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: metropolitan area networks (MANs), various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks (WLANs), and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1004 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0246] Display screen 1005 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1005 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1001 for processing. In this case, display screen 1005 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1005, serving as the front panel of terminal 1000; in other embodiments, there may be at least two display screens, respectively disposed on different surfaces of terminal 1000 or in a folded design; in still other embodiments, display screen 1005 may be a flexible display screen, disposed on a curved or folded surface of terminal 1000. Furthermore, display screen 1005 may also be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1005 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0247] The camera assembly 1006 is used to acquire images or videos. Optionally, the camera assembly 1006 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1006 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0248] The audio circuit 1007 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 1001 for processing, or input to the radio frequency circuit 1004 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal 1000. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1001 or the radio frequency circuit 1004 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1007 may also include a headphone jack.

[0249] The positioning component 1008 is used to determine the current geographical location of the terminal 1000 in order to enable navigation or LBS (Location Based Service). The positioning component 1008 can be a positioning component based on the US GPS (Global Positioning System), China's BeiDou system, Russia's Granas system, or the EU's Galileo system.

[0250] The power supply 1009 is used to power the various components in the terminal 1000. The power supply 1009 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When the power supply 1009 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.

[0251] In some embodiments, the terminal 1000 further includes one or more sensors 1010. The one or more sensors 1010 include, but are not limited to: an accelerometer 1011, a gyroscope 1012, a pressure sensor 1013, a fingerprint sensor 1014, an optical sensor 1015, and a proximity sensor 1016.

[0252] Accelerometer 1011 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by terminal 1000. For example, accelerometer 1011 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 1001 can control touchscreen 1005 to display the user interface in landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1011. Accelerometer 1011 can also be used for games or for acquiring user motion data.

[0253] The gyroscope sensor 1012 can detect the orientation and rotation angle of the terminal 1000. The gyroscope sensor 1012, in conjunction with the accelerometer sensor 1011, can collect 3D motion data from the user on the terminal 1000. Based on the data collected by the gyroscope sensor 1012, the processor 1001 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0254] The pressure sensor 1013 can be disposed on the side bezel of the terminal 1000 and / or on the lower layer of the touch display screen 1005. When the pressure sensor 1013 is disposed on the side bezel of the terminal 1000, it can detect the user's grip signal on the terminal 1000, and the processor 1001 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 1013. When the pressure sensor 1013 is disposed on the lower layer of the touch display screen 1005, the processor 1001 can control the operable controls on the UI interface based on the user's pressure operation on the touch display screen 1005. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0255] The fingerprint sensor 1014 is used to collect a user's fingerprint. The processor 1001 identifies the user based on the fingerprint collected by the fingerprint sensor 1014, or vice versa. When the user's identity is identified as trusted, the processor 1001 authorizes the user to perform relevant sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 1014 can be located on the front, back, or side of the terminal 1000. When the terminal 1000 has physical buttons or a manufacturer's logo, the fingerprint sensor 1014 can be integrated with the physical buttons or manufacturer's logo.

[0256] An optical sensor 1015 is used to collect ambient light intensity. In one embodiment, the processor 1001 can control the display brightness of the touch screen 1005 based on the ambient light intensity collected by the optical sensor 1015. Specifically, when the ambient light intensity is high, the display brightness of the touch screen 1005 is increased; when the ambient light intensity is low, the display brightness of the touch screen 1005 is decreased. In another embodiment, the processor 1001 can also dynamically adjust the shooting parameters of the camera assembly 1006 based on the ambient light intensity collected by the optical sensor 1015.

[0257] The proximity sensor 1016, also known as a distance sensor, is typically mounted on the front panel of the terminal 1000. The proximity sensor 1016 is used to detect the distance between the user and the front of the terminal 1000. In one embodiment, when the proximity sensor 1016 detects that the distance between the user and the front of the terminal 1000 is gradually decreasing, the processor 1001 controls the touchscreen display 1005 to switch from a screen-on state to a screen-off state; when the proximity sensor 1016 detects that the distance between the user and the front of the terminal 1000 is gradually increasing, the processor 1001 controls the touchscreen display 1005 to switch from a screen-off state to a screen-on state.

[0258] Those skilled in the art will understand that Figure 10 The structure shown does not constitute a limitation on terminal 1000 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0259] If the electronic device is provided as a terminal and server, see [link to relevant documentation]. Figure 11 , Figure 11 This is a schematic diagram of a server structure provided in an embodiment of this application. The server 1100 can vary considerably due to different configurations or performance. It may include one or more central processing units (CPUs) 1101 and one or more memories 1102. The memory 1102 stores at least one line of program code, which is loaded and executed by the processor 1001 to implement the methods provided in the above-described method embodiments. Of course, the server 1100 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 1100 may also include other components for implementing device functions, which will not be elaborated here.

[0260] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including program code, the instructions of which can be executed by a processor in an electronic device to perform the intent information determination method in the above embodiments. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0261] In an exemplary embodiment, a computer program product is also provided, which stores at least one piece of program code, which is loaded and executed by a processor to implement the intent information determination method in the embodiments of this application.

[0262] In some embodiments, the computer program involved in the present application embodiments may be deployed and executed on a computer device, or executed on multiple computer devices located in one location, or executed on multiple computer devices distributed in multiple locations and interconnected through a communication network. Multiple computer devices distributed in multiple locations and interconnected through a communication network may constitute a blockchain system.

[0263] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0264] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for determining intent information, characterized in that, The method includes: Acquire multiple first audio pairs, where each first audio pair is the audio of a dialogue between a first dialogue object and a second dialogue object; Determine an audio tag for each first audio pair, the audio tag being used to represent the intent parameter of the first audio pair; Based on the audio tags of each first audio pair, filter out first audio pairs whose intent parameters do not meet the preset parameters from the plurality of first audio pairs to obtain at least one second audio pair; Based on the at least one second audio pair, determine the target intent information of the first dialogue object; determining the audio tag for each first audio pair includes: Determine the slot vector of each slot in the plurality of slots and the context vector of the plurality of first audio pairs, wherein the context vector is used to represent the context of the plurality of first audio pairs; For each first audio pair, a first audio vector is determined for the first audio pair, the first audio vector being used to represent the association between the first audio pair and the slot; The second audio vector is determined based on the slot vector of each slot, the context vector of the plurality of first audio pairs, and the first audio vector of the first audio pair; The audio tag of the first audio pair is determined based on the dimension of the second audio vector.

2. The method according to claim 1, characterized in that, The process of determining the first audio vector of the first audio pair includes: The first audio pair is encoded to obtain the third audio vector; Based on the third audio vector and the slot vector of each slot, determine the probability distribution of each slot in each vector dimension of the first audio pair; The first audio vector is determined based on the probability distribution of each slot in each vector dimension of the first audio pair and the third audio vector.

3. The method according to claim 1, characterized in that, The process of determining the context vectors of the plurality of first audio pairs includes: Based on the word vector of each word in the first audio pair, determine the probability distribution of each word in the first audio pair; Based on the probability distribution of each word in the first audio pair, determine the sentence vector of the first audio pair; Based on the sentence vector of the first audio pair, determine the probability distribution of each of the plurality of first audio pairs; based on the probability distribution of each first audio pair and its corresponding sentence vector, determine the context vector of the plurality of first audio pairs.

4. The method according to claim 1, characterized in that, The process of determining the slot vector of the slot includes: Determine the domain vector of the domain to which the slot belongs and the slot vector of the slot to which the slot belongs; The sum of the neighborhood vector and the slot position vector of the slot is determined to obtain the slot vector of the slot.

5. The method according to claim 1, characterized in that, Determining the audio tag of the first audio pair based on the dimension of the second audio vector includes: If the dimension of the second audio vector is the first dimension, the audio tag of the first audio pair is determined to be the first tag. The first tag is used to indicate that the intent parameters of the first audio pair do not meet the preset parameters. If the dimension of the second audio vector is the second dimension, the audio tag of the first audio pair is determined to be the second tag, and the second tag is used to indicate that the intent parameters of the first audio pair satisfy the preset parameters.

6. The method according to claim 1, characterized in that, Determining the target intent information of the first dialogue object based on the at least one second audio pair includes: Determine the target audio vector based on the first audio vector of the at least one second audio pair; Obtain the slot vector and the first dialogue audio for each of the multiple slots, wherein the first dialogue audio includes the multiple first audio pairs; Based on the target audio vector, the slot vector of the slot, and the fourth audio vector of the first dialogue audio, the intent information of the slot is determined; The target intent information is determined based on the intent information of each slot.

7. The method according to claim 6, characterized in that, The step of determining the intent information of the slot based on the target audio vector, the slot vector of the slot, and the fourth audio vector of the first dialogue audio includes: Based on the target audio vector, the slot vector of the slot, and the fourth audio vector of the first dialogue audio, the intent vector and the predicted intent value of the slot are determined. Based on the intent vector of the slot, determine the intent type of the slot; Based on the intent type and predicted intent value of the slot, the actual intent value of the slot is determined, and the intent information of the slot is obtained.

8. A method for determining intent information, characterized in that, The method includes: Acquire multiple first audio pairs, where each first audio pair is the audio of a dialogue between a first dialogue object and a second dialogue object; The multiple first audio pairs are input into the intent information determination model to obtain the target intent information of the first dialogue object; the intent information determination model is used to determine the audio tag of each first audio pair, the audio tag being used to represent the intent parameter of the first audio pair; based on the audio tag of each first audio pair, first audio pairs whose intent parameters do not meet preset parameters are filtered out from the multiple first audio pairs to obtain at least one second audio pair; based on the at least one second audio pair, the target intent information is determined; the process of determining the intent information determination model includes: Acquire multiple sample audio pairs and sample intent information, wherein the sample audio pairs are dialogue audio between a first sample object and a second sample object; Based on the multiple sample audio pairs and the sample intent information, a model is trained to obtain the intent information determination model.

9. An intent information determination device, characterized in that, The device includes: The first acquisition module is used to acquire multiple first audio pairs, wherein the first audio pair is the audio of a dialogue between a first dialogue object and a second dialogue object; A first determining module is configured to determine an audio tag for each first audio pair, the audio tag representing the intent parameter of the first audio pair, wherein determining the audio tag for each first audio pair includes: Determine the slot vector of each slot in the plurality of slots and the context vector of the plurality of first audio pairs, wherein the context vector is used to represent the context of the plurality of first audio pairs; For each first audio pair, a first audio vector is determined for the first audio pair, the first audio vector being used to represent the association between the first audio pair and the slot; The second audio vector is determined based on the slot vector of each slot, the context vector of the plurality of first audio pairs, and the first audio vector of the first audio pair; Based on the dimension of the second audio vector, determine the audio tag of the first audio pair; The filtering module is used to filter out first audio pairs whose intent parameters do not meet preset parameters from the plurality of first audio pairs based on the audio tags of each first audio pair, so as to obtain at least one second audio pair; The second determining module is used to determine the target intent information of the first dialogue object based on the at least one second audio pair.

10. An intent information determination device, characterized in that, The device includes: The second acquisition module is used to acquire multiple first audio pairs, wherein the first audio pair is the dialogue audio between the first dialogue object and the second dialogue object; The input module is used to obtain the target intent information of the first dialogue object from the plurality of first audio pairs input intent information determination models, and the process of determining the intent information determination model includes: Acquire multiple sample audio pairs and sample intent information, wherein the sample audio pairs are dialogue audio between a first sample object and a second sample object; Based on the multiple sample audio pairs and the sample intent information, a model is trained to obtain the intent information determination model; The intent information determination model is used to determine the audio tag of each first audio pair. The audio tag is used to represent the intent parameter of the first audio pair. Based on the audio tag of each first audio pair, first audio pairs whose intent parameters do not meet the preset parameters are filtered out from the plurality of first audio pairs to obtain at least one second audio pair. Based on the at least one second audio pair, the target intent information is determined.

11. An electronic device, characterized in that, The electronic device includes a processor and a memory, wherein the memory stores at least one piece of program code, which is loaded and executed by the processor to implement the intent information determination method as described in any one of claims 1 to 7 or the intent information determination method as described in claim 8.

12. A computer-readable storage medium, characterized in that, The storage medium stores at least one piece of program code, which is loaded and executed by a processor to implement the intent information determination method as described in any one of claims 1 to 7 or the intent information determination method as described in claim 8.

13. A computer program product, characterized in that, The computer program product stores at least one piece of program code, which is loaded and executed by a processor to implement the intent information determination method as described in any one of claims 1 to 7 or the intent information determination method as described in claim 8.

Citation Information

Patent Citations

  • Information interaction method and device based on intention recognition, equipment and storage medium

    CN111104495A

  • Intention recognition model processing method and device, computer equipment and storage medium

    CN113220828A