Speech processing device, method and program

The utterance processing device uses intention and search condition estimations to enhance the accuracy of recognizing invalid utterances, enabling more flexible dialogue management by combining deep learning models.

JP7775028B2Active Publication Date: 2025-11-25KK TOSHIBA
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2021178503
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-01
Publication Date
2025-11-25
Estimated Expiration
2041-11-01

AI Technical Summary

Technical Problem

Existing dialogue systems struggle to accurately recognize invalid utterances, as collecting example utterances for training is difficult and time-consuming, and existing methods based on general rules or machine learning are insufficient.

Method used

An utterance processing device that includes a receiving unit, a first estimation unit for intention estimation, a second estimation unit for search condition estimation, and a determination unit to determine whether an utterance is invalid based on the results of both estimations, using deep learning models to enhance accuracy.

Benefits of technology

The device achieves higher accuracy in recognizing invalid utterances by combining intention and search condition estimations, allowing for more flexible and accurate dialogue management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007775028000001
    Figure 0007775028000001
  • Figure 0007775028000002
    Figure 0007775028000002
  • Figure 0007775028000003
    Figure 0007775028000003
Patent Text Reader

Abstract

To recognize invalid speech with high accuracy.SOLUTION: A speech sentence processor according to an embodiment is equipped with a reception unit, a first estimation unit, a second estimation unit, and a determination unit. The reception unit accepts a speech sentence from a user. The first estimation unit estimates intention expressed by the user based on the speech sentence. The second estimation unit estimates search conditions specified by the user based on the speech sentence. The determination unit determines whether the speech sentence is an invalid speech or not in each of a first estimation result concerning the estimated intention and a second estimation result concerning the estimated search condition.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] An embodiment of the present invention relates to an utterance processing apparatus, method, and program. [Background technology]

[0002] Dialogue systems that perform actions desired by a user by repeating a dialogue with the user multiple times are expected to be used in many situations. To ensure smooth dialogue progression and avoid malfunctions, a dialogue system should correctly distinguish between utterances from a user and those that belong to the domain targeted by the system and are capable of handling them (hereinafter also referred to as "valid utterances"), and those that do not belong to the domain and are not capable of handling them (hereinafter also referred to as "invalid utterances"). In particular, it is desirable for a dialogue system to respond to invalid utterances by, for example, asking the user about the intention of the utterance or ambiguities, or informing the user that it is not capable of handling them, without performing a predetermined action. Therefore, a dialogue system is required to recognize invalid utterances with high accuracy.

[0003] To achieve the above objective, a first method is to train a dialogue system using a large number of example utterances related to invalid utterances collected in advance by an engineer (e.g., an AI engineer). However, collecting such example utterances is practically difficult, and it is difficult to achieve the objective using methods based on general rules or machine learning. A second method is to have an engineer collect examples of invalid utterances from the dialogue system's logs and train the dialogue system using the collected examples of invalid utterances. However, this method is not practical in view of the time and physical costs required by the engineer. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent No. 6667504 Summary of the Invention [Problem to be solved by the invention]

[0005] The problem to be solved by the present invention is to recognize invalid utterances with high accuracy. [Means for solving the problem]

[0006] An utterance sentence processing device according to an embodiment includes a receiving unit, a first estimation unit, a second estimation unit, and a determination unit. The receiving unit receives an utterance sentence from a user. The first estimation unit estimates an intention expressed by the user based on the utterance sentence. The second estimation unit estimates search conditions specified by the user based on the utterance sentence. The determination unit determines whether the utterance sentence is an invalid utterance for each of a first estimation result related to the estimated intention and a second estimation result related to the estimated search conditions. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 is a block diagram showing an example of the functional configuration of an utterance sentence processing apparatus according to a first embodiment. [Figure 2] FIG. 2 is a flowchart showing an example of the operation of the speech sentence processing apparatus according to the first embodiment. [Figure 3] FIG. 10 is a block diagram showing an example of the functional configuration of an utterance sentence processing apparatus according to a second embodiment. [Figure 4] FIG. 10 is a flowchart showing an example of the operation of the speech sentence processing apparatus according to the second embodiment. [Figure 5] FIG. 10 is a diagram showing an example of a dialogue realized by the speech sentence processing apparatus according to the second embodiment. [Figure 6] FIG. 10 is a diagram showing an example of a dialogue realized by the speech sentence processing apparatus according to the second embodiment. [Figure 7] FIG. 10 is a diagram showing an example of a dialogue realized by the speech sentence processing apparatus according to the second embodiment. [Figure 8] FIG. 11 is a block diagram showing an example of the functional configuration of an utterance sentence processing apparatus according to a third embodiment. [Figure 9] FIG. 11 is a flowchart showing an example of the operation of the speech sentence processing apparatus according to the third embodiment. [Figure 10] FIG. 1 is a block diagram showing an example of the hardware configuration of an utterance sentence processing device according to first to third embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0008] Hereinafter, an utterance processing device, method, and program according to the embodiment will be described with reference to the drawings. In the following embodiments, parts with the same reference numerals perform similar operations, and redundant explanations will be omitted as appropriate.

[0009] (First embodiment) FIG. 1 is a block diagram showing an example of the functional configuration of a spoken sentence processing device 1 according to the first embodiment. The spoken sentence processing device 1 is a device that performs various operations by processing utterances from a user and is an example of a dialogue system. For example, the spoken sentence processing device 1 is a search dialogue system that performs an information search about restaurants using utterances from a user. When an utterance intended to search for information about restaurants (e.g., "Looking for a Japanese restaurant," "Are there any Japanese restaurants?", "I want to eat Japanese food," or "Japanese food is good") is input, the utterance processing device 1 searches for restaurants serving Japanese food from among multiple restaurants and notifies the user of the search results, including the restaurant name and location information. On the other hand, when an utterance not intended to search for information (e.g., "By the way, a Japanese restaurant has opened nearby," or "A store that sells Japanese clothing") is input, the utterance processing device 1 does not perform an information search about restaurants and instead notifies the user, for example, that it cannot respond. To realize such operations, the utterance processing device 1 according to the first embodiment includes a receiving unit 111, a first estimation unit 112, a second estimation unit 113, and a determination unit 114.

[0010] The reception unit 111 receives an utterance from a user. The utterance may be text input by the user via an input means such as a keyboard, or text converted from the user's speech by speech recognition. The reception unit 111 then outputs the received utterance to each of the first estimation unit 112 and the second estimation unit 113.

[0011] The first estimation unit 112 estimates the intention expressed by the user based on the uttered sentence. Specifically, the first estimation unit 112 estimates the intention of the uttered sentence by applying the intention estimation model 121 to the uttered sentence received from the reception unit 111. Subsequently, the first estimation unit 112 outputs the estimation result by the intention estimation model 121 (intention estimation result; first estimation result) to the determination unit 114.

[0012] The intention estimation model 121 estimates, as an estimation result (first estimation result), a probability (first probability) that the utterance sentence belongs to a class (intention class) corresponding to each of a plurality of candidates related to the intention. Specifically, the intention estimation model 121 is a classification model based on deep learning, and outputs the probability that the input utterance sentence belongs to each of a plurality of intention classes set in advance. The parameters of the intention estimation model 121 may be trained in advance using training data. Note that the intention estimation model 121 may employ an existing intention estimation method (e.g., mood analysis) based on the entire sentence of the utterance sentence.

[0013] First, if the utterance sentence is, for example, "Looking for a Japanese restaurant," "Are there any Japanese restaurants?", "I want to eat Japanese food," or "Japanese food is good," the intention of the utterance sentence is "Specifying a search condition." Second, if the utterance sentence is, for example, "Hello" or "Thank you," the intention of the utterance sentence is "Greeting." Third, if the utterance sentence is, for example, "Yes" or "No," the intention of the utterance sentence is "Affirmative or Negative." Fourth, if the utterance sentence is, for example, "Start over" or "Show me the previous restaurant again," the intention of the utterance sentence is "System operation." As such, since each utterance sentence from a user contains a different intention, the intention estimation model 121 only needs to have an intention class corresponding to each of the intentions of the utterance sentences assumed in advance.

[0014] When an engineer trains the intention estimation model 121, he or she assigns an intention class to each of K (K is a natural number) pre-specified intentions, such as intention class 1 for "specifying search conditions," intention class 2 for "greeting," intention class 3 for "affirmative or negative," intention class 4 for "system operation," and so on. At this time, the engineer considers "invalid utterance" as a single intention, just like other intentions, and assigns intention class K+1 to "invalid utterance." Next, the engineer assigns one of intention classes 1 to K+1 to each of multiple example utterances collected in advance, and creates labeled training data. The engineer then trains the parameters of the intention estimation model 121 using the labeled training data. Thus, the intention estimation model 121 is trained to classify multiple example utterances collected in advance into one of intention classes 1 to K+1. The trained intention estimation model 121 outputs the probability that a user's utterance corresponds to each of intention classes 1 to K, as well as the probability that the utterance corresponds to intention class K+1.

[0015] In the above example, the engineer does not need to set the intent class K+1 for "invalid utterance." In this case, the engineer assigns one of intent classes 1 to K to each of multiple example utterances collected in advance, and creates labeled training data. The engineer then trains the parameters of the intention estimation model 121 using the labeled training data. In this way, the intention estimation model 121 is trained to classify multiple example utterances collected in advance into one of the intent classes 1 to K. The trained intention estimation model 121 outputs the probability that an utterance from a user corresponds to each of the intent classes 1 to K.

[0016] The second estimation unit 113 estimates search criteria specified by the user based on the uttered sentence. Specifically, the second estimation unit 113 estimates search criteria of the uttered sentence by applying at least one search criteria estimation model 122 to the uttered sentence received from the reception unit 111. Subsequently, the second estimation unit 113 outputs estimation results (search criteria estimation results; second estimation results) by each search criteria estimation model 122 to the determination unit 114. At this time, it is sufficient that the number of search criteria estimation models 122 applied is the same as the number of attributes used as search criteria.

[0017] The search condition estimation model 122 estimates the probability (second probability) that the uttered sentence corresponds to each of a plurality of values ​​related to the search conditions as an estimation result (second estimation result). Specifically, the search condition estimation model 122 is a model based on deep learning, and outputs the probability that the input uttered sentence corresponds to each of a plurality of values ​​set in advance. Note that the search condition estimation model 122 may focus on keywords included in the uttered sentence and output the probability that each value corresponds.

[0018] In this embodiment, it is assumed that multiple values ​​corresponding to one attribute are set in advance in one search criteria estimation model 122. For example, the search criteria estimation model 122 for the cuisine genre has multiple values ​​corresponding to the cuisine genre (e.g., Japanese, Western, Chinese, Asian / ethnic). On the other hand, the search criteria estimation model 122 for the location has multiple values ​​corresponding to the location (e.g., around station A, shopping district B, suburbs of city C, within X minutes' walking distance, within Ym's straight-line distance). On the other hand, the search criteria estimation model 122 for the price range has multiple values ​​corresponding to the price range (e.g., high, medium, low). Each value may be extracted from a list of multiple restaurants to be searched. Furthermore, each value is not limited to a specific value specifying a predetermined value, and may include a null value that does not specify a predetermined value. In other words, a null value is a value that the search criteria estimation model 122 does not assume as a search target.

[0019] The determination unit 114 determines whether or not the uttered sentence is an invalid utterance in each of the estimation result regarding the estimated intention (first estimation result) and the estimation result regarding the estimated search condition (second estimation result). Specifically, the determination unit 114 receives the first estimation result from the first estimation unit 112 and the second estimation result from the second estimation unit 113, and determines whether or not the uttered sentence is an invalid utterance in each estimation result.

[0020] In this embodiment, the determination unit 114 calculates multiple types of determination results. First, the determination unit 114 calculates a determination result (first determination result) indicating that the uttered sentence has been determined to be not an invalid utterance in each of the first estimation result and the second estimation result. Second, the determination unit 114 calculates a determination result (second determination result) indicating that the uttered sentence has been determined to be an invalid utterance in either the first estimation result or the second estimation result. Third, the determination unit 114 calculates a determination result (third determination result) indicating that the uttered sentence has been determined to be an invalid utterance in each of the first estimation result and the second estimation result. Of these, the second determination result includes a determination result (fourth determination result) indicating that the uttered sentence has been determined to be an invalid utterance in the first estimation result and that the uttered sentence has been determined to be not an invalid utterance in the second estimation result. Furthermore, the second judgment result includes a judgment result (fifth judgment result) indicating that the utterance sentence was determined to be not an invalid utterance in the first judgment result and that the utterance sentence was determined to be an invalid utterance in the second estimation result.

[0021] The determination unit 114 determines whether the user's utterance sentence is an invalid utterance based on the first estimation result. For example, the determination unit 114 identifies the intention class with the highest probability among the probabilities corresponding to each of the intention classes 1 to K+1 output from the first estimation unit 112. If the identified intention class is intention class K+1 (i.e., the intention class corresponding to the invalid utterance), the determination unit 114 determines that the user's utterance sentence is an invalid utterance. Conversely, if the identified intention class is not intention class K+1, the determination unit 114 determines that the user's utterance sentence is not an invalid utterance (i.e., a valid utterance).

[0022] For example, assume that the probability that an utterance sentence corresponds to intention class 1 "specifying a search condition" is 0.1, the probabilities that it corresponds to intention classes 2 to K are all 0, and the probability that it corresponds to intention class K+1 "invalid utterance" is 0.9. In this case, the intention class assigned the highest probability is intention class K+1, and therefore the determination unit 114 determines that the user's utterance sentence is an invalid utterance.

[0023] Alternatively, the determination unit 114 determines that the user's utterance is an invalid utterance if the probability of each of the intention classes 1 to K output from the first estimation unit 112 being equal to or less than a predetermined threshold. The predetermined threshold can be set to an arbitrary value by an engineer managing the utterance processing device 1. Conversely, if at least one of the probabilities of each of the intention classes 1 to K exceeds the predetermined threshold, the determination unit 114 determines that the user's utterance is not an invalid utterance (i.e., a valid utterance). Note that the determination unit 114 may determine that the utterance is an invalid utterance if all of the probabilities are less than the predetermined threshold, and may determine that the utterance is not an invalid utterance if at least one of the probabilities is equal to or greater than the predetermined threshold.

[0024] For example, assume that the probability of falling into any of the intention classes 1 to K is less than 0.5 and the predetermined threshold is 0.5. In this case, since the probability of falling into any of the intention classes 1 to K is equal to or less than the predetermined threshold, the determination unit 114 determines that the user's utterance is an invalid utterance.

[0025] As described above, the determination unit 114 may determine whether the user's utterance sentence in the first estimation result is an invalid utterance based on the probability that the utterance sentence falls into the intention class K+1, or may determine based on the probability that the utterance sentence falls into the intention classes 1 to K. Both approaches may be applied in combination with each other, or only one of them may be applied.

[0026] On the other hand, the determination unit 114 determines whether the user's utterance sentence is an invalid utterance based on the second estimation result. For example, if the probability of each of the multiple values ​​related to the search condition output from the second estimation unit 113 corresponding to the utterance sentence is equal to or less than a predetermined threshold, the determination unit 114 determines that the user's utterance sentence is an invalid utterance. The predetermined threshold can be set to an arbitrary value by an engineer managing the utterance sentence processing device 1. Conversely, if at least one of the probabilities corresponding to the multiple values ​​related to the search condition exceeds the predetermined threshold, the determination unit 114 determines that the user's utterance sentence is not an invalid utterance (i.e., a valid utterance). Note that the determination unit 114 may determine that the utterance sentence is an invalid utterance if all the probabilities are less than the predetermined threshold, and may determine that the utterance sentence is not an invalid utterance if at least one of the probabilities is equal to or greater than the predetermined threshold.

[0027] For example, the determination unit 114 focuses on the probability that each of multiple values ​​related to the cuisine genre output from the search condition estimation model 122 for the cuisine genre corresponds to. Here, it is assumed that the search condition estimation model 122 for the cuisine genre outputs the probability that the search conditions of the uttered sentence correspond to each of four types of values ​​(Japanese cuisine, Western cuisine, Chinese cuisine, and Asian / ethnic). For example, it is assumed that the probabilities that each of "Japanese cuisine" to "Asian / ethnic" corresponds to is less than 0.5, and the predetermined threshold is 0.5. In this case, since the probabilities that each value corresponds to are all equal to or less than the predetermined threshold, the determination unit 114 determines that the sentence uttered by the user is an invalid utterance.

[0028] Next, the determination unit 114 performs the same determination as above for the search criteria estimation models 122 other than the search criteria estimation model 122 for the cuisine genre. Specifically, the determination unit 114 determines whether the user's utterance is an invalid utterance for each search criteria estimation model 122, including the search criteria estimation model 122 for location and the search criteria estimation model 122 for price range. If the utterance is determined to be an invalid utterance in each search criteria estimation model 122, the determination unit 114 determines that the user's utterance is an invalid utterance. Conversely, if the utterance is determined to be not an invalid utterance in at least one search criteria estimation model 122, the determination unit 114 determines that the user's utterance is not an invalid utterance.

[0029] Alternatively, the determination unit 114 identifies the value assigned the highest probability among the respective probabilities corresponding to multiple values ​​related to the search criteria. If the identified value is a null value (i.e., a value that does not specify a specific value), the determination unit 114 determines that the user's utterance is an invalid utterance. An operation similar to the above can be performed for each of the multiple search criteria estimation models 122. If a null value is identified for all attributes targeted by the multiple search criteria estimation models 122, the determination unit 114 determines that the user's utterance is an invalid utterance. Conversely, if a null value is not identified in at least one search criteria estimation model 122, the determination unit 114 determines that the user's utterance is not an invalid utterance.

[0030] 2 is a flow diagram showing an operation example of the spoken sentence processing device 1 according to the first embodiment. The start point in this operation example may be the point at which the spoken sentence processing device 1 receives an utterance sentence from a user for the first time, or the point at which the spoken sentence processing device 1 receives the next utterance sentence after receiving multiple utterance sentences from a user.

[0031] In step S101, the receiving unit 111 receives an utterance sentence from a user, and outputs the received utterance sentence to the first estimating unit 112 and the second estimating unit 113, respectively.

[0032] In step S102, the first estimation unit 112 estimates the user's intention of the utterance by applying the intention estimation model 121 to the utterance sentence received from the reception unit 111. Subsequently, the first estimation unit 112 outputs the estimation result (first estimation result) to the determination unit 114.

[0033] In step S103, the second estimation unit 113 estimates the user's search criteria by applying at least one search criteria estimation model 122 to the utterance received from the reception unit 111. Subsequently, the second estimation unit 113 outputs the estimation result (second estimation result) to the determination unit 114.

[0034] In step S104, the determination unit 114 receives the estimation result of the intention from the first estimation unit 112 and the estimation result of the search condition from the second estimation unit 113, and determines whether the sentence uttered by the user is an invalid utterance for each of the received estimation results. Subsequently, the determination unit 114 calculates multiple types of determination results (first determination result, second determination result, third determination result).

[0035] Note that steps S102 and S103 may be executed simultaneously, or step S103 may be executed prior to step S102. Furthermore, after step S104, the utterance sentence processing device 1 may end a series of processes and then wait for the reception of the next utterance sentence.

[0036] The above describes the utterance sentence processing device 1 according to the first embodiment. The utterance sentence processing device 1 according to the first embodiment determines whether an utterance sentence from a user is an invalid utterance by using an estimator based on intention estimation and an estimator based on search condition estimation. Here, assume that one estimation result determines that the utterance sentence is not an invalid utterance, while the other estimation result determines that the utterance sentence is an invalid utterance (i.e., a case where two types of estimation results are contradictory). In this way, even if one estimator fails to detect an invalid utterance, the utterance sentence processing device 1 can recognize the invalid utterance by using the estimation result from the other estimator. Therefore, the utterance sentence processing device 1 can recognize invalid utterances with higher accuracy than when using only one type of estimator.

[0037] Furthermore, in the above case, the utterance sentence processing apparatus 1 according to the first embodiment can obtain information about which of the estimator based on intention estimation and the estimator based on search condition estimation failed to detect the invalid utterance. The utterance sentence processing apparatus 1 according to the second embodiment generates a response sentence based on this information, and therefore can advance a more flexible dialogue compared to when only one type of estimator is used.

[0038] (Second embodiment) 3 is a block diagram showing an example of a functional configuration of the utterance sentence processing device 1 according to the second embodiment. The utterance sentence processing device 1 according to the second embodiment further includes a response generation unit 115 and an output unit 116 in addition to the components (reception unit 111, first estimation unit 112, second estimation unit 113, and determination unit 114) included in the utterance sentence processing device 1 according to the first embodiment.

[0039] The response generation unit 115 generates response sentences corresponding to the first, second, and third judgment results, respectively. As described above, the first judgment result is a judgment result indicating that the utterance sentence was determined to be a valid utterance by both estimators, the second judgment result is a judgment result indicating that the utterance sentence was determined to be a valid utterance by one estimator and an invalid utterance by the other estimator, and the third judgment result is a judgment result indicating that the utterance sentence was determined to be an invalid utterance by both estimators. Furthermore, the fourth judgment result included in the second judgment result is a judgment result indicating that the utterance sentence was determined to be an invalid utterance by the intention estimator and a valid utterance by the search condition estimator. Meanwhile, the fifth judgment result included in the second judgment result is a judgment result indicating that the utterance sentence was determined to be a valid utterance by the intention estimator and an invalid utterance by the search condition estimator.

[0040] Specifically, the response generation unit 115 generates a response sentence by referring to a response sentence template 123 corresponding to each determination result output from the determination unit 114. Subsequently, the response generation unit 115 outputs the generated response sentence to the output unit 116. Note that the response generation unit 115 may generate the response sentence by further referring to the estimation result from the first estimation unit 112 (first estimation result) and the estimation result from the second estimation unit 113 (second estimation result).

[0041] The response sentence templates 123 are a plurality of standard sentences prepared in advance for each determination result. First, as the response sentence template 123 corresponding to the first determination result, for example, a standard sentence including a placeholder (XX) into which predetermined search results can be embedded (e.g., "How about XX restaurant?") is prepared. The predetermined search results include search results for each of cuisine genre, location, and price range.

[0042] Second, the response sentence template 123 corresponding to the fourth judgment result among the second judgment results includes a placeholder (XX) into which a predetermined search result can be embedded, and is a standard phrase (e.g., "Is the cuisine genre XX?") that prompts the user to confirm the search conditions estimated by the second estimation unit 113.

[0043] On the other hand, as the response sentence template 123 corresponding to the fifth determination result among the second determination results, a fixed phrase according to the intention estimated by the first estimation unit 112 is prepared. For example, if the estimated intention is "specify search conditions," a fixed phrase (e.g., "What cuisine genre would you like?") that prompts the user to specify at least one search condition is prepared. This is because it is assumed that the user has specified an attribute of the search conditions that is not supported by the utterance sentence processing device 1. Instead, a fixed phrase (e.g., "Search not possible with those conditions") that informs the user that a search cannot be performed using the search conditions entered by the user may be prepared. This is because it is assumed that the user has specified an unknown value that corresponds to an attribute of the search conditions that is supported by the utterance sentence processing device 1.

[0044] Thirdly, as the response sentence template 123 corresponding to the third determination result, a fixed phrase (for example, "Is there any other information you would like to know?") is prepared to convey that the utterance sentence processing device 1 cannot handle the request.

[0045] The output unit 116 outputs the response sentence. Specifically, the output unit 116 receives the response sentence from the response generation unit 115 and outputs the received response sentence to the user. The output unit 116 may output the response sentence to a display device (e.g., a display or a smartphone), and the display device may display the response sentence to the user. Alternatively, the output unit 116 may output the response sentence to an audio device (e.g., a speaker or a smartphone), and the audio device may convert the response sentence into voice and play it back to the user. According to any one of these output modes, the user can visually or audibly recognize the response sentence to the spoken sentence.

[0046] 4 is a flow diagram showing an example of the operation of the utterance sentence processing device 1 according to the second embodiment. Steps S201 to S204 according to the second embodiment are the same as steps S101 to S104 according to the first embodiment. It is assumed that in step S204, the determination unit 114 has calculated any one of the first, second, and third determination results described above for the utterance sentence from the user.

[0047] In step S205, the response generation unit 115 branches the process depending on the type of the determination result received from the determination unit 114. First, if the response generation unit 115 receives a first determination result (i.e., if the utterance sentence is determined to be a valid utterance in both estimation results), the process proceeds to step S206. Second, if the response generation unit 115 receives a second determination result (i.e., if the utterance sentence is determined to be an invalid utterance in only one of the estimation results), the process proceeds to step S207. Third, if the response generation unit 115 receives a third determination result (i.e., if the utterance sentence is determined to be an invalid utterance in both estimation results), the process proceeds to step S208.

[0048] In step S206, the response generation unit 115 generates a response sentence for presenting the search results based on the estimated intention, the estimated search condition, and the response sentence template 123. For example, the response generation unit 115 searches for information that meets specific conditions based on the estimated search conditions and calculates the search results. Next, the response generation unit 115 generates a response sentence by embedding the calculated search results in the response sentence template 123 that corresponds to the search results.

[0049] In step S207, the response generation unit 115 generates a response sentence for prompting the user for confirmation based on the intention estimation result, the search condition estimation result, and the response sentence template 123. For example, the response generation unit 115 generates a response sentence by selecting a response sentence template 123 that explains the situation from among the multiple response sentence templates 123.

[0050] In step S208, the response generation unit 115 generates a response sentence indicating that the user's utterance is an invalid utterance, based on the intention estimation result, the search condition estimation result, and the response sentence template 123. For example, the response generation unit 115 generates the response sentence by selecting the response sentence template 123 indicating that the response is not possible.

[0051] In step S209, the output unit 116 receives the response sentence from the response generation unit 115 and outputs the received response sentence to an output device such as a display device or an audio device. The output device outputs the response sentence to the user.

[0052] After step S209, the utterance sentence processing apparatus 1 may end the series of processes and wait for the reception of the next utterance sentence.

[0053] 5, 6, and 7 are diagrams showing an example of a dialogue realized by the utterance sentence processing apparatus 1 according to the second embodiment. FIG. 5 illustrates three utterance sentences from a user and response sentences from the utterance sentence processing apparatus 1 to each utterance sentence. The utterance sentence processing apparatus 1 receives utterance sentences U1, U2, and U3 from the user, estimates intentions I1, I2, and I3 of the utterance sentences and search conditions C1, C2, and C3, and then outputs response sentences S1, S2, and S3. Each intention is estimated by an intention estimation model 121. Each search condition consists of a combination of three attributes (cuisine genre, location, and price range), and each value is estimated by a search condition estimation model 122 corresponding to each attribute.

[0054] For example, in the first turn (U1-S1) of the dialogue, the utterance processing device 1 estimates the intention I1 "Specify search conditions" for the utterance U1 and the search conditions C1 "Cuisine genre: Japanese cuisine, Location: North side of the station, Price range:" based on the utterance U1 from the user, where price range, an attribute related to the search conditions C1, is null. Therefore, based on the estimated intention I1 and the search conditions C1, the utterance processing device 1 outputs a response sentence S1 "Do you have a preference for the price range?" that asks the user about the search conditions related to the price range. In this way, in a general situation, if the estimated intention is "Specify search conditions," a specific value is estimated for any one of the attributes related to the search conditions.

[0055] Meanwhile, in the third turn (U3-S3) of the dialogue, the utterance processing device 1 estimates the intention I3 "invalid utterance" for the utterance U3 and the search criteria C3 "cuisine genre:, location:, price range:" based on the utterance U3 from the user "It's close from here, so I think I'll go there." Here, each attribute related to the search criteria C3 has an empty value. Therefore, the utterance processing device 1 outputs a response sentence S3 "Is there any other information you would like to know?" indicating that the response is not possible.

[0056] In the upper part of FIG. 6 (FIG. 6a), based on the user's utterance U11 "My favorite food is Japanese food," the utterance processing device 1 estimates the intention I11 "invalid utterance" and the search criteria C11 "food genre: Japanese food, location:, price range:" for the utterance U11. Next, the utterance processing device 1 determines that the utterance U11 is an invalid utterance as a result of the intention estimation, and determines that the utterance U11 is a valid utterance as a result of the search criteria estimation (corresponding to the fourth determination result). Therefore, the utterance processing device 1 outputs a response sentence S11 "Is it correct that the food genre is Japanese food?" to prompt the user to confirm the estimated search criteria "Japanese food."

[0057] In the lower part of Figure 6 (Figure 6b), we consider a case in which a conventional utterance processing device only has an estimator for intention estimation and does not have an estimator for search condition estimation. Based on an utterance U12 similar to an utterance U11, the device estimates the intention I12 "invalid utterance" for the utterance U12. The device then outputs a response S12 "Is there any other information you would like to know?" informing the user that the device is unable to respond. In other words, even in the same situation, a device with only one type of estimator cannot proceed with a flexible dialogue, such as generating a response prompting the user for confirmation, as in the utterance processing device 1 with two types of estimators.

[0058] In the upper part of Figure 7 (Figure 7a), the spoken sentence processing device 1 estimates the intention I21 "Specify search conditions" and the search conditions C21 "Cuisine genre:, Location:, Price range:" for the spoken sentence U21 based on the user's spoken sentence U21 "I want to eat something." Next, the spoken sentence processing device 1 determines that the utterance U21 is a valid utterance as a result of the intention estimation, and determines that the utterance U21 is an invalid utterance as a result of the search condition estimation (corresponding to the fifth determination result). Therefore, the spoken sentence processing device 1 outputs a response sentence S21 "What cuisine genre would you like?" to prompt the user to specify at least one search condition.

[0059] In the lower part of Figure 7 (Figure 7b), we consider a case in which a conventional utterance processing device only has an estimator for estimating search conditions, but does not have an estimator for estimating intentions. Based on an utterance U22 similar to an utterance U21, the device estimates search conditions C22 for the utterance U22, "Cuisine genre:, Location:, Price range:." The device then outputs a response S22, "Is there any other information you would like to know?", informing the user that the device is unable to respond. In other words, even in the same situation, a device with only one type of estimator cannot proceed with a flexible dialogue, such as generating a response that prompts the user to specify at least one search condition, as in the utterance processing device 1 with two types of estimators.

[0060] The above has described the utterance sentence processing apparatus 1 according to the second embodiment. The utterance sentence processing apparatus 1 according to the second embodiment determines whether an utterance sentence from a user is an invalid utterance by using an estimator based on intention estimation and an estimator based on search condition estimation. Next, the utterance sentence processing apparatus 1 generates and outputs an appropriate response sentence based on the estimation results obtained from each estimator.

[0061] For example, if the intention estimation estimates that an utterance is an invalid utterance, while the search condition estimation estimates that "Cuisine genre: Japanese cuisine" is the search condition of the utterance, the utterance processing device 1 can output a response sentence asking the user about the search condition. If the user responds affirmatively to the response sentence, the utterance processing device 1 can recognize that the intention estimation has failed. Conversely, if the user responds negatively to the response sentence, the utterance processing device 1 can recognize that the search condition estimation has failed. In other words, the utterance processing device 1 can classify the cause of the estimation failure in which estimator. Furthermore, by using a template for the response sentence corresponding to the cause of the failure, the utterance processing device 1 can return a response sentence corresponding to each cause. Therefore, the utterance processing device 1 can proceed with a more flexible dialogue.

[0062] (Third embodiment) 8 is a block diagram showing an example of a functional configuration of the utterance sentence processing device 1 according to the third embodiment. The utterance sentence processing device 1 according to the third embodiment further includes a storage unit 117 and a sample storage unit 124 in addition to the components (reception unit 111, first estimation unit 112, second estimation unit 113, and determination unit 114) included in the utterance sentence processing device 1 according to the first embodiment.

[0063] The first estimation unit 112 estimates, as a result of intention estimation (first estimation result), a similarity between a feature of a sentence uttered by the user and a feature of the sentence uttered in each of the multiple samples stored in the sample storage unit 124. Specifically, the first estimation unit 112 extracts a feature of the sentence uttered by the user by applying the intention estimation model 121 to the sentence uttered received from the reception unit 111. Next, the first estimation unit 112 estimates a similarity between the extracted feature and a feature of the sentence uttered in each of the multiple samples stored in the sample storage unit 124. Thereafter, the first estimation unit 112 outputs each estimated similarity to the determination unit 114. Note that the first estimation unit 112 may output the feature of the sentence uttered by the user to the storage unit 117.

[0064] The similarity is, for example, the cosine similarity between vectors representing the features of the utterance. The vectors are hidden vectors calculated by a deep learning-based model, such as the intention estimation model 121. If the respective cosine similarities are equal to or less than a predetermined threshold, the determination unit 114 determines that the user's utterance is an invalid utterance in the first estimation result. Alternatively, the similarity may be the Euclidean distance between the vectors. In this case, if the respective Euclidean distances are greater than a predetermined threshold, the determination unit 114 determines that the user's utterance is an invalid utterance in the first estimation result. Alternatively, the determination unit 114 may identify the features of samples near the features of the utterance from the user based on a nearest neighbor method, and thereby identify the intention label to which the sample belongs. Incidentally, the determination unit 114 determines whether the user's utterance is an invalid utterance in the estimation result of the search condition (second estimation result) using a method similar to the method described in the first embodiment.

[0065] The storage unit 117 receives the determination result from the determination unit 114 and stores the feature quantity of the sentence uttered by the user in the sample storage unit 124 according to the type of the received determination result. Specifically, the storage unit 117 stores the feature quantity of the sentence uttered by the user according to the second determination result (i.e., a determination result indicating that the first estimation result from the first estimation unit 112 and the second estimation result from the second estimation unit 113 are contradictory). In particular, the storage unit 117 stores the feature quantity of the sentence uttered by the user when the intention estimation estimates "search condition specification" and no specific value has been estimated for any attribute in the search condition estimation. Alternatively, the storage unit 117 may store the feature quantity of the sentence uttered by the user when the intention estimation estimates "search condition specification" and a specific value has been specified in the search condition estimation but the probability that the value corresponds to the specified value is lower than a preset threshold.

[0066] The sample storage unit 124 stores multiple samples, each of which associates an intention label corresponding to multiple intention candidates with a feature of the utterance corresponding to the intention label. In other words, each sample is a pair of an intention label and a feature of the utterance, similar to the intention class. Each intention label represents a pre-determined intention. That is, the intention label is one of K labels, including “search condition specification,” “greeting,” “affirmative or negative,” “system operation,” etc., and K+1 labels, including one “invalid utterance.” An engineer managing the utterance processing device 1 may apply the intention estimation model 121 to multiple utterance examples collected in advance to extract feature of the utterance from each example, and assign at least one extracted feature to each intention label.

[0067] 9 is a flowchart showing an example of the operation of the speech sentence processing apparatus 1 according to the third embodiment. Steps S301 to S304 according to the third embodiment are the same as steps S101 to S104 according to the first embodiment, except for step S302.

[0068] In step S302, the first estimation unit 112 extracts features by applying the intention estimation model 121 to a sentence uttered by a user. Then, the first estimation unit 112 calculates a similarity between the extracted features and features of at least one sentence uttered stored in the sample storage unit 124.

[0069] In step S305, the storage unit 117 selects whether to store the user's utterance sentence based on the determination result of the determination unit 114. If the storage unit 117 selects to store the utterance sentence (Yes in step S305), the process proceeds to step S306. On the other hand, if the storage unit 117 selects not to store the utterance sentence (No in step S305), the series of operations ends.

[0070] In step S306, the storage unit 117 stores the feature amount of the user's utterance sentence in the sample storage unit 124. After step S306, the utterance sentence processing apparatus 1 may end the series of operations and wait for the reception of the next utterance sentence.

[0071] The above describes the spoken sentence processing apparatus 1 according to the third embodiment. The spoken sentence processing apparatus 1 according to the third embodiment determines whether an utterance from a user is an invalid utterance by using an estimator based on intention estimation and an estimator based on search condition estimation. Specifically, the spoken sentence processing apparatus 1 extracts features of the utterance by applying the intention estimation model 121 to the utterance from the user, and calculates a similarity between the extracted features and features of at least one utterance stored in the sample storage unit 124. Next, the spoken sentence processing apparatus 1 performs intention estimation based on the calculated similarity. Thereafter, the spoken sentence processing apparatus 1 stores the features of the utterance from the user according to the determination result.

[0072] According to the third embodiment, even if the intention estimation model 121 assigns a high probability to an erroneous estimation result, the intention estimation error can be automatically corrected based on the estimation result from a different perspective, i.e., search condition estimation. As the utterance sentence processing device 1 is operated, the number of cases like those described above and the number of samples increase, and therefore the types of invalid utterances that can be correctly recognized increase. As a result, the utterance sentence processing device 1 can improve the robustness of intention estimation.

[0073] The utterance sentence processing device 1 according to the third embodiment is applicable not only to the intention "specifying a search condition" exemplified above, but also to other intentions (e.g., greeting, affirmative or negative, system operation). For example, the utterance sentence processing device 1 according to the third embodiment is also applicable to cases where the result of intention estimation is "greeting," "affirmative or negative," or "system operation," but the result of search condition estimation does not match any of the values ​​assumed in advance.

[0074] Furthermore, engineers and others who manage the utterance processing device 1 can improve the detection accuracy of the system by reusing the extracted features of the user's invalid utterances in another dialogue system different from the utterance processing device 1.

[0075] The utterance sentence processing apparatus 1 according to the third embodiment may further include the response generation unit 115 and the output unit 116 according to the second embodiment. The utterance sentence processing apparatus 1 according to a combination of both embodiments can store the feature amount of the utterance sentence from the user according to the determination result, and generate and output a response sentence based on the determination result.

[0076] 10 is a block diagram showing an example of the hardware configuration of the spoken sentence processing device 1 according to the first to third embodiments. The spoken sentence processing device 1 includes, as hardware, a processing circuit 11, a memory 12, a display 13, an input interface 14, and a communication interface 15.

[0077] The processing circuitry 11 controls the operation of the utterance sentence processing device 1. The processing circuitry 11 has processors such as a CPU (Central Processing Unit), an MPU (Micro Processing Unit), and a GPU (Graphics Processing Unit) as hardware. The processing circuitry 11 executes each program deployed in the memory 12 via at least one processor, thereby realizing each unit corresponding to each program (e.g., a reception unit 111, a first estimation unit 112, a second estimation unit 113, a determination unit 114, a response generation unit 115, an output unit 116, and a storage unit 117). Note that each unit can be realized by a processing circuit 11 consisting of a single processor or a processing circuit 11 combining multiple processors.

[0078] The memory 12 stores information such as data and programs used by the processing circuit 11. The memory 12 has a semiconductor memory element such as a random access memory (RAM) as hardware. The memory 12 may be a drive device that reads and writes information from and to an external storage device such as a magnetic disk (floppy disk, hard disk), a magneto-optical disk (MO), an optical disk (CD, DVD, Blu-ray), a flash memory (USB flash memory, memory card, SSD), or a magnetic tape. The storage area of ​​the memory 12 may be located inside or in an external storage device of the utterance sentence processing device 1. The memory 12 may store an intention estimation model 121, a search condition estimation model 122, a response sentence template 123, and a sample storage unit 124. Furthermore, the memory 12 and the sample storage unit 124 may be configured as separate entities.

[0079] The display 13 displays information such as data generated by the processing circuit 11 and data stored in the memory 12. For example, a cathode ray tube (CRT) display, a liquid crystal display (LCD), a plasma display, an organic electroluminescence display (OLED), a tablet terminal, or other displays can be used as the display 13. Incidentally, the display 13 may also display a response message.

[0080] The input interface 14 accepts input from a user who uses the utterance sentence processing device 1, converts the accepted input into an electrical signal, and outputs it to the processing circuit 11. The input interface 14 can be a physical operation part such as a mouse, keyboard, trackball, switch, button, joystick, touchpad, touch panel display, or microphone. The input interface 14 may also be a device that accepts input from an external input device separate from the utterance sentence processing device 1, converts the accepted input into an electrical signal, and outputs it to the processing circuit 11. The input interface 14 may also accept input regarding a predetermined threshold value, etc., from an engineer who manages the utterance sentence processing device 1.

[0081] The communication interface 15 transmits and receives data to and from an external device. Any communication standard can be used between the communication interface 15 and the external device. The communication interface 15 may also receive utterances from the user.

[0082] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, as well as within the scope of the invention and its equivalents as defined in the claims. [Explanation of symbols]

[0083] 1. Speech processing device, 11. Processing circuit, 12. Memory, 13. Display, 14. Input interface, 15. Communication interface, 111. Reception unit, 112. First estimation unit, 113. Second estimation unit, 114. Determination unit, 115. Response generation unit, 116. Output unit, 117. Storage unit, 121. Intention estimation model, 122. Search condition estimation model, 123. Response sentence template, 124. Sample storage unit

Claims

1. a reception unit that receives an utterance from a user; a first estimation unit that estimates an intention expressed by the user based on the utterance; a second estimation unit that estimates a search condition designated by the user based on the utterance; a determination unit that calculates, with respect to a first estimation result related to the estimated intention and a second estimation result related to the estimated search condition, a fourth determination result indicating that the utterance sentence is determined to be an invalid utterance in the first estimation result and that the utterance sentence is determined to be not an invalid utterance in the second estimation result, and calculates a fifth determination result indicating that the utterance sentence is determined to be not an invalid utterance in the first estimation result and that the utterance sentence is determined to be an invalid utterance in the second estimation result; a response generation unit that generates a response sentence corresponding to the fourth determination result, the response sentence prompting the user to confirm the estimated search condition, and generates a response sentence corresponding to the fifth determination result, the response sentence prompting the user to specify at least one search condition; An utterance sentence processing device comprising:

2. the determination unit calculates a first determination result indicating that the utterance sentence is determined to be not an invalid utterance in each of the first estimation result and the second estimation result, and a third determination result indicating that the utterance sentence is determined to be an invalid utterance in each of the first estimation result and the second estimation result. The speech processing device according to claim 1 .

3. the response generation unit generates response sentences corresponding to the first determination result, the fourth determination result, the fifth determination result, and the third determination result, respectively; further comprising an output unit that outputs the response sentence. The speech processing device according to claim 2 .

4. the output unit outputs the response sentence to a display device or an audio device. The speech processing device according to claim 3 .

5. the first estimation unit estimates, as the first estimation result, a first probability that the utterance sentence corresponds to an intention class corresponding to each of a plurality of candidates related to an intention; The speech processing device according to any one of claims 1 to 4.

6. When the intention class includes an intention class corresponding to an invalid utterance, the determination unit determines that the utterance sentence in the first estimation result is an invalid utterance when the intention class to which the highest first probability is assigned among the respective first probabilities is an intention class corresponding to the invalid utterance. The speech processing device according to claim 5 .

7. When the intention class does not include an intention class corresponding to an invalid utterance, the determination unit determines that the utterance sentence in the first estimation result is an invalid utterance when each of the first probabilities is equal to or smaller than a threshold. The speech processing device according to claim 5 .

8. The speech recognition system further includes a sample storage unit configured to store a plurality of samples in which intention labels corresponding to a plurality of candidates for intention are associated with features of utterances corresponding to the intention labels, the first estimation unit estimates, as the first estimation result, a similarity between a feature of a sentence uttered by the user and a feature of a sentence uttered in each of the plurality of samples; The speech processing device according to any one of claims 1 to 4.

9. The similarity is a cosine similarity between vectors representing features of utterances, the determination unit determines that the sentence uttered by the user in the first estimation result is an invalid utterance when each of the cosine similarities is equal to or smaller than a threshold. The speech processing device according to claim 8 .

10. a storage unit configured to store a feature quantity of the utterance sentence from the user in the sample storage unit according to the fourth determination result and the fifth determination result, The speech processing device according to claim 8 or 9.

11. the second estimation unit estimates, as the second estimation result, a second probability that the utterance sentence corresponds to each of a plurality of values ​​related to a search condition; the determination unit determines that the utterance sentence in the second estimation result is an invalid utterance when each of the second probabilities is equal to or smaller than a threshold. The speech processing device according to any one of claims 1 to 10.

12. A computer comprising: Accepts an utterance from the user, Inferring the intention expressed by the user based on the spoken sentence; Inferring search criteria designated by the user based on the spoken sentence; calculating a fourth determination result indicating that the utterance sentence is determined to be an invalid utterance in the first estimation result and that the utterance sentence is determined not to be an invalid utterance in the second estimation result, based on a first estimation result related to the estimated intention and a second estimation result related to the estimated search condition; calculating a fifth determination result indicating that the utterance sentence is determined not to be an invalid utterance in the first estimation result and that the utterance sentence is determined to be an invalid utterance in the second estimation result; generating a response sentence corresponding to the fourth determination result, the response sentence prompting the user to confirm the estimated search condition; generating a response sentence that prompts the user to specify at least one search condition as a response sentence corresponding to the fifth determination result; Spoken sentence processing method.

13. On the computer, a reception function for receiving an utterance from a user; a first estimation function for estimating an intention expressed by the user based on the utterance; a second estimation function for estimating a search condition designated by the user based on the spoken sentence; a determination function that calculates, with respect to a first estimation result related to the estimated intention and a second estimation result related to the estimated search condition, a fourth determination result indicating that the utterance sentence is determined to be an invalid utterance in the first estimation result and that the utterance sentence is determined to be not an invalid utterance in the second estimation result, and calculates a fifth determination result indicating that the utterance sentence is determined to be not an invalid utterance in the first estimation result and that the utterance sentence is determined to be an invalid utterance in the second estimation result; a response generation function that generates a response sentence corresponding to the fourth determination result, the response sentence prompting the user to confirm the estimated search conditions, and generates a response sentence corresponding to the fifth determination result, the response sentence prompting the user to specify at least one search condition; A speech processing program that achieves this.

Citation Information

Patent Citations

  • Information terminal

    JP2006317573A

  • Dialogue system, and method for updating dialogue flow and program

    JP2011215742A

  • Generation apparatus, generation method and generation program

    JP2018194902A

  • Dialogue system, dialogue method, and dialogue program

    JP2019086679A

  • Wireless terminal, management server, intention interpretation server, control method thereof, and program

    JP2019109752A