Dialog processing device, vehicle including dialog processing device, and dialog processing method
By introducing buffers and controllers into the dialogue processing device, the problem of inaccurate utterance end detection is solved, and the accuracy of the dialogue processing device in recognizing user intent and outputting responses is improved, as well as user convenience.
Patent Information
- Application Number
- CN202010373462.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-06
- Filing Date
- 2020-05-06
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2040-05-06
AI Technical Summary
Traditional dialogue processing devices struggle to accurately identify the end point of a user's speech when silent segments are detected, leading to a failure to correctly recognize the user's intent and outputting a mismatched response.
By introducing a buffer and controller into the dialogue processing device, the speech end time point is detected, the speech recognition result after the speech end time point is generated, and the accurate user intent recognition result is generated based on the speech signal in the buffer, and the corresponding response is output.
It improves the accuracy of dialogue processing devices in recognizing user intent in their speech, ensuring that the output response matches the user's intent and increasing user convenience.
Smart Images

Figure CN112347233B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The disclosure relates to a conversation processing device for recognizing an intention of a user through a conversation with the user and providing information or a service desired by the user, a vehicle including the same, and a conversation processing method thereof. BACKGROUND
[0002] The conversation processing device is an apparatus for having a conversation with a user. The conversation processing device operates to recognize utterance of the user, recognize an intention of the user through a result of recognizing the utterance, and output a response for providing a desired information or service to the user.
[0003] When recognizing the utterance of the user, it is necessary to determine a portion in which the user actually speaks a sentence, and for this, it is difficult to detect a point at which the user finishes the sentence. Such detection is called end point detection (EPD).
[0004] The conventional conversation processing device recognizes that the user has finished the sentence when a non-speech portion of greater than or equal to a predetermined time is detected. The conventional conversation processing device generates an utterance recognition result based on utterance data obtained until the sentence is finished. In this case, when the user does not speak for the predetermined time without an intention to finish the sentence, a next utterance input by the user is not considered in generating the utterance recognition result. Accordingly, a response that does not match the intention of the user is output. SUMMARY
[0005] Accordingly, an object of the disclosure is to provide a conversation processing device capable of receiving an utterance signal of a user and outputting a response corresponding to the utterance signal of the user, a vehicle including the same, and a conversation processing method thereof.
[0006] Other aspects of the disclosure will be apparent from the following description and from the description taken in conjunction with the accompanying drawings, part of which will be described in the following description.
[0007] Accordingly, an aspect of the disclosure is to provide a conversation processing device. The conversation processing device includes an utterance input device configured to receive an utterance signal of a user, a first buffer configured to store the received utterance signal therein, an output device, and a controller. The controller is configured to detect a sentence finish time point based on the stored utterance signal, generate a second utterance recognition result corresponding to the utterance signal after the sentence finish time point based on whether an intention of the user is recognized from a first utterance recognition result corresponding to the utterance signal before the sentence finish time point, and control the output device to output a response corresponding to the intention of the user determined based on at least one of the first utterance recognition result or the second utterance recognition result.
[0008] When the user's intention cannot be recognized from the first utterance recognition result, a second utterance recognition result corresponding to the utterance signal after the utterance end time point can be generated.
[0009] The dialogue processing device can further include a second buffer. When the user's intention cannot be recognized according to the first utterance recognition result, the controller can store the first utterance recognition result in the second buffer.
[0010] When the user's intention cannot be recognized from the first utterance recognition result, the controller can generate a second utterance recognition result based on the number of times of utterance recognition.
[0011] When the number of times of utterance recognition is less than a predetermined reference value, the controller can generate a second utterance recognition result based on the utterance signal after the utterance end time point.
[0012] When the number of times of utterance recognition is greater than or equal to the predetermined reference value, the controller can delete the data stored in the first buffer and generate a response corresponding to a case where the user's intention cannot be recognized.
[0013] When the response corresponding to the user's utterance signal is output, the controller can set the number of times of utterance recognition to an initial value.
[0014] When the second utterance recognition result is generated, the controller can determine an intention candidate group for determining the user's intention based on at least one of the first utterance recognition result or the second utterance recognition result. The controller can further determine one selected from the determined intention candidate group as the user's intention.
[0015] The controller can determine the accuracy of the intention candidate group, and determine an intention candidate having the highest accuracy among the intention candidate group as the user's intention.
[0016] When the utterance end time point is detected, the controller can be configured to delete the data stored in the first buffer and store the utterance signal input after the utterance end time point in the first buffer.
[0017] When the user's intention can be recognized from the first utterance recognition result, the controller can delete the data stored in the first buffer.
[0018] Another aspect of the disclosure is to provide a vehicle including a speech input device configured to receive a user's speech signal, a first buffer configured to store the received speech signal therein, an output device, and a controller. The controller is configured to detect an utterance end time point based on the stored speech signal, generate a second speech recognition result corresponding to the speech signal after the utterance end time point based on whether an intention of the user is recognized from a first speech recognition result corresponding to the speech signal before the utterance end time point, and control the output device to output a response corresponding to the intention of the user determined based on at least one of the first speech recognition result or the second speech recognition result.
[0019] The controller can generate the second speech recognition result corresponding to the speech signal after the utterance end time point when the intention of the user is not recognized from the first speech recognition result.
[0020] When the second speech recognition result is generated, the controller can determine an intention candidate group for determining the intention of the user based on at least one of the first speech recognition result or the second speech recognition result. The controller can further determine one selected from the determined intention candidate group as the intention of the user.
[0021] Another aspect of the disclosure is to provide a dialogue processing method. The processing method includes receiving a user's speech signal, storing the received speech signal therein, detecting an utterance end time point based on the stored speech signal, generating a second speech recognition result corresponding to the speech signal after the utterance end time point based on whether an intention of the user is recognized from a first speech recognition result corresponding to the speech signal before the utterance end time point, and outputting a response corresponding to the intention of the user determined based on at least one of the first speech recognition result or the second speech recognition result.
[0022] Generating the second speech recognition result corresponding to the speech signal after the utterance end time point can include generating the second speech recognition result corresponding to the speech signal after the utterance end time point when the intention of the user is not recognized from the first speech recognition result.
[0023] Generating the second speech recognition result corresponding to the speech signal after the utterance end time point can include storing the first speech recognition result in a second buffer when the intention of the user is not recognized from the first speech recognition result.
[0024] Generating the second speech recognition result corresponding to the speech signal after the utterance end time point can include generating the second speech recognition result based on a number of times of speech recognition when the intention of the user is not recognized from the first speech recognition result.
[0025] Generating the second speech recognition result corresponding to the speech signal after the end time point of the utterance can include generating the second speech recognition result based on the speech signal after the end time point of the utterance when the number of times of speech recognition is less than a predetermined reference value.
[0026] The dialogue processing method can further include, when the second speech recognition result is generated, determining an intent candidate group for determining an intent of the user based on at least one of the first speech recognition result or the second speech recognition result, and determining one selected from the determined intent candidate group as the intent of the user. BRIEF DESCRIPTION OF DRAWINGS
[0027] These and / or other aspects of the disclosure will become apparent and more readily appreciated from the following description of the embodiments, taken in conjunction with the accompanying drawings in which:
[0028] Figure 1 is a control block diagram illustrating a dialogue processing apparatus according to an embodiment.
[0029] Figure 2 is a diagram for describing an operation of a dialogue processing apparatus according to an embodiment.
[0030] Figure 3 is a diagram for describing an operation of a dialogue processing apparatus according to an embodiment.
[0031] Figure 4 is a flowchart illustrating a dialogue processing method according to an embodiment.
[0032] Figure 5A and Figure 5B is a flowchart illustrating a dialogue processing method according to another embodiment. DETAILED DESCRIPTION
[0033] Throughout the specification, the same numbers refer to the same elements throughout. Not all elements of the embodiments of the disclosure are described. Descriptions overlapping with each other in the embodiments are omitted. The terms used throughout the specification, such as "part", "module", "member", "block", etc., can be implemented in software and / or hardware, and a plurality of "parts", "modules", "members", or "blocks" can be implemented in a single element, or a single "part", "module", "member", or "block" can include a plurality of elements.
[0034] It should be further understood that the term "connected" or its derivatives refer both to direct and indirect connections. Indirect connections include connections through wireless communication networks.
[0035] It should be further understood that the terms "comprises" and / or "comprising", when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0036] Although the terms "first", "second", "A", "B", etc. can be used in describing various components, these terms do not limit the corresponding components, but are used for the purpose of distinguishing one component from another only.
[0037] As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0038] Reference numerals used for the steps of the method are only for the convenience of explanation, not to limit the order of the steps. Therefore, unless the context clearly indicates otherwise, the written command can be performed in writing.
[0039] Hereinafter, the operating principle and embodiments of the present disclosure are described below with reference to the accompanying drawings.
[0040] Figure 1 is a control block diagram illustrating a dialogue processing apparatus according to an embodiment.
[0041] Referring to Figure 1 , the dialogue processing apparatus 100 according to an embodiment includes a speech input device 110, a communicator 120, a controller 130, an output device 140, and a storage 150.
[0042] The speech input device 110 can receive a command in the form of speech of a user. In other words, the speech input device 110 can receive a speech signal of a user. To this end, the speech input device 110 can include a microphone that receives a sound and converts the sound into an electrical signal.
[0043] The storage 150 can store various types of data used directly or indirectly by the dialogue processing apparatus 100 to output a response corresponding to the speech of the user.
[0044] In addition, the storage 150 can include a first buffer 151 and a second buffer 152. The first buffer 151 can store an input speech signal, and the second buffer 152 can store a result of speech recognition.
[0045] The memory 150 can be implemented using at least one of a nonvolatile storage device such as a cache, a read only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), and a flash memory, a volatile storage device such as a random access memory (RAM), or another storage medium such as a hard disk drive (HDD), a CD-ROM, etc., but the implementation of the memory 150 is not limited thereto. The memory 150 can be a memory implemented as a chip separate from a processor (described below in connection with the controller 130), or a single chip integrated with the processor.
[0046] The controller 130 includes an input processor 131 configured to recognize an input speech signal to generate a speech recognition result, and a dialogue manager 132 configured to recognize an intention of a user based on the speech recognition result and determine an action corresponding to the intention of the user, and a result processor 133 configured to generate a dialogue response for performing the determined action.
[0047] The input processor 131 can recognize an input speech signal of a user, and can convert the speech signal of the user into a text type utterance. The input processor 131 can apply a natural language understanding algorithm to the utterance text to recognize an intention of a user.
[0048] At least one of the utterance text or the intention of the user can be output by the input processor 131 as a speech recognition result. The input processor 131 can transmit the speech recognition result to the dialogue manager 132.
[0049] To this end, the input processor 131 can include a speech recognition module, and can be implemented using a processor (not shown) that performs an operation for processing an input speech.
[0050] The speech processing operation of the input processor 131 can be performed when an input speech is input, or can be performed when an input speech is input and a specific condition is satisfied.
[0051] In detail, the input processor 131 can perform the above-described speech processing operation when a predetermined call word is input or when a speech recognition start command is received from a user.
[0052] Further, the input processor 131 can generate a speech recognition result corresponding to a speech signal input for a specific portion. In detail, the input processor 131 can generate a speech recognition result corresponding to a speech signal from an utterance start time point to an utterance end time point.
[0053] To this end, the input processor 131 can detect a speech start time point or a speech end time point based on the input speech signal. The input processor 131 can determine a time point at which the input predetermined call word or a speech recognition start command is received from the user as the speech start time point. In this case, the user can input the speech recognition start command by saying the predetermined call word or through a separate button, and the input processor 131 can recognize the user's speech from the speech start time point.
[0054] In addition, when there is a silent portion greater than or equal to a predetermined time in the input speech signal, the input processor 131 can recognize that the user's speech is terminated. In this case, the input processor 131 can determine a time point at which a predetermined time elapses from a start time of the silent portion as the speech end time point.
[0055] On the other hand, when the user stops the speech for a short time without intending to terminate the speech, the mute portion during which the user stops the speech for a short time can cause the user's speech to be mistakenly recognized as having been terminated. Therefore, the speech signal of the user after the speech is stopped can not be considered when generating a speech recognition result. In this case, when it is difficult to recognize the user's intention based on only the speech signal before the user stops the speech, the user's intention can not be accurately recognized. Therefore, it is difficult to output a dialog response suitable for the user's situation.
[0056] To this end, the input processor 131 can determine whether to recognize the speech signal after the speech end time point based on whether the user's intention corresponding to the speech signal before the speech end time point can be recognized. Details thereof are described below.
[0057] On the other hand, in addition to the above-described operations, the input processor 131 can perform the following operations: removing noise of the input speech signal and recognizing the user corresponding to the input speech signal.
[0058] The dialog manager 132 can determine the user's intention based on the speech recognition result received from the input processor 131 and determine an action corresponding to the user's intention.
[0059] The result processor 133 can provide a specific service or output a system speech to continue the dialog according to the output result of the dialog manager 132. The result processor 133 can generate a dialog response and a command required to perform the received action and can output the command. The dialog response can be output as text, an image, or audio. When the command is output, a service corresponding to the output command, such as vehicle control and external content provision, can be performed.
[0060] The control unit 130 can include a memory (not shown) for storing data regarding an algorithm or a program representing the algorithm for controlling the operation of the components of the dialogue processing device 100. The control unit 130 can also include a processor (not shown) that uses the data stored in the memory to perform the above-described operations. In this case, the memory and the processor can be implemented as separate chips. Alternatively, the memory and the processor can be implemented as one chip.
[0061] Alternatively, the input processor 131, the dialogue manager 132, and the result processor 133 can be integrated into one processor or can be implemented as separate processors.
[0062] The output device 140 can output a response generated by the controller 130 visually or aurally. To this end, the output device 140 can include a display (not shown) or a speaker (not shown). The display (not shown) and the speaker (not shown) can not only output a response to the user's utterance, a query to the user, or information to be provided to the user, but also output a confirmation of the user's intention and respond to the user's utterance in a visual or aural manner at the other end.
[0063] The communicator 120 can communicate with an external device such as a server. To this end, the communicator 120 can include one or more components enabling communication with the external device, for example, at least one of a short distance communication module such as a Bluetooth module, an infrared communication module, and a radio frequency identification (RFID) communication module; a wired communication module such as a controller area network (CAN) communication module and a local area network (LAN) module; and a wireless communication module such as a Wi-Fi module and a wireless broadband module.
[0064] When the communicator 120 includes a wireless communication module, the wireless communication module can include a wireless communication interface including an antenna and a transmitter for transmitting a signal. In addition, the wireless communication module can further include a signal conversion module for converting a digital control signal output from the dialogue processing device 100 through the wireless communication interface into an analog type wireless signal under the control of the dialogue processing device 100.
[0065] The wireless communication module can include a wireless communication interface including an antenna and a receiver for receiving a signal. In addition, the wireless communication module can further include a signal conversion module for demodulating an analog type wireless signal received through the wireless communication interface into a digital control signal.
[0066] At least one component can be added or omitted to correspond to Figure 1The performance of the components of the illustrated dialog processing device 100. Furthermore, the mutual positions of the components can be changed to correspond to the performance or structure of the system.
[0067] Figure 1 Each of the illustrated components can refer to a software component and / or a hardware component, such as a field programmable gate array (FPGA) and an application specific integrated circuit (ASIC).
[0068] Figure 1 All or some of the components of the illustrated dialog processing device 100 can be included in a vehicle, can recognize speech of a user including a driver and a passenger of the vehicle, and can provide an appropriate response.
[0069] Figure 2 is a diagram for describing the operation of a dialog processing device according to an embodiment.
[0070] Reference Figure 2 According to an embodiment, the dialog processing device 100 can receive a speech signal Sin including the utterance "Santa Fe... let me know fuel efficiency." In this case, it is assumed that the user stops the utterance for a short time after "Santa Fe" and then utters "let me know fuel efficiency."
[0071] The input processor 131 can determine a time point t1 at which the user utters the predetermined call word or inputs the speech recognition start command as the utterance start time point. The input processor 131 can also store the speech signal input after the utterance start time point in the first buffer 151.
[0072] The input processor 131 can detect the utterance end time point based on the speech signal stored in the first buffer 151.
[0073] In detail, when there is a silent portion TS1 corresponding to a time point t2 at which the user stops inputting the speech to a time point t3 at which a predetermined time has elapsed after the time point t2 in the stored speech signal, the input processor 131 can determine the time point at which the silent portion TS1 ends as the utterance end time point.
[0074] When the utterance end time point is detected, the input processor 131 can recognize the speech signal input in a portion (hereinafter, referred to as a first portion A) before the utterance end time point. The input processor 131 can generate the utterance text X[i]: "Santa Fe" as a speech recognition result corresponding to the speech signal.
[0075] At this time, the input processor 131 can initialize the first buffer 151 by deleting the speech signal of the first portion A stored in the first buffer 151. Furthermore, the input processor 131 can store the speech signal input after the utterance end time point of the first portion A in the first buffer 151.
[0076] In addition, the input processor 131 can generate a speech recognition result corresponding to the speech signal input after the utterance end time point, according to whether the user's intention can be recognized based on at least one of the speech recognition result of the first portion A or the speech recognition result of the second portion B.
[0077] In detail, when the user's intention cannot be recognized using only the speech signal input in the first portion A, the input processor 131 can generate a speech recognition result corresponding to the speech signal input in a portion (hereinafter referred to as a second portion B) after the utterance end time point.
[0078] In this case, the input processor 131 can store the speech recognition result of the first portion A in the second buffer 152.
[0079] Thereafter, when the utterance end time point of the second portion B is detected by the presence of another silent portion TS2, the input processor 131 can recognize the speech signal input in the second portion B stored in the first buffer 151. The input processor 131 can generate the utterance text X[i+1]: "let me know fuel efficiency" as a speech recognition result corresponding to the speech signal.
[0080] At this time, the input processor 131 can initialize the first buffer 151 by deleting the speech signal of the second portion B stored in the first buffer 151. In this case, the input processor 131 can store the speech signal input after the utterance end time point of the second portion B in the first buffer 151.
[0081] In addition, the input processor 131 can generate a speech recognition result corresponding to the speech signal input after the utterance end time point, according to whether the user's intention can be recognized based on at least one of the speech recognition result of the first portion A or the speech recognition result of the second portion B.
[0082] In this case, when the user's intention cannot be recognized using only the speech recognition result corresponding to the second portion B, the utterance text X[i+1]: "let me know fuel efficiency", the input processor 131 can combine the utterance text X[i]: "Santa Fe" as the speech recognition result of the first portion A with the utterance text X[i+1]: "let me know fuel efficiency" as the speech recognition result of the second portion B. The input processor 131 can further recognize the user's intention based on the combined speech recognition result.
[0083] When the user's intention cannot be recognized even in combination with the speech recognition result of the first part A and the speech recognition result of the second part B, the input processor 131 can generate a speech recognition result corresponding to the speech signal input in the part after the utterance end time point of the second part B. Thereafter, the above-mentioned subsequent operations can be repeated.
[0084] On the other hand, when the user's intention is recognizable based on the speech recognition result of at least one part, the input processor 131 can transmit the determined user's intention to the dialogue manager 132. The dialogue manager 132 can determine an action corresponding to the user's intention. The result processor 133 can generate a dialogue response for performing the action received from the dialogue manager 132. The response of the result processor 133 can be output through the output device 140.
[0085] In addition, when the user's intention is recognizable based on the speech recognition result of at least one part, the input processor 131 can block the input of the speech signal by switching the speech input device 110 to an off state. In addition, the input processor 131 can initialize at least one of the first buffer 151 or the second buffer 152 by deleting data stored in at least one of the first buffer 151 or the second buffer 152.
[0086] The input processor 131 can perform a processing operation according to whether the user's intention is recognizable for the speech signal after the utterance end time point. The input processor 131 can not only use the speech recognition result after the utterance end time point, but also use the speech recognition result before the utterance end time point as a basis for controlling the operation of recognizing the user's intention. Accordingly, the user's intention is accurately recognized, a response suitable for the user is output, and the convenience of the user is increased.
[0087] Figure 3 is a diagram for describing the operation of a dialogue processing device according to another embodiment.
[0088] Reference Figure 3 , according to another embodiment, the dialogue processing device 100 can receive a speech signal Sin' including the utterance "Santa Fe... no, Sonata, let me know the fuel efficiency." In this case, it is assumed that the user stops the utterance for a short time after "Santa Fe" and then says "no, Sonata, let me know the fuel efficiency."
[0089] As described above with reference to Figure 2 , the input processor 131 can determine the time point t1' at which the user says the predetermined call word or inputs the speech recognition start command as the utterance start time point. The input processor 131 can store the speech signal input after the utterance start time point in the first buffer 151.
[0090] The input processor 131 can detect an utterance end time point based on the speech signal stored in the first buffer 151. Specifically, when there is a silent portion TS3 from a time point t2' at which the user stops inputting speech to a time point t3' at which a predetermined time has elapsed after the time point t2' in the stored speech signal, the input processor 131 can determine the time point at which the silent portion TS3 ends as the utterance end time point.
[0091] When the utterance end time point is detected, the input processor 131 can recognize the speech signal input in a portion (hereinafter, referred to as a first portion C) before the utterance end time point. The input processor 131 can generate an utterance text X[i']:"Santa Fe" as a speech recognition result corresponding to the speech signal.
[0092] At this time, the input processor 131 can initialize the first buffer 151 by deleting the speech signal of the first portion C stored in the first buffer 151. In addition, the input processor 131 can store the speech signal input after the utterance end time point of the first portion C in the first buffer 151.
[0093] In addition, the input processor 131 can generate a speech recognition result corresponding to the speech signal after the utterance end time point based on whether the user's intention can be recognized from the speech recognition result corresponding to the speech signal input in the first portion C.
[0094] In detail, when the user's intention cannot be recognized using only the speech signal input in the first portion C, the input processor 131 can generate a speech recognition result corresponding to the speech signal input in a portion (hereinafter, referred to as a second portion D) after the utterance end time point.
[0095] In this case, the input processor 131 can store the speech recognition result of the first portion C in the second buffer 152.
[0096] Thereafter, when the utterance end time point of the second portion D is detected through the presence of another silent portion TS4, the input processor 131 can recognize the speech signal input in the second portion D stored in the first buffer 151. The input processor 131 can generate an utterance text X[i'+1]:"No, Sonata, let me know the fuel efficiency" as a speech recognition result corresponding to the speech signal.
[0097] At this time, the input processor 131 can initialize the first buffer 151 by deleting the speech signal of the second portion D stored in the first buffer 151. In addition, the input processor 131 can store the speech signal input after the utterance end time point of the second portion D in the first buffer 151.
[0098] In addition, the input processor 131 can generate a speech recognition result corresponding to the speech signal after the utterance end time point, according to whether the user's intention can be recognized based on at least one of the speech recognition result of the first part C or the speech recognition result of the second part D.
[0099] In this case, when the user's intention can be recognized using only the utterance text X[i'+1]: "No, Sonata, let me know the fuel efficiency" which is the speech recognition result corresponding to the second part D, the input processor 131 can transmit the user's intention to the dialogue manager 132. When the user's intention is transmitted to the result processor 133 via the dialogue manager 132, the result processor 133 generates a dialogue response.
[0100] Alternatively, the input processor 131 can determine an intention candidate group for determining the user's intention based on at least one of the utterance text X[i'+1]: "No, Sonata, let me know the fuel efficiency" which is the speech recognition result of the second part D or the utterance text X[i']: "Santa Fe" which is the speech recognition result of the first part C. The input processor 131 can determine one selected from the intention candidate group as the user's intention.
[0101] In detail, the input processor 131 can determine the accuracy with respect to the intention candidate group, and can determine an intention candidate having the highest accuracy among the intention candidate group as the user's intention. In this case, the accuracy of the intention candidate group can be calculated as a probability value. The input processor 131 can determine an intention candidate having the highest probability value as the user's intention.
[0102] For example, the input processor 131 can determine a first intention candidate based on the utterance text X[i'+1]: "No, Sonata, let me know the fuel efficiency" which is the speech recognition result of the second part D. In addition, the input processor 131 can determine a second intention candidate based on a result value of combining the speech recognition results of the first part C and the second part D: "Santa Fe, No, Sonata, let me know the fuel efficiency". The input processor 131 can determine an intention candidate having the highest accuracy among the first intention candidate and the second intention candidate as the user's intention.
[0103] Thereafter, the input processor 131 can transmit the determined user's intention to the dialogue manager 132. When the user's intention is transmitted to the result processor 133 through the dialogue manager 132, the result processor 133 can generate a dialogue response.
[0104] Also, when determining the user's intention, the input processor 131 can block input of the speech signal by switching the speech input device 110 to an off state. Also, the input processor 131 can initialize at least one of the first buffer 151 or the second buffer 152 by deleting data stored in the at least one.
[0105] The input processor 131 can determine an intention candidate set of the user by combining speech recognition results of at least one speech recognition portion divided based on the utterance end time point. The input processor 131 can further determine a final intention of the user based on accuracy of the intention candidate set. Accordingly, the user's intention is accurately recognized, a response suitable for the user is output, and the user's convenience is increased.
[0106] Figure 4 is a flowchart illustrating a dialogue processing method according to an embodiment.
[0107] Reference Figure 4 The dialogue processing device 100 can identify whether a call command is recognized (401). In this case, the call command can be set as a predetermined call word, and the user can issue the call command by saying the predetermined call word or inputting a speech recognition start command (for example, by manipulating a button).
[0108] When the call command is recognized (YES in operation 401), the dialogue processing device 100 can store the input speech signal in the first buffer 151 (402). In this case, the dialogue processing device 100 can store the speech signal input in real time in the first buffer 151.
[0109] The dialogue processing device 100 can generate a speech recognition result based on the speech signal stored in the first buffer (403).
[0110] The dialogue processing device 100 can identify whether the utterance end time point is detected (404). In detail, when there is a silent portion corresponding to a time point at which the user stops inputting speech to a time point at which a predetermined time has elapsed in the stored speech signal, the dialogue processing device 100 can determine the time point at which the silent portion ends as the utterance end time point.
[0111] When the utterance end time point is detected (YES in operation 404), the dialogue processing device 100 can initialize the first buffer 151 by deleting the speech signal (the speech signal input in the nth speech recognition portion) stored in the first buffer 151 (405). The dialogue processing device 100 can store the speech signal (the speech signal input in the n+1th speech recognition portion) input after the utterance end time point in the first buffer 151 (406).
[0112] Thereafter, the dialogue processing device 100 can identify whether the user's intention is identifiable using the utterance recognition result generated based on the utterance signal before the utterance end time point (407). In this case, the utterance recognition result can represent the utterance recognition result generated in operation 403, i.e., the utterance recognition result corresponding to the utterance signal input in the nthutterance recognition section.
[0113] When the user's intention is not identifiable (NO in operation 407), the dialogue processing device 100 can store the generated utterance recognition result in the second buffer 152 (410). In other words, the dialogue processing device 100 can store the utterance recognition result generated in operation 403 (the utterance recognition result corresponding to the utterance signal input in the nthutterance recognition section) in the second buffer 152. Thereafter, the dialogue processing device 100 can generate an utterance recognition result based on the utterance signal stored in the first buffer 151 (the utterance signal input in the (n+1)thutterance recognition section) (403). In this case, the utterance signal stored in the first buffer 151 can represent the utterance signal input after the utterance end time point (i.e., the utterance signal input in the (n+1)thutterance recognition section).
[0114] Thereafter, the dialogue processing device 100 can identify whether the utterance end time point of the (n+1)thutterance recognition section is detected (404). When the utterance end time point is detected (YES in operation 404), the dialogue processing device 100 can perform operations 405 and 406 as described above. Thereafter, the dialogue processing device 100 can check whether the user's intention is identifiable based on at least one of the utterance recognition result of the (n+1)thutterance recognition section or the utterance recognition result of the nthutterance recognition section stored in the second buffer 152 (407). Thereafter, the above-described subsequent processing can be repeated.
[0115] As another example, when the user's intention is identifiable (YES in operation 407), the dialogue processing device 100 can block the input of the utterance signal and initialize the first buffer 151 (408). In detail, the dialogue processing device 100 can block the input of the utterance signal by switching the utterance input device 110 to an off state, and can initialize the first buffer 151 by deleting the data stored in the first buffer 151. In this case, the dialogue processing device 100 can initialize the second buffer 152 by deleting the data stored in the second buffer 152.
[0116] The dialogue processing device 100 can generate and output a response corresponding to the user's intention (409).
[0117] The input processor 131 can perform a processing operation on the utterance signal after the utterance end time point according to whether the user's intention is identifiable. The input processor 131 can use not only the utterance recognition result after the utterance end time point but also the utterance recognition result before the utterance end time point as a basis for controlling the operation of recognizing the user's intention. Accordingly, the user's intention is accurately recognized, a response suitable for the user is output, and the user's convenience is increased.
[0118] Referring to Figure 4 Operation 404 is performed after operation 403, but operation 403 and operation 404 can be simultaneously performed, and operation 403 can be performed after operation 404. However, operation 403 is performed after operation 404 means that when the utterance end time point is detected (Yes in operation 404), operation 403 is performed to generate an utterance recognition result based on the utterance signal stored in the first buffer 151.
[0119] Figure 5A And 5B is a flowchart illustrating a conversation processing method according to another embodiment.
[0120] Referring to Figure 5A And Figure 5B The conversation processing device 100 according to an embodiment can identify whether a call command is recognized (501). When the call command is recognized (Yes in operation 501), the conversation processing device 100 can store the input utterance signal in the first buffer 151 (502). In this case, the conversation processing device 100 can store the utterance signal input in real time in the first buffer 151.
[0121] The conversation processing device 100 can generate an utterance recognition result based on the utterance signal stored in the first buffer (503).
[0122] The conversation processing device 100 can identify whether the utterance end time point is detected (504).
[0123] When the utterance end time point is detected (Yes in operation 504), the conversation processing device 100 can initialize the first buffer 151 by deleting the utterance signal (the utterance signal input in the nth utterance recognition section) stored in the first buffer 151 (505). The conversation processing device 100 can store the utterance signal (the utterance signal input in the (n+1)th utterance recognition section) input after the utterance end time point in the first buffer 151 (506).
[0124] Thereafter, the dialogue processing device 100 can check whether the user's intention can be recognized using the utterance recognition result generated based on the utterance signal before the utterance end time point (507). In this case, the utterance recognition result can represent the utterance recognition result generated in operation 503, i.e., the utterance recognition result corresponding to the utterance signal input in the nthutterance recognition section.
[0125] When the user's intention cannot be recognized (NO in operation 507), the dialogue processing device 100 can identify whether a result value exists in the utterance recognition result (510).
[0126] In this case, when the utterance text exists in the utterance recognition result, the dialogue processing device 100 can confirm that the result value exists. In other words, when the utterance text is not generated in the utterance recognition result, for example, because the user does not utter the utterance, the dialogue processing device 100 can determine that the result value does not exist in the utterance recognition result.
[0127] When the result value exists in the utterance recognition result (YES in operation 510), the dialogue processing device 100 can store the generated utterance recognition result in the second buffer 152 (511). In other words, the dialogue processing device 100 can store the utterance recognition result generated in operation 503 (the utterance recognition result corresponding to the utterance signal input in the nthutterance recognition section) in the second buffer 152. Thereafter, the dialogue processing device 100 can generate the utterance recognition result based on the utterance signal stored in the first buffer 151 (the utterance signal input in the nthutterance recognition section) (503). In this case, the utterance signal stored in the first buffer 151 can represent the utterance signal input after the utterance end time point (i.e., the utterance signal input in the nthutterance recognition section).
[0128] Thereafter, the dialogue processing device 100 can identify whether the utterance end time point of the nthutterance recognition section is detected (504). When the utterance end time point is detected (YES in operation 504), the dialogue processing device 100 can perform operations 505 and 506 as described above. Thereafter, the dialogue processing device 100 can check whether the user's intention can be recognized based on at least one of the utterance recognition result of the nthutterance recognition section or the utterance recognition result of the nthutterance recognition section stored in the second buffer 152 (507). Thereafter, the above-described subsequent processes can be repeated.
[0129] In another example, when there is no result value in the speech recognition result (NO in operation 510), the dialog processing device 100 can determine whether the number of speech recognition is greater than or equal to a reference value (512). In this case, the number of speech recognition can represent the number of times the speech recognition result is generated. In addition, the reference value for the number of speech recognition can represent the maximum number of speech recognition obtained in consideration of the storage capacity of the memory 150.
[0130] When the number of speech recognition is less than the reference value (NO in operation 512), the dialog processing device 100 can generate a speech recognition result based on the speech signal stored in the first buffer 151 (503). In this case, the speech signal stored in the first buffer 151 can represent a speech signal input after the utterance end time point (i.e., a speech signal input in the n+1 speech recognition section). Thereafter, the above-described subsequent process can be repeated.
[0131] In another example, when the number of speech recognition is equal to or greater than the reference value (YES in operation 512), or when the user's intent is recognizable (YES in operation 507), the dialog processing device 100 can block the input of the speech signal and initialize the first buffer 151 (508). In detail, the dialog processing device 100 can block the input of the speech signal by switching the speech input device 110 to an off state, and can initialize the first buffer 151 by deleting data stored in the first buffer 151. In this case, the dialog processing device 100 can initialize the second buffer 152 by deleting data stored in the second buffer 152.
[0132] The dialog processing device 100 can generate and output a response corresponding to the user's intent (509). In this case, the dialog processing device 100 can set the number of speech recognition to an initial value. When a call command is recognized after the number of speech recognition is set to the initial value, the dialog processing device 100 can generate a speech recognition result corresponding to the first speech recognition section.
[0133] The input processor 131 can perform a processing operation on the speech signal after the utterance end time point according to whether the user's intent is recognizable. The input processor 131 can use not only the speech recognition result after the utterance end time point but also the speech recognition result before the utterance end time point as a basis for controlling the operation of recognizing the user's intent. Accordingly, the user's intent is accurately recognized, a response suitable for the user is output, and the user's convenience is increased.
[0134] In addition, since the speech recognition result for the speech input by the user is generated according to the number of speech recognition, efficient speech recognition can be performed in consideration of the storage capacity.
[0135] Referring to Figure 5A , operation 504 is performed after operation 503, but operation 503 and operation 504 can be performed simultaneously, and operation 503 can also be performed after operation 504. However, operation 503 is performed after operation 504 indicates that when the end-of-utterance time point is detected (Yes in operation 504), operation 503 is performed to generate an utterance recognition result based on the utterance signal stored in the first buffer 151.
[0136] The disclosed embodiments can be embodied in the form of a recording medium storing instructions executable by a computer. The instructions can be stored in the form of program codes, and when executed by a processor, can generate program modules to perform the operations of the disclosed embodiments. The recording medium can be embodied as a computer-readable recording medium.
[0137] The computer-readable recording medium includes various recording media in which instructions decodable by a computer are stored, such as read-only memory (ROM), random access memory (RAM), magnetic tape magnetic disk, flash memory, optical data storage device, etc.
[0138] As is apparent from the above, the conversation processing device, the vehicle including the same, and the conversation processing method thereof can improve the accuracy of recognizing a user's utterance and processing a conversation. By accurately recognizing a user's intention, the convenience of the user can be improved.
[0139] Although the embodiments of the disclosure have been described for illustrative purposes, it will be understood by those of ordinary skill in the art that various modifications, additions and substitutions can be made without departing from the scope and spirit of the disclosure. Therefore, the embodiments of the disclosure are not described for the purpose of limitation.
Claims
1. A dialog processing device comprising: a speech input device configured to receive a speech signal of a user; a first buffer configured to store the received speech signal in the first buffer; an output device; and a controller configured to detect a speech end time point based on the stored speech signal, generate a second speech recognition result corresponding to a speech signal after the speech end time point based on whether an intention of the user is recognized from a first speech recognition result corresponding to a speech signal before the speech end time point, and control the output device to output a response corresponding to an intention of the user determined based on a combination of the first speech recognition result and the second speech recognition result, wherein the second speech recognition result corresponding to the speech signal after the speech end time point is generated when the intention of the user is not recognized from the first speech recognition result. 2.The dialog processing device of claim 1, further comprising a second buffer, wherein the controller stores the first speech recognition result in the second buffer when the intention of the user is not recognized from the first speech recognition result. wherein wherein the controller generates the second speech recognition result based on a number of speech recognitions when the intention of the user is not recognized from the first speech recognition result.
3. The dialog processing device according to claim 1, wherein wherein the controller generates the second speech recognition result based on the speech signal after the speech end time point when the number of speech recognitions is less than a predetermined reference value.
4. The dialog processing device according to claim 3, wherein wherein the controller deletes data stored in the first buffer and generates a response corresponding to a case where the intention of the user is not recognized when the number of speech recognitions is greater than or equal to the predetermined reference value.
5. The dialog processing device according to claim 3, wherein wherein the controller sets the number of speech recognitions to an initial value when a response corresponding to the speech signal of the user is output.
6. The dialog processing device according to claim 1, wherein wherein the controller determines an intention candidate group for determining the intention of the user based on a combination of the first speech recognition result and the second speech recognition result when the second speech recognition result is generated, and determines one intention candidate selected from the determined intention candidate group as the intention of the user.
7. The dialog processing device according to claim 1, wherein wherein the controller determines accuracy of the intention candidate group, and determines an intention candidate having the highest accuracy among the intention candidate group as the intention of the user.
8. The dialog processing device according to claim 7, wherein wherein the controller is configured to delete data stored in the first buffer and store a speech signal input after the speech end time point in the first buffer when the speech end time point is detected.
9. The dialog processing device according to claim 1, wherein wherein the controller deletes data stored in the first buffer when the intention of the user is recognized from the first speech recognition result.
10. The dialog processing device according to claim 1, wherein 11.A vehicle comprising: a speech input device configured to receive a speech signal of a user; a first buffer configured to store the received speech signal in the first buffer; an output device; and a controller configured to detect a speech end time point based on the stored speech signal, generate a second speech recognition result corresponding to a speech signal after the speech end time point based on whether an intention of the user is recognized from a first speech recognition result corresponding to a speech signal before the speech end time point, and control the output device to output a response corresponding to an intention of the user determined based on a combination of the first speech recognition result and the second speech recognition result, wherein the second speech recognition result corresponding to the speech signal after the speech end time point is generated when the intention of the user is not recognized from the first speech recognition result. a controller configured to detect a speech end time point based on the stored speech signal, generate a second speech recognition result corresponding to a speech signal after the speech end time point based on whether an intention of the user is recognized from a first speech recognition result corresponding to a speech signal before the speech end time point, and the controller configured to control the output device to output a response corresponding to the intention of the user determined based on a combination of the first speech recognition result and the second speech recognition result, wherein, when the intention of the user is not recognized from the first speech recognition result, the controller generates the second speech recognition result corresponding to the speech signal after the speech end time point.
12. The vehicle of claim 11, wherein, When the second speech recognition result is generated, the controller determines an intention candidate group for determining the intention of the user based on a combination of the first speech recognition result and the second speech recognition result, and determines one intention candidate selected from the determined intention candidate group as the intention of the user.
13. A dialogue processing method, comprising: receiving a speech signal of a user; storing the received speech signal in a first buffer; detecting a speech end time point based on the stored speech signal; generating a second speech recognition result corresponding to a speech signal after the speech end time point based on whether an intention of the user is recognized from a first speech recognition result corresponding to a speech signal before the speech end time point; and outputting a response corresponding to the intention of the user determined based on a combination of the first speech recognition result and the second speech recognition result, wherein generating the second speech recognition result corresponding to a speech signal after the speech end time point includes: when the intention of the user is not recognized from the first speech recognition result, generating the second speech recognition result corresponding to the speech signal after the speech end time point.
14. The dialog processing method of claim 13, wherein, Generating the second speech recognition result corresponding to a speech signal after the speech end time point includes: when the intention of the user is not recognized from the first speech recognition result, storing the first speech recognition result in a second buffer.
15. The dialog processing method of claim 13, wherein, Generating the second speech recognition result corresponding to a speech signal after the speech end time point includes: when the intention of the user is not recognized from the first speech recognition result, generating the second speech recognition result based on a number of times of speech recognition.
16. The dialog processing method of claim 15, wherein, Generating the second speech recognition result corresponding to a speech signal after the speech end time point includes: when the number of times of speech recognition is less than a predetermined reference value, generating the second speech recognition result based on the speech signal after the speech end time point.
17. The dialog processing method of claim 13, further comprising: When the second speech recognition result is generated, an intention candidate group for determining the intention of the user is determined based on a combination of the first speech recognition result and the second speech recognition result, and one intention candidate selected from the determined intention candidate group is determined as the intention of the user.
Citation Information
Patent Citations
Voice identification method and device
CN109360551A
Speech recognition method and apparatus
US20180173494A1