Information processing device, information processing method, and information processing program
The information processing system addresses ambiguous user intents through a three-level classification model, ensuring accurate responses by distinguishing between clear tasks, casual conversations, and ambiguous utterances, thereby enhancing user interaction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-04-20
- Publication Date
- 2026-03-27
AI Technical Summary
Conventional interaction systems struggle to appropriately respond to user utterances with ambiguous intentions, leading to mismatched responses that fail to meet user needs.
An information processing system employing machine learning to classify user utterances into clear tasks, casual conversations, or ambiguous intents, using a three-level classification model to detect and respond appropriately to ambiguous intents.
Enables appropriate responses to user utterances with ambiguous intentions, improving user experience by clarifying user intents and providing relevant feedback.
Smart Images

Figure 0007836785000001 
Figure 0007836785000002 
Figure 0007836785000003
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and an information processing program.
Background Art
[0002] Conventionally, when interacting with an agent, an interaction system that can avoid interaction failures as much as possible has been disclosed.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, the above conventional technology is such that two agents interact with a person according to a script (dialog data). In an interaction system, there are also user utterances with ambiguous intentions. If a response is made as it is to a user utterance with an ambiguous intention, there is a risk of making a response that does not match the user's intention. Therefore, in an interaction system, a means for appropriately responding to a user utterance with an ambiguous intention is required.
[0005] The present application has been made in view of the above, and an object thereof is to appropriately respond to a user utterance with an ambiguous intention.
Means for Solving the Problems
[0006] The information processing apparatus according to the present application, for information to be classified, The intention is clear and it can be considered a task. in the case of task the label of The intention is clear and it can be taken as casual conversation. in the case of Casual chat the label of, and when the intention is ambiguous It's hard to tell if it's a task or just casual conversation. in the case of an ambiguous label Along with labels for each type of informationA learning unit that constructs a model using machine learning with the labeled data obtained by attaching the label, and a unit that inputs information obtained from the user into the model, task , Casual chat Perform a three-level classification that categorizes the subject into either , , or ambiguous. In addition, if it is ambiguous, classify it into one of the following categories. This means that information with an ambiguous intention is being used. By type The detection unit that detects, When information with an ambiguous intent is detected, a selection unit selects a response content for each type of information with an ambiguous intent, and a response unit responds to the user with the selected response content for each type of information with an ambiguous intent. It is characterized by being equipped with [the following features]. [Effects of the Invention]
[0007] According to one embodiment, it is possible to respond appropriately even to user utterances with ambiguous intent. [Brief explanation of the drawing]
[0008] [Figure 1] Figure 1 is an explanatory diagram showing an overview of the information processing method according to the embodiment. [Figure 2] Figure 2 is an explanatory diagram illustrating an overview of responses to user utterances with ambiguous intent. [Figure 3] Figure 3 shows an example of the configuration of an information processing system according to the embodiment. [Figure 4] Figure 4 shows an example of the configuration of a terminal device according to this embodiment. [Figure 5] Figure 5 shows an example of the configuration of a server device according to this embodiment. [Figure 6] Figure 6 shows an example of a user information database. [Figure 7] Figure 7 shows an example of a historical information database. [Figure 8] Figure 8 shows an example of a speech information database. [Figure 9] Figure 9 is a flowchart showing the processing procedure according to the embodiment. [Figure 10] Figure 10 shows an example of a hardware configuration. [Modes for carrying out the invention]
[0009] Hereinafter, embodiments for implementing the information processing apparatus, information processing method, and information processing program according to the present application (hereinafter referred to as "embodiments") will be described in detail with reference to the drawings. Note that the information processing apparatus, information processing method, and information processing program according to the present application are not limited by this embodiment. Also, in the following embodiments, the same parts are denoted by the same reference numerals, and duplicate explanations are omitted.
[0010] 〔1. Outline of Information Processing Method〕 First, referring to FIG. 1, an outline of the information processing method performed by the information processing apparatus according to the embodiment will be described. FIG. 1 is an explanatory diagram showing an outline of the information processing method according to the embodiment. In FIG. 1, a case where an appropriate response is made to user utterances with ambiguous intentions will be described as an example.
[0011] As shown in FIG. 1, the information processing system 1 includes a terminal device 10 and a server device 100. The terminal device 10 and the server device 100 are connected to each other via a network N (see FIG. 3) so as to be communicable with each other by wire or wirelessly. In the present embodiment, the terminal device 10 cooperates with the server device 100.
[0012] The terminal device 10 is a smart device such as a smartphone or a tablet terminal used by a user U, and is a portable terminal device capable of communicating with an arbitrary server device via a wireless communication network such as 5G (Generation) or LTE (Long Term Evolution). Further, the terminal device 10 has a screen such as a liquid crystal display and has a screen having a touch panel function, and receives various operations on display data such as content, such as a tap operation, a slide operation, and a scroll operation, from the user U using a finger or a stylus. Note that an operation performed on a region of the screen where content is displayed may also be regarded as an operation on the content. Also, the terminal device 10 may be an information processing device such as a desktop PC (Personal Computer) or a notebook PC, in addition to a smart device.
[0013] The server device 100 is an information processing device that cooperates with the terminal device 10 of each user U and provides various APIs (Application Programming Interfaces) services and various data to the terminal device 10 of each user U. It is realized by a computer, a cloud system, or the like.
[0014] Also, the server device 100 may be an information processing device that provides some kind of web service online to the terminal device 10 of each user U. For example, as a web service, the server device 100 may provide services such as Internet connection, search service, SNS (Social Networking Service), e-commerce (EC: Electronic Commerce), electronic payment, online game, online banking, online trading, accommodation / ticket reservation, video / music distribution, news, map, route search, route guidance, route information, operation information, weather forecast, etc. In fact, the server device 100 may cooperate with various servers that provide the above web services and mediate the web services, or be responsible for the processing of the web services.
[0015] Note that the server device 100 can obtain user information regarding the user U. For example, the server device 100 obtains information regarding the attributes of the user U, such as the gender, age, and residential area of the user U. Then, the server device 100 stores and manages the information regarding the attributes of the user U together with the identification information (such as user ID) indicating the user U.
[0016] Furthermore, the server device 100 acquires various historical information (log data) indicating user U's actions from user U's terminal device 10, or from various servers based on the user ID, etc. For example, the server device 100 acquires location history, which is the history of user U's location and date and time, from the terminal device 10. The server device 100 also acquires search history, which is the history of search queries entered by user U, from the search server (search engine). The server device 100 also acquires browsing history, which is the history of content viewed by user U, from the content server. The server device 100 also acquires purchase history (payment history), which is the history of user U's product purchases and payment processing, from the e-commerce server or payment processing server. The server device 100 may also acquire listing history and sales history, which are the history of user U's listings on the marketplace, from the e-commerce server or payment processing server. The server device 100 also acquires posting history, which is the history of user U's posts, from posting servers that provide word-of-mouth posting services or SNS servers. The various servers mentioned above may also be the server device 100 itself. In other words, the server device 100 may function as the various servers mentioned above.
[0017] [1-1. Setting ambiguous labels] In intelligent conversational assistants (conversational agents) that handle both casual conversation and tasks through dialogue, there are user utterances with ambiguous intentions that could be interpreted as either casual conversation or tasks, such as "I'm hungry" or "My back hurts badly." If the intention of such utterances is definitively inferred, the response may not meet the request, resulting in a poor user experience.
[0018] For example, if a user utterance has a clear (not ambiguous) intent, such as "Tell me where Tokyo Station is," the app can respond in a way that aligns with the user's intent, such as "We will display the location of Tokyo Station in the app." However, with ambiguous user utterances, such as simply saying "Tokyo Station," it can be difficult to determine whether the user is asking for the location of Tokyo Station or requesting a route search to Tokyo Station. Responding directly to such ambiguous user utterances may result in a response that does not match the user's intent. Therefore, it is desirable to determine the ambiguity of the user's utterance in advance and address it in subsequent responses if necessary.
[0019] In this embodiment, as shown in Figure 1(a), the dialogue system labels user utterances as "task," "small talk," or "ambiguous," and applies an utterance intent classifier trained on the constructed dataset. As shown in Figure 2, this enables the detection of user utterances with ambiguous intent and allows the system to ask for clarification or provide responses with multiple intentions depending on the ambiguous intent. Figure 2 is an explanatory diagram illustrating an overview of responses to user utterances with ambiguous intent.
[0020] Furthermore, as an advanced form, as shown in Figure 1(b), the analysis results of the dialogue system regarding user utterances may also be input into the utterance intent classifier for learning. For example, if the analysis results of the dialogue system show that when the user utters only "Tokyo Station," and the Softmax function score is 0.6, indicating that the user is asking for "the location of Tokyo Station" and is seeking a response related to a map, then the response "Map," confidence level "0.6," and location "Tokyo" may be input into the utterance intent classifier for learning.
[0021] For example, as shown in Figure 1, the server device 100 collects multiple user utterances via the network N (see Figure 3) (step S1). In this embodiment, the server device 100 collects pairs consisting of user utterances and system responses from the dialogue log of the intelligent dialogue assistant.
[0022] Next, the server device 100 constructs a training dataset by assigning labels to user utterances (step S2). In this embodiment, the server device 100 performs crowdsourcing labeling on the log data of the intelligent dialogue assistant.
[0023] Specifically, the server device 100 presents the obtained conversation to a crowdsourced worker and labels the conversational utterances by requesting the worker to select from three labels: "Task" for user utterances with a task-oriented intention, "Casual Conversation" for non-task-oriented intentions, and "Ambiguous" for utterances with an ambiguous intention.
[0024] Next, the server device 100 constructs a model using machine learning with the training dataset (step S3). Model construction includes model updates. In this embodiment, the server device 100 first identifies trends in what kinds of utterances have ambiguous intentions from the obtained labeled data, and then constructs a supervised learning model (utterance intention classifier) using BERT (Bidirectional Encoder Representations from Transformers). That is, the server device 100 constructs a supervised classifier using the constructed dataset.
[0025] Next, the server device 100 receives user utterances from the user U's terminal device 10 via the network N (see Figure 3) (step S4).
[0026] Next, the server device 100 classifies the user utterance using the constructed model (step S5). In this embodiment, the server device 100 inputs the user utterance into the model (utterance intent classifier) and performs a three-level classification to determine whether it is a "task," "small talk," or "ambiguous." As a result, the server device 100 detects utterances with ambiguous intent.
[0027] Next, the server device 100 selects a response according to the classification of the user's utterance (step S6). In this embodiment, when the server device 100 detects that the user's utterance is "ambiguous," it selects a response that asks for clarification or details the user's intention accordingly.
[0028] Next, the server device 100 responds to the user U's terminal device 10 via the network N (see Figure 3) with the selected response content (step S7). In this embodiment, when the server device 100 detects that the user's utterance is "ambiguous," it takes appropriate actions such as asking for clarification or requesting further details of the user's intent.
[0029] As described above, the server device 100 according to this embodiment constructs a model that learns whether the information to be classified (user utterance) is a first classification target (task), a second classification target (small talk), or an ambiguous classification target (ambiguous). When the server device 100 acquires information from the user, it inputs the acquired information into the constructed model, classifies it into one of the first classification targets, second classification targets, or ambiguous classification targets, and returns a response according to the classification.
[0030] Furthermore, the information to be classified and the information obtained from the user may be strings extracted from the audio. Based on these strings, it may be determined whether the audio is a task instruction, casual conversation, or an ambiguous voice.
[0031] Furthermore, the server device 100 constructs a supervised speech intent classifier that has learned the features of learning information to which labels have been assigned indicating whether it is a first-classification target, a second-classification target, or an ambiguous classification target.
[0032] At this time, the server device 100 generates learning information through crowdsourcing. It shows the string created from the speech (spoken string) to people and asks them to assign one of three labels: it is a task, it is casual conversation, or it is unclear whether it belongs to a task or casual conversation and the intention is ambiguous. The server device 100 then trains its model to output the assigned label when it receives a spoken string as input.
[0033] The server device 100 also acquires utterances to be classified. When a user speaks, it converts the utterance into a string using speech recognition technology and inputs it into the model. Based on the model's classification results, it generates a response to the utterance string and provides it to the user.
[0034] Analysis of randomly selected information labeled "ambiguous" revealed that it could be broadly classified into "speech recognition errors," "nouns," "questions," "self-disclosure," "requests / commands," "points of criticism," and "other." Speech recognition errors and nouns were found to be particularly prevalent. Speech recognition errors included errors in kana-kanji conversion and word omissions, resulting in many instances where the intent was unclear and ambiguous. Nouns, requests / commands, and questions are generally used with the intention of information retrieval or terminal operation, but not always interpreted in this way. Information disclosure is often used in casual conversation, but some instances can be interpreted as implicit requests for tasks.
[0035] Furthermore, the server device 100 may further label and learn the utterances that have been classified as "ambiguous," according to their type, and use this information in its responses. Some utterances have an "ambiguous" intent. Accuracy can be improved by appropriately setting these as "classification targets."
[0036] For example, if the server device 100 assigns the label "ambiguous" to each of multiple user utterances, it may further assign labels such as speech recognition error, noun, question, self-disclosure, request / command, comment, or other, and then use the resulting labeled data to build a model using machine learning. The server device 100 may then input the user utterances into the model and further classify them into one of the following categories: speech recognition error, noun, question, self-disclosure, request / command, comment, or other.
[0037] Furthermore, the server device 100 may add a "meaningless" label in addition to the "ambiguous" label.
[0038] Furthermore, the server device 100 may construct a training dataset by assigning "negative," "positive," and "ambiguous" labels when classifying news, and then construct a supervised learning model using BERT. Alternatively, the learning model may be a model that classifies sentiment analysis, social media, etc., into three categories: "positive," "negative," and "ambiguous."
[0039] Furthermore, the server device 100 may add an "ambiguous" function as an intelligent dialogue assistant (conversation agent) function, in addition to the "small talk" function and the "task" function. For example, the server device 100 may determine which function to use to respond to a user utterance by performing a three-level classification that determines whether it is a "task," "small talk," or "ambiguous" using a learning model (classifier).
[0040] Furthermore, if additional functions are added to the intelligent dialogue assistant, the server device 100 may use a learning model (classifier) to perform a four-class classification to determine whether the response is a "task," "casual conversation," "ambiguous," or "additional function," thereby determining which of these functions to use in response.
[0041] For example, a learning model (classifier) learns about a service with multiple functions, and in addition to those functions, it also learns about functions whose classification is ambiguous.
[0042] [2. Example of an information processing system configuration] Next, the configuration of the information processing system 1, which includes the server device 100 according to the embodiment, will be described using Figure 3. Figure 3 is a diagram showing an example of the configuration of the information processing system 1 according to the embodiment. As shown in Figure 3, the information processing system 1 according to the embodiment includes a terminal device 10 and a server device 100. These various devices are connected to each other via a network N, either by wire or wireless communication. The network N is, for example, a LAN (Local Area Network) or a WAN (Wide Area Network) such as the Internet.
[0043] Furthermore, the number of devices included in the information processing system 1 shown in Figure 3 is not limited to those illustrated. For example, in Figure 3, only one terminal device 10 is shown for the sake of illustration, but this is merely an example and not limiting; there may be two or more.
[0044] Terminal device 10 is an information processing device used by user U. For example, terminal device 10 may be a smart device such as a smartphone or tablet, a mobile phone such as a feature phone, a PC (Personal Computer), a PDA (Personal Digital Assistant), a game console or AV equipment with communication functions, an information appliance or digital appliance, a car navigation system, a wearable device such as a smartwatch or head-mounted display, a smart speaker, smart glasses, etc. Alternatively, terminal device 10 may be a house or building compatible with the Internet of Things (IoT), a car, a home appliance, an electronic device, etc.
[0045] Furthermore, the terminal device 10 can connect to the network N via wireless communication networks such as LTE (Long Term Evolution), 4G (4th Generation), and 5G (5th Generation), or via short-range wireless communication such as Bluetooth (registered trademark) and Wi-Fi (Local Area Network), and communicate with the server device 100.
[0046] The server device 100 is, for example, a computer such as a PC or blade server, or a mainframe or workstation. The server device 100 may also be implemented through cloud computing.
[0047] [3. Example of terminal device configuration] Next, the configuration of the terminal device 10 will be explained using Figure 4. Figure 4 is a diagram showing an example of the configuration of the terminal device 10. As shown in Figure 4, the terminal device 10 comprises a communication unit 11, a display unit 12, an input unit 13, a positioning unit 14, a sensor unit 20, a control unit 30 (controller), and a storage unit 40.
[0048] (Communications Section 11) The communication unit 11 is connected to the network N (see Figure 3) by wire or wireless connection and transmits and receives information to and from the server device 100 via the network N. For example, the communication unit 11 can be implemented using a NIC (Network Interface Card) or an antenna.
[0049] (Display section 12) The display unit 12 is a display device that displays various information such as location information. For example, the display unit 12 may be a liquid crystal display (LCD) or an organic electro-luminescent display (OLED). The display unit 12 may also be a touch panel display, but is not limited to this.
[0050] (Input section 13) The input unit 13 is an input device that receives various operations from the user U. For example, the input unit 13 has buttons for inputting characters, numbers, etc. The input unit 13 may also be an input / output port (I / O port) or a USB (Universal Serial Bus) port. If the display unit 12 is a touch panel display, a part of the display unit 12 functions as the input unit 13. The input unit 13 may also be a microphone that receives voice input from the user U. The microphone may be wireless.
[0051] (Positioning unit 14) The positioning unit 14 receives signals (radio waves) transmitted from GPS (Global Positioning System) satellites and, based on the received signals, acquires position information (e.g., latitude and longitude) indicating the current position of the terminal device 10. In other words, the positioning unit 14 determines the position of the terminal device 10. Note that GPS is just one example of a GNSS (Global Navigation Satellite System).
[0052] Furthermore, the positioning unit 14 can determine its position using various methods other than GPS. For example, the positioning unit 14 may use various communication functions of the terminal device 10 to determine its position as an auxiliary positioning means for position correction, etc., as described below.
[0053] (Wi-Fi positioning) For example, the positioning unit 14 determines the location of the terminal device 10 by utilizing the Wi-Fi® communication function of the terminal device 10 and the communication network provided by each telecommunications company. Specifically, the positioning unit 14 determines the location of the terminal device 10 by performing Wi-Fi communication, etc., and determining the distance to nearby base stations and access points.
[0054] (Beacon positioning) Furthermore, the positioning unit 14 may determine the location using the Bluetooth® function of the terminal device 10. For example, the positioning unit 14 determines the location of the terminal device 10 by connecting to a beacon transmitter connected via the Bluetooth® function.
[0055] (Geomagnetic positioning) Furthermore, the positioning unit 14 determines the position of the terminal device 10 based on the geomagnetic pattern of the structure, which has been measured in advance, and the geomagnetic sensor provided by the terminal device 10.
[0056] (RFID positioning) Furthermore, if, for example, the terminal device 10 is equipped with an RFID (Radio Frequency Identification) tag function equivalent to that of a contactless IC card used at a train station ticket gate or in a store, or if it is equipped with a function to read RFID tags, the location where it was used will be recorded along with the information on the payment or other transactions made by the terminal device 10. The positioning unit 14 may determine the location of the terminal device 10 by acquiring such information. Alternatively, the location may be determined by an optical sensor or infrared sensor equipped in the terminal device 10.
[0057] The positioning unit 14 may, if necessary, determine the position of the terminal device 10 using one or a combination of the positioning means described above.
[0058] (Sensor unit 20) The sensor unit 20 includes various sensors mounted on or connected to the terminal device 10. The connection can be wired or wireless. For example, the sensors may be detection devices other than the terminal device 10, such as wearable devices or wireless devices. In the example shown in Figure 4, the sensor unit 20 includes an acceleration sensor 21, a gyro sensor 22, a barometric pressure sensor 23, a temperature sensor 24, a sound sensor 25, a light sensor 26, a magnetic sensor 27, and an image sensor (camera) 28.
[0059] The sensors 21-28 described above are merely examples and not limiting. In other words, the sensor unit 20 may be configured to include some of the sensors 21-28, or it may include other sensors such as humidity sensors in addition to or instead of the sensors 21-28.
[0060] The acceleration sensor 21 is, for example, a 3-axis acceleration sensor and detects the physical movement of the terminal device 10, such as its direction of movement, velocity, and acceleration. The gyro sensor 22 detects the physical movement of the terminal device 10, such as its tilt in the three axes, based on its angular velocity. The barometric pressure sensor 23 detects the atmospheric pressure around the terminal device 10, for example.
[0061] Since the terminal device 10 is equipped with the acceleration sensor 21, gyroscope 22, barometric pressure sensor 23, etc., it becomes possible to determine the position of the terminal device 10 using technologies such as pedestrian dead-reckoning (PDR) that utilize these sensors 21 to 23. This makes it possible to obtain indoor location information that is difficult to obtain with positioning systems such as GPS.
[0062] For example, a pedometer using an accelerometer 21 can calculate the number of steps, walking speed, and distance walked. Additionally, a gyroscope 22 can be used to determine the user U's direction of movement, gaze direction, and body tilt. Furthermore, the barometric pressure detected by the barometric pressure sensor 23 can be used to determine the altitude and floor number of the user U's terminal device 10.
[0063] The temperature sensor 24 detects, for example, the ambient temperature around the terminal device 10. The sound sensor 25 detects, for example, the ambient sound around the terminal device 10. The light sensor 26 detects the ambient illumination around the terminal device 10. The magnetic sensor 27 detects, for example, the Earth's magnetic field around the terminal device 10. The image sensor 28 captures an image of the area around the terminal device 10.
[0064] The aforementioned pressure sensor 23, temperature sensor 24, sound sensor 25, light sensor 26, and image sensor 28 can detect the surrounding environment and conditions of the terminal device 10 by detecting atmospheric pressure, temperature, sound, and illuminance, respectively, and by capturing images of the surroundings. Furthermore, it becomes possible to improve the accuracy of the location information of the terminal device 10 based on the surrounding environment and conditions.
[0065] (Control Unit 30) The control unit 30 includes, for example, a microcomputer having a CPU (Central Processing Unit), ROM (Read Only Memory), RAM, input / output ports, and various circuits. Alternatively, the control unit 30 may be composed of hardware such as an integrated circuit (ASIC) or FPGA (Field Programmable Gate Array). The control unit 30 includes a transmission unit 31, a reception unit 32, and a processing unit 33.
[0066] (Transmitter 31) The transmission unit 31 can transmit various information, such as information input by the user U using the input unit 13, various information detected by sensors 21-28 mounted on or connected to the terminal device 10, and location information of the terminal device 10 determined by the positioning unit 14, to the server device 100 via the communication unit 11.
[0067] (Receiving unit 32) The receiving unit 32 can receive various information provided by the server device 100 and requests for various information from the server device 100 via the communication unit 11.
[0068] (Processing 33) The processing unit 33 controls the entire terminal device 10, including the display unit 12. For example, the processing unit 33 can output and display various information transmitted by the transmission unit 31 and various information received from the server device 100 by the reception unit 32 to the display unit 12.
[0069] (Storage unit 40) The storage unit 40 is implemented by, for example, semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or by storage devices such as HDD (Hard Disk Drive), SSD (Solid State Drive), and optical discs. Various programs and various data are stored in this storage unit 40.
[0070] [4. Example of Server Device Configuration] Next, the configuration of the server device 100 according to the embodiment will be described using Figure 5. Figure 5 is a diagram showing an example of the configuration of the server device 100 according to the embodiment. As shown in Figure 5, the server device 100 includes a communication unit 110, a storage unit 120, and a control unit 130.
[0071] (Communications Department 110) The communication unit 110 is implemented, for example, by a NIC (Network Interface Card). The communication unit 110 is connected to the network N (see Figure 3) by wire or wireless connection.
[0072] (Storage unit 120) The storage unit 120 is implemented by, for example, semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or by storage devices such as HDDs, SSDs, and optical discs. As shown in Figure 5, the storage unit 120 includes a user information database 121, a history information database 122, and a speech information database 123.
[0073] (User Information Database 121) The user information database 121 stores user information about user U. For example, the user information database 121 stores various information such as user U's attributes. Figure 6 shows an example of the user information database 121. In the example shown in Figure 6, the user information database 121 has items such as "User ID (Identifier)", "Age", "Gender", "Home", "Workplace", and "Interests".
[0074] "User ID" refers to identification information used to identify user U. Note that "User ID" may be user U's contact information (telephone number, email address, etc.) or identification information used to identify user U's terminal device 10.
[0075] Furthermore, "Age" indicates the age of user U, identified by the user ID. Note that "Age" may be information indicating user U's specific age (e.g., 35 years old), or information indicating user U's age group (e.g., 30s), or "Age" may be information indicating user U's date of birth, or information indicating user U's generation (e.g., born in the 1980s). Furthermore, "Gender" indicates the gender of user U, identified by the user ID.
[0076] Furthermore, "Home" indicates the location information of user U's home, which is identified by the user ID. In the example shown in Figure 6, "Home" is represented by an abstract code such as "LC11," but it could also be latitude and longitude information, etc. Also, for example, "Home" could be a regional name or address.
[0077] Furthermore, "Workplace" indicates the location information of the workplace (or school in the case of a student) of user U, identified by the user ID. In the example shown in Figure 6, "Workplace" is illustrated with an abstract code such as "LC12," but it may also be latitude and longitude information, etc. Also, for example, "Workplace" may be a regional name or address.
[0078] Furthermore, "Interests" indicate the interests of user U, who is identified by their user ID. In other words, "Interests" indicate the subjects of high interest to user U, who is identified by their user ID. For example, "Interests" may be search queries (keywords) that user U enters into a search engine. In the example shown in Figure 6, one "Interest" is shown for each user U, but there may be multiple interests.
[0079] For example, in the example shown in Figure 6, user U, identified by user ID "U1", is in their 20s and is male. Also, for example, user U, identified by user ID "U1", has their home address at "LC11". Furthermore, for example, user U, identified by user ID "U1", has their workplace at "LC12". Finally, for example, user U, identified by user ID "U1", is interested in "sports".
[0080] In the example shown in Figure 6, abstract values such as "U1," "LC11," and "LC12" are used to illustrate the information, but it is assumed that "U1," "LC11," and "LC12" actually store specific strings, numbers, or other information. In the following diagrams relating to other information, abstract values may also be used to illustrate the information.
[0081] The user information database 121 is not limited to the above and may store various types of information depending on the purpose. For example, the user information database 121 may store various types of information about user U's terminal device 10. In addition, the user information database 121 may store information about user U's demographic, psychographic, geographic, and behavioral attributes. For example, the user information database 121 may store information such as name, family structure, place of origin (hometown), occupation, job title, income, qualifications, type of residence (detached house, apartment, etc.), whether or not a car is owned, commuting time, commuting route, commuter pass section (station, line, etc.), frequently used stations (other than the nearest station to home / workplace), lessons / classes (location, time, etc.), hobbies, interests, and lifestyle.
[0082] (History Information Database 122) The history information database 122 stores various information related to the history information (log data) that shows the user U's actions. Figure 7 shows an example of the history information database 122. In the example shown in Figure 7, the history information database 122 has items such as "User ID", "Location History", "Search History", "Browsing History", "Purchase History", and "Posting History".
[0083] "User ID" indicates identification information used to identify user U. "Location History" indicates the location history, which is the history of user U's location and movements. "Search History" indicates the search history, which is the history of search queries entered by user U. "Browsing History" indicates the browsing history, which is the history of content viewed by user U. "Purchase History" indicates the purchase history, which is the history of purchases made by user U. "Posting History" indicates the posting history, which is the history of posts made by user U. Note that "Posting History" may include questions about user U's possessions.
[0084] For example, in the example shown in Figure 7, user U, identified by user ID "U1", moves as described in "Location History #1", searches as described in "Search History #1", views content as described in "Browsing History #1", purchases specified goods at specified stores as described in "Purchase History #1", and posts as described in "Posting History #1".
[0085] In the example shown in Figure 7, abstract values such as "U1", "Location History #1", "Search History #1", "Browsing History #1", "Purchase History #1", and "Posting History #1" are used for illustration. However, it is assumed that "U1", "Location History #1", "Search History #1", "Browsing History #1", "Purchase History #1", and "Posting History #1" will actually store specific strings, numbers, and other information.
[0086] The history information database 122 is not limited to the above and may store various types of information depending on the purpose. For example, the history information database 122 may store the usage history of user U for a specified service. The history information database 122 may also store the visit history of user U to a physical store or a facility. The history information database 122 may also store the payment history of user U using the terminal device 10 for payments (electronic payments).
[0087] (Speech information database 123) The speech information database 123 stores various information about user utterances and responses obtained from the dialogue logs of the intelligent dialogue assistant. Figure 8 shows an example of the speech information database 123. In the example shown in Figure 8, the speech information database 123 has items such as "utterance," "response," "label," and "score."
[0088] "Utterance" refers to user utterances obtained from the dialogue log of the intelligent dialogue assistant. The content of the utterance may be a string extracted from the speech. "Response" refers to the response to the user utterance. The response may be a process corresponding to the content of the utterance. "Label" refers to the label assigned to the user utterance. For example, it may be one of the labels "Task," "Small Talk," or "Ambiguous." "Score" refers to the score of the label. For example, the score may be a confidence score or a score from the softmax function.
[0089] For example, in the example shown in Figure 8, the data for which the response "Map" is returned in response to the utterance "Tokyo Station" is labeled "Ambiguous" and its label score is indicated as "0.6".
[0090] In practice, one could use a classifier that takes a single utterance as input and outputs a softmax function score for each of the labels "task," "small talk," and "ambiguous," and then determine that the utterance corresponds to the label with the highest score.
[0091] The speech information database 123 is not limited to the above and may store various types of information depending on the purpose. For example, the speech information database 123 may store identification information to identify the user U who made the utterance. In addition, the speech information database 123 may store the context of user U at the time of the utterance, linked to the utterance.
[0092] (Control unit 130) Returning to Figure 5, let's continue the explanation. The control unit 130 is a controller, and is realized by various programs (corresponding to an example of an information processing program) stored in the internal memory of the server device 100, such as a CPU (Central Processing Unit), MPU (Micro Processing Unit), ASIC (Application Specific Integrated Circuit), or FPGA (Field Programmable Gate Array), executing them using a memory area such as RAM as the working area. In the example shown in Figure 5, the control unit 130 has an acquisition unit 131, a generation unit 132, a learning unit 133, a detection unit 134, a selection unit 135, and a response unit 136.
[0093] (Acquisition part 131) The acquisition unit 131 acquires the search query entered by the user U. For example, when the user U enters a search query into a search engine or the like and performs a keyword search, the acquisition unit 131 acquires the search query via the communication unit 110. In other words, the acquisition unit 131 acquires the keyword entered by the user U into the search box of a search engine, website, or application via the communication unit 110.
[0094] Furthermore, the acquisition unit 131 acquires user information about user U via the communication unit 110. For example, the acquisition unit 131 acquires identification information (such as user ID), location information, and attribute information of user U from user U's terminal device 10. The acquisition unit 131 may also acquire identification information and attribute information of user U when user U is registered. The acquisition unit 131 then registers the user information in the user information database 121 of the storage unit 120.
[0095] Furthermore, the acquisition unit 131 acquires various historical information (log data) indicating the user U's actions via the communication unit 110. For example, the acquisition unit 131 acquires various historical information indicating the user U's actions from the user U's terminal device 10, or from various servers based on the user ID, etc. The acquisition unit 131 then registers the various historical information in the history information database 122 of the storage unit 120.
[0096] Furthermore, the acquisition unit 131 receives and acquires information input by the user U via the communication unit 110. For example, the acquisition unit 131 receives and acquires the user's utterances via the communication unit 110. The acquisition unit 131 also collects pairs consisting of user utterances and system responses from the dialogue log of the intelligent dialogue assistant. (Generation unit 132) The generation unit 132 constructs a dataset in which the information entered by the user is labeled with a first-category label if it represents a first-category statement, a second-category label if it represents a second-category statement, and an ambiguous label if the intention is ambiguous. For example, the generation unit 132 constructs a dataset in which the user's utterances are labeled as either task, small talk, or ambiguous.
[0097] (Learning Section 133) The learning unit 133 constructs a model using machine learning with labeled data obtained by assigning a label for the first classification in the case of the first intention, a label for the second classification in the case of the second intention, and an ambiguous label in the case of an ambiguous intention to the information to be classified.
[0098] For example, the learning unit 133 constructs a model using machine learning with labeled data obtained by assigning a label for the first category if there is a first intention, a label for the second category if there is a second intention, and an ambiguous label if the intention is ambiguous, to each of multiple user utterances.
[0099] At this time, the learning unit 133 constructs a model by machine learning using labeled data obtained by assigning a task label to each of the multiple user utterances if the user utterance has a task-oriented intention, a casual conversation label if the intention is not task-oriented, and an ambiguous label if the intention is ambiguous and does not fall under either task-oriented or non-task-oriented.
[0100] For example, the learning unit 133 uses labeled data obtained by extracting strings from the audio of each of multiple user utterances, assigning a label for the first classification in the case of the first intention, a label for the second classification in the case of the second intention, and an ambiguous label in the case of an ambiguous intention, to construct a model using machine learning.
[0101] Furthermore, the learning unit 133 constructs a model using machine learning with the labeled data obtained by adding labels such as speech recognition error, noun, question, self-disclosure, request / command, comment, and other labels to each of the multiple user utterances.
[0102] From another perspective, the learning unit 133 uses the dataset constructed by the generation unit 132 to build a classifier that performs ternary classification, classifying user utterances into either tasks, small talk, or ambiguity, using supervised learning.
[0103] (Detection unit 134) The detection unit 134 inputs information obtained from the user into a model and performs a three-level classification, classifying it into either the first category, the second category, or ambiguous, thereby detecting information with ambiguous intent.
[0104] For example, the detection unit 134 inputs the utterance obtained from the user into a model and performs a three-level classification that classifies it into either the first category, the second category, or ambiguity, thereby detecting utterances with ambiguous intent.
[0105] At this time, the detection unit 134 inputs the utterance acquired from the user into the model and performs a three-level classification to determine whether it is a voice giving instructions for a task, a voice of casual conversation, or a voice with an ambiguous intent, thereby detecting utterances with an ambiguous intent.
[0106] For example, the detection unit 134 inputs a string extracted from the speech audio obtained from the user into a model and performs a three-level classification that categorizes it into either the first category, the second category, or ambiguity, thereby detecting utterances with ambiguous intent.
[0107] Furthermore, the detection unit 134 inputs the utterance acquired from the user into the model and further classifies it into one of the following categories: speech recognition error, noun, question, self-disclosure, request / command, point, or other.
[0108] From another perspective, the detection unit 134 inputs the user's utterance into the classifier to detect utterances with ambiguous intent.
[0109] (Selection Department 135) If the selection unit 135 detects information with an ambiguous intent, it selects a response to the information with an ambiguous intent. The response unit 136 may also select processing content according to the information entered by the user. For example, the response unit 136 may select an application to launch according to the content of the user's utterance. The selection unit 135 may also select processing content for information with an ambiguous intent if it detects such information.
[0110] (Response section 136) The response unit 136 responds to the user who input the information via the communication unit 110. For example, the response unit 136 responds to the user's utterance via the communication unit 110. In this case, if the response unit 136 detects information with an ambiguous intent, it will ask for clarification or elaborate on the intent according to the information. Alternatively, the response unit 136 will respond to the user with a response selected by the selection unit 135.
[0111] Furthermore, the response unit 136 may perform processing according to the information entered by the user. For example, the response unit 136 may launch an application corresponding to the content of the user's utterance. In this case, if the response unit 136 detects information with an ambiguous intent, it will ask for clarification or clarify the intent according to the information, and then perform processing according to the information entered by the user. Alternatively, the response unit 136 will perform processing according to the processing content selected by the selection unit 135.
[0112] [5. Processing Procedure] Next, the processing procedure by the server device 100 according to the embodiment will be described using Figure 9. Figure 9 is a flowchart of the processing procedure according to the embodiment. Note that the processing procedure shown below is repeatedly executed by the control unit 130 of the server device 100.
[0113] For example, as shown in Figure 9, the acquisition unit 131 of the server device 100 acquires user utterances from the dialogue log of the intelligent dialogue assistant via the communication unit 110 (step S101).
[0114] Next, the generation unit 132 of the server device 100 constructs a dataset in which each user utterance included in the dialogue log is labeled with either "Task" (indicating a task-oriented intention), "Small Talk" (indicating a non-task-oriented intention), or "Ambiguous" (indicating an ambiguous intention). (Step S102)
[0115] Next, the learning unit 133 of the server device 100 uses the dataset constructed by the generation unit 132 to construct a classifier that performs ternary classification, classifying user utterances into tasks, casual conversation, or ambiguity, using supervised learning (step S103).
[0116] Next, the detection unit 134 of the server device 100 inputs the user's utterance into the classifier and performs a three-level classification to classify the user's utterance into one of three categories: task, small talk, or ambiguity, thereby detecting utterances with ambiguous intent (step S104).
[0117] Next, if the selection unit 135 of the server device 100 detects an utterance with an ambiguous intent, it selects a response to the utterance with an ambiguous intent (step S105).
[0118] Next, the response unit 136 of the server device 100 responds to the user with the response content selected by the selection unit 135, thereby requesting clarification of the user's utterance or clarifying their intentions (step S106).
[0119] [6. Variant Example] The terminal device 10 and server device 100 described above may be implemented in various other forms besides those of the embodiment described above. Therefore, the following describes modifications of the embodiment.
[0120] In the above embodiment, some or all of the processing performed by the server device 100 may actually be performed by the terminal device 10. For example, the processing may be completed in a standalone manner (by the terminal device 10 alone). In this case, the terminal device 10 is assumed to have the functions of the server device 100 in the above embodiment. Furthermore, in the above embodiment, since the terminal device 10 is in cooperation with the server device 100, from the perspective of the user U, it appears as if the processing of the server device 100 is also being performed by the terminal device 10. In other words, from another perspective, it can be said that the terminal device 10 is equipped with the server device 100.
[0121] Furthermore, in the above embodiment, if the server device 100 detects a user utterance with an ambiguous intent, it may use the model to estimate an appropriate response to the user utterance with an ambiguous intent. For example, the server device 100 may construct a training dataset by associating confidence levels, likelihood levels, or scores with pairs of user utterances with ambiguous intent and responses, and then construct a model using machine learning with the constructed dataset. The server device 100 then inputs the user utterance with an ambiguous intent into the model and provides a response to the user utterance with an ambiguous intent based on the response obtained as the output of the model and the confidence levels, likelihood levels, or scores.
[0122] [7. Effects] As described above, the information processing device (terminal device 10 and server device 100) according to the present application comprises a learning unit 133 that constructs a model by machine learning using labeled data obtained by assigning a label for the first classification in the case of a first intention, a label for the second classification in the case of a second intention, and an ambiguous label in the case of an ambiguous intention to the information to be classified, and a detection unit 134 that inputs information obtained from a user U into the model and performs a three-class classification to classify it into the first classification, the second classification, or ambiguous, thereby detecting information with an ambiguous intention.
[0123] For example, the learning unit 133 builds a model using machine learning with labeled data obtained by assigning a label for the first category if there is a first intention, a label for the second category if there is a second intention, and an ambiguous label if the intention is ambiguous to each of multiple user utterances. The detection unit 134 inputs the utterances obtained from the user into the model and detects utterances with ambiguous intentions by performing ternary classification, classifying them into either the first category, the second category, or ambiguous.
[0124] Furthermore, the learning unit 133 builds a model using machine learning with the labeled data obtained by assigning a task label to each of multiple user utterances if the user utterance has a task-oriented intention, a casual conversation label if the intention is not task-oriented, and an ambiguous label if the intention is ambiguous and does not fall into either task-oriented or non-task-oriented categories. The detection unit 134 inputs the utterances obtained from the user into the model and performs a three-class classification to determine whether it is a voice instructing a task, a voice of casual conversation, or a voice with an ambiguous intention, thereby detecting utterances with ambiguous intentions.
[0125] Furthermore, the learning unit 133 constructs a model using machine learning with labeled data obtained by extracting strings from the audio of each of multiple user utterances and assigning a label for the first category if there is a first intention, a label for the second category if there is a second intention, and an ambiguous label if the intention is ambiguous. The detection unit 134 inputs the strings extracted from the audio of utterances obtained from the user into the model and performs ternary classification to classify them into either the first category, the second category, or ambiguous, thereby detecting utterances with ambiguous intentions.
[0126] Furthermore, the learning unit 133, when labeling each of multiple user utterances as ambiguous, constructs a model using machine learning with the resulting labeled data, which is further labeled with categories such as speech recognition error, noun, question, self-disclosure, request / command, comment, and other. The detection unit 134 inputs the utterances obtained from the user into the model and further classifies them into one of the following categories: speech recognition error, noun, question, self-disclosure, request / command, comment, or other.
[0127] Furthermore, the information processing device according to the present invention further includes a response unit 136 that, when it detects information with an ambiguous intent, asks for clarification or clarifies the intent in accordance with the information.
[0128] Alternatively, the information processing device according to the present invention further includes a selection unit 135 that selects a response to information with an ambiguous intent when it detects such information, and a response unit 136 that responds to the user with the selected response.
[0129] From another perspective, the information processing device according to the present invention comprises: a generation unit 132 that constructs a dataset in which user utterances are labeled as tasks, small talk, or ambiguous; a learning unit 133 that constructs a classifier that performs ternary classification using the dataset to classify user utterances as tasks, small talk, or ambiguous through supervised learning; and a detection unit 134 that inputs user utterances into the classifier and detects utterances with ambiguous intent.
[0130] Through any or a combination of the above-described processes, the information processing device according to the present invention can appropriately respond even to user utterances with ambiguous intent.
[0131] [8. Hardware Configuration] Furthermore, the terminal device 10 and server device 100 according to the above-described embodiment are realized by a computer 1000 having a configuration such as that shown in Figure 10. The following explanation will use the server device 100 as an example. Figure 10 shows an example of the hardware configuration. The computer 1000 is connected to an output device 1010 and an input device 1020, and has a configuration in which an arithmetic unit 1030, a primary storage device 1040, a secondary storage device 1050, an output interface 1060, an input interface 1070, and a network interface 1080 are connected by a bus 1090.
[0132] The arithmetic unit 1030 operates based on programs stored in the primary storage device 1040 and the secondary storage device 1050, as well as programs read from the input device 1020, and executes various processes. The arithmetic unit 1030 can be implemented using, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), an ASIC (Application Specific Integrated Circuit), or an FPGA (Field Programmable Gate Array).
[0133] The primary storage device 1040 is a memory device, such as RAM (Random Access Memory), that temporarily stores data used by the arithmetic unit 1030 for various calculations. The secondary storage device 1050 is a storage device where data used by the arithmetic unit 1030 for various calculations and various databases are registered, and can be implemented using ROM (Read Only Memory), HDD (Hard Disk Drive), SSD (Solid State Drive), flash memory, etc. The secondary storage device 1050 may be internal storage or external storage. The secondary storage device 1050 may also be a removable storage medium such as USB (Universal Serial Bus) memory or SD (Secure Digital) memory card. The secondary storage device 1050 may also be cloud storage (online storage), NAS (Network Attached Storage), file server, etc.
[0134] The output I / F 1060 is an interface for transmitting information to be output to output devices 1010, such as displays, projectors, and printers, and is implemented using connectors of standards such as USB (Universal Serial Bus), DVI (Digital Visual Interface), and HDMI (High Definition Multimedia Interface). The input I / F 1070 is an interface for receiving information from various input devices 1020, such as mice, keyboards, keypads, buttons, and scanners, and is implemented using, for example, USB.
[0135] Furthermore, the output interface 1060 and input interface 1070 may be wirelessly connected to the output device 1010 and input device 1020, respectively. In other words, the output device 1010 and input device 1020 may be wireless devices.
[0136] Furthermore, the output device 1010 and the input device 1020 may be integrated as a touch panel. In this case, the output I / F 1060 and the input I / F 1070 may also be integrated as an input / output I / F.
[0137] The input device 1020 may also be a device that reads information from, for example, an optical recording medium such as a CD (Compact Disc), DVD (Digital Versatile Disc), or PD (Phase Change Rewritable Disk), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.
[0138] The network interface 1080 receives data from other devices via network N and sends it to the computing unit 1030, and also transmits data generated by the computing unit 1030 to other devices via network N.
[0139] The arithmetic unit 1030 controls the output device 1010 and the input device 1020 via the output interface 1060 and the input interface 1070. For example, the arithmetic unit 1030 loads a program from the input device 1020 or the secondary storage device 1050 onto the primary storage device 1040 and executes the loaded program.
[0140] For example, when computer 1000 functions as a server device 100, the arithmetic unit 1030 of computer 1000 realizes the functions of the control unit 130 by executing a program loaded onto the primary storage device 1040. Alternatively, the arithmetic unit 1030 of computer 1000 may load a program obtained from another device via the network interface 1080 onto the primary storage device 1040 and execute the loaded program. Furthermore, the arithmetic unit 1030 of computer 1000 may cooperate with other devices via the network interface 1080 and call and use program functions, data, etc., from other programs on other devices.
[0141] [9. Other] Although embodiments of the present invention have been described above, the present invention is not limited by the content of these embodiments. Furthermore, the aforementioned components include those that can be easily conceived by those skilled in the art, those that are substantially the same, and those that fall within the so-called equivalent range. Moreover, the aforementioned components can be combined as appropriate. Furthermore, various omissions, substitutions, or modifications of the components can be made without departing from the gist of the embodiments described above.
[0142] Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically by known methods. In addition, the processing procedures, specific names, and information including various data and parameters shown in the above document and drawings can be arbitrarily changed unless otherwise specified. For example, the various information shown in each figure is not limited to the information shown.
[0143] Furthermore, the components of each illustrated device are functionally conceptual and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions.
[0144] For example, the server device 100 described above may be implemented using multiple server computers, and the configuration can be flexibly changed, such as by calling external platforms via APIs (Application Programming Interfaces) or network computing depending on the function.
[0145] Furthermore, the embodiments and modifications described above can be combined as appropriate, provided that the processing content is not inconsistent.
[0146] Furthermore, the terms "section, module, unit" mentioned above can be replaced with "means" or "circuit," etc. For example, the acquisition unit can be replaced with acquisition means or acquisition circuit. [Explanation of Symbols]
[0147] 1. Information Processing System 10 Terminal devices 100 Server Devices 110 Communications Department 120 Storage section 121 User Information Database 122 History Information Database 123 Speech Information Database 130 Control Unit 131 Acquisition Department 132 Generation part 133 Learning Department 134 Detection unit 135 Selection Department 136 Response section
Claims
1. The learning unit constructs a model using machine learning with labeled data obtained by assigning labels to the information to be classified: a task label if the intention is clear and it can be considered a task, a chat label if the intention is clear and it can be considered chat, and an ambiguous label if the intention is unclear and it is difficult to determine whether it is a task, chat, or not. A detection unit that inputs information obtained from a user into the model and performs a three-level classification into tasks, casual conversation, or ambiguity, and in the case of ambiguity, classifies it into one of the categories, thereby detecting information with ambiguous intent for each category. When information with an ambiguous intent is detected, a selection unit selects the appropriate response content for each type of information with an ambiguous intent, A response unit that responds to the user with selected response content for each type of information with an ambiguous intent, An information processing device characterized by comprising:
2. The learning unit constructs a model using machine learning with labeled data obtained by assigning a task label to each of multiple user utterances if the intention is clear and it can be considered a task, a chat label if the intention is clear and it can be considered chat, and an ambiguous label if the intention is unclear and it is difficult to determine whether it is a task, chat, or not. The detection unit inputs the utterance obtained from the user into the model and performs a three-level classification into task, casual conversation, or ambiguity. In the case of ambiguity, it classifies it into one of the categories, thereby detecting utterances with ambiguous intent for each category. When the selection unit detects an utterance with an ambiguous intent, it selects a response for each type of utterance with an ambiguous intent. The response unit responds to ambiguous utterances with selected response content for each category to the user. The information processing apparatus according to feature 1.
3. The learning unit constructs a model using machine learning with labeled data obtained by assigning a task label to each of multiple user utterances if the user utterance has a task-oriented intention, a casual conversation label if the intention is not task-oriented, and an ambiguous label if the intention is ambiguous and does not fall under either task-oriented or non-task-oriented, along with labels for each type of utterance. The detection unit inputs the utterance obtained from the user into the model and performs a three-level classification to determine whether it is a voice giving a task instruction, a voice of casual conversation, or a voice with an ambiguous intent. If the voice has an ambiguous intent, it determines which type of voice it is, thereby detecting utterances with ambiguous intent by type. When the selection unit detects an utterance with an ambiguous intent, it selects a response for each type of utterance with an ambiguous intent. The response unit responds to ambiguous utterances with selected response content for each category to the user. The information processing apparatus according to feature 1.
4. The learning unit constructs a model using machine learning with labeled data obtained by extracting strings from the audio of multiple user utterances, assigning a task label if the intention is clear and it can be considered a task, a chat label if the intention is clear and it can be considered chat, and an ambiguous label if the intention is unclear and it is difficult to determine whether it is a task or chat, along with labels for each type of string. The detection unit inputs the string extracted from the speech audio obtained from the user into the model and performs a three-level classification into task, casual conversation, or ambiguity. In the case of ambiguity, it classifies it into one of the categories, thereby detecting utterances with ambiguous intent for each category. When the selection unit detects an utterance with an ambiguous intent, it selects a response for each type of utterance with an ambiguous intent. The response unit responds to ambiguous utterances with selected response content for each category to the user. The information processing apparatus according to feature 1.
5. The learning unit, when assigning the aforementioned ambiguity label to each of multiple user utterances, further assigns labels such as speech recognition error, noun, question, self-disclosure, request / command, comment, and other labels to the resulting labeled data, and uses machine learning to construct a model. The detection unit inputs the utterance obtained from the user into the model and further classifies it into one of the following categories: speech recognition error, noun, question, self-disclosure, request / command, comment, or other. The information processing apparatus according to feature 1.
6. When the response unit detects information with an ambiguous intent, it will ask for clarification or clarify the intent in accordance with the information. The information processing apparatus according to feature 1.
7. The learning unit constructs a model by machine learning using labeled data obtained by assigning the labels Task, casual conversation, ambiguity, or meaningless to the information to be classified. The detection unit inputs the information to be classified into the model and performs a four-level classification into one of the following categories: task, casual conversation, ambiguity, or meaningless. In the case of ambiguity, it classifies it into one of these categories, thereby detecting utterances with ambiguous intent for each category. The information processing apparatus according to feature 1.
8. The learning unit constructs a model by machine learning using labeled data obtained by labeling utterances acquired from the user with the functions of an intelligent dialogue assistant, namely tasks, small talk, ambiguity, or additional functions, The detection unit inputs the user's utterances regarding the functions of the intelligent dialogue assistant into the model and performs a four-level classification to classify them into one of the following: task, small talk, ambiguity, or additional functions. A selection unit that selects which function to use to respond to the user's utterance: task, casual conversation, ambiguity, or additional function, The response unit responds to the user with a selected function as a response from the intelligent dialogue assistant. The information processing apparatus according to feature 1.
9. An information processing method performed by an information processing device, The learning process involves using labeled data obtained by assigning labels to the information to be classified: a task label if the intent is clear and it can be considered a task, a chat label if the intent is clear and it can be considered casual conversation, and an ambiguous label if the intent is unclear and it is difficult to determine whether it is a task, casual conversation, or casual conversation, along with labels for each type of information, to build a model using machine learning. A detection process in which information obtained from a user is input into the model and classified into one of three categories: task, casual conversation, or ambiguous, and if it is ambiguous, it is classified into one of the categories, thereby detecting information with ambiguous intent for each category. If information with an ambiguous intent is detected, a selection process is performed to select the appropriate response content for each type of information with an ambiguous intent. A response process that responds to information with an ambiguous intent by providing a response content for each selected category to the user, An information processing method characterized by including
10. The learning procedure involves using labeled data obtained by assigning labels to the information to be classified: a task label if the intent is clear and it can be considered a task, a chat label if the intent is clear and it can be considered chat, and an ambiguous label if the intent is unclear and it is difficult to determine whether it is a task, chat, or not, along with labels for each type of information, to build a model using machine learning. A detection procedure that inputs information obtained from a user into the model and performs a three-level classification into either a task, casual conversation, or ambiguity, and in the case of ambiguity, classifies it into one of the categories, thereby detecting information with ambiguous intent for each category. If information with an ambiguous intent is detected, a selection procedure is used to select the appropriate response content for each type of information with an ambiguous intent. A response procedure that responds to the user with selected response content for each type of information with an ambiguous intent, An information processing program characterized by causing a computer to execute it.
Citation Information
Patent Citations
Orphan speech detection system and method
JP2017534941A
Classification device, classification method, and classification program
JP2018151786A
Dialogue system and domain determining method
JP2019070957A
Dialogue systems and programs
JP2023001299A