Information processing device, information processing method, mobile body control device, mobile body control method, and program

By associating machine learning models with specific usage scenarios, the device efficiently classifies user intentions for controlling a mobile object through speech, addressing the need for extensive training data and improving accuracy.

JP7733465B2Active Publication Date: 2025-09-03HONDA MOTOR CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021058446
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-03-30
Publication Date
2025-09-03
Estimated Expiration
2041-03-30

AI Technical Summary

Technical Problem

Existing technologies require a large amount of training data to accurately classify various user intentions for controlling a mobile object through speech, leading to inefficiencies or inaccuracies in intent classification.

Method used

An information processing device that identifies the usage scenario of a user and selects a specific machine learning model tailored to that scenario, allowing for smaller-scale learning and improved intent classification by associating different machine learning models with distinct usage scenes such as before, during, and after boarding a vehicle.

Benefits of technology

Enables accurate and efficient classification of speech intentions using smaller-scale learning, reducing the need for extensive training data and improving recognition accuracy by utilizing models appropriate for each scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007733465000002
    Figure 0007733465000002
  • Figure 0007733465000003
    Figure 0007733465000003
  • Figure 0007733465000004
    Figure 0007733465000004
Patent Text Reader

Abstract

To provide an information processing apparatus which provides utterance intent classification by a model constructed by smaller-scale learning in control of a mobile object by utterance, an information processing method, a mobile object control device, a mobile object control method, and a program.SOLUTION: In an information processing system including a server that acquires position information of each of vehicles to control travel of the vehicles, the server 110 which is an information processing apparatus capable of controlling a mobile object on the basis of an instruction by an utterance of a user. A control unit 404 of the server includes: a scene identification unit 414 which identifies a use scene of a target user out of a plurality of use scenes in using the mobile object; a user data acquisition unit 413 which acquires utterance information of the target user; a model selection unit 415 which selects a different machine learning model according to the identified use scene of the target user; and an utterance intent estimation unit 416 which estimates an intent of an utterance of the target user by using the selected machine learning model.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, a control device for a mobile body, a control method for a mobile body, and a program. [Background technology]

[0002] In recent years, the development of man-machine interfaces using natural language has progressed. Non-Patent Document 1 proposes a technology that realizes intent classification and slot filling in spoken sentences using a language representation model called BERT. Intent classification in spoken sentences is a technology that estimates the user's intention, for example, in the user's instructions or questions (also called queries), while slot filling is a technology that recognizes information already provided by the user or information that is missing, and asks clarifying questions or supplements the information. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Qian Chen and two others, BERT for Joint Intent Classification and Slot Filling, February 28, 2019, https: / / arxiv.org / pdf / 1902.10909.pdf Summary of the Invention [Problem to be solved by the invention]

[0004] Non-Patent Document 1 proposes a technology that simultaneously performs intent classification and slot filling using a single model implemented with BERT, which requires training using a huge amount of data in order to classify utterances into one of many intent classes.

[0005] In order for a single classifier model to classify user intentions, it is necessary to solve a classification problem for a large number of intention classes that encompass all possible scenarios. When it is assumed that a user controls a mobile object through speech, a large number of user intentions may exist, such as an inquiry about availability to summon a nearby mobile object, a route instruction for the mobile object, an instruction regarding the vehicle's operation (e.g., an instruction to accelerate), and an instruction to forward the mobile object after riding. In other words, when controlling a mobile object through speech, a large model is required to classify various utterance intentions, from an inquiry about availability to an instruction to forward the object. As a result, a huge amount of training data may be required, or an intention classification result with the desired accuracy may not be obtained.

[0006] The present invention has been made in consideration of the above-mentioned problems, and its purpose is to realize a technology that can provide classification of speech intentions using a model constructed through smaller-scale learning when controlling a moving object through speech. [Means for solving the problem]

[0007] According to the present invention, An information processing device capable of controlling a moving object based on a user's spoken instruction, an identification means for identifying which usage scene of a target user is associated with a plurality of usage scenes related to instructions from the user when using a mobile object; an acquisition means for acquiring utterance information of the target user; a selection means for selecting a different machine learning model depending on the usage scene of the identified target user; An estimation means for estimating the intention of the target user's utterance using the selected machine learning model. death, The identification means uses information acquired from a moving body associated with the target user as to whether the target user is riding in the moving body and information as to whether the target user has gotten on or off the moving body within a predetermined time to identify whether the scene is before the target user uses the moving body, while the target user is using the moving body, or after the target user has used the moving body. The present invention provides an information processing device characterized by: [Effects of the Invention]

[0008] According to the present invention, in controlling a moving object by speech, it is possible to provide classification of speech intentions using a model constructed through smaller-scale learning. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of an information processing system according to an embodiment of the present invention. [Figure 2] FIG. 1 is a block diagram showing an example of the hardware configuration of a vehicle according to an embodiment of the present invention. [Figure 3] FIG. 1 is a block diagram showing an example of the functional configuration of a vehicle according to an embodiment of the present invention; [Figure 4] FIG. 1 is a block diagram illustrating an example of the functional configuration of a server according to the present embodiment. [Figure 5A] FIG. 1 is a diagram illustrating an example of a usage scenario when using a vehicle, an intention class of an utterance associated with the usage scenario, and an utterance corresponding to the intention class, according to the present embodiment. [Figure 5B] FIG. 10 is a diagram illustrating an example in which a hidden Markov model is applied to a usage scene before boarding in this embodiment. [Figure 5C] FIG. 10 is a diagram illustrating how the intention of an utterance is estimated in successive usage scenes according to the present embodiment. [Figure 6] 1 is a flowchart showing a series of operations in a speech intention estimation process according to the present embodiment. [Figure 7] FIG. 10 is a diagram illustrating an example of an information processing system according to another embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, the embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention as claimed, and not all combinations of features described in the embodiments are necessarily essential to the invention. Two or more of the features described in the embodiments may be arbitrarily combined. Furthermore, the same reference numerals are used for the same or similar components, and redundant explanations will be omitted.

[0011] (Configuration of information processing system) The configuration of an information processing system 1 according to this embodiment will be described with reference to Fig. 1. The information processing system 1 includes a vehicle 100 as an example of a moving object, a server 110 as an example of an information processing device, and a communication device 120.

[0012] In this information processing system 1, the user 130 can interact with the vehicle 100 or control the operation of the vehicle 100 by speaking in natural language. When the communication device 120 receives an utterance from the user 130 directed to the vehicle 100, it transmits the utterance information to the server 110. The server 110 estimates the user's intention from the user's utterance information. Based on the estimated intention, the server 110 performs slot filling (recognizing information already provided by the user or information that is missing, and providing the user 130 with a question for clarification as needed). When the server 110 recognizes the user's intention, it identifies information necessary from the utterance information to identify specific instructions and accepts the user's instructions. When the user's utterance is an instruction directed to the vehicle 100 (e.g., "Pick me up at my current location immediately"), the server 110 transmits a control instruction corresponding to the instruction to the vehicle 100.

[0013] Vehicle 100 is an example of a moving body, and is, for example, an ultra-compact mobility vehicle that is equipped with a battery and moves mainly by motor power. An ultra-compact mobility vehicle is an ultra-compact vehicle that is more compact than a typical automobile and has a passenger capacity of about one or two people. In this embodiment, vehicle 100 is, for example, a four-wheeled vehicle. Note that in the following embodiments, the moving body is not limited to a vehicle, and may include a small mobility vehicle that runs alongside a walking user to carry luggage or lead a person, and may also include other moving bodies that are capable of autonomous movement (for example, a walking robot).

[0014] The vehicle 100 connects to the network 140 via wireless communication such as Wi-Fi or fifth-generation mobile communication. The vehicle 100 can measure conditions inside and outside the vehicle (such as the vehicle's position, driving status, and surrounding object landmarks) using various sensors and transmit the measured data to the server 110. The data collected and transmitted in this manner is generally referred to as floating data, probe data, traffic information, etc. Information about the vehicle is transmitted to the server 110 at regular intervals or in response to the occurrence of a specific event. The vehicle 100 can travel autonomously even when the user 130 is not on board. The vehicle 100 receives information such as control commands provided by the server 110 or controls the operation of the vehicle using data measured by the vehicle itself.

[0015] The server 110 is an example of an information processing device. The server 110 is configured with one or more server devices, and is capable of acquiring vehicle-related information transmitted from the vehicle 100, utterance information transmitted from the communication device 120, and respective position information via the network 111, and controlling the traveling of the vehicle 100. The server 110 executes a user intention estimation process (described later) from the user utterance information to estimate the user's intention in the utterance.

[0016] The communication device 120 is, for example, a smartphone, but is not limited to this and may be an earphone-type communication terminal, a personal computer, a tablet terminal, a game console, etc. The communication device 120 connects to the network 140 via wireless communication such as Wi-Fi or fifth-generation mobile communication. The communication device 120 receives speech from the user 130 and transmits the received speech information (audio information) to the server 110.

[0017] The network 111 includes a communication network such as the Internet or a mobile phone network, and transmits information between the server 110 and the vehicle 100 and information between the server 110 and the communication device 120 .

[0018] (Vehicle configuration) Next, with reference to FIG. 2, a configuration of a vehicle 100 as an example of a vehicle according to this embodiment will be described.

[0019] Fig. 2(A) shows a side view of the vehicle 100 according to this embodiment, and Fig. 2(B) shows the internal configuration of the vehicle 100. In the figure, arrow X indicates the longitudinal direction of the vehicle 100, with F indicating the front and R indicating the rear. Arrows Y and Z indicate the width direction (left-right direction) and up-down direction of the vehicle 100.

[0020] The vehicle 100 is an electric autonomous vehicle equipped with a propulsion unit 12 and using a battery 13 as its main power source. The battery 13 is, for example, a secondary battery such as a lithium-ion battery, and the vehicle 100 is propelled by the propulsion unit 12 using power supplied from the battery 13. The propulsion unit 12 is a four-wheeled vehicle equipped with a pair of left and right front wheels 20 and a pair of left and right rear wheels 21. The propulsion unit 12 may be in another form, such as a tricycle. The vehicle 100 is equipped with a seat 14 for one or two people. The seat 14 transmits, for example, a pressure sensor or the like, to the control unit 30 whether or not a passenger is riding in the seat 14.

[0021] The traveling unit 12 includes a steering mechanism 22. The steering mechanism 22 is a mechanism that uses a motor 22a as a drive source to change the steering angle of the pair of front wheels 20. By changing the steering angle of the pair of front wheels 20, the traveling direction of the vehicle 100 can be changed. The traveling unit 12 also includes a drive mechanism 23. The drive mechanism 23 is a mechanism that uses a motor 23a as a drive source to rotate the pair of rear wheels 21. By rotating the pair of rear wheels 21, the vehicle 100 can move forward or backward.

[0022] The vehicle 100 is equipped with detection units 15 to 17 that detect targets around the vehicle 100. The detection units 15 to 17 are a group of external sensors that monitor the periphery of the vehicle 100, and in the present embodiment, each is an imaging device that captures an image of the periphery of the vehicle 100, and includes, for example, an optical system such as a lens and an image sensor. However, instead of or in addition to the imaging device, it is also possible to employ radar or lidar (Light Detection and Ranging).

[0023] Two detection units 15 are arranged at the front of the vehicle 100, spaced apart in the Y direction, and mainly detect targets in front of the vehicle 100. Detection units 16 are arranged on the left and right sides of the vehicle 100, respectively, and mainly detect targets on the sides of the vehicle 100. Detection unit 17 is arranged at the rear of the vehicle 100, and mainly detects targets behind the vehicle 100.

[0024] 3 is a block diagram of a control system of the vehicle 100. The vehicle 100 includes a control unit (ECU) 30. The control unit 30 includes a processor such as a CPU, a storage device such as a semiconductor memory, an interface with external devices, etc. The storage device stores programs executed by the processor and data used by the processor for processing. A plurality of sets of processors, storage devices, and interfaces may be provided for different functions of the vehicle 100 and configured to be able to communicate with each other. Voice recognition processing may be performed on input voice, and image recognition processing may be performed on images captured by a detection unit.

[0025] The control unit 30 executes corresponding processing in response to the detection results of the detection units 15 to 17, input information from the operation panel 31, audio information input from the audio input device 33, control commands from the server 110, etc. The control unit 30 controls the motors 22a and 23a (travel control of the traveling unit 12), controls the display on the operation panel 31, and outputs audio alerts and information to the occupants of the vehicle 100.

[0026] The voice input device 33 can collect voices of passengers of the vehicle 100. The control unit 30 can recognize the input voices and execute corresponding processing. The GNSS (Global Navigation Satellite system) sensor 34 receives GNSS signals and detects the current position of the vehicle 100.

[0027] The storage device 35 is a large-capacity storage device that stores map data including information on routes that the vehicle 100 can travel, landmarks such as buildings, stores, etc. The storage device 35 may also store programs executed by the processor, data used by the processor for processing, etc. The storage device 35 may also store various parameters of machine learning models for speech recognition and image recognition executed by the control unit 30 (for example, trained parameters of a deep neural network, etc.).

[0028] The communication device 36 is a communication device that can connect to the network 140 via wireless communication such as Wi-Fi or fifth generation mobile communication.

[0029] (Server configuration) Next, with reference to FIG. 4, the configuration of the server 110 as an example of an information processing device according to this embodiment will be described.

[0030] The control unit 404 includes a processor such as a CPU, a storage device such as a semiconductor memory, an interface with an external device, etc. The storage device stores programs executed by the processor, data used by the processor for processing, etc. A plurality of sets of processors, storage devices, and interfaces may be provided for different functions of the server 110 and configured to be able to communicate with each other. The control unit 404 executes programs to perform various operations of the server 110, a user intention estimation process (described later), control of the vehicle 100, etc. In addition to the CPU, the control unit 404 may further include a GPU or dedicated hardware suitable for executing processing of a machine learning model such as a neural network.

[0031] The user data acquisition unit 413 acquires utterance information of the user 130 transmitted from the communication device 120. The user data acquisition unit 413 also acquires floating data information (such as the vehicle's position and the presence or absence of an occupant) transmitted from the vehicle 100. The user data acquisition unit 413 may store the acquired utterance information, position information, etc. in the storage unit 403. The utterance information acquired by the user data acquisition unit 413 is input to a trained model in the inference stage (which has already been trained), but may also be used as training data for training a machine learning model executed by the server 110.

[0032] The scene identification unit 414 identifies the current situation (scene) in which the user is placed. For example, the scene identification unit 414 identifies whether the user's scene is before getting on, during getting on, or after getting off. An example of a scene identification method will be described later.

[0033] The model selection unit 415 selects a machine learning model for the scene identified by the scene identification unit 414. As will be described later, there are multiple machine learning models, and each machine learning model is associated with, for example, before getting on, during getting on, or after getting off. That is, each machine learning model estimates a different intention class for each associated usage scene, and is configured to output the likelihood of the intention class for any of the scenes.

[0034] The speech intention estimation unit 416 estimates the intention of the user's utterance using the machine learning model selected by the model selection unit 415. The method for estimating the intention of the utterance will be described later.

[0035] The speech information processing unit 417 identifies necessary information from the speech information to identify specific instruction content based on the recognized speech intention. For example, if the intention of the user's speech information is a request to be picked up, the speech information processing unit 417 identifies information such as where and what time to be picked up. After identifying the necessary information, the speech information processing unit 417 accepts the user's instruction. The speech information processing unit 417 may further include slot filling processing. The speech information processing unit 417 may include multiple machine learning models different from the machine learning model for estimating the speech intention, and each machine learning model may be composed of, for example, a deep neural network (DNN). The DNN becomes trained by performing a learning stage process, and new utterance information can be input into the trained DNN to process the new utterance information (inference stage process).

[0036] Vehicle control unit 418 controls the operation of vehicle 100 based on the content of the utterance recognized by voice information processing unit 417. For example, when information such as where and what time to pick up the user is identified from the user's utterance information, vehicle control unit 418 identifies a route based on the current positions of the user and vehicle, map information, etc., and causes the vehicle to travel along that route.

[0037] The server 110 generally has more abundant computational resources available than the vehicle 100. This allows the server 110 to provide computational results faster than if each vehicle 100 were equipped with computational resources for executing a machine learning model, and can also contribute to reducing vehicle costs. Furthermore, the server 110 receives and stores speech information from various users, allowing it to collect learning data containing a wide variety of speech information, enabling more robust inference processing.

[0038] The communication unit 401 is a communication device including, for example, a communication circuit, and communicates with external devices such as the vehicle 100 and the communication device 120. The communication unit 401 receives position information and occupant presence information from the vehicle 100, and speech information and position information from the communication device 120, as well as transmits control commands to the vehicle 100 and speech information to the communication device 120.

[0039] The power supply unit 402 supplies power to each unit in the server 110. The storage unit 403 is a non-volatile memory such as a hard disk or semiconductor memory.

[0040] (Overview of user intent estimation process) As described above, assuming that a user controls a mobile object by speaking, there may be many user intentions, such as, for example, an inquiry about the availability of a mobile object to call a nearby mobile object, instructions on a route for the mobile object, instructions regarding the vehicle's operation (e.g., instructions to accelerate), and instructions to return a mobile object after a ride has ended.

[0041] However, before using a vehicle, there is a possibility that an utterance with the intention of inquiring about the availability of the vehicle or calling the vehicle may be made, but there is a low possibility that an utterance with the intention of instructing the vehicle to be forwarded after use will be made. In other words, some of the conversational intentions made in each scene before boarding, during boarding, and after getting off may not appear in other scenes.

[0042] For this reason, in this embodiment, an intention class is compiled for each assumed scene (usage scene) when controlling a moving object by speech, and a machine learning model is associated with each scene. Each machine learning model estimates only the intention class of the associated usage scene. In this way, a model appropriate for each scene can be used, and each model can be made smaller in size and built with smaller learning than when a single model classifies multiple intention classes. Furthermore, improved recognition accuracy can be expected.

[0043] Hereinafter, the relationship between usage scenes and utterance intention classes according to this embodiment, and the intention estimation algorithm will be described with reference to FIGS. 5A to 5C.

[0044] 5A shows examples of usage scenarios when using a vehicle, intention classes of utterances associated with the usage scenarios, and utterances corresponding to the intention classes. As shown in FIG. 5A, usage scenario 501 is divided into, for example, a state before getting into vehicle 100 (pre-boarding state), a state while getting into vehicle 100 (boarding state), and a state after getting off vehicle 100 (post-boarding state). Note that "always" in FIG. 5A does not refer to a specific usage scenario, but rather means that the three intention classes belonging to "always" are included in all usage scenarios.

[0045] The utterance intention class 502 represents the intention of the user's utterance. Seven intention classes, such as inquiry, pick-up request, greeting, destination instruction, agreement, denial, and asking to reflect, are associated with the usage scene "before boarding." Seven intention classes, such as route instruction, stop instruction, acceleration instruction, deceleration instruction, agreement, denial, and asking to reflect, are associated with the usage scene "during boarding." Similarly, the seven intention classes shown in FIG. 5A are associated with the usage scene "after getting off."

[0046] Example utterances 503 show examples of utterances corresponding to each intent class. For example, an utterance such as "Can I get in now?" corresponds to the intent of "enquiry."

[0047] In this way, a predetermined number of consecutive scenes are defined as possible scenarios when using a vehicle, and only a portion of all the intent classes that can be associated with the multiple usage scenes are associated with each usage scene. In this way, the machine learning model only needs to output inference results for a portion of all the intent classes that can be associated with the multiple usage scenes. Therefore, the machine learning model that estimates the intent classes can be made smaller and trained using a small amount of training data.

[0048] In the example shown in FIG. 5A, the same number of intention classes are associated with all usage scenarios, but a different number of intention classes may be associated with each usage scenario. Furthermore, the intention classes are not limited to the above examples, and may include other intention classes, or some of the intention classes shown in FIG. 5A may not be included. For example, the usage scenario before getting into the vehicle may further include a "catch-up request" that requests the vehicle to catch up with the user's destination. A possible utterance for the catch-up request is, for example, "catch up."

[0049] Furthermore, the utterance examples are examples of utterances in which the user 130 speaks to the vehicle 100. However, the utterances are not limited to examples of utterances in which the user 130 speaks to the vehicle 100, and utterances in which the user 130 speaks to a human concierge (who mediates vehicle control) may also be used.

[0050] Next, an example of applying a hidden Markov model to a usage scenario before boarding will be described with reference to Fig. 5B. A Markov model is a probabilistic model that follows a stochastic process in which the probability distribution of a state at any time depends only on the immediately preceding state. In this embodiment, a hidden Markov model (also called an HMM) is applied to solve the problem of estimating the hidden state (utterance intention) behind an observable state (utterance information) when it is given.

[0051] Numerals 510 to 513 in FIG. 5B represent hidden states in a hidden Markov model, which correspond to intent classes. Note that in the example shown in FIG. 5B, only four intent classes are shown to avoid complicating the diagram. The numerical values ​​written within the intent class circles indicate initial state probabilities. That is, they indicate the probability (likelihood) of the intent class that may occur immediately after the usage scenario "before boarding" has occurred. Furthermore, each arrow indicates a transition between intent classes (states), and the numerical values ​​attached to the arrows (e.g., "0.aa") indicate state transition probabilities. The distributions of the initial state probabilities and the state transition probabilities may be determined in advance. For example, the transition probabilities between each intent class and the initial state probabilities of the correct answer data included in the training data can be calculated and used.

[0052] Furthermore, an example of the intention estimation process according to this embodiment will be described with reference to Fig. 5C. In the example shown in Fig. 5, "before boarding" is defined as the first usage scene, and "during boarding" is defined as the second usage scene, and the utterance intention is estimated for each usage scene. 510 to 513 in the figure correspond to the intention classes shown in Fig. 5B, and the bar graphs in the probability distribution show the probability (likelihood) of each intention class. 520 to 523 in the figure correspond to the "during boarding" intention classes (route instruction, stop instruction, acceleration instruction, deceleration instruction), respectively, and the bar graphs in the probability distribution show the probability (likelihood) of these intention classes.

[0053] In the initial state probability distribution 530 for the first usage scenario, as shown in FIG. 5B , the probabilities of inquiries and greetings are higher than the other probabilities. After the start of the first usage scenario, the user makes an utterance (e.g., "Can I get in?"). Then, the server 110 calculates the probability (likelihood) of the intention class using a machine learning model associated with the first usage scenario, and calculates the probability (likelihood) of the intention class by taking into account the initial state probability distribution 530 (probability distribution 540). The probability distribution 540 for the first usage scenario indicates that the probability (likelihood) of the intention of an inquiry is high. Furthermore, when the user makes a next utterance, the server 110 calculates the probability (likelihood) of the intention class using the same machine learning model, and calculates the probability (likelihood) of the intention class by taking into account the state transition probability. In this way, by calculating the likelihood of the intention using the machine learning model and also taking into account the probability of transitioning from one intention state to the next intention state, the final utterance intention can be estimated by taking into account the likelihood of the intention and the ease of transition in the probability distribution.

[0054] Thereafter, when the usage scene changes, the server 110 calculates the probability (likelihood) of the intent class using the machine learning model associated with the second usage scene, the initial state probability distribution of the second usage scene, and the state transition probability distribution of the second usage scene.

[0055] In this embodiment, the server 110 calculates the likelihood of the intention class according to the following formula: The following intention class is calculated for each usage scene as described above.

number

[0056] (A series of operations for user intention estimation processing) Next, a series of operations in the user intention estimation process in the server 110 will be described with reference to FIG. 6. This process is realized by the control unit 404 executing a program. The machine learning model executed in this series of operations is in a state where it has been trained (inference stage) using training data. In the following explanation, for simplicity, it will be described as if the control unit 404 executes each process, but the corresponding process is executed by each part of the control unit 404 (described above in FIG. 4).

[0057] In S601, the control unit 404 receives a start trigger from the communication device 120. The start trigger indicates, for example, the start of use of a service that controls a vehicle based on a user's utterance. This start trigger is transmitted from the communication device 120 in response to, for example, the user 130 starting an application for using the service on the communication device 120, or uttering a predetermined term indicating the start of use of the service.

[0058] In S602, the control unit 404 identifies a vehicle to be associated with the user 130. The control unit 404 identifies the vehicle 100 that is closest to the user 130, for example, based on the current positions of various vehicles that are constantly known from floating data transmitted from the vehicles and the current position of the user 130. This method is not limited to this, and a vehicle specified by the user on the communication device 120 may be identified as the vehicle to be associated.

[0059] In S603, the control unit 404 acquires information for determining the usage scene from the identified vehicle 100. The information for determining the usage scene includes, for example, information on whether or not an occupant is in the vehicle and information on whether or not the user 130 has boarded within a predetermined time. The information on whether or not an occupant is in the vehicle is obtained, for example, from the seat of the vehicle. The information may further include information on the occupant recognized by an imaging device installed in the vehicle.

[0060] Note that this step may be omitted if this information is included in the floating data transmitted from vehicle 100 to server 110. In this case, control unit 404 simply acquires information about the identified vehicle 100 from the floating data. Although not explicitly shown in FIG. 6, if another user is in vehicle 100, control unit 404 returns the process to S602 and identifies another vehicle.

[0061] In S604, the control unit 404 identifies the user's usage scene for the vehicle 100. The usage scene is identified from the above-mentioned {before boarding, during boarding, after disembarking}. If there is no occupant in the vehicle and the user 130 has not boarded within a predetermined time, the control unit 404 identifies the current usage scene as before boarding. If there is an occupant in the vehicle and the occupant is the user 130, the control unit 404 determines the current usage scene as during boarding. Furthermore, if there is no occupant in the vehicle and the user 130 boards within a predetermined time, the control unit 404 identifies the current usage scene as after disembarking.

[0062] In S605, the control unit 404 selects a machine learning model according to the identified usage scenario. Each machine learning model is trained using different training data for each corresponding usage scenario. For example, the training data is assigned a label indicating the correct utterance intention for the user's utterance information, and is further assigned a label indicating the corresponding usage scenario. In other words, when the control unit 404 trains a machine learning model for each usage scenario, only the training data for the corresponding usage scenario can be input to the machine learning model for training.

[0063] In S606, the control unit 404 determines whether the user's utterance information has been acquired. If the control unit 404 has acquired the user's utterance information from the communication device 120, the control unit 404 proceeds to S607; if not, the control unit 404 returns to S606 and waits for the acquisition of the user's utterance information.

[0064] In S607, the control unit 404 estimates the intention of the utterance using the machine learning model selected in S605. Specifically, the control unit 404 performs a calculation according to the above-mentioned formula to estimate the output intention class argmax b(c t ) is calculated. At this time, if the user's utterance information is utterance information immediately after a new usage scene is identified, the calculation for t=1 is performed, and if not, the calculation for t≧2 is performed.

[0065] In S608, the control unit 404 transmits a control command to the vehicle according to the intention of the utterance. For example, as described above, the control unit 40 identifies necessary information from the utterance information to identify specific instruction content based on the estimated intention of the utterance. For example, if the intention of the user's utterance information is a request to be picked up, the control unit 40 identifies information such as where and what time the vehicle 100 will be picked up. The voice information processing unit 417 may further include slot fulfillment processing. Furthermore, the control unit 404 transmits a control command to the vehicle 100 to control the operation of the vehicle 100 based on the recognized utterance content. For example, if information such as where and what time the vehicle 100 will be picked up is identified from the user's utterance information, the control unit 404 identifies a route based on the current positions of the user and vehicle, map information, etc., and transmits a control command to the vehicle 100 to drive along that route.

[0066] In S609, the control unit 404 determines whether the user operation has ended. The control unit 404 determines, for example, whether information indicating the end has been received from the communication device 120. The information indicating the end is transmitted from the communication device 120 in response to, for example, the user 130 uttering a predetermined term indicating the end of use of the service at the communication device 120. If the control unit 404 determines that the user operation has ended, it ends this series of processes; if not, it returns to S603 and repeats the processes from S603 onwards.

[0067] In the above-described embodiment, the server 110 identifies a usage scene based on information from the vehicle 100. However, the server 110 may identify a usage scene based on other information. For example, the server 110 may identify a usage scene based on information from the communication device 120. For example, the communication device 120 may identify a usage scene by receiving a start trigger transmitted from the communication device 120 and information indicating the occurrence of proximity to a vehicle. As described above, the usage trigger is transmitted from the communication device 120 in response to, for example, the user 130 starting an application for using the above-described service on the communication device 120 or uttering a predetermined term indicating the start of use of the service. Furthermore, for example, the user may bring the communication device 120 close to the vehicle 100 when getting in and out of the vehicle 100. When the communication device 120 detects proximity to the vehicle via close-proximity wireless communication or the like, the communication device 120 transmits information indicating the occurrence of proximity to the vehicle. For example, after receiving the start trigger, if the server 110 has not received information indicating the occurrence of proximity, the server 110 may identify the usage scene as "before boarding," and if the server 110 subsequently receives information indicating the occurrence of proximity, the server 110 may identify the usage scene as "during boarding." Furthermore, if the server 110 receives information indicating the occurrence of proximity, the server 110 may identify the usage scene as "after disembarking." Note that instead of the server 110 identifying the usage scene, the communication device 120 may identify the usage scene and transmit the identified usage scene to the server 110 in response to a change in the usage scene.

[0068] As described above, in the above embodiment, an information processing device capable of controlling a vehicle based on a user's spoken instructions is configured to first identify which of multiple usage scenarios for vehicle use corresponds to the target user. After identifying the usage scenario, a different machine learning model is selected depending on the identified usage scenario for the target user, and the selected machine learning model is used to estimate the target user's utterance intention. In this way, when controlling a moving object based on utterances, it is possible to provide classification of utterance intentions using a model constructed through smaller-scale learning.

[0069] (Variation) Modifications of the present invention will now be described. In the above embodiment, an example has been described in which the utterance intention estimation process is executed in the server 110. However, the above-mentioned utterance intention estimation process can also be executed on the vehicle side. In this case, as shown in FIG. 7, an information processing system 900 is configured with a vehicle 710 and a communication device 120. User utterance information is transmitted from the communication device 120 to the vehicle 710. The configuration of the vehicle 710 may be the same as that of the vehicle 100, except that the control unit 30 is capable of executing the utterance intention estimation process. The control unit 30 of the vehicle 710 operates as a control device for the vehicle 710 and executes a stored program to execute the above-mentioned utterance intention estimation process. In the series of operations shown in FIG. 6, communication between the server and the vehicle may be performed inside the vehicle (for example, inside the control unit 30). Other processes can be executed in the same way as in the server.

[0070] In this way, in a control device capable of controlling a vehicle based on user's spoken instructions, first, it is determined which of multiple usage scenarios for vehicle use corresponds to the target user. After determining the usage scenario, a different machine learning model is selected according to the identified usage scenario of the target user, and the selected machine learning model is used to estimate the target user's utterance intention. In this way, when controlling a moving object through speech, it is possible to provide classification of utterance intention using a model constructed through smaller-scale learning.

[0071] <Summary of the embodiment> 1. The information processing device (e.g., 110) of the above embodiment is An information processing device capable of controlling a moving object (e.g., 120) based on a user's spoken instruction, Identification means (e.g., 414) for identifying which of a plurality of usage scenarios when using a mobile object the target user is in; Acquisition means (e.g., 413) for acquiring speech information of a target user; A selection means (e.g., 415) for selecting different machine learning models according to the usage scenario of the identified target user; and an estimation means (e.g., 416) for estimating the intention of the target user's utterance using the selected machine learning model.

[0072] According to this embodiment, in controlling a moving object by speech, it is possible to provide classification of speech intentions using a model constructed through smaller-scale learning.

[0073] 2. In the information processing device of the above embodiment, The machine learning model estimates different intent classes (for example, 501, 502) for each usage scenario to which the machine learning model is associated.

[0074] According to this embodiment, it is possible to use a machine learning model that is suitable for each scene.

[0075] 3. In the information processing device of the above embodiment, The estimation means estimates the intention of the target user using a machine learning model that outputs likelihoods for only some of all intention classes that can be associated with a plurality of usage scenes.

[0076] According to this embodiment, a model that outputs a small number of intention classes that are appropriate for each scene can be used, which makes it possible to easily learn the model.

[0077] 4. In the information processing device of the above embodiment, The estimation means estimates the intention of the target user's utterance by adding a calculation using the initial state probability distribution set for the intention class as a prior distribution to the output of the selected machine learning model.

[0078] This embodiment makes it possible to reflect common sense about speech that depends on the context of the scene (it is unlikely that the first word in a communication is intended to be a denial, or that it is intended to be a greeting during a conversation).

[0079] 5. In the information processing device of the above embodiment, The initial state probability distribution set as the prior distribution is determined separately for each usage scenario.

[0080] According to this embodiment, it is possible to reflect common sense about speech for each scene.

[0081] 6. In the information processing device of the above embodiment, The estimation means estimates the intention of the target user's utterance by adding a calculation using the state transition probability distribution between the intention classes to the output of the selected machine learning model.

[0082] According to this embodiment, it is possible to estimate an intention taking into consideration the order of transitions of intentions in an actual dialogue.

[0083] 7. In the information processing device of the above embodiment, The state transition probability distribution is determined separately for each scene.

[0084] This embodiment allows for separate definition of state transition probabilities between scene intent classes.

[0085] 8. In the information processing device of the above embodiment, When estimating the intention of the utterance at time t, the estimation means estimates the intention of the target user's utterance by taking into account the output of the selected machine learning model and the estimation result estimated for the utterance immediately before the utterance at time t.

[0086] According to this embodiment, the probability distribution for utterances up to time t-1 can be recursively considered.

[0087] 9. In the information processing device of the above embodiment, Each machine learning model is trained using different training data for each corresponding usage scenario, and the training data includes labels indicating the usage scenario.

[0088] According to this embodiment, the machine learning model can be trained using training data that is scaled down for each usage scenario, and the training data can be easily assigned to a usage scenario according to its label and used for training.

[0089] 10. In the information processing device of the above embodiment, The usage scene of the target user is identified based on information from a mobile object associated with the target user. According to this embodiment, the usage scene can be determined with high accuracy by determining the usage scene based on information provided by the mobile object that is the target of use.

[0090] 11. In the above embodiment, the control device (e.g., 30) of the moving body (e.g., 710) A control device for a moving object that can be controlled based on a user's spoken instruction, Identification means (e.g., 30, S604) for identifying which of a plurality of usage scenarios when using a mobile object the target user is in; Acquiring means (e.g., 30, S606) for acquiring speech information of a target user; A selection means (e.g., 30, S605) for selecting different machine learning models according to the usage scenario of the identified target user; and an estimation means (e.g., 30, S607) for estimating the intention of the target user's utterance using the selected machine learning model.

[0091] According to this embodiment, in controlling a moving object by speech, it is possible to provide classification of speech intentions using a model constructed through smaller-scale learning.

[0092] The invention is not limited to the above-described embodiment, and various modifications and variations are possible within the scope of the gist of the invention. [Explanation of symbols]

[0093] 100... vehicle, 110... server, 120... communication device, 404... control unit, 413... user data acquisition unit, 414... scene identification unit, 415... model selection unit, 416... speech intention estimation unit, 417... speech information processing unit, 418... vehicle control unit

Claims

1. An information processing device capable of controlling a moving object based on a user's spoken instruction, an identification means for identifying which usage scene of a target user is associated with a plurality of usage scenes related to instructions from the user when using a mobile object; an acquisition means for acquiring utterance information of the target user; a selection means for selecting a different machine learning model depending on the usage scene of the identified target user; an estimation means for estimating the intention of the target user's utterance using the selected machine learning model; The information processing device is characterized in that the identification means uses information obtained from a mobile body associated with the target user as to whether the target user is riding in the mobile body and information as to whether the target user has boarded or disembarked from the mobile body within a predetermined period of time to identify whether the scene is before the target user uses the mobile body, while using the mobile body, or after the target user has used the mobile body.

2. An information processing device capable of controlling a moving object based on a user's spoken instruction, an identification means for identifying whether a target user's usage scene is a scene before the target user uses the mobile object, a scene while the target user uses the mobile object, or a scene after the target user uses the mobile object, among a plurality of usage scenes when the target user uses the mobile object; an acquisition means for acquiring utterance information of the target user; a selection means for selecting a different machine learning model depending on the usage scene of the identified target user; an estimation means for estimating the intention of the target user's utterance using the selected machine learning model; The information processing device is characterized in that the identification means uses information obtained from a mobile body closest to the target user's current location or specified by the target user as to whether the target user is riding in the mobile body, and information as to whether the target user has boarded or disembarked from the mobile body within a specified period of time, to identify whether the scene is before the target user uses the mobile body, while the target user is using the mobile body, or after the target user has used the mobile body.

3. The information processing device according to claim 1 , wherein the machine learning model estimates a different intent class for each usage scene associated with the machine learning model.

4. 4. The information processing device according to claim 3, wherein the estimation means estimates the intention of the target user using the machine learning model that outputs likelihoods for only some of all intention classes that can be associated with the plurality of usage scenes.

5. 5. The information processing device according to claim 1, wherein the estimation means estimates the intention of the utterance of the target user by using an output of the selected machine learning model and an initial state probability distribution set for an intention class as a prior distribution.

6. 6. The information processing apparatus according to claim 5, wherein the initial state probability distribution set as the prior distribution is determined separately for each usage scenario.

7. 7. The information processing device according to claim 1, wherein the estimation means estimates the intention of the utterance of the target user by using an output of the selected machine learning model and a state transition probability distribution between intention classes.

8. 8. The information processing apparatus according to claim 7, wherein the state transition probability distribution is determined separately for each scene.

9. The information processing device according to any one of claims 1 to 8, characterized in that, when estimating the intention of the utterance at time t, the estimation means estimates the intention of the utterance of the target user using the output of the selected machine learning model and an estimation result estimated for an utterance immediately before the utterance at time t.

10. 10. The information processing device according to claim 1, wherein each of the machine learning models is trained using different training data for each corresponding usage scenario, and the training data includes a label indicating the usage scenario.

11. 11. The information processing device according to claim 1, wherein the identification unit identifies which usage scene of the target user is based on information from a moving object associated with the target user.

12. An information processing method for an information processing device capable of controlling a moving object based on a user's spoken instruction, comprising: an identification step of identifying which usage scene of the target user is among a plurality of usage scenes related to instructions from the user when using the mobile object; an acquisition step of acquiring utterance information of the target user; a selection step of selecting a different machine learning model depending on the usage scenario of the identified target user; an estimation step of estimating the intention of the target user's utterance using the selected machine learning model; The identification process is an information processing method characterized by using information obtained from a mobile body associated with the target user as to whether the target user is riding in the mobile body, and information as to whether the target user has boarded or disembarked from the mobile body within a predetermined period of time, to identify whether the scene is before the target user uses the mobile body, while the target user is using the mobile body, or after the target user has used the mobile body.

13. A control device for a moving object that can be controlled based on a user's spoken instruction, an identification means for identifying which usage scene of a target user is associated with a plurality of usage scenes related to instructions from the user when using a mobile object; an acquisition means for acquiring utterance information of the target user; a selection means for selecting a different machine learning model depending on the usage scene of the identified target user; an estimation means for estimating the intention of the target user's utterance using the selected machine learning model; The control device is characterized in that the identification means uses information on whether the target user is riding in the mobile vehicle and information on whether the target user has boarded or disembarked from the mobile vehicle within a predetermined period of time to identify whether the scene is before the target user uses the mobile vehicle, while using the mobile vehicle, or after the target user has used the mobile vehicle.

14. A method for controlling a moving object that can be controlled based on a user's spoken instruction, comprising: an identification step of identifying which usage scene of the target user is among a plurality of usage scenes related to instructions from the user when using the mobile object; an acquisition step of acquiring utterance information of the target user; a selection step of selecting a different machine learning model depending on the usage scenario of the identified target user; an estimation step of estimating the intention of the target user's utterance using the selected machine learning model; The identification process is a control method characterized by using information on whether the target user is riding in the moving body and information on whether the target user has boarded or disembarked from the moving body within a specified period of time to identify whether the scene is before the target user uses the moving body, while using the moving body, or after using the moving body.

15. A program for causing a computer to function as each of the means of the information processing device according to any one of claims 1 to 11.

16. A program for causing a computer to function as each of the means of the control device according to claim 13.

Citation Information

Patent Citations

  • Driver intention recognition method based on probability correction

    CN111717217A

  • Customer intention identification method and device based on artificial intelligence, and computer equipment

    CN112036550A

  • Request estimation apparatus

    JP2000020090A

  • Intention estimation device and model learning method

    JP2015230384A

  • Estimation device, estimation method, and estimation program

    JP2018045332A