Processing system, learning method, processing method, and program

The processing system addresses the challenge of selecting appropriate actions for mobile objects by using a learned model to process voice, peripheral, and movement information, resulting in improved control and interaction with enhanced system simplicity.

WO2025120851A1PCT designated stage expired Publication Date: 2025-06-12HONDA MOTOR CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/044066
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-08
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing systems struggle to select appropriate actions in various usage scenarios for mobile objects, leading to ineffective control and interaction.

Method used

A processing system that acquires voice information, peripheral information, and movement information, and uses a learned model to determine the appropriate action for a mobile object based on feature information derived from these inputs.

Benefits of technology

Enables the selection of appropriate actions for mobile objects, improving control and interaction by considering multiple information sources, thereby simplifying system construction and enhancing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023044066_12062025_PF_FP_ABST
    Figure JP2023044066_12062025_PF_FP_ABST
Patent Text Reader

Abstract

This processing system comprises: an acquisition unit that acquires speech information relating to speech uttered by a user, surrounding information relating to a situation surrounding an object to be controlled, and movement information relating to the movement of the object; and a processing unit that, when input information based on the speech information, the surrounding information, and the movement information is inputted, determines the behavior of the object on the basis of behavior information relating to the behavior of the object corresponding to the input information, said behavior information being obtained by inputting the input information to a trained model that outputs the behavior information.
Need to check novelty before this filing date? Find Prior Art

Description

Processing system, learning method, processing method, and program

[0001] The present invention relates to a processing system, a learning method, a processing method, and a program.

[0002] Conventionally, an information processing device has been disclosed that includes an identification means for identifying which of a plurality of usage scenarios when using a mobile object a target user's usage scenario is, an acquisition means for acquiring speech information of the target user, a selection means for selecting a different machine learning model depending on the identified usage scenario of the target user, and an estimation means for estimating the intention of the target user's speech using the selected machine learning model (see, for example, Patent Document 1).

[0003] Japanese Patent Application Laid-Open No. 2022-155107

[0004] However, the above techniques sometimes fail to select an appropriate action.

[0005] The present invention has been made in consideration of the above circumstances, and one of its objects is to provide a processing system, a learning method, a processing method, and a program that are capable of selecting an appropriate action.

[0006] The processing system, learning method, processing method, and program according to the present invention employ the following configuration: (1): A processing system for a moving object according to one aspect of the present invention includes an acquisition unit that acquires voice information related to a voice uttered by a user, peripheral information related to a situation around a target to be controlled, and movement information related to the movement of the target, and a processing unit that, when input information based on the voice information, the peripheral information, and the movement information is input, determines the behavior of the target based on the behavior information obtained by inputting the input information into a trained model that outputs behavior information related to the behavior of the target according to the input information.

[0007] (2) In the aspect (1) above, the input information is first feature information indicating features based on the voice information, the surrounding information, and the movement information.

[0008] (3): In the aspect (2) above, the processing unit generates second feature information indicating features of the audio information, third feature information indicating features of the peripheral information, and fourth feature information indicating features of the movement information, and derives the first feature information based on the generated second feature information, third feature information, and fourth feature information.

[0009] (4): In any of the aspects of (3) above, the input information, the voice information, the surrounding information, and the movement information are text data, and the first feature information, the second feature information, the third feature information, and the fourth feature information are data obtained by vectorizing the text data.

[0010] (5): In the above aspect (1), the trained model is a model in which training information has been trained, and the training information is information in which input information based on the voice information, the surrounding information, and the movement information is associated with behavioral information corresponding to the input information, and the trained model is a model trained to output the behavioral information associated with the input information when the input information is input.

[0011] (6): In the aspect (1) above, the acquisition unit further acquires interaction information between the user and the target, and the processing unit, upon receiving input information based on the voice information, the surrounding information, the movement information, and the interaction information, determines the target's behavior based on the behavior information obtained by inputting the input information into a trained model that outputs behavior information regarding the target's behavior according to the input information.

[0012] (7) In any of the above aspects (1) to (6), the processing unit controls the moving object that is the object of control based on the behavior information.

[0013] (8): In any of the above aspects (1) to (6), the surrounding information includes information indicating the position of an object detected by the target moving body and the type of the object, and the movement information includes information indicating the position of the moving body, information indicating the orientation of the moving body, and information indicating the speed of the moving body.

[0014] (9): In any of the above aspects (1) to (6), the processing unit estimates a speech intention from audio information related to the speech spoken by the user, and when input information based on the speech intention, the peripheral information, and the movement information, or input information based on the speech intention, the audio information, the peripheral information, and the movement information, inputs the input information into a trained model that outputs behavioral information related to the behavior of the target according to the input information, and determines the behavior of the target based on the behavioral information obtained by inputting the input information.

[0015] (10): In any of the above aspects (1) to (6), the behavior of the target mobile object includes at least one of movement and speech of the mobile object.

[0016] (11): In any of the above aspects (1) to (6), the behavior of the target moving body includes at least one of behaviors that change the speed, moving direction, moving distance, stopping position, and stopping time of the moving body.

[0017] (12): In any of the above aspects (1) to (6), the utterances of the target mobile object include at least one of an inquiry to the user, a suggestion to the user, a reconsideration, a response, a thank you, an apology, an alternative, an arrival location, an arrival time, information about the mobile object and its surrounding environment, and a confirmation utterance to the user.

[0018] (13): In another aspect of the present invention, a learning method is provided in which a computer acquires learning information that associates audio information about the voice spoken by a user, peripheral information about the surrounding situation of a controlled object, and movement information about the movement of the object with information about the object's behavior, and uses the acquired learning information to train a learning model so that, when input information based on the audio information, the peripheral information, and the movement information is input, it outputs information about the object's behavior that is associated with the audio information, the peripheral information, and the movement information, thereby generating a trained model.

[0019] (14): Another aspect of the present invention is a processing method in which a computer acquires voice information about a voice spoken by a user, peripheral information about the situation around a target to be controlled, and movement information about the movement of the target, and when input information based on the voice information, the peripheral information, and the movement information is input, the computer inputs the input information into a trained model that outputs behavioral information about the behavior of the target according to the input information, and determines the behavior of the target based on the behavioral information obtained by inputting the input information.

[0020] (15): A program according to another aspect of the present invention causes a computer to acquire voice information about a voice spoken by a user, peripheral information about the situation around a target of control, and movement information about the movement of the target, and, when input information based on the voice information, the peripheral information, and the movement information is input, causes a trained model that outputs behavioral information about the behavior of the target according to the input information to determine the behavior of the target based on the behavioral information obtained by inputting the input information.

[0021] According to the above aspects (1) to (15), it is possible to select an appropriate action.

[0022] According to the above aspect (7), the moving body can be controlled more appropriately.

[0023] According to the above aspect (8), by utilizing the surrounding environment of the moving object and the state of the moving object, it is possible to select a more appropriate action.

[0024] FIG. 1 is a diagram showing an example of the configuration of a processing system 1 including a mobile object; FIG. 2 is a diagram showing an example of the functional configuration of a mobile object; FIG. 3 is a diagram showing an example of the functional configuration of a server device; FIG. 4 is a diagram for explaining an overview of processing; FIG. 5 is a diagram for explaining processing for generating input information; FIG. 6 is a diagram showing an example of the contents of behavior correspondence information; FIG. 7 is a diagram showing an example of control information; FIG. 8 is a diagram showing an example of user behavior and behavior of a mobile object; FIG. 9 is a flowchart showing an example of the flow of processing executed by a server device; FIG. 10 is a diagram for explaining utterance intention; FIG. 11 is a diagram showing an example of learning information.

[0025] Hereinafter, embodiments of a processing system, a learning method, a processing method, and a program according to the present invention will be described with reference to the drawings.

[0026] The processing system of the present invention determines the behavior of an object and controls the object to perform a predetermined behavior. The object may be a moving object (an automobile, a motorcycle, a light vehicle, a micromobility vehicle, or an aircraft), a robot (e.g., an interactive robot), or a mechanically operated object (e.g., an interactive robot arm or an interactive robot hand). In the following description, the object is assumed to be a small moving object, such as a small vehicle.

[0027] 1 is a diagram showing an example of the configuration of a processing system 1 including a mobile object 100. The processing system 1 includes, for example, one or more terminal devices 2, one or more mobile objects 100, a server device 300, and a learning device 500. These communicate with each other via, for example, a network NW. The network NW is, for example, any network such as a LAN, a WAN, or an internet line.

[0028] [Terminal Device] The terminal device 2 is, for example, a computer device such as a smartphone or a tablet terminal. The terminal device 2 requests the provision of authorization to use the moving object 100 and obtains information indicating that the use has been permitted, for example, based on a user operation. The terminal device 2 provides the server device 300 with the voice input by the user into the microphone of the terminal device 2.

[0029] [Mobile Body] The mobile body 100 can autonomously move in areas where vehicles can travel and areas where pedestrians can move. For example, the mobile body 100 can travel in areas where vehicles cannot pass. The mobile body 100 can move in areas where pedestrians can pass, such as roadways and sidewalks. For example, the mobile body 100 may be used in indoor or outdoor facilities or private land, such as shopping centers, airports, parks, and theme parks, and can move in areas where pedestrians can pass.

[0030] 2 is a diagram showing an example of the functional configuration of the mobile object 100. The mobile object 100 includes a first wheel 120, a first motor 122, a second wheel 130, a second motor 132, a battery 134, a braking device 136, a steering device 138, a detection unit 180, a communication unit 182, an HMI 184, a GNSS (Global Navigation Satellite System) unit 186, and a control device 200. Some of the above functional configurations may be omitted.

[0031] The first motor 122 and the second motor 132 are operated by power supplied to a battery 134. The first motor 122 drives the first wheel 120, and the second motor 132 drives the second wheel 130. The first motor 122 may be an in-wheel motor provided in the wheel of the first wheel 120, and the second motor 132 may be an in-wheel motor provided in the wheel of the second wheel 130.

[0032] The brake device 136 outputs a brake torque to each wheel based on an instruction from the control device 200. The steering device 138 includes an electric motor. The electric motor applies a force to a rack and pinion mechanism based on an instruction from the control device 200, for example, to change the direction of the first wheel 120 or the second wheel 130, thereby changing the course of the mobile object 100.

[0033] The detection unit 180 is, for example, one or more cameras that capture images of the scenery in the target area. The detection unit 180 captures images of the scenery at predetermined intervals. Instead of (or in addition to) a camera, the detection unit 180 may be a sensor capable of detecting objects, such as a radar or a lidar (Light Detection and Racing).

[0034] The communication unit 182 is a communication interface for communicating with the terminal device 2 or the server device 300 .

[0035] The HMI 184 presents various information to the user and receives input operations from the user, and includes various display devices, a speaker, a microphone, a buzzer, a touch panel, switches, keys, and the like.

[0036] The GNSS unit 186 includes a position determination device. The position determination device is a device that determines the position of the mobile body 100. The position determination device includes, for example, a GNSS receiver that determines the position of the mobile body 100 based on signals received from GNSS (Global Navigation Satellite System) satellites.

[0037] In addition to the above configuration, the moving body 100 may be provided with a sensor or the like that detects the state of the moving body 100, such as the speed or direction of the moving body 100. Information held by the moving body 100 may be provided to the server device 300. For example, the moving body 100 provides the server device 300 with information necessary for processing by the server device 300, such as the direction in which the moving body 100 is facing and the speed of the moving body 100.

[0038] [Control Device] The control device 200 includes, for example, a recognition unit 220 and a control unit 240. The recognition unit 220 and the control unit 240 are realized by, for example, a hardware processor such as a CPU (Central Processing Unit) executing a program (software). Some or all of these components may be realized by hardware (including circuitry) such as an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), a GPU (Graphics Processing Unit), or an SOC (System On Chip), or may be realized by a combination of software and hardware. The program may be stored in advance in a storage device (a storage device with a non-transitory storage medium) such as an HDD (Hard Disk Drive) or flash memory, or may be stored in a removable storage medium (a non-transitory storage medium) such as a DVD or CD-ROM, and installed by inserting the storage medium into a drive device.

[0039] The storage unit 260 is realized by a storage device such as a HDD, flash memory, RAM (Random Access Memory), etc. The storage unit 260 stores programs executed by the control device 200, information used for various processes, etc. The storage unit 260 may also store information on areas in which the mobile object 100 can move, information on landmarks such as buildings, information on stores, etc.

[0040] The recognition unit 220 recognizes the type, position, speed, acceleration, and other status of objects around the mobile object 100 based on, for example, images captured by the detection unit 180. Objects include traffic participants and obstacles in facilities or on roads. The recognition unit 220 may recognize and track the user of the mobile object 100. For example, the recognition unit 220 tracks the user based on an image (e.g., a facial image or an overall image of the user) captured by the user registered via the HMI 184 when the user uses the mobile object 100, or a facial image of the user (or feature values ​​obtained from the user's facial image) provided by the terminal device 2 or the server device 300. If the detection unit 180 is a radar device or LIDAR, the recognition unit 220 may recognize the situation around the mobile object 100 using the detection results of the radar device or LIDAR instead of (or in addition to) images.

[0041] The control unit 240 generates a trajectory to a destination specified by the user. The destination may be the location of a facility, a predetermined landmark, a predetermined sign, or the like. The trajectory is a trajectory that allows the user to reach the destination reasonably without interfering with objects. For example, the distance to the destination, the time required to reach the destination, the ease of travel of the trajectory, and the like are scored, and a trajectory is derived in which each score and the combined score of the scores are equal to or greater than a threshold.

[0042] The control unit 240 controls the motors (first motor 122, second motor 132), the brake device 136, and the steering device 138 so that the moving body 100 travels along a trajectory that satisfies a preset standard. The control unit 240 controls the HMI to output sound from the HMI.

[0043] 3 is a diagram showing an example of the functional configuration of a server device. Server device 300 includes, for example, an information management unit 310, an information processing unit 320, and an information providing unit 330. Information processing unit 320, or the combined functional configuration of information processing unit 320 and information providing unit 330, is an example of a "processing unit."

[0044] The information management unit 310, the information processing unit 320, and the information providing unit 330 are realized, for example, by a hardware processor such as a CPU executing a program (software). Some or all of these components may be realized by hardware (including circuitry) such as an LSI, ASIC, FPGA, GPU, or SOC, or may be realized by a combination of software and hardware. The program may be stored in advance in a storage device such as an HDD or flash memory (a storage device with a non-transitory storage medium), or may be stored in a removable storage medium (a non-transitory storage medium) such as a DVD or CD-ROM, and installed by inserting the storage medium into a drive device.

[0045] The storage unit 360 is realized by a storage device such as a HDD, flash memory, or RAM (Random Access Memory). The storage unit 260 stores a trained model 362, behavior correspondence information 364, control information 366, and the like. The trained model 362 is, for example, a decision tree algorithm such as LightGBM or other machine learning models. Details of each piece of information will be described later.

[0046] The information management unit 310 acquires voice information related to the voice spoken by the user, peripheral information related to the situation around the target of control (e.g., the mobile object 100), and movement information related to the movement of the target. The information management unit 310 acquires from the mobile object 100 information (voice information, peripheral information, movement information) acquired by the control device of the mobile object 100.

[0047] The information processing unit 320 determines the behavior of the target based on behavior information obtained by inputting input information based on audio information, surrounding information, and movement information into the trained model 362.

[0048] The information providing unit 330 controls the mobile object 100 so that the mobile object 100 behaves in accordance with the determined action. The information providing unit 330 provides the mobile object 100 with instruction information that instructs the operation of the mobile object 100. The instruction information may be information that defines each action (e.g., a control scenario) or may be detailed control values.

[0049] A part or all of the functional configuration of the server device 300 described above may be installed in the mobile object 100. A part of the functional configuration included in the mobile object 100 described above may be installed in the server device 300.

[0050] [Summary] When the server device 300 receives input information based on voice information, surrounding information, and movement information, the server device 300 determines the behavior of the target based on the behavior information obtained by inputting the input information into the trained model 362, which outputs behavior information regarding the target's behavior according to the input information.

[0051] The server device 300 may input input information, which is the above information plus interaction information, to the trained model 362, and determine the target behavior based on the behavior information obtained by inputting the input information. In the following example, it is assumed that interaction information is used.

[0052] FIG. 4 is a diagram for explaining an overview of the processing. The server device 300 extracts feature information indicating these features based on the voice information, surrounding information, movement information, and interaction information, and selects an action using information based on the extracted feature information as input information. The server device 300 provides information on the selected action to the mobile object 100. The action is one or both of movement and speech. The mobile object 100 moves and speaks based on the control information of the server device 300.

[0053] The voice information is, for example, information on a user's utterance obtained by voice recognition. The voice information is, for example, text data of the utterance. The information processing unit 320 analyzes the obtained utterance and converts it into text data to generate the voice information. The voice information is information on the user's utterance, and may be provided from the mobile object 100 or from the terminal device 2.

[0054] The surrounding information is, for example, the detection results of the surrounding environment of the moving body 100, such as the detection results of objects detected or recognized by the moving body 100 and the detection results of landmarks. The detection results include the type of object (e.g., a person, an object, etc.), the number, color, coordinates, distance from the moving body 100, and confidence level. The surrounding information is, for example, text data obtained from the detection results of the detection unit 180. The text data obtained from the detection results is, for example, text data indicating the surrounding environment, such as text indicating the type of object (e.g., "person" for a person) or the color of the object (e.g., "black" if the object is black). For example, if there is a person five meters ahead of the moving body 100, the surrounding information is text data indicating that "a person is located five meters ahead of the moving body 100."

[0055] The information processing unit 320 acquires the detection result from the detection unit 180 or the recognition result from the recognition unit 220, analyzes the acquired detection or recognition result, and converts the detection or recognition result into text data. The information processing unit 320 uses a predetermined algorithm to identify the type of object included in the detection result, its position, distance from the mobile object 100, confidence level, etc., and generates surrounding information as text data based on the identified results.

[0056] The movement information includes information related to the movement of the mobile object 100, such as the position information of the mobile object 100, the direction of the mobile object 100, the speed of the mobile object 100, the distance from the current location of the mobile object 100 to the destination, the difference in coordinates between the location of the mobile object 100 at the previous user utterance point and the current location of the mobile object 100, and traffic conditions (traffic jam, one-way street, road closure, no movement). The movement information is, for example, text data related to the movement of the mobile object 100. The movement information is, for example, text data indicating that the mobile object 100 is traveling at 5 km / h, that the distance to the destination is 20 m, etc.

[0057] The traffic conditions are information derived by the server device 300 based on the location of the mobile object 100, map information held by the server device 300, and congestion information. For example, the server device 300 refers to the map information and congestion information corresponding to the location of the mobile object 100, and recognizes whether or not there is congestion, road signs included in the map information, road types, and the like, to acquire the traffic conditions. The traffic conditions may be treated as information included in the surrounding information (surrounding information detected by the detection unit 180 or surrounding information recognized by the recognition unit 220) instead of the movement information.

[0058] The information processing unit 320 acquires information relating to the movement of the mobile object 100, analyzes the acquired information relating to the movement, and converts the information relating to the movement into text data. The information processing unit 320 converts information such as speed information, distance, coordinate difference, and traffic conditions into text data based on a predetermined algorithm.

[0059] The interaction information includes, for example, a sentence spoken by the mobile object 100, the elapsed time from the previous user utterance to the current utterance, the number of utterances made by the mobile object 100 or the user, the number of times each type of action has been performed, the names of landmarks passed, etc. The interaction information is not limited to the above and may be any information relating to the interaction between the mobile object 100 and the user.

[0060] The information processing unit 320 acquires the interaction information, analyzes the acquired interaction information, and converts the interaction information into text data. The information processing unit 320 converts information such as the utterance of the mobile object 100, the elapsed time, and the number of utterances into text data based on a predetermined algorithm.

[0061] [Generation of Input Information] FIG. 5 is a diagram for explaining the process of generating input information. The information processing unit 320 converts each of the above-described text data into feature information (feature vectors). The feature information is obtained by vectorizing the text data using a distributed representation technique. For example, the information processing unit 320 generates a vector of a first predetermined dimension such as [0, 1, 0, 0] from the text data, and compresses the generated vector of a second predetermined dimension. For example, the information processing unit 320 performs principal component analysis (PCA) to perform compression.

[0062] The values ​​forming the vector are not limited to those mentioned above, and may be any value between minus 1 and plus 1. The information processing unit 320 generates, for example, a feature vector obtained by vectorizing the text data of the audio information (second feature information), a feature vector obtained by vectorizing the text data of the peripheral information (third feature information), and a feature vector obtained by vectorizing the text data of the movement information (fourth feature information).

[0063] The information processing unit 320 concatenates each piece of feature information (feature vector) to generate input information. For example, the concatenation order is a predetermined order. The input information is an example of "first feature information indicating features based on audio information, peripheral information, and movement information." When the input information is input, the trained model 362 outputs output information. The output information is, for example, identification information for identifying the behavior of the behavior correspondence information 364.

[0064] FIG. 6 is a diagram showing an example of the contents of the behavior correspondence information 364. The behavior correspondence information 364 is information in which identification information and behaviors are associated with each other. In this embodiment, the behavior correspondence information 364 defines behaviors for identification information 1 to 30. The following is an example, and other examples may be included. The behavior may be only speaking, only moving, or a combination of these. The processing system 1, for example, controls the speaking of the mobile object 100 or moves the mobile object 100.

[0065] "Control of speech" includes, for example, at least one of control of speech to the user, inquiries to the user, suggestions to the user, asking for clarification, responses, thanks, apologies, alternatives, arrival location, arrival time, information about the mobile body 100 and the surrounding environment of the mobile body 100, and confirmation to the user.

[0066] The "control of movement" includes at least one of control to change the speed, movement direction, movement distance, stopping position, and stopping time of the moving body 100.

[0067] 1. Move to the user's vicinity 2. Move to the user's vicinity, no speech 3. Move to the vicinity of a specified location 4. Move to a specified location 5. Move a short distance 6. Stop on the spot 7. Pause at a nearby landmark 8. Confirm whether to move to the specified location 9. Inquire about the stopping location 10. Inquire about the details of the stopping location

[0068] 11. Suggest an alternative stopping location 12. Speak the reason for the failure to identify a stopping location 13. Explain that stopping is prohibited 14. Speak nearby landmarks 15. Speak the arrival time 16. Provide traffic information 17. Speak the remaining distance to the stopping location 18. Speak the speed 19. Check the color of the user's clothing 20. Check the color of an object.

[0069] 21. Affirmative response 22. Ask again 23. Thank you 24. Apology 25. Register user information 26. Register user information, no speech 27. Register stopping location in slot 28. Provide vehicle specifications 29. Chat 30. No action

[0070] The processing system 1 (one or both of the server device 300 and the control device 200 of the mobile object 100) controls the mobile object 100 based on the above actions and the control information 366. FIG. 7 is a diagram showing an example of the control information 366. The control information 366 is information that defines a control scenario and a processing procedure for each action. The control information 366 may define only utterances, only actions (movements), or a combination of utterances and actions. For example, when the identification information "1. Move to the vicinity of the user" in the action correspondence information 364 is identified, processes 1, 2, 3, etc. defined in the control information 366 are executed. Process 1 is for the processing system 1 to recognize the user. Process 2 is for the processing system 1 to generate a trajectory to the vicinity of the user. Process 3 is for the processing system 1 to speak that the mobile object 100 will move to the vicinity of the user. In this manner, the processing system 1 controls the mobile object 100 based on the control information 366.

[0071] FIG. 8 is a diagram illustrating an example of the behavior of user U and the behavior of mobile object 100. Assume that at time T, user U utters, for example, "Stop at a location near the vending machine." The content of the utterance (or voice information or information that is the source of the voice information) is provided to server device 300. Furthermore, surrounding information, movement information, and interaction information (or information that is the source of the surrounding information, movement information, and interaction information) are provided from mobile object 100 to server device 300. The interaction information may be stored in server device 300. Based on the acquired information, server device 300 determines the behavior of mobile object 100 and controls mobile object 100.

[0072] At time T+1, server device 300 identifies the action and processing procedure and moves mobile object 100 to a position near vending machine B. At this time, if an obstacle is present, mobile object 100 recognizes the obstacle and approaches vending machine B while avoiding the obstacle. In the example of FIG. 8 , user U also moves near vending machine B at time T+1, and user U can meet mobile object 100 near vending machine B.

[0073] As described above, the server device 300 can select an appropriate action and control the mobile object 100 to perform the appropriate action by using the various information as described above.

[0074] 9 is a flowchart showing an example of the flow of processing executed by the server device 300. First, the server device 300 acquires voice information, peripheral information, movement information, and interaction information (steps S100, S102, S104, and S106). Next, the server device 300 extracts feature information indicating the features of each piece of acquired information from each piece of information (step S108). For example, feature information of the voice information, feature information of the peripheral information, feature information of the movement information, and feature information of the interaction information are extracted.

[0075] Next, the server device 300 integrates the respective pieces of feature information to generate input information (step S110). Next, the server device 300 inputs the input information to the trained model 362 (step S112). Next, the server device 300 acquires the output information output by the trained model 362 (step S114).

[0076] Next, the server device 300 identifies an action and control information based on the action correspondence information 364 and the output information (step S116). Next, the server device 300 controls the moving object 100 based on the identified action (step S118). This completes one routine of the flowchart. As described above, the processing system 1 can select an appropriate action.

[0077] Here, for example, when an action is selected using only utterances, an appropriate action may not be selected. Also, for example, if a trained model is prepared for each scene and the trained model is used depending on the scene, the construction of the system may become complicated.

[0078] In this embodiment, as described above, by utilizing information other than speech, it is possible to select more appropriate actions by taking into account the state of the moving body 100 and the surrounding environment, and furthermore, since there is no need to prepare multiple trained models, the system can be constructed more easily.

[0079] [Use of Speech Intention] When generating feature information from audio information, the information processing unit 320 may generate feature information (second feature information) by using the speech intention. FIG. 10 is a diagram for explaining speech intentions. As shown in FIG. 10 , speech intentions 2 are associated with a plurality of utterances, and speech intentions 1 are associated with a plurality of utterance intentions 2. Speech intention 1 is a superordinate concept of speech intention 2. For example, speech intention 1 is a major item such as "instruction to move" or "backchannel," and speech intention 2 corresponding to the "instruction to move" of speech intention 1 is a medium item such as "direction," and an utterance for "direction" is "turn right." The information processing unit 320 may generate feature information based on at least one or more items of speech intention 1, speech intention 2, and the utterance, or may generate feature information based on the utterance and one or both of speech intentions 1 and 2. When the information processing unit 320 receives input information based on speech intention, peripheral information, and movement information, or input information based on speech intention, audio information, peripheral information, and movement information, it inputs the input information into a trained model 362 that outputs behavioral information regarding the behavior of the mobile body 100 according to the input information, and determines the behavior of the target based on the behavioral information obtained by inputting the input information.

[0080] As described above, by utilizing the intention of the utterance, feature information that better reflects the type of utterance is generated.

[0081] [Learning Device] The learning device 500 learns the learning information to generate a trained model 362. FIG. 11 is a diagram showing an example of the learning information 520. The learning information 520 is information in which input information for learning is associated with behavioral information. The input information is information based on feature information of speech information, feature information of peripheral information, feature information of movement information, and feature information of interaction information (input information generated by concatenating each feature vector). The behavioral information is utterance, behavior, or a combination of utterance and behavior. The behavioral information is behavioral information that is considered appropriate in the environment or situation in which the input information was obtained. When the input information of the learning information 520 is input, the learning device 500 generates a trained model 362 that has been trained to output behavioral information associated with the input information of the learning information 520.

[0082] Note that the feature information may include the location of the mobile object 100 on a map and information from each sensor equipped on the mobile object 100. In this case, the learning device 500 learns learning information in which input information including the feature information is associated with behavioral information. The server device 300 may input the input information including the feature information to the trained model 362 to obtain output information. The learning information 520 may also include the above-mentioned speech intention information. In this case, when input information based on a speech intention, peripheral information, and movement information, or input information based on a speech intention, audio information, peripheral information, and movement information, is input, the trained model 362 is trained to output behavioral information regarding the behavior of the mobile object 100 according to the input information. The trained model 362 then outputs behavioral information according to the input information.

[0083] According to the embodiment described above, the processing system 1 can select appropriate behavior by inputting input information based on the audio information, the surrounding information, and the movement information into a trained model that outputs behavioral information regarding the behavior of the target according to the input information, and determining the behavior of the target based on the behavioral information obtained by inputting the input information.

[0084] The above-described embodiment can be expressed as follows: A processing system comprising: a storage medium for storing computer-readable instructions; and a processor connected to the storage medium, wherein the processor executes the computer-readable instructions to: acquire audio information related to a voice spoken by a user, peripheral information related to a situation surrounding a control target, and movement information related to the movement of the target; and, when input information based on the audio information, the peripheral information, and the movement information is input, a trained model outputs behavior information related to the behavior of the target in accordance with the input information, and the input information is input into a trained model, which determines the behavior of the target based on the behavior information obtained.

[0085] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the gist of the present invention.

[0086] REFERENCE SIGNS LIST 1 Processing system 2 Terminal device 100 Mobile object 180 Detection unit 200 Control device 220 Recognition unit 240 Control unit 260 Storage unit 300 Server device 310 Information management unit 320 Information processing unit 330 Information provision unit 360 Storage unit 362 Learned model 364 Action correspondence information 366 Control information

Claims

1. An acquisition unit that acquires voice information regarding the voice spoken by a user, peripheral information regarding the situation around the control target, and movement information regarding the movement of the target; and a processing unit that, when inputting input information based on the voice information, the peripheral information, and the movement information, determines the action of the target based on the action information regarding the action of the target corresponding to the input information obtained by inputting the input information into a learned model that outputs the action information. A processing system comprising:

2. The processing system according to claim 1, wherein the input information is first feature information indicating features based on the voice information, the peripheral information, and the movement information.

3. The processing unit according to claim 2, generates second feature information indicating the features of the voice information, third feature information indicating the features of the peripheral information, and fourth feature information indicating the features of the movement information, and derives the first feature information based on the generated second feature information, third feature information, and fourth feature information. The processing system described.

4. The input information, the voice information, the peripheral information, and the movement information are text data, and the first feature information, the second feature information, the third feature information, and the fourth feature information are data obtained by vectorizing the text data. The processing system according to claim 3.

5. The learned model is a model in which learning information has been learned, the learning information is information in which the voice information, input information based on the peripheral information and the movement information, and action information corresponding to the input information are associated, and the learned model is When the input information is input, it is a model learned to output the action information associated with the input information. The processing system according to claim 1.

6. The acquisition unit further acquires interaction information between the user and the target, and the processing unit, when inputting input information based on the voice information, the peripheral information, the movement information, and the interaction information, inputs the input information into a learned model that outputs action information regarding the action of the target corresponding to the input information. Based on the obtained action information, the action of the target is determined. The processing system according to claim 1.

7. The processing unit controls the moving object to be controlled based on the action information, and the processing system according to any one of claims 1 to 6.

8. The surrounding information includes the position of an object detected by the target moving object and information indicating the type of the object, and the movement information includes the position information of the moving object, information indicating the direction of the moving object, and information indicating the speed of the moving object, and the processing system according to any one of claims 1 to 6.

9. The processing unit estimates a speech intention from speech information regarding the speech uttered by the user, and when inputting input information based on the speech intention, the surrounding information, and the movement information, or input information based on the speech intention, the speech information, the surrounding information, and the movement information, the processing unit inputs the input information to a learned model that outputs action information regarding the action of the target according to the input information, and determines the action of the target based on the action information obtained. The processing system according to any one of claims 1 to 6.

10. The action of the target moving object includes at least one of the movement and speech of the moving object, and the processing system according to any one of claims 1 to 6.

11. The action of the target moving object includes at least one of the actions of changing the speed, moving direction, moving distance, stop position, and stop time of the moving object, and the processing system according to any one of claims 1 to 6.

12. The speech of the target moving object includes at least one of an inquiry to the user, a proposal to the user, a repetition, an answer, gratitude, an apology, an alternative, an arrival location, an arrival time, information on the moving object and the surrounding environment of the moving object, and a confirmation speech to the user, and the processing system according to any one of claims 1 to 6.

13. A learning method in which a computer acquires learning information in which speech information regarding speech uttered by a user, surrounding information regarding the situation around the object to be controlled, movement information regarding the movement of the object, and information regarding the action of the object are associated, and uses the acquired learning information to learn a learning model so that when input information based on the speech information, the surrounding information, and the movement information is input, information regarding the action of the object associated with the speech information, the surrounding information, and the movement information is output, and a learned model is generated.

14. A processing method in which a computer obtains voice information regarding the voice spoken by a user, peripheral information regarding the situation around the object to be controlled, and movement information regarding the movement of the object, and when inputting input information based on the voice information, the peripheral information, and the movement information, determines the action of the object based on the action information regarding the action of the object corresponding to the input information, which is obtained by inputting the input information into a learned model that outputs the action information.

15. A program for causing a computer to obtain voice information regarding the voice spoken by a user, peripheral information regarding the situation around the object to be controlled, and movement information regarding the movement of the object, and when inputting input information based on the voice information, the peripheral information, and the movement information, causing the computer to determine the action of the object based on the action information regarding the action of the object corresponding to the input information, which is obtained by inputting the input information into a learned model that outputs the action information.

Citation Information

Patent Citations

  • Information processing apparatus, information processing method, mobile object control device, mobile object control method, and program

    JP2022155107A

  • Information processing device, information processing method, and program

    JP2020187282A

  • Automatic operation vehicle control device, vehicle allocation system, and vehicle allocation method

    JP2021177283A

  • Control device for mobile object, control method for mobile object, mobile object, information processing method, and program

    WO2023187890A1