Information processing system, information processing device, and information processing method
The information processing system addresses the challenge of collecting comprehensive learning data for robots by using a trigger detection unit, data collection unit, and license acquisition unit to enhance the recognition engine's performance through user-generated feedback.
Patent Information
- Application Number
- PCT/JP2025/019423
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-14
- Filing Date
- 2025-05-29
- Publication Date
- 2025-12-18
AI Technical Summary
Existing technologies face challenges in collecting comprehensive learning data for robots across all usage environments before shipment, making it difficult to effectively improve the performance of recognition engines.
An information processing system that includes a trigger detection unit to collect feedback data, a data collection unit to gather data used by the recognition engine, and a license acquisition unit to obtain user permission for using the feedback data, enabling efficient retraining of the recognition engine post-shipment.
Facilitates the efficient collection and utilization of user-generated data for relearning the recognition engine, enhancing its performance by adapting to diverse usage environments.
Smart Images

Figure JP2025019423_18122025_PF_FP_ABST
Abstract
Description
Information processing system, information processing device, and information processing method
[0001] The present technology relates to an information processing system, an information processing device, and an information processing method, and in particular to an information processing system, an information processing device, and an information processing method that are suitable for use when collecting data for re-learning a recognition engine of an autonomous moving body.
[0002] In recent years, robots using learning models generated by machine learning have been increasingly used (see, for example, Patent Literature 1). For example, entertainment robots that can interact with users use recognition engines, which are learning models for various types of recognition.
[0003] Japanese Patent Application Laid-Open No. 2022-175570
[0004] However, it is difficult to collect learning data that covers all the usage environments of all robots before shipping and to execute the learning process for the recognition engine.
[0005] In order to improve the performance of the recognition engine, it is effective to retrain the recognition engine using training data generated using data collected from each user after shipment, and then distribute the improved recognition engine.
[0006] The present technology has been made in light of such circumstances, and makes it possible to efficiently collect data to be used for relearning the recognition engine of an autonomous mobile object such as an entertainment robot.
[0007] An information processing system according to one aspect of the present technology includes a trigger detection unit that detects a trigger for collecting feedback data, a data collection unit that collects the feedback data including data used in processing by a recognition engine that recognizes at least one of a situation surrounding an autonomous moving body and a situation of the autonomous moving body around the time the trigger is detected, and a license acquisition unit that acquires a user's license to use the feedback data.
[0008] An information processing device according to one aspect of the present technology includes a trigger detection unit that detects a trigger for collecting feedback data, a data collection unit that collects the feedback data including data used in processing by a recognition engine that recognizes at least one of a situation surrounding an autonomous moving body and a situation of the autonomous moving body around the time the trigger is detected, and a license acquisition unit that acquires a user's license to use the feedback data.
[0009] An information processing method according to one aspect of the present technology includes an information processing system detecting a trigger for collecting feedback data, collecting the feedback data including data used in processing by a recognition engine that recognizes at least one of the situation surrounding an autonomous moving body and the situation of the autonomous moving body around the time the trigger was detected, and obtaining a user's permission to use the feedback data.
[0010] In one aspect of the present technology, a trigger for collecting feedback data is detected, the feedback data includes data used in processing by a recognition engine that recognizes at least one of the situation around the autonomous moving body and the situation of the autonomous moving body around the time the trigger is detected, and a user's permission to use the feedback data is obtained.
[0011] 1 is a block diagram showing an embodiment of an information processing system to which the present technology is applied; FIG. 2 is a front view of an autonomous moving body; FIG. 3 is a rear view of the autonomous moving body; FIG. 4 is a perspective view of the autonomous moving body; FIG. 5 is a side view of the autonomous moving body; FIG. 6 is a top view of the autonomous moving body; FIG. 7 is a bottom view of the autonomous moving body; FIG. 8 is a schematic diagram for explaining the internal structure of the autonomous moving body; FIG. 9 is a schematic diagram for explaining the internal structure of the autonomous moving body; FIG. 10 is a block diagram showing an example of the functional configuration of the autonomous moving body; FIG. 11 is a block diagram showing an example of the functional configuration implemented by a control unit of the autonomous moving body; FIG. 12 is a block diagram showing an example of the functional configuration of an information processing terminal; FIG. 13 is a block diagram showing an example of the functional configuration of an information processing server; FIG. 14 is a flowchart for explaining feedback data collection processing; FIG. 15 is a diagram showing an example of a conversation history screen; FIG. 16 is a diagram showing an example of a conversation history screen; FIG. 17 is a diagram showing an example of a license acquisition screen; FIG. 18 is a flowchart for explaining recognition engine improvement processing; FIG. 19 is a diagram showing an example of the configuration of a computer.
[0012] Hereinafter, embodiments of the present technology will be described. The description will be made in the following order: 1. Embodiment 2. Modification 3. Other
[0013] <<1. Embodiment>> An embodiment of the present technology will be described with reference to FIGS. 1 to 18 .
[0014] <Configuration Example of Information Processing System 1> FIG. 1 is a block diagram showing an embodiment of an information processing system 1 to which the present technology is applied.
[0015] The information processing system 1 includes autonomous mobile bodies 11-1 to 11-n, information processing terminals 12-1 to 12-n, and an information processing server 13.
[0016] In the following, when there is no need to distinguish between the autonomous mobile bodies 11-1 to 11-n, they will simply be referred to as the autonomous mobile body 11. In the following, when there is no need to distinguish between the information processing terminals 12-1 to 12-n, they will simply be referred to as the information processing terminals 12.
[0017] Communication is possible between each autonomous mobile body 11 and the information processing server 13, between each information processing terminal 12 and the information processing server 13, between each autonomous mobile body 11 and each information processing terminal 12, between each autonomous mobile body 11, and between each information processing terminal 12 via the network 21. In addition, direct communication is also possible between each autonomous mobile body 11 and each information processing terminal 12, between each autonomous mobile body 11, and between each information processing terminal 12 without going through the network 21.
[0018] The autonomous mobile object 11 is an information processing device that performs autonomous operation without being controlled by the information processing terminal 12 and the information processing server 13, or by being controlled by the information processing terminal 12 or the information processing server 13. The autonomous mobile object 11 is also an agent device that enables interaction with a user (e.g., conversation, physical contact, etc.) to be realized more naturally and effectively.
[0019] The autonomous mobile body 11 is an information processing device that recognizes its own and its surrounding situations based on collected sensing data, etc., and autonomously selects and executes various actions according to the situation. Unlike a robot that simply acts in accordance with user instructions, one of the features of the autonomous mobile body 11 is that it autonomously executes appropriate actions according to the situation.
[0020] The autonomous mobile body 11 can, for example, perform user recognition, object recognition, etc. based on captured images, and perform various autonomous actions according to the recognized user, object, etc. The autonomous mobile body 11 can also, for example, perform voice recognition based on the user's speech, and perform actions based on the user's instructions, etc.
[0021] Furthermore, the autonomous mobile body 11 performs pattern recognition learning to acquire the ability to recognize users and objects. In this case, the autonomous mobile body 11 can perform pattern recognition learning related to objects, etc., not only by supervised learning based on given learning data, but also by dynamically collecting learning data based on instructions from a user, etc.
[0022] Furthermore, the autonomous moving body 11 can be disciplined by a user. Here, the discipline of the autonomous moving body 11 is broader than, for example, general discipline in which the autonomous moving body 11 is taught rules and prohibited actions and made to memorize them, and refers to changes in the autonomous moving body 11 that the user can sense as a result of the user's interaction with the autonomous moving body 11.
[0023] Furthermore, the autonomous mobile body 11 can express its own state and communicate with a user or other autonomous mobile bodies by outputting output sounds. The output sounds of the autonomous mobile body 11 include operation sounds that are output in response to the state of the autonomous mobile body 11, and speech sounds for communicating with a user, other autonomous mobile bodies, etc.
[0024] The shape, capabilities, desires, and other levels of the autonomous mobile body 11 can be designed as appropriate depending on the purpose and role. For example, the autonomous mobile body 11 is configured as an autonomous mobile robot that autonomously moves within a space and performs various actions.
[0025] Specifically, for example, the autonomous mobile body 11 may be composed of various types of robots, such as a running type, a walking type, a flying type, a swimming type, etc. For example, the autonomous mobile body 11 may be composed of an autonomous mobile robot such as an entertainment robot that has a shape and behavioral capabilities that imitate those of an animal such as a human or a dog. For example, the autonomous mobile body 11 may be composed of a vehicle or other device that has the ability to interact with a user.
[0026] The information processing terminal 12 is, for example, a smartphone, a tablet terminal, a PC (personal computer), or the like, and is used by the user of the autonomous mobile body 11. The information processing terminal 12 realizes various functions by executing a predetermined application program (hereinafter simply referred to as an application). For example, the information processing terminal 12 manages and customizes the autonomous mobile body 11 by executing the predetermined application.
[0027] For example, the information processing terminal 12 communicates with the information processing server 13 via the network 21 or directly with the autonomous mobile body 11 to collect various data related to the autonomous mobile body 11, present the data to the user, or give instructions to the autonomous mobile body 11.
[0028] The information processing server 13, for example, collects various types of data from each autonomous mobile body 11 and each information processing terminal 12, provides various types of data to each autonomous mobile body 11 and each information processing terminal 12, and controls the behavior of each autonomous mobile body 11. Furthermore, for example, the information processing server 13 can perform pattern recognition learning and processing corresponding to user discipline, similar to the autonomous mobile body 11, based on the data collected from each autonomous mobile body 11 and each information processing terminal 12. Furthermore, for example, the information processing server 13 supplies each information processing terminal 12 with the above-mentioned applications and various types of data related to each autonomous mobile body 11.
[0029] Here, for example, various recognition functions of the autonomous mobile body 11 are realized by a recognition engine, which is an AI model generated by machine learning. The recognition engine is, for example, pre-installed when the autonomous mobile body 11 is shipped.
[0030] As will be described later, the information processing server 13 can improve the recognition engine by performing re-learning using data collected from each autonomous mobile body 11. Each autonomous mobile body 11 can improve its recognition performance by acquiring an improved recognition engine from the information processing server 13 and installing it (updating the recognition engine).
[0031] The recognition engine may be divided into multiple engines depending on, for example, the recognition function or the recognition target. In this case, the information processing server 13 may improve all the recognition engines or only some of the recognition engines.
[0032] The network 21 may be composed of, for example, public networks such as the Internet, telephone networks, and satellite communication networks, various LANs (Local Area Networks) including Ethernet (registered trademark), and WANs (Wide Area Networks). The network 21 may also include dedicated network such as an IP-VPN (Internet Protocol-Virtual Private Network). The network 21 may also include wireless communication networks such as Wi-Fi (registered trademark) and Bluetooth (registered trademark).
[0033] The configuration of the information processing system 1 can be flexibly changed depending on the specifications, operation, etc. For example, the autonomous mobile body 11 may communicate information with various external devices in addition to the information processing terminal 12 and the information processing server 13. The above external devices can include, for example, servers that transmit weather, news, and other service information, and various home appliances owned by the user.
[0034] Furthermore, for example, the autonomous mobile bodies 11 and the information processing terminals 12 do not necessarily have to have a one-to-one relationship, and may have, for example, a many-to-many, many-to-one, or one-to-many relationship. For example, one user can use one information processing terminal 12 to check data related to multiple autonomous mobile bodies 11, or can use multiple information processing terminals to check data related to one autonomous mobile body 11.
[0035] <Configuration Example of Autonomous Mobile Body 11> Next, a configuration example of the autonomous mobile body 11 will be described with reference to FIGS. 2 to 11. The autonomous mobile body 11 can be various devices that perform autonomous operation based on environmental recognition. In the following, a case will be described where the autonomous mobile body 11 is an agent-type robot device with an oblong shape that travels autonomously on wheels. The autonomous mobile body 11 performs autonomous operation according to, for example, its surroundings (including the user) and its own situation, thereby realizing various interactions including the presentation of information. The autonomous mobile body 11 is, for example, a small robot that is large and heavy enough to be easily lifted by a user with one hand.
[0036] <Examples of Exterior of Autonomous Moving Body 11> First, examples of exterior of the autonomous moving body 11 will be described with reference to FIGS.
[0037] Fig. 2 is a front view of the autonomous mobile body 11, and Fig. 3 is a rear view of the autonomous mobile body 11. Figs. 4A and 4B are perspective views of the autonomous mobile body 11. Fig. 5 is a side view of the autonomous mobile body 11. Fig. 6 is a top view of the autonomous mobile body 11. Fig. 7 is a bottom view of the autonomous mobile body 11.
[0038] 2 to 6, the autonomous moving body 11 has eye units 101L and 101R corresponding to the left and right eyes on the upper part of the main body. The eye units 101L and 101R are realized by, for example, a single or two independent OLEDs (Organic Light Emitting Diodes), LEDs, etc., and can express gaze, blinking, etc.
[0039] The autonomous moving body 11 also includes cameras 102L and 102R above the eye units 101L and 101R. The cameras 102L and 102R have the function of capturing images of the user and the surrounding environment. In this case, the autonomous moving body 11 may implement simultaneous localization and mapping (SLAM) based on the images captured by the cameras 102L and 102R.
[0040] The eye 101L, the eye 101R, the camera 102L, and the camera 102R are disposed on a substrate (not shown) disposed inside the exterior surface. The exterior surface of the autonomous mobile body 11 is basically formed using an opaque material, but a head cover 104 made of a transparent or translucent material is provided in the portion corresponding to the substrate on which the eye 101L, the eye 101R, the camera 102L, and the camera 102R are disposed. This allows the user to recognize the eye 101L and the eye 101R of the autonomous mobile body 11, and the autonomous mobile body 11 can capture images of the outside world.
[0041] 2, 4, and 7, the autonomous moving body 11 is equipped with a ToF (Time of Flight) sensor 103 at the bottom of the front face. The ToF sensor 103 is configured, for example, by 1D-ToF, and has the function of detecting the distance to an object or the ground in front of the autonomous moving body 11. The ToF sensor 103 can, for example, accurately detect the distance to various objects, detect steps, etc., and prevent the autonomous moving body 11 from falling or tipping over.
[0042] 3, 5, etc., the autonomous moving body 11 is provided on its rear surface with a connection terminal 105 for an external device and a power switch 106. The autonomous moving body 11 can connect to an external device via the connection terminal 105 and perform information communication, for example.
[0043] 7, the autonomous mobile body 11 is provided with wheels 107L and 107R on its bottom surface. The wheels 107L and 107R are driven by different motors (not shown). This allows the autonomous mobile body 11 to perform moving operations such as moving forward, backward, turning, and rotating.
[0044] The wheels 107L and 107R can be stored inside the main body and can be protruded to the outside. For example, the autonomous mobile body 11 can perform a jumping motion by forcefully protruding the wheels 107L and 107R to the outside. Note that FIG. 7 shows a state in which the wheels 107L and 107R are stored inside the main body.
[0045] In the following, when there is no need to distinguish between the eye unit 101L and the eye unit 101R, they will simply be referred to as the eye unit 101. In the following, when there is no need to distinguish between the camera 102L and the camera 102R, they will simply be referred to as the camera 102. In the following, when there is no need to distinguish between the wheel 107L and the wheel 107R, they will simply be referred to as the wheel 107.
[0046] <Example of Internal Structure of Autonomous Moving Body 11> FIGS. 8 and 9 are schematic diagrams showing the internal structure of the autonomous moving body 11. FIG.
[0047] 8, the autonomous mobile body 11 includes an IMU 121 and a communication device 122 arranged on an electronic board. The IMU 121 detects three-dimensional acceleration and angular velocity of the autonomous mobile body 11. The communication device 122 is configured to realize wireless communication with the outside, and includes, for example, a Bluetooth or Wi-Fi antenna.
[0048] The autonomous moving body 11 also includes a speaker 123, for example, inside the side surface of the main body. The autonomous moving body 11 can output various sounds using the speaker 123.
[0049] 9 , the autonomous mobile body 11 is provided with a microphone 124L, a microphone 124M, and a microphone 124R inside the upper part of the main body. The microphones 124L, 124M, and 124R collect the user's speech and surrounding environmental sounds. By providing the autonomous mobile body 11 with multiple microphones 124L, 124M, and 124R, it is possible to collect sounds generated in the surrounding area with high sensitivity and detect the position of the sound source.
[0050] 8 and 9, the autonomous mobile body 11 is equipped with motors 125A to 125E (however, motor 125E is not shown). Motor 125A and motor 125B, for example, drive a substrate on which the eye unit 101 and camera 102 are disposed in the vertical and horizontal directions. Motor 125C realizes a forward tilt posture of the autonomous mobile body 11. Motor 125D drives wheel 107L. Motor 125E drives wheel 107R. Motors 125A to 125E enable the autonomous mobile body 11 to express a wide range of movements.
[0051] Hereinafter, when there is no need to distinguish between the microphones 124L to 124R, they will be simply referred to as microphones 124. Hereinafter, when there is no need to distinguish between the motors 125A to 125E, they will be simply referred to as motors 125.
[0052] 10 shows an example of the functional configuration of the autonomous mobile body 11. The autonomous mobile body 11 includes a control unit 201, a sensor unit 202, an input unit 203, a storage unit 204, a light source 205, a sound output unit 206, a drive unit 207, and a communication unit 208.
[0053] The control unit 201 has a function of controlling each component of the autonomous mobile body 11. For example, the control unit 201 controls the activation and deactivation of each component. The control unit 201 also supplies control signals received from the information processing server 13 to the light source 205, sound output unit 206, and drive unit 207.
[0054] The sensor unit 202 has a function of collecting various data related to the surrounding situation (including the user). For example, the sensor unit 202 includes the above-mentioned camera 102, ToF sensor 103, inertial sensor 121, microphone 124, etc. In addition to the above-mentioned sensors, the sensor unit 202 may also include various other sensors, such as a geomagnetic sensor, a touch sensor, various optical sensors including an IR (infrared) sensor, a temperature sensor, a humidity sensor, etc. The sensor unit 202 supplies sensor data output from each sensor to the control unit 201.
[0055] The input unit 203 includes, for example, buttons and switches such as the power switch 106 described above, and detects physical input operations by the user.
[0056] The storage unit 204 stores various programs and data necessary for the processing of the autonomous moving body 11 .
[0057] The light source 205 includes, for example, the above-described eye unit 101 and expresses the eye movement of the autonomous moving body 11 .
[0058] The sound output unit 206 includes, for example, the above-mentioned speaker 123 and an amplifier, and outputs an output sound based on the output sound data supplied from the control unit 201 .
[0059] The drive unit 207 includes, for example, the wheels 107 and the motor 125 described above, and is used to express the body movement of the autonomous mobile body 11 .
[0060] The communication unit 208 includes, for example, the above-mentioned connection terminal 105 and communication device 122, and communicates with the information processing terminal 12, the information processing server 13, and other external devices.
[0061] <Example of Functional Configuration of Control Unit 201> FIG. 11 shows an example of a functional configuration realized by the control unit 201 of the autonomous moving body 11 executing a predetermined control program.
[0062] The information processing unit 251 includes a recognition unit 261 , an action planning unit 262 , an operation control unit 263 , a trigger detection unit 264 , and a data collection unit 265 .
[0063] The recognition unit 261 has the function of using a recognition engine to recognize the situation around the autonomous mobile body 11 and the situation of the autonomous mobile body 11 based on sensor data supplied from each sensor of the sensor unit 202 and control data supplied from the operation control unit 263.
[0064] For example, the recognition unit 261 performs user identification, recognition of the user's facial expression and gaze, object recognition, color recognition, shape recognition, marker recognition, obstacle recognition, step recognition, brightness recognition, recognition of stimuli to the autonomous mobile body 11, etc. For example, the recognition unit 261 performs emotion recognition related to the user's voice, word understanding, recognition of the position of a sound source, etc. For example, the recognition unit 261 recognizes the ambient temperature, the presence of an animal, the posture and movement of the autonomous mobile body 11, etc. For example, the recognition unit 261 recognizes a device combined with the autonomous mobile body 11 (hereinafter referred to as a combined device).
[0065] Examples of the autonomous mobile body 11 being combined with a combination device include when one of the autonomous mobile body 11 and the combination device is attached to the other, when one of the autonomous mobile body 11 and the combination device rides on the other, when the autonomous mobile body 11 and the combination device are combined, etc. Examples of the combination device include parts that can be attached to and detached from the autonomous mobile body 11 (hereinafter referred to as optional parts), a vehicle on which the autonomous mobile body 11 can ride (hereinafter referred to as a boarding vehicle), and a device to which the autonomous mobile body 11 can be attached and detached (hereinafter referred to as an attachment device).
[0066] Possible optional parts include, for example, parts that resemble parts of an animal's body (e.g., eyes, ears, nose, mouth, beak, horns, tail, wings, etc.), costumes, mascot costumes, parts that extend the functions and capabilities of the autonomous mobile body 11 (e.g., medals, weapons, etc.), wheels, caterpillar tracks, etc. Possible ride-on mobile bodies include, for example, cars, drones, robot vacuum cleaners, etc. Possible attachment devices include, for example, combined robots composed of multiple parts including the autonomous mobile body 11.
[0067] The combined device does not necessarily have to be a device dedicated to the autonomous moving body 11, but may be, for example, a general-purpose device.
[0068] Furthermore, the recognition unit 261 has a function of estimating and understanding the situation in which the autonomous moving body 11 is placed, based on the recognized information. In this case, the recognition unit 261 may perform comprehensive situation estimation using environmental knowledge stored in advance.
[0069] The recognition unit 261 supplies data indicating the recognition result to the action planning unit 262 , the operation control unit 263 , the trigger detection unit 264 , and the data collection unit 265 .
[0070] The recognition unit 261 temporarily stores data (for example, sensor data, control data, etc.) used in the recognition process of the recognition engine in the data storage unit 252 .
[0071] The behavior planning unit 262 sets an operation mode that defines the operation of the autonomous mobile body 11 based on the recognition result by the recognition unit 261, for example, the recognition result of the combined device by the recognition unit 261. The behavior planning unit 262 also has a function of planning an action to be taken by the autonomous mobile body 11 based on, for example, the recognition result by the recognition unit 261, the operation mode, and learning knowledge. Furthermore, the behavior planning unit 262 executes the behavior plan using, for example, a machine learning algorithm such as deep learning. The behavior planning unit 262 supplies data indicating the operation mode and the behavior plan to the operation control unit 263.
[0072] The operation control unit 263 controls the operation of the autonomous mobile body 11 by controlling the drive unit 207, the light source 205, and the sound output unit 206 based on the recognition result by the recognition unit 261, the action plan by the action planning unit 262, and the operation mode. For example, the operation control unit 263 causes the autonomous mobile body 11 to move while maintaining a forward-leaning posture, or to perform forward and backward movement, turning movement, rotational movement, etc. For example, the operation control unit 263 causes the autonomous mobile body 11 to actively perform an inducement action that induces interaction between the user and the autonomous mobile body 11. For example, the operation control unit 263 controls the output of various output sounds. The operation control unit 263 supplies control data related to the operation being performed by the autonomous mobile body 11 to the recognition unit 261.
[0073] The control data includes, for example, the rotation angle and rotation speed of the motor 125, the operation mode of the autonomous moving body 11, the angle of each joint, the content of speech, and the like.
[0074] The operation control unit 263 controls the output sound from the sound output unit 206 based on the recognition result by the recognition unit 261, the action plan by the action planning unit 262, and the operation mode. For example, the operation control unit 263 sets a control method for the output sound based on the operation mode, etc., and controls the output sound (for example, control of the content of the output sound to be generated and the output timing of the output sound, etc.) based on the set control method. The operation control unit 263 then generates output sound data for outputting the output sound and supplies it to the sound output unit 206.
[0075] In addition, the operation control unit 263 transmits information regarding the operation of the autonomous mobile body 11 (e.g., the operation history of the autonomous mobile body 11) to the information processing terminal 12 and the information processing server 13 via the communication unit 208 and, if necessary, the network 21.
[0076] The trigger detection unit 264 detects a trigger (hereinafter referred to as a collection trigger) for collecting feedback data that can be used for relearning the recognition engine, based on the recognition result by the recognition unit 261. When the trigger detection unit 264 detects a collection trigger, it notifies the data collection unit 265 that a collection trigger has been detected.
[0077] The data collection unit 265 collects feedback data including data used in processing by the recognition engine (e.g., sensor data, control data, etc.) from the data storage unit 252 around the time when the collection trigger is detected. The data collection unit 265 transmits the collected feedback data to the information processing terminal 12 via the communication unit 208.
[0078] The data storage unit 252 stores data (for example, sensor data, control data, etc.) used in the recognition process of the recognition engine for a predetermined period of time, and deletes the data after the predetermined period of time has elapsed.
[0079] <Example of Functional Configuration of Information Processing Terminal 12> FIG. 12 shows an example of the functional configuration of the information processing terminal 12. As shown in FIG.
[0080] The information processing terminal 12 includes an input unit 301 , a communication unit 302 , an information processing unit 303 , an output unit 304 , and a storage unit 305 .
[0081] The input unit 301 includes an input device for a user to input data, instructions, etc. For example, the input unit 301 includes a touch panel, buttons, switches, etc.
[0082] The communication unit 302 communicates with the autonomous mobile body 11 and the information processing server 13 via the network 21 , and also communicates directly with the autonomous mobile body 11 without going through the network 21 .
[0083] The information processing unit 303 includes a mobile object control unit 311 , a permission acquisition unit 312 , and an output control unit 313 .
[0084] The mobile object control unit 311 includes a recognition unit 321, a behavior planning unit 322, and a motion control unit 323. The recognition unit 321, the behavior planning unit 322, and the motion control unit 323 have the same functions as the recognition unit 261, the behavior planning unit 262, and the motion control unit 263 of the autonomous mobile object 11. In other words, the recognition unit 321, the behavior planning unit 322, and the motion control unit 323 can perform various processes in place of the recognition unit 261, the behavior planning unit 262, and the motion control unit 263 of the autonomous mobile object 11.
[0085] This allows the information processing terminal 12 to remotely control the autonomous moving body 11 , and the autonomous moving body 11 can perform various operations under the control of the information processing terminal 12 .
[0086] The license acquisition unit 312 receives feedback data from the autonomous mobile body 11 via the network 21 and the communication unit 302, and temporarily stores the feedback data in the storage unit 305. The license acquisition unit 312 executes processing to acquire a user's license for the feedback data received from the autonomous mobile body 11. For example, the license acquisition unit 312 instructs the output control unit 313 to display a screen for obtaining a license for using the feedback data from the user (hereinafter referred to as a license acquisition screen). For example, the license acquisition unit 312 acquires information indicating whether or not the feedback data can be used, which is input by the user via the input unit 301. When a license for using the feedback data has been obtained, the license acquisition unit 312 transmits the feedback data to the information processing server 13 via the communication unit 302 and the network 21.
[0087] The output control unit 313 controls the output of various types of information (for example, visual information, auditory information, tactile information, etc.) from the output unit 304.
[0088] The output unit 304 includes an output device that outputs various types of information (for example, visual information, auditory information, tactile information, etc.) For example, the output unit 304 includes a display device, a speaker, a haptic element, etc.
[0089] The storage unit 305 stores programs, data, etc. required for processing by the information processing terminal 12 .
[0090] <Example of Functional Configuration of Information Processing Server 13> FIG. 13 shows an example of the functional configuration of the information processing server 13. As shown in FIG.
[0091] The information processing server 13 includes a communication unit 401 , an information processing unit 402 , a learning data storage unit 403 , a recognition engine storage unit 404 , and a storage unit 405 .
[0092] The communication unit 401 communicates with the autonomous moving body 11 and the information processing terminal 12 via the network 21 .
[0093] The information processing unit 402 includes a mobile object control unit 411 , a learning data generation unit 412 , a learning unit 413 , and a distribution unit 414 .
[0094] The mobile object control unit 411 includes a recognition unit 421, a behavior planning unit 422, and a motion control unit 423. The recognition unit 421, the behavior planning unit 422, and the motion control unit 423 have the same functions as the recognition unit 261, the behavior planning unit 262, and the motion control unit 263 of the autonomous mobile object 11. In other words, the recognition unit 421, the behavior planning unit 422, and the motion control unit 423 can perform various processes in place of the recognition unit 261, the behavior planning unit 262, and the motion control unit 263 of the autonomous mobile object 11.
[0095] This allows the information processing server 13 to remotely control the autonomous mobile body 11, and the autonomous mobile body 11 can perform various operations under the control of the information processing server 13.
[0096] The training data generation unit 412 generates training data to be used for retraining the recognition engine, using feedback data received from each information processing terminal 12. The training data generation unit 412 includes a sampling unit 431 and an annotation unit 432.
[0097] The sampling unit 431 receives feedback data from each information processing terminal 12 via the network 21 and the communication unit 401, and stores the received feedback data in the training data storage unit 403. The sampling unit 431 samples the feedback data stored in the training data storage unit 403, and extracts feedback data to be used for re-training the recognition engine. The sampling unit 431 supplies information about the extracted feedback data to the annotation unit 432.
[0098] The annotation unit 432 annotates the feedback data extracted by the sampling unit 431 and generates training data to be used for retraining the recognition engine. The annotation unit 432 stores the generated training data in the training data storage unit 403.
[0099] The learning unit 413 improves the recognition engine by re-learning the recognition engine through a predetermined machine learning method using the learning data stored in the learning data storage unit 403. The learning unit 413 stores the improved recognition engine in the recognition engine storage unit 404.
[0100] The distribution unit 414 distributes the recognition engine stored in the recognition engine storage unit 404 to the autonomous mobile body 11 and the information processing terminal 12 by, for example, publishing the recognition engine on a website or the like.
[0101] The learning data storage unit 403 stores the feedback data received from each information processing terminal 12 and the learning data generated by the learning data generation unit 412 .
[0102] The recognition engine storage unit 404 stores the recognition engines generated or updated (re-learned) by the learning unit 413 .
[0103] The storage unit 405 stores various programs and data required for the processing of the information processing server 13 .
[0104] <Feedback Data Collection Process> Next, the feedback data collection process executed by the autonomous moving body 11 and the information processing terminal 12 will be described with reference to the flowchart of FIG.
[0105] In step S1, the trigger detection unit 264 determines whether a collection trigger has been detected. Specifically, the trigger detection unit 264 acquires information indicating the situation around the autonomous mobile body 11 and the recognition result of the situation of the autonomous mobile body 11 from the recognition unit 261. The trigger detection unit 264 determines whether a collection trigger has been detected based on at least one of the situation around the autonomous mobile body 11 and the situation of the autonomous mobile body 11.
[0106] The collection trigger includes, for example, a trigger based on the situation around the autonomous mobile body 11 and a trigger based on the situation of the autonomous mobile body 11. The trigger based on the situation around the autonomous mobile body 11 includes, for example, a trigger based on the situation of a user around the autonomous mobile body 11 and a trigger based on an event around the autonomous mobile body 11.
[0107] For example, the trigger detection unit 264 detects a collection trigger based on at least one of the user's speech content, gestures (including poses and signs), facial expressions, and interactions with the autonomous moving body 11.
[0108] Specifically, for example, the trigger detection unit 264 detects a predetermined gesture of the user (for example, a clenched hand pose) as a collection trigger. For example, the trigger detection unit 264 may recognize that the collection trigger is continuing while the user continues the predetermined gesture (for example, while the index finger is raised).
[0109] For example, the trigger detection unit 264 detects a facial expression that expresses a predetermined emotion of the user (for example, a sad facial expression) as a collection trigger.
[0110] For example, the trigger detection unit 264 detects a predetermined utterance content of the user (for example, "Show me a gesture, 3, 2, 1, yes") as a collection trigger.
[0111] For example, the trigger detection unit 264 detects a predetermined sound (such as the sound of the front door lock opening and closing) as a collection trigger.
[0112] For example, the trigger detection unit 264 detects a predetermined utterance content of the autonomous moving body 11 (for example, "Show me this gesture") as a collection trigger.
[0113] The determination process in step S1 is repeatedly executed until it is determined that a collection trigger has been detected, and if it is determined that a collection trigger has been detected, the process proceeds to step S2.
[0114] In step S2, the data collection unit 265 collects feedback data.
[0115] Specifically, the trigger detection unit 264 notifies the data collection unit 265 that a collection trigger has been detected.
[0116] The data collection unit 265 collects feedback data including data used in processing by the recognition engine around the time when the collection trigger is detected. For example, the data collection unit 265 extracts sensor data and control data used in processing by the recognition engine from the data storage unit 252 during a predetermined period before and after the collection trigger, and generates feedback data including the extracted data.
[0117] For example, the feedback data includes at least one of sensor data indicating the situation around the autonomous mobile body 11, sensor data indicating the situation of the autonomous mobile body 11, and control data of the autonomous mobile body 11. For example, the feedback data includes video data captured by the camera 102. For example, the feedback data may include one or more of sensor data acquired by the ToF sensor 103 (hereinafter referred to as ToF data), sensor data acquired by the IMU 121 (hereinafter referred to as IMU data), and audio data collected by the microphone 124. For example, the feedback data may include at least a portion of the control data output from the operation control unit 263.
[0118] For example, the data collection unit 265 may assign correct answer data to the feedback data based on the content of the user's utterance recognized by the recognition unit 261. For example, when the user is explaining a gesture that he or she is making, the data collection unit 265 may assign the content of the gesture to the feedback data as correct answer data.
[0119] For example, the autonomous mobile body 11 may move to a location where it is easy to collect feedback data and collect the feedback data. For example, the behavior planning unit 262 may create a behavior plan to move to the entrance based on the knowledge that the sound of locks opening and closing is easily obtained at the entrance, and the operation control unit 263 may move in the direction of the entrance based on the behavior plan. Then, the data collection unit 265 may collect feedback data including the sound of locks opening and closing near the entrance.
[0120] The data collection unit 265 transmits the collected feedback data to the information processing terminal 12 via the communication unit 208 .
[0121] In response to this, the permission acquisition unit 312 of the information processing terminal 12 receives the feedback data via the communication unit 208 and temporarily stores it in the storage unit 305 .
[0122] In step S3, the information processing terminal 12 confirms whether the feedback data can be used. For example, the output control unit 313, in accordance with an instruction from the license acquisition unit 312, causes the output unit 304 to display a screen for obtaining a license to use the feedback data from the user (hereinafter referred to as a license acquisition screen). In this case, some or all of the feedback data may be presented.
[0123] Here, examples of display of the license acquisition screen will be described with reference to Figures 15 to 17. Figures 15 to 17 show examples of application screens that allow the user to interact with the autonomous moving body 11, etc.
[0124] FIG. 15 shows an example of a conversation history screen that displays a conversation history as a history of interactions between the user and the autonomous moving body 11.
[0125] In the autonomous mobile body column 501 on the left side of the conversation history screen, the content of the speech of the autonomous mobile body 11 is displayed in chronological order.
[0126] The user column 502 on the right side of the conversation history screen displays the user's utterances in chronological order. In addition to the user's utterances, the user column also displays the recognition results of the user's actions (e.g., signs) recognized by the autonomous mobile body 11 in chronological order.
[0127] For example, when one of the recognition results of the user's behavior is selected, an input field 503 for feedback on the selected recognition result is displayed, as shown in Fig. 16. A thumbs-up icon 503a and a thumbs-down icon 503b are displayed in the input field 503. For example, if the recognition result is correct, the thumbs-up icon 503a is pressed, and if the recognition result is incorrect, the thumbs-down icon 503b is pressed.
[0128] When either the thumbs-up icon 503a or the thumbs-down icon 503b is pressed, the license input screen of FIG. 17 is displayed.
[0129] The license input screen may be displayed by a predetermined operation, even if the thumbs-up icon 503a and the thumbs-down icon 503b are not pressed.
[0130] The license input screen includes a recognition result field 521 , a video playback area 522 , a license button 523 , and a questionnaire field 524 .
[0131] The recognition result column 521 displays the recognition result of the user's motion to be evaluated.
[0132] The video playback area 522 is an area for displaying a video based on the video data included in the feedback data corresponding to the recognition result to be evaluated. That is, in the video playback area 522, a video based on the video data used for the recognition result to be evaluated is displayed.
[0133] The permission button 523 is a button for permitting the use of feedback data corresponding to the recognition result to be evaluated. In other words, the permission button 523 is a button for permitting the use of feedback data including video data corresponding to the video displayed in the video playback area 522. For example, while watching the video in the video playback area, the user determines whether or not to permit the use of the feedback data, and presses the permission button 523 if they permit the use of the feedback data.
[0134] The questionnaire field 524 is a field where the user can input improvements that they would like to see made to the recognition function (recognition engine) of the autonomous mobile body 11. For example, in this example, it is possible to input improvements to the recognition engine, such as the engine recognizing voices arbitrarily, producing incorrect voice recognition results, or failing to recognize the speaker (user).
[0135] For example, the function of playing this video may be provided to the user as a function of peeking into the mind of the autonomous mobile body 11 (the recognition state of the autonomous mobile body 11). This allows the user to grant permission to use feedback data including video data that represents the inside of the mind of the autonomous mobile body 11, which is expected to lower the psychological barrier to granting permission.
[0136] In step S4, the license acquisition unit 312 determines whether or not a license to use the feedback data has been obtained. For example, if the license button 523 is pressed on the license input screen of Fig. 17, the license acquisition unit 312 determines that a license to use the feedback data has been obtained, and the process proceeds to step S5.
[0137] In step S5 , the permission acquisition unit 312 transmits the feedback data to the information processing server 13 via the communication unit 302 and the network 21 .
[0138] In response to this, the sampling unit 431 of the information processing server 13 receives the feedback data via the network 21 and the communication unit 401 and stores it in the learning data storage unit 403 .
[0139] Thereafter, the process returns to step S1, and the processes from step S1 onwards are executed.
[0140] On the other hand, if it is determined in step S4 that permission to use the feedback data has not been obtained, the process returns to step S1, and the processes from step S1 onwards are executed. That is, in this case, the feedback data is not transmitted to the information processing server 13.
[0141] <Recognition Engine Improvement Process> Next, the recognition engine improvement process executed by the information processing server 13 will be described with reference to the flowchart of FIG.
[0142] This process may be initiated, for example, when it is time to improve the recognition engine.
[0143] The timing for improving the recognition engine can be set arbitrarily. For example, the recognition engine may be improved when the amount of accumulated feedback data exceeds a predetermined threshold. For example, the recognition engine may be improved when a predetermined period of time has passed since the previous version of the recognition engine was released.
[0144] In step S51, the sampling section 431 samples the feedback data.
[0145] Here, using all of the feedback data stored in the learning data storage unit 403 as learning data is inefficient, requiring a great deal of cost and time.
[0146] In response to this, for example, the sampling unit 431 extracts feedback data to be used for re-training the recognition engine from the feedback data stored in the training data storage unit 403 based on the degree of contribution to improving the performance of the recognition engine.
[0147] This extracts feedback data that is expected to contribute greatly to improving the performance of the recognition engine, such as feedback data that is likely to be misrecognized by the current recognition engine and feedback data with unusual patterns.
[0148] Specifically, for example, feedback data in which the recognition results by the recognition engine for multiple types of data included in the feedback data are different is extracted. For example, feedback data in which the recognition results for video data and IMU data are different (for example, whether or not the user's hand claps are recognized is different) is extracted.
[0149] For example, the sampling unit 431 may extract the feedback data automatically or manually, for example, in accordance with instructions from a user (hereinafter referred to as a developer) involved in relearning the recognition engine.
[0150] When extracting feedback data manually, for example, the sampling unit 431 presents candidate feedback data to be extracted via a UI or audio, and extracts the feedback data based on feedback from the developer.
[0151] The sampling unit 431 supplies information about the extracted feedback data to the annotation unit 432 .
[0152] In step S52, the annotation unit 432 annotates the feedback data, i.e., the annotation unit 432 annotates the feedback data extracted in the process of step S51, and generates learning data.
[0153] The annotation unit 432 may perform the annotation automatically or manually. In the latter case, for example, the annotation unit 432 adds annotations to the feedback data in accordance with instructions from the developer.
[0154] For example, the annotation unit 432 may align the time axes of different types of time series data included in the feedback data, and use time series data with high recognizability (e.g., visibility or readability) (e.g., video data or audio data) to annotate time series data with low recognizability (e.g., IMU data).
[0155] For example, the mobile object control unit 411 uses the feedback data to cause the autonomous mobile object 11 to reproduce a movement. Then, a developer may input annotations for the feedback data while observing the movement of the autonomous mobile object 11.
[0156] For example, the annotation unit 432 uses the feedback data to present a video that recreates the field of view of the autonomous mobile body 11 to a developer wearing virtual reality (VR) goggles. Then, for example, the developer may input annotations of the feedback data while watching the presented video, with the feeling that he or she is operating the autonomous mobile body 11.
[0157] As a result, a correct label indicating a correct recognition result (e.g., the type of user's signature to be recognized, the type of voice to be recognized, etc.) is assigned to the feedback data, and training data is generated. The correct label may be assigned, for example, by an explanatory sentence in natural language.
[0158] The annotation unit 432 stores the generated learning data in the learning data storage unit 403 .
[0159] In step S53, the learning unit 413 re-trains the recognition engine. Specifically, the learning unit 413 re-trains the recognition engine by a predetermined machine learning method using the training data stored in the training data storage unit 403. This improves the recognition engine.
[0160] The machine learning method is not particularly limited as long as it is, for example, supervised learning.
[0161] The learning unit 413 may also improve multiple recognition engines.
[0162] For example, when the recognition engine differs depending on the recognition function or recognition target, the learning unit 413 may re-learn and improve each recognition engine.
[0163] For example, the learning unit 413 may divide the learning data based on the user's attributes (e.g., gender, age, occupation, etc.) and retrain and improve multiple recognition engines for users with each attribute.
[0164] The learning unit 413 stores the improved recognition engine in the recognition engine storage unit 404 .
[0165] In step S54, the distribution unit 414 distributes the recognition engine. For example, the distribution unit 414 publishes the improved recognition engine on a website or the like. At this time, if multiple recognition engines are improved, each recognition engine is published. Furthermore, for example, each version of the recognition engine before and after improvement may be published.
[0166] In response to this, the user downloads a desired recognition engine from a website using the information processing terminal 12 and installs it in the autonomous mobile body 11 and the information processing terminal 12. This improves the recognition performance of the autonomous mobile body 11.
[0167] Furthermore, for example, the user may provide feedback on the evaluation of the improved recognition engine using the autonomous mobile body 11 or the information processing terminal 12. For example, the user may input the evaluation of the improved recognition engine by talking to the autonomous mobile body 11 in which the improved recognition engine is installed.
[0168] For example, if the autonomous moving body 11 is successful in recognition, the user may have a conversation including positive evaluations such as "You understood well" or "You've become smart." On the other hand, if the autonomous moving body 11 fails in recognition, the user may have a conversation including negative evaluations such as "You're wrong" or "Is it still difficult?"
[0169] In response to this, for example, the autonomous mobile body 11 feeds back the user's evaluation of the improved recognition engine to the information processing server 13 via the information processing terminal 12 .
[0170] For example, the user may input a numerical value or the like to evaluate the improved recognition engine into the information processing terminal 12 .
[0171] In response to this, for example, the information processing terminal 12 feeds back to the information processing server 13 the user's evaluation of the improved recognition engine.
[0172] Then, the recognition engine improvement process ends.
[0173] In this manner, data to be used for relearning the recognition engine of the autonomous mobile body 11 can be collected efficiently.
[0174] For example, by having the autonomous mobile body 11 automatically collect feedback data based on a collection trigger, it becomes possible to easily collect multimodal data that is effective for retraining the recognition engine.
[0175] For example, because a user's permission to use each piece of feedback data can be obtained, it becomes easy to collect data that can be used safely without risk of privacy violations, etc. Furthermore, because users can easily check the recognition results and the contents of the feedback data, it becomes possible to provide data with peace of mind that does not pose a risk of privacy violations, etc.
[0176] For example, sampling of feedback data allows for efficient retraining of the recognition engine using carefully selected training data.
[0177] For example, by performing annotation using multimodal feedback data, it is possible to obtain more accurate correct labels for training data, thereby improving training accuracy.
[0178] Furthermore, by distributing an improved recognition engine, the recognition performance of the autonomous mobile body 11 is improved. This makes it possible to give the user the impression that the autonomous mobile body 11 is growing, for example.
[0179] Furthermore, for example, through the above processing, it becomes possible to realize MLOps involving users.
[0180] <<2. Modifications>> Modifications of the above-described embodiments of the present technology will now be described.
[0181] <Modifications Regarding Allocation of Processing> For example, it is possible to appropriately change the allocation of processing in the information processing system 1. For example, some or all of the processing of the autonomous mobile body 11, the information processing terminal 12, or the information processing server 13 may be executed by another device.
[0182] For example, the autonomous mobile body 11 may execute part or all of the processing of the information processing terminal 12. For example, the autonomous mobile body 11 may transmit feedback data directly to the information processing server 13 without going through the information processing terminal 12.
[0183] In this case, for example, the autonomous mobile body 11 may present at least a part of the feedback data to the user. For example, if the autonomous mobile body 11 is equipped with a display or a projector, the autonomous mobile body 11 may present to the user a video based on video data included in the feedback data. For example, the autonomous mobile body 11 may present to the user a sound based on audio data included in the feedback data.
[0184] In response to this, for example, the user may grant permission to use the feedback data by having a conversation (voice interaction) with the autonomous mobile body 11. For example, the user may grant permission to use the feedback data by speaking to the autonomous mobile body 11 with content such as "You may use the data."
[0185] Furthermore, for example, the autonomous mobile body 11 may obtain permission to use the feedback data by asking the user whether or not to use the feedback data. For example, the autonomous mobile body 11 may ask the user, "May I use this data?", and the user may respond with "Yes," thereby granting permission to use the feedback data.
[0186] Then, for example, the autonomous mobile body 11 may directly transmit the feedback data for which a license has been obtained to the information processing server 13 via the network 21 .
[0187] Furthermore, for example, the autonomous mobile body 11 or the information processing terminal 12 may transmit all collected feedback data to the information processing server 13. In this case, for example, the information processing server 13 may retain feedback data for which a user's permission to use has been obtained, use the data for retraining the recognition engine, and delete feedback data for which a user's permission to use has not been obtained.
[0188] For example, the autonomous moving body 11 may transmit sensor data and control data to the information processing terminal 12, and the information processing terminal 12 may collect feedback data.
[0189] <Application Example of the Present Technology> Although the above describes an example in which the present technology is applied to an entertainment robot, the present technology can be applied to any autonomous mobile body equipped with a recognition function. For example, the present technology can be applied to vehicles and drones that use recognition functions for self-driving.
[0190] <<3. Others>> <Example of Computer Configuration> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs that make up the software are installed on a computer. Here, the computer includes a computer built into dedicated hardware, and a general-purpose personal computer, for example, that can execute various functions by installing various programs.
[0191] FIG. 19 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.
[0192] In the computer 1000 , a CPU (Central Processing Unit) 1001 , a ROM (Read Only Memory) 1002 , and a RAM (Random Access Memory) 1003 are interconnected by a bus 1004 .
[0193] An input / output interface 1005 is further connected to the bus 1004. An input unit 1006, an output unit 1007, a storage unit 1008, a communication unit 1009, and a drive 1010 are connected to the input / output interface 1005.
[0194] The input unit 1006 includes input switches, buttons, a microphone, an image sensor, etc. The output unit 1007 includes a display, a speaker, etc. The storage unit 1008 includes a hard disk, a non-volatile memory, etc. The communication unit 1009 includes a network interface, etc. The drive 1010 drives removable media 1011 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0195] In the computer 1000 configured as described above, the CPU 1001 performs the above-described series of processes by, for example, loading a program recorded in the memory unit 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004 and executing it.
[0196] The program executed by the computer 1000 (CPU 1001) can be provided by being recorded on a removable medium 1011 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0197] In the computer 1000, the program can be installed in the storage unit 1008 via the input / output interface 1005 by inserting the removable medium 1011 into the drive 1010. The program can also be received by the communication unit 1009 via a wired or wireless transmission medium and installed in the storage unit 1008. Alternatively, the program can be installed in the ROM 1002 or the storage unit 1008 in advance.
[0198] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.
[0199] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are housed in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.
[0200] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.
[0201] For example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.
[0202] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.
[0203] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0204] <Examples of Combinations of Configurations> The present technology can also have the following configurations.
[0205] (1) An information processing system comprising: a trigger detection unit that detects a trigger for collecting feedback data; a data collection unit that collects the feedback data including data used in processing by a recognition engine that recognizes at least one of a surrounding situation of an autonomous mobile body and a situation of the autonomous mobile body around the time the trigger is detected; and a license acquisition unit that acquires a user's license for the feedback data. (2) The information processing system described in (1), further comprising: a recognition unit that recognizes at least one of a surrounding situation of the autonomous mobile body and a situation of the autonomous mobile body using the recognition engine. (3) The information processing system described in (2), wherein the trigger detection unit detects the trigger based on at least one of a surrounding situation of the autonomous mobile body and a situation of the autonomous mobile body recognized by the recognition unit. (4) The information processing system described in (3), wherein the trigger detection unit detects the trigger based on a user's situation around the autonomous mobile body recognized by the recognition unit. (5) The information processing system according to (4), wherein the trigger detection unit detects the trigger based on at least one of a gesture, a speech content, and a facial expression of a user around the autonomous moving body recognized by the recognition unit. (6) The information processing system according to any of (2) to (5), wherein the license acquisition unit acquires the license based on a voice interaction between the user recognized by the recognition unit and the autonomous moving body. (7) The information processing system according to any of (1) to (6), further comprising an output control unit that controls presentation of at least one of a recognition result by the recognition engine and at least a portion of the feedback data to the user.(8) The information processing system according to any one of (1) to (7), further comprising: an output control unit that controls presentation of a history of interaction between a user and the autonomous moving body and a recognition result by the recognition engine in the history, wherein the license acquisition unit acquires the license in response to user feedback on the recognition result. (9) The information processing system according to any one of (1) to (8), further comprising: an annotation unit that performs annotation of the feedback data for which the license has been obtained and generates training data to be used for retraining the recognition engine. (10) The information processing system according to (9), wherein the feedback data includes first time-series data and second time-series data, and the annotation unit performs annotation of the second time-series data based on a recognition result by the recognition engine based on the first time-series data. (11) The information processing system according to (9) or (10), further comprising: a sampling unit that extracts the feedback data to be used for retraining the recognition engine from the feedback data for which the license has been obtained; and the annotation unit performs annotation of the feedback data extracted by the sampling unit to generate the training data. (12) The information processing system according to (11), wherein the sampling unit extracts the feedback data based on a contribution to performance improvement of the recognition engine. (13) The information processing system according to (12), wherein the feedback data includes first data and second data; and the sampling unit extracts the feedback data for which the recognition result based on the first data and the recognition result based on the second data differ from each other. (14) The information processing system according to any of (9) to (13), further comprising: a training unit that performs retraining of the recognition engine using the training data. (15) The information processing system according to (14), further comprising: a distribution unit that distributes the recognition engine obtained by retraining.(16) The information processing system according to any of (1) to (15), wherein the feedback data includes at least one of sensor data indicating a situation around the autonomous mobile body, sensor data indicating a situation of the autonomous mobile body, and control data of the autonomous mobile body. (17) An information processing device comprising: a trigger detection unit that detects a trigger for collecting feedback data; a data collection unit that collects the feedback data including data used in processing by a recognition engine that recognizes at least one of the situation around the autonomous mobile body and the situation of the autonomous mobile body around the time the trigger is detected; and a license acquisition unit that acquires a user's license for the feedback data. (18) An information processing method, wherein an information processing system includes: detecting a trigger for collecting feedback data; collecting the feedback data including data used in processing by a recognition engine that recognizes at least one of the situation around the autonomous mobile body and the situation of the autonomous mobile body around the time the trigger is detected; and acquiring a user's license for the feedback data.
[0206] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0207] 1 Information processing system, 11-1 to 11-n Autonomous moving body, 12-1 to 12-n Information processing terminal, 113 Information processing server, 101L, 101R Eye unit, 102L, 102R Camera, 103 ToF sensor, 104 Head cover, 121 IMU, 123 Speaker, 124L to 124R Microphone, 125A to 125E Motor, 201 Control unit, 202 Sensor unit, 206 Sound output unit, 207 Drive unit, 251 Information processing unit, 252 Data accumulation unit, 261 Recognition unit, 262 Action planning unit, 263 Operation control unit, 264 Trigger detection unit, 265 Data extraction unit, 303 Information processing unit, 312 Permission acquisition unit, 313 Output control unit, 304 Output unit, 402 Information processing unit, 403 Learning data storage unit, 404 Recognition engine storage unit, 412 Learning data generation unit, 413 Learning unit, 414 Distribution unit, 431 Sampling unit, 432 Annotation unit
Claims
1. An information processing system comprising: a trigger detection unit that detects a trigger for collecting feedback data; a data collection unit that collects the feedback data including data used in processing by a recognition engine that recognizes at least one of the situation around an autonomous moving body and the situation of the autonomous moving body around the time the trigger is detected; and a license acquisition unit that acquires a user's license to use the feedback data.
2. The information processing system according to claim 1, further comprising a recognition unit that uses the recognition engine to recognize at least one of the surrounding situation of the autonomous moving body and the situation of the autonomous moving body.
3. The information processing system according to claim 2, wherein the trigger detection unit detects the trigger based on at least one of the surrounding situation of the autonomous moving body recognized by the recognition unit and the situation of the autonomous moving body.
4. The information processing system according to claim 3, wherein the trigger detection unit detects the trigger based on the situation of the user around the autonomous moving body recognized by the recognition unit.
5. The information processing system according to claim 4, wherein the trigger detection unit detects the trigger based on at least one of a gesture, speech content, and facial expression of a user around the autonomous moving body recognized by the recognition unit.
6. The information processing system according to claim 2, wherein the license acquisition unit acquires the license based on a voice interaction between the user recognized by the recognition unit and the autonomous moving body.
7. The information processing system of claim 1, further comprising an output control unit that controls the presentation of at least one of the recognition result by the recognition engine and at least a portion of the feedback data, wherein the permission acquisition unit acquires the permission to use after at least one of the recognition result by the recognition engine and at least a portion of the feedback data has been presented to the user.
8. The information processing system according to claim 1, further comprising an output control unit that controls presentation of a history of interactions between a user and the autonomous mobile body and the recognition results by the recognition engine in the history, wherein the permission acquisition unit acquires the permission to use in the form of user feedback on the recognition results.
9. The information processing system according to claim 1, further comprising an annotation unit that performs annotation of the feedback data for which the license has been obtained and generates training data to be used for retraining the recognition engine.
10. The information processing system of claim 9, wherein the feedback data includes first time series data and second time series data, and the annotation unit performs annotation of the second time series data based on the recognition result of the recognition engine based on the first time series data.
11. An information processing system as described in claim 9, further comprising a sampling unit that extracts the feedback data to be used for retraining the recognition engine from the feedback data for which the license has been obtained, and the annotation unit performs annotation of the feedback data extracted by the sampling unit to generate the training data.
12. The information processing system according to claim 11, wherein the sampling unit extracts the feedback data based on the degree of contribution to improving the performance of the recognition engine.
13. The information processing system of claim 12, wherein the feedback data includes first data and second data, and the sampling unit extracts the feedback data in which the recognition result based on the first data by the recognition engine differs from the recognition result based on the second data.
14. An information processing system according to claim 9, further comprising a learning unit that uses the training data to retrain the recognition engine.
15. The information processing system according to claim 14, further comprising a distribution unit that distributes the recognition engine obtained by relearning.
16. The information processing system according to claim 1, wherein the feedback data includes at least one of sensor data indicating the situation around the autonomous moving body, sensor data indicating the situation of the autonomous moving body, and control data for the autonomous moving body.
17. An information processing device comprising: a trigger detection unit that detects a trigger for collecting feedback data; a data collection unit that collects the feedback data including data used in processing by a recognition engine that recognizes at least one of the situation around an autonomous moving body and the situation of the autonomous moving body around the time the trigger is detected; and a license acquisition unit that acquires a user's license to use the feedback data.
18. An information processing method, comprising: an information processing system detecting a trigger for collecting feedback data; collecting the feedback data including data used in processing by a recognition engine that recognizes at least one of the situation around an autonomous moving body and the situation of the autonomous moving body around the time the trigger was detected; and obtaining a user's permission to use the feedback data.
Citation Information
Patent Citations
Information processing equipment, information processing method and program
JP2012038240A
Orchestrating the execution of a sequence of actions requested to be performed through an automation assistant
JP2021533399A
Object detection apparatus and object detection method
JP2024066084A