Information processing device, mobile object, control method thereof, program, and storage medium
The information processing device uses image recognition to generate efficient questions for target user estimation by calculating impurity, addressing the inefficiencies of conventional methods.
Patent Information
- Application Number
- JP2022041683
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-16
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-03-16
AI Technical Summary
Conventional technologies for identifying a target user from multiple people do not effectively utilize image recognition features, leading to inefficient question generation and estimation.
An information processing device that utilizes image recognition to detect and extract feature amounts from captured images, calculates impurity to determine the separation degree of targets, and generates questions to minimize the number of questions needed for accurate estimation.
Efficiently estimates the target user by generating optimal questions based on image recognition features, reducing the number of interactions required for identification.
Smart Images

Figure 0007725399000001 
Figure 0007725399000002 
Figure 0007725399000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, a mobile object, a control method thereof, a program, and a storage medium. [Background technology]
[0002] In recent years, small mobile objects such as electric vehicles with a seating capacity of one to two people, called ultra-compact mobility (also called micromobility), and mobile interactive robots that provide various services to people have become known. Such mobile objects identify an arbitrary object from a group of people or buildings as a target object (hereinafter referred to as a target) and provide various services. In order to identify a user who is the target object, the mobile object interacts with the user to narrow down the candidates.
[0003] Regarding questions to the user, Patent Document 1 proposes a technology for generating a decision tree for question ordering, which allows the number of questions to be asked to the user to be reduced even if the user's answers are incorrect, when the user is asked multiple questions through dialogue and candidates for classification results are narrowed down from the user's answers. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2018-5624 Summary of the Invention [Problem to be solved by the invention]
[0005] However, the above-mentioned conventional technology has the following problems. The above-mentioned conventional technology reduces the number of questions posed to the user, while also taking into consideration cases where the user's answers are incorrect when narrowing down classification result or search result candidates from the answers. However, the above-mentioned conventional technology narrows down classification result candidates from answers to multiple questions posed to the user, and does not effectively utilize information other than the user's answers. In particular, when estimating a target user from among multiple people, the feature amounts of captured images around the user's location are very significant information.
[0006] The present invention has been made in view of the above-mentioned problems, and aims to generate efficient questions by utilizing features of image recognition and to estimate a target object. [Means for solving the problem]
[0007] According to the present invention, for example, an information processing device is characterized by comprising: an acquisition means for acquiring a captured image; an extraction means for detecting multiple targets included in the captured image and extracting multiple feature amounts for each of the detected multiple targets; an acquisition means for acquiring, for each feature amount extracted by the extraction means, an impurity indicating the degree to which a specific target cannot be separated from the multiple targets when a question is asked to a user to estimate a specific target from the multiple targets based on each feature amount; and a generation means for generating the questions based on the feature amounts extracted by the extraction means and the impurity for each feature amount, in order to reduce the number of questions asked to minimize the impurity.
[0008] Furthermore, according to the present invention, for example, a moving body is characterized by comprising: an acquisition means for acquiring a captured image; an extraction means for detecting a plurality of targets included in the captured image and extracting a plurality of feature amounts for each of the detected targets; an acquisition means for acquiring, for each feature amount extracted by the extraction means, an impurity indicating the degree to which a predetermined target cannot be separated from the plurality of targets when a question is asked to a user to estimate a predetermined target from the plurality of targets based on each feature amount; and a generation means for generating the questions based on the feature amounts extracted by the extraction means and the impurity for each feature amount in order to reduce the number of questions asked to minimize the impurity. [Effects of the Invention]
[0009] According to the present invention, it is possible to generate efficient questions by utilizing image recognition features and estimate the target object. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a diagram illustrating an example of a system according to an embodiment of the present invention. [Figure 2] FIG. 1 is a block diagram showing an example of the hardware configuration of a mobile body according to an embodiment of the present invention; [Figure 3] FIG. 1 is a block diagram showing an example of the functional configuration of a moving body according to an embodiment of the present invention; [Figure 4] FIG. 1 is a block diagram showing an example of the configuration of a server and a communication device according to an embodiment of the present invention. [Figure 5] FIG. 1 is a diagram illustrating image acquisition according to the present embodiment. [Figure 6] FIG. 1 is a diagram illustrating image analysis according to the present embodiment. [Figure 7] FIG. 10 is a diagram illustrating question generation according to the present embodiment. [Figure 8] FIG. 10 is a diagram comparing questions according to the present embodiment with questions in a comparative example. [Figure 9] 1 is a flowchart showing a series of operations in a user estimation process using speech and images according to the present embodiment. [Figure 10]10 is a flowchart showing a series of operations in a user estimation process (S106) using speech and captured images according to the present embodiment. [Figure 11] A flowchart showing a series of detailed processing operations in S206 according to the present embodiment. [Figure 12] FIG. 10 is a diagram illustrating an example of a system according to another embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, the embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention as claimed, and not all combinations of features described in the embodiments are necessarily essential to the invention. Two or more of the features described in the embodiments may be arbitrarily combined. Furthermore, the same reference numerals are used for the same or similar components, and redundant explanations will be omitted.
[0012] <System configuration> The configuration of a system 1 according to this embodiment will be described with reference to FIG. 1. The system 1 includes a vehicle (mobile object) 100, a server 110, and a communication device (communication terminal) 120. In this embodiment, the server 110 estimates the user 130 using speech information from the user 130 and captured images of the area around the vehicle 100, and causes the user 130 to merge with the vehicle 100. The user interacts with the server 110 via a predetermined application running on the communication device 120 held by the user, and moves to a meeting point designated by the user (for example, a nearby red postbox that serves as a landmark) while providing information such as the user's location through speech. The server 110 estimates the user and the meeting point, and controls the vehicle 100 to move to the estimated meeting point. Each component will be described in detail below.
[0013] The vehicle 100 is equipped with a battery and is, for example, an ultra-compact mobility vehicle that moves mainly by motor power. An ultra-compact mobility vehicle is more compact than a typical automobile and has a passenger capacity of approximately one or two people. In this embodiment, the vehicle 100 is described as an ultra-compact mobility vehicle, but this is not intended to limit the present invention, and the vehicle may be, for example, a four-wheeled vehicle or a saddle-ride vehicle. Furthermore, the vehicle of the present invention is not limited to vehicles, but may also be a vehicle that carries luggage and runs alongside a person walking, or a vehicle that leads a person. Furthermore, the present invention is not limited to four-wheeled or two-wheeled vehicles, but may also be applied to walking robots that are capable of autonomous movement. In other words, the present invention can be applied to moving bodies such as these vehicles and walking robots, and the vehicle 100 is an example of a moving body.
[0014] The vehicle 100 connects to the network 140 via wireless communication such as Wi-Fi or fifth-generation mobile communication. The vehicle 100 can measure conditions inside and outside the vehicle (such as the vehicle's position, driving status, and surrounding object landmarks) using various sensors and transmit the measured data to the server 110. The data collected and transmitted in this manner is generally referred to as floating data, probe data, traffic information, etc. Information about the vehicle is transmitted to the server 110 at regular intervals or in response to the occurrence of a specific event. The vehicle 100 can travel autonomously even when the user 130 is not on board. The vehicle 100 receives information such as control commands provided by the server 110 or controls the operation of the vehicle using data measured by the vehicle itself.
[0015] The server 110 is an example of an information processing device, and is configured with one or more server devices, and is capable of acquiring information about the vehicle transmitted from the vehicle 100, and speech information and position information transmitted from the communication device 120, via the network 140, estimating the user 130, and controlling the driving of the vehicle 100. The driving control of the vehicle 100 includes a process of adjusting the merging position between the user 130 and the vehicle 100.
[0016] The communication device 120 is, for example, a smartphone, but is not limited to this and may be an earphone-type communication terminal, a personal computer, a tablet terminal, a game console, etc. The communication device 120 connects to the network 140 via wireless communication such as Wi-Fi or fifth generation mobile communication.
[0017] The network 140 includes a communication network such as the Internet or a mobile phone network, and transmits information between the server 110 and the vehicle 100 or the communication device 120. In this system 1, when a user 130 and the vehicle 100, who are located far apart, approach each other close enough to visually identify a target (a visual landmark), the system estimates the user's location using speech information and image information captured by the vehicle 100, and adjusts the merging position. Note that in this embodiment, an example is described in which a camera capturing images of the surroundings of the vehicle 100 is installed on the vehicle itself, but a camera or the like does not necessarily have to be installed on the vehicle 100. For example, images captured by a surveillance camera or the like already installed around the vehicle 100 may be used, or both may be used. This allows images captured at a more optimal angle to be used when identifying the user's location. For example, when a user verbally describes their position relative to a landmark, the system analyzes an image captured by a camera close to the estimated position of the landmark, thereby more accurately identifying the user requesting to merge with the micromobility.
[0018] Before the user 130 and the vehicle 100 approach close enough to visually recognize a target, the server 110 first moves the vehicle 100 to a general area that includes the user's current location or the user's predicted location. When the vehicle 100 reaches the general area, the server 110 transmits to the communication device 120 visual landmarks and voice information inquiring about user-related information (e.g., "Are there any stores nearby?" or "Are your clothes black?") based on an image of the user 130 that is expected to have been taken. A location associated with a visual landmark may include, for example, a location name included in map information. Here, a visual landmark refers to a physical object visible to the user, and may include various objects such as a building, a traffic light, a river, a mountain, a statue, or a sign. The server 110 receives user speech information (e.g., "There is a building with a xx coffee shop here") from the communication device 120, the voice information including the location associated with the visual landmark. Then, the server 110 acquires the position of the relevant location from the map information and moves the vehicle 100 to the vicinity of the location (i.e., the vehicle approaches close enough for the user to visually confirm the target object, etc.). After that, according to this embodiment, efficient questions that reduce the number of questions are generated from captured images of the user's surroundings based on feature amounts predicted by an image recognition model, and the user is estimated from the user's answers to the questions. Details of the question generation method will be described later. Note that in this embodiment, a case in which a human user is estimated will be described, but other targets instead of humans may also be estimated. For example, a signboard, building, etc. designated by the user as a landmark may be estimated. In this case, the questions will target the other targets.
[0019] <Configuration of moving body> Next, the configuration of a vehicle 100 as an example of a moving body according to this embodiment will be described with reference to Fig. 2. Fig. 2(A) shows a side view of the vehicle 100 according to this embodiment, and Fig. 2(B) shows the internal configuration of the vehicle 100. In the figure, arrow X indicates the longitudinal direction of the vehicle 100, F indicates the front, and R indicates the rear. Arrows Y and Z indicate the width direction (left-right direction) and up-down direction of the vehicle 100.
[0020] Vehicle 100 is an electric autonomous vehicle equipped with a propulsion unit 12 and using a battery 13 as its main power source. Battery 13 is a secondary battery such as a lithium-ion battery, and vehicle 100 is propelled by propulsion unit 12 using power supplied from battery 13. Propulsion unit 12 is a four-wheeled vehicle equipped with a pair of left and right front wheels 20 and a pair of left and right rear wheels 21. Propulsion unit 12 may be in another form, such as a tricycle. Vehicle 100 is equipped with seating 14 for one or two people.
[0021] The traveling unit 12 includes a steering mechanism 22. The steering mechanism 22 is a mechanism that uses a motor 22a as a drive source to change the steering angle of the pair of front wheels 20. By changing the steering angle of the pair of front wheels 20, the traveling direction of the vehicle 100 can be changed. The traveling unit 12 also includes a drive mechanism 23. The drive mechanism 23 is a mechanism that uses a motor 23a as a drive source to rotate the pair of rear wheels 21. By rotating the pair of rear wheels 21, the vehicle 100 can move forward or backward.
[0022] The vehicle 100 is equipped with detection units 15 to 17 that detect targets around the vehicle 100. The detection units 15 to 17 are a group of external sensors that monitor the periphery of the vehicle 100, and in the present embodiment, each is an imaging device that captures an image of the periphery of the vehicle 100, and includes, for example, an optical system such as a lens and an image sensor. However, instead of or in addition to the imaging device, it is also possible to employ radar or lidar (Light Detection and Ranging).
[0023] Two detection units 15 are arranged at the front of the vehicle 100, spaced apart in the Y direction, and mainly detect targets in front of the vehicle 100. Detection units 16 are arranged on the left and right sides of the vehicle 100, respectively, and mainly detect targets on the sides of the vehicle 100. Detection unit 17 is arranged at the rear of the vehicle 100, and mainly detects targets behind the vehicle 100.
[0024] <Control structure of moving object> FIG. 3 is a block diagram of a control system of vehicle 100, which is a moving body. Here, the configuration necessary for implementing the present invention will be mainly described. Therefore, other configurations may be included in addition to the configurations described below. Vehicle 100 is equipped with a control unit (ECU) 30. Control unit 30 includes a processor represented by a CPU, a storage device such as a semiconductor memory, an interface with an external device, etc. The storage device stores programs executed by the processor and data used by the processor for processing, etc. Multiple sets of processors, storage devices, and interfaces may be provided for different functions of vehicle 100 and configured to be able to communicate with each other.
[0025] The control unit 30 acquires the detection results of the detection units 15 to 17, input information from the operation panel 31, audio information input from the audio input device 33, and control commands (such as transmission of captured images and current location) from the server 110, and executes corresponding processes. The control unit 30 controls the motors 22a and 23a (travel control of the traveling unit 12), controls the display on the operation panel 31, and outputs audio alerts and information to the occupants of the vehicle 100.
[0026] The voice input device 33 collects the voices of the occupants of the vehicle 100. The control unit 30 can recognize the input voices and execute corresponding processing. The GNSS (Global Navigation Satellite system) sensor 34 receives GNSS signals to detect the current position of the vehicle 100. The storage device 35 is a large-capacity storage device that stores map data including information on routes the vehicle 100 can travel, landmarks such as buildings, stores, etc. The storage device 35 may also store programs executed by the processor and data used by the processor for processing. The storage device 35 may store various parameters (e.g., trained parameters and hyperparameters of a deep neural network) of machine learning models for voice recognition and image recognition executed by the control unit 30. The communication unit 36 is a communication device that can be connected to the network 140 via wireless communication such as Wi-Fi or fifth-generation mobile communication.
[0027] <Server and communication device configuration> Next, a configuration example of the server 110 and the communication device 120 as an example of an information processing device according to this embodiment will be described with reference to Fig. 4. Note that the functions of the server 110 described below may be implemented in the vehicle 100 as shown in a modified example described later. In this case, a control unit 404 of the server 110 described later is implemented in a form integrated with the control unit 30 of the mobile body.
[0028] (Server configuration) First, an example configuration of the server 110 will be described. Here, the configuration necessary for implementing the present invention will be mainly described. Therefore, in addition to the configuration described below, other configurations may be included. The control unit 404 includes a processor such as a CPU, a storage device such as a semiconductor memory, an interface with external devices, etc. The storage device stores programs executed by the processor and data used by the processor for processing. Multiple sets of processors, storage devices, and interfaces may be provided for different functions of the server 110 and configured to be able to communicate with each other. The control unit 404 executes programs to perform various operations of the server 110 and the adjustment process for the merging position, which will be described later. In addition to the CPU, the control unit 404 may further include a GPU or dedicated hardware suitable for executing the processing of machine learning models such as neural networks.
[0029] The user data acquisition unit 413 acquires image and position information transmitted from the vehicle 100. The user data acquisition unit 413 also acquires at least one of utterance information of the user 130 transmitted from the communication device 120 and position information of the communication device 120. The user data acquisition unit 413 may store the acquired image and position information in the storage unit 403. The image and utterance information acquired by the user data acquisition unit 413 is input to a trained model in the inference stage to obtain an inference result, but may also be used as training data for training a machine learning model executed by the server 110.
[0030] The speech information processing unit 414 includes a machine learning model that processes speech information and executes learning-stage processing and inference-stage processing for the machine learning model. The machine learning model of the speech information processing unit 414 performs, for example, calculations of a deep learning algorithm using a deep neural network (DNN) to recognize place names, landmark names such as buildings, store names, and landmark names included in the speech information. The landmarks may include passersby, signs, road signs, outdoor facilities such as vending machines, building components such as windows and entrances, roads, vehicles, motorcycles, and the like included in the speech information. The DNN becomes trained by performing learning-stage processing, and by inputting new speech information into the trained DNN, recognition processing (inference-stage processing) for the new speech information can be performed. Note that, in this embodiment, an example is described in which the server 110 executes speech recognition processing; however, speech recognition processing may also be executed in a vehicle or a communication device, and the recognition results may be transmitted to the server 110.
[0031] The image information processing unit 415 includes a machine learning model that processes image information and executes the learning stage and inference stage processes of the machine learning model. The machine learning model of the image information processing unit 415 performs processing to recognize targets included in the image information, for example, by performing calculations of a deep learning algorithm using a deep neural network (DNN). Targets may include passersby, signs, road signs, outdoor facilities such as vending machines, building components such as windows and entrances, roads, vehicles, motorcycles, etc. included in the image. For example, the machine learning model of the image information processing unit 415 is an image recognition model that extracts features of passersby included in the image (for example, objects near passersby, the color of their clothes, the color of their bag, whether they are wearing a mask, whether they are wearing a smartphone, etc.).
[0032] The question generation unit 416 acquires impurity for each feature based on multiple feature values extracted by an image recognition model from an image captured by the vehicle 100 and their reliability, and recursively generates a set of questions that minimizes impurity in the shortest time based on the derived impurity. Impurity indicates the degree to which a target cannot be separated (from other targets) within a group of targets. The user estimation unit 417 estimates a user based on the user's answer to the generated question. Here, user estimation refers to estimating a user (target) who requests merging with the vehicle 100, and estimates the requesting user from one or more people within a specified area. The merging position estimation unit 418 executes adjustment processing for the merging position between the user 130 and the vehicle 100. Details of the impurity acquisition processing, user estimation processing, and merging position adjustment processing will be described later.
[0033] The server 110 generally has more abundant computational resources available than the vehicle 100. Furthermore, by receiving and storing image data captured by various vehicles, it is possible to collect learning data in a wide variety of situations, enabling learning to be performed in a wider range of situations. An image recognition model is generated from this stored information, and the image recognition model is used to extract features from the captured image.
[0034] The communication unit 401 is a communication device including, for example, a communication circuit, and communicates with external devices such as the vehicle 100 and the communication device 120. The communication unit 401 receives image information and position information from the vehicle 100, and at least one of speech information and position information from the communication device 120, and also transmits control commands to the vehicle 100 and speech information to the communication device 120. The power supply unit 402 supplies power to each component within the server 110. The storage unit 403 is a non-volatile memory such as a hard disk or semiconductor memory.
[0035] (Configuration of communication device) Next, the configuration of the communication device 120 will be described. The communication device 120 refers to a portable device such as a smartphone owned by the user 130. Here, the configuration necessary for implementing the present invention will be mainly described. Therefore, other configurations may be included in addition to the configurations described below. The communication device 120 includes a control unit 501, a storage unit 502, an external communication device 503, a display operation unit 504, a microphone 507, a speaker 508, and a speed sensor 509. The external communication device 503 includes a GPS 505 and a communication unit 506.
[0036] The control unit 501 includes a processor such as a CPU. The storage unit 502 stores programs executed by the processor, data used by the processor for processing, and the like. The storage unit 502 may be incorporated inside the control unit 501. The control unit 501 is connected to other components 502, 503, 504, 508, and 509 via signal lines such as buses, can send and receive signals, and controls the entire communication device 120.
[0037] The control unit 501 can communicate with the communication unit 401 of the server 110 via the network 140 using the communication unit 506 of the external communication device 503. The control unit 501 also acquires various information via the GPS 505. The GPS 505 acquires the current location of the communication device 120. This makes it possible to provide the server 110 with location information along with user speech information, for example. Note that the GPS 505 is not an essential component in the present invention, and the present invention provides a system that can be used even in facilities, such as indoors, where location information from the GPS 505 cannot be acquired. Therefore, the location information from the GPS 505 is treated as supplementary information when estimating the user.
[0038] The display operation unit 504 is, for example, a touch panel type liquid crystal display, and can display various information and accept user operations. The display operation unit 504 displays information such as the content of an inquiry from the server 110 and the merging position with the vehicle 100. When an inquiry is received from the server 110, the user's speech can be captured by the microphone 507 of the communication device 120 by operating a selectable microphone button. The microphone 507 captures the user's speech as audio information. The microphone may be activated by, for example, pressing a microphone button displayed on the operation screen, and may capture the user's speech. The speaker 508 outputs an audio message when making an inquiry to the user in accordance with an instruction from the server 110 (e.g., "Is the bag red?"). If the inquiry is audio, the communication device 120 can communicate with the user even if it has a simple configuration such as a headset without a display screen. Furthermore, even if the user does not have the communication device 120 in hand, the user can hear the inquiry from the server 110 through, for example, earphones. In the case of a text inquiry, the inquiry from the server 110 is displayed on the display operation unit of the communication device 120, and the user can obtain the user's response by pressing a button displayed on the operation screen or by entering text in a chat window. In this case, unlike when making an inquiry by voice, the inquiry can be made without being affected by surrounding environmental sounds (noise). The speed sensor 509 is an acceleration sensor that detects acceleration in the forward / backward, left / right, and up / down directions of the communication device 120. The output values indicating the acceleration output from the speed sensor 509 are stored in a ring buffer of the storage unit 502, and the oldest records are overwritten. The server 110 may acquire this data and use it to detect the direction of movement of the user.
[0039] <Overview of question generation using speech and images> 5 to 8, an overview of question generation using speech and images executed in server 110 will be described. Here, a process of generating efficient questions for identifying a target user or landmark such as a signboard from a captured image acquired by vehicle 100 will be described.
[0040] (Captured image) FIG. 5 is a diagram illustrating an example of a captured image acquired by the vehicle 100. In FIG. 5, the vehicle 100 has moved to a general location based on the user's speech information and location information. After moving to the general location, the vehicle 100 captures an image of the area around the estimated location of the target user using at least one of the detection units 15 to 17. The captured image 600 includes pedestrians A, B, C, and D, a building 601, a utility pole 602, and crosswalks 603 and 604 on the road. After acquiring the captured image 600, the vehicle 100 transmits the captured image 600 to the server 110. If the vehicle 100 has an image recognition model, the vehicle 100 may extract features from the captured image. If the vehicle 100 does not have an imaging function, the vehicle 100 may acquire images captured using cameras installed in other vehicles or buildings in the vicinity. Furthermore, image analysis may be performed using these multiple captured images.
[0041] (Feature extraction) FIG. 6 is a diagram showing feature amounts extracted from a captured image 600 by the image recognition model in the server 110. Reference numeral 610 denotes extracted features (hereinafter referred to as feature amounts). The image information processing unit 415 of the server 110 first detects people using the image recognition model. Here, four passersby A to D are detected in the captured image 600. The image information processing unit 415 then extracts feature amounts for each detected person. As shown in 610, features related to the detected people include, for example, objects located near the detected people, the color and type of the detected people's clothing, the color of their pants, and the color of their bag. Furthermore, the behavior of the detected people, for example, whether they are looking at their smartphone, whether they are wearing a mask, whether they are standing, and the direction they are facing, are detected. As shown in 610, feature amounts are extracted for each of the detected passersby A to D. Furthermore, if the target object is a building or a signboard, features may include objects located near the detected object, the color and type of the detected object, and text or patterns displayed on the object. (Question generation according to impurity) FIG. 7 illustrates a question generation method using impurity according to this embodiment. First, the question generation unit 416 of the server 110 extracts one or more features using an image recognition model, and then acquires the feature values, their reliability, and the weights of the features themselves. The reliability is a value indicating, for example, how confident the image recognition model is in predicting the feature values. The weight is a value indicating how much the feature is reflected in the impurity calculation. The reliability and weight may be values that are updated as needed by machine learning. The weights of the feature values can also be set heuristically for each feature value. Furthermore, the question generation unit 416 recursively generates optimal and efficient questions based on the acquired features, their weights, and reliability. It is desirable that the generated questions be questions that humans can answer with a yes / no, thereby reducing the variety of answers. In other words, this has the secondary effect of reducing the difficulty of speech understanding and speech recognition by computers.
[0042] The example shown in Figure 7 will be explained. As shown in 610, feature amounts are extracted for passersby A to D from a captured image 600. Among these, as shown in 701, the target user who has requested merging is designated as B. As mentioned above, impurity indicates the degree to which a target cannot be separated (from other targets) within a group of targets. Therefore, when all passersby A to D are included, the impurity is calculated as "4.8" using the impurity calculation model described below.
[0043] Here, if the weights and reliability of all feature quantities are equal, the question generation unit 416 generates a question that minimizes impurity in the shortest time possible, i.e., a question asking about features possessed by only one user, such as "Is the bag red?" Of course, if there is no feature possessed by only one user, multiple questions may be generated. In that case, the questions may be asked sequentially, or a question may be asked based on features possessed by a user that is deemed more likely based on other information, such as the user's location information. In the example of 610, if the user answers "Yes" to the above question, passerby B can be estimated as the target user. On the other hand, if the user answers "No," the set is narrowed down to passersby A, C, and D, and the next question is generated.
[0044] On the other hand, if the weight and reliability of the bag color are low, the question generation unit 416 generates a question, for example, "Are you looking at your smartphone?", using other features with high weights and reliability. If the user answers "Yes," the set is narrowed down to passersby A and B, and the impurity becomes "1.9." Next, the question generation unit 416 generates the question "Are you wearing a mask?". This makes it possible to estimate the target user regardless of whether the user answers "Yes" or "No." In this way, the question generation unit 416 generates optimal and efficient questions by taking into account the weights of the feature values and the reliability of the feature values.
[0045] The impurity calculation model can be formulated in various ways. For example, it can be formulated heuristically or by function approximation using neural networks. As mentioned above, the weights of the features can be set heuristically or learned from data using machine learning.
[0046] An example of an impurity calculation model is shown in 702 in Figure 7. 703 indicates the number of objects other than the target included in the set. For example, if the target is a person, it indicates the number of objects other than the specified person included in a set of multiple people. The smaller N is, the smaller the impurity. 704 indicates the penalty based on the weight of the feature and the reliability of the feature value. The smaller the penalty, the smaller the impurity. 705 indicates the content of each variable. Furthermore, F indicates the set of each feature (set of feature values), and M indicates the number of dimensions of the feature. f k indicates the set of feature values that each object has for the kth feature. * k indicates the feature values of the target user. N indicates the number of objects. w indicates the set of weights for each feature. C fk indicates the reliability of the k-th feature obtained from the image recognition result for each object. Note that the impurity calculation model 702 is merely an example and is not intended to limit the present invention. For example, instead of simply calculating the sum of each term 702 and 703, it is possible to introduce a coefficient or normalization based on the number of objects. Also, for the penalty term, instead of simply calculating the inverse of the weight or reliability, it is possible to introduce other operations or functions. Furthermore, function approximation using a neural network or the like may be introduced depending on the amount of data collected.
[0047] (Efficient generated questions) FIG. 8 shows an example of an efficient question according to this embodiment and a comparative example. In this comparative example, questions are sequentially generated using the extracted features shown in 610 to narrow down the target user. Therefore, multiple questions are likely to be generated. As shown in FIG. 8, questions such as "Are there any buildings nearby?", which are characteristic of all passersby A to D, and "Are your clothes black?", which are characteristic of passersby A and B, may be generated. On the other hand, according to the present invention, as described above with reference to FIG. 7, a question such as "Are your shoes red?" is generated using features possessed by as few passersby as possible. For example, if passerby B is the target user, a "Yes" answer is accepted, allowing the target user to be identified with a single question. As described above, this embodiment minimizes impurity in the shortest possible time, thereby minimizing the number of interactions required to identify the target user.
[0048] <Merge control processing steps> Next, a series of operations of merging control in the server 110 according to this embodiment will be described with reference to FIG. 9. This processing is realized by the control unit 404 executing a program. In the following description, for simplicity's sake, it is assumed that the control unit 404 executes each process, but the corresponding process is executed by each part of the control unit 404. Here, a flow in which a user and a vehicle finally merge will be described, but a characteristic configuration of the present invention is a configuration related to estimation (identification) of the user, and a configuration for estimating the merging position is not an essential configuration. In other words, although a processing procedure including control related to estimation of the merging position will be described below, control may be performed to execute only the processing procedure related to estimation of the user.
[0049] In S101, the control unit 404 receives a request (a merge request) to start merging with the vehicle 100 from the communication device 120. In S102, the control unit 404 acquires user position information from the communication device 120. The user position information is position information acquired by the GPS 505 of the communication device 120. The position information may also be received simultaneously with the request of S101. In S103, the control unit 404 identifies a rough area for merging (also simply referred to as a merging area or a predetermined region) based on the user position acquired in S102. The merging area is, for example, an area with a radius of a predetermined distance (for example, several hundred meters) centered on the current position of the user 130 (communication device 120).
[0050] In S104, the control unit 404 tracks the movement of the vehicle 100 heading toward the meeting area, for example, based on location information periodically transmitted from the vehicle 100. Note that the control unit 404 can select, for example, from among a plurality of vehicles located around the current location of the user 130 (or a destination point after a predetermined time), the vehicle closest to the current location as the vehicle 100 that will meet with the user 130. Alternatively, if information specifying a specific vehicle 100 is included in the meeting request, the control unit 404 may select the vehicle 100 as the vehicle 100 that will meet with the user 130.
[0051] In S105, the control unit 404 determines whether the vehicle 100 has reached the merging area. For example, if the distance between the vehicle 100 and the communication device 120 is within the radius of the merging area, the control unit 404 determines that the vehicle 100 has reached the merging area and proceeds to S106. If not, the server 110 returns the process to S105 and waits for the vehicle 100 to reach the merging area.
[0052] In S106, the control unit 404 estimates the user using the user's utterance and the captured image. Details of the user estimation process using the user's utterance and the captured image will be described later. Subsequently, in S107, the control unit 404 further estimates a junction position based on the user estimated in S106. For example, by estimating the user in the captured image, if the user utters something like "a nearby red postbox" as the junction position, the junction position can be more accurately estimated by searching for a red postbox closest to the estimated user. Thereafter, in S108, the control unit 404 transmits position information of the junction position to the vehicle. That is, the control unit 404 transmits the junction position estimated in the process of S107 to the vehicle 100, thereby moving the vehicle 100 to the junction position. After transmitting the junction position to the vehicle 100, the control unit 404 terminates the series of operations.
[0053] Next, a series of operations of the user estimation process (S106) using utterances and captured images in the server 110 will be described with reference to Fig. 10. Note that this process is realized by the control unit 404 executing a program, similar to the process shown in Fig. 9.
[0054] In S201, the control unit 404 acquires an image captured by the vehicle 100. Note that images may also be acquired from a surveillance camera installed in a vehicle other than the vehicle 100 or in a building in the vicinity where the target user is thought to be located.
[0055] In S202, the control unit 404 uses an image recognition model to detect one or more people included in the acquired captured image. Subsequently, in S203, the control unit 404 uses the image recognition model to extract features for each detected person. As a result of the processing in S202 and S203, for example, people and their respective features shown in 610 in FIG. 6 are extracted. Note that here, weights and reliability are also assigned to the extracted feature amounts.
[0056] Next, in S204, the control unit 404 obtains the impurity for each feature extracted in S203 using the above-mentioned formula. Subsequently, in S205, the control unit 404 generates a question that minimizes the number of questions based on the impurity.
[0057] In S206, the control unit 404 sends a question to the user according to the generated question, and repeats the question according to the user's answer until the user can be identified, and then ends the processing of this flowchart. The detailed processing will be described later with reference to FIG. 11.
[0058] The detailed processing of S206 will be described with reference to Fig. 11. Note that this processing is realized by the control unit 404 executing a program, similar to the processing shown in Fig. 9.
[0059] In S301, the control unit 404 transmits the question in the generated question set that has been asked the least number of times based on the weights and reliability of the features related to each question and the number of times the question has been asked to the communication device 120. Here, the question set is a set that includes one or more questions, and indicates a set from which a target user can be estimated by having a dialogue with the user based on the questions in the question set.
[0060] Next, in S302, the control unit 404 determines whether a user answer to the question sent in S301 has been received from the communication device 120. If it has been received, the process proceeds to S303; if not, the process waits in S302 until it is received. Note that if a predetermined time or more has passed since the question was sent but the user answer has not been received, the question may be sent again, or the process may end with an error.
[0061] In S303, the control unit 404 determines whether the target user can be narrowed down based on the user's answers. That is, if the user can be estimated, the process proceeds to S304, and if not, the process returns to S301 to send the next question. In S304, the control unit 404 estimates the target user and ends the process of this flowchart.
[0062] <Modification> Modifications of the present invention will be described below. In the above embodiment, an example has been described in which merging control including user estimation is performed by the server 110. However, the above-described processing can also be performed by a moving body such as a vehicle or a walking robot. In this case, as shown in FIG. 12, the system 1200 includes a vehicle 1210 and a communication device 120. User utterance information is transmitted from the communication device 120 to the vehicle 1210. Image information captured by the vehicle 1210 is processed by a control unit within the vehicle instead of being transmitted via a network. The configuration of the vehicle 1210 may be the same as that of the vehicle 100, except that the control unit 30 can perform merging control. The control unit 30 of the vehicle 1210 operates as a control device for the vehicle 1210 and executes a stored program to perform the above-described processing. In the series of operations shown in FIGS. 9 to 11, communication between the server and the vehicle may be performed within the vehicle (for example, within the control unit 30 or between the control unit 30 and the detection unit 15). Other processing can be performed in the same manner as with the server.
[0063] <Summary of the embodiment> 1. The information processing device (e.g., 110) of the above embodiment is An acquisition means (401) for acquiring a captured image; an extraction means (415, S203) for detecting a plurality of targets included in the captured image and extracting a plurality of feature amounts for each of the detected plurality of targets; an acquisition means (415, S204) for acquiring, for each feature extracted by the extraction means, an impurity indicating a degree to which the predetermined target cannot be separated from the plurality of targets when a question for estimating a predetermined target from the plurality of targets based on each feature is asked to a user; and a generation means (416, S205) for generating questions based on the feature extracted by the extraction means and the impurity for each feature, in order to reduce the number of questions asked to minimize the impurity.
[0064] According to this embodiment, it is possible to generate efficient questions using the feature amounts of image recognition and estimate the target object.
[0065] 2. In the information processing device of the above embodiment, the extraction means extracts the features using an image recognition model (S203), and the generation means generates the question that minimizes the impurity in the shortest time based on the features and the impurity, as well as the reliability and weight of the features extracted using the image recognition model (S205).
[0066] According to this embodiment, feature extraction can be performed efficiently using a trained image recognition model, and optimal questions can be generated according to the reliability and weighting of the model.
[0067] 3. In the information processing device of the above embodiment, the reliability indicates the reliability of feature values indicating the values of the feature values extracted by the image recognition model for each of the plurality of targets (FIG. 7). Furthermore, the weight is set heuristically or based on machine learning for each feature (FIG. 7).
[0068] According to this embodiment, feature extraction can be performed efficiently using a trained image recognition model, optimal questions can be generated according to the reliability and weight of the feature, and the weight of each feature can be set appropriately.
[0069] 4. In the information processing device of the above embodiment, the impurity is obtained based on at least one of the number of targets other than the specified target included in the set of targets, and a penalty based on the weight and / or reliability of the feature (Figure 7).
[0070] According to this embodiment, impurity is derived while taking into consideration the reliability and weight of each feature amount, and efficient question generation can be performed.
[0071] 5. The information processing device of the above embodiment further includes a transmitting means (401, S301) that transmits the question generated by the generating means to a communication device owned by the user, a receiving means (401, S302) that receives an answer to the question from the communication device, and an estimating means (417, S304) that estimates the specified target from among the multiple targets according to the answer received by the receiving means.
[0072] According to this embodiment, targets such as a user can be efficiently estimated according to a query generated so as to minimize impurity in the shortest time possible.
[0073] 6. In the information processing device of the above embodiment, the acquisition means acquires location information from a communication device owned by the user, and acquires captured images of the area around the location information from outside (401, 413).
[0074] According to this embodiment, the user's approximate location can be identified, and captured images of the surrounding area can be used to generate questions.
[0075] 7. In the information processing device of the above embodiment, the acquisition means acquires an image captured by a vehicle into which the user requests merging from the vehicle (15 to 17, S201).
[0076] According to this embodiment, it is possible to estimate the target more accurately and meet up with the target user.
[0077] 8. In the information processing device of the above embodiment, the acquisition means acquires, from a camera installed in the vicinity of the location information, an image captured by the camera.
[0078] According to this embodiment, even if the vehicle does not have an imaging function, it is possible to acquire an image of the surroundings of the target user.
[0079] 9. In the information processing device of the above embodiment, when the target is a person, the feature is at least one piece of information indicating nearby objects, the color and type of clothes, the color of a bag, whether the person is looking at a communication device, and whether the person is wearing a mask (FIG. 8). Also, the feature is at least one piece of information among the color, type, characters, and designs displayed on the target.
[0080] According to this embodiment, it is possible to efficiently estimate the target (including the user who is the target) based on various feature amounts.
[0081] 10. The mobile unit (e.g., 1210) of the above embodiment is An acquisition means (401) for acquiring a captured image; an extraction means (415, S203) for detecting a plurality of targets included in the captured image and extracting a plurality of feature amounts for each of the detected plurality of targets; an acquisition means (415, S204) for acquiring, for each feature extracted by the extraction means, an impurity indicating a degree to which the predetermined target cannot be separated from the plurality of targets when a question for estimating a predetermined target from the plurality of targets based on each feature is asked to a user; and a generation means (416, S205) for generating questions based on the feature extracted by the extraction means and the impurity for each feature, in order to reduce the number of questions asked to minimize the impurity.
[0082] According to this embodiment, it is possible to generate efficient questions and estimate targets in a moving body using image recognition features without going through a server. [Explanation of symbols]
[0083] 100, 1210...vehicle, 110...server, 120...communication device, 404...control unit, 413...user data acquisition unit, 414...voice information processing unit, 415...image information processing unit, 416...merging position estimation unit, 417...user estimation unit
Claims
1. An information processing device, an acquisition means for acquiring a captured image; an extraction means for detecting a plurality of targets included in the captured image and extracting a plurality of feature amounts for each of the detected targets; an acquisition means for acquiring, for each feature extracted by the extraction means, an impurity indicating a degree to which the predetermined target cannot be separated from the plurality of targets when a question for estimating a predetermined target from the plurality of targets based on each feature is asked to a user; a generation means for generating the questions based on the feature amounts extracted by the extraction means and the impurity for each feature amount, so as to reduce the number of questions asked to minimize the impurity; An information processing device comprising:
2. the extraction means extracts the feature amount using an image recognition model; The information processing device according to claim 1, characterized in that the generation means generates the question that minimizes the impurity in the shortest time based on the feature and the impurity, as well as the reliability and weight of the feature extracted using the image recognition model.
3. The information processing apparatus according to claim 2 , wherein the reliability indicates a reliability of a feature value indicating a value of a feature extracted by the image recognition model for each of the plurality of targets.
4. The information processing apparatus according to claim 2 , wherein the weight is set heuristically or based on machine learning for each feature amount.
5. The information processing device according to any one of claims 2 to 4, characterized in that the impurity is obtained according to at least one of the number of targets other than the specified target included in the set of multiple targets and a penalty based on the weight and / or reliability of the feature.
6. a transmission means for transmitting the question generated by the generation means to a communication device owned by the user; receiving means for receiving an answer to the question from the communication device; an estimation means for estimating the predetermined target from among the plurality of targets in accordance with the response received by the receiving means; 6. The information processing apparatus according to claim 1, further comprising:
7. 7. The information processing apparatus according to claim 1, wherein the acquiring means acquires the location information from a communication device owned by the user, and acquires an image of an area around the location information from an external device.
8. 8. The information processing apparatus according to claim 7, wherein the acquisition means acquires an image captured by a vehicle into which the user is requesting to merge from the vehicle.
9. 8. The information processing apparatus according to claim 7, wherein the acquisition means acquires an image captured by a camera installed in the vicinity of the location information from the camera.
10. 10. The information processing device according to claim 1, wherein the feature is at least one piece of information indicating, when the target is a person, nearby objects, the color of clothes, the type of clothes, the color of a bag, the type of bag, whether the person is looking at a communication device, and whether the person is wearing a mask.
11. 11. The information processing apparatus according to claim 1, wherein the feature amount is at least one of information on the color, type, character, and design of the target.
12. A mobile object, an acquisition means for acquiring a captured image; an extraction means for detecting a plurality of targets included in the captured image and extracting a plurality of feature amounts for each of the detected targets; an acquisition means for acquiring, for each feature extracted by the extraction means, an impurity indicating a degree to which the predetermined target cannot be separated from the plurality of targets when a question for estimating a predetermined target from the plurality of targets is asked to a user based on each feature; a generation means for generating the questions based on the feature amounts extracted by the extraction means and the impurity for each feature amount, so as to reduce the number of questions for minimizing the impurity; A moving object comprising:
13. A control method for an information processing device, comprising: an acquisition step of acquiring a captured image; an extraction step of detecting a plurality of targets included in the captured image and extracting a plurality of feature amounts for each of the detected plurality of targets; an acquisition step of acquiring, for each feature extracted in the extraction step, an impurity indicating a degree to which the predetermined target cannot be separated from the plurality of targets when a question for estimating a predetermined target from the plurality of targets based on each feature is asked to a user; a generating step of generating questions based on the feature amounts extracted in the extracting step and the impurity for each feature amount, so as to reduce the number of questions for minimizing the impurity; 10. A method for controlling an information processing device, comprising:
14. A method for controlling a moving object, comprising: an acquisition step of acquiring a captured image; an extraction step of detecting a plurality of targets included in the captured image and extracting a plurality of feature amounts for each of the detected plurality of targets; an acquisition step of acquiring, for each feature extracted in the extraction step, an impurity indicating a degree to which the predetermined target cannot be separated from the plurality of targets when a question for estimating a predetermined target from the plurality of targets based on each feature is asked to a user; a generating step of generating questions based on the feature amounts extracted in the extracting step and the impurity for each feature amount, so as to reduce the number of questions for minimizing the impurity; A method for controlling a moving object, comprising:
15. A program for causing a computer to function as each of the means of the information processing device according to any one of claims 1 to 11.
16. A program for causing a computer to function as each means of a mobile body according to claim 12.
17. A storage medium storing a program for causing a computer to function as each of the means of the information processing device according to any one of claims 1 to 11.
18. A storage medium storing a program for causing a computer to function as each of the means of the mobile body according to claim 12.
Citation Information
Patent Citations
Device for constructing interactive dictionary
JP1993158980A
Information retrieving device
JP1998187739A
Decision tree generation device, decision tree generation method, decision tree generation program and query system
JP2018005624A
Recognizing passengers assigned to autonomous vehicles
JP2022016448A
Generating labels for images associated with a user
US20170185670A1