Mobile device, control device, and control method
The mobile device employs gesture-driven tutorials and voice guidance to intuitively register user faces, addressing the limitations of conventional methods by improving user interaction and accuracy in face registration processes.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-04-01
- Publication Date
- 2026-03-25
AI Technical Summary
Conventional face registration processes for electronic devices lack intuitiveness, relying on displayed characters and images for guiding face orientation, which can be cumbersome and less user-friendly.
A mobile device and control method that utilize gesture guidance voice and gesture-driven tutorials to guide face orientation during registration, using a mobile object with a hemispherical face part and camera to capture face images at various angles, accompanied by voice prompts to facilitate intuitive face registration.
Enables intuitive and efficient face registration without the need for displays, improving user experience and accuracy by guiding facial movements through gestures and voice, reducing errors and enhancing face recognition capabilities.
Smart Images

Figure 0007835091000001 
Figure 0007835091000002 
Figure 0007835091000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a mobile body, a control device, and a control method, and particularly to a mobile body, a control device, and a control method that enable intuitive face registration.
Background Art
[0002] In recent years, in various electronic devices such as smartphones, a method of using face authentication based on face information of a pre-registered user to unlock when it is identified that the face is the user's own face has become widespread.
[0003] For example, Patent Document 1 discloses a face recognition device that identifies an input face image by evaluating the similarity with a registered face image registered in a registered face group selected based on the input result of an image or voice, and confirms that the identified input face image is the person of the registered face image.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] By the way, in the conventional face registration process for pre-registering user's face information, the orientation of the user's face is guided using characters and images displayed on the display, but there is a demand for more intuitive face registration.
[0006] The present disclosure has been made in view of such a situation, and enables intuitive face registration.
Means for Solving the Problems
[0007] One aspect of the present disclosure of a mobile device and control device includes a gesture control unit that controls a gesture drive that expresses the movement of the user's face when streaming shooting is performed in a later stage, during a tutorial of a face registration process in which the user's face is registered in advance, and a guidance voice control unit that controls the output of a gesture guidance voice that matches the gesture, together with the gesture drive.
[0008] One aspect of the control method of this disclosure includes controlling a gesture drive that expresses the movements of the user's face when streaming footage is taken in a later stage, during a tutorial of the face registration process in which the user's face is registered in advance, and controlling the output of a gesture guidance voice that matches the gesture, together with the gesture drive.
[0009] In one aspect of this disclosure, during the tutorial for the face registration process in which the user's face is registered in advance, gesture driving is controlled to represent the user's facial movements during streaming recording in a later stage of processing. Along with the gesture driving, the output of gesture guidance audio corresponding to the gesture is also controlled. [Brief explanation of the drawing]
[0010] [Figure 1] This figure shows an example of how this technology is being used in mobile devices. [Figure 2] This diagram illustrates an example of gesture-guided voice guidance and gesture activation. [Figure 3] This is a block diagram showing an example configuration of one embodiment of a mobile body. [Figure 4] This diagram illustrates the feature vector and the center vector. [Figure 5] This is a diagram explaining thresholds. [Figure 6] This is a flowchart explaining the face registration process. [Figure 7] This is a block diagram showing an example configuration of one embodiment of a computer to which this technology is applied. [Modes for carrying out the invention]
[0011] Hereinafter, specific embodiments to which the present technology is applied will be described in detail with reference to the drawings.
[0012] <Usage example of the mobile object> FIG. 1 is a diagram for explaining the usage situation of a mobile object to which the present technology is applied. FIG. 1 shows an example of a situation where a user is using a mobile object 11 placed on a table as viewed from the side.
[0013] For example, the mobile object 11 is an agent-type robot device capable of autonomous driving, enabling more natural and effective communication with the user. Also, the mobile object 11 is, for example, a small robot configured to be of a size and weight that can be easily lifted by the user with one hand.
[0014] The mobile object 11 has a hemispherical face part 13 provided on the upper part of an elliptical main body 12 in the vertical direction. A camera 14 and an eye part 15 are provided on the front side of the face part 13 (the side facing right in FIG. 1), and tires 16 are provided on the bottom surface of the main body 12.
[0015] The face part 13 is configured to be able to freely change its orientation in the vertical and horizontal directions by a drive mechanism built into the mobile object 11.
[0016] The camera 14 photographs the direction facing the front of the face part 13 and acquires a still image or a moving image. For example, as shown in FIG. 1, when the user is facing the face part 13 directly, the camera 14 can acquire a face image by photographing the user's face.
[0017] The eye part 15 is composed of, for example, an LED (Light Emitting Diode) or an organic EL (Electro Luminescence), etc., and can express a line of sight, a blink, etc. In FIG. 1, only one eye part 15 is shown, but as shown in FIG. 2, two eye parts 15L and 15R are provided side by side left and right when viewing the face part 13 from the front.
[0018] The tire 16 can rotate freely by a drive mechanism built in the moving body 11, and realizes moving operations such as the forward movement, backward movement, turning, and rotation of the moving body 11.
[0019] The moving body 11 configured as described above can pre-register the user's face information in a face database, and can realize communication suitable for each user by performing face authentication processing when the user uses it.
[0020] When performing face authentication processing, the moving body 11 outputs, for example, a guidance voice "Look at me" from a speaker (not shown). In response to this, when the user approaches the face so as to look at the face part 13 of the moving body 11, the moving body 11 acquires face information from a face image obtained by photographing the user's face with the camera 14. Then, the moving body 11 can identify the face of each individual user by performing face authentication processing for evaluating the similarity between the face information of that user and each of the plurality of face information registered in the face database.
[0021] For example, in face authentication processing using a feature vector as face information, the feature vector obtained from the user's face photographed by the camera 14 is used as the identification target, and the distance (cosine distance (similarity), Euclidean distance, etc.) between the feature vector of the identification target and the registered feature vector is calculated one by one. Then, in face authentication processing, it can be identified that the face of the registered feature vector whose distance is greater than or equal to a predetermined threshold and the face of the user photographed by the camera 14 are the same person.
[0022] By the way, in the face registration process, which involves pre-registering the user's facial information, it is necessary to guide the user to change the direction of their face towards the camera in the following directions: straight ahead, to the right, to the left, up, and down. For example, when the face registration process is performed on a smartphone, the user is guided to change the direction of their face using text and images displayed on the screen.
[0023] In contrast, the mobile device 11 is configured to perform a face registration process that is intuitive and easy for the user to understand, without using a display, by using gesture guidance voice and gesture-driven tutorials.
[0024] Referring to Figure 2, an example of gesture guidance audio and gesture activation in a tutorial when the mobile device 11 performs face registration processing will be explained.
[0025] For example, gesture-driven operation expresses the user's facial movements (speed and orientation) during streaming recording using gestures of the face portion 13 of the mobile device 11, and gesture guidance audio outputs a constant rhythm corresponding to the speed of the face portion 13 of the mobile device 11 in accordance with those gestures.
[0026] First, as shown in Figure 2A, the mobile unit 11, with its face unit 13 facing forward, outputs a pre-guidance voice message saying, "Move your face slowly in time with the sound," to explain how to move the face, and then starts the tutorial.
[0027] For example, in the tutorial, as shown in Figure 2B, the mobile unit 11 outputs a gesture guidance voice at a constant rhythm, saying "One, two, three, four, five!", and performs a gesture drive that rotates the face unit 13 to face to the right at a constant speed. Similarly, as shown in Figure 2B, the mobile unit 11 outputs a gesture guidance voice at a constant rhythm, saying "One, two, three, four, five!", and performs a gesture drive that rotates the face unit 13 to face to the left at a constant speed.
[0028] Furthermore, as shown in Figure 2D, the mobile body 11 outputs a gesture guidance voice with a constant rhythm, saying "One, two, three, four, five!", and performs a gesture drive that rotates the face part 13 so that it faces upwards at a constant speed. Similarly, as shown in Figure 2E, the mobile body 11 outputs a gesture guidance voice with a constant rhythm, saying "One, two, three, four, five!", and performs a gesture drive that rotates the face part 13 so that it faces downwards at a constant speed.
[0029] After the tutorial is completed, the mobile unit 11 starts streaming recording, as will be described later with reference to the flowchart in Figure 6, and outputs a voice prompt that guides the user's face direction, similar to the gesture guidance voice prompt used when gesture driving is performed, saying "One, two, three, four, five!". At this time, the mobile unit 11 streams recording of the user's face while keeping the face unit 13 fixed facing forward, and guides the user to change the direction of their face in the right, left, up, and down directions in sequence. The mobile unit 11 also guides the user to ensure that their face is always facing forward while changing the direction of their face to the right, left, up, and down.
[0030] Note that in the tutorial, it is not necessary to perform gestures in all four directions: right, left, up, and down. Performing gestures in at least one direction is sufficient. For example, the moving object 11 may perform gestures in one direction (left / right) and one direction (up / down) in the tutorial.
[0031] <Example of a mobile device configuration> Figure 3 is a block diagram showing an example configuration of one embodiment of a mobile body.
[0032] As shown in Figure 3, the mobile unit 11 is configured to include an audio output unit 21, a drive unit 22, an imaging unit 23, a storage unit 24, a face registration processing unit 25, and a threshold setting unit 26. The face registration processing unit 25 performs face registration processing and includes a guidance voice control unit 31, a gesture control unit 32, a feature vector extraction unit 33, and a center vector calculation unit 34.
[0033] The audio output unit 21 is configured, for example, with a speaker, and outputs guidance audio necessary for guiding the user during the face registration process, in accordance with the control of the guidance audio control unit 31.
[0034] The drive unit 22 is composed of, for example, a motor, and performs gesture driving to rotate the face unit 13 according to the control of the gesture control unit 32, as described with reference to Figure 2, so that the face unit 13 faces to the right, left, up, and down.
[0035] The imaging unit 23 is composed of, for example, an image sensor in the camera 14, and can acquire an image by photographing a subject in front of the face unit 13. For example, it acquires a face image by streaming the user's face and supplies it to the feature vector extraction unit 33.
[0036] The memory unit 24 is composed of non-volatile memory such as flash memory, and registers the center vector calculated by the center vector calculation unit 34 in the face registration process in the face database.
[0037] The threshold setting unit 26 sets a threshold used in the face recognition process, which evaluates the similarity with the center vector registered in the face database, and stores it in the storage unit 24. The threshold set by the threshold setting unit 26 will be described later with reference to Figure 5.
[0038] During the tutorial, the guidance voice control unit 31 controls the output of gesture guidance voices that match the gestures of the face part 13 of the mobile body 11, that is, the output of pre-guidance voices and gesture guidance voices as explained with reference to Figure 2. Furthermore, the guidance voice control unit 31 controls the output of start guidance voices, face-direction guidance voices, end guidance voices, etc., as explained later with reference to the flowchart in Figure 6, and outputs the guidance voices from the voice output unit 21.
[0039] During the tutorial, the gesture control unit 32 controls the gesture drive of the face portion 13 of the mobile body 11 to represent the user's facial movements during streaming, that is, the speed and direction of the user's facial movements. In other words, as explained with reference to Figure 2, the gesture control unit 32 controls the drive unit 22 to perform a gesture drive that rotates the face portion 13 to face right, left, up, and down at a constant speed in accordance with the rhythm of the gesture guidance voice.
[0040] The feature vector extraction unit 33 extracts multiple feature vectors from face images at various angles acquired by streaming imaging by the imaging unit 23 and supplies them to the center vector calculation unit 34.
[0041] The center vector calculation unit 34 calculates the center vector which is the center of all the feature vectors supplied from the center vector calculation unit 34.
[0042] Here, we will explain the feature vector and the center vector with reference to Figure 4.
[0043] As described above, the mobile device 11 performs streaming photography during the face registration process, and the feature vector extraction unit 33 extracts multiple feature vectors from face images at various angles acquired by the streaming photography. Figure 4 shows an image of multiple feature vectors, but in reality, the feature vectors are vectors on a 512-dimensional hypersphere.
[0044] The center vector calculation unit 34 then calculates the center vector using all the face images acquired through streaming. In other words, the result is stable by using the center of the feature vector extracted from face images at various angles. The center vector calculation unit 34 may also calculate the center vector when a predetermined number (for example, 50) of feature vectors have been accumulated. The feature vector extraction unit 33 is pre-trained so that the facial features of the same person are close together, and the facial features of different people are far apart.
[0045] Furthermore, as shown in Figure 5, the threshold setting unit 26 sets a fixed threshold D θ , maximum distance threshold D R , and a fixed threshold D θ and the threshold D of the maximum distance R The maximum value among these is set as the threshold used in the facial recognition process. For example, a fixed threshold D θ This is a value determined during the design of the mobile body 11, and is the maximum distance threshold D. R This value corresponds to the distance from the center vector to the feature vector located furthest away.
[0046] By using thresholds set in this manner, the mobile device 11 can relatively easily verify and implement facial recognition processing, and can also reduce the processing load.
[0047] When the mobile device 11 performs facial recognition processing, it calculates the distance between the feature vector extracted by the feature vector extraction unit 33 and the registered center vector, and identifies the center vector with the closest distance that also falls within the threshold range set by the threshold setting unit 26 as the same.
[0048] Here, the mobile device 11 is equipped with a feature extractor that has been trained to minimize the distance between the center vector, or representative vector, of each face class and the feature vector of the same face class. In other words, this feature extractor distributes the feature vectors of a person's face at various angles so that they spread out from the center vector, or representative vector. Therefore, the center of the feature vectors of face images at various angles collected by streaming capture can capture approximately the center of the distribution of feature vectors of that face. Furthermore, by normalizing the feature vectors output by the feature extractor and adding the constraint that they lie on a hypersphere, it is possible to avoid being affected by the length of the vectors, so the center of the feature vectors of face images collected by streaming capture can be expected to be stable. On the other hand, a feature extractor that learns to simply bring the feature vectors of the same face closer together and move the feature vectors of different faces further apart, rather than minimizing the distance between the center vector or representative vector of each face class and the feature vector of the same face class, does not guarantee that the shape of the distribution will spread out concentrically (strictly speaking, not a circle but a hypersphere because it is multidimensional), and good accuracy cannot be expected when performing face recognition using the centers of the feature vectors of face images collected by streaming.
[0049] <Example of face registration processing> The face registration process performed by the face registration processing unit 25 will be explained with reference to the flowchart shown in Figure 6.
[0050] For example, when a user uses the mobile device 11 for the first time and says "Let's be friends," the facial registration process is initiated as a result of the speech recognition processing performed on that voice.
[0051] In step S11, the guidance voice control unit 31 controls the output of the pre-guidance voice and causes the voice output unit 21 to output the pre-guidance voice. For example, the voice output unit 21 outputs the pre-guidance voice, "Let me remember your face for a moment as a memento of our meeting," and then outputs the pre-guidance voice, "I'll remember you from various angles," to explain to the user that the face will be photographed from multiple angles. If the user is the owner of the mobile device 11, the user's name may be registered in advance, and the user's name may be confirmed at the start of the face registration process. Then, the voice output unit 21 outputs the pre-guidance voice, "I'll show you how it's done first!" to explain to the user that the tutorial will begin, and the process proceeds to step S12.
[0052] In step S12, the guidance voice control unit 31 controls the output of the gesture guidance voice, and the gesture control unit 32 controls the gesture drive. As a result, as described above with reference to Figure 2, the tutorial is performed by the voice output unit 21 outputting the gesture guidance voice while the drive unit 22 performs the gesture drive.
[0053] In step S13, the guidance voice control unit 31 controls the output of the start guidance voice and causes the voice output unit 21 to output the start guidance voice. For example, the voice output unit 21 explains how to move the face by outputting the start guidance voice, "Got it? Move your face as slowly as possible like this," and then declares that it will start taking pictures of the face by outputting the start guidance voice, "Okay, now we'll start memorizing your face." Then, the voice output unit 21 outputs the start guidance voice, "Stare intently at my face," to get the user to face forward, and the process proceeds to step S14.
[0054] In step S14, streaming imaging by the imaging unit 23 is started, and facial images are sequentially supplied from the imaging unit 23 to the feature vector extraction unit 33.
[0055] In step S15, the guidance voice control unit 31 controls the output of the face-direction guidance voice and causes the voice output unit 21 to start outputting the face-direction guidance voice. As a result, the voice output unit 21 starts outputting face-direction guidance voices such as, "Turn to the right from there. One, two, three, four, five!", "Look at my face again, and this time turn to the left. One, two, three, four, five!", "Look at my face again, and this time turn upwards. One, two, three, four, five!", and "Look at my face again, and this time turn downwards. One, two, three, four, five!".
[0056] In step S16, the feature vector extraction unit 33 detects the user's face from the face image supplied by the imaging unit 23. Here, if the size of the detected user's face is small, the feature vector extraction unit 33 will not proceed to step S17, as it will be difficult to detect the part points in step S17. For example, in this case, the wheels 16 can be driven to move the mobile body 11 so that a face image of an appropriate size is captured.
[0057] In step S17, the feature vector extraction unit 33 detects feature points for each part, such as the eyes, nose, and mouth, from the user's face detected in step S16, and estimates the face orientation (yaw, pitch, roll) based on these feature points. If the estimated face orientation is outside the specified range, or if the face orientation could not be estimated, the unit 33 does not proceed to step S18.
[0058] In step S18, the feature vector extraction unit 33 adjusts the position using the part points estimated in step S17, then extracts the feature vector of the user's face and supplies it to the center vector calculation unit 34.
[0059] In step S19, the face registration processing unit 25 determines whether or not the face orientation guidance has been completed. For example, the face registration processing unit 25 determines that the face orientation guidance has been completed when the output of the face orientation guidance voice that was started in step S15 has finished, that is, when all the guidance to turn the user's face to the right, left, up, and down has been given.
[0060] If the face registration processing unit 25 determines in step S19 that face orientation guidance is not yet complete, the process returns to step S16, and the same process is repeated thereafter. On the other hand, if the face registration processing unit 25 determines in step S19 that face orientation guidance is complete, the process proceeds to step S20.
[0061] In step S20, streaming imaging by the imaging unit 23 is completed, and the supply of face images from the imaging unit 23 to the feature vector extraction unit 33 is stopped. At this time, the center vector calculation unit 34 has accumulated multiple feature vectors supplied from the feature vector extraction unit 33 during the period in which streaming imaging was performed.
[0062] In step S21, the center vector calculation unit 34 calculates the center of multiple feature vectors supplied from the feature vector extraction unit 33 to obtain a center vector, and registers the center vector in the face database of the storage unit 24.
[0063] In step S22, the guidance voice control unit 31 controls the output of the termination guidance voice and causes the voice output unit 21 to output the termination guidance voice. Here, the processing in step S22 can be performed in the time required to process steps S20 and S21. For example, the voice output unit 21 may output termination guidance voices such as "Okay, I'll try not to forget your face, just a moment" or "I'm remembering now, just a moment" during the time required to process steps S20 and S21. If the user is using the mobile device 11 for the first time, the user's name may be registered. When the processing in steps S20 and S21 is completed, the voice output unit 21 outputs termination guidance voice "I've remembered!" and the process is terminated.
[0064] As described above, the mobile device 11 can complete the registration of the user's face information through voice guidance and gesture drive without using a display. Specifically, in the tutorial, the mobile device 11 performs gesture drive to represent the user's facial movements during streaming shooting using gestures based on the speed and orientation (range of facial movement) of the mobile device 11's face unit 13, and outputs gesture guidance voice at a constant rhythm corresponding to the speed of the mobile device 11's face unit 13 in accordance with the gesture, so that the user can easily understand how to move their face when streaming shooting. Therefore, the user can move the orientation of their face without getting lost by following the face orientation guidance voice.
[0065] Furthermore, the mobile device 11 can avoid situations in the face registration process using streaming photography where, for example, the feature vector when the face is not facing forward deviates from the feature vector of the face facing forward, leading to the face being treated as a different person. The mobile device 11 can also avoid situations where the face registration process is not completed due to the inability to detect that the face is facing a specific direction due to the accuracy of the face angle estimation, or where multiple captures are required before the face registration process is completed. In addition, the mobile device 11 can provide a more robust and accurate face recognition function by calculating the central vector that is the center of multiple feature vectors extracted from face images at various angles and registering it in the face database. In other words, the mobile device 11 can improve the accuracy of face recognition by using face images at various angles acquired through streaming photography.
[0066] <Example of computer configuration> Figure 7 is a block diagram showing an example of the hardware configuration of a computer that executes the series of processes described above by a program.
[0067] In a computer, the CPU (Central Processing Unit) 101, ROM (Read Only Memory) 102, RAM (Random Access Memory) 103, and EEPROM (Electronically Erasable and Programmable Read Only Memory) 104 are interconnected by a bus 105. An input / output interface 106 is further connected to the bus 105, and this interface 106 is connected externally. The feature extraction process can be performed by the CPU, as well as by a GPU (Graphics Processing Unit), DSP (Digital Signal Processor), FPGA (Field Programmable Gate Array), etc.
[0068] In a computer configured as described above, the CPU 101 loads programs stored in ROM 102 and EEPROM 104, for example, into RAM 103 via bus 105 and executes them, thereby performing the series of processes described above. In addition, programs executed by the computer (CPU 101) can be pre-written to ROM 102, or installed or updated from an external source via input / output interface 106 into EEPROM 104.
[0069] In this specification, the processes performed by a computer according to a program do not necessarily have to be performed chronologically in the order described in the flowchart. That is, the processes performed by a computer according to a program include processes that are executed in parallel or individually (e.g., parallel processing or object-based processing).
[0070] Furthermore, the program may be processed by a single computer (processor), or it may be processed in a distributed manner by multiple computers. Moreover, the program may be transferred to a remote computer for execution.
[0071] Furthermore, in this specification, a system means a collection of multiple components (devices, modules (parts), etc.), regardless of whether all components are located in the same enclosure or not. Therefore, multiple devices housed in separate enclosures and connected via a network, and a single device in which multiple modules are housed in one enclosure, are both considered systems.
[0072] Furthermore, for example, the configuration described as a single device (or processing unit) may be divided and configured as multiple devices (or processing units). Conversely, the configurations described above as multiple devices (or processing units) may be combined and configured as a single device (or processing unit). It is also possible to add configurations other than those described above to the configuration of each device (or each processing unit). Moreover, if the overall system configuration and operation are substantially the same, a part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit).
[0073] Furthermore, for example, this technology can be configured as cloud computing, where a single function is shared and processed collaboratively by multiple devices via a network.
[0074] Furthermore, for example, the program described above can be executed on any device. In that case, the device should have the necessary functions (such as functional blocks) and be able to obtain the necessary information.
[0075] Furthermore, each step described in the flowchart above can be executed by a single device or shared among multiple devices. Additionally, if a single step includes multiple processes, these processes can be executed by a single device or shared among multiple devices. In other words, multiple processes within a single step can be executed as multiple steps. Conversely, processes described as multiple steps can be combined and executed as a single step.
[0076] Furthermore, the program executed by the computer may be executed in a chronological order according to the sequence of steps described herein, or it may be executed in parallel or individually at necessary times, such as when a call is made. In other words, as long as no inconsistencies arise, the processing of each step may be executed in an order different from the sequence described above. Moreover, the processing of the steps of this program may be executed in parallel with the processing of other programs, or it may be executed in combination with the processing of other programs.
[0077] Furthermore, the technologies described in this specification can be implemented independently, as long as they do not create a contradiction. Of course, any multiple technologies can also be implemented in combination. For example, some or all of the technologies described in one embodiment can be combined with some or all of the technologies described in another embodiment. In addition, some or all of the above-mentioned technologies can be implemented in combination with other technologies not mentioned above.
[0078] <Examples of configuration combinations> Furthermore, this technology can also be configured as follows. (1) During the tutorial for the face registration process, in which the user's face is registered in advance, a gesture control unit controls the gesture drive that expresses the user's facial movements as gestures during the subsequent streaming recording process, and Along with the gesture drive, a guidance voice control unit controls the output of a gesture guidance voice that matches the gesture. A mobile device equipped with [the following features]. (2) The gesture control unit causes the user's facial movements during streaming to be represented by gestures of the moving body's face. The mobile body described in (1) above. (3) The aforementioned voice guidance control unit outputs a constant rhythm corresponding to the speed as the gesture guidance voice. The mobile body described in (2) above. (4) The guidance voice control unit controls the output of face direction guidance voice that guides the user's facial movements during the streaming recording that takes place after the tutorial. A mobile body as described in any of (1) to (3) above. (5) A feature vector extraction unit extracts multiple feature vectors from user face images at various angles, which are acquired when the user changes the direction of their face according to the guidance of the face direction guidance voice while the aforementioned streaming recording is being performed. A center vector calculation unit calculates a center vector that is the center of multiple feature vectors and registers it in a face database. The mobile body described in (4) above, further comprising the above. (6) Threshold setting unit for setting thresholds used in face recognition processing that evaluates similarity with the center vector registered in the face database. To further enhance The mobile body described in (5) above. (7) The threshold setting unit sets one of the following as the threshold: a first threshold value determined during the design phase, a second threshold value corresponding to the distance to the feature vector furthest from the center vector, or a third threshold value which is the maximum value of the first threshold and the second threshold. The mobile body described in (6) above. (8) During the tutorial for the face registration process, in which the user's face is registered in advance, a gesture control unit controls the gesture drive that expresses the user's facial movements as gestures during the subsequent streaming recording process, and Along with the gesture drive, a guidance voice control unit controls the output of a gesture guidance voice that matches the gesture. A control device equipped with the following features. (9) The control device During the tutorial for the face registration process, which involves pre-registering the user's face, the subsequent process controls the gesture drive that represents the user's facial movements during streaming recording. Along with the gesture drive, the output of gesture guidance voice corresponding to the gesture is controlled. A control method including
[0079] It should be noted that this embodiment is not limited to the embodiment described above, and various modifications are possible without departing from the spirit of this disclosure. Furthermore, the effects described herein are merely illustrative and not limiting, and other effects may also exist. [Explanation of symbols]
[0080] 11 Mobile unit, 12 Main body, 13 Face unit, 14 Camera, 15 Eye unit, 16 Tires, 21 Audio output unit, 22 Drive unit, 23 Imaging unit, 24 Memory unit, 25 Face registration processing unit, 26 Threshold setting unit, 31 Guidance voice control unit, 32 Gesture control unit, 33 Feature vector extraction unit, 34 Center vector calculation unit
Claims
1. During the tutorial for the face registration process, in which the user's face is registered in advance, a gesture control unit controls the gesture drive that expresses the user's facial movements as gestures during the subsequent streaming recording process, and Along with the gesture drive, a guidance voice control unit controls the output of a gesture guidance voice that matches the gesture. A mobile device equipped with [the following features].
2. The gesture control unit causes the user's facial movements during streaming to be represented by gestures of the moving body's face. The mobile body according to claim 1.
3. The aforementioned voice guidance control unit outputs a constant rhythm corresponding to the speed as the gesture guidance voice. The mobile body according to claim 2.
4. The guidance voice control unit controls the output of face direction guidance voice that guides the user's facial movements during the streaming recording that takes place after the tutorial. The mobile body according to claim 1.
5. A feature vector extraction unit extracts multiple feature vectors from user face images at various angles, which are acquired when the user changes the direction of their face according to the guidance of the face direction guidance voice while the aforementioned streaming recording is being performed. A center vector calculation unit calculates a center vector that is the center of multiple feature vectors and registers it in a face database. The mobile body according to claim 4, further comprising
6. Threshold setting unit for setting thresholds used in face recognition processing that evaluates similarity with the center vector registered in the face database. The mobile body according to claim 5, further comprising:
7. The threshold setting unit sets one of the following as the threshold: a first threshold value determined during the design phase, a second threshold value corresponding to the distance to the feature vector at the furthest point from the center vector, or a third threshold value which is the maximum value of the first threshold and the second threshold. The mobile body according to claim 6.
8. During the tutorial for the face registration process, in which the user's face is registered in advance, a gesture control unit controls the gesture drive that expresses the user's facial movements as gestures during the subsequent streaming recording process, and Along with the gesture drive, a guidance voice control unit controls the output of a gesture guidance voice that matches the gesture. A control device equipped with the following features.
9. The control device During the tutorial for the face registration process, which involves pre-registering the user's face, the subsequent process controls the gesture drive that represents the user's facial movements during streaming recording. Along with the gesture drive, the output of gesture guidance voice corresponding to the gesture is controlled. A control method including
Citation Information
Patent Citations
Registering device and method for person recognizer
JP2000259834A
Individual recognition device and passage control device
JP2003141541A
Face identification device, face identification method, recording medium and robot device
JP2004302644A
Information processor
JP2015090662A
Response device, response system, response method, and recording medium
WO2017217314A1