Robot-implemented speech recognition method

The humanoid robot uses omnidirectional capture and image processing to enhance speech recognition, addressing the challenge of multiple users and noise, ensuring reliable interaction without positional constraints.

WO2026087447A1PCT designated stage Publication Date: 2026-04-30ENCHANTED TOOLS
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ENCHANTED TOOLS
Filing Date
2025-10-20
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Existing humanoid robots face limitations in speech recognition when interacting with multiple users in noisy environments, requiring users to adopt specific positions and struggling to distinguish target sound signals from noise, especially when users speak at low volumes.

Method used

A humanoid robot employs omnidirectional microphone arrays and cameras to capture sound and image data, determining arrival direction angles and implementing a sound signal separation algorithm with automatic speech recognition, allowing reliable interaction without requiring a preferred orientation or position.

Benefits of technology

The method effectively identifies the intended user and separates their sound waves from others, enabling reliable speech recognition even in noisy environments, allowing users to interact freely without specific positioning and reducing accidental detections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025080237_30042026_PF_FP_ABST
    Figure EP2025080237_30042026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method for achieving speech recognition by a robot (2), comprising: ● capturing sound signals emitted by users (1) ● recording an image of the users (1) ● transmitting the sound signals and the image to a processing unit of the robot (2), which processing unit is configured to determine a first angle made between each user (1) and the robot (2); ● detecting in the image a sign made by one of the users (1), the processing unit determining a second angle made between the user (1) having made the sign and the robot (2); ● comparing the second angle with each first angle, the processing unit determining the first angle closest to the second angle; ● applying a speech recognition algorithm to the sound signal of the user (1) associated with the first angle closest to the second angle.
Need to check novelty before this filing date? Find Prior Art

Description

Speech recognition process implemented by a robot

[0001] The present invention relates to the technical field of robots, in particular humanoid robots.

[0002] More specifically, the invention relates to the technical field of social interactions between one or more users and such robots, in particular social interactions requiring processing of sound and visual signals by such robots. STATE OF THE ART

[0003] A robot, particularly a humanoid one, configured to interact socially with one or more users is known from the prior art. For a user wishing to interact with such a robot to, for example, give it instructions in natural language, a robot is known to include a sound signal capture device, for example, an array of omnidirectional microphones, commonly referred to as a "microphone array," the robot being configured to implement an automatic recognition algorithm.

[0004] In an environment exposing the robot to sound signals other than those of the user wishing to interact with the robot, a robot is known to be configured to implement an automatic speech recognition algorithm that allows for the dissociation of a target sound signal associated with the user from other sound signals, which are then considered noise as understood in the field of signal processing. To achieve this, it is known to determine, relative to a reference axis of the capture axis, an arrival angle for each sound wave source; the algorithm thus preferentially processes sound waves associated with a predetermined arrival angle.This method forces the user to adopt a preferred position in relation to the robot's sensing organ, and has limitations that are particularly noticeable when several users wish to interact simultaneously with the robot, especially when a user speaks at a sound level lower than the ambient sound level.

[0005] The invention aims to resolve all or part of the drawbacks of the prior art, by proposing a robot configured to implement a more reliable speech recognition process, in particular to identify a user wishing to interact with the robot in a noisy environment. PRESENTATION OF THE INVENTION

[0006] More specifically, the invention relates to a speech recognition method implemented by a robot, in particular a humanoid robot, configured to interact with users, the method comprising: omnidirectional multichannel capture of sound signals emitted by users via at least one set of microphones of the robot; recording of an image of at least some of the users by at least one camera of the robot; transmission of the sound signals and the image to a processing unit of the robot, the processing unit being configured to determine a first angle of arrival direction formed between a direction in which each user who emitted a sound signal is located relative to the robot and a reference axis of the robot;the detection on the image of a first activation sign made by one of the users, the processing unit determining from said image a second arrival direction angle formed between the direction in which the user who made the first activation sign is located relative to the robot and the robot's reference axis; the comparison of the second arrival direction angle with each first arrival direction angle, the processing unit determining the first arrival direction angle closest to the second arrival direction angle; the implementation of a sound signal separation algorithm, followed by an automatic speech recognition algorithm on the user's sound signal associated with the first arrival direction angle closest to the second arrival direction angle.

[0007] The determination of the first directions of arrival angles from sound signals is commonly referred to as "Direction Of Arrival (DOA) estimation" by the person in the trade.

[0008] Thanks to this combination of features, this method makes it possible to determine with a very high degree of reliability which user actually wishes to interact with the robot, and to more effectively separate the sound waves emitted by that user from other sound waves, particularly in an environment with multiple users and ambient noise. This method is especially advantageous for allowing a user with a low volume, for example, a user who does not wish to draw attention to themselves, to interact with the robot. Furthermore, this method also allows any user within the robot's camera's field of view to interact with the robot without requiring a preferred orientation or position relative to the robot.

[0009] Advantageously, the processing unit delimits a first frame on the image for each user, this first frame including the user's face. In such a configuration, the determination of the second angle of arrival direction is made between the user's delimited frame and the robot's reference axis. Thus, the determination of this angle is more precise, and the robot can also store the user's face in memory for a subsequent iteration of the process.

[0010] Advantageously, the robot's camera image is centered on the first frame of the user for which the automatic speech recognition algorithm has been implemented. In such a configuration, the robot can track the face of the user currently interacting with it by centering the camera's field of view on their face, particularly when the user moves relative to the robot and / or when the robot moves relative to the user.

[0011] Advantageously, the processing unit delineates a second frame on the image for each user, this second frame including one of the user's hands, and the detection of the first activation signal is preferentially performed within this second frame. In such a configuration, the first activation signal can only be performed by one of the user's hands, so any other gesture, including a gesture similar to the first activation signal with the other hand, is not recognized in the image by the robot's processing unit. Thus, the risk of accidental, unintentional detection is significantly reduced.

[0012] Advantageously, the first sign of activation is formed by a gesture of one of the user's hands.

[0013] Advantageously, the method includes, simultaneously with the detection of the first activation signal from one of the users in the image, the detection of a second activation signal among the captured audio signals. In such a configuration, the reliability of the method is further improved, as the user must simultaneously perform two activation signals. Thus, the risk of accidental, unintentional detection is significantly reduced. Moreover, such a second signal can be used to orient the camera so that the user who performed the second signal is at least partially within the camera's field of view for the potential detection of the first activation signal.

[0014] According to another aspect of the invention, it relates to a robot, in particular a humanoid robot, configured to interact with users, the robot comprising: at least one microphone configured to omnidirectionally capture sound signals emitted by users; at least one camera configured to record an image of at least some of the users; a processing unit; the robot being configured to implement a speech recognition process as described above. PRESENTATION OF THE FIGURES

[0015] The invention will be better understood upon reading the following description, given solely by way of example, and referring to the accompanying drawings given by way of non-limiting examples, in which identical references are given to similar objects and on which:

[0016] This is a schematic representation for a robot interacting with users according to a first aspect of the invention;

[0017] This is a logic diagram of an automatic speech recognition process implemented by the robot according to another aspect of the invention;

[0018] This is a schematic representation of the principle of carrying out the steps of the process of the;

[0019] This is a schematic representation in principle of the implementation of steps in the process according to another embodiment;

[0020] It should be noted that the figures set out the invention in detail to enable implementation of the invention; although not limiting, said figures serve in particular to better define the invention where appropriate. DETAILED DESCRIPTION OF THE INVENTION

[0021] The invention relates in particular to a robot 2, configured to interact with users 1 as illustrated in the figure. The robot 2 is preferably a humanoid robot, for example a ballbot, that is to say, a robot moving on a single spherical wheel. The term "interact" refers to the ability of the robot 2 to generate one or more social logistics actions following a request from at least one user 1, for example, through instructions in natural language addressed directly to the robot 2 and processed by a processing unit (not shown) of the robot 2. The robot 2 includes in particular at least one set of microphones 4. The set of microphones 4 preferably consists of at least two omnidirectional microphones configured to capture sound signals emitted by at least one user 1, in particular through their voice. The robot 2 also includes at least one camera 6.Camera 6 is configured to record in real time an image of at least some of the users 1. Preferably, camera 6 is an RGB-D type camera configured to determine a depth map on an image.

[0022] According to another aspect of the invention, robot 2 is configured to implement a speech recognition process as illustrated in the.

[0023] The process includes a first step consisting of the omnidirectional multichannel capture E1 of sound signals emitted by users 1, via the set of microphones 4 of the robot 2.

[0024] The method also includes a second step consisting of recording an image E2 of at least some of the users 1 by at least the camera 6 of the robot 2. The first step E1 and the second step E2 are preferably carried out simultaneously and continuously, as long as the robot 2 detects a non-zero sound level. The recorded image contains depth information for each object or user 1 seen.

[0025] The method also includes a third step consisting of transmitting the sound signals and the image E3 to a processing unit of the robot 2. The processing unit is configured to determine a first arrival direction angle 7 formed between a direction in which each user 1 who emitted a sound signal is located relative to the robot and a reference axis 8 of the robot 2. This step is illustrated, among other things, in Figure 1 and Figure 2. The term "reference axis" refers to an axis linked to the frame of reference of the robot 2 from which angles can be determined. The reference axis 8 is preferably an axis belonging to a plane of symmetry of the robot, for example, an axis included in the sagittal plane of the robot 2. Naturally, it is possible that the set of microphones 4 of the robot 2 may be offset from the reference axis 2 by a predetermined distance, which is taken into account in determining the first arrival direction angles 7.

[0026] The method also includes a fourth step consisting of detecting on the image E4 a first activation sign 5 made by one of the users 1, the processing unit determining from said image a second arrival direction angle 9 formed between the direction in which the user 1 who made the first activation sign 5 is located relative to the robot 2 and the reference axis 8 of the robot 2. As explained previously, it is possible that the camera 6 of the robot 2 is offset from the reference axis 2 by a predetermined distance, taken into account in the determination of the second arrival direction angle 7. Preferably, the camera 6 of the robot 2 is mounted so that the reference axis 8 is a bisector of the viewing angle of the camera 6, thus minimizing this offset.

[0027] The process also includes a fifth step consisting of comparing E5 the second direction of arrival angle 9 with each first direction of arrival angle, the processing unit determining the first direction of arrival angle 7 closest to the second direction of arrival angle 9.

[0028] The process further includes a sixth step consisting of implementing a sound signal separation algorithm, followed by a speech recognition algorithm E6 on the user's sound signal 1 associated with the first arrival direction angle 7 closest to the second arrival direction angle 9. Preferably, the speech recognition algorithm includes an MVDR beamforming technique, more commonly referred to by those skilled in the art as "Minimum Variance Distortion Response beamforming technique". Alternatively, the speech recognition algorithm may consist of a beamforming technique based on a neural network trained to perform multichannel source separation at a specific angle.

[0029] The method details the principle of implementation. Here, two initial arrival direction angles 7 are determined from the auditory signals, one for each user 1. The left-hand user 1 makes a first activation signal, formed here by a gesture with one of the user's hands. Thus, a second arrival direction angle 9 is calculated from the direction of the left-hand user 1, that is, the one who made the first activation signal. Subsequently, the second arrival direction angle is compared in turn to each of the first arrival direction angles. It is then possible to identify the auditory signals on which to implement the speech separation techniques described.

[0030] Thanks to this combination of features, this method makes it possible to determine with a very high degree of reliability which user 1 actually wishes to interact with robot 2, also known as the user of interest, and to more effectively separate the sound waves emitted by said user 1 from other sound waves, particularly in an environment with multiple users 1 and ambient noise. This method is especially advantageous for allowing a user 1 with a low noise level, for example, a user of interest 1 who does not wish to attract the attention of other users 1, to interact with robot 2. Furthermore, this method also allows any user 1 within the field of view of robot 2's camera 6 to interact with robot 2 without requiring a preferred orientation or position relative to robot 2.

[0031] Finally, such a process makes it possible to prevent the risk of performing speech recognition on an individual other than the user of interest 1 who would be fortuitously aligned with said user, in front of or behind the latter in relation to robot 2.

[0032] The diagram details the principle of implementation of the process according to another embodiment. Here, the processing unit of robot 2 delimits a first frame 10 for each user 1. The first frame 10 is delimited around the face of each user 1. In such a configuration, the determination of the second arrival direction angle 9 is made between the first frame 10 delimited for user 1 and the reference axis 8 of robot 2. Thus, the determination of said angle is more precise, the robot 2 also being able to store the user's face in memory to accelerate a subsequent iteration of the process.

[0033] Similarly, the processing unit of robot 2 delimits a second frame 12 on the image for each user 1, the second frame 12 including one of the hands of user 1 and the detection of the first activation sign 5 being preferentially carried out within the second frame 12. In such a configuration, the first activation sign 5 can only be performed by one of the hands of user 1, so that any other gesture, including a gesture similar to the first activation sign 5 of the other hand, is not recognized on the image by the processing unit of robot 2. Thus, the risk of accidental involuntary detection is significantly reduced.

[0034] Preferably in this embodiment, the camera 6 of robot 2's field of view, i.e., the image, is centered on the first frame 10 of user 1 for which the automatic speech recognition algorithm has been implemented. In such a configuration, robot 2 can track the face of user 1 currently interacting with it by centering the camera 6's field of view on their face, particularly when user 1 is moving relative to robot 2 and / or when robot 2 is moving relative to user 1. Thus, the determination of the first direction-of-arrival angles 7 and the second direction-of-arrival angle 9 is dynamic, so that the comparison remains relevant in the event of relative movement of robot 2 and users 1.

[0035] It should also be noted that the invention is not limited to the embodiments described above. Indeed, it will be apparent to a person skilled in the art that various modifications can be made to the embodiments described above, in light of the information just provided.

[0036] For example, regardless of the embodiment, the method may include, simultaneously with the detection in the image of an activation signal from one of the users 1, the detection among the captured sound signals of a second activation signal (not shown). In such a configuration, the reliability of the method is further improved, as user 1 must simultaneously perform two activation signals. Thus, the risk of accidental, unintentional detection is significantly reduced. Moreover, such a second signal can be used to orient the camera 6 so that the user 1 who performed the second signal is at least partially within the field of view of the camera 6 for possible detection of the first activation signal as described above.

[0037] In the detailed presentation of the invention given above, the terms used shall not be interpreted as limiting the invention to the embodiments set forth in this description, but shall be interpreted as including all equivalents which can be foreseen by a person skilled in the art by applying their general knowledge to the implementation of the teaching which has just been disclosed to them.

Claims

A speech recognition method implemented by a robot (2), in particular a humanoid robot, configured to interact with users (1), the method comprising: omnidirectional multichannel capture (E1) of sound signals emitted by the users (1) via at least one set of microphones (4) of the robot (2); recording an image (E2) of at least some of the users (1) by at least one camera (6) of the robot (2); transmission of the sound signals and the image (E3) to a processing unit of the robot (2), the processing unit being configured to determine a first direction of arrival angle (7) formed between a direction in which each user (1) who emitted a sound signal is located relative to the robot (2) and a reference axis (8) of the robot (2);the detection on the image (E4) of a first activation sign (5) made by one of the users (1), the processing unit determining from said image a second arrival direction angle (7) formed between the direction in which the user (1) who made the first activation sign (5) is located with respect to the robot (2) and the reference axis (8) of the robot (2); the comparison (E5) of the second arrival direction angle (7) with each first arrival direction angle (7), the processing unit determining the first arrival direction angle (7) closest to the second arrival direction angle (7); the implementation of a sound signal separation algorithm, followed by an automatic speech recognition algorithm (E6) on the sound signal of the user (1) associated with the first arrival direction angle (7) closest to the second arrival direction angle (7). Method according to claim 1 wherein the processing unit delimits on the image a first frame (10) for each user (1), the first frame including the face (3) of the user (1). Method according to claim 2, wherein the image from the camera (6) of the robot (2) is centered on the first frame (10) of the user (1) for which the automatic speech recognition algorithm has been implemented. Method according to claim 3, wherein the processing unit delimits on the image a second frame (10) for each user (1), the second frame (12) including one of the hands of the user (1) and the detection of the first sign of activation (5) being carried out preferentially within the second frame (12). A method according to any one of the preceding claims, wherein the first activation sign (5) is formed by a gesture of one of the user's hands (1). A method according to any one of the preceding claims, comprising, concurrently with the detection on the image of the first sign of activation (5) of one of the users (1), the detection among the captured sound signals of a second sign of activation. Robot (2), in particular humanoid, configured to interact with users (1), the robot (2) comprising: at least one set of microphones (4) configured to omnidirectionally capture sound signals emitted by the users (1); at least one camera (6) configured to record an image of at least a part of the users (1); a processing unit; the robot (2) being configured to implement a speech recognition method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Control method and device of intelligent equipment, equipment and medium

    CN110187766A

  • Robot and robot control method

    JP6845121B2

  • Control device, control method, and program

    JP7452363B2

  • Input Determination Method

    US20180181197A1

  • Conversation facilitating method and electronic device using the same

    US20230041272A1