Noise reduction in robot-human communication

A gesture library paired with noise profiles addresses the issue of mechanical noise interference in robot speech recognition, enhancing the robot's communication capabilities by improving signal-to-noise ratio and speech recognition performance.

JP7893607B2Inactive Publication Date: 2026-07-22MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
MICROSOFT TECHNOLOGY LICENSING LLC
Filing Date
2020-03-17
Publication Date
2026-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The presence of mechanical noise generated by a robot during gestural interactions reduces the effectiveness of speech recognition systems, impairing the robot's ability to understand and respond to user utterances.

Method used

A gesture library is utilized to pair predetermined robot gestures with noise profiles, allowing for the application of these profiles to cancel out mechanical noise from incoming audio, thereby improving the signal-to-noise ratio and enhancing speech recognition performance.

Benefits of technology

The implementation of a gesture library with noise profiles effectively reduces mechanical noise, improving the robot's ability to recognize and respond to user speech during gestural interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007893607000001
    Figure 0007893607000001
  • Figure 0007893607000002
    Figure 0007893607000002
  • Figure 0007893607000003
    Figure 0007893607000003
Patent Text Reader

Abstract

Noise reduction in robotic systems involves the use of a gesture library that pairs noise profiles with gestures that can be performed by the robot. A gesture to be performed by the robot is obtained, and the robot performs the gesture. The performance of the gesture by the robot generates noise, and when a user speaks to the robot while the robot is performing the gesture, the incoming audio includes both user speech and robot noise. The noise profile associated with the gesture is obtained from the gesture library and applied to remove the robot noise from the incoming audio.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] background

[0001] This disclosure relates in general to robot-human communication, and more particularly to noise reduction in robot-human communication. [Background technology]

[0002]

[0002] A robot is generally an electromechanical device guided by a computer or electronic programming. Robots can be used in a variety of applications, and are often envisioned in relation to their use in industrial applications. Recently, the use of robots in the field of human-robot interaction has been increasing, and the quality of human-robot interaction may be influenced by several factors, such as the robot's ability to recognize speech uttered by the user, as well as the robot's ability to interpret the speech and respond to it in an appropriate manner.

[0003]

[0003] In order to provide a more natural communication environment for human-robot interaction, it may be desirable for the robot to provide gestures along with the spoken utterances to realize a more natural communication process. Adding gestures to the robot's capabilities presents further challenges that may affect the robot system's ability to recognize and appropriately interpret the utterances spoken by the user.

[0004] [0004 The description of the embodiments is in relation to these and other general considerations. Also, while certain issues are discussed relatively, the embodiments are not limited to solving the specific issues identified in the background. [Overview of the project]

[0005] overview

[0005] The following is a simplified overview to provide a basic understanding of some aspects described herein. This overview is not a comprehensive overview of the claimed subject matter. It is not intended to identify any major or important elements of the claimed subject matter, nor to define its scope. Its sole purpose is to present some concepts in a simple form as a prelude to the more detailed descriptions presented below.

[0006]

[0006] According to one aspect of the present disclosure, a method for noise reduction in a robot system includes acquiring a gesture performed by the robot, receiving an incoming audio including a voice from a user and noise generated by the robot's performance of the gesture, acquiring a noise profile associated with the gesture from a gesture library, and applying the noise profile to remove robot noise from the incoming audio.

[0007]

[0007] In an embodiment, the gesture library includes a plurality of predetermined gestures that the robot may be expected to perform. Each of the predetermined gestures is paired with a noise profile to remove robot noise if incoming audio, including user voice, is received while the robot is performing the gesture.

[0008]

[0008] In another embodiment, a device for noise reduction within a robot system includes a processor and a memory coupled to the processor and storing instructions for execution by the process. When the instructions are executed by the process, the device causes the device to acquire a gesture to be performed by the robot, to receive an incoming audio including voice from the user and robot noise generated by the robot's execution of the gesture, to acquire a noise profile associated with the gesture from a gesture library, and to apply the noise profile to remove the robot noise from the incoming audio.

[0009]

[0009] In another embodiment, the computer-readable medium includes a computer-executable instruction which, when executed by a computer, causes the computer to perform a method of noise reduction in a robotic system in which a robot performs a gesture. The method includes receiving notification that a robot is performing a gesture; receiving an incoming voice, the incoming voice including a user utterance mixed with mechanical robot noise generated by the robot performing the gesture; obtaining a noise profile associated with a gesture from a gesture library which includes a plurality of predetermined gestures paired with noise profiles; and applying the noise profile to the incoming voice in order to remove the mechanical robot noise generated by the robot performing the gesture.

[0010] Brief Description of the Drawings (Non-Limited Embodiments of the Present Disclosure)

[0010] Various embodiments of the present disclosure will be described with reference to the following drawings. [Brief explanation of the drawing]

[0011] [Figure 1]

[0011] This is a schematic diagram showing a robotic system in which an exemplary implementation of the subject matter described herein may be implemented. [Figure 2]

[0012] A flowchart illustrating a method for noise reduction within a robot system according to an embodiment of this disclosure is shown. [Figure 3A]

[0013] A schematic diagram of a symbolic representation of one type of gesture according to an embodiment of the present disclosure is shown. [Figure 3B]

[0014] Exemplary symbols of parts of a robot's body are shown. [Figure 4]

[0015] This is a schematic diagram illustrating the generation of gesture-noise profile pairs for a gesture library according to embodiments of the present disclosure. [Figure 5]

[0016] This is a schematic diagram illustrating exemplary components of a computing device suitable for implementing noise reduction within a robotic system according to embodiments of the present disclosure. [Modes for carrying out the invention]

[0012] Detailed description of the embodiment

[0017] The following detailed description refers to accompanying drawings, which form part of this description and in which specific embodiments are shown for illustrative purposes. These embodiments are described in sufficient detail to enable those skilled in the art to carry out the art. Other embodiments may be available, and structural, logical, and electrical modifications may be made without departing from the spirit and scope of this disclosure. Accordingly, the following detailed description should not be construed as limiting, and only the accompanying claims and their equivalents define the scope. Identical reference numerals in the figures refer to identical components, which should be apparent in connection with use.

[0013]

[0018] Figure 1 shows a schematic diagram of a robot system 400 according to an embodiment of the present disclosure. As shown in Figure 1, the robot system 400 generally includes a robot 100, a device 10, and a server 300. The device 10 can control the robot to perform various gestures, for example, by sending commands to the robot 100 to control motors / actuators 110 that orient parts of the robot's body in a particular manner.

[0014]

[0019] In addition to performing gestures, the robot 100 can be a chat robot that is associated with the gesture in the voice uttered by the robot 100 in order to provide a more natural, extensive, and effective communication environment between, for example, the user 50 and the robot 100. In the communication between the robot and the human, the user 50 can interact with the robot 100 by supplying a message through speech / voice or other expressions. The incoming voice including the voice of the user 50 is received by the robot through the microphone 30, and the microphone 30 may or may not be embedded in the robot 200. The server 300 can include a voice recognition module 310 for processing the voice of the user. The server 300 can be in the form of a cloud-based computer, for example, chat intelligence having voice recognition capabilities in the case of a chat robot that interacts with the speech / voice of the user 50.

[0015]

[0020] [[ID=�]]The device 10 has the ability to control the robot 100 to execute a predetermined number of different gestures. The device 10 receives the processed information from the server 300 and interprets the processed information to control the robot 100 to execute a specific gesture. The device 10 includes a motion control module 14 that receives the processed information from the server 300 and generates a command to control the robot 100 to move one or more parts of the robot body in a specific direction to execute a gesture. The command can be, for example, a series of joint angles that indicate how to orient the movable body part of the robot 100.

[0016]

[0021] Robot 100 receives commands from device 10 and executes commands for performing gestures by operating a plurality of motors / actuators 110. The motors / actuators 110 direct the body parts of the robot according to the instructions of device 100. In addition, robot 100 can also have movement control capabilities beyond those related to performing gestures from a predetermined number of different gestures. For example, robot 100 can have the ability to balance when unexpected movements occur during the execution of a gesture. This additional movement control capability can be realized through the movement control module 14 of device 10 or can be movement control that is executed independently of device 10 by the internal movement system of robot 100 itself. Motors / actuators 110 (e.g., servo and / or stepper mode), transformers, chassis flexures and contacts, hydraulic devices, chamber echoes inside the robot, gears, etc. involved in providing gesture, operation, or movement functions generate mechanical noise 20.

[0017]

[0022] In the process of interaction between a robot and a human, it is natural for user 50 to want to communicate with robot 100 while robot 100 is performing a gesture. For example, during a normal chat process between user 50 and robot 100, user 50 may first express or ask a question to robot 100 and then expect to receive a response. Robot 100's response may include a gesture performed by robot 100. While robot 100 is performing this gesture, user 50 may want to speak to robot 100 (for example, to ask a follow-up question). To achieve a more natural and smooth communication process between robot 100 and user 50, the robot system 400 should be able to respond to utterances from user 50 issued while robot 100 is performing a gesture. However, if user 50 speaks while robot 100 is performing a gesture, the incoming sound picked up by microphone 30, including the user's utterance / speech signal, will be mixed with the mechanical noise 20 generated by robot 100's gesture performance. The presence of mechanical noise 20 in the incoming speech reduces the performance of the speech recognition service provided by the speech recognition module 310 of the server 300, thereby reducing the robot system 400's ability to understand and respond to the user's utterances.

[0018]

[0023] According to various embodiments of the present disclosure, the speech recognition performance of the robot system 100 is improved by reducing the relative level of internal mechanical noise to the utterance / speech signal detected by the robot's microphone 30, resulting in an increased signal-to-noise ratio with respect to the speech content. Various embodiments of the present disclosure provide a gesture library in which gestures expected to be performed by the robot 100 are paired with noise profiles. Knowledge of the gestures that the robot 100 may be instructed to perform allows for the retrieval of the corresponding noise profile, which can then be used to cancel out mechanical noise components mixed with the user's utterances 50.

[0019]

[0024] Exemplary implementations of the subject matter described herein will be described with reference to robot system 400. However, robot system 400 is described for illustrative purposes only and does not imply any limitation on the scope of the subject matter described herein. For example, concepts and principles are equally applicable to standalone machines.

[0020]

[0025] Figure 2 shows a flowchart of a method for noise reduction within a robot system 400 according to an embodiment of the present disclosure. The method can be performed, for example, on the device 10 shown in Figure 1. The illustrated device 10 may be a client device or a cloud-based device, or it may be part of the server 300 or robot 100 shown in Figure 1. The method may include further actions not shown, and / or illustrated steps may be omitted. The scope of the subject matter described herein is not limited to this embodiment.

[0021]

[0026] The method will be described with reference to Figures 1 and 2. In 201, a gesture performed by the robot 100 is acquired. The gesture may be one of a plurality of predetermined gestures that the robot 100 is capable of performing. The gesture may be represented, for example, by a symbolic representation of the gesture. The symbolic representation may be a digital signal format in which the orientation of a part of the robot 100's body is represented by symbols that can be interpreted by the device 100 to generate instructions for the robot 100 to orient that part of its body in a particular way.

[0022]

[0027] In embodiments of this disclosure, each gesture performed by the robot 100 may be represented using a gesture language in which symbols are used to represent the orientation of a part of the robot body. The gesture language is preferably machine-independent (or hardware-independent) in that the language can be interpreted and compiled independently of the type of robot 100 performing the gesture. A particular gesture performed by the robot 100 can be determined by the server 300, for example, through a gesture language module 320. The server 300 can then provide the device 10 with a symbolic representation of the gesture performed by the robot 100.

[0023]

[0028] The server 300 can utilize, for example, a library that pairs a set of predetermined gestures that may be formed by the robot 100 with symbolic representations of those gestures. Thus, the gesture language module 320 can determine the appropriate gesture to be performed by the robot 100 and send the symbolic representation of the gesture to the device 10. However, this disclosure is not limited to this method. For example, the device 10 itself could perform this function instead.

[0024]

[0029] One exemplary gesture language that can be used by the robot system 400 is Labanotation. Figures 3A and 3B show typical Labanotation for performing gestures. Labanotation is a notation used to record human movements, where symbols define the orientation of various parts of the body. Specifically, in this specification, Labanotation defines the orientation of at least one body part of robot 100 with respect to multiple time slots 301. Labanotation is machine-independent and can therefore be implemented by multiple different types of hardware (or robots). In addition, as a concise symbolic representation, Labanotation is easy to transmit between the robot and a cloud computer (e.g., server 300) through a limited communication channel. Also, Labanotation generally requires less memory than other types of representation.

[0025]

[0030] In some embodiments, it is possible to determine the orientation of at least one body part of the robot 100 in multiple time slots 301 through continuously captured and / or recorded gestures, and then obtain a symbol corresponding to the orientation. The symbol can then be saved as part of a lab annotation, associated with the corresponding time slot 301.

[0026]

[0031] In some embodiments, at least one body part includes multiple body parts, and the labannotation includes a first dimension corresponding to multiple time slots 301 and a second dimension corresponding to multiple body parts. Figure 3A shows such a labannotation representing a particular gesture. In this labannotation, each column corresponds to one specific body part, such as the left hand, left arm, torso, right arm, right hand, or head. Each row corresponds to a time slot having a given duration. Furthermore, the symbols represent the direction to which the body part is oriented at that moment. It should be noted that the sample labannotation in Figure 3A is shown for illustrative purposes only and does not imply any limitation on the scope of the subject matter described herein. In other words, more complex labannotations involving more body parts are possible.

[0027]

[0032] In 202, the device 10 can cause the robot to perform a gesture. Once the device 10 obtains a gesture to be performed, the device 10 instructs the robot 100 to orient its body parts to perform the specific gesture. For example, the device 10 may receive a symbolic representation of the gesture from the server 300, determine joint angles based on the symbolic representation, and instruct the robot 100 to control its motors 110 to specific joint angles. The various motors 110 of the robot 100 move specific parts of the robot 100 so that the robot 100 can perform the gesture.

[0028]

[0033] In the process of executing commands provided by the device, the motors 110 and mechanical parts of the robot 100 involved in providing gestures generate mechanical noise 20 that can be picked up by the microphone 30. This noise 20 becomes problematic when the microphone 30 receives incoming speech, which includes user speech (e.g., user utterances) with which the robot system 400 should interact. In such cases, the incoming speech from the user may be speech in which speech recognition is performed to determine how the robot system 400 should respond to the user's utterances. If the incoming speech is received while the mechanical noise 20 is being generated by the motors 110 and movable parts of the robot 100, the mechanical noise 20 will mix with the incoming speech. The presence of mechanical noise 20 in the incoming speech may reduce the performance of the speech recognition service provided by the speech recognition module 310 of the server 300, which is used by the robot system 400 to understand the meaning of the user's utterances and respond to them.

[0029]

[0034] In order to reduce the mechanical robot noise 20 picked up by the microphone 30 and mixed with the incoming voice, a noise profile INM is used in 204 to remove the mechanical robot noise 20 from the incoming voice. N The noise profile is obtained. The noise profile is ultimately used to cancel out the mechanical noise 20 associated with the gesture execution by the robot 100 when the noise profile is applied in S205. By canceling out the mechanical noise 20 from the incoming speech, the signal-to-noise ratio of the incoming signal is improved, which in turn improves the ability of the speech recognition module 310 to recognize, convert, and respond effectively to the user's utterances contained within the incoming signal.

[0030]

[0035] In embodiments of this disclosure, the noise profile is a gesture (LA1, LA2, ..., LA N ) is the noise profile (INM1, INM2, ..., INM N) is obtained from the gesture library 12 that pairs with it. The gesture library 12 contains a finite number of gestures (i.e., a plurality of predetermined gestures (LA1, LA2,..., LA N )) that are expected to be executed by the robot 100 to interact with the user 50. For each of these gestures, the gesture library 12 contains a noise profile LNM N for canceling out the mechanical noise 20 generated by the execution of the gesture LA N by the robot 100. In 205 where the noise profile INM N is applied to remove the mechanical noise 20 from the incoming voice picked up by the microphone 30, the noise signal associated with the execution of the gesture can be mixed with the incoming voice in a phase-shifted state, for example, to obtain a clearer voice signal that better represents the utterance spoken by the user 50 while the robot 100 is executing the gesture.

[0031]

[0036] In an embodiment where the gesture is represented by a symbolic expression such as the lavanotation LA N , the gesture library 12 can index each of the noise profiles (INM1, INM2,..., INM N ) with respect to the lavanotation representing the gesture that generated the noise 20 for which the noise profile INM N was generated. In such a case, when the server 300 provides a specific lavanotation LA N , the device can retrieve the appropriate noise profile INM N from the gesture library 12 based on the lavanotation LA N received from the server 300.

[0032]

[0037] In an exemplary embodiment, the noise profiles (INM1, INM2,..., INM NEach of these can be an inverse noise model that can be mixed with the audio signal picked up by the microphone 30 to perform noise cancellation. Inverse Noise Model INM N The robot uses the inverse noise model INM. N This is the reciprocal of the noise signal generated by the robot 100 when performing the associated gesture. Therefore, the inverse noise model can be mixed with the audio signal picked up by the microphone 30 during the performance of the gesture by the robot 100 by adding the inverse noise model to the audio signal.

[0033]

[0038] The sequence of steps described above is not limited to the specific order in which they are described, but can be performed in any suitable order or simultaneously. For example, obtaining a noise profile from the gesture library 12 may occur simultaneously with, or before, having the robot perform a gesture and receiving incoming audio from the user.

[0034]

[0039] Furthermore, the actions described herein may be computer-executable instructions that can be implemented by one or more processors and / or stored on one or more computer-readable media. Computer-executable instructions may include routines, subroutines, programs, execution threads and / or similar. Furthermore, the results of the actions of the method may be stored in computer-readable media, displayed on a display device, and / or similar actions may be performed. Computer-readable media may be any suitable computer-readable storage device such as memory, hard drives, CDs, DVDs, flash drives or similar. The term “computer-readable media” as used herein is not intended to include propagating signals.

[0035]

[0040] Figure 4 is a schematic diagram illustrating the generation of gesture-noise profile pairs contained within the gesture library 12 according to an embodiment of the present disclosure. In the embodiment shown in Figure 4, labanonotation is used as a machine-independent gesture language that symbolically represents gestures performed by the robot 100. However, the gesture language used to generate gesture-noise pairs is not limited to labanonotation.

[0036]

[0041] Figure 4 shows one embodiment in which mechanical noise 20 generated by the execution of a gesture by the robot 100 is recorded in order to generate a noise profile. In this case, the gesture library 12 includes a noise profile based on the pre-recorded noise signal.

[0037]

[0042] Upon acquiring lavanonotation, the robot controller module 220 of the device 100 controls the robot 100 to perform the gesture, for example, by sending commands to the robot 100 to orient one or more parts of the robot body in a specific way. When performing the gesture, the robot 100 generates mechanical noise 20, which is caused by, for example, the robot's motors 110 (e.g., servos and / or stepper motors), transformers, chassis flex and contacts, hydraulic systems, internal chamber echoes of the robot, gears, etc. The mechanical noise 20 is recorded, and a noise profile is generated based on the pre-recorded noise 20. The generated noise profile is then paired with the gesture (in this example, the lavanonotation representation of the gesture). The noise profiles stored in the library 12 may be, for example, a digital recording of the pre-recorded noise signal, the reciprocal of the pre-recorded noise, or another noise profile generated based on the pre-recorded noise.

[0038]

[0043] As described above, the noise profile is used to cancel out mechanical robot noise picked up by the microphone 30 and mixed with the incoming user voice. In one exemplary embodiment, the noise profile may include a pre-recorded signal mixed with the incoming user voice in phase, or the reciprocal of a pre-recorded noise signal added to the incoming user voice. The process of generating gesture-noise pairs is iterated for each gesture / labannotation contained in the gesture library 12. Since there is a finite number of gestures / labannotations that the robot 100 is expected to perform, the robot system 400 can provide a robot 100 that can perform a predetermined number of gestures using the gesture language, independently of the robot 100, while also providing the ability to perform noise cancellation of noise specific to the robot's particular hardware. Thus, the system 400 can ultimately provide gesture services to multiple different types of robots, independently of the hardware and software implemented by the robot 100, while also having the ability to perform noise cancellation of noise specific to the motors, mechanical components, etc., of each of the different types of robots.

[0039]

[0044] In one embodiment, the same microphone 30 used to capture incoming user voice is used to generate the gesture library 12. The use of the same microphone 30 is beneficial in that the hardware component used to pre-record noise signals is the same as the one that picks up noise signals when the robot system is operating, thereby further ensuring that the noise signals in the gesture library 12 are an accurate representation of the noise picked up by the microphone 30 when the robot 200 performs the associated gestures.

[0040]

[0045] In one embodiment, when generating the gesture library 12, pre-recorded robot noise audio signals are synchronized with the corresponding gestures so that noise cancellation occurs at the appropriate time. As shown in Figure 3A, when performing labannotation, time progresses from bottom to top in 301, and specific combinations of various symbols indicating different orientations of multiple body parts in a given time slot 301 are performed so that the robot 100 can perform the corresponding motions sequentially with respect to time. When constructing the gesture library 12 according to the method described above, since the robot 100 itself is used to generate the pre-recorded robot noise, it can be assumed that the pre-recorded noise signals of the noise profile are synchronized with the specific movements of the robot in the time slots in which they occur. To synchronize the incoming audio with the selected noise profile, the device 10 can timestamp the point in time when motor control begins, and then synchronize this timestamp with the noise model associated with the gesture being performed by the robot so that the start time when the microphone 30 receives the incoming audio is synchronized with the start time when the noise profile is applied.

[0041]

[0046] The embodiment described in relation to Figure 4 shows an example in which the gesture library 12 is generated based on a pre-recorded noise signal picked up by the microphone 30, but the generation of this gesture library 12 is not limited to this. For example, in one embodiment, the noise profile associated with a particular gesture can be obtained from an alternative source without requiring the robot 100 itself to record mechanical noise 20. In addition, in another embodiment of the present invention, the noise profiles of gestures in the gesture library 12 can be generated using a physical model that represents the noise generated by the robot when the robot performs the gesture. The physical model can predict the propagation of sound that occurs when the robot performs the gesture. Unlike the use of data collected from a sound sensor or similar (as in the case of generating a noise profile using noise acquired from the microphone 30), the physical model includes predictions such as motor waveforms, chassis sound emulations, and sound reflection patterns.

[0042]

[0047] Embodiments of the present disclosure may also include an overlay model that can integrate unexpected sounds with an existing gesture library 12. The overlay model can be computed, for example, by using an augmented noise record that may be generated in real time or according to a physical model. Unexpected sounds from received motor movements may, for example, be the result of the robot 100 recovering on its own or resisting an external unexpected force that occurs while the robot 100 is performing a gesture. The overlay model for unexpected sounds can be applied together with a pre-recorded noise model for a particular gesture to facilitate additional noise cancellation if additional unexpected movements occur when the robot is performing a gesture.

[0043]

[0048] In addition, in embodiments of this disclosure, an environmental noise physical model can also be generated to represent the environmental noise that may be picked up by the microphone 30 while the user 50 interacts with the robot 100. The physical model for environmental noise predicts the noise generated by the environment in which the robot interacts. The physical model for environmental noise can be added to the gesture library 12 and can be further mixed with the incoming speech in a phase-shifted manner to reduce the environmental noise picked up by the microphone 30. The gesture library 12 may include multiple environment models, each modeling a different environment in which the robot may exist.

[0044]

[0049] In 205, once the noise model is applied to the incoming speech signal, the noise-canceled speech signal can be transmitted to the speech recognition module 310. The speech recognition module 310 converts the noise-canceled speech signal into oral interaction elements used and provided to the device 10. For example, the speech recognition module 310 can perform analysis based on the content of the noise-canceled speech signal and prepare utterances to be spoken by the robot 100 as responses or answers to user utterances contained within the noise-canceled speech signal. Furthermore, the gesture language module 320 can also determine gestures to be performed by the robot 100 based on the output of the speech recognition module 310. The gestures may accompany utterances spoken by the robot 100, or, instead, the speech recognition module 310 may decide that utterances are not performed by the robot, and the gesture language module 320 may determine gestures to be performed by the robot 100 without any utterances from the robot.

[0045]

[0050] When determining an appropriate gesture for robot 100 in conjunction with the robot's speech, server 300 may, for example, extract a concept from the speech uttered by the robot and retrieve a gesture from a library corresponding to the extracted concept. A concept may be a representative example extracted from a cluster of words, such as "hello," "good," "thank you," or "hungry." However, this disclosure is not limited to any particular method of selecting a gesture to be performed by robot 100.

[0046]

[0051] Once a gesture is acquired, the robot system 400 can again perform the method shown in Figure 4 to remove robot noise from any incoming sounds received by the microphone 30 while the gesture is being performed.

[0047]

[0052] Figure 5 is a block diagram of apparatus 10 suitable for implementing one or more implementations of the subject matter described herein. For example, apparatus 10 can function as described above with reference to Figure 1. However, since various implementations can be implemented in a variety of general-purpose or special-purpose computing environments, apparatus 10 is not intended to imply any limitation on the scope of use or function of the subject matter described herein.

[0048]

[0053] As shown in the figure, the device 10 includes at least one processor 120 and memory 140. The processor 120 executes computer executable instructions and may be a real or virtual processor. In a multiprocessing system, multiple processors execute computer executable instructions to increase processing power. The memory 130 may be volatile memory (e.g., registers, cache, RAM), non-volatile memory (e.g., ROM, EEPROM, flash memory), or any combination thereof. The memory 130 and its associated computer-readable media provide storage for data, data structures, computer executable instructions, etc. of the device 10.

[0049]

[0054] In accordance with the implementation of the subject matter described herein, memory 130 is coupled to processor 120 and stores instructions executed by processor 120. When executed by processor 120, these instructions cause the device to: acquire a gesture performed by the robot; receive incoming audio, which includes voice from the user along with robot noise generated by the robot's execution of the gesture; obtain a noise profile associated with the gesture from the gesture library to remove robot noise generated by the robot's execution from the incoming audio; and apply the noise profile to remove robot noise from the incoming audio.

[0050]

[0055] In the example shown in Figure 5, the device 10 further includes one or more communication connections 140. Interconnection mechanisms such as buses, controllers, or networks interconnect the components of the device 10. Typically, operating system software provides an operating environment for other software running within the device 10 and coordinates the activities of the components of the device 10.

[0051]

[0056] The communication connection 140 enables communication over a communication medium with another computing entity. In addition, the functionality of the components of the device 10 can be implemented by a single or multiple computing machines that can communicate over the communication connection. Thus, the device 10 can operate within a networked environment (e.g., a robotic system environment 400) by using logical connections to one or more other servers, network PCs, or other common network nodes. Without limitation, the communication medium includes, for example, wired or wireless network connection techniques.

[0052]

[0057] Implementations of the subject matter described herein include a computer-readable medium containing computer-executable instructions. When executed by a computer, these instructions cause the computer to perform a method of noise reduction in a robotic system in which a robot performs a gesture, the method comprising: receiving notification that a robot is performing a gesture; receiving an incoming voice, the incoming voice including a user utterance mixed with mechanical robot noise generated by the robot's performance of the gesture; obtaining a noise profile associated with a gesture from a gesture library containing a plurality of predetermined gestures paired with noise profiles; and applying the noise profile to the incoming voice in order to remove the mechanical robot noise generated by the robot's performance of the gesture.

[0053]

[0058] Computer storage media include, without limitation, RAM, ROM, EPROM, EEPROM, flash memory or other semiconductor memory technologies, CD-ROM, DVD or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other media used to store desired information and accessible by a computer.

[0054]

[0059] The terminology used herein is intended solely to describe specific embodiments and is not intended to be limiting. The singular forms “a,” “an,” and “it” as used herein are also intended to include the plural form unless the context clearly indicates otherwise. The terms “include,” “contain,” “have,” “include,” “contain,” and / or “have,” as used herein, specify the presence of the described feature, complete, step, action, element, and / or component, but do not exclude the presence or addition of one or more other features, complete, step, action, element, component, and / or group thereof.

[0055]

[0060] The corresponding structures, materials, actions, or equivalents of all means or step-plus-function elements in the attached claims are intended to include any structures, materials, or actions that perform a function in combination with other claimed elements, if any, as specifically claimed. This description is presented for illustrative and explanatory purposes only and is not intended to be exhaustive or limiting in the form disclosed. Many variations and modifications will become apparent to those skilled in the art without departing from the scope and spirit of the art. The embodiments are selected and described to best illustrate the principles of the art and its practical applications and to enable those skilled in the art to understand the art for various embodiments with various modifications to suit specific intended uses.

[0056]

[0061] Although specific embodiments are described, those skilled in the art will understand that other embodiments equivalent to those described exist. Therefore, the art is not limited to the specific illustrated embodiments, but is limited only by the scope of the appended claims.

[0057]

[0062] According to one aspect of the present disclosure, a method for noise reduction within a robot system includes receiving a gesture performed by the robot, receiving an incoming voice which includes a voice from a user and robot noise generated by the robot performing a gesture, obtaining a noise profile associated with the gesture from a gesture library for removing robot noise generated by the robot performing a gesture from the incoming voice, and applying the noise profile to remove the robot noise from the incoming voice.

[0058]

[0063] In this embodiment, the noise profile may be an inverse noise model, and applying the noise profile to remove robot noise from the incoming speech includes applying the inverse noise model to the incoming speech.

[0059]

[0064] In this embodiment, the noise profile may include a pre-recorded noise signal from a robot performing a gesture, and applying the noise profile to remove robot noise may include mixing the pre-recorded noise signal with the incoming speech in a phase-shifted manner.

[0060]

[0065] In this embodiment, the gesture library may include a plurality of predetermined gestures performed by the robot, and each of the predetermined gestures is paired with a noise profile for removing robot noise.

[0061]

[0066] In this embodiment, the method may further include generating a gesture library, which may include having a robot perform a predetermined gesture and recording robot noise generated by the robot's performance of the gesture for each of the predetermined gestures in order to generate a noise profile.

[0062]

[0067] In this embodiment, incoming sound may be received by a robot microphone, and recording robot noise generated by the robot performing the gesture in order to generate a noise profile for each of the predetermined gestures may include recording robot noise generated by the robot performing the gesture using the robot microphone for each of the predetermined gestures.

[0063]

[0068] In this embodiment, the gesture library may include multiple symbolic representations of gestures performed by the robot, and each of these symbolic representations is paired with a noise profile for removing robot noise.

[0064]

[0069] In this embodiment, obtaining a symbolic representation of a gesture performed by the robot may involve obtaining a labannotation that defines the orientation of at least one body part of the robot over multiple time slots.

[0065]

[0070] In this embodiment, at least one body part comprises multiple body parts, and causing the robot to perform a gesture may involve performing labnotation to trigger the multiple body parts to perform the gesture according to their individual orientations in multiple time slots.

[0066]

[0071] According to another aspect of the present invention, a device for noise reduction in a robot system includes a processor and a memory coupled to the processor and storing instructions for execution by the processor, the instructions, when executed by the processor, cause the device to: acquire a gesture to be performed by the robot; receive an incoming sound, the incoming sound includes a voice from the user and robot noise generated by the robot's execution of the gesture; obtain a noise profile associated with the gesture from a gesture library for removing robot noise generated by the robot's execution of the gesture from the incoming sound; and apply the noise profile to remove robot noise from the incoming sound.

[0067]

[0072] In this embodiment, the noise profile may be an inverse noise model, and applying the noise profile to remove robot noise from the incoming speech may include applying the inverse noise model to the incoming speech.

[0068]

[0073] In this embodiment, the noise profile includes a pre-recorded noise signal from a robot performing a gesture, and applying the noise profile to remove robot noise may include mixing the pre-recorded noise signal with the incoming speech in a phase-shifted manner.

[0069]

[0074] In this embodiment, the gesture library may include a plurality of predetermined gestures performed by the robot, and each of the predetermined gestures is paired with a noise profile for removing robot noise.

[0070]

[0075] In this embodiment, when an instruction is executed by the processor, it may cause the device to further generate a gesture library, which includes causing the robot to perform a predetermined gesture and recording robot noise generated by the robot's performance of the gesture for each of the predetermined gestures in order to generate a noise profile.

[0071]

[0076] In this embodiment, incoming sound may be received by a robot microphone, and recording robot noise generated by the robot performing the gesture in order to generate a noise profile for each of the predetermined gestures may include recording robot noise generated by the robot performing the gesture using the robot microphone for each of the predetermined gestures.

[0072]

[0077] In this embodiment, obtaining a gesture performed by the robot may include obtaining a symbolic representation of the gesture performed by the robot, and the instruction, when executed by the processor, may further cause the device to cause the robot to perform the gesture, which includes controlling the orientation of at least one body part of the robot according to the symbolic representation.

[0073]

[0078] In this embodiment, the gesture library may include multiple symbolic representations of gestures performed by the robot, and each of these symbolic representations is paired with a noise profile for removing robot noise.

[0074]

[0079] In this embodiment, obtaining a symbolic representation of a gesture performed by the robot may involve obtaining a labannotation that defines the orientation of at least one body part of the robot over multiple time slots.

[0075]

[0080] In this embodiment, at least one body part may include multiple body parts, and causing the robot to perform a gesture may include performing labnotation to trigger the multiple body parts to perform the gesture according to their individual orientations in multiple time slots.

[0076]

[0081] According to another aspect of the present invention, the computer-readable medium includes a computer-executable instruction which, when executed by a computer, causes the computer to perform a method of noise reduction in a robotic system in which a robot performs a gesture, the method comprising receiving notification that a robot is performing a gesture; receiving an incoming voice, the incoming voice including a user's utterance mixed with mechanical robot noise generated by the robot's performance of the gesture; obtaining a noise profile associated with a gesture from a gesture library which includes a plurality of predetermined gestures paired with noise profiles; and applying the noise profile to the incoming voice in order to remove the mechanical robot noise generated by the robot's performance of the gesture.

[0077]

[0082] In this embodiment, the multiple noise profiles in the gesture library may include pre-recorded noise signals from a robot performing a predefined gesture, and applying a noise profile to remove robot noise may include mixing the pre-recorded signals of the noise profile associated with the gesture with the incoming speech in a phase-shifted manner.

[0078]

[0083] In this embodiment, a plurality of predetermined gestures may be represented in the gesture library by a plurality of symbolic representations of the predetermined gesture, each symbolic representation defining the orientation of at least one body part of the robot when performing the gesture.

Claims

1. A method for noise reduction within a robot system, wherein the apparatus is To obtain the gestures performed by the robot, Receiving incoming audio, wherein the incoming audio includes the voice from the user and robot noise generated by the robot performing the gesture, Obtaining a noise profile associated with a gesture from a gesture library for removing robot noise generated by the robot's execution of the gesture from the incoming audio, wherein the noise profile is generated using a physical model representing the noise generated by the robot when it performs the gesture. Applying the noise profile to remove the robot noise from the incoming sound Includes, A method further comprising integrating any additional unexpected sounds with the gesture library if such unexpected movements occur when the robot performs the gesture.

2. The method according to claim 1, wherein the physical model includes predicting the propagation of sound that occurs when the robot performs a gesture.

3. The method according to claim 2, wherein the prediction of sound propagation includes any one of the following: a motor waveform, chassis sound emulation, and / or a prediction of a sound reflection pattern.

4. The method according to claim 1, wherein the unexpected sound includes the sound of the robot righting itself, or the sound of the robot resisting an external unexpected force that occurs while performing the gesture.

5. The method according to claim 1, wherein the gesture library includes a plurality of predetermined gestures performed by the robot, each of the predetermined gestures being paired with a noise profile for removing robot noise.

6. Obtaining a gesture performed by a robot includes obtaining a symbolic representation of the gesture performed by the robot. The method according to any one of claims 1 to 5, further comprising causing the robot to perform the gesture, wherein causing the robot to perform the gesture includes controlling the orientation of at least one body part of the robot in accordance with the symbolic representation.

7. The method according to claim 6, wherein the gesture library includes a plurality of symbolic representations of gestures performed by the robot, each of which is paired with a noise profile for removing robot noise.

8. The method according to claim 7, wherein obtaining a symbolic representation of a gesture performed by a robot includes obtaining a lavanotation that defines the orientation of at least one body part of the robot over a plurality of time slots.

9. A device for noise reduction within a robot system, Processor and A memory coupled to the processor, which stores instructions for execution by the processor. The instruction includes, and when executed by the processor, the device, To obtain the gestures performed by the robot, Receiving incoming audio, wherein the incoming audio includes the voice from the user and robot noise generated by the robot performing the gesture, Obtaining a noise profile associated with a gesture from a gesture library for removing robot noise generated by the robot's execution of the gesture from the incoming audio, wherein the noise profile is generated using a physical model representing the noise generated by the robot when it performs the gesture. Applying the noise profile to remove the robot noise from the incoming sound Have them do it, A device that, if an additional unexpected movement occurs when the robot performs the gesture, further causes the device to integrate such unexpected sound with the gesture library.

10. A computer-readable storage medium containing computer-executable instructions, wherein, when executed by a computer, the instructions cause the computer to perform a method for noise reduction in a robotic system in which a robot performs gestures, and the method is Receiving notification that the robot is performing the gesture, Receiving an incoming sound, wherein the incoming sound includes the user's utterance mixed with mechanical robot noise generated by the robot's execution of the gesture, Obtaining a noise profile associated with a gesture from a gesture library containing a plurality of predetermined gestures paired with a noise profile, wherein the noise profile is generated using a physical model representing the noise generated by the robot when the robot performs the gesture. In order to remove the mechanical robot noise generated by the execution of the gesture by the robot, the noise profile is applied to the incoming speech. Includes, A computer-readable storage medium further includes integrating any additional unexpected sounds with the gesture library if such unexpected movements occur when the robot performs the gesture.