Noise reduction in robot-human communication
By establishing a gesture library in the robot system and using noise profiles to eliminate robot noise, the problem of noise interference when the robot executes gestures is solved, the speech recognition performance is improved, and a more natural human-computer interaction is achieved.
Patent Information
- Application Number
- CN202080034072.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-08
- Filing Date
- 2020-03-17
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2040-03-17
AI Technical Summary
The mechanical noise generated by the robot when performing gestures interferes with voice communication with the user, degrades the performance of the speech recognition system, and affects the robot system's ability to understand and respond to the user's utterances.
By establishing a posture library in the robot system, recording and storing the correspondence between each posture and noise profile, the noise profile is used to eliminate robot noise from the input audio and improve the signal-to-noise ratio of speech recognition.
It effectively eliminates the noise generated when the robot performs gestures, improves the performance of the speech recognition system, enhances the robot's ability to understand and respond to user speech, and achieves more natural human-computer interaction.
Smart Images

Figure CN113826160B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to robot-to-human communication, and particularly to noise reduction in robot-to-human communication. Background Art
[0002] A robot is an electromechanical machine typically guided by a computer or electronic program. Robots can be used in a wide range of applications and are generally considered to be used in industrial applications. Recently, the use of robots in the field of human-robot interaction has increased, and many human-robot interactions can be affected by various factors, such as the robot's ability to recognize the words spoken by the user, and the robot's ability to interpret the words and respond with appropriate behavior.
[0003] In order to provide a more natural environment for human-robot interaction, it may be desirable to provide the robot with gestures as well as verbal expressions to enable a more natural communication process. Adding gestures to the robot's capabilities introduces additional challenges that can affect the robot system's ability to recognize the utterances spoken by the user and interpret them appropriately.
[0004] The embodiments have been described with respect to these and other general considerations.Although relatively specific problems have also been discussed, the embodiments should not be limited to solving the specific problems identified in the background. Summary of the Invention
[0005] The following is presented in a simplified summary to provide a basic understanding of some aspects described herein. This summary is not a comprehensive overview of the claimed subject matter. It is intended to identify key points or important elements of the claimed subject matter rather than to delineate its scope. Its sole purpose is to present some concepts in a simplified form as a prelude to the more detailed description that follows.
[0006] According to one aspect of the present disclosure, a method for noise reduction in a robotic system includes: obtaining a gesture to be performed by a robot; receiving input audio, the input audio including audio from a user and robot noise caused by the robot performing the gesture; retrieving a noise profile associated with the gesture from a gesture library; and applying the noise profile to remove the robot noise from the input audio.
[0007] In some embodiments, the gesture library includes a plurality of predetermined gestures that the robot may be expected to perform. Each predetermined gesture is paired with a noise profile for removing robot noise when input audio including user audio is received while the robot performs the gesture.
[0008] According to another aspect, an apparatus for noise reduction in a robotic system includes a processor coupled to the processor and configured to process instructions for execution by the processor. The instructions, when executed by the processor, cause the apparatus to: obtain a gesture to be performed by a robot; receive input audio comprising audio from a user and robot noise caused by the robot performing the gesture; retrieve a noise profile associated with the gesture from a gesture library; and apply the noise profile to remove the robot noise from the input audio.
[0009] According to another aspect, a computer-readable medium includes computer-executable instructions that, when executed by a computer, cause the computer to perform a method for noise reduction in a computer system, wherein a robot performs a gesture. The method includes receiving an indication that the robot is performing a gesture; obtaining input audio, the input audio including user speech mixed with mechanical robotic noise caused by the robot performing the gesture; retrieving a noise profile associated with the gesture from a library of gestures including a plurality of predetermined gestures paired with noise profiles; and applying the noise profile to the audio input to remove the mechanical robotic noise caused by the robot performing the gesture. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Various embodiments according to the present invention will be described with reference to the accompanying drawings, in which:
[0011] Figure 1 An exemplary schematic diagram of a robotic system is shown in which the subject events described herein may be implemented.
[0012] Figure 2 A flowchart of a method for noise reduction in a robot system according to an embodiment of the present disclosure is shown.
[0013] Figure 3A A schematic diagram illustrating types of symbolic representations of gestures according to an embodiment of the present disclosure.
[0014] Figure 3B Example symbols showing the main body parts of a robot.
[0015] Figure 4 A schematic diagram indicating a method for creating gesture-noise profile pairs for a gesture library according to an embodiment of the present disclosure is shown.
[0016] Figure 5 A high-level diagram of exemplary components of a computer device suitable for implementing noise reduction in a robotic system according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0017] In the following detailed description, reference is made to the accompanying drawings, which are incorporated herein by way of example, and in which specific embodiments are shown. These embodiments are described in sufficient detail to enable those skilled in the art to practice these techniques. Other embodiments may be utilized, and structural, logical, and electrical changes may be made without departing from the spirit and scope of the present invention. Therefore, the following detailed description should not be considered limiting, and the scope is limited only by the appended claims and their equivalents. It will be apparent from the context of use that like reference numerals in the drawings represent like parts.
[0018] Figure 1 FIG. 4 shows a schematic diagram of a robot system 400 according to an embodiment of the present disclosure. Figure 1 As shown, the robot system 400 generally includes a robot 100, an apparatus 10, and a server 300. The apparatus 10 can control the robot to perform various postures by, for example, sending commands to the robot 100 to control motors or actuators 110 so that the main body of the robot 100 is oriented in a specific manner.
[0019] In addition to performing gestures, the robot 100 may be, for example, a chatbot, whose gestures accompany the speech spoken by the robot 100 to provide a more natural, comprehensive, and effective communication environment between the user 50 and the robot 100. During robot-human communication, the user 50 may interact with the robot 100 by delivering information through speech / utterances or other expressions. Input audio, including the speech of the user 50, is received by the robot through a microphone 30, which may or may not be embedded in the robot 200. The server 300 may include a speech recognition module 310 for processing the user's speech. The server 300 may be in the form of a cloud-based computer, for example, having speech recognition capabilities, and is used for conversational intelligence in the verbal / speech interaction between the chatbot and the user 50.
[0020] The device 10 is capable of controlling the robot 100 to execute a predetermined number of different poses. The device 10 receives processed information from the server 300 and interprets the processed information to control the robot 100 to execute a specific pose. The device 10 includes a motion control module 14, which receives the processed information from the server 300 and generates commands to control the robot 100 to move one or more robot body parts in a specific direction to execute the pose. For example, the command may be a series of joint angles that instruct the robot 100 how to orient the moving body part.
[0021] The robot 100 receives commands from the device 10 and executes them to execute gestures operated by a plurality of motors or actuators 110. The motors or actuators 110 orient the robot's main body in a manner directed by the device 10. Furthermore, the robot 100 may have motion control capabilities beyond those involved in executing gestures from a predetermined number of different gestures. For example, the robot 100 may have balancing capabilities in case unexpected motion occurs during gesture execution. Additional motion control capabilities may be implemented by the motion control module 14 of the device 10, or may be motion control performed independently of the device 10 by the robot 100's internal motion system. The motors or actuators 110 (e.g., servo and / or stepper motors), transformers, flexible chassis, connections, hydraulics, cavity echoes within the robot, gears, etc., may be involved in generating noise 20 (e.g., mechanical noise or mechanical robot noise) during gesture control, manipulation, or movement functions.
[0022] During a robot-human interaction, it's natural for the user 50 to desire to communicate with the robot 100 while the robot is performing a gesture. For example, during a normal conversation between the user 50 and the robot 500, the user 50 may begin to make an expression or ask a question to the robot 100, then expect a response. The robot 100's response may include a gesture performed by the robot 100. While the robot 100 is performing the gesture, the user 50 may desire to speak to the robot 100 (e.g., to ask a follow-up question). To achieve a more natural and smoother communication process between the robot 100 and the user 50, the robot system 400 should be able to respond to the speech uttered by the user 50 while the robot 100 is performing the gesture. However, if the user speaks while the robot 100 is performing the gesture, the input audio received by the microphone 30, including the user's speech / language signal, is mixed with the mechanical noise caused by the robot 100 performing the gesture. The presence of noise 20 in the input audio reduces the performance of the speech recognition service provided by the speech recognition module 310 of the server 300, thereby reducing the robot system 400's ability to understand and respond to the user's speech.
[0023] According to various embodiments of the present disclosure, the speech recognition system of the robot system 400 improves speech / language signals perceived by the robot's microprocessor (e.g., processor 120) by reducing internal mechanical noise, resulting in an increase in the signal-to-noise ratio of the audio content. Various embodiments of the present disclosure provide a gesture library in which gestures that the robot 100 is expected to perform are paired with noise profiles. Knowing the gestures that the robot 100 can be commanded to perform, the corresponding noise profile can be retrieved to remove the mechanical noise components mixed with the user's 50 voice.
[0024] Exemplary implementations of the subject matter described herein will be described with reference to a robotic system 400, however the robotic system 400 is described for illustrative purposes only and does not imply any limitations on the subject matter as described herein. For example, the ideas and principles are also applicable to stand-alone machines.
[0025] Figure 2 A method for noise reduction in a robot system 400 according to an embodiment of the present disclosure is shown. For example, the method may be used in accordance with Figure 1 The device 10 shown in the figure can be a client device or a cloud-based device, or it can be a Figure 1 The server 300 or a portion of the robot 100 is shown. The method may also include additional actions not shown and / or omit exemplary steps. The scope of the subject matter described herein is not limited in this respect.
[0026] The method will refer to Figure 1 as well as Figure 2 As described, at 201, a gesture to be performed by the robot 100 is obtained. The gesture may be one of predetermined gestures that the robot 100 is capable of performing. For example, the gesture may be represented by a symbolic representation of the gesture. The symbolic representation may be in a digital signal format, where the orientation of a main body portion of the robot 100 is represented by a symbol that can be interpreted by the device 10 to generate an instruction for the robot 100 to orient its main body portion in a specific manner.
[0027] In embodiments of the present disclosure, each gesture performed by the robot 100 may be represented using gesture speech, wherein symbols are used to represent the orientation of the robot's body parts. The gesture speech is preferably machine-independent (or hardware-independent), wherein the speech can be interpreted and compiled regardless of the type of gesture performed by the robot 100. The specific gesture to be performed by the robot 100 may be determined, for example, by the server 300 via the gesture speech module 320. The server 300 may then provide a symbolic representation of the gesture to be performed by the robot to the device 10.
[0028] For example, the server 300 may utilize a library that pairs a plurality of predetermined gestures that can be performed by the robot 100 with symbolic representations of the gestures. Thus, the gesture speech module 320 may determine the appropriate gesture to be performed by the robot 100 and send the symbolic representation of the gesture to the device 10. However, the present disclosure is not limited in this manner. For example, the device 10 itself may alternatively perform this function.
[0029] One exemplary gesture utterance that the robotic system 400 may use is labanotation. Figure 3A3B shows a typical Laban notation for performing a gesture. Laban notation is a symbolic system for recording human body movements, in which symbols define the orientation of various body parts. Specifically, Laban notation herein defines the orientation of at least one body part of the robot 100 relative to a plurality of time slots 301. Laban notation is machine-independent and can therefore be implemented by a variety of different types of hardware (or robots). Furthermore, as a concise symbolic representation, Laban notation is easily transferred between the robot and a cloud computer (e.g., server 300) via limited communication channels. Laban notation also generally requires less memory than other types of representations.
[0030] In some embodiments, the orientation of at least one body part of the robot 100 in a plurality of time slots 301 can be determined by continuously capturing / recording poses, and then symbols corresponding to the orientations can be obtained. Thereafter, the symbols associated with the corresponding time slots 301 as part of Laban notation can be saved.
[0031] In some embodiments, the at least one body portion comprises a plurality of body portions, and the Laban notation comprises a first dimension corresponding to the plurality of time slots 301 and a second dimension corresponding to the plurality of body portions. Figure 3A Such notation for a particular posture is shown. In this Laban notation, each column corresponds to a particular body part, such as left hand, left arm, support, right arm, right hand, head. Each row corresponds to a time slot of a given duration. In addition, the symbol indicates in which direction the body part is oriented at that time. Note that Figure 3A The sample Laban notation in FIG is shown for illustrative purposes only and does not imply any limitation on the scope of the subject matter described herein. In other words, more complex Laban notations involving more subject matter are also possible.
[0032] At 202, the device 10 may cause the robot 100 to perform a gesture. Once the device 10 has acquired the gesture to be performed, the device 10 instructs the robot 100 to orient its body portion to perform the specific gesture. For example, the device 10 may receive a symbolic representation of the gesture from the server 300, determine joint angles based on the symbolic representation, and instruct the robot 100 to control its motors 110 to the specific joint angles. The various motors or actuators 110 of the robot 100 move the specific portion of the robot 100, causing the robot 100 to perform the gesture.
[0033] By executing instructions provided by the device, the motors 110 and mechanical components of the robot 100 involved in providing the gesture are involved in providing the gesture that generates a noise 20 that can be received by the microphone 30. This noise 20 becomes problematic when the microphone 30 receives input audio (203) including user audio (e.g., user speech) with which the robotic system 400 should interact. In this case, the input audio from the user may be audio for performing speech recognition so that the robotic system 400 can determine how it should respond to the user's speech. When the input audio is received and the noise 20 is generated by the motors or brakes 110 and the moving body parts of the robot 100, the noise 20 is mixed with the input audio. The presence of mechanical noise in the input audio can reduce the performance of the speech recognition service. The presence of noise 20 in the input audio can reduce the performance of the speech recognition service provided by the speech recognition module 310 of the server 300, which the robotic system 400 uses to understand and respond to the meaning of the user's speech.
[0034] In order to reduce the noise 20 received by the microphone 30 and mixed with the input signal, in 204, a noise profile INM for removing the noise 20 from the input audio is generated. N The noise profile is ultimately used to eliminate the noise 20 associated with the execution of the gesture by the robot 100 when the noise profile is applied in S205. By eliminating the noise 20 from the input audio, the signal-to-noise ratio of the input audio is improved, which enhances the ability of the speech recognition module 310 to recognize, translate, and effectively respond to the user's utterances contained in the input audio.
[0035] In an embodiment of the present disclosure, the noise profile is retrieved from the gesture library 12, where the gestures (LA1, LA2, ..., LA N ) and noise profiles (INM1, INM2, ..., INM N ) pairing. The gesture library 12 includes a limited number of gestures (ie, a plurality of predetermined gestures (LA1, LA2, ..., LA N For each of these postures, the posture library 12 includes a noise profile INM for eliminating the noise 20 caused by the execution posture LAN of the robot 100. N . In Application 205 Noise Profile INM N The robot 100 will perform posture LA N When the noise 20 is induced, the noise signal associated with the execution of the gesture can be mixed out of phase with the input audio, for example, to obtain a clearer audio signal that better represents the words spoken by the user 50 when the robot 100 performs the gesture.
[0036] In gestures represented by symbols, such as Laban notation LA N In the embodiment shown, the gesture library 12 can store each noise profile (INM1, INM2, ..., INM N ) indexed to the Laban notation representation of the gesture causing the noise 20, the noise profile INM N is created for the noise 20. In this case, when the server 300 provides a specific Laban notation LA N When the device can receive the Laban notation LA from the server 300, N , retrieve the appropriate noise profile LNM from the posture library 12 N .
[0037] In an exemplary embodiment, each noise profile (INM1, INM2, ..., INM N ) may be an inverse noise model that may be mixed with the audio signal received by microphone 30 to perform noise cancellation. Inverse noise model LNM N When the robot performs with the inverse noise model LNM N The associated gesture is the inverse of the noise signal caused by the robot 100. Therefore, the inverse noise signal can be mixed with the audio signal received through the microphone 30 during the execution of the gesture by the robot 100 by adding the inverse noise signal to the audio signal.
[0038] The order of the above steps is not limited to the particular order in which they are described, and may be performed in any suitable order or simultaneously. For example, the retrieval of the noise profile from the gesture library 12 may occur simultaneously with, before, or after causing the robot to perform a gesture and receiving input audio from the user.
[0039] Furthermore, the actions described herein may be computer-executable instructions that can be implemented by one or more processors and / or stored on one or more computer-readable media. Computer-executable instructions may include routines, subroutines, programs, execution threads, and the like. Furthermore, the results of the actions of the methods may be stored on a computer-readable medium, displayed on a display device, and the like. The computer-readable medium may be any suitable computer-readable storage device, such as a memory, a hard drive, a CD, a DVD, a flash drive, and the like. As used herein, the term "computer-readable medium" is not intended to include propagated signals.
[0040] Figure 4 FIG. 1 is a schematic diagram for explaining the creation of a pair of gesture noise profiles to be included in the gesture library 12 according to an embodiment of the present disclosure. Figure 4 In the illustrated embodiment, Laban notation is used as an independent gesture utterance for symbolically representing a machine for a gesture performed by the robot. However, the gesture utterance used to create a gesture-noise pairing is not limited to Laban notation.
[0041] Figure 4 An embodiment is shown in which the machine noise 20 generated by the execution of gestures by the robot 100 is recorded to create a noise profile. In this case, the gesture library 12 includes a noise profile based on a pre-recorded noise signal.
[0042] Upon acquiring the Laban notation, the robot controller module 220 of the robot 200 controls the robot to execute a gesture, for example, by sending instructions to the robot 100 to orient one or more robot body parts in a specific manner. While executing the gesture, the robot 100 generates noise 20, for example, caused by the robot's motors 110 (e.g., servo and / or stepper motors), transformers, chassis flexure and contact, hydraulics, cavity echoes within the robot, gears, and the like. The noise 20 is recorded, and a noise profile is created based on the predetermined noise 20. The created noise profile is then paired with the gesture (in this example, the Laban notation represents the gesture). The noise profile stored in the library 12 can, for example, be a digital recording of a pre-recorded noise signal, the inverse of a pre-recorded noise signal, or other noise profiles created based on pre-recorded noise.
[0043] As described above, the noise profile will be used to eliminate mechanical robot noise that is picked up by microphone 30 and mixed with the input user audio. In an exemplary embodiment, the noise profile can include a pre-recorded noise signal that is mixed with the input user audio at different phases or the inverse of the pre-recorded noise signal that is added to the input user audio. The process of creating a gesture noise pair is repeated for each gesture / Laban notation contained in the gesture library 12. Because there are a limited number of gestures / Laban notations that the robot is expected to perform, gesture utterances that are independent of the robot 100 can be used for the robot system 400 to provide a predetermined number of gestures that can be performed, while also providing the ability to perform noise cancellation for noise specific to the specific hardware of the robot. Therefore, the robot system 400 can ultimately provide gesture services to multiple different types of robots that are independent of the hardware and software implemented by the robot 100, while having the ability to perform noise cancellation for noise of motors, mechanical components, etc. that are specific to each type of robot.
[0044] In an embodiment, the same microphone 30 used to capture input user audio is used to create the gesture library 12. Using the same microphone 30 can be beneficial because the hardware components used to pre-record the noise signal are the same hardware components used to receive the noise signal of the gesture library 12 during operation of the robotic system, thereby also ensuring that the noise signal of the gesture library 12 is an accurate representation of the noise that will be received by the microphone 30 when the robot 200 performs the associated gesture.
[0045] In an embodiment, when creating the gesture library 12, the pre-recorded robot noise audio signals are synchronized with the corresponding gestures so that the noise cancellation occurs at the appropriate time. Figure 3A As shown, when Laban notation is executed, time passes from bottom to top in 301, and specific combinations of various symbols indicating various orientations of multiple body parts will be executed at a given time slot 301, so that the robot 100 can continuously perform corresponding actions related to time. When constructing the gesture library 12 according to the method described above, the robot 100 itself is used to generate pre-recorded robot noise, and therefore, the pre-recorded noise signal of the noise profile can be assumed to be synchronized with the robot's specific actions at the time slot when the specific action occurs. In order to synchronize the input audio with the selected noise profile, the device 10 can set a timestamp at the point where motor control begins, and then synchronize this timestamp with the noise model associated with the gesture performed by the robot, so that the start point of the microphone 30 receiving the input audio and the start point of the noise profile applied at the start point are synchronized.
[0046] Although the reference Figure 4 The described embodiment shows an example of creating a gesture library 12 based on a pre-recorded noise signal received by the microphone 30, but the creation of the gesture library 12 is not limited to this. For example, in an embodiment, the noise profile associated with a specific gesture can be obtained from an alternative source without requiring the robot 100 to record the noise 20 itself. In addition, in another embodiment of the disclosed invention, the noise profile of the gesture in the gesture library 12 can be created using a physical model that represents the noise created by the robot when the robot performs the gesture. The physical model can predict the sound propagation that occurs when the robot performs the gesture. Unlike data collected from acoustic sensors, etc. (such as in the case of creating a noise profile using a noise signal obtained from the microphone 30), the physical model includes predictions of motor waveforms, chassis sound simulations, sound reflection patterns, etc.
[0047] Embodiments of the present disclosure may also include an overlay model that can integrate unexpected sounds with the existing gesture library 12. The overlay model can be calculated, for example, based on a physical model, or using an extended noise recording that can be generated in real time. Unexpected sounds from received motor actions can, for example, be the result of the robot 100 correcting itself or eliminating external unexpected forces that occur when the robot performs a gesture. The overlay model for unexpected sounds can be applied together with the pre-recorded noise model if additional unexpected activity occurs during the execution of the robot gesture.
[0048] Furthermore, in embodiments of the present disclosure, an environmental noise physics model may also be created to represent the environmental noise that may be received by the microphone 30 when the user 50 interacts with the robot. The physics model for environmental noise predicts the noise created by the environment in which the robot interacts. The physics model for environmental noise may be added to the gesture library 12 and may be mixed out of phase with the input audio to reduce the environmental noise received by the microphone 30. The gesture library 12 may include multiple environmental models, each modeling a different environment in which the robot may be present.
[0049] Once the noise model has been applied to the input audio signal in 205, the noise-canceled audio signal can be sent to the speech recognition module 310. The speech recognition module 310 translates the noise-canceled audio signal into a spoken interaction element that is used by and provided to the device 10. For example, the speech recognition module 310 can perform an analysis based on the content of the noise-canceled audio signal and can prepare an utterance to be spoken by the robot 100 as a response or answer to the user utterance contained in the noise-canceled audio signal. In addition, the gesture utterance module 320 can determine a gesture to be performed by the robot 100 based on the speech recognition module 310. The gesture can be accompanied by an utterance to be spoken by the robot 100, or alternatively, the speech recognition module 310 can determine that the utterance is not to be performed by the robot, and the gesture utterance module 320 can determine the gesture to be performed by the robot 100 without an accompanying robot utterance.
[0050] When determining an appropriate gesture for the robot 100 to accompany the robot's utterance, the server 300 may, for example, extract concepts from the utterance to be spoken by the robot and retrieve gestures corresponding to the extracted concepts from a library. A concept may be a representative extracted from a vocabulary cluster, and such concepts may include, for example, "hello," "good," "thank you," "hungry," and the like. However, the present disclosure is not limited to any particular method for selecting gestures to be performed by the robot 100.
[0051] Once the pose is acquired, the robotic system 400 can then execute Figure 4 The method shown in is to remove the robotic noise from any input audio received by microphone 30 when a gesture is performed.
[0052] Figure 5 is a block diagram of an apparatus 10 suitable for implementing one or more embodiments of the subject matter described herein. For example, the apparatus 10 may be as discussed above with reference to Figure 1 However, the apparatus 10 is not intended to suggest any limitation as to the scope of use or functionality of the subject matter described herein, as various implementations may be implemented in distributed general-purpose or special-purpose computer environments.
[0053] As shown, device 10 includes at least one processor 120 and memory 130. Processor 120 executes computer-executable instructions and can be a real or virtual processor. In a multi-processing system, multiple processors execute computer-executable instructions to increase processing power. Memory 130 can be volatile memory (e.g., registers, cache, RAM), non-volatile memory (e.g., ROM, EPROM, flash memory), or a combination thereof. Memory 130 and its associated computer-readable media provide storage for data, data structures, computer-executable instructions, and the like for device 10.
[0054] According to an implementation of the subject matter described herein, the memory 130 is coupled to the processor 120 and stores instructions for execution by the processor 120. These instructions, when executed by the processor 120, cause the apparatus to: obtain a gesture to be performed by the robot, receive input audio comprising audio from a user and noise caused by the robot performing the gesture; retrieve a noise profile associated with the gesture from a gesture library for removing noise caused by the robot performing the gesture from the input audio; and apply the noise profile to remove the robot noise from the input audio.
[0055] exist Figure 5 In the example shown in FIG, device 10 also includes one or more communication connections 140. An operating system is a mechanism for interconnecting components of device 10, such as a bus, controller, or network. Typically, an operating system provides an operating environment for other software executed in device 10 and coordinates the activities of the components of device 10.
[0056] The communication connection 140 enables communication with another computing entity over a communication medium. In addition, the functionality of the components of the device 10 can be implemented in a single computer machine or in multiple computer machines capable of communicating over the communication connection. Thus, the device 10 can operate in a network environment (e.g., the environment of the robotic system 400) using logical connections to one or more other servers, network PCs, or other public network nodes. By way of example and not limitation, the communication medium includes wired or wireless connection network technology.
[0057] Implementations of the subject matter described herein include a computer-readable medium containing computer-executable instructions that, when executed by a computer, cause the computer to perform a method for noise reduction in a robotic system in which a robot performs a gesture, the method comprising: receiving an indication that the robot performs a gesture; receiving input audio comprising user speech mixed with mechanical robotic noise caused by the robot's performance of the gesture; retrieving a noise profile associated with the gesture from a gesture library comprising a plurality of predetermined gestures paired with noise profiles; and applying the noise profile to the input audio to remove the mechanical robotic noise caused by the robot's performance of the gesture.
[0058] Computer storage media includes, but is not limited to, RAM, ROM, EPROM, EEPROM, flash memory or other solid-state memory technology, CD-ROM, DVD or other optical storage, cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by the computer.
[0059] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. When used in this specification, the terms "comprises," "comprising," "having," "including," "containing," and / or "having" specify the presence of stated features, integers, steps, operations, elements, and / or parts, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, parts, and / or groups thereof.
[0060] The corresponding structures, materials, acts, and equivalents of all parts or steps plus functional elements, if any, in the claims below are intended to include any structure, material, or act for performing that function in combination with other claimed elements as specifically claimed. This specification is presented for purposes of illustration and description and is not intended to be exhaustive or to limit the forms disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the technology. The embodiments are chosen and described in order to best explain the principles of the technology and practical applications, and to enable others of ordinary skill in the art to understand the technology in various embodiments with various modifications as are suited to the particular use contemplated.
[0061] Although specific embodiments have been described, those skilled in the art will recognize that there are other embodiments that are equivalent to the described embodiments. Accordingly, the present technology is not to be limited by the specific illustrated embodiments, but only by the scope of the appended claims.
[0062] According to one aspect of the present disclosure, a method for noise reduction in a robotic system includes obtaining a gesture to be performed by a robot; receiving input audio, the input audio comprising audio from a user and robot noise caused by the robot performing the gesture; retrieving a noise profile associated with the gesture from a gesture library, the noise profile being used to remove noise caused by the robot performing the gesture from the input audio; and applying the noise profile to remove the robot noise from the input audio.
[0063] In this aspect, the noise profile may be an inverse noise model, and applying the noise profile to remove the robotic noise from the input audio may include applying the inverse noise model to the input audio.
[0064] In this aspect, the noise profile may include a pre-recorded noise signal of the robot performing a gesture, and applying the noise profile to remove the robot noise may include mixing the pre-recorded noise signal out of phase with the input signal.
[0065] In this aspect, the gesture library may include a plurality of predetermined gestures to be performed by the robot, and each of the predetermined gestures is paired with a noise profile for removing robot noise.
[0066] In this aspect, the method may further include creating a gesture library, wherein creating the gesture library may include: causing the robot to perform predetermined gestures; and for each of the predetermined gestures, recording robot noise caused by the robot performing the gesture to create a noise profile.
[0067] In this aspect, the input audio can be received by a robot microphone, and the weapon can record the robot noise caused by the robot performing the gesture for each predetermined posture in the predetermined posture to create a noise profile, which may include: using the robot microphone to record the robot noise caused by the robot performing the gesture for each predetermined posture in the predetermined posture.
[0068] In this aspect, the gesture library includes a plurality of symbolic representations of gestures that may be performed by the robot, and each of the symbolic representations is paired with a noise profile for removing robot noise.
[0069] In this aspect, obtaining a symbolic representation of a gesture to be performed by the robot may include obtaining Laban notation defining an orientation of at least one body part of the robot with respect to a plurality of time slots.
[0070] In this aspect, the at least one body portion includes a plurality of body portions, and causing the robot to perform a gesture may include executing Laban notation to trigger the plurality of body portions to perform the gesture according to respective orientations in the plurality of time slots.
[0071] According to another aspect of the present disclosure, an apparatus for noise reduction in a robotic system includes a processor, and a memory coupled to the processor and storing instructions for execution by the processor, the instructions, when executed by the processor, causing the apparatus to: obtain a gesture to be performed by the robot; receive input audio, the input audio including audio from a user and robot noise caused by the robot performing the gesture; retrieve a noise profile associated with the gesture from a gesture library, the noise profile being used to remove the robot noise caused by the robot performing the gesture from the input audio; and apply the noise profile to remove the robot noise from the input audio.
[0072] From this aspect, the noise profile can be an inverse noise model, and applying the noise profile to remove the robotic noise from the input audio can include applying the inverse noise model to the input audio.
[0073] In this aspect, the noise profile includes a pre-recorded noise signal of the robot performing a gesture, and applying the noise profile to remove the robot noise may include mixing the pre-recorded noise signal out of phase with the input audio.
[0074] In this aspect, the gesture library may include a plurality of predetermined gestures performed by the robot, and each of the predetermined gestures is paired with a noise profile for removing robot noise.
[0075] In this aspect, when the instructions are executed by the processor, the device may further create a gesture library, wherein creating the gesture library includes: causing the robot to perform predetermined gestures; and for each predetermined gesture in the predetermined gestures, recording the robot noise caused by the user performing the gesture to create a noise profile.
[0076] In this aspect, the input audio may be received by a robot microphone; and for each predetermined posture, recording the robot noise caused by the robot performing the posture to create may include: using the robot microphone, for each predetermined posture in the predetermined postures, recording the robot noise caused by the robot performing the posture to create a noise profile.
[0077] In this aspect, obtaining a gesture to be performed by the robot may include: obtaining a symbolic representation of the gesture to be performed by the robot, and when the instruction is executed by the processor, may also include: causing the device to cause the robot to perform the gesture, including controlling the orientation of at least one body part of the robot according to the symbolic representation.
[0078] In this aspect, the gesture library may include symbolic representations of a plurality of gestures to be performed by the robot, and each symbolic representation is paired to a noise profile for removing robot noise.
[0079] In this aspect, obtaining a symbolic representation of a gesture to be performed by the robot may include obtaining Laban notation defining an orientation of at least one body portion of the robot relative to a plurality of time slots.
[0080] In this aspect, the at least one body portion may include a plurality of body portions and causing the robot to perform a gesture may include executing Laban notation to trigger the plurality of body portions to perform the gesture according to respective orientations in a plurality of time slots.
[0081] According to another aspect of the invention, a computer-readable storage medium includes computer-executable instructions that, when executed by a computer, cause the computer to perform a method for noise reduction in a robotic system in which a robot performs a gesture, the method comprising: receiving an indication that the robot is performing a gesture; receiving input audio comprising user speech mixed with mechanical robotic noise caused by the robot performing the gesture; retrieving a noise profile associated with the gesture from a gesture library comprising a plurality of predetermined gestures paired with noise profiles; and applying the noise profile to the input audio to remove the mechanical robotic noise caused by the robot performing the gesture.
[0082] In this aspect, the plurality of noise profiles of the gesture library may include pre-recorded noise signals of the robot performing predetermined gestures, and applying the noise profiles to remove robot noise may include mixing the pre-recorded noise signals associated with the gestures out of phase with the input signal.
[0083] In this aspect, the plurality of predetermined gestures may be represented in the gesture library by a plurality of symbolic representations of the predetermined gestures, wherein the symbolic representations define an orientation of at least one body portion of the robot when performing the gestures.
Claims
1. A method for noise reduction in a robotic system, the method comprising: Get the pose to be performed by the robot; receiving input audio, the input audio comprising audio from a user, robot noise caused by the robot performing the gesture, and unexpected noise during the robot performing the gesture; retrieving a noise profile associated with the gesture from a gesture library, the noise profile being used to remove, from the input audio, the robot noise caused by the robot performing the gesture, wherein the noise profile is an inverse noise model, and applying the noise profile to remove the robot noise from the input audio comprises: applying the inverse noise model to the input audio; applying the noise profile to remove the robotic noise from the input audio; and applying a coverage model for undesired noise to remove the undesired noise from the input audio, Wherein the noise profile comprises a pre-recorded noise signal of the robot performing the gesture, and applying the noise profile to remove the robot noise from the input audio comprises mixing the pre-recorded noise signal out of phase with the input audio. 2 . The method of claim 1 , wherein the gesture library comprises a plurality of predetermined gestures performed by the robot, and each of the predetermined gestures is paired with a noise profile for removing robot noise.
3. The method according to claim 2, further comprising creating the gesture library, wherein creating the gesture library comprises: causing the robot to perform the predetermined posture; as well as For each of the predetermined gestures, a robot noise caused by the robot performing the gesture is recorded to create a noise profile.
4. The method of claim 3 , wherein the input audio is received by a robot microphone, and for each of the predetermined gestures, recording robot noise caused by the robot performing the gesture to create a noise profile comprises: For each of the predetermined gestures, a robot noise caused by the robot performing the gesture is recorded using the robot microphone.
5. The method according to claim 1, wherein: Obtaining a gesture to be performed by the robot comprises obtaining a symbolic representation of the gesture to be performed by the robot, and The method also includes causing the robot to perform the gesture, wherein causing the robot to perform the gesture includes controlling an orientation of at least one body portion of the robot according to the symbolic representation. 6 . The method of claim 5 , wherein the gesture library comprises a plurality of symbolic representations of gestures to be performed by the robot, and each of the symbolic representations is paired with a noise profile for removing robot noise.
7. The method of claim 6 , wherein obtaining a symbolic representation of a gesture to be performed by the robot comprises: Laban notation defining an orientation of at least one body portion of the robot relative to a plurality of time slots is obtained.
8. A device for reducing noise in a robot system, comprising processor; a memory coupled to the processor and storing instructions for execution by the processor, the instructions, when executed by the processor, causing the apparatus to: Get the pose to be performed by the robot; receiving input audio, the input audio comprising audio from a user, robot noise caused by the robot performing the gesture, and unexpected noise during the robot performing the gesture; retrieving a noise profile associated with the gesture from a gesture library, the noise profile being used to remove, from the input audio, the robot noise caused by the robot performing the gesture, wherein the noise profile is an inverse noise model, and applying the noise profile to remove the robot noise from the input audio comprises: applying the inverse noise model to the input audio; applying the noise profile to remove the robotic noise from the input audio; and applying a coverage model for undesired noise to remove the undesired noise from the input audio, Wherein the noise profile comprises a pre-recorded noise signal of the robot performing the gesture, and applying the noise profile to remove the robot noise from the input audio comprises mixing the pre-recorded noise signal out of phase with the input audio. 9 . The apparatus of claim 8 , wherein the gesture library includes a plurality of predetermined gestures performed by the robot, and each of the predetermined gestures is paired with a noise profile for removing robot noise.
10. The apparatus of claim 9, wherein the instructions, when executed by the processor, further cause the apparatus to create the gesture library, wherein creating the gesture library comprises: causing the robot to perform the predetermined posture; as well as For each of the predetermined gestures, a robot noise caused by the robot performing the gesture is recorded to create a noise profile.
11. The apparatus according to claim 8, wherein: Obtaining a gesture to be performed by the robot comprises obtaining a symbolic representation of the gesture to be performed by the robot, and The instructions, when executed by the processor, further cause the apparatus to cause the robot to perform the gesture including controlling an orientation of at least one body portion of the robot according to the symbolic representation.
12. The apparatus of claim 11, wherein the gesture library comprises a plurality of symbolic representations of gestures to be performed by the robot, and each of the symbolic representations is paired with a noise profile for removing robot noise.
13. The apparatus of claim 12 , wherein obtaining a symbolic representation of a gesture to be performed by the robot comprises: Laban notation defining an orientation of at least one body portion of the robot relative to a plurality of time slots is obtained.
14. A computer-readable storage medium comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a computer, cause the computer to perform a method for reducing noise in a robotic system in which a robot performs a gesture, the method comprising: receiving an indication that the robot is performing the gesture; receiving input audio comprising user speech mixed with mechanical robot noise caused by the robot performing the gesture and unexpected noise during the robot performing the gesture; as well as retrieving a noise profile associated with the gesture from a gesture library comprising a plurality of predetermined gestures paired with noise profiles, wherein the noise profile is an inverse noise model, and applying the noise profile to remove the robotic noise from the input audio comprises: applying the inverse noise model to the input audio; applying the noise profile to the input audio to remove the mechanical robot noise caused by the robot performing the gesture; and applying a coverage model for undesired noise to remove the undesired noise from the input audio, Wherein the noise profile comprises a pre-recorded noise signal of the robot performing the gesture, and applying the noise profile to remove the robot noise from the input audio comprises mixing the pre-recorded noise signal out of phase with the input audio.
15. The computer-readable storage medium of claim 14, wherein the plurality of noise profiles of the gesture library include pre-recorded noise signals of the robot performing the predetermined gestures, and applying the noise profiles to remove the robot noise from the input audio comprises: A pre-recorded noise signal of the noise profile associated with the gesture is mixed out of phase with the input audio.
16. The computer-readable storage medium of claim 14, wherein the plurality of predetermined poses are represented in the pose library by a plurality of symbolic representations of the predetermined poses, wherein the symbolic representations define an orientation of at least one body portion of the robot when performing a pose.
17. The computer-readable storage medium of claim 16, the method further comprising: Obtaining a gesture to be performed by the robot comprises obtaining a symbolic representation of the gesture to be performed by the robot, and The robot is caused to perform the gesture, wherein causing the robot to perform the gesture comprises controlling an orientation of at least one body portion of the robot according to the symbolic representation.
Citation Information
Patent Citations
Portable hearing examination device
CN106308812A
Acoustic data processor and acoustic data processing method
US20100299145A1