Mobile robot with audio sensing system

By using microphone arrays and machine learning models, semantic audio scene data is generated, which solves the problem of mobile robots' difficulty in navigation in complex environments, and achieves more efficient environment perception and operation capabilities.

CN120122633APending Publication Date: 2025-06-10ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411797623.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-07
Filing Date
2024-12-09
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing mobile robots are difficult to navigate effectively in some scenarios, such as in the event of insufficient lighting, camera failure or object obscurity.

Method used

Audio signals are received through the microphone array, audio characteristic data of acoustic activity is extracted, and data of arrival directions of acoustic activity is generated. Using machine learning models and knowledge graphs, semantic audio scene data is generated to guide the movement of mobile robots.

Benefits of technology

It realizes effective navigation and identification of sound sources in complex environments, and improves the operational capabilities of mobile robots in a variety of scenarios, including in the event of insufficient lighting or sensor failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120122633A_ABST
    Figure CN120122633A_ABST
Patent Text Reader

Abstract

A mobile robot with an audio perception system is provided. The mobile robot includes a microphone array having a set of microphones. The microphone array is at least partially disposed on the mobile robot. The mobile robot receives audio signals from the microphone array. Audio feature data of the acoustic activity is extracted from the audio signal. Direction of arrival (DOA) data of acoustic activity is generated based on the audio signal. The machine learning model is configured to generate audio event data using the audio feature data. The audio event data identifies at least one sound source of the audio feature data. The knowledge graph is queried using the audio event data to obtain entity data. The entity data has a predetermined relationship with the audio event data. Semantic audio scene data is generated using the audio event data, the DOA data, and the entity data. The mobile robot performs an action based on the semantic audio scene data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to mobile robots, and more particularly, to mobile robots having digital audio processing for semantic perception of audio scenes. Background Art

[0002] Mobile robots are sometimes used in the residential domain. For example, a vacuuming robot can be used to vacuum a room with minimal user interaction. To this end, these vacuuming robots use sensors to sense their environment and navigate around various obstacles. For example, the sensors can include cameras, motion sensors, collision sensors, etc. However, these vacuuming robots may have difficulty navigating around objects in many scenarios, such as when there is insufficient lighting, a camera malfunction, an object occlusion, etc. Summary of the Invention

[0003] The following is an overview of certain embodiments described in detail below. The presented aspects are merely provided to give the reader a brief overview of these certain embodiments, and the description of these aspects is not intended to limit the scope of the present disclosure. Indeed, the present disclosure may cover various aspects that may not be explicitly set forth below.

[0004] According to at least one aspect, a computer-implemented method involves controlling a mobile robot in an environment. The method includes receiving an audio signal via a microphone array. The microphone array is at least partially disposed on the mobile robot. The method includes extracting audio feature data of acoustic activity from the audio signal. The method includes generating direction-of-arrival (DOA) data of acoustic activity based on the audio signal. The method includes generating audio event data using the audio feature data via at least one machine learning model. The audio event data identifies at least one sound source of the audio feature data. The method includes extracting entity data by querying a knowledge graph using the audio event data. The entity data has a relationship with the audio event data. The method includes generating semantic audio scene data using the audio event data, the DOA data, and the entity data. The method includes performing an action of the mobile robot based on the semantic audio scene data.

[0005] According to at least one aspect, a mobile robot includes at least a microphone array, one or more processors, and one or more memories. The one or more processors communicate data with the microphone array. The one or more memories communicate data with the one or more processors. The one or more memories include computer-readable data stored thereon, the computer-readable data including instructions that, when executed by the one or more processors, perform a method. The method includes receiving an audio signal via the microphone array. The microphone array is at least partially disposed on the mobile robot. The method includes extracting audio feature data of acoustic activity from the audio signal. The method includes generating DOA data of acoustic activity based on the audio signal. The method includes generating audio event data using the audio feature data via at least one machine learning model. The audio event data identifies at least one sound source of the audio feature data. The method includes extracting entity data by querying a knowledge graph using the audio event data. The entity data has a relationship with the audio event data. The method includes generating semantic audio scene data using the audio event data, the DOA data, and the entity data. The method includes performing an action of the mobile robot based on the semantic audio scene data.

[0006] According to at least one aspect, one or more non-transitory computer-readable media have computer-readable data stored thereon, the computer-readable data including instructions that, when executed by one or more processors, cause the one or more processors to perform a method for controlling a mobile robot in an environment. The method includes receiving an audio signal via the microphone array. The microphone array is at least partially disposed on the mobile robot. The method includes extracting audio feature data of acoustic activity from the audio signal. The method includes generating DOA data of acoustic activity based on the audio signal. The method includes generating audio event data using the audio feature data via at least one machine learning model. The audio event data identifies at least one sound source of the audio feature data. The method includes extracting entity data by querying a knowledge graph using the audio event data. The entity data has a relationship with the audio event data. The method includes generating semantic audio scene data using the audio event data, the DOA data, and the entity data. The method includes performing an action of the mobile robot based on the semantic audio scene data.

[0007] These and other features, aspects, and advantages of the present invention are discussed in the following detailed description with reference to the drawings, in which like characters represent like or identical parts throughout. Additionally, the drawings are not necessarily to scale, as some features may be enlarged or minimized to show details of particular components. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 is a diagram of an example of a mobile robot including an audio perception system according to an example embodiment of the present disclosure.

[0009] Figure 2 It is a diagram of an example of an audio perception system according to an exemplary embodiment of the present disclosure.

[0010] Figure 3 It is a block diagram of an example of a signal processing system according to an exemplary embodiment of the present disclosure.

[0011] Figure 4 It is a flowchart of an example of an audio perception system according to an exemplary embodiment of the present disclosure.

[0012] Figure 5 It is a flowchart of an example of a process related to an audio perception system according to an exemplary embodiment of the present disclosure.

[0013] Figure 6 It is a block diagram of an example of a system including a mobile robot according to an exemplary embodiment of the present disclosure. Detailed Description

[0014] Embodiments are described herein, which have been illustrated and described by way of example, and by the foregoing description, many of their advantages will be understood, and it will be apparent that various changes can be made in the form, construction, and arrangement of the components without departing from the disclosed subject matter or sacrificing one or more of its advantages. In fact, the description of these embodiments is merely explanatory. These embodiments are susceptible to various modifications and alternative forms, and the appended claims are intended to cover and include such changes and are not limited to the specific forms disclosed, but cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosure.

[0015] Figure 1 It is a diagram of a non - limiting example of a mobile robot 100 and a docking station 102 according to an exemplary embodiment. More specifically, the mobile robot 100 is configured to connect or engage with the docking station 102, and the docking station 102 includes at least a power supply. The power supply is configured to charge the battery of the mobile robot 100 when the mobile robot 100 is connected or engaged with the docking station 102. In addition, the mobile robot 100 is configured to disconnect or disengage from the docking station 102 and move around its environment autonomously or at least partially autonomously.

[0016] In addition, the mobile robot 100 is configured to perform one or more actions. For example, in Figure 1In this case, the mobile robot 100 is a cleaning robot, which includes a cleaning device (e.g., a vacuum assembly, etc.). The mobile robot 100 is configured to perform actions related to navigating and moving around its environment. The mobile robot 100 is also configured to perform at least one action related to cleaning a surface (e.g., a floor, a carpet, etc.) when it is stationary in its environment, when it is moving around the environment, or when transitioning between being stationary and moving (and vice versa). Additionally, the mobile robot 100 is configured to perform one or more other actions (e.g., sending a notification, sending an alert, etc.) with respect to specific detected, specifically identified audio events, and / or specific semantic audio scene data, as described later in this disclosure.

[0017] In addition, as Figure 1 shown, the mobile robot 100 includes a microphone array 104. The microphone array 104 may include a set of microphones. In some implementations and / or embodiments, the microphone array may be spherical, binaural, or any compatible type. The microphone array 104 may include one or more omnidirectional microphones, which are configured to receive sound or acoustic information from all directions.

[0018] The microphone array 104 includes multiple microphones, which are configured to detect external sounds of the mobile robot 100, internal sounds of the mobile robot 100, or both internal and external sounds of the mobile robot 100. In some implementations and / or embodiments, the microphone array 104 is only provided on or carried by the mobile robot 100. In this regard, the microphone array 104 may be located on or carried by various surfaces of the mobile robot 100 itself to detect acoustic activities in its environment, its own noise, or a combination thereof. In other implementations and / or embodiments, the microphone array 104 includes a first subset of microphones provided on or carried by the mobile robot 100, and a second subset of microphones strategically located elsewhere (i.e., not provided on and / or carried by the mobile robot 100 itself). For example, in Figure 1 this case, the mobile robot 100 includes a microphone array 104, which includes a first subset 104A of microphones located on the mobile robot 100 and a second subset 104B of microphones located on the docking station 102. The first subset 104A is configured to detect acoustic activities in the environment of the mobile robot, its own noise, or a combination thereof. The second subset 104B includes one or more microphones, which can be used to monitor the environmental area near and around the docking station 102. In Figure 1 the example shown, the number of microphones in the first subset 104A is greater than the number of microphones in the second subset 104B.

[0019] Regarding the placement of microphones on the mobile robot 100, the microphone array 104 can include microphones located on or carried by various surfaces of the mobile robot 100. For example, the microphone array 104 can include multiple microphones that are disposed on one or more surfaces (e.g., top surface, upper surface, outer surface, etc.) of the mobile robot 100. The microphone array 104 can include one or more microphones that are positioned along one or more circumferential or peripheral portions of the mobile robot 100. The microphone array 104 can include one or more microphones that are positioned along one or more sides of the mobile robot 100. For example, in Figure 1 the microphone array 104 includes multiple microphones that are located in a peripheral portion on the outer surface portion of the mobile robot 100. In this example, the microphones are evenly spaced around the circumference of the upper or top surface portion of the mobile robot 100.

[0020] In addition, one or more surfaces of the mobile robot 100 can further include a power button 108 and a light indicator 110. For example, in Figure 1 the power button 108 and the light indicator 110 are located on the upper / top, outer surface of the mobile robot 100. The status or color of the light indicator 110 can indicate the status of the mobile robot 100 (e.g., battery status, operating status, etc.).

[0021] Regarding the placement of microphones on the docking station 102, the microphone array 104 can be located on or carried by various surfaces of the docking station 102. For example, the microphone array 104 can include multiple microphones that are disposed on one or more surfaces (e.g., top surface, upper surface, outer surface, etc.) of the docking station 102. The microphone array 104 can include one or more microphones that are positioned along one or more circumferential or peripheral portions of the docking station 102. The microphone array 104 can include one or more microphones that are positioned along one or more sides of the docking station 102. For example, in Figure 1 the microphone array 104 includes multiple microphones that are located on the top outer surface portion of the docking station 102. As Figure 1 shown, the microphones are located at opposite ends of the docking station 102. In this example, the microphones are aligned along the longitudinal axis of the docking station 102.

[0022] The mobile robot 100 is configured to obtain acoustic information of its environment via the microphone array 104, while also collecting other information from one or more other sensors 106 (e.g., cameras, LIDAR, etc.) of the sensor system 608. More specifically, with respect to obtaining acoustic information, the mobile robot 100 is configured to receive multi-channel audio signals from the microphone array 104. The mobile robot 100 includes an audio perception system 200 that receives and processes the audio signals obtained from the microphone array 104. The mobile robot 100 is configured to generate semantic audio scene data based on the original audio signal. The mobile robot 100 is configured to use semantic audio scene data, other information (e.g., digital images, map data, motion sensor data, metadata, other sensor data, etc.), semantic maps, or any multiple and combination thereof when operating in its environment.

[0023] The mobile robot 100 is configured to operate in a variety of different modes associated with the audio perception system 200. For example, the mobile robot 100 is configured to operate in an audio monitoring mode, in which the mobile robot 100 is configured to listen for acoustic activity without performing certain sound-inducing actions (e.g., moving, vacuuming, etc.) and / or with minimal or no self-noise, so that the mobile robot 100 can most effectively detect and identify acoustic activity by remaining stationary in its environment. As another example, the mobile robot 100 is configured to operate in a patrol mode, in which the mobile robot 100 minimizes self-noise by not performing certain actions (e.g., vacuuming), so that the mobile robot 100 can better detect and identify acoustic activity in its environment. The mobile robot 100 is also able to identify its own internal noise or self-noise that is not related to certain actions such as cleaning. More specifically, in patrol mode, the mobile robot 100 can move and / or stay around its environment while disabling some actions (e.g., debris suction, vacuuming, etc.), so that the mobile robot 100 can better identify acoustic activities in the environment than when the mobile robot 100 is cleaning (e.g., vacuuming). In addition, as yet another example, the mobile robot 100 is configured to operate in a normal operating mode. In normal operating mode, the mobile robot 100 is configured to perform at least one action (e.g., vacuuming) while stationary or moving in its environment. In normal operating mode, the mobile robot 100 is configured to detect and identify acoustic activities in its environment, and also identify acoustic activities related to the mobile robot 100 itself (e.g., internal noise, self-noise, etc.). As discussed above, the mobile robot 100 is more effectively controlled by using acoustic information obtained from audio signals via the audio perception system 200.

[0024] Figure 2 , Figure 3 and Figure 4Shows an example aspect of an audio perception system 200 according to an example embodiment. More specifically, as Figure 2 and Figure 4 shown, the audio perception system 200 is configured to process audio signals captured by a microphone array 104. The audio perception system 200 is configured to filter and detect acoustic information of interest. The audio perception system 200 is configured to identify one or more sound patterns of interest, one or more sound sources, and the corresponding location of each sound source. The audio perception system 200 is configured to generate semantic audio scene data using the raw audio signals. For example, in Figure 2 and Figure 4 , the audio perception system 200 includes at least a signal processing system 202, a machine learning (ML) system 204, and a knowledge graph (KG) 206.

[0025] Referring to Figure 3 , the signal processing system 202 is configured to receive multi-channel audio signals from the microphone array. The signal processing system 202 is configured to process the multi-channel audio signals and generate audio feature data of these audio signals for the ML system 204. The signal processing system 202 includes a plurality of modules. For example, in Figure 3 , the set of modules includes a non-silence detection module 300, a voice activity detection module 302, an audio segmentation module 304, a noise cancellation module 306, an environment learning module 308, a DOA estimation module 310, a dereverberation module 312, a room impulse response (RIR) estimation module 314, and an acoustic echo cancellation module 316. The signal processing system 202 may include any number of modules from this set of modules and combinations thereof.

[0026] The non-silence detection module 300 is configured to detect non-silence regarding the audio signals. In this regard, for example, the non-silence detection module 300 is configured to detect acoustic activity from the audio signals. The signal processing system 202 is configured to extract or generate audio feature data related to the non-silence detection or acoustic activity of the audio signals.

[0027] The voice activity detection module 302 is configured to identify and distinguish audio signal segments containing human or non-human speech. In this regard, for example, the voice activity detection module 302 is configured to detect voice activity from the audio signals, filter out the voice activity from the audio signals to obtain non-voice activity, and then generate audio feature data related to these non-voice activities. Thus, the voice activity detection module 302 is configured to generate audio feature data related to voice activity and non-voice activity regarding the audio signals.

[0028] In addition, the signal processing system 202 may include a voice filtering module that detects voice (e.g., human speech) in the audio signal. The signal processing system 202 may filter and detect voice information of interest. For example, in some implementations and / or embodiments, after detecting that an audio segment is human speech, the audio perception system 200 is configured to classify the audio segment with respect to one of the control commands (e.g., stop vacuuming, start vacuuming, return to docking station 102, etc.) for controlling the mobile robot 100. In addition, the signal processing system 202 is configured to remove or encrypt voice detections related to privacy issues before performing further downstream tasks. Such removal or encryption is configured to address user privacy concerns.

[0029] The audio segmentation module 304 is configured to provide audio feature data related to one or more segments of the audio signal. These segments may include different durations. For example, these segments may include a first segment of the acoustic activity of the audio signal having a first time length, a second segment of another acoustic activity of the audio signal having a second time length, and so on. As an example, for instance, the length of a segment may be determined by the duration of the detected acoustic activity (e.g., voice activity, etc.) in the audio signal.

[0030] The noise cancellation module 306 is configured to reduce unwanted sound by adding another sound specifically designed to cancel the unwanted sound. The noise cancellation module 306 is configured to provide audio feature data related to the noise cancellation performed on the audio signal.

[0031] The environmental learning module 308 is configured to detect and monitor dynamic changes in the acoustic environmental conditions. The environmental learning module 308 is configured to extract knowledge from these acoustic environmental conditions, such as identifying various noise types and mapping the distribution of the background acoustics. The environmental learning module 308 serves as a valuable resource for enhancing the functionality of other signal processing components as well as the ML system 204. The environmental learning module 308 is configured to generate audio feature data related to the acoustic environmental conditions.

[0032] The DOA estimation module 310 is configured to perform direction of arrival (DOA) estimation by beamforming, a deep learning framework, any suitable DOA estimation technique, or any combination thereof. In addition, the signal processing system 202 is configured to generate DOA data based on the DOA estimation. The DOA data provides the relative direction of each sound source detected by the mobile robot 100. The DOA estimation module 310 is configured to generate DOA data related to the estimation of the relative direction of each sound source (e.g., 30 degrees, 45 degrees, etc.). In addition, the DOA estimation module 310 is configured to generate audio feature data related to the DOA data of the audio signal.

[0033] The dereverberation module 312 is configured to cancel the reverberation effect in an enclosed space (such as a room and a hall). Such sound reflection is a natural acoustic phenomenon, which may cause audio quality degradation. The dereverberation module 312 is configured to enhance the clarity and intelligibility of the audio signal. The dereverberation module 312 is configured to analyze the characteristics of the reverberation components in the audio signal and apply corrective measures to reduce or eliminate them. The dereverberation module 312 is configured to generate audio feature data related to such dereverberation of the audio signal.

[0034] The RIR estimation module 314 is configured to perform RIR estimation for the audio signal. The RIR estimation module 314 is configured to provide audio feature data related to the RIR estimation, which measures and models the way sound interacts with its environment, including reflection, reverberation, and echo. The audio feature data provides information and / or sound behavior in a specific environment. More specifically, for example, the audio feature data related to the RIR estimation provides insights into room characteristics, such as room surface data, room geometry data, room material data, etc. Such audio feature data supports the interpretation of the sound reflection pattern within a given space. As a non-limiting example, for instance, the RIR estimation module 314 is configured to perform RIR estimation by emitting a sharp sound and analyzing the subsequently recorded return signal. The dereverberation module 312 is configured to generate audio feature data related to the RIR estimation.

[0035] The acoustic echo cancellation module 316 is designed to eliminate or reduce the presence of acoustic echo in the audio signal. Acoustic echo is generated when the sound emitted by a speaker is captured by a microphone and re-introduced into the audio signal, resulting in unwanted feedback. The module is configured to identify the echo components in the audio signal and apply corrective measures to eliminate or reduce their presence. Common techniques such as adaptive filtering can be used to dynamically adapt to the characteristics of the echo. The acoustic echo cancellation module 316 is customized to enhance the overall intelligibility and quality of the audio signal. The acoustic echo cancellation module 316 is configured to generate audio feature data related to the acoustic echo cancellation.

[0036] As discussed above, the signal processing system 202 includes a plurality of modules configured to receive raw audio signals from the microphone array 104 of the microphone group. The raw audio signals are multi-channel audio signals. The group of modules can process the audio signals sequentially (e.g., one or more modules execute simultaneously, sequentially, etc.), which effectively generates audio feature data. For example, the signal processing system 202 is configured to perform signal processing to remove irrelevant information (e.g., noise, reverberation, echo, etc.) from the raw audio signals and extract meaningful audio feature data from the raw audio signals. In this regard, the signal processing system 202 can process the audio signals via a noise cancellation module 306, a dereverberation module 312, an acoustic echo cancellation module 316, or any combination thereof, prior to other modules (e.g., an audio segmentation module 304, a voice activity detection module, etc.) of the signal processing system 202.

[0037] The signal processing system 202 is configured to perform signal processing via filtering, signal transformation, machine learning (e.g., autoencoders, pre-trained ML models, etc.), or any audio processing means. The signal processing system 202 uses signal processing techniques to generate or extract audio feature data using the raw audio signals. The audio feature data can include one or more audio feature vectors, embedding data, any suitable ML format for the ML system 204, or any combination of them. The audio feature data can include DOA data, metadata, etc. As a non-limiting example, the metadata can include, for example, environmental information such as the distribution mean and standard deviation of the last ten minutes from some parts of the audio signal. The signal processing system 202 provides the audio feature data to the ML system 204.

[0038] As Figure 2 and Figure 4 shown, the ML system 204 is configured to receive the audio feature data from the signal processing system 202. The ML system 204 can include a support vector machine (SVM), a Gaussian mixture model (GMM), a hidden Markov model (HMM), a neural network, any suitable machine learning model, or any combination of them. As a non-limiting example, the ML system 204 can include a convolutional neural network (CNN), a recurrent neural network (RNN), a long short-term memory (LSTM), any suitable neural network, or any combination of them. The ML system 204 can include one or more pre-trained ML models. Alternatively, for the signal processing system 202 and the ML system 204, the audio perception system 200 can include an end-to-end machine learning system (e.g., a neural network architecture) that extracts audio feature data from the audio signals and identifies each sound pattern or acoustic activity of the audio signals.

[0039] The ML system 204 is configured to receive audio feature data from the signal processing system 202. The ML system 204 is configured to identify one or more interesting sound patterns regarding the audio feature data. The ML system 204 is configured to identify at least one acoustic event of the audio feature data and generate audio event data using the audio feature data. The audio event data identifies at least one sound source of the audio feature data. The audio event data can identify various sound sources. As a non-limiting example, the audio event data can identify acoustic activities related to baby crying, glass breaking, dog barking, doorbell ringing, knocking on the door, kitchen sounds (e.g., cooking, frying, cutting food, opening cabinets, etc.), sirens, gunshots, radio, television, sneezing, coughing, screaming, etc. Refer to Figure 4 , as a non-limiting example, the ML system 204 is configured to classify the audio feature data as "TV" with a probability data of 0.95. Given the probability data of "TV" for the audio feature data (or the confidence score data of 0.95), the audio perception system 200 is then configured to select and use "TV" as the audio event data regarding the multi-channel audio signal.

[0040] In some implementations and / or embodiments, the ML system 204 is configured to identify surface characteristics. In this regard, for example, the ML system 204 is configured to identify whether the floor is carpet, concrete, hardwood, linoleum, or another floor material. The ML system 204 can be trained or pre-trained for the sounds of the mobile robot 100 moving on these different types of floors. The ML system 204 can include a classifier that is configured to generate audio event data (e.g., floor classification data such as carpet, tile, hardwood, etc.) using the audio feature data. Using this audio event data and / or semantic audio scene data, the mobile robot 100 can be configured to perform a predetermined action, such as vacuuming only when the mobile robot 100 identifies that it is on the carpet. As another example, the mobile robot 100 can perform an action based on the semantic audio scene data, such as switching from one mode (e.g., vacuuming mode) to another mode (e.g., audio monitoring mode, etc.). The semantic audio scene data helps the mobile robot 100 identify, monitor, and track various events and / or changes in its environment.

[0041] In some implementations and / or embodiments, the ML system 204 is configured to identify acoustic activities related to when the mobile robot 100 begins to interact (e.g., vacuum) with any unacceptable items (e.g., clothes, curtains, cables, etc.). For example, the ML system 204 may include a pre-trained ML model that identifies one or more internal sounds indicating when the dust / debris collector is full or substantially full. When the ML system 204 generates audio event data indicating a potential problem with the cleaning assembly (e.g., the suction assembly, the dust / debris collector, etc.), the mobile robot 100 may activate an alarm to notify the user of the potential problem via an I / O device, a mobile communication device 604, etc.

[0042] In addition, in some implementations and / or embodiments, the ML system 204 is configured to identify abnormal sounds, which may include internal sounds of the mobile robot 100, external sounds of the mobile robot 100, or a combination of internal and external sounds of the mobile robot 100. Regarding internal sounds from the mobile robot 100 itself, the ML system 204 is configured to learn the normal operating sounds of the mobile robot 100 and then use these normal operating sounds to detect abnormal sounds and / or anomalies (e.g., component failures, unacceptable behaviors, etc.) of the mobile robot 100. In addition, in some implementations and / or embodiments, the distribution (e.g., clustering, etc.) of the normal operating sound patterns of the mobile robot 100 may be used to detect other sound patterns that deviate beyond a predefined threshold. The ML system 204 may label and / or generate audio event data indicating that these other sound patterns are abnormal or abnormal sounds.

[0043] In some implementations and / or embodiments, the ML system 204 is configured to identify sounds that may require a user's attention and interaction. For example, in some implementations and / or embodiments, the ML system 204 includes at least one pre-trained ML model to detect knocking sounds, whistling sounds, and / or clicking sounds on a window, as these sounds can indicate (i) the window being open during inclement weather, (ii) weather stripping wear, (iii) another window problem, or (iv) any combination of them. In some implementations and / or embodiments, the ML system 204 includes at least one pre-trained ML model to detect vibration sounds or electrical noises in a residential wall, which indicate an overloaded circuit breaker, a loose power outlet, etc. In some implementations and / or embodiments, the ML system 204 includes at least one pre-trained ML model to identify clinking sounds from a plumbing system or the sound of water flowing in a residential wall as an indication of a plumbing or piping problem. In some implementations and / or embodiments, the ML system 204 includes at least one pre-trained ML model to detect abnormal or persistent sounds from a furnace. In some implementations and / or embodiments, the ML system 204 includes at least one pre-trained ML model to detect acoustic activities (e.g., humming noises, electrical noises, etc.) from appliances (e.g., refrigerators, dishwashers, dryers, power outlets, etc.) or malfunctioning appliances. In some implementations and / or embodiments, the ML system 204 may include at least one pre-trained ML model to detect jumping and scratching sounds of unwanted wildlife (e.g., rodents, raccoons, birds, etc.) in the interior walls or attic of a house. In any of these implementations, the audio perception system 200 is configured to transmit an alert to notify the user of the audio event data generated by the ML system 204.

[0044] Next, the audio perception system 200 uses the audio event data to query the KG 206. The KG 206 includes a knowledge base that includes interrelated descriptions of entities and the encoding of the semantics or relationships behind these entities. More specifically, as an example, the KG 206 captures the spatial relationships between (i) object / object, (ii) object / region, and (iii) region / region. In this regard, the audio perception system 200 is configured to use the reasoning of the KG 206 to determine the region of the detected sound source, such that the mobile robot 100 can navigate towards or away from the region with the detected object sound according to the situation. As a non-limiting example, in Figure 4In this case, the audio perception system 200 is configured to query the KG 206 using the audio event data of "TV". Based on this query, the KG 206 can determine that "TV" has a relationship (e.g., "is located in") with the "living room" with display probability data of "0.7", and has another relationship (e.g., "is located in") with the "kitchen" with display probability data of "0.2", and so on. The audio perception system 200 can then use the probability data itself and / or together with the DOA data to determine that the audio event of "TV" is most likely to occur in the "living room", and then select the "living room" as the entity data of the semantic audio scene data.

[0045] The KG 206 includes knowledge of room prototypes, room configurations, residential structures, and other relevant data, which are pre-prepared based on the geographical area of the mobile robot 100 and / or the culture of the geographical area. In this regard, there may be different probabilities associated with specific objects appearing in certain locations of the residence. For example, in some regions of Asia, the kitchen area of some residences may be separate and located outside the residence. For these users, the KG 206 can include knowledge, prototypes, and / or templates compatible with this geographical area.

[0046] In addition, when the mobile robot 100 is first introduced into a new environment (e.g., a new residence), the ML system 204 can include at least one pre-trained ML model that has not been fully adapted to this new environment. In the case of including the KG 206, the mobile robot 100 is configured to identify sounds based on the following knowledge: (i) room prototypes, (ii) common sense knowledge of the probability of sounds generated by objects and the probability of objects appearing in specific areas. As a non-limiting example, for instance, the mobile robot 100 is configured to detect and identify the chopping sound most likely coming from the kitchen, the rumbling or vibrating sound of the washing machine most likely coming from the laundry area, etc.

[0047] Reference Figure 4 Regarding the KG 206, the audio perception system 200 is configured to generate entity data using the audio event data. More specifically, the audio perception system 200 is configured to query the KG 206 using the audio event data and extract the entity data that has the maximum probability relationship with the audio event data. As a non-limiting example, in Figure 4In this case, the audio perception system 200 is configured to query the KG 206 using "TV" (i.e., the first identified audio event). The KG 206 is configured to provide context information that enhances the understanding of the detected audio event "TV". Based on this query of the KG 206, the audio perception system 200 determines that "living room" is the entity data (e.g., location data) that has the highest probability of a relationship (e.g., "located in") with "TV" (i.e., the audio event data). Additionally or alternatively, the audio perception system 200 can be configured to extract other entity data based on other relationships connected to the audio event data via the KG 206, thereby providing additional context information for the audio event (in addition to or instead of the location data). This seamless interaction between the ML system 204 and the KG 206 enhances the context relevance of the audio event data, thus contributing to a more robust and intelligent decision-making process for the mobile robot 100.

[0048] The audio perception system 200 generates semantic audio scene data, which includes multiple audio feature data (e.g., DOA data, metadata, etc.), audio event data (e.g., sound source data), and related entity data (e.g., location data of the sound source). As a non-limiting example, in Figure 4 this case, the audio perception system 200 is configured to generate semantic audio scene data based on an audio signal, where the semantic audio scene data includes: (i) first audio scene data indicating a first audio event of a TV in the 30-degree direction in the living room; (ii) second audio scene data indicating a second audio event of a conversation (e.g., speech activity) between humans in the 45-degree direction in the living room; and (iii) so on (if detected and available). In one example, given a scenario where the TV is on in the living room, the first audio scene data represents the most likely audio scene data for a given multi-channel audio signal, and the second audio scene data represents the next most likely audio scene data for the same multi-channel audio signal. In another example, given another scenario where two people are having a conversation in the living room while the TV in the living room is on, the first audio scene data and the second audio scene data represent the two audio events with the highest probability of occurring simultaneously detected from the same multi-channel audio signal.

[0049] Refer to Figure 4 , as an example, the mobile robot 100 can perform one or more actions (e.g., perform vacuuming, stop vacuuming, etc.) based on the semantic audio scene data. For example, in Figure 4In this case, when the TV transitions from the off state to the on state and / or when two people are having a conversation in the living room, the mobile robot 100 is configured to generate semantic audio scene data reflecting the current state of the living room and perform a navigation action away from the living room (or navigate to other areas of its environment before going to the living room) to avoid disturbing the activities in the living room with its cleaning at the current time. As discussed above, the mobile robot 100 is configured to use the semantic audio scene data in many applications and / or downstream tasks. In addition, the mobile robot 100 is configured to use the semantic audio scene data to advantageously enhance its environmental layout data or map data.

[0050] Figure 5 is a flowchart of an example of a process 500 of a mobile robot 100 related to an audio perception system 200 according to an example embodiment. In this example, the process 500 includes a plurality of steps associated with generating and / or updating a semantic map. The semantic map can be used to control the mobile robot 100. For example, the mobile robot 100 can use the semantic map to navigate around its environment, avoid obstacles, and perform one or more actions. In addition, the process 500 can include more or fewer steps than those Figure 5 discussed, as long as these modifications provide the same functions and / or objectives as the Figure 5 process 500.

[0051] In step 502, according to an example, the process 500 includes receiving sensor data from one or more non-audio sensors of the mobile robot 100. For example, other sensors can include image sensors (e.g., cameras), LIDAR sensors (e.g., 3D LIDAR sensors), any relevant sensors (e.g., infrared, radar, etc.), or any combination of them. As a non-limiting example, for instance, the other sensor data at least includes 3D LIDAR data. After receiving the other sensor data from one or more other sensors of the sensor system 608, the process 500 proceeds to step 506.

[0052] In step 504, according to an example, the process 500 includes receiving an audio signal from the microphone array 104. The raw audio signal is transmitted from the microphone array 104 to the audio perception system 200. The audio perception system 200 is configured to receive the raw audio signal. As previously discussed, the audio perception system 200 includes a signal processing system 202, an ML system 204, and a KG 206. After receiving the multi-channel audio signal from the microphone array 104, the process 500 proceeds to step 508.

[0053] In step 506, according to an example, process 500 includes generating simultaneous localization and mapping (SLAM) or point cloud data from other sensor data. In this example, the other sensor data may include LIDAR data from the LIDAR sensor of sensor system 608. As another example, additionally or alternatively, the other sensor data may include image data from an image sensor. Sensor system 608 and / or processing system 606 are configured to process the other sensor data and generate processed sensor data (e.g., SLAM data, point cloud data, etc.). Next, after completing step 506, process 500 proceeds to step 510.

[0054] In step 508, according to an example, process 500 includes performing audio scene recognition and DOA estimation. More specifically, in response to receiving a multi-channel audio signal from microphone array 104, audio perception system 200 generates semantic audio scene data. As described above, audio perception system 200 processes the multi-channel audio signal and generates semantic audio scene data based on the multi-channel audio signal. The semantic audio scene data at least includes audio event data, DOA data, and location data. The semantic audio scene data is useful for providing semantic information about object / landmark detection obtained by other sensors (e.g., LIDAR). For example, the semantic audio scene data can be used to generate labels identifying the objects detected in the point cloud. After generating the semantic audio scene data, process 500 proceeds to step 510.

[0055] In step 510, according to an example, process 500 includes combining the SLAM data and the semantic audio scene data. This combining step may include fusing and optimizing the SLAM data and the semantic audio scene data. In this regard, for example, audio perception system 200 can provide semantic audio scene data for map construction to support SLAM (simultaneous localization and mapping). Audio perception system 200 is configured to reduce any ambiguity of the detected objects, which will require less real-time computing memory / resources than either LIDAR-based or vision-based SLAM, especially on an embedded platform. After fusing and / or optimizing the SLAM data and the semantic audio scene data, process 500 proceeds to step 512.

[0056] In step 512, according to an example, process 500 includes generating or updating a semantic map based on optimization of sensor fusion data (e.g., semantic audio scene data and SLAM data). Generally, a semantic map at least includes semantic labels of detected objects and identified regions associated with specific orientations and locations in the environment. For example, a semantic map may include semantic labels of detected objects or detected sound sources (e.g., microwave oven, TV, washing machine, dryer, dog, toilet, etc.) associated with location data of the environment, and semantic labels of identified regions (e.g., kitchen, laundry room, living room, bathroom, bedroom, corridor, etc.). The semantic map provides higher accuracy and information for the mobile robot 100 by using hybrid data related to location data (e.g., at least LIDAR data and audio scene data). The semantic map may include layout data of the environment. After generating or updating the semantic map, process 500 is configured to proceed to step 510 when the mobile robot 100 receives new sensor data of the environment, after receiving new SLAM / point cloud data and acoustic scene data.

[0057] Figure 6 is a block diagram of an example of a system 600 according to an example embodiment, the system including a mobile robot 100. More specifically, in this example, the system 600 at least includes a mobile robot 100 and a docking station 102. Additionally, the system 600 is configured to include a remote computing system 602 and one or more mobile communication devices 604, the one or more mobile communication devices 604 communicating with the mobile robot 100 and / or the remote computing system 602. For example, the mobile communication device 604 may include a mobile phone (e.g., a smart phone, etc.), a tablet computer, a laptop computer, etc. The mobile communication device 604 may provide alerts, notifications, sensor data (e.g., digital images, audio data, etc.), or any combination of them regarding audio event data and / or audio scene data.

[0058] As Figure 6 shown, the mobile robot 100 at least includes a processing system 606 having at least one processing device. For example, the processing system 606 at least includes an electronic processor, a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a microprocessor, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), any suitable processing technology, or any combination of them. The processing system 606 is operable to provide the functions described herein.

[0059] The mobile robot 100 is configured to include at least one sensor system 608. The sensor system 608 senses the environment and generates sensor data based thereon. The sensor system 608 communicates data with the processing system 606. The sensor system 608 also communicates data directly or indirectly with the memory system 610. The sensor system 608 includes a plurality of sensors. As described above, the sensor system 608 includes a microphone array or a set of microphones. In addition, the sensor system 608 includes motion sensors. The sensor system 608 includes image sensors, light detection and ranging (LIDAR) sensors, or any plurality and combination thereof. In addition, the sensor system 608 may include thermal sensors, ultrasonic sensors, infrared sensors, radar sensors, collision sensors, satellite-based radio navigation sensors (e.g., GPS sensors), any suitable sensors, or any plurality and combination thereof. In this regard, the sensor system 608 includes a set of sensors that enable the mobile robot 100 to sense its environment and use the sensed information to operate effectively in its environment.

[0060] The mobile robot 100 includes a memory system 610 that is operatively connected to the processing system 606. In an example embodiment, the memory system 610 includes at least one non-transitory computer-readable storage medium that is configured to store various data and provide access to various data so that at least the processing system 606 can perform the operations and functions disclosed herein. The memory system 610 includes a single memory device or multiple memory devices. The memory system 610 may include electrical, electronic, magnetic, optical, semiconductor, electromagnetic, or any suitable storage technology operable with the mobile robot 100. For example, the memory system 610 includes random access memory (RAM), read-only memory (ROM), flash memory, disk drives, memory cards, optical storage devices, magnetic storage devices, memory modules, any suitable type of memory device, or any plurality and combination thereof.

[0061] The memory system 610 includes at least a control program 612, an audio perception system 200, and other related data 614, which are stored on the memory system 610, and each of which includes computer-readable data. The computer-readable data may include instructions, code, routines, various related data, any software technology, or any plurality and combination thereof. When executed by the processing system 606, these instructions are configured to perform at least the functions described in this disclosure.

[0062] The control program 612 is configured to directly or indirectly control the mobile robot 100 based on various data (e.g., user commands, sensor data, semantic audio scene data, semantic maps, etc.). The audio perception system 200 is configured to generate semantic audio scene data based on audio signals. Meanwhile, other relevant data 614 provides various data (e.g., operating systems, etc.) that are related to one or more components of the mobile robot 100 and enable the mobile robot 100 to perform the functions discussed herein.

[0063] In addition, the mobile robot 100 includes other functional modules 616. The other functional modules 616 may include a power source (e.g., one or more batteries, etc.) that can be charged by the power source of the docking station 102. The other functional modules 616 may include one or more I / O devices (e.g., display devices, speaker devices, etc.). The one or more I / O devices may provide alerts, notifications, sensor data (e.g., digital images, audio data, etc.) regarding audio event data and / or audio scene data, or any plurality and combination thereof. In addition, the other functional modules 616 may include any relevant hardware, software, or combination thereof that aids or facilitates the functions of the mobile robot 100.

[0064] The mobile robot 100 also includes communication technologies 618 (e.g., wired communication technologies, wireless communication technologies, or a combination thereof) that enable the components of the mobile robot 100 to (i) communicate with each other, (ii) communicate with a remote computing system 602 (e.g., a cloud computing system, a server, etc.), (iii) communicate with one or more mobile communication devices 604, or (iv) communicate with any plurality or combination thereof. The communication technologies 618 may communicate with one or more communication / computer networks.

[0065] In addition, the mobile robot 100 includes an attachment assembly 620. The attachment assembly 620 is configured to perform tasks. For example, in Figure 1 , the mobile robot 100 is a cleaning robot, and the attachment assembly 620 includes a cleaning device. In this example, the attachment assembly 620 includes a plurality of vacuuming components (e.g., a vacuum suction system) that enable the mobile robot 100 to perform one or more actions related to vacuuming various surfaces (e.g., carpets, etc.). The attachment assembly 620 may include a plurality of cleaning components (e.g., brushes, rags, etc.).

[0066] In addition, the mobile robot 100 includes a set of actuators 622. The set of actuators 622 includes one or more actuators that are involved in enabling the mobile robot 100 to perform one or more actions and functions of the mobile robot 100 as described herein. For example, the set of actuators may include one or more actuators that are involved in driving the wheels of the mobile robot 100 such that the mobile robot 100 is configured to move around its environment. The set of actuators may include one or more actuators that are involved in steering the mobile robot 100. The set of actuators may include one or more actuators that are involved in a braking system that stops the movement of the wheels of the mobile robot 100. The set of actuators may include one or more actuators that are involved in controlling or driving the attachment assembly 620. In this regard, the set of actuators may include one or more actuators that are involved in other actions and / or functions of the mobile robot 100.

[0067] As described in the present disclosure, the mobile robot 100 provides several advantages and benefits. For example, the mobile robot 100 includes an audio perception system 200 that advantageously provides the mobile robot 100 with semantic perception of the audio scene in its environment. Using the audio perception system 200, the mobile robot 100 is configured to perform one or more actions using the semantic audio scene data of its environment. The semantic audio scene data provides the mobile robot 100 with context and semantic information about its environment, as well as one or more acoustic activities occurring in its environment. Using the semantic audio scene data, the mobile robot 100 is configured to identify sound sources and their corresponding locations such that the mobile robot 100 can identify and navigate towards or away from certain objects, certain events, certain regions, or any combination thereof. In this regard, the mobile robot 100 can be controlled to maintain a predetermined distance between the mobile robot 100 and a particular object and / or a particular region in its environment. The mobile robot 100 is also configured to advantageously provide audio surveillance of itself and its environment. The mobile robot 100 is also configured to selectively and effectively perform at least one action (e.g., vacuuming a carpet, sending an alert to the mobile communication device 604, etc.) based on the semantic audio scene data.

[0068] In addition, the mobile robot 100 is configured to detect and identify static objects (e.g., refrigerators, washing machines, etc.) and dynamic objects (e.g., barking dogs, crying children, etc.) in its environment, as well as their directions (e.g., DOA) and positions (e.g., kitchen, laundry room, etc.) relative to the mobile robot 100 and / or its environment. Further, using the audio perception system 200, the mobile robot 100 is configured to detect and identify objects that may not be detected by other sensors (e.g., cameras, LIDAR, etc.) due to malfunctions, object occlusions, etc. In this regard, for example, the advantage of the mobile robot 100 is that it is configured to operate effectively in a variety of scenarios, such as when there is insufficient lighting, camera malfunctions, object occlusions, etc.

[0069] In addition, the foregoing description is intended to be illustrative and not restrictive, and is provided in the context of a particular application and its requirements. Those skilled in the art will realize from the foregoing description that the present invention can be implemented in various forms, and various embodiments can be implemented individually or in combination. Thus, although embodiments of the present invention have been described in connection with specific examples of the present invention, the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of the described embodiments, and the true scope of the embodiments and / or methods of the present invention is not limited to the embodiments shown and described, as various modifications will become apparent to those skilled in the art after studying the drawings, the specification, and the appended claims. Additionally or alternatively, components and functions can be separated or combined in a manner different from the various described embodiments, and different terms can be used to describe them. These and other variations, modifications, additions, and improvements may fall within the scope of the present disclosure as defined in the appended claims.

Claims

1. A computer-implemented method for controlling a mobile robot in an environment, the computer-implemented method comprising: receiving an audio signal via a microphone array, the microphone array comprising a set of microphones disposed at least in part on the mobile robot; Extracting audio feature data of acoustic activity from audio signals; generating direction of arrival (DOA) data of acoustic activity based on the audio signal; generating audio event data using the audio feature data via at least one pre-trained machine learning model, the audio event data identifying at least one sound source of the audio feature data; Extracting entity data by querying a knowledge graph using the audio event data, the entity data having a relationship with the audio event data; Generate semantic audio scene data using audio event data, DOA data and entity data; as well as Execute actions of mobile robots based on semantic audio scene data.

2. The computer-implemented method of claim 1 , further comprising: generating layout data of an environment using one or more sensors of the mobile robot, the one or more sensors including an image sensor; as well as Generate a semantic map by combining layout data with semantic audio scene data, The actions include actuating the actuators of the mobile robot based on the semantic map.

3. The computer-implemented method of claim 1 , wherein: Entity data includes location data; and The location data and the audio event data are connected by a relationship of the knowledge graph, which has the largest probability among other relationships in the knowledge graph.

4. The computer-implemented method of claim 1 , further comprising: generating filtered data of acoustic activity by performing noise cancellation on the audio signal to remove self-noise of the mobile robot; extracting speech data from the filtered data; as well as Extracting non-speech data from the filtered data, The audio feature data includes speech data and non-speech data.

5. The computer-implemented method of claim 1 , further comprising: Generate room impulse response data using an audio signal, The audio feature data includes room impulse response data. 6 . The computer-implemented method of claim 1 , wherein the action comprises sending a message to a mobile communication device to provide notification of the semantic audio scene data.

7. The computer-implemented method of claim 1 , wherein: The mobile robot is configured to be coupled to a docking station; The docking station includes a power source for moving the robot; The mobile robot includes a cleaning device; and The set of microphones includes a first subset of microphones disposed on the mobile robot and a second subset of microphones disposed on the docking station.

8. A mobile robot comprising: Microphone array; one or more processors in data communication with the microphone array; as well as One or more memories in data communication with one or more processors, the one or more memories including computer readable data stored thereon, which when executed by the one or more processors, performs a method including the following operations: receiving an audio signal via a microphone array, the microphone array comprising a set of microphones disposed at least in part on the mobile robot; Extracting audio feature data of acoustic activity from audio signals; generating direction of arrival (DOA) data of acoustic activity based on the audio signal; generating audio event data using the audio feature data via at least one pre-trained machine learning model, the audio event data identifying at least one sound source of the audio feature data; Extracting entity data by querying a knowledge graph using the audio event data, the entity data having a relationship with the audio event data; Generate semantic audio scene data using audio event data, DOA data and entity data; as well as Execute actions of mobile robots based on semantic audio scene data.

9. The mobile robot according to claim 8, wherein the method further comprises: generating layout data of the environment using one or more sensors of the mobile robot, the one or more sensors including a light detection and ranging (LIDAR) sensor; as well as Generate a semantic map by combining layout data with semantic audio scene data, The actions include actuating the actuators of the mobile robot based on the semantic map.

10. The mobile robot according to claim 8, wherein: Entity data includes location data; and The location data and the audio event data are connected by a relationship of the knowledge graph, which has the largest probability among other relationships in the knowledge graph.

11. The mobile robot according to claim 8, wherein the method further comprises: generating filtered data of acoustic activity by performing noise cancellation on the audio signal to remove self-noise of the mobile robot; extracting speech data from the filtered data; as well as Extracting non-speech data from the filtered data, The audio feature data includes speech data and non-speech data.

12. The mobile robot according to claim 8, wherein the method further comprises: Generate room impulse response data using an audio signal, The audio feature data includes room impulse response data.

13. The mobile robot of claim 8, wherein the action comprises sending a message to a mobile communication device to provide notification of the semantic audio scene data.

14. The mobile robot according to claim 8, wherein: The mobile robot is configured to be coupled to a docking station; The docking station includes a power source for moving the robot; The mobile robot includes a cleaning device; and The set of microphones includes a first subset of microphones disposed on the mobile robot and a second subset of microphones disposed on the docking station.

15. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform a method for controlling a mobile robot in an environment, the method comprising: receiving an audio signal via a microphone array, the microphone array comprising a set of microphones disposed at least in part on the mobile robot; Extracting audio feature data of acoustic activity from audio signals; generating direction of arrival (DOA) data of acoustic activity based on the audio signal; generating audio event data using the audio feature data via at least one pre-trained machine learning model, the audio event data identifying at least one sound source of the audio feature data; Extracting entity data by querying a knowledge graph using the audio event data, the entity data having a relationship with the audio event data; Generate semantic audio scene data using audio event data, DOA data and entity data; as well as Execute actions of mobile robots based on semantic audio scene data.

16. The one or more non-transitory computer-readable media of claim 15, wherein the method further comprises: generating layout data of the environment using one or more sensors of the mobile robot, the one or more sensors including a light detection and ranging (LIDAR) sensor; as well as Generate a semantic map by combining layout data with semantic audio scene data, The actions include actuating the actuators of the mobile robot based on the semantic map.

17. The one or more non-transitory computer-readable media of claim 15, wherein: Entity data includes location data; and The location data and the audio event data are connected by a relationship of the knowledge graph, which has the largest probability among other relationships in the knowledge graph.

18. The one or more non-transitory computer-readable media of claim 15, further comprising: generating filtered data of acoustic activity by performing noise cancellation on the audio signal to remove self-noise of the mobile robot; extracting speech data from the filtered data; as well as Extracting non-speech data from the filtered data, The audio feature data includes speech data and non-speech data.

19. The one or more non-transitory computer-readable media of claim 15, further comprising: Generate room impulse response data using an audio signal, The audio feature data includes room impulse response data.

20. The one or more non-transitory computer-readable media of claim 15, wherein: The mobile robot is configured to be coupled to a docking station; The docking station includes a power source for moving the robot; The mobile robot includes a cleaning device; and The set of microphones includes a first subset of microphones disposed on the mobile robot and a second subset of microphones disposed on the docking station.