A robot action execution method, electronic device and computer storage medium
By constructing a multi-microphone array and sound field modeling technology, the robot can accurately extract and locate sound source signals in complex environments, solving the problem of insufficient acoustic information processing in existing robot auditory perception systems and achieving more efficient action execution and intelligent interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN WEIAI INTELLIGENT CO LTD
- Filing Date
- 2025-05-13
- Publication Date
- 2026-04-28
AI Technical Summary
Existing robotic auditory perception systems struggle to acquire and process comprehensive acoustic information in complex environments, making it difficult to perform corresponding actions.
Multiple microphone arrays with different characteristics are constructed to collect environmental sound signals. The target sound source signals are extracted through signal processing and sound field modeling techniques, and the robot is controlled to perform actions based on the sound field model.
It improves the robot's perception and response capabilities in complex environments, enables dynamic interaction between the robot and the surrounding acoustic environment, and enhances its level of intelligence.
Smart Images

Figure CN120244987B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot perception technology, and in particular to a robot motion execution method, electronic device, and computer storage medium. Background Technology
[0002] In the current development of robotics technology, auditory perception is one of the key components, used to acquire sound information input from the environment.
[0003] Current robots' auditory perception is typically limited to unidirectional sound source localization or simple speech recognition tasks, lacking comprehensive perception capabilities for complex sound fields in three-dimensional space. This makes it difficult for existing robots to perform corresponding actions based on sound information. In real-world applications such as smart homes, smart security, and industrial production, sound information is rich and complex. Relying solely on traditional auditory perception, robots cannot effectively acquire and process omnidirectional acoustic information, making it difficult to meet practical needs.
[0004] The above description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention
[0005] The purpose of this application is to provide a robot action execution method, electronic device, and computer storage medium, which aims to solve the problem that robots equipped with existing auditory perception have difficulty performing corresponding actions based on sound information, and improve the robot's perception and response capabilities in complex environments.
[0006] To achieve the above objectives:
[0007] In a first aspect, embodiments of this application provide a method for executing robot actions, the method comprising:
[0008] Based on the robot's body characteristics, construct multiple sets of multi-microphone arrays with different properties;
[0009] Control the multi-microphone array to collect sound signals from different directions in the environment;
[0010] The sound signal is processed to obtain multiple sound source signals, and the target sound source signal is extracted from the multiple sound source signals;
[0011] The target sound source signal is located, and a sound field model is constructed based on the located target sound source signal;
[0012] The robot is controlled to perform corresponding actions based on the sound field model.
[0013] In one embodiment, the sound signals collected in the environment include:
[0014] The robot dynamically collects sound signals from the environment based on its motion posture.
[0015] In one embodiment, before processing the sound signal to obtain multiple sound source signals, the method further includes:
[0016] The audio signal is subjected to noise reduction, frame segmentation, and frequency domain transformation processing in sequence.
[0017] In one embodiment, processing the sound signal to obtain multiple sound source signals includes:
[0018] The speech, environmental noise, and mechanical noise in the sound signal are separated according to a preset multi-microphone array algorithm to obtain multiple sound source signals with directional characteristics.
[0019] In one embodiment, extracting the target sound source signal from the plurality of sound source signals includes:
[0020] The sound source signal is matched with a database trained based on samples of speech, environmental noise and mechanical self-noise, and the speech in the sound source signal that meets the preset target features is extracted to obtain the target sound source signal.
[0021] In one embodiment, after extracting the target sound source signal from the plurality of sound source signals, the method further includes:
[0022] The beam is dynamically adjusted according to the direction of the target sound source signal to suppress interference signals in other directions;
[0023] The reverberation caused by environmental reflections in the target sound source signal is removed.
[0024] In one embodiment, constructing a sound field model based on the target sound source signal includes:
[0025] Acquire the sound source parameters of the target sound source signal, wherein the sound source parameters include at least one of the following: the location, intensity, frequency distribution, phase information, radiation pattern, time-varying characteristics, type, phase coherence, and nonlinear characteristics of the target sound source signal;
[0026] The sound field model is constructed based on the sound source parameters.
[0027] In one embodiment, after constructing the sound field model based on the target sound source signal, the method further includes:
[0028] Real-time acquisition of the robot's motion state data;
[0029] The motion change trend of the robot is obtained based on the motion state data;
[0030] The sound field model is dynamically updated based on the trend of motion change.
[0031] Secondly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of robot action execution as described above.
[0032] Thirdly, embodiments of this application provide a computer storage medium storing a computer program, which, when executed by the processor, implements the steps of robot action execution as described above.
[0033] This application acquires and processes sound signals from the robot's surrounding environment to obtain multiple sound source signals, extracts the target sound source signal, locates the target sound source signal, constructs a sound field model based on it, and finally controls the robot to perform corresponding actions according to the sound field model. In this way, dynamic interaction between the robot and the surrounding acoustic environment is realized, which significantly improves the robot's perception, response capabilities and intelligence level in complex environments. Attached Figure Description
[0034] Figure 1 This is a flowchart illustrating a robot action execution method provided in an embodiment of the present invention.
[0035] Figure 2 This is a schematic diagram of a robot motion execution system and a robot equipped with the system, provided as an embodiment of the present invention.
[0036] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.
[0037] Explanation of reference numerals in the attached figures:
[0038] 10. Terminal; 20. Multi-microphone array module; 21. Scalable linear array; 30. Signal processing module; 40. Sound field modeling module; 50. Interactive response module; 60. Display module.
[0039] 310. Processor; 311. Memory; 312. Network interface; 313. Bus system. Detailed Implementation
[0040] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0041] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.
[0042] It should be understood that although the terms first, second, third, etc., may be used herein to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this document, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if," as used herein, can be interpreted as "when," "when," or "in response to determination." Furthermore, as used herein, the singular forms "a," "an," and "the" are intended to also include the plural forms unless the context indicates otherwise. It should be further understood that the terms "comprising," "including," indicate the presence of the stated feature, step, operation, element, component, item, kind, and / or group, but do not exclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, kinds, and / or groups. The terms "or" and "and / or" as used herein are to be interpreted as inclusive, or mean any one or any combination thereof. Therefore, "A, B, or C" or "A, B, and / or C" means "any one of the following: A; B; C; A and B; A and C; B and C; A, B, and C". Exceptions to this definition will only occur if the combination of elements, functions, steps, or operations is inherently mutually exclusive in some way.
[0043] It should be understood that although the steps in the flowcharts of this application's embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.
[0044] It should be noted that step designations such as S1 and S2 are used in this document for the purpose of more clearly and concisely describing the corresponding content, and do not constitute a substantial limitation on the order. In specific implementation, those skilled in the art may execute S2 first and then S1, etc., but these should all be within the protection scope of this application.
[0045] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0046] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustration and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.
[0047] The following is a brief explanation of the terms that may be used in this application:
[0048] Blind source separation (BSS) is a technique for reconstructing the original signal source by observing the signal without knowing the source or transmission channel. It is of great significance in signal processing and is widely used in areas such as speech processing, image processing, and biomedical signal processing.
[0049] Adaptive beamforming is an advanced signal processing technique widely used in radar, sonar, wireless communication, and voice enhancement. It enhances the signal in the direction of the target while suppressing interference signals from other directions by dynamically adjusting the direction and shape of the beam.
[0050] De-reverberation is a technique used to reduce or eliminate the impact of reverberation on audio signals. Reverberation is a complex echo effect formed when sound reflects and attenuates multiple times in an enclosed space. It affects the clarity and intelligibility of audio signals, especially in scenarios such as voice communication, speech recognition, audio recording, and sound field reconstruction. De-reverberation techniques analyze and process audio signals to remove reverberation components, thereby restoring a clearer, dry sound signal.
[0051] The wave superposition method is a method based on the fundamental principles of wave theory used to analyze and calculate the composite wave field when multiple wave sources interact in space. It is widely applied in acoustics, optics, electromagnetics, seismology, and other fields to study phenomena such as wave interference, diffraction, and scattering.
[0052] The equivalent source method is a numerical method for analyzing acoustic and wave problems, primarily used to solve complex sound field problems such as sound source localization, sound field reconstruction, and sound radiation analysis. It approximates the distribution of the actual sound field by introducing a set of virtual equivalent sources within the boundary or region of a known sound field. This method combines the concepts of the boundary element method (BEM) and point source superposition, offering both efficiency and flexibility.
[0053] The acoustic inverse problem refers to the problem of inversely deducing the characteristics of the sound source (such as the location, intensity, and frequency distribution of the sound source) or the propagation path of the sound field from known sound field observation data (such as sound pressure and sound intensity).
[0054] First Embodiment
[0055] Please refer to Figure 1 The diagram shows a flowchart of a robot action execution method, which includes:
[0056] S1. Based on the robot's body characteristics, construct multiple sets of multi-microphone arrays with different characteristics. Specifically, in this embodiment, regarding the deployment strategy, to achieve omnidirectional sound acquisition, the robot's external body characteristics can be scanned first through a terminal. Then, based on the robot's size and target frequency band, an optimized design is performed to determine the appropriate array spacing. Multiple miniature microphone arrays are adaptively deployed on the robot's body surface and movable parts (such as the robotic arm) to form a multi-layered three-dimensional acquisition network. In terms of array configuration, a multi-dimensional microphone array is adopted, covering both horizontal and vertical directions, which can capture sound signals from various angles, thereby improving the accuracy and comprehensiveness of sound acquisition.
[0057] S2. Control the multi-microphone array to collect sound signals from different directions in the environment. Specifically, in this embodiment, after the terminal constructs the microphone array distribution pattern adapted to the robot in step S1, the robot is manually assembled with the multi-microphone array according to the construction result. In actual implementation, the multi-microphone array module configured inside the robot body controls the multi-microphone array to collect sound signals from different directions in the environment in real time.
[0058] S3. Process the sound signal to obtain multiple sound source signals, and extract the target sound source signal from the multiple sound source signals. Specifically, in this embodiment, in the front-end processing stage, the signal processing module configured inside the robot body comprehensively uses blind source separation technology to process the mixed sound signals collected in step S2, thereby extracting multiple independent sound source signals. In the back-end processing stage, the signal processing module uses a trained deep learning model to accurately extract the target sound source signal from the multiple sound source signals.
[0059] S4. Locate the target sound source signal and construct a sound field model based on the located target sound source signal. Specifically, in this embodiment, during the sound source localization stage, the signal processing module combines time delay estimation (TDOA) with a deep learning model. Through time delay estimation, it accurately calculates the time difference between the arrival times of the target sound source signal at different microphones. Then, utilizing the powerful feature extraction and analysis capabilities of the deep learning model, it outputs the three-dimensional coordinates and confidence level of the target sound source signal, achieving accurate sound source localization. In the acoustic reconstruction algorithm, the sound field modeling module configured inside the robot body selects either the wave superposition method or the equivalent source method. Both algorithms can accurately construct a sound field model based on the target sound source signal.
[0060] S5. Control the robot to perform corresponding actions based on the sound field model. Specifically, in this embodiment, based on the sound field characteristics reconstructed in step S4, the interaction response module configured inside the robot body controls the robot to perform corresponding actions. For example, in a voice interaction scenario, the robot accurately responds and interacts based on the voice content and speaker's location identified by the sound field model; in a navigation and obstacle avoidance scenario, the robot judges the location and danger level of obstacles based on the sound field model, plans a reasonable movement path, avoids noise sources or dangerous areas, and achieves safe and efficient movement.
[0061] In this way, by executing methods S1 to S5, the robot achieves dynamic interaction with the surrounding acoustic environment, which significantly improves the robot's perception and response capabilities in complex environments and enhances the robot's intelligence level.
[0062] Optionally, controlling the multi-microphone array to collect sound signals from different directions in the environment in step S2 includes: dynamically collecting sound signals in the environment according to the robot's motion posture. Specifically, in this embodiment, the step of dynamically collecting sound signals in the environment according to the robot's motion posture is achieved through the collaborative work of a scalable linear array and a multi-dimensional microphone array. The multi-microphone array includes a scalable linear array, which can adaptively adjust according to the robot's motion posture to ensure efficient sound collection under different motion states. In terms of deployment strategy, users can optimize the design based on the robot's size and target frequency band to determine a suitable array spacing, thereby improving the accuracy and comprehensiveness of sound collection.
[0063] Optionally, before processing the sound signal to obtain multiple sound source signals as described in step S3, the method further includes: sequentially performing noise reduction, framing, and frequency domain transformation processing on the sound signal. Specifically, in this embodiment, before processing the sound signal, preprocessing of the acquired sound signal is required. The acquired sound signal is sequentially subjected to noise reduction, framing, and frequency domain transformation (FFT) processing. Noise reduction improves the sound signal quality by removing environmental noise and noise generated by the robot's own hardware. Framing divides the continuous sound signal into fixed-length frames, facilitating analysis and processing in subsequent steps. Frequency domain transformation converts the time-domain signal into a frequency-domain signal, extracts the frequency features of the sound signal, and uses the differences in frequency features to separate different sound sources, providing richer information for sound source separation and localization in subsequent steps.
[0064] Optionally, the processing of the sound signal to obtain multiple sound source signals in step S3 includes: separating speech, environmental noise, and mechanical self-noise in the sound signal according to a preset multi-microphone array algorithm to obtain multiple sound source signals with directional features. Specifically, in this embodiment, firstly, a deep neural network (such as a fully connected network, convolutional neural network, or recurrent neural network) is used to learn the mapping relationship between reverberation signals and dry sound signals. Secondly, the powerful pattern recognition capability of the trained neural network model is used to accurately distinguish speech, environmental noise, and mechanical self-noise. Finally, combined with a geometric algorithm, the direction and position of the sound source are accurately calculated based on information such as the time difference of sound arriving at different microphones, achieving effective separation and precise localization of multiple sound sources in complex environments. The preset multi-microphone array algorithm is preferably a trained neural network model, and the trained neural network model is preferably a convolutional neural network (CNN).
[0065] Optionally, extracting the target sound source signal from multiple sound source signals in step S3 includes: matching the sound source signals with a database trained based on speech, environmental noise, and mechanical self-noise samples, and extracting speech from the sound source signals that conforms to preset target features to obtain the target sound source signal. Specifically, in this embodiment, since convolutional neural networks can extract local features of signals, they are suitable for frequency domain signal processing. Specifically, the convolutional neural network is trained with a large number of labeled speech, environmental noise, and mechanical self-noise samples to learn the features of different sound sources and generate a trained database. In actual implementation, the target sound source signal is obtained by matching the sound source signals with the database and extracting the successfully matched speech.
[0066] In one scenario, when multiple people are speaking simultaneously, the robot can identify the voice features and voiceprint information of different speakers, and determine the sound source signal that matches the preset target features (such as voiceprint matching with the robot's preset interaction object, voice command keyword matching, etc.) as the target sound source signal, while other mismatched sound source signals are determined as interference signals.
[0067] Optionally, after extracting the target sound source signal from multiple sound source signals as described in step S3, the method further includes: dynamically adjusting the beam according to the direction of the target sound source signal to suppress interference signals in other directions; and removing reverberation caused by environmental reflection in the target sound source signal. Specifically, in this embodiment, after extracting the target sound source signal, adaptive beamforming technology is first used to dynamically adjust the beam according to the direction of the target sound source signal to enhance the signal strength of the target sound source signal while suppressing interference signals in other directions. Then, dereverberation technology is used to remove the reverberation caused by environmental reflection in the target sound source signal, restoring a clear and clean sound signal.
[0068] Optionally, the construction of the sound field model based on the located target sound source signal in step S4 includes: acquiring the sound source parameters of the target sound source signal, wherein the sound source parameters include at least one of the following: position, intensity, frequency distribution, phase information, radiation pattern, time-varying characteristics, type, phase coherence, and nonlinear characteristics of the target sound source signal; and constructing the sound field model based on the sound source parameters. Specifically, in one embodiment, either the wave superposition method or the equivalent source method is used to reconstruct the sound pressure distribution or particle velocity distribution of the robot's surrounding environment space, i.e., the sound field model, based on the sound source parameters of the target sound source signal. In another embodiment, the sound field model of the robot's surrounding environment space can also be reconstructed based on the principle of solving the inverse acoustic problem. By reconstructing the sound field model, the layout of the microphone array can be optimized, the effect of sound source localization and speech enhancement can be improved, and more accurate acoustic environment information can be provided to the robot.
[0069] Optionally, after constructing the sound field model based on the located target sound source signal in step S4, the method further includes:
[0070] S41. Real-time acquisition of robot motion state data. Specifically, in this embodiment, the robot relies on its own sensors to collect motion state data. The inertial measurement unit (IMU) acquires the robot's attitude information in real time, such as roll angle, pitch angle, and yaw angle, accurately reflecting the robot's posture changes. The position encoder is used to record the robot's position data, including coordinate information in space. In indoor robot movement scenarios, the IMU can monitor the robot's turning, bending, and other actions, while the position encoder can determine its specific position in space. This data is used to provide a basis for subsequent adjustments.
[0071] S42. Obtain the robot's motion change trend based on motion state data. Specifically, in this embodiment, the signal processing module receives and analyzes the collected motion state data. By comparing the posture and position data at different times, the robot's motion trend and magnitude of change are determined. When the robot moves rapidly or its posture changes significantly, it means that its surrounding sound field environment may have changed significantly, requiring timely adjustment of the sound field model. If the motion state change is small, the adjustment frequency is appropriately reduced to balance computational resources and model accuracy.
[0072] S43. Dynamically update the sound field model based on the trend of motion changes. Specifically, in this embodiment, the sound field modeling module initiates a dynamic update mechanism based on the analysis results of motion state changes. If the wave superposition method or equivalent source method is used for sound field modeling, the parameters of each point source or equivalent source in the model will be recalculated according to the robot's position changes. When the robot approaches the sound source, the relevant parameters are adjusted to reflect the influence of distance changes on sound pressure and particle velocity distribution. For posture changes, the direction and angle of sound acquisition by the microphone array are redefined, thereby updating the sound field model to ensure that the model is synchronized with the actual sound field in real time, providing the robot with accurate acoustic environment information.
[0073] Optionally, after constructing the sound field model based on the located target sound source signal in step S4, to intuitively display the sound field information, visualization technology is used through the display module on the robot. For example, drawing acoustic force maps and generating sound pressure isosurfaces are displayed on the display module, allowing operators to clearly understand the distribution of the sound field. Thus, during the initial manual training of the robot, sound field visualization allows operators to clearly obtain the sound field distribution, facilitating adjustments to the training plan and subsequent scene training based on the training results.
[0074] Second Embodiment
[0075] Please refer to Figure 2 , Figure 2 A schematic diagram of a robot motion execution system and a robot equipped with the system is shown. The system includes:
[0076] Terminal 10 is used to construct multiple sets of multi-microphone arrays with different characteristics based on the robot's body features.
[0077] The multi-microphone array module 20 is used to control the multi-microphone array to collect sound signals from different directions in the environment.
[0078] The signal processing module 30 is electrically connected to the multi-microphone array module 20. It is used to process the sound signal to obtain multiple sound source signals, extract the target sound source signal from the multiple sound source signals, and locate the target sound source signal.
[0079] The sound field modeling module 40 is electrically connected to the signal processing module 30 and is used to construct a sound field model based on the located target sound source signal.
[0080] The interactive response module 50 is electrically connected to the sound field modeling module 40 and is used to control the robot to perform corresponding actions based on the sound field model.
[0081] The terminal 10 can be a computer, laptop, tablet, smartphone, smart bracelet, 3D scanner, etc., while the multi-microphone array module 20, signal processing module 30, sound field modeling module 40 and interactive response module 50 are configured inside the robot body.
[0082] In terms of microphone array deployment, existing technologies mostly employ planar or linear layouts. This traditional layout limits the viewing angle of the microphone array, resulting in significant limitations in the acoustic data collected and making it difficult to accurately reconstruct complex three-dimensional sound fields. For example, in indoor environments, sound undergoes multiple reflections and refractions, making it difficult for planar or linear arrays to capture changes in sound signals from all directions, leading to substantial errors in sound field reconstruction.
[0083] Optionally, the multi-microphone array module 20 includes a retractable linear array 21. The retractable linear array 21 can be deployed on the robot's body surface and movable parts, covering width, depth, and height. When the robot moves, the robotic arm's movements change its position and posture. The retractable linear array 21 can then adjust its length and angle accordingly to the robotic arm's movement, capturing sound signals from various angles thanks to its omnidirectional coverage. In scenarios where security robots patrol, even when the robot turns, its posture changes, allowing the retractable linear array 21 to still collect sound signals from different directions without blind spots caused by movement, ensuring comprehensive and accurate sound signal acquisition. In scenarios where service robots provide services to customers, when the robotic arm extends to retrieve items, the retractable linear array 21 can extend, expanding the sound acquisition range and adjusting its angle to better capture sound signals from the surrounding environment. When the robotic arm retracts towards the robot body, the retractable linear array 21 also retracts, avoiding unnecessary interference due to excessive length and ensuring efficient sound acquisition in different movement states.
[0084] Furthermore, the signal processing module 30 is also used to preprocess the acquired sound signals. The acquired sound signals are sequentially subjected to noise reduction, framing, and frequency domain transform (FFT) processing. Noise reduction improves the sound signal quality by removing environmental noise and noise generated by the robot's own hardware. Framing divides the continuous sound signal into fixed-length frames, facilitating subsequent analysis and processing. Frequency domain transform converts the time-domain signal into a frequency-domain signal, extracts the frequency features of the sound signal, and uses the differences in frequency features to separate different sound sources, providing richer information for sound source separation and localization in subsequent steps.
[0085] Understandably, the interaction response module 50 is responsible for translating the sound field model into the robot's behavioral strategy. When a specific sound source is detected, such as a user's call or a danger signal, the interaction response module 50 controls the robot to turn towards the sound source. If a sudden noise, such as a collision sound or an alarm sound, is encountered, the interaction response module 50 guides the robot to avoid it in time. In this way, intelligent interaction between the robot and its surrounding acoustic environment is achieved, enabling it to make reasonable action decisions based on the reconstructed sound field model.
[0086] Optionally, the robot motion execution system also includes a display module 60, which is electrically connected to the sound field modeling module 40. The display module 60 displays image information such as plotted acoustic force maps and generated sound pressure isosurfaces, allowing operators to clearly understand the sound field distribution. Thus, during the initial manual training of the robot, sound field visualization enables operators to clearly obtain the sound field distribution, facilitating adjustments to the training plan and subsequent scene training based on the training results.
[0087] In one scenario, taking a security robot as an example, multiple six-microphone circular arrays are evenly installed on the robot's head and body, while multiple four-microphone linear arrays are installed on the sides and arms. A GPU-accelerated equivalent source method is used for sound field modeling, with the model updated every 100ms. When an abnormal sound source, such as the sound of breaking glass or unusual footsteps, is detected within the monitoring area, the system accurately locates the sound source and assesses its hazard level based on the dynamically updated sound field model. Once a dangerous situation is identified, an alarm mechanism is immediately triggered, and an alarm notification containing the sound source's location information is sent to the security personnel's terminal devices, enabling real-time security monitoring of the monitored area.
[0088] In another scenario, taking a service robot as an example, a multi-dimensional microphone array is installed on the robot's head and body, and a retractable linear array 21 is configured on the movable robotic arm. When a guest calls the robot in the hotel lobby, the multi-microphone array quickly collects sound signals. After signal preprocessing, sound source separation and localization, the system quickly calculates the guest's three-dimensional position coordinates. The sound field modeling module 40 reconstructs the sound field in real time, providing the robot with more accurate acoustic environment information. The interaction response module 50, based on the positioning information and sound field characteristics, controls the robot to turn towards the guest, while simultaneously controlling the robotic arm to make a friendly welcoming gesture. During interaction with the guest, the robot utilizes precise sound source localization and noise suppression functions to clearly recognize the guest's voice commands, providing guidance, consultation, and other services, achieving an efficient and natural human-computer interaction experience.
[0089] Based on the same inventive concept as the foregoing embodiments, this invention provides an electronic device, such as... Figure 3As shown, the device includes: a processor 310 and a memory 311 storing a computer program; wherein, Figure 3 The processor 310 shown in the diagram does not indicate that there is only one processor 310, but only indicates the positional relationship of the processor 310 relative to other devices. In practical applications, there can be one or more processors 310; similarly, Figure 3 The memory 311 illustrated herein has the same meaning, that is, it is only used to indicate the positional relationship of memory 311 relative to other devices. In practical applications, there can be one or more memories 311. When the processor 310 runs the computer program, the method applied to the above-mentioned device is implemented.
[0090] The device may also include at least one network interface 312. The various components of the device are coupled together via a bus system 313. It is understood that the bus system 313 is used to implement communication between these components. In addition to a data bus, the bus system 313 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 3 The general designated all buses as Bus System 313.
[0091] The memory 311 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memory 311 described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0092] The memory 311 in this embodiment of the invention is used to store various types of data to support the operation of the device. Examples of this data include: any computer programs used to operate on the device, such as operating systems and applications; contact data; phonebook data; messages; pictures; videos, etc. The operating system includes various system programs, such as the framework layer, core library layer, driver layer, etc., used to implement various basic services and handle hardware-based tasks. Applications can include various applications, such as media players, browsers, etc., used to implement various application services. Here, the program implementing the method of this embodiment of the invention can be included in the application.
[0093] Based on the same inventive concept as the foregoing embodiments, this embodiment also provides a computer-readable storage medium storing a computer program. The computer-readable storage medium can be a magnetic random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM), etc.; it can also be various devices including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc. When the computer program stored in the computer-readable storage medium is run by a processor, it implements the above method. For the specific steps implemented when the computer program is executed by the processor, please refer to [link to relevant documentation]. Figure 1 The description of the illustrated embodiments will not be repeated here.
[0094] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0095] In this document, the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, which includes not only the elements listed but also other elements not expressly listed.
[0096] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for executing robot actions, characterized in that, The method includes: Based on the robot's body characteristics, construct multiple sets of multi-microphone arrays with different properties; Control the multi-microphone array to collect sound signals from different directions in the environment; The sound signal is processed to obtain multiple sound source signals, and the target sound source signal is extracted from the multiple sound source signals; The target sound source signal is located, and a sound field model is constructed based on the located target sound source signal; The robot is controlled to perform corresponding actions based on the sound field model. The control of the multi-microphone array to collect sound signals from different directions in the environment includes: The robot dynamically acquires sound signals from the environment based on its motion posture; The step of constructing a sound field model based on the located target sound source signal then includes: Real-time acquisition of the robot's motion state data; The motion change trend of the robot is obtained based on the motion state data; The sound field model is dynamically updated based on the trend of motion change.
2. The method according to claim 1, characterized in that, Before processing the sound signal to obtain multiple sound source signals, the method further includes: The audio signal is subjected to noise reduction, frame segmentation, and frequency domain transformation processing in sequence.
3. The method according to claim 1, characterized in that, The process of processing the sound signal to obtain multiple sound source signals includes: The speech, environmental noise, and mechanical noise in the sound signal are separated according to a preset multi-microphone array algorithm to obtain multiple sound source signals with directional characteristics.
4. The method according to claim 3, characterized in that, Extracting the target sound source signal from the plurality of sound source signals includes: The sound source signal is matched with a database trained based on samples of speech, environmental noise and mechanical self-noise, and the speech in the sound source signal that meets the preset target features is extracted to obtain the target sound source signal.
5. The method according to claim 1, characterized in that, After extracting the target sound source signal from the plurality of sound source signals, the method further includes: The beam is dynamically adjusted according to the direction of the target sound source signal to suppress interference signals in other directions; The reverberation caused by environmental reflections in the target sound source signal is removed.
6. The method according to claim 1, characterized in that, The step of constructing a sound field model based on the target sound source signal includes: Acquire the sound source parameters of the target sound source signal, wherein the sound source parameters include at least one of the following: the location, intensity, frequency distribution, phase information, radiation pattern, time-varying characteristics, type, phase coherence, and nonlinear characteristics of the target sound source signal; The sound field model is constructed based on the sound source parameters.
7. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the robot action execution method as described in any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the robot action execution method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Sound source positioning method and device
CN117289208A
Substation inspection method and system based on quadruped robot
CN119567253A
Robot and control method thereof
US20230201396A1