Robot action execution method, electronic equipment and computer storage medium
By constructing multi-microphone array and sound field models, the robot can accurately locate the target sound source and perform corresponding actions in complex environments, solving the problem of insufficient acoustic information processing in the existing robot auditory perception system and improving perception and response capabilities.
Patent Information
- Application Number
- CN202510614434.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-13
AI Technical Summary
Existing robot auditory perception systems are difficult to fully acquire and process all-round acoustic information in complex environments, making it difficult to perform corresponding actions based on sound information.
Build multiple sets of multi-microphone arrays with different characteristics, collect ambient sound signals, and reconstruct the sound field model through signal processing and sound field model, accurately locate the target sound source signal, and control the robot to perform actions according to the sound field model.
It realizes dynamic interaction between the robot and the surrounding acoustic environment, improves the perception and response capabilities in complex environments, and enhances the degree of intelligence.
Smart Images

Figure CN120244987A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot perception, and in particular, to a robot action execution method, an electronic device, and a computer storage medium. Background Art
[0002] In the current development process of robot technology, auditory perception, as one of the key links, is used to obtain the input sound information in the environment.
[0003] The auditory perception of existing robots usually can only complete the sound source localization in a single direction, or is limited to simple speech recognition tasks, lacking the comprehensive perception ability for complex sound fields in three-dimensional space, resulting in the difficulty for existing robots to execute corresponding actions according to sound information. In actual application scenarios, such as smart home, intelligent security, industrial production and other environments, the sound information is rich and complex. Relying only on traditional auditory perception, robots cannot effectively obtain and process omnidirectional acoustic information, making it difficult to meet actual needs.
[0004] The above description is to provide general background information and does not necessarily constitute prior art. Summary of the Invention
[0005] The purpose of this application is to provide a robot action execution method, an electronic device, and a computer storage medium, aiming to solve the problem that robots equipped with existing auditory perception are difficult to execute corresponding actions according to sound information, and improve the perception and response ability of robots in complex environments.
[0006] To achieve the above purpose: In a first aspect, an embodiment of this application provides a robot action execution method, and the method includes: Construct multiple groups of multi-microphone arrays with different characteristics according to the body characteristics of the robot; Control the multi-microphone arrays to collect sound signals from different directions in the environment; Process the sound signals to obtain multiple sound source signals, and extract a target sound source signal from the multiple sound source signals; Locate the target sound source signal, and construct a sound field model according to the located target sound source signal; Control the robot to execute corresponding actions according to the sound field model.
[0007] In an embodiment, the collecting of the sound signals in the environment includes: Dynamically collect the sound signals in the environment according to the motion posture of the robot.
[0008] In an embodiment, before the processing of the sound signals to obtain multiple sound source signals, the method further includes: Perform noise reduction, framing, and frequency domain transformation on the sound signal in sequence.
[0009] In one embodiment, the processing of the sound signal to obtain multiple sound source signals includes: Separate the speech, environmental noise, and mechanical self-noise in the sound signal according to a preset set of microphone array algorithms to obtain multiple sound source signals with direction characteristics.
[0010] In one embodiment, the extraction of the target sound source signal from multiple sound source signals includes: Match the sound source signal with a database trained based on speech, environmental noise, and mechanical self-noise samples, and extract the speech in the sound source signal that conforms to the preset target characteristics to obtain the target sound source signal.
[0011] In one embodiment, after extracting the target sound source signal from multiple sound source signals, the method further includes: Dynamically adjust the beam according to the direction of the target sound source signal to suppress interference signals in other directions; Remove the reverberation formed by environmental reflection in the target sound source signal.
[0012] In one embodiment, the construction of the sound field model according to the target sound source signal includes: Obtain the sound source parameters of the target sound source signal, where the sound source parameters at least include at least one of the position, intensity, frequency distribution, phase information, radiation pattern, time-varying characteristics, type, phase coherence, and non-linear characteristics of the target sound source signal; Construct the sound field model according to the sound source parameters.
[0013] In one embodiment, after constructing the sound field model according to the target sound source signal, the method further includes: Real-time collect the motion state data of the robot; Obtain the motion change trend of the robot according to the motion state data; Dynamically update the sound field model according to the motion change trend.
[0014] In a second aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the robot action execution described above are implemented.
[0015] In a third aspect, an embodiment of the present application provides a computer storage medium storing a computer program, which when executed by a processor implements the steps of the robot action execution as described above.
[0016] In the present application, sound signals in the surrounding environment of the robot are collected and processed to obtain a plurality of sound source signals, a target sound source signal is extracted therefrom, the target sound source signal is located, and an acoustic field model is constructed based on it. Finally, the robot is controlled to execute corresponding actions according to the acoustic field model. In this way, the dynamic interaction between the robot and the surrounding acoustic environment is realized, and the perception, response ability and intelligence level of the robot in a complex environment are significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic flowchart of a method for robot action execution provided by an embodiment of the present invention.
[0018] Figure 2 It is a schematic structural diagram of a robot action execution system provided by an embodiment of the present invention and a robot equipped with the system.
[0019] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention.
[0020] DESCRIPTION OF REFERENCE NUMERALS: 10, terminal; 20, multi-microphone array module; 21, retractable linear array; 30, signal processing module; 40, acoustic field modeling module; 50, interaction response module; 60, display module.
[0021] 310, processor; 311, memory; 312, network interface; 313, bus system. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0023] It should be noted that in this document, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising such element. In addition, components, features, and elements with the same name in different embodiments of this application may have the same meaning or may have different meanings, and their specific meanings need to be determined based on their explanations in the specific embodiments or further in combination with the context of the specific embodiments.
[0024] It should be understood that although the terms first, second, third, etc. may be used herein to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this document, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining". Furthermore, as used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context indicates otherwise. It should be further understood that the terms "comprising", "including" indicate the presence of the stated features, steps, operations, elements, components, items, kinds, and / or groups, but do not exclude the presence, occurrence or addition of one or more other features, steps, operations, elements, components, items, kinds, and / or groups. The terms "or" and "and / or" used herein are interpreted as inclusive, or meaning any one or any combination. Thus, "A, B or C" or "A, B and / or C" means "any one of the following: A; B; C; A and B; A and C; B and C; A, B and C". An exception to this definition only occurs when the combination of elements, functions, steps or operations are mutually exclusive in some way.
[0025] It should be understood that although the steps in the flowchart in the embodiments of the present application are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit and can be executed in other orders. Moreover, at least a part of the steps in the figure may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0026] It should be noted that in this article, step codes such as S1 and S2 are used. The purpose is to more clearly and briefly express the corresponding content and do not constitute a substantial limitation in order. Those skilled in the art may execute S2 first and then S1 during specific implementation, etc., but these should all be within the protection scope of the present application.
[0027] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0028] In the subsequent description, suffixes such as "module", "component" or "unit" used to represent elements are only for the convenience of explaining the present application and have no specific meaning in themselves. Therefore, "module", "component" or "unit" can be used interchangeably.
[0029] First, the nouns that the present application may involve will be explained as follows: Blind Source Separation (BSS) is a technology that can recover the original signal source only by observing signals without knowing the signal source and the transmission channel. It is of great significance in the field of signal processing and is widely used in fields such as speech processing, image processing, and biomedical signal processing.
[0030] Adaptive beamforming technology (Beamforming) is an advanced signal processing technology widely used in fields such as radar, sonar, wireless communication, and speech enhancement. It dynamically adjusts the direction and shape of the beam to enhance the signal in the target direction while suppressing interference signals in other directions.
[0031] Dereverberation is a technology used to reduce or eliminate the impact of reverberation on audio signals. Reverberation is a complex echo effect formed by multiple reflections and attenuations of sound in an enclosed space. It affects the clarity and intelligibility of audio signals, especially in scenarios such as voice communication, speech recognition, audio recording, and sound field reconstruction. Dereverberation technology analyzes and processes audio signals to remove the reverberation component, thereby restoring a clearer dry sound signal.
[0032] The Wave Superposition Method is a method based on the basic principles of wave theory, used to analyze and calculate the synthetic wave field when multiple wave sources interact in space. It is widely applied in multiple fields such as acoustics, optics, electromagnetics, and seismology, for studying phenomena such as wave interference, diffraction, and scattering.
[0033] The Equivalent Source Method is a numerical method used for the analysis of acoustic and wave problems, mainly for solving complex sound field problems such as sound source localization, sound field reconstruction, and sound radiation analysis. It approximates the distribution of the actual sound field by introducing a set of virtual equivalent sources within the boundary or region of the known sound field. This method combines the ideas of the Boundary Element Method (BEM) and point source superposition, and has high efficiency and flexibility.
[0034] The acoustic inverse problem refers to the problem of inferring the characteristics of the sound source (such as the sound source position, intensity, frequency distribution, etc.) or the sound field propagation path, etc., starting from the known sound field observation data (such as sound pressure, sound intensity, etc.).
[0035] The first embodiment Please refer to Figure 1 The flowchart of a robot action execution method shown, the method includes: S1. Construct multiple microphone arrays with different characteristics according to the body characteristics of the robot. Specifically, in this embodiment, in terms of the deployment strategy, to achieve omnidirectional sound collection, the terminal can first scan the external body characteristics of the robot, and then optimize the design according to the size of the robot and the target frequency band to determine the appropriate array spacing, and adaptively deploy multiple miniature microphone arrays on the body surface and movable parts (such as robotic arms) of the robot to form a multi-layer three-dimensional collection network. In terms of the array configuration, a multi-dimensional microphone array is adopted, which covers the horizontal and vertical directions and can capture sound signals from all angles, thereby improving the accuracy and comprehensiveness of sound collection.
[0036] S2. Control the multi-microphone array to collect sound signals from different directions in the environment. Specifically, in this embodiment, after the terminal constructs the microphone array distribution mode adapted to the robot in step S1, the robot is assembled with a multi-microphone array manually according to the construction result. In actual implementation, the multi-microphone array module configured inside the robot body controls the multi-microphone array to collect different sound signals in the environment in real time.
[0037] S3. Process the sound signals to obtain multiple sound source signals, and extract the target sound source signal from the multiple sound source signals. Specifically, in this embodiment, in the front-end processing stage, the signal processing module configured inside the robot body comprehensively uses the blind source separation technology to process the mixed sound signals collected in step S2, and then extracts multiple independent sound source signals. In the back-end processing stage, the signal processing module uses the trained deep learning large model to accurately extract the target sound source signal from the multiple sound source signals.
[0038] S4. Locate the target sound source signal and construct a sound field model based on the located target sound source signal. Specifically, in this embodiment, in the sound source localization link, the signal processing module combines time delay estimation (TDOA) with a deep learning model, accurately calculates the time difference of the target sound source signal arriving at different microphones through time delay estimation, and then uses the powerful feature extraction and analysis capabilities of the deep learning model to output the three-dimensional coordinates and confidence of the target sound source signal, realizing accurate sound source localization. In the sound field reconstruction algorithm, the sound field modeling module configured inside the robot body selects any one of the wave superposition method or the equivalent source method, and these two algorithms can accurately construct a sound field model based on the target sound source signal.
[0039] S5. Control the robot to execute corresponding actions according to the sound field model. Specifically, in this embodiment, according to the sound field characteristics reconstructed in step S4, the interaction response module configured inside the robot body controls the robot to execute corresponding actions. Exemplarily, in a voice interaction scenario, the robot accurately responds and interacts based on the voice content and the speaker's position recognized by the sound field model; in a navigation and obstacle avoidance scenario, the robot judges the position and danger level of the obstacle based on the sound field model, plans a reasonable movement path, avoids noise sources or dangerous areas, and realizes safe and efficient movement.
[0040] In this way, by executing methods S1~S5, the dynamic interaction between the robot and the surrounding acoustic environment is realized, significantly improving the robot's perception and response capabilities in complex environments and enhancing the robot's intelligence level.
[0041] Optionally, controlling the multi-microphone array to collect sound signals from different directions in the environment in step S2 includes: dynamically collecting sound signals in the environment according to the motion posture of the robot. Specifically, in this embodiment, through the collaborative work of the telescopic linear array and the multi-dimensional microphone array, the step of dynamically collecting sound signals in the environment according to the motion posture of the robot is realized. Among them, the multi-microphone array includes a telescopic linear array, and the telescopic linear array can be adaptively adjusted according to the motion posture of the robot to ensure efficient sound collection in different motion states. In terms of the deployment strategy, the user can optimize the design according to the size of the robot and the target frequency band, determine the appropriate array spacing, so as to improve the accuracy and comprehensiveness of sound collection.
[0042] Optionally, before processing the sound signal to obtain multiple sound source signals in step S3, the method further includes: performing noise reduction, framing, and frequency domain transformation processing on the sound signal in sequence. Specifically, in this embodiment, before performing the processing of the sound signal, it is also necessary to preprocess the collected sound signal. The collected sound signal is sequentially subjected to noise reduction, framing, and frequency domain transformation (FFT) processing. Among them, the noise reduction processing improves the quality of the sound signal by removing environmental noise and the noise generated by the robot's own hardware. The framing processing divides the continuous sound signal into frames of a fixed length to facilitate the analysis and processing of subsequent steps. The frequency domain transformation processing converts the time domain signal into a frequency domain signal, extracts the frequency characteristics of the sound signal, and uses the differences in the frequency characteristics to separate different sound sources, providing richer information for the subsequent sound source separation and localization steps.
[0043] Optionally, processing the sound signal to obtain multiple sound source signals in step S3 includes: separating the speech, environmental noise, and mechanical self-noise in the sound signal according to a preset multi-group microphone array algorithm to obtain multiple sound source signals with direction characteristics. Specifically, in this embodiment, first, use a deep neural network (such as a fully connected network, a convolutional neural network, a recurrent neural network) to learn the mapping relationship between the reverberant signal and the dry sound signal. Secondly, use the powerful pattern recognition ability of the trained neural network model to accurately distinguish speech, environmental noise, and mechanical self-noise. Finally, combined with the geometric algorithm, according to information such as the time difference of sound arriving at different microphones, accurately calculate the direction and position of the sound source, and realize the effective separation and precise localization of multiple sound sources in a complex environment. Among them, the preset multi-group microphone array algorithm is preferably a trained neural network model, and the trained neural network model is preferably a convolutional neural network (CNN).
[0044] Optionally, extracting the target sound source signal from multiple sound source signals as described in step S3 includes: matching the sound source signal with a database trained based on speech, environmental noise, and mechanical self-noise samples, and extracting the speech in the sound source signal that conforms to the preset target features to obtain the target sound source signal. Specifically, in this embodiment, since the convolutional neural network can extract the local features of the signal, it is applicable to frequency-domain signal processing. Specifically, the convolutional neural network is trained with a large number of labeled speech, environmental noise, and mechanical self-noise samples so that it learns the features of the sounds of different sound sources and generates a trained database. In actual implementation, the sound source signal is matched with the database, and the successfully matched speech is extracted to obtain the target sound source signal.
[0045] In one scenario, when the robot faces a scene where multiple people are speaking simultaneously, it can identify the speech features, voiceprint information, etc. of different speakers, and determine the sound source signal that conforms to the preset target features (such as matching the voiceprint of the robot's preset interaction object, matching the keywords of the voice command, etc.) as the target sound source signal, and determine other unmatched sound source signals as interference signals.
[0046] Optionally, after extracting the target sound source signal from multiple sound source signals as described in step S3, the method further includes: dynamically adjusting the beam according to the direction of the target sound source signal to suppress interference signals in other directions; removing the reverberation formed by environmental reflection in the target sound source signal. Specifically, in this embodiment, after the extraction of the target sound source signal is completed, first, the adaptive beamforming technology (Beamforming) is used to dynamically adjust the beam according to the direction of the target sound source signal, so as to enhance the signal strength of the target sound source signal while suppressing interference signals in other directions. Then, the reverberation removal technology is adopted to remove the reverberation formed by environmental reflection in the target sound source signal and restore a clear and clean sound signal.
[0047] Optionally, constructing a sound field model based on the located target sound source signal in step S4 includes: obtaining the sound source parameters of the target sound source signal, where the sound source parameters at least include at least one of the position, intensity, frequency distribution, phase information, radiation pattern, time-varying characteristics, type, phase coherence, and nonlinear characteristics of the target sound source signal; constructing a sound field model according to the sound source parameters. Specifically, in one embodiment, any one of the wave superposition method or the equivalent source method is used to reconstruct the sound pressure distribution or particle velocity distribution of the space around the robot based on the sound source parameters of the target sound source signal, that is, the sound field model. In another embodiment, the sound field model of the space around the robot can also be reconstructed based on the principle of solving acoustic inverse problems. By reconstructing the sound field model, the layout of the microphone array can be optimized, the effects of sound source localization and speech enhancement can be improved, and more accurate acoustic environment information can be provided for the robot.
[0048] Optionally, after constructing the sound field model based on the located target sound source signal in step S4, the method further includes: S41. Real-time collect the motion state data of the robot. Specifically, in this embodiment, the robot relies on the on-body sensors to collect the motion state data. The inertial measurement unit (IMU) obtains the robot's attitude information in real time, such as roll angle, pitch angle, and yaw angle, etc., which can accurately reflect the posture change of the robot. The position encoder is used to record the position data of the robot, covering the coordinate information in space. In the indoor moving scenario of the robot, the inertial measurement unit can monitor the actions of the robot such as turning and bending, and the position encoder can determine its specific position in space. These data are used as the basis for subsequent adjustments.
[0049] S42. Obtain the motion change trend of the robot according to the motion state data. Specifically, in this embodiment, after receiving the collected motion state data, the signal processing module performs analysis. By comparing the attitude and position data at different times, the motion trend and change amplitude of the robot are judged. When the robot moves quickly or its posture changes significantly, it means that the surrounding sound field environment may change significantly, and the sound field model needs to be adjusted in time. If the change in the motion state is small, the adjustment frequency is appropriately reduced to balance the computing resources and the model accuracy.
[0050] S43. Dynamically update the sound field model according to the motion change trend. Specifically, in this embodiment, according to the analysis result of the motion state change, the sound field modeling module starts the dynamic update mechanism. If the wave superposition method or the equivalent source method is used for sound field modeling, the parameters of each point source or equivalent source in the model will be recalculated according to the change of the robot's position. When the robot approaches the sound source, the relevant parameters are adjusted to reflect the influence of the distance change on the sound pressure and particle velocity distribution. For the attitude change, the direction and angle of the sound collected by the microphone array are re-determined, and then the sound field model is updated to ensure that the model is synchronized with the actual sound field in real time, providing accurate acoustic environment information for the robot.
[0051] Optionally, after constructing the sound field model according to the located target sound source signal in step S4, in order to visually display the sound field information, through the display module carried by the robot, visualization technology is used, such as drawing a sound intensity diagram, generating a sound pressure isosurface, etc. to be displayed on the display module, so that the operator can clearly understand the distribution of the sound field. In this way, during the preliminary training of the robot by humans, the sound field visualization can enable the operator to clearly obtain the sound field distribution, facilitating the adjustment of the training plan according to the training results and subsequent scenario training.
[0052] Second Embodiment Please refer to Figure 2 , Figure 2 which shows a schematic structural diagram of a robot motion execution system and a robot equipped with this system. The system includes: A terminal 10, which is used to construct multiple groups of multi-microphone arrays with different characteristics according to the body characteristics of the robot.
[0053] A multi-microphone array module 20, which is used to control the multi-microphone array to collect sound signals from different directions in the environment.
[0054] A signal processing module 30, which is electrically connected to the multi-microphone array module 20, and is used to process the sound signal to obtain multiple sound source signals, extract the target sound source signal from the multiple sound source signals, and is also used to locate the target sound source signal.
[0055] A sound field modeling module 40, which is electrically connected to the signal processing module 30, and is used to construct a sound field model according to the located target sound source signal.
[0056] An interaction response module 50, which is electrically connected to the sound field modeling module 40, and is used to control the robot to perform corresponding actions according to the sound field model.
[0057] Among them, the terminal 10 can be a computer, a laptop, a tablet computer, a smart phone, a smart bracelet, a 3D scanner, etc., and the multi-microphone array module 20, the signal processing module 30, the sound field modeling module 40, and the interaction response module 50 are configured inside the robot body.
[0058] From the perspective of the microphone array deployment method, the prior art mostly adopts a planar or linear layout. This traditional layout form limits the perspective of the microphone array, and the acoustic data collected has great limitations, making it difficult to accurately reconstruct a complex three-dimensional sound field. For example, in an indoor environment, sound will be reflected and refracted multiple times, and it is difficult for a planar or linear array to capture the changes in sound signals in all directions, resulting in a large error in sound field reconstruction.
[0059] Optionally, the multi-microphone array module 20 includes a telescopic linear array 21. The telescopic linear array 21 can be deployed on the surface and movable parts of the robot body, covering three directions: width, depth, and height. When the robot moves, the movement of the robotic arm will change its position and posture. At this time, the telescopic linear array 21 can correspondingly change its own length and angle according to the movement state of the robotic arm. With its all-round coverage characteristics, it can capture sound signals from all angles. In the scenario where a security robot conducts patrols, when the robot turns, even if its body posture changes, the telescopic linear array 21 can still collect sound signals from different directions, and there will be no blind area in sound collection due to movement, ensuring the comprehensiveness and accuracy of sound signal collection. In the scenario where a service robot provides services to customers, when the robotic arm extends to pick up an item, the telescopic linear array 21 can extend to expand the sound collection range and adjust the angle at the same time to better capture the sound signals from the surrounding environment. When the robotic arm contracts and approaches the robot body, the telescopic linear array 21 also contracts accordingly, avoiding unnecessary interference due to being too long and ensuring efficient sound collection in different movement states.
[0060] Furthermore, the signal processing module 30 is also used to preprocess the collected sound signals. The collected sound signals are successively subjected to noise reduction, framing, and frequency domain transformation (FFT) processing. Among them, the noise reduction processing improves the quality of the sound signal by removing environmental noise and the noise generated by the robot's own hardware. The framing processing divides the continuous sound signal into frames of a fixed length, facilitating the analysis and processing of subsequent steps. The frequency domain transformation processing converts the time-domain signal into a frequency-domain signal, extracts the frequency characteristics of the sound signal, and uses the differences in frequency characteristics to separate different sound sources, providing richer information for sound source separation and localization in subsequent steps.
[0061] It can be understood that the interaction response module 50 is responsible for converting the sound field model into the behavior strategy of the robot. When a specific sound source is detected, such as a user's call or a danger signal, the interaction response module 50 is used to control the robot to turn towards the direction of the sound source. In case of sudden noises, such as collision sounds or alarm sounds, the interaction response module 50 is used to guide the robot to avoid in time. In this way, the intelligent interaction between the robot and the surrounding acoustic environment is realized, enabling it to make reasonable action decisions based on the reconstructed sound field model.
[0062] Optionally, the robot action execution system further includes a display module 60. The display module 60 is electrically connected to the sound field modeling module 40 and is used to display image information such as drawing sound force diagrams and generating sound pressure isosurfaces, so that the operator can clearly understand the distribution of the sound field. In this way, during the initial manual training of the robot, sound field visualization can enable the operator to clearly obtain the sound field distribution, facilitating the adjustment of the training plan according to the training results and subsequent scenario training.
[0063] In one scenario, taking a security robot as an example, multiple six-microphone circular arrays are evenly installed on the head and body of the security robot, and multiple four-microphone linear arrays are installed on the side and arm of the body respectively. The equivalent source method accelerated by GPU is used for sound field modeling, and the sound field model is updated every 100 ms. When abnormal sound sources appear in the monitoring area, such as the sound of broken glass or abnormal footsteps, the system determines the position of the sound source through accurate sound source localization and judges the danger level of the sound source in combination with the dynamically updated sound field model. Once a dangerous situation is determined, the alarm mechanism is immediately triggered, and at the same time, an alarm notification containing the sound source position information is sent to the terminal device of the security personnel, realizing real-time security monitoring of the monitoring area.
[0064] In another scenario, taking a service robot as an example, multi-dimensional microphone arrays are installed on the head and body of the service robot, and a telescopic linear array 21 is configured on the movable robotic arm. When a guest calls the robot in the hotel lobby, the multi-microphone array quickly collects sound signals. After signal preprocessing, sound source separation and localization, the system quickly calculates the three-dimensional position coordinates of the guest. The sound field modeling module 40 reconstructs the sound field in real time, providing more accurate acoustic environment information for the robot. The interaction response module 50 controls the robot to turn towards the direction of the guest according to the positioning information and sound field characteristics, and at the same time controls the robotic arm to make a friendly welcome gesture. During the interaction with the guest, the robot uses accurate sound source localization and noise suppression functions to clearly identify the voice commands of the guest and provide services such as guidance and consultation for the guest, realizing an efficient and natural human-machine interaction experience.
[0065] Based on the same inventive concept as the foregoing embodiments, an embodiment of the present invention provides an electronic device, such as Figure 3As shown, the device includes: a processor 310 and a memory 311 storing a computer program; wherein, Figure 3 The processor 310 shown in Figure 3 does not refer to the number of processors 310 being one, but only refers to the positional relationship of the processor 310 relative to other components. In practical applications, the number of processors 310 can be one or more; similarly, Figure 3 The memory 311 shown in Figure 3 has the same meaning, that is, it only refers to the positional relationship of the memory 311 relative to other components. In practical applications, the number of memories 311 can be one or more. When the processor 310 runs the computer program, the method applied to the above device is implemented.
[0066] The device may further include: at least one network interface 312. Each component in the device is coupled together through a bus system 313. It can be understood that the bus system 313 is used to realize the connection and communication between these components. In addition to including a data bus, the bus system 313 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 3 all kinds of buses are labeled as the bus system 313.
[0067] Among them, the memory 311 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM, Random Access Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as a static random access memory (SRAM, Static Random Access Memory), a synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), a dynamic random access memory (DRAM, Dynamic Random Access Memory), a synchronous dynamic random access memory (SDRAM, Synchronous Dynamic Random Access Memory), a double data rate synchronous dynamic random access memory (DDR SDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), an enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random Access Memory), a sync link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and a direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 311 described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0068] The memory 311 in the embodiments of the present invention is used to store various types of data to support the operation of the device. Examples of such data include: any computer programs for operating on the device, such as an operating system and application programs; contact data; phone book data; messages; pictures; videos, etc. Among them, the operating system contains various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application programs can include various application programs, such as a Media Player, a Browser, etc., for implementing various application services. Here, the program for implementing the method of the embodiments of the present invention can be included in the application programs.
[0069] Based on the same inventive concept as the foregoing embodiments, this embodiment also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. The computer-readable storage medium can be a ferromagnetic random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM), etc.; or it can be various devices including one or any combination of the above memories, such as a mobile phone, a computer, a tablet device, a personal digital assistant, etc. When the computer program stored in the computer-readable storage medium is run by a processor, the above method is implemented. For the specific step flow implemented when the computer program is executed by the processor, please refer to Figure 1 the description of the illustrated embodiments, which will not be repeated here.
[0070] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0071] In this document, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, in addition to the listed elements, and may also include other elements not specifically listed.
[0072] As described above, it is only a specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claimed rights.
Claims
1. A method for a robot to execute actions, characterized in that The method includes: Constructing multiple multi-microphone arrays with different characteristics according to the physical characteristics of the robot; Controlling the multi-microphone arrays to collect sound signals from different directions in the environment; Processing the sound signals to obtain multiple sound source signals, and extracting a target sound source signal from the multiple sound source signals; Locating the target sound source signal and constructing a sound field model based on the located target sound source signal; Controlling the robot to perform corresponding actions according to the sound field model.
2. The method according to claim 1, wherein Collecting the sound signals in the environment includes: Dynamically collecting the sound signals in the environment according to the motion posture of the robot.
3. The method according to claim 1, characterized in that, Before processing the sound signals to obtain multiple sound source signals, the method further includes: Performing noise reduction, framing, and frequency domain transformation processing on the sound signals in sequence.
4. The method according to claim 1, characterized in that Processing the sound signals to obtain multiple sound source signals includes: Separating the speech, environmental noise, and mechanical self-noise in the sound signals according to a preset multi-group microphone array algorithm to obtain multiple sound source signals with direction characteristics.
5. The method according to claim 4, characterized in that Extracting the target sound source signal from the multiple sound source signals includes: Matching the sound source signal with a database trained based on speech, environmental noise, and mechanical self-noise samples, and extracting the speech in the sound source signal that conforms to the preset target characteristics to obtain the target sound source signal.
6. The method according to claim 1, wherein After extracting the target sound source signal from the multiple sound source signals, the method further includes: Dynamically adjusting the beam according to the direction of the target sound source signal to suppress interference signals in other directions; Removing the reverberation formed by environmental reflection in the target sound source signal.
7. The method according to claim 1, characterized in that, Constructing the sound field model according to the target sound source signal includes: Obtaining the sound source parameters of the target sound source signal, where the sound source parameters at least include at least one of the position, intensity, frequency distribution, phase information, radiation pattern, time-varying characteristics, type, phase coherence, and nonlinear characteristics of the target sound source signal; Constructing the sound field model according to the sound source parameters.
8. The method according to claim 1, wherein After constructing the sound field model according to the target sound source signal, the method further includes: Real-time collecting the motion state data of the robot; Obtaining the motion change trend of the robot according to the motion state data; Dynamically updating the sound field model according to the motion change trend.
9. An electronic device, characterized in that, Including a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the steps of the robot action execution method according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the readable storage medium, and when the processor executes the computer program, the steps of the robot action execution method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Spherical following robot and following control method thereof
CN108297108A
Voice processing method and device, electronic equipment and storage medium
CN111383629A
Speaker positioning method, device and equipment
CN115424633A
Sound source positioning method and device
CN117289208A
Customer service robot and dialogue processing system and method based on deep learning
CN117995187A