A robot pickup method, device and medium based on a microphone array

By combining a stereo microphone array with lidar and thermal imager, the problem of blind spots caused by poor microphone array positioning on the robot was solved, enabling effective sound pickup from people of different directions and heights, and improving the robot's proactive greeting and interaction effects.

CN115767378BActive Publication Date: 2026-04-21XIAMEN MEIYA PICO INFORMATION CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAMEN MEIYA PICO INFORMATION CO LTD
Filing Date
2022-12-09
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

When existing microphone arrays are poorly positioned on the robot, they are difficult to pick up sound effectively, especially for people of different directions and heights, which affects the robot's ability to proactively approach and interact with guests.

Method used

A stereo microphone array solution is adopted. The first microphone array searches for audio signals in all directions, calculates the angle information of the sound source, and combines the LiDAR and thermal imager to determine the object. The robot is then controlled to move closer to the sound source and switch to the second microphone array to pick up the sound, achieving complementary switching to enhance the sound pickup capability.

Benefits of technology

It achieves blind-spot-free voice pickup, effectively adapts to people of different heights, improves voice interaction, and allows the robot to actively approach and pick up voices at the optimal angle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115767378B_ABST
    Figure CN115767378B_ABST
Patent Text Reader

Abstract

This application proposes a robot sound pickup method, device, and medium based on a microphone array. The method includes: S1, activating a first microphone array to search for audio signals in all directions; S2, determining the strongest target audio signal and its direction by comparison, and calculating the sound source angle information of the target audio signal; S3, controlling the robot to rotate to face the sound source and approach the sound source according to the sound source angle information; S4, switching to a second microphone array to pick up sound in the direction of the sound source. This application, by employing a stereo microphone array scheme and using a complementary switching method between the first and second microphone arrays, achieves blind-spot-free sound pickup by the robot, greatly improving the effectiveness of the robot's proactive approach and interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of signal processing, specifically to a method, device, and medium for robot sound pickup based on a microphone array. Background Technology

[0002] With the popularization of computer technology, people's lives have gradually entered the intelligent era, and artificial intelligence technology is applied to computers, mobile phones, tablets, smart TVs, etc. However, most portable devices such as mobile phones and tablets currently use near-field voice recognition technology, which has relatively poor ability to cope with complex environments, affecting human-computer interaction. At the same time, existing robot solutions mainly realize the function of turning around in place to greet guests, and the function of actively greeting guests is lacking or ineffective.

[0003] Therefore, some solutions introduce planar microphone arrays for robot human-robot interaction, providing functions such as voice wake-up, far-field sound pickup, sound source localization, and echo cancellation. This planar microphone array approach can meet the needs of small robots. However, for large robots, if the microphone array is placed too low, it is easily obstructed by structural components, preventing it from processing audio from different directions; if the microphone array is placed too high, the sound pickup effect for shorter people is still not ideal. Therefore, the superior performance of the microphone array is limited, and the effectiveness of the robot proactively greeting and interacting with visitors remains unsatisfactory. Summary of the Invention

[0004] To address the aforementioned technical problems, this application proposes a robot sound pickup method, device, and medium based on a microphone array.

[0005] According to the first aspect of this application, a method for robot sound pickup based on a microphone array is proposed, comprising the following steps:

[0006] S1. Activate the first microphone array to search for audio signals in all directions;

[0007] S2. By comparison, determine the target audio signal with the strongest signal and its direction, and calculate the sound source angle information of the target audio signal;

[0008] S3. Control the robot to rotate to face the sound source and move closer to the sound source based on the sound source angle information;

[0009] S4. Switch to the second microphone array to pick up sound in the direction of the sound source.

[0010] Preferably, the calculation of the sound source angle information in step S2 specifically includes:

[0011] At least two microphones in the first microphone array are selected based on the direction of the target audio signal;

[0012] The sound source angle information is calculated based on the time difference between the reception of the target audio signal by at least two selected microphones.

[0013] Preferably, step S3 further includes:

[0014] The robot uses a single-line lidar to detect the object emitting the sound and determines whether the object is a person. If so, it controls the robot to move to a preset distance from the object.

[0015] Preferably, determining whether the object is a person includes:

[0016] S3a. Use a thermal imager to measure the temperature of the object and determine whether the temperature measurement result is within the preset range. If so, proceed to step S3b.

[0017] S3b. Determine whether the shape of the object conforms to human morphology. If so, the object is a person.

[0018] Preferably, step S3b specifically includes:

[0019] S3b1. Based on the point cloud information scanned by the single-line lidar, determine whether the object has two arc-shaped surfaces at a preset height above the ground. If so, proceed to step S3b2.

[0020] S3b2. Determine whether the widths of the two arc-shaped surfaces and the distance between them are within a preset threshold. If so, the object is a person.

[0021] Preferably, when the first microphone array or the second microphone array is picking up sound, it forms a directional sound pickup beam based on the calculated sound source angle information.

[0022] Preferably, the first microphone array is arranged circumferentially on the robot, and the second microphone array is arranged vertically on the robot.

[0023] According to a second aspect of this application, a microphone array-based robot sound pickup device is provided, disposed on the robot, comprising:

[0024] The first microphone array is configured to search for audio signals in all directions;

[0025] The microphone array core module is configured to determine the target audio signal with the strongest signal and its direction by comparison, and to calculate the sound source angle information of the target audio signal.

[0026] The robot main control module is configured to control the robot to rotate to face the sound source and move closer to the sound source based on the sound source angle information;

[0027] A second microphone array is configured to pick up sound in the direction of the sound source;

[0028] An audio switching module is configured to switch between the first microphone array and the second microphone array.

[0029] Preferably, when the first microphone array or the second microphone array is picking up sound, it forms a directional sound pickup beam based on the calculated sound source angle information.

[0030] According to a third aspect of this application, a computer-readable storage medium is proposed that stores a computer program, which, when executed by a processor, implements the microphone array-based robotic sound pickup method as described in the first aspect of this application.

[0031] This application proposes a robot sound pickup method, device, and medium based on a microphone array. It employs a stereo microphone array scheme, using a complementary switching between a first microphone array and a second microphone array. The first microphone array compensates for the angular limitations of the second microphone array, while the second microphone array compensates for the vertical angle deficiencies of the first microphone array. This significantly enhances the robot's sound pickup capability in the vertical direction, achieving blind-spot-free sound pickup. During voice interaction, users will not experience poor interaction due to differences in microphone pickup caused by age or height. Furthermore, the robot uses a thermal imager and a lidar module to determine whether the sound source is a person or an object. Simultaneously, it uses lidar ranging to locate the sound source, actively approaches it, and adjusts to the optimal distance and angle for optimal sound pickup during voice interaction. Attached Figure Description

[0032] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated in and constitute a part of this specification. The drawings illustrate embodiments and, together with the description, serve to explain the principles of this application. Other embodiments and many anticipated advantages of these embodiments will be readily recognized as they become better understood through reference to the following detailed description. Elements in the drawings are not necessarily to scale. The same reference numerals refer to corresponding similar parts.

[0033] Figure 1 This is a flowchart of a robot sound pickup method based on a microphone array according to an embodiment of this application;

[0034] Figure 2This is a schematic diagram of a microphone array-based robotic sound pickup device according to an embodiment of this application.

[0035] Explanation of reference numerals in the attached diagram: 1. Power amplifier / recapture signal module; 2. First microphone array; 3. Second microphone array; 4. Microphone array core module; 5. Audio switching switch module; 6. Controller; 7. AGC gain filtering module; 8. Robot main control module; 9. Power amplifier module; 10. Audio output module. Detailed Implementation

[0036] The features and exemplary embodiments of various aspects of this application will now be described in detail. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only configured to explain this application and are not configured to limit this application. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples of this application.

[0037] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0038] According to the first aspect of this application, a method for robotic sound pickup based on a microphone array is proposed. Figure 1 A flowchart of a microphone array-based robot sound pickup method according to an embodiment of this application is shown, as follows: Figure 1 As shown, the method includes the following steps:

[0039] S1. Activate the first microphone array to search for audio signals in all directions.

[0040] Specifically, the first microphone array can contain N microphones, where N ≥ 1. These N microphones can be arranged horizontally, vertically, or randomly. The first microphone array can be positioned on the top or neck of the robot, or on a higher part of the robot. The first microphone array searches for audio signals in all directions, which can be 360° horizontally or 360° vertically. In this embodiment, the first microphone array contains 6 microphones, which are horizontally spaced apart on the top of the robot, thereby searching for audio signals in 360° horizontally.

[0041] It's important to note that the microphone array's sensitivity can be adjusted before activation. Adjusting the pickup sensitivity according to the operating environment will affect the echo cancellation capability. Increasing the sensitivity in one direction can easily cause the microphone array to misinterpret echoes. Therefore, in quieter environments, lowering the sensitivity can achieve better sound pickup.

[0042] S2. By comparison, determine the target audio signal with the strongest signal and its direction, and calculate the sound source angle information of the target audio signal.

[0043] In a specific embodiment, the strongest target audio signal and its direction are determined by comparing the audio signal strengths of multiple microphones in the first microphone array. Then, at least two microphones in the first microphone array that receive the target audio signal are selected, and the sound source angle information of the target audio signal can be calculated by comparing the time difference of the target audio signal arriving at different microphones using an algorithm.

[0044] S3. Based on the sound source angle information, control the robot to rotate to face the sound source and move closer to the sound source.

[0045] In a specific embodiment, after acquiring the sound source angle information, the robot first rotates to face the sound source, then determines whether the source is a person or an object, and simultaneously uses a single-line lidar to detect the distance to the object. For example, if the object is determined to be an object, the robot remains stationary, and the first microphone continues to search for audio signals in all directions; if the object is determined to be a person, it further determines whether the distance to the object is within a preset range (e.g., 2 meters). If so, the robot remains stationary; otherwise, the robot moves to a preset distance (e.g., 1.2 meters) from the person based on radar positioning and stops.

[0046] Determining whether the source of a sound is a person includes the following steps:

[0047] S3a. Use a thermal imager to measure the temperature of the object and determine whether the temperature measurement result is within the preset range. If so, proceed to step S3b.

[0048] S3b. Determine whether the shape of the object conforms to human morphology. If so, the object is a person.

[0049] In a specific embodiment, the robot uses an intelligent thermal imager to measure the temperature of the object emitting the sound source. If the temperature measurement result is within a preset range (e.g., 36-40 degrees Celsius), it indicates that the object is alive, but it is still uncertain whether it is a person or an animal, so further judgment is needed. In this embodiment, human morphological matching is specifically used to determine whether the object is a person, and the specific steps are as follows:

[0050] S3b1. Based on the point cloud information from the single-line lidar scan, determine whether the object has two arc-shaped surfaces at a preset height above the ground. If so, proceed to step S3b2.

[0051] S3b2. Determine whether the widths of the two curved surfaces and the distance between them are within a preset threshold. If so, the object is a person.

[0052] In a specific embodiment, based on the point cloud information from a previous single-line lidar scan, it can be determined whether the object has two curved surfaces at a preset height above the ground (e.g., about 15cm). If so, it can be preliminarily considered that these two curved surfaces conform to the morphology of two legs. Then, it is further determined whether the width of these two curved surfaces and the distance between them are within a preset threshold. For example, if it is determined that the width of these two curved surfaces is within 25cm and the distance between them is within 45cm, it means that the diameter and leg distance of these "two legs" basically conform to the morphology of the human body, and at this point, it can be determined that the object is a person.

[0053] S4. Switch to the second microphone array to pick up sound from the direction of the sound source.

[0054] In a specific embodiment, people of different ages correspond to different heights. In order to enable the robot to have a good sound pickup effect for people of different heights, a second microphone array is added in this embodiment. The second microphone array is set at a lower height than the microphone. Thus, the second microphone array makes up for the shortcomings of the first microphone array in the vertical angle, and the first microphone makes up for the shortcomings of the second microphone in the horizontal direction. The two complement each other and achieve sound pickup without blind spots.

[0055] Specifically, the second microphone array may contain M microphones, where N≥1. The M microphones may be arranged at intervals, in an array, or randomly in the vertical direction. The second microphone array picks up sound in a fixed direction. In this embodiment, the second microphone array contains 4 microphones, which are arranged at intervals along the vertical direction on the robot's body, thereby picking up sound in a fixed direction in the vertical direction.

[0056] Thus, the first microphone array is used to search for and locate audio signals in all directions, while the second microphone array is used to pick up sound in the location direction. The first and second microphone arrays achieve seamless sound pickup by switching and complementing each other.

[0057] In a preferred embodiment, after obtaining the sound source angle information, the first microphone array or the second microphone array can be controlled to form a directional sound pickup beam during the sound pickup process, thereby enhancing the sound pickup capability in that direction, suppressing the sound pickup capability in other directions, and reducing noise interference.

[0058] In summary, the implementation principle of this embodiment is as follows:

[0059] During sound pickup, the robot first uses its top-mounted first microphone array to search for audio signals in a 360° horizontal direction, identifying the strongest target audio signal and its direction. An algorithm then calculates the sound source angle information of this target audio signal. Based on this angle information, the robot rotates to face the sound source and uses a thermal imager and single-line LiDAR to identify the object emitting the sound. When a person is identified, the robot approaches them. Finally, it switches to the second microphone array to pick up sound from the direction of the person. Thus, the first and second microphone arrays complement each other to form a stereo microphone, achieving blind-spot-free sound pickup. The robot can proactively approach and interact with visitors with satisfactory results.

[0060] It should be noted that in this embodiment, regardless of the height of the person emitting the sound, the sound source angle information is ultimately obtained through the first microphone array for localization, and then the system switches to the second microphone array for sound pickup. However, in other embodiments, the first microphone array can also further obtain the height information of the sound source. Thus, when the height information of the sound source matches the height of the first microphone array, the robot continues to use the first microphone array for sound pickup after moving closer to the sound source; conversely, when the height information of the sound source does not match the height information of the first microphone array, the robot switches to the second microphone array for sound pickup after moving closer to the sound source.

[0061] This application proposes a microphone array-based robot sound pickup method. It employs a stereo microphone array scheme, using a complementary switching between a first microphone array and a second microphone array. The first microphone array compensates for the angular limitations of the second microphone array, while the second microphone array compensates for the vertical angle deficiencies of the first microphone array. This significantly enhances the robot's vertical sound pickup capability, achieving blind-spot-free sound pickup. During voice interaction, users will not experience poor interaction due to differences in microphone pickup caused by age or height. Furthermore, the robot uses a thermal imager and a lidar module to identify the sound source (person or object), and simultaneously uses lidar ranging to locate the sound source, actively approaching and adjusting to the optimal distance and angle for optimal sound pickup during voice interaction.

[0062] According to a second aspect of this application, based on the same concept, a microphone array-based robotic sound pickup device is also proposed, which is mounted on the aforementioned robot. Figure 2 A schematic diagram of a microphone array-based robotic sound pickup device according to an embodiment of this application is shown, as follows: Figure 2 As shown, the device includes:

[0063] First microphone array 2, configured to search for audio signals in all directions;

[0064] The microphone array core module 4 is configured to determine the target audio signal with the strongest signal and its direction by comparison, and to calculate the sound source angle information of the target audio signal.

[0065] Robot main control module 8 is configured to control the robot to rotate to face the sound source and move closer to the sound source based on the sound source angle information;

[0066] The second microphone array 3 is configured to pick up sound in the direction of the sound source;

[0067] Audio switching module 5 is configured to switch between the first microphone array 2 and the second microphone array 3;

[0068] The microphone array core module 4, based on its integrated speech algorithm, utilizes the spatial filtering characteristics of the microphone array and calculates the sound source angle information to enable either the first microphone array 2 or the second microphone array 3 to form a directional sound pickup beam during sound pickup. This enhances the sound pickup capability in that direction, suppresses sound pickup capabilities in other directions, and reduces noise interference. The robot main control module 8 is also used for ASR, NLP, and TTS processing of the audio signal.

[0069] It should be noted that the robot's main control module 8 has a motion structure at its bottom, which can realize movement functions such as turning, moving forward, and moving backward.

[0070] In a preferred embodiment, the device further includes:

[0071] Controller 6 is configured to control the audio switching module 5. In this embodiment, controller 6 is an MCU. The MCU communicates with the microphone array core module 4 via the I2C protocol. After obtaining the sound source angle information and the person identification result, it controls the audio switching module to switch between the first microphone array 2 and the second microphone array 3. At the same time, it controls the robot's steering and movement by controlling the robot main control module 8.

[0072] The power amplifier / echo sampling module 1 is configured to connect the power amplifier / echo cancellation reference input signal to the host computer's audio output interface via a 3.5mm audio cable. The microphone array core module 4 performs echo cancellation on the audio backend signal to prevent feedback from various signals received by the microphone array. For better echo cancellation, the placement distance between the speakers and microphones should be as far as possible, and sound insulation cotton should be used for isolation.

[0073] The AGC gain filtering module 7 is configured to send the audio signal collected by the microphone array into the AGC gain filtering circuit for amplification and filtering. Furthermore, the controller 6 adjusts the digitally adjustable resistor in the AGC gain filtering circuit to set the microphone array's pickup sensitivity. Adjusting the pickup sensitivity according to the microphone array's operating environment affects its echo cancellation capability. A unilateral increase in sensitivity can easily cause the microphone array to misinterpret echoes. Therefore, in quieter environments, reducing the sensitivity can achieve better pickup results.

[0074] Power amplifier module 9 is configured to send the audio signal processed by ASR, NLP, and TTS to a 100W power amplifier for audio amplification.

[0075] The audio output module 10, consisting of left and right channel speakers, is configured to output the audio signal processed by the microphone array core module 4 and the power amplifier module 9, ensuring that the robot can interact with people through voice with optimal sound pickup.

[0076] According to a third aspect of this application, based on the same concept, a computer-readable storage medium is further proposed that stores a computer program which, when executed by a processor, implements the microphone array-based robotic sound pickup method as described in the first aspect of this application.

[0077] In the embodiments of this application, it should be understood that the disclosed technical content can be implemented in other ways. The device / system / method embodiments described above are merely illustrative. For example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0078] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0079] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0080] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0081] It is obvious that those skilled in the art can make various modifications and alterations to the embodiments of this application without departing from the spirit and scope of this application. In this way, this application also aims to cover such modifications and alterations if they fall within the scope of the claims and their equivalents. The word "comprising" does not exclude the presence of other elements or steps not listed in the claims. The simple fact that certain measures are described in mutually different dependent claims does not indicate that a combination of these measures cannot be used for profit. Any reference numerals in the claims should not be considered limiting in scope.

Claims

1. A method for robot sound pickup based on a microphone array, characterized in that, Includes the following steps: S1. Activate the first microphone array to search for audio signals in all directions. The first microphone array is arranged at circumferential intervals on the robot. S2. By comparison, determine the target audio signal with the strongest signal and its direction, and calculate the sound source angle information of the target audio signal; S3. Detect the object emitting the sound source using a single-line lidar, and simultaneously determine whether the object is a person. If so, control the robot to move to a preset distance from the object; wherein... The determination of whether the object is a person includes: S3a. Use a thermal imager to measure the temperature of the object and determine whether the temperature measurement result is within the preset range. If so, proceed to step S3b1. S3b1. Based on the point cloud information scanned by the single-line lidar, determine whether the object has two arc-shaped surfaces at a preset height above the ground. If so, proceed to step S3b2. S3b2. Determine whether the widths of the two arc-shaped surfaces and the distance between them are within a preset threshold. If so, the object is a person. S4. Switch to the second microphone array to pick up sound from the direction of the sound source. The second microphone array is arranged vertically at intervals on the robot.

2. The method according to claim 1, characterized in that, The calculation of the sound source angle information in step S2 specifically includes: At least two microphones in the first microphone array are selected based on the direction of the target audio signal; The sound source angle information is calculated based on the time difference between the reception of the target audio signal by at least two selected microphones.

3. The method according to claim 1, characterized in that, When the first microphone array or the second microphone array is picking up sound, it forms a directional sound pickup beam based on the calculated sound source angle information.

4. A microphone array-based robot sound pickup device, disposed on the robot, characterized in that, include: A first microphone array, configured to search for audio signals in all directions, is arranged circumferentially spaced on the robot. The microphone array core module is configured to determine the target audio signal with the strongest signal and its direction by comparison, and to calculate the sound source angle information of the target audio signal. The robot's main control module is configured to use a single-line lidar to detect the object emitting the sound source and determine whether the object is a person. If so, the module controls the robot to move to a preset distance from the object. The determination of whether the object is a person includes using a thermal imager to measure the temperature of the object and determining whether the temperature measurement result is within a preset range. If so, based on the point cloud information scanned by the single-line lidar, the module determines whether the object has two arc-shaped surfaces at a preset height above the ground. If so, the module determines whether the width of the two arc-shaped surfaces and the distance between them are within a preset threshold. If so, the object is a person. A second microphone array, configured to pick up sound in the direction of the sound source, is arranged vertically at intervals on the robot. An audio switching module is configured to switch between the first microphone array and the second microphone array.

5. The apparatus according to claim 4, characterized in that, When the first microphone array or the second microphone array is picking up sound, it forms a directional sound pickup beam based on the calculated sound source angle information.

6. A computer-readable storage medium storing a computer program that, when executed by a processor, performs the method as described in any one of claims 1-3.

Citation Information

Patent Citations

  • An audio frequency acquisition method and device based on a microphone array

    CN106098075A

  • Robot, robot welcome movement method and readable storage medium

    CN113359753A

  • Self-moving robot control method, device and equipment and readable storage medium

    CN113787517A