Intelligent pickup

CN224774998UActive Publication Date: 2026-09-18SHENZHEN NANTIAN DONGHUA TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202521840173.4
Authority / Receiving Office
CN · China
Patent Type
Utility models(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2026-09-18
Estimated Expiration
2035-08-28

AI Technical Summary

Technical Problem

[0003]本实用新型的主要目的是提出一种智能拾音器,旨在解决现有的拾音器要求声源定位功能全时运行,能耗高,且依赖麦克风阵列的时延差进行声源粗定位,误判率高的技术问题

Benefits of technology

本实用新型的技术方案能够通过麦克风阵列、摄像头、控制电路板以及声源追踪机构的结合,分两步实现精确的声源定位操作,相对于现有技术传统技术仅依赖麦克风阵列的时延差进行声源定位操作,精度更高,也不容易引发定位混淆,因此可降低误判率。另外,该设计可在使用者说话的间隙令麦克风阵列进入低功耗监听模式并关闭摄像头,避免二者长期处于高负载状态,从而减少能耗、保证设备寿命并避免器件性能衰减而影响长期定位精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN224774998U_ABST
    Figure CN224774998U_ABST
Patent Text Reader

Abstract

The utility model discloses an intelligent pickup, including microphone array, control circuit board, camera and sound source tracking mechanism, control circuit board is connected with microphone array, camera and sound source tracking mechanism electricity respectively, and microphone array and camera all set up on sound source tracking mechanism. The utility model discloses the technical scheme can realize accurate sound source positioning operation in two steps through the combination of microphone array, camera, control circuit board and sound source tracking mechanism, and the precision is higher, and positioning confusion is not easy to cause, thereby can reduce the misjudgment rate. In addition, the design can make microphone array enter low -power monitoring mode and close camera in the interval of the user speaking, avoids both long -term in high load state, thereby reduces energy consumption, guarantees equipment life and avoids device performance attenuation and influences long -term positioning precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This utility model relates to the field of microphone technology, and in particular to an intelligent microphone. Background Technology

[0002] A microphone, also known as a listening device, is a device that integrates a microphone and audio amplification circuitry to collect ambient sound and transmit it to backend equipment. Currently, some microphones possess sound source localization and tracking capabilities. However, existing solutions generally require this function to operate continuously, resulting in the device being under high load for extended periods. This not only significantly increases energy consumption and shortens the device's lifespan but also affects long-term positioning accuracy due to device performance degradation. Furthermore, traditional technologies rely solely on the time delay difference of the microphone array for sound source localization, which can easily lead to localization confusion in dense sound source scenarios (such as when two people are close together), increasing the misjudgment rate. Utility Model Content

[0003] The main purpose of this invention is to propose an intelligent microphone, which aims to solve the technical problems of existing microphones that require the sound source localization function to operate continuously, have high energy consumption, and rely on the time delay difference of the microphone array for coarse sound source localization, resulting in a high misjudgment rate.

[0004] To achieve the above objectives, the present invention proposes an intelligent microphone, which includes a microphone array and a control circuit board, wherein the microphone array is electrically connected to the control circuit board; it also includes a camera and a sound source tracking mechanism, wherein the control circuit board is electrically connected to the camera and the sound source tracking mechanism respectively, and the microphone array and the camera are both mounted on the sound source tracking mechanism.

[0005] Optionally, it also includes a base, on which the sound source tracking mechanism is mounted, the sound source tracking mechanism including a first rotating frame, a first swing frame and a tracking power source.

[0006] Optionally, the first rotating frame includes a sleeve, and the base includes a vertical rod, with the sleeve rotatably fitted onto the top of the vertical rod.

[0007] Optionally, the bottom of the first swing frame is hinged to the top of the first rotating frame; The tracking power source is a dual-axis motor, which includes a vertical output shaft and a horizontal output shaft. The first rotating frame also includes a mounting cavity, in which the dual-axis motor is mounted. The vertical output shaft passes through a sleeve and is connected to a vertical rod, and the horizontal output shaft is connected to the first swing frame.

[0008] Optionally, the first swing frame includes a support rod and a mounting base, the microphone array is mounted on the mounting base, and the camera is disposed at the end of the support rod.

[0009] Optionally, it also includes a base, on which the sound source tracking mechanism is mounted. The sound source tracking mechanism includes a second rotating frame, a second swing frame, and a tracking power source. The second rotating frame is rotatably mounted on the top of the base, and the second swing frame is swingably mounted on the second rotating frame. The tracking power source is a rotating motor and a swing motor.

[0010] Optionally, the control circuit board integrates a wake-up word detection module, which is used to identify and detect the voice information collected by the microphone array.

[0011] Optionally, the control circuit board also integrates a processor unit, which integrates a sound source localization algorithm, a live face recognition localization algorithm, a mouth movement analysis algorithm, motor control logic, and an adaptive beamforming algorithm. The sound source localization algorithm is used to calculate the time delay difference of each microphone in the microphone array to locate the speaker's position. The live face recognition localization algorithm is used to identify faces. The mouth movement analysis algorithm is used to extract facial key points and calculate the mouth opening degree to determine the speaker. The motor control logic is used to achieve precise, efficient and safe motion management. The adaptive beamforming algorithm is used to dynamically adjust the pickup beam according to the orientation of the microphone array.

[0012] The technical solution of this utility model has the following beneficial effects: This invention achieves precise sound source localization in two steps through the combination of a microphone array, camera, control circuit board, and sound source tracking mechanism. Compared to existing technologies that rely solely on the time delay difference of the microphone array for sound source localization, this method offers higher accuracy and is less prone to localization confusion, thus reducing the false positive rate. Furthermore, this design allows the microphone array to enter a low-power listening mode and the camera to shut down during user speech intervals, preventing both from operating under high load for extended periods. This reduces energy consumption, extends device lifespan, and avoids performance degradation that could affect long-term positioning accuracy. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of this utility model or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this utility model. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0014] Figure 1 This is a schematic diagram of the overall structure of a first embodiment of the intelligent microphone of this utility model; Figure 2This is a schematic diagram of the overall structure of the dual-axis motor in Embodiment 1 of the intelligent microphone of this utility model; Figure 3 This is a circuit diagram of the dual-axis motor in Embodiment 1 of the intelligent microphone of this utility model; Figure 4 This is a schematic diagram of the overall structure of a second embodiment of the intelligent microphone of this utility model.

[0015] The attached figures illustrate the following: Example 1: Base 1, Vertical rod 11, First rotating frame 2, Sleeve 21, First swing frame 3, Support rod 31, Mounting base 311, Microphone 4, Dual-axis motor 5, Vertical output shaft 51, Horizontal output shaft 52.

[0016] Example 2: Base 1, vertical rod 11, first rotating frame 6, second swing frame 7, swing hinge shaft 71.

[0017] The realization of the purpose, functional features and advantages of this utility model will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0018] The technical solutions of the present utility model will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present utility model, and not all embodiments. Based on the embodiments of the present utility model, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present utility model.

[0019] It should be noted that all directional indicators (such as up, down, left, right, front, back, etc.) in this utility model embodiment are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicator will also change accordingly.

[0020] Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by this utility model.

[0021] This utility model proposes an intelligent microphone. Example 1

[0022] like Figures 1 to 3As shown in Embodiment 1 of this utility model, the intelligent microphone includes a microphone array, a control circuit board, a camera, and a sound source tracking mechanism. The control circuit board is electrically connected to the microphone array, the camera, and the sound source tracking mechanism, respectively. The microphone array and the camera are both mounted on the sound source tracking mechanism. In this embodiment, the microphone array is mounted on the microphone 4, and the camera is a high-definition infrared live camera with a supplementary lighting module for illumination. In use, the microphone array first collects surrounding voice information, then the control circuit board analyzes and processes the voice information, and then controls the sound source tracking mechanism to rotate and / or swing according to the analysis and processing results, so that the microphone array and the camera face the speaker, thereby achieving preliminary sound source localization. Secondly, the camera collects image information of the current orientation, then the control circuit board analyzes and processes the image information, and then controls the sound source tracking mechanism to further rotate and / or swing according to the analysis and processing results, so that the microphone array and the camera accurately face the speaker, thereby achieving precise sound source localization.

[0023] like Figure 1 and Figure 2 As shown, the intelligent microphone also includes a base 1, on which the sound source tracking mechanism is mounted. Specifically, the sound source tracking mechanism includes a first rotating frame 2, a first swing frame 3, and a tracking power source. In this embodiment, the base 1 includes a vertical rod 11, the first rotating frame 2 includes a sleeve 21 and a mounting cavity, the first swing frame 3 includes a support rod 31 and a mounting base 311, and the tracking power source is a dual-axis motor 5, which includes a vertical output shaft 51 and a horizontal output shaft 52. Further, the sleeve 21 of the first rotating frame 2 is rotatably fitted onto the top of the vertical rod 11, the bottom of the first swing frame 3 is hinged to the top of the first rotating frame 2, the dual-axis motor 5 is fixedly mounted in the mounting cavity, and the vertical output shaft 51 of the dual-axis motor 5 passes through the sleeve 21 and connects to the top of the vertical rod 11, while the horizontal output shaft 52 is connected to the first swing frame 3. Furthermore, the microphone array is mounted on the mounting base 311, and the camera is located at the end of the support rod 31.

[0024] It is worth noting that: transmission mechanisms are provided between the vertical output shaft 51 and the base 1 and between the horizontal output shaft 52 and the first swing frame 3. The vertical output shaft 51 and the horizontal output shaft 52 are respectively connected to the vertical rod 11 and the first swing frame 3 through two transmission mechanisms. Since the principle between the dual-axis motor 5, the vertical output shaft 51, the horizontal output shaft 52, the vertical rod 11 and the first swing frame 3 is a mature existing technology, it will not be described in detail in this embodiment.

[0025] When the control circuit board controls the sound source tracking mechanism to rotate based on the collected voice information, specifically, the dual-axis motor 5 drives the vertical output shaft 51 to rotate, thereby driving the first rotating frame 2 and the first swing frame 3 to rotate around the vertical rod 11 through the transmission mechanism. When the control circuit board controls the sound source tracking mechanism to swing based on the collected voice information, specifically, the dual-axis motor 5 drives the horizontal output shaft 52 to rotate, thereby driving the swing frame to swing around the horizontal output shaft 52 through the transmission mechanism.

[0026] like Figure 3 As shown, a wake-up word detection module and a processor unit are integrated on the control circuit board. Specifically: the microphone array, wake-up word detection module, processor unit, and sound source tracking mechanism are electrically connected in sequence; the camera, processor unit, and sound source tracking mechanism are also electrically connected in sequence.

[0027] In this embodiment, the wake-word detection module is used to identify and detect the voice information collected by the microphone array. Additionally, the processor unit integrates a sound source localization algorithm, a live face recognition localization algorithm, a mouth motion analysis algorithm, motor control logic, and an adaptive beamforming algorithm. Specifically, the sound source localization algorithm calculates the time delay difference between each microphone in the microphone array to locate the speaker's position; the live face recognition localization algorithm identifies faces; the mouth motion analysis algorithm extracts facial key points and calculates the mouth opening degree to determine the speaker; the motor control logic achieves accurate, efficient, and safe motion management; and the adaptive beamforming algorithm dynamically adjusts the pickup beam according to the orientation of the microphone array. Since the sound source localization algorithm, live face recognition localization algorithm, mouth motion analysis algorithm, motor control logic, and adaptive beamforming algorithm are all mature existing technologies, their principles will not be elaborated further. Example 2

[0028] like Figure 4 As shown, the difference between this embodiment and Embodiment 1 is that the sound source tracking mechanism includes a second rotating frame, a second swing frame 7, and a tracking power source. The second rotating frame is also sleeved on the top of the vertical rod 11, but a mounting groove is provided on the upper part of the second rotating frame. The second swing frame 7 is mounted in the mounting groove through a swing hinge shaft 71. The aforementioned camera and microphone array are both located on the front of the second swing frame 7. In this embodiment, the tracking power source is a rotation motor and a swing motor. The rotation motor is used to drive the second rotating frame to rotate around the vertical rod 11, and the swing motor is used to drive the second swing frame 7 to swing around the swing hinge shaft 71.

[0029] Specifically, the working principle and process of this utility model are as follows: First, preliminary sound source localization is triggered by the wake-word detection module: when no one is speaking, the microphone array is in low-power monitoring mode, and the camera is also off. When the speech captured by the microphone array is recognized by the wake-word detection module, and the wake-word module detects a preset wake-word (such as "The meeting is now starting," "Hello," etc.) from the recognized speech information, the processor unit, based on the detection result of the wake-word detection module, enables the microphone array to start full-power speech information acquisition mode. Then, the sound source localization algorithm in the processor unit calculates the approximate location of the sound source (i.e., the speaker) based on the time difference of address (TDOA) and phase difference of the speech signal received by the microphone array in full-power acquisition mode. The processor unit then controls the sound source tracking mechanism to turn and / or swing according to the calculation result of the sound source localization algorithm.

[0030] Secondly, the processor unit identifies dynamic faces to confirm the speaker: the camera collects video and / or image information of the coarse localization area, and then the live face recognition localization algorithm identifies the video and / or image information to extract the live face (avoiding interference from photos, portraits, etc. to locate the sound source), and extracts the face bounding box coordinates (x,y,w,h) to segment the mouth area of ​​each live frontal target; then the mouth movement analysis algorithm calculates the mouth movement to lock the current speaker.

[0031] The position and orientation of the microphone array and camera are precisely adjusted again by the sound source tracking mechanism to achieve accurate sound source localization: based on the face bounding box coordinates (x,y,w,h) and the camera's intrinsic parameter matrix, the speaker's three-dimensional coordinates (X,Y,Z) are located. Then, the motor control logic controls the dual-axis motor 5 to adjust the position and orientation of the microphone array and camera, so that the central axis of the microphone array is aligned with the speaker's direction. The adaptive beamforming algorithm is activated to form a directional sound pickup beam based on the orientation of the microphone array and camera. At the same time, adaptive noise cancellation (ANC) is used to eliminate low-frequency ambient noise.

[0032] Finally, the wake-word detection module and the processor unit determine whether the microphone array has entered a low-power listening mode and whether the camera has been turned off. If the microphone array does not record voice information within a preset time, the processor unit determines whether to put the microphone array into a low-power listening mode and turn off the camera. This reduces the power of the microphone array and camera during the user's speech intervals, avoids both being in a high-load state for a long time, thereby reducing energy consumption, ensuring device lifespan, and preventing device performance degradation that could affect long-term positioning accuracy.

[0033] It is worth noting that when the smart microphone is no longer needed (such as when the meeting ends), the user can directly turn it off without maintaining the microphone array in low-power monitoring mode.

[0034] This invention achieves precise sound source localization in two steps through the combination of a microphone array, camera, control circuit board, and sound source tracking mechanism. Compared to existing technologies that rely solely on the time delay difference of the microphone array for sound source localization, this method offers higher accuracy and is less prone to localization confusion, thus reducing the false positive rate. Furthermore, this design allows the microphone array to enter a low-power listening mode and the camera to shut down during user speech intervals, preventing both from operating under high load for extended periods. This reduces energy consumption, extends device lifespan, and avoids performance degradation that could affect long-term positioning accuracy.

[0035] The above description is only a preferred embodiment of the present utility model and does not limit the patent scope of the present utility model. All equivalent structural transformations made under the inventive concept of the present utility model using the contents of the present utility model specification and drawings, or direct / indirect applications in other related technical fields, are included within the patent protection scope of the present utility model.

Claims

1. An intelligent pickup, comprising a microphone array and a control circuit board, the microphone array being electrically connected with the control circuit board; characterized in that, It also includes a camera and a sound source tracking mechanism. The control circuit board is electrically connected to the camera and the sound source tracking mechanism respectively. The microphone array and the camera are both mounted on the sound source tracking mechanism.

2. The smart pickup of claim 1, wherein, It also includes a base, on which the sound source tracking mechanism is mounted. The sound source tracking mechanism includes a first rotating frame, a first swing frame, and a tracking power source.

3. The smart pickup of claim 2, wherein, The first rotating frame includes a sleeve, and the base includes a vertical rod, with the sleeve rotatably fitted onto the top of the vertical rod.

4. The smart pickup of claim 3, wherein, The bottom of the first swing frame is hinged to the top of the first rotating frame; The tracking power source is a dual-axis motor, which includes a vertical output shaft and a horizontal output shaft. The first rotating frame also includes a mounting cavity, in which the dual-axis motor is mounted. The vertical output shaft passes through a sleeve and is connected to a vertical rod, and the horizontal output shaft is connected to the first swing frame.

5. The smart pickup of claim 3, wherein, The first swing frame includes a support rod and a mounting base, the microphone array is mounted on the mounting base, and the camera is disposed at the end of the support rod.

6. The smart pickup of claim 1, wherein, It also includes a base, on which the sound source tracking mechanism is mounted. The sound source tracking mechanism includes a second rotating frame, a second swing frame, and a tracking power source. The second rotating frame is rotatably mounted on the top of the base, and the second swing frame is swingably mounted on the second rotating frame. The tracking power source is a rotating motor and a swing motor.

7. The smart pickup of claim 1, wherein, The control circuit board integrates a wake-up word detection module, which is used to identify and detect the voice information collected by the microphone array.

8. The smart pickup of claim 1, wherein, The control circuit board also integrates a processor unit, which integrates a sound source localization algorithm, a live face recognition localization algorithm, a mouth movement analysis algorithm, motor control logic, and an adaptive beamforming algorithm.