Microphone device and method capable of positioning sound source and digital human device
By integrating infrared laser ranging sensors in the microphone array, combining TDOA algorithm and signal processing module, high-precision three-dimensional positioning of sound sources is achieved, solving the problem of insufficient three-dimensional positioning accuracy and robustness in the existing technology, and improving the real-time and anti-interference ability of the sound source positioning system.
Patent Information
- Application Number
- CN202510793402.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-25
AI Technical Summary
The existing acoustic positioning system is difficult to achieve high-precision sound source positioning in three-dimensional space. The data acquisition of sensor systems is not synchronized and calibration error accumulation leads to limited spatial coordinate fusion accuracy, insufficient multipath interference suppression ability in dynamic sound source scenarios, and easy drift in complex acoustic environments.
The ring array matrix structure consisting of multiple sound pickup microphones and infrared laser ranging sensors is adopted, combined with the TDOA algorithm and signal processing module, sound and distance signals are synchronized, and the three-dimensional coordinates of the sound source are obtained through the infrared laser ranging sensor to achieve high-precision and real-time positioning.
It realizes high-precision positioning of three-dimensional coordinates of sound sources, improves the accuracy of spatial and temporal data matching, enhances the robustness of sound sources in complex environments, reduces the system's computing complexity and power consumption, and supports real-time tracking in dynamic scenarios.
Smart Images

Figure CN120370261A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and acoustic positioning, and particularly relates to a three-dimensional space sound capture and processing system based on a microphone array and a sound source localization algorithm, which is particularly applicable to digital human devices, intelligent voice devices, and human-machine collaboration systems that require precise sound source direction localization and intelligent interaction. Background Art
[0002] With the rapid development of artificial intelligence technology, digital human devices with natural human-computer interaction capabilities are increasingly widely used in fields such as education, customer service, and entertainment. As the most natural interaction method, face-to-face communication requires digital humans to accurately perceive the spatial position of the sound source, so as to realize anthropomorphic interaction behaviors such as line-of-sight guidance and voice response.
[0003] Sound source localization technology, as a key research direction in the field of intelligent perception, has important application values in scenarios such as robot interaction, intelligent security, and remote conferencing systems. With the popularization of intelligent devices, the market's demand for high-precision and low-latency acoustic positioning is increasing day by day. Especially, the ability to realize sound source trajectory tracking in three-dimensional space has become one of the core technical challenges for improving the human-computer interaction experience.
[0004] Current mainstream acoustic positioning systems mostly rely on time-delay estimation algorithms based on microphone arrays, and calculate the sound source azimuth by analyzing the time difference of sound waves arriving at different microphones. Typical solutions include beamforming technology for linear arrays and angle-of-arrival estimation methods for planar arrays. Some improved algorithms can improve the horizontal plane positioning accuracy by optimizing the array layout. There are also studies that attempt to introduce auxiliary positioning means, such as combining visual sensors or inertial navigation modules, but the system integration complexity increases significantly.
[0005] The existing technology system still faces three challenges: First, traditional acoustic positioning methods lack effective measurement means in the vertical dimension and are difficult to meet the complete solution requirements of three-dimensional space coordinates; second, multi-modal sensor systems have problems such as asynchronous data acquisition and cumulative calibration errors, resulting in limited spatial coordinate fusion accuracy; in addition, the ability to suppress multipath interference in dynamic sound source scenarios is insufficient, and complex acoustic environments are likely to cause positioning result drift. These technical defects seriously restrict the popularization and deployment of sound source localization systems in industrial application scenarios. Summary of the Invention
[0006] In view of this, embodiments of the present invention provide a microphone device, method, and digital human device that can localize sound sources to solve the above technical problems.
[0007] To achieve the above object, in a first aspect, a microphone device that can localize sound sources is provided, which includes:
[0008] Multiple basic units, each of the basic units includes a plurality of pickup microphones arranged in a circular and evenly spaced manner, and an infrared laser ranging sensor located at the center of the pickup microphones, and the plurality of pickup microphones are fixedly installed around the infrared laser ranging sensor; the plurality of basic units are arranged to form a ranging microphone array matrix;
[0009] A signal processing module, configured to synchronously collect the sound signals of the pickup microphones and the distance signals of the infrared laser ranging sensor, and perform source localization on the sound source based on the time difference of arrival (TDOA) algorithm to obtain the two-dimensional plane coordinates of the sound source; determine the three-dimensional coordinates of the sound source according to the two-dimensional plane coordinates and the distance signals of the infrared laser ranging sensor.
[0010] In a second aspect, a digital human device is provided, which includes:
[0011] The microphone device described above, configured to determine the three-dimensional coordinates of the sound source;
[0012] An interaction module, connected to the signal processing module of the microphone device, configured to adjust the display posture of the digital human to the direction of the sound source according to the three-dimensional coordinates, and perform face-to-face interaction with the sound source object.
[0013] In a third aspect, a method for sound source localization based on the microphone device capable of localizing a sound source described in the first aspect is provided, which includes:
[0014] Synchronously collect sound signals through the pickup microphones, and synchronously collect distance signals through the infrared laser ranging sensor;
[0015] Process the sound signals based on the time difference of arrival (TDOA) algorithm to obtain the two-dimensional plane coordinates of the sound source;
[0016] Determine the three-dimensional coordinates of the sound source according to the two-dimensional plane coordinates and the distance signals.
[0017] The above technical solutions have the following beneficial technical effects:
[0018] The present invention realizes the high-precision and real-time positioning of the three-dimensional coordinates of the sound source by integrating the infrared laser ranging sensor and the sound pickup microphone in a circular array and constructing a matrix structure of multiple basic units, effectively solving the problems such as difficult acquisition of distance information and low sensor cooperation efficiency in the prior art. Specifically, the beneficial effects of the present invention are as follows: First, by integrating the infrared laser ranging sensor at the center of the sound pickup microphone in each basic unit, a spatially close cooperation layout is formed, which not only simplifies the system structure but also can directly obtain the distance of the sound source by using the high-precision characteristics of laser ranging, avoiding the limitation that the traditional TDOA algorithm can only provide two-dimensional coordinates; Second, the signal processing module synchronously collects the sound and distance signals and performs fusion processing based on a unified time reference, improving the matching accuracy of spatio-temporal data, especially enabling real-time tracking of the sound source position in a dynamic scenario; Third, the circularly and evenly distributed sound pickup microphone array enhances the ability to capture sound signals in all directions, combined with the central ranging sensor to form an all-round perception system, effectively improving the robustness of sound source positioning in complex environments; Fourth, the matrix structure of multiple basic units can further optimize the positioning algorithm through space diversity technology, improve the positioning accuracy by using redundant information, and at the same time support a distributed processing architecture, reducing the system calculation complexity and power consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:
[0020] Figure 1 is a functional block diagram of a microphone device capable of positioning a sound source according to an embodiment of the present invention;
[0021] Figure 2 is a schematic structural diagram of a basic unit according to an embodiment of the present invention;
[0022] Figure 3 is a functional block diagram of a digital human device according to an embodiment of the present invention;
[0023] Figure 4 is a schematic structural diagram of a computer system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The following describes exemplary embodiments of the present invention with reference to the drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, the description below omits the description of well-known functions and structures.
[0025] Embodiment 1
[0026] As Figure 1As shown, this embodiment provides a microphone device capable of locating a sound source, which includes:
[0027] Multiple basic units ( Figure 1 Only one basic unit is drawn exemplarily), each of the basic units includes a plurality of pickup microphones arranged in a ring shape with uniform spacing, and an infrared laser ranging sensor located at the center of the pickup microphones, and the plurality of pickup microphones are fixedly installed around the infrared laser ranging sensor; the plurality of basic units are arranged to form a ranging microphone array matrix;
[0028] The signal processing module is used to synchronously collect the sound signal of the pickup microphone and the distance signal of the infrared laser ranging sensor, and locate the sound source based on the TDOA (Time Difference of Arrival) algorithm to obtain the two-dimensional plane coordinates of the sound source; and determine the three-dimensional coordinates of the sound source according to the two-dimensional plane coordinates and the distance signal of the infrared laser ranging sensor.
[0029] like Figure 2 As shown, in some embodiments, each of the basic units includes six pickup microphones 101, which are arranged in a ring with uniform spacing to form a regular hexagonal structure, and the infrared laser ranging sensor 102 is located at the center of the regular hexagon.
[0030] In some embodiments, the sound pickup microphones 101 are spaced at equal distances from each other, and the emission direction of the infrared laser ranging sensor 102 is perpendicular to the plane where the sound pickup microphones 101 are located and is arranged outward.
[0031] In some embodiments, the TDOA algorithm includes the following steps: calculating the time difference Δt of the sound signal reaching different pickup microphones through signal correlation analysis; converting the time difference into a distance difference according to the distance difference formula Δd=c·Δt, where c is the sound propagation speed; converting each of the distance differences into a hyperbolic equation, and constructing a nonlinear equation group containing multiple hyperbolic equations based on a receiver formed by at least 3 basic units; using a numerical method to solve the nonlinear equation group to obtain the two-dimensional plane coordinates of the sound source; the numerical method includes: least squares method, Chan algorithm or Taylor series expansion method.
[0032] In some embodiments, the signal processing module is specifically used to determine the corresponding target basic unit according to the two-dimensional plane coordinates, call the distance signal of the infrared laser ranging sensor in the target basic unit, and calculate the vertical distance between the sound source and the ranging microphone array matrix in combination with the two-dimensional plane coordinates to form the three-dimensional coordinates of the sound source.
[0033] In some embodiments, the signal processing module determines a corresponding target basic unit according to the two-dimensional plane coordinates and retrieves the distance signal of the infrared laser ranging sensor in the target basic unit, which specifically includes: based on the position mapping relationship between the two-dimensional plane coordinates and each basic unit in the ranging microphone array matrix, determining the number of the target basic unit where the sound source is located; retrieving the distance signal of the infrared laser ranging sensor corresponding to the number of the target basic unit.
[0034] In some embodiments, each basic unit of the ranging microphone array matrix achieves time synchronization through a synchronization method; the sound pickup microphone and the infrared laser ranging sensor are integrated on the same substrate, and a sound transmission structure for acoustic wave conduction and a light transmission window for infrared signal emission are arranged on the surface of the substrate.
[0035] Specifically, the ranging microphone array matrix adopts a master-slave hardware synchronization architecture to achieve time synchronization of each basic unit. The reference unit serves as the master node, and a high-precision clock module is built in to provide a stable reference clock, and transmits the clock signal to other basic units as slave nodes through a dedicated synchronization bus; each slave node is configured with a phase-locked loop circuit to phase-lock the local clock with the master node clock to form a unified time reference. The master node periodically sends timestamp calibration information, and the slave node automatically adjusts the receiving parameters according to the signal reception status to compensate for the bus transmission delay, ensuring that the sampling clocks of each unit maintain high-precision synchronization and providing a reliable time reference for the subsequent time difference measurement of sound source localization.
[0036] Specifically, the sound pickup microphone and the infrared laser ranging sensor are integrated on the same substrate through a modular design, and functional partitions are arranged on the substrate surface to be compatible with different signal transmission requirements. Multiple sound pickup units are arranged in the microphone area, and a sound transmission structure is opened on the surface and covered with a sound transmission material, which allows acoustic waves to be efficiently conducted while blocking foreign impurities from entering; a light transmission window is correspondingly arranged in the infrared sensor area, and a filter film matching the infrared signal wavelength is covered to ensure that the infrared emission and reception signals pass through without attenuation and filter out environmental light interference. The signal processing circuit is integrated on the back of the substrate, and the electromagnetic shielding design is adopted to reduce the mutual interference between the microphone pre-amplification circuit and the laser sensor drive circuit, realizing the integrated integration of the acoustic wave acquisition and infrared ranging functions.
[0037] In some embodiments, the signal processing module includes a preprocessing unit, a positioning calculation unit, and a coordinate fusion unit; the preprocessing unit is configured to perform band-pass filtering denoising and timestamp synchronization calibration on the sound signal and the distance signal to obtain a calibrated sound signal sequence and a distance signal sequence; the positioning calculation unit is configured to perform a TDOA algorithm based on the calibrated sound signal sequence to calculate the two-dimensional plane coordinates of the sound source; the coordinate fusion unit is configured to generate the three-dimensional coordinates of the sound source according to the distance signal corresponding to the target basic unit in the two-dimensional plane coordinates and the calibrated distance signal sequence.
[0038] The specific working process of the three-dimensional positioning of the device is as follows: When the sound source emits a sound, the signal processing module first synchronously collects the sound signals (sampling rate 48 kHz) of all the sound pickup microphones 101 and the distance signals (measurement frequency 50 Hz) of the infrared laser ranging sensors 102 through a high-precision clock source, and adds accurate timestamps with an error less than 10 ns to the two types of signals; then performs preprocessing on the sound signals, including band-pass filtering with 300 Hz - 8 kHz and spectral subtraction denoising, calculates the TDOA values of the sound source arriving at each microphone based on the generalized cross-correlation - phase transform algorithm, and constructs a hyperbolic equation system in combination with the geometric parameters of the microphone array; then uses the TDOA measurement data of at least 4 basic units, solves the nonlinear equation system through the Chan algorithm to obtain the preliminary two-dimensional plane coordinates (x0, y0) of the sound source; further, by calculating the Euclidean distance between (x0, y0) and the center coordinates of each basic unit, selects the basic unit with the closest distance and retrieves the real-time measurement value z0 of its infrared laser ranging sensor, combines z0 with the two-dimensional coordinates to form three-dimensional coordinates (x0, y0, z0); finally, performs Kalman filtering processing on the three-dimensional coordinate sequence of 10 consecutive frames, outputs the smoothed final three-dimensional coordinates, and realizes the positioning accuracy of horizontal ±5°, vertical ±3°, and distance ±5 mm.
[0039] When the microphone device of this embodiment is applied to a digital human system, the signal processing module transmits the three-dimensional coordinate data to the digital human control unit through the USB 3.0 interface. The control unit calculates the orientation angles (horizontal yaw angle and vertical pitch angle) of the digital human according to the coordinates, and drives the digital human model to perform corresponding posture adjustments on the display interface. At the same time, the system dynamically adjusts the voice response volume and expression richness of the digital human according to the distance information, realizing a more natural face-to-face interaction experience. In a multi-sound source scenario, the system can distinguish the main sound source and the background sound source based on the energy threshold and the spatial clustering algorithm, and preferentially perform positioning and response on the main sound source to improve the interaction performance in complex environments.
[0040] Embodiment 2
[0041] As Figure 3 shown, this embodiment provides a digital human device, which includes:
[0042] The microphone device described in Embodiment 1 is used to determine the three-dimensional coordinates of a sound source;
[0043] An interaction module, connected to the signal processing module of the microphone device, is configured to adjust the display posture of the digital human to the sound source direction according to the three-dimensional coordinates and perform face-to-face interaction with the sound source object.
[0044] The microphone device in this embodiment adopts the structure described in Embodiment 1, and realizes accurate sound source localization through the design and combination of multiple basic units. Each basic unit contains 6 pick-up microphones, which are arranged in a regular hexagon structure with uniform spacing in a ring shape, and the infrared laser range finder is located at the center of the regular hexagon. The advantage of this regular hexagon layout is that 6 pick-up microphones can be evenly distributed on the same circumference at equal intervals, ensuring high balance and comprehensiveness in the acquisition of sound source signals in all directions, and effectively reducing the blind area of signal acquisition. Multiple such basic units are arranged to form a ranging microphone array matrix. By reasonably planning the positions and orientations of each basic unit, the coverage range of sound source localization can be expanded, and the accuracy and reliability of localization can be improved. The signal processing module synchronously acquires the sound signals of the pick-up microphones and the distance signals of the infrared laser range finder, locates the sound source based on the TDOA algorithm, determines the coordinates of the sound source in the two-dimensional plane using the time difference of the sound signals collected by multiple pick-up microphones, and then combines the distance signals provided by the infrared laser range finder to accurately calculate the three-dimensional coordinates of the sound source. This microphone device combines the basic unit of the regular hexagon structure with the array matrix, giving full play to the synergistic effect of the pick-up microphones and the infrared laser range finder, and realizing high-precision three-dimensional sound source localization.
[0045] In this embodiment, the interaction module establishes a connection with the signal processing module of the microphone device and can receive the three-dimensional coordinate data of the sound source in real time. It accurately adjusts the display posture of the digital human to the direction of the sound source according to the three-dimensional coordinates and realizes face-to-face interaction with the sound source object. Specifically, the interaction module integrates an attitude adjustment algorithm and a drive control unit internally. After receiving the three-dimensional coordinates, the attitude adjustment algorithm quickly calculates the angles, directions, and attitude parameters of each part of the body that the digital human needs to rotate. The drive control unit then controls the display device of the digital human, such as a display screen or a virtual image generation module, according to these parameters, so that the postures such as the face orientation and body orientation of the digital human are adjusted to the direction of the sound source in real time. When interacting face-to-face with the sound source object, the interaction module can not only realize the posture adjustment of the digital human, but also combine the preset interaction logic and speech recognition and synthesis technologies to enable the digital human to actively voice a response to the sound source object, realizing natural and smooth two-way interaction. For example, when the sound source object is located in the front left of the digital human, the interaction module will control the digital human's head to turn left and the body to lean forward slightly, presenting a posture of concentrating on listening. At the same time, it generates a corresponding answer according to the speech content of the sound source object and plays it through the speech output module of the digital human. This interaction module realizes the full process automation and intelligence from sound source positioning to digital human posture adjustment and interaction. Through accurate coordinate parsing and efficient control algorithms, it ensures that the digital human can interact face-to-face with the sound source object in real time and accurately, improving the realism and experience of the interaction.
[0046] In some embodiments, the interaction module includes: a display control unit for calculating the orientation parameters of the digital human in the display interface according to the three-dimensional coordinates; a rendering engine connected to the display control unit for generating a three-dimensional model of the digital human facing the direction of the sound source based on the orientation parameters; and a voice processing unit for performing recognition processing on the sound source voice signal collected by the microphone device, generating interactive text content, and converting the interactive text content into a voice signal through speech synthesis technology and outputting it through the speaker of the digital human device.
[0047] Specifically, the display control unit of this embodiment avoids the limitations of traditional two-dimensional plane orientation calculation and constructs a multi-dimensional space mapping model. First, the unit converts the three-dimensional coordinates (X, Y, Z) of the sound source output by the microphone device into local coordinate system parameters with the digital human model as the origin. By introducing a spatial rotation algorithm based on quaternions, it accurately calculates the Euler angles (pitch angle, yaw angle, roll angle) of the digital human in the display interface. At the same time, it performs dynamic compensation in combination with the physical parameters of the display device (such as the curvature of the curved screen, the geometric relationship of multi-screen splicing). In particular, for the situation where there are multiple sound sources in a complex scene, the display control unit adopts a weighted fusion algorithm, dynamically assigns priorities according to the signal intensity and distance parameters of each sound source, and ensures that the digital human posture remains natural and coordinated in multi-target interaction. This design realizes the non-linear mapping from physical space coordinates to display space parameters. By introducing a device calibration matrix and a dynamic weight strategy, it solves the perspective deviation problem of traditional plane orientation calculation in three-dimensional space interaction, enabling the digital human to maintain an accurate visual orientation on special-shaped display devices such as curved screens and holographic projections.
[0048] The device calibration matrix is a mathematical conversion matrix established in advance for the physical characteristics of the display device (such as the curvature radius of the curved screen, the spatial geometric relationship formed by multi-screen splicing, the optical path refraction parameters of holographic projection, etc.). It stores the mapping relationship parameters between the actual physical space of the display device and the coordinate space of the digital human model. By operating the three-dimensional coordinates of the sound source collected by the microphone device with this matrix, geometric compensation for device characteristics such as curved surface distortion, perspective shift, and multi-screen seams can be achieved, ensuring the physical consistency between the digital human orientation parameters and the actual presentation effect of the display device.
[0049] The dynamic weight strategy is to construct a weight function according to the signal intensity (such as the sound pressure level amplitude collected by the pick-up microphone) and distance parameters (the straight-line distance measured by the infrared laser ranging sensor) of each sound source in a multi-sound source interaction scenario. Through the Gaussian attenuation model or the inverse distance weighted algorithm, the priority weights of each sound source are dynamically calculated, enabling the display control unit to fuse the coordinate information of multiple sound sources according to the weight ratio when calculating the digital human posture, avoiding posture mutations caused by a single strong signal sound source or posture jitters caused by multi-sound source competition, and thus realizing natural and smooth multi-target interaction coordination control.
[0050] Specifically, the rendering engine module adopts a collaborative rendering technology that combines dynamic skeleton driving and ambient light adaptation to build an intelligent visual presentation system oriented to the sound source direction. After receiving the orientation parameters output by the display control unit, the rendering engine first performs real-time redirection on the skeleton system of the digital human three-dimensional model, and dynamically adjusts the joint angles of the head, torso, and limbs by optimizing the IK (Inverse Kinematics) algorithm, so that the central axis of the torso is optimally aligned with the sound source direction vector. Different from the traditional method of calling a fixed pose library, the rendering engine of this embodiment integrates a physically based real-time shadow calculation module, which can dynamically generate ambient occlusion (AO) and soft shadow effects according to the relative position of the sound source and the digital human model. For example, when the sound source is located in the front left of the digital human, the ambient light on the right side of the model is automatically enhanced to simulate the real light and shadow logic. In addition, for high-frame-rate interaction scenarios, the rendering engine adopts a rendering strategy of enhancing the orientation area based on depth buffer to perform super-resolution reconstruction on the facial details in the sound source direction, improving the visual realism of the key interaction area while ensuring the overall rendering efficiency. The advantage of this technical solution is that it deeply integrates the spatial positioning parameters into the three-dimensional rendering pipeline, and through dynamic skeleton driving and intelligent light and shadow adaptation, it realizes the physical consistency between the digital human visual presentation and the sound source orientation, avoiding the limitation of the expressiveness of the traditional preset pose library in complex interaction scenarios.
[0051] Specifically, the speech processing unit constructs an intelligent interaction engine based on spatial audio features, realizing the deep fusion processing from speech signal acquisition to multi-modal response. This unit first performs time-frequency domain analysis on the sound source speech signal output by the microphone device, combines the distance parameter (Z-axis coordinate) provided by the infrared laser ranging sensor and the azimuth angle calculated by the TDOA algorithm to establish a spatial audio model including distance attenuation and air conduction delay, and performs pre-emphasis and reverberation compensation processing on the speech signal, solving the problem of the decrease in signal-to-noise ratio in far-field sound pickup in traditional speech recognition. In the semantic understanding stage, the speech processing unit introduces a context-aware attention mechanism, integrates the three-dimensional coordinate information of the sound source as a position embedding vector into the natural language processing (NLP) model, so that the generated interactive text content can perceive the spatial attributes of the sound source. For example, when the sound source is located directly behind the digital human, the interactive text will automatically include guiding statements such as "Please come in front of me". The speech synthesis link adopts a personalized voiceprint generation technology based on variational autoencoder (VAE), and dynamically adjusts the volume, speech rate, and spectral characteristics of the synthesized speech according to the distance and azimuth of the sound source, simulating the Doppler effect of sound propagation in the real space. This embodiment converts the three-dimensional space coordinates into the core parameters of speech processing, and through constructing a closed-loop processing link of spatial audio, semantics, and synthesis, realizes the dual adaptation of interactive speech to the sound source orientation in terms of acoustic characteristics and semantic content, improving the immersion and intelligence of multi-modal interaction.
[0052] In some embodiments, the face-to-face interaction includes dialogue interaction, and the interaction module is further configured to: calculate the horizontal azimuth angle and vertical pitch angle of the sound source object based on the three-dimensional coordinates; send a steering instruction to the display control unit to control the virtual head or body model of the digital human in the display interface to rotate to the direction corresponding to the azimuth angle and pitch angle; continuously collect the voice signal of the sound source object through the microphone device, convert it into text input through the voice recognition module, generate an interaction response text through the natural language processing model, and finally generate and output the response voice of the digital human through the voice synthesis module.
[0053] In some embodiments, the face-to-face interaction includes dialogue interaction, and the interaction module is further configured to perform the following specific operations:
[0054] First, calculate the horizontal azimuth angle and vertical pitch angle of the sound source object based on the three-dimensional coordinates. Specifically, the interaction module can obtain signal parameters such as the time difference or phase difference of the sound emitted by the sound source object through a microphone array, combine the preset three-dimensional space coordinate system, and use the principles of geometric acoustics and triangulation algorithms to calculate the horizontal azimuth angle and vertical pitch angle of the sound source object in this coordinate system. Among them, the horizontal azimuth angle takes the front of the digital human model in the display interface as the 0° reference, and the angle value rotated clockwise or counterclockwise represents the position of the sound source in the horizontal direction; the vertical pitch angle takes the horizontal line of sight of the digital human model as the 0° reference, and the angle value tilted upward or downward represents the position of the sound source in the vertical direction.
[0055] Second, send a steering instruction to the display control unit to control the virtual head or body model of the digital human in the display interface to rotate to the direction corresponding to the azimuth angle and pitch angle. After calculating the horizontal azimuth angle and vertical pitch angle, the interaction module generates a steering instruction containing the angle information according to the preset instruction format and sends it to the display control unit through the communication interface of the system. After receiving the steering instruction, the display control unit converts the horizontal azimuth angle and vertical pitch angle into the rotation angle and rotation speed of the virtual head or body model around the corresponding axis (for example, the horizontal azimuth angle corresponds to rotation around the vertical axis, and the vertical pitch angle corresponds to rotation around the horizontal axis) according to the kinematic parameters of the digital human model, and then controls the model to rotate in the specified direction and rate, so that the virtual head or body model of the digital human accurately faces the direction of the sound source object.
[0056] Furthermore, the voice signal of the sound source object is continuously collected by the microphone device and converted into text input through the voice recognition module. The microphone device receives the voice sound waves emitted by the sound source object in real time, converts them into electrical signals and transmits them to the voice recognition module. The voice recognition module uses voice recognition algorithms based on deep learning, such as deep neural network (DNN), recurrent neural network (RNN), or convolutional neural network (CNN), etc., to preprocess the input voice signal, extract features (such as Mel Frequency Cepstral Coefficients MFCC), and match and recognize the acoustic model and language model, convert the continuous voice signal into the corresponding text sequence, and perform noise reduction, error correction, etc. on the text to generate accurate text input.
[0057] Then, the interactive response text is generated through the natural language processing model. After receiving the text input output by the voice recognition module, the natural language processing module first performs preprocessing such as word segmentation, part-of-speech tagging, and syntactic analysis of the text to understand the grammatical structure and semantic information of the text. Then, using a preset dialogue model (such as a rule-based dialogue system, a statistical machine learning model, or an end-to-end dialogue model based on deep learning), combined with the preset role, dialogue scenario, and context information of the digital human, analyze the user's intentions and needs, and generate an interactive response text that conforms to the dialogue logic and context. During the generation process, sentiment analysis and style adjustment can also be performed on the response text to make the response more natural and appropriate.
[0058] Finally, the response voice of the digital human is generated by the speech synthesis module and output. After the speech synthesis module obtains the interactive response text generated by the natural language processing model, it first performs text normalization processing on the text, converting numbers, abbreviations, etc. into the corresponding pronunciation forms. Then, using speech synthesis technology (such as the STRAIGHT algorithm based on parametric synthesis, the unit selection method based on concatenative synthesis, or the neural speech synthesis model based on deep learning), according to the preset voice characteristics of the digital human (such as timbre, pitch, speech rate, etc.), convert the text into the corresponding voice signal. The generated voice signal is processed through audio processing (such as noise reduction, gain adjustment, etc.) and then played through a speaker or other audio output device to realize the voice response of the digital human to the sound source object.
[0059] In some embodiments, the digital human device further includes an image acquisition component, which is used to determine the orientation of the image acquisition component according to the three-dimensional coordinates of the sound source, so as to ensure that the image acquisition component and the microphone device synchronously face the sound source direction for image acquisition and sound localization.
[0060] In some embodiments, the interactive module is connected to the signal processing module of the microphone device through a software interface to obtain the three-dimensional coordinate data of the sound source in real time.
[0061] This embodiment enhances the environmental perception and response capabilities of the digital human by integrating a high-precision sound source localization and intelligent interaction system. Its advantages are as follows: It adopts a multi-modal data fusion mechanism to achieve sub-centimeter spatial positioning accuracy, and combines a real-time pose solution algorithm to ensure that the virtual image maintains natural eye contact and body orientation with the sound source object; Through a hardware-level synchronous acquisition and hierarchical signal processing architecture, it effectively suppresses multi-path interference in complex acoustic environments and can achieve an end-to-end response delay of less than 200 ms in dynamic scenarios; The modular design takes into account both system scalability and energy consumption optimization, supports multi-object tracking and continuous interaction sessions, and provides a reliable spatial perception foundation for immersive human-computer interaction.
[0062] Embodiment Three
[0063] This embodiment provides a sound source localization method based on the microphone device described in Embodiment One, which includes:
[0064] Synchronously collect sound signals through a pickup microphone and synchronously collect distance signals through the infrared laser range finder sensor;
[0065] Process the sound signals based on the TDOA algorithm to obtain the two-dimensional plane coordinates of the sound source;
[0066] Determine the three-dimensional coordinates of the sound source according to the two-dimensional plane coordinates and the distance signals.
[0067] In some embodiments, the synchronous collection of sound signals and distance signals specifically includes:
[0068] Use a high-precision clock source to provide a unified clock reference for all sensors to ensure that the timestamp synchronization error between the sound signal and the distance signal is less than a preset threshold;
[0069] Collect sound signals at a preset sampling rate, read the distance signals of the infrared laser range finder sensor at a preset frequency, and add timestamps containing time information to both types of signals.
[0070] In some embodiments, the processing of sound signals based on the TDOA algorithm specifically includes:
[0071] Perform band-pass filtering and noise reduction preprocessing on the collected sound signals;
[0072] Use the generalized cross-correlation algorithm to calculate the time difference of the sound signal arriving at different pickup microphones;
[0073] According to the geometric parameters of the microphone array, convert each time difference into a distance difference and construct a hyperbolic equation system, and use a non-linear optimization algorithm to solve the equation system to obtain the two-dimensional plane coordinates (x, y) of the sound source.
[0074] In some embodiments, determining the three-dimensional coordinates based on the two-dimensional plane coordinates and the distance signal specifically includes:
[0075] Establish a position mapping table for each basic unit in the ranging microphone array matrix, and store the central coordinates and unique numbers of each basic unit;
[0076] Calculate the Euclidean distance between the two-dimensional plane coordinates (x, y) and the central coordinates of each basic unit, and select the basic unit with the smallest distance as the target basic unit;
[0077] Retrieve the real-time distance signal z of the infrared laser ranging sensor in the target basic unit, and combine the two-dimensional coordinates to form the three-dimensional coordinates (x, y, z) of the sound source.
[0078] In some embodiments, it further includes a filtering process for the three-dimensional coordinates:
[0079] Perform filtering on the three-dimensional coordinate sequences of multiple consecutive frames to eliminate random measurement noise;
[0080] Calculate the motion trajectory of the sound source based on the filtered three-dimensional coordinates. When the trajectory change rate exceeds the preset threshold, trigger the dynamic calibration program of the positioning system.
[0081] In some embodiments, when constructing the hyperbolic equation system, at least the measurement data of four basic units are used, and the hyperbolic equation corresponding to each basic unit is:
[0082]
[0083] In the formula, (x i , y i ) and (x j , y j ) are the central coordinates of any two basic units, Δt ij is the time difference for the sound signal to reach the pickup microphone in the corresponding basic unit, and c is the sound propagation speed.
[0084] In some embodiments, during the process of retrieving the distance signal of the target basic unit, a signal validity verification mechanism is adopted:
[0085] When the fluctuations of consecutive multiple distance measurement values of the same basic unit exceed the preset ratio, trigger the sensor self-check program;
[0086] If the self-check finds that the ranging error exceeds the preset range, automatically switch to the distance signal of the adjacent basic unit for fusion calculation.
[0087] Embodiment 4
[0088] Such as Figure 2As shown, the hollow circles are the smallest units of the array microphone device, with 6 in a circle and evenly spaced in a ring. The solid circles are infrared laser ranging sensors for measuring distances outward from the vertical screen. Six microphones are fixedly installed around one infrared laser ranging sensor in the middle, forming an integral device containing 7 sensors. The one in the middle is an infrared laser ranging sensor, and the surrounding microphones for picking up sound form a regular hexagon structure when closely arranged in a circle. This device is the smallest unit, and various ranging microphone array matrices can be arranged with this device. For example, a 10*10 square matrix of 100 matrices, or a strip matrix of 16*2 = 32 units.
[0089] Six microphones are installed around one infrared laser ranging sensor in the middle to form an integral device. The distances between the microphones are equal. When someone approaches, the approaching object is detected by the infrared laser ranging sensor. When a sound is emitted, the microphone samples the sound. The sound sampling and distance sampling are carried out synchronously. After taking multiple such devices, for example 100, and installing them in the style of a 10*10 microphone matrix, when someone stands in front of the microphone matrix and makes a sound, the sound source point can be detected through the sound source. The microphone array is based on the sound source localization method: the microphone array algorithm based on TDOA (Time Difference of Arrival). The TDOA algorithm is used to determine which microphone the sound source reaches first. In the case of the same sound wave, it is certain that the microphone closest to the sound source senses it first, and the microphone farthest from the sound source senses it last. Based on this principle, the distance between the sound source and which microphone is the closest can be determined. The time for the signal to be transmitted from the target (such as a mobile phone, sound source, etc.) to multiple receivers (such as a base station, microphone array, etc.) is different. By calculating the time difference (TDOA) between two receivers, it can be converted into a distance difference (Δd = c·Δt, where c is the signal propagation speed). Each time difference corresponds to a hyperbola (two-dimensional) or a hyperboloid (three-dimensional), and the intersection point of multiple hyperbolas / hyperboloids is the target position. Assuming that the time difference between the signals received by receivers A(x1, y1) and B(x2, y2) is Δt, the distance difference is:
[0090] At least 3 receivers (two-dimensional) or 4 receivers (three-dimensional) are required to determine the unique intersection point.
[0091] The algorithm steps are as follows: By analyzing the signal correlation (such as the cross-correlation algorithm), accurately calculate the time difference Δt of the signal arriving at different receivers. Convert each Δt into a distance difference equation and construct a system of nonlinear equations. Use numerical methods (such as the least squares method, Chan algorithm, Taylor series expansion method) to solve the system of equations to obtain the target coordinates (x, y) or (x, y, z). Its advantage is that it does not require the target and the receiver to be time-synchronized, only strict synchronization between receivers (such as through GPS or wired connection) is required, reducing the system complexity. Compared with positioning based on signal strength (RSSI), TDOA is less affected by environmental attenuation. The time difference measurement needs to reach the nanosecond level (for example, a 1ns error results in an approximate distance error of about 0.3 meters).
[0092] The sound source position obtained by the microphone array algorithm is further used to obtain the physical distance of the sound source here by using an infrared laser ranging sensor. After determining the two-dimensional coordinates in the plane relative to the microphone array, the left-right, up-down positions of the person are determined according to the two-dimensional coordinates, thereby determining which microphone component the person is in, such as the microphone component numbered 5. At this time, by retrieving the infrared laser sensor of the No. 5 microphone component, the vertical distance between the person here and the matrix wall can be obtained. From the above two values (the two-dimensional plane coordinates and the vertical distance between the sound source and the microphone matrix), the three-dimensional positioning of the sound source person or object in front of the microphone matrix is determined.
[0093] In the embodiment of the present invention, while using the microphone to pick up sound, the distance of the object in front of the sound source of the microphone is obtained, providing a basic support for the three-dimensional space positioning of the sound source.
[0094] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be repeated here.
[0095] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements any one of the above methods.
[0096] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiments of the method, the present invention can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. Of course, there are other ways of readable storage media, such as quantum memory, graphene memory, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0097] The present invention also provides an electronic device. The electronic device according to an embodiment of the present invention includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided by the present invention.
[0098] Next, refer to Figure 4 , which shows a schematic structural diagram of a computer system 400 suitable for implementing the electronic device according to an embodiment of the present invention. Figure 4 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0099] As Figure 4 shown, the computer system 400 includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage section 408 into the random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the computer system 400 are also stored. The CPU 401, ROM 402, and RAM 403 are connected to each other through a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.
[0100] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as required. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as required so that a computer program read therefrom is installed into the storage section 408 as required.
[0101] Specifically, according to the embodiments disclosed in the present invention, the process described in the above main step diagram can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the method shown in the main step diagram. In the above embodiment, the computer program can be downloaded and installed from a network through the communication section 409, and / or installed from the removable medium 411. When the computer program is executed by the central processing unit 801, the above functions defined in the system of the present invention are executed.
[0102] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0103] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions. The units described in the embodiments of the present invention can be implemented in software or in hardware. The described units can also be provided in a processor.
[0104] The above specific embodiments do not limit the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A microphone device capable of locating a sound source, characterized in that, Including: A plurality of basic units, each of the basic units includes a plurality of pick-up microphones arranged in a ring at equal intervals, and an infrared laser ranging sensor located at the center of the pick-up microphones. The plurality of pick-up microphones are fixedly installed around the infrared laser ranging sensor; the plurality of basic units are arranged to form a ranging microphone array matrix. A signal processing module, configured to synchronously collect the sound signals of the pick-up microphones and the distance signals of the infrared laser ranging sensor, and perform sound source localization based on the time difference of arrival (TDOA) algorithm to obtain the two-dimensional plane coordinates of the sound source; determine the three-dimensional coordinates of the sound source according to the two-dimensional plane coordinates and the distance signals of the infrared laser ranging sensor.
2. The microphone device according to claim 1, wherein Each of the basic units contains 6 pick-up microphones, which are arranged in a ring at equal intervals to form a regular hexagon structure, and the infrared laser ranging sensor is located at the center of the regular hexagon.
3. The microphone device according to claim 1, characterized in that, The intervals between the pick-up microphones are equal, and the emission direction of the infrared laser ranging sensor is perpendicular to the plane where the pick-up microphones are located and is set outward.
4. The microphone device according to claim 1, characterized in that, The TDOA algorithm includes the following steps: calculating the time difference Δt of the sound signal arriving at different pick-up microphones through signal correlation analysis; converting the time difference into a distance difference according to the distance difference formula Δd = c·Δt, where c is the sound propagation speed; converting each of the distance differences into a hyperbolic equation, and constructing a non-linear equation set containing a plurality of hyperbolic equations based on the receivers formed by at least 3 of the basic units; using a numerical method to solve the non-linear equation set to obtain the two-dimensional plane coordinates of the sound source; the numerical method includes: the least squares method, the Chan algorithm or the Taylor series expansion method.
5. The microphone device according to claim 4, wherein The signal processing module is specifically configured to determine the corresponding target basic unit according to the two-dimensional plane coordinates, retrieve the distance signal of the infrared laser ranging sensor in the target basic unit, and calculate the vertical distance between the sound source and the ranging microphone array matrix in combination with the two-dimensional plane coordinates to form the three-dimensional coordinates of the sound source.
6. The microphone device according to claim 5, characterized in that, The signal processing module determines the corresponding target basic unit according to the two-dimensional plane coordinates and retrieves the distance signal of the infrared laser ranging sensor in the target basic unit, specifically including: determining the target basic unit number where the sound source is located based on the position mapping relationship between the two-dimensional plane coordinates and each basic unit in the ranging microphone array matrix; retrieving the distance signal of the infrared laser ranging sensor corresponding to the target basic unit number.
7. The microphone device according to claim 1, wherein, Each basic unit of the ranging microphone array matrix realizes time synchronization through a synchronous method. The pick-up microphones and the infrared laser ranging sensor are integrated on the same substrate, and the surface of the substrate is provided with a sound transmission structure for sound wave conduction and a light transmission window for infrared signal emission.
8. The microphone device according to claim 1, wherein, The signal processing module includes a preprocessing unit, a positioning calculation unit and a coordinate fusion unit. The preprocessing unit is configured to perform band-pass filtering denoising and timestamp synchronization calibration on the sound signal and the distance signal to obtain a calibrated sound signal sequence and a distance signal sequence. The positioning calculation unit is configured to execute the TDOA algorithm based on the calibrated sound signal sequence to calculate the two-dimensional plane coordinates of the sound source; The coordinate fusion unit is configured to generate the three-dimensional coordinates of the sound source according to the two-dimensional plane coordinates and the distance signal of the corresponding target basic unit in the calibrated distance signal sequence.
9. A digital human device, characterized in that, Comprising: The microphone device according to any one of claims 1-8, configured to determine the three-dimensional coordinates of the sound source; The interaction module is connected to the signal processing module of the microphone device and is configured to adjust the display posture of the digital human to the sound source direction according to the three-dimensional coordinates and perform face-to-face interaction with the sound source object.
10. A method for sound source localization of the microphone device capable of localizing a sound source according to claim 1, characterized in that, Comprising: The sound signal is synchronously collected by the pickup microphone, and the distance signal is synchronously collected by the infrared laser ranging sensor; The sound signal is processed based on the TDOA algorithm to obtain the two-dimensional plane coordinates of the sound source; The three-dimensional coordinates of the sound source are determined according to the two-dimensional plane coordinates and the distance signal.
Citation Information
Patent Citations
Phase-locked-amplifier-based sound location method under strong disturbance
CN103176167A
Location method, location system and terminal device
CN107643509A
Robot pickup method and device based on microphone array and medium
CN115767378A
Unmanned aerial vehicle positioning device and method based on Kalman filtering algorithm
CN116381717A
Sound source positioning device and trolley
CN216848142U
Cited By
Positioning method and related equipment
CN121208752A