An acoustic device autonomous mapping method for voice localization

By collaborating with a robotic vacuum cleaner and a smart voice device, the robot constructs an electronic map and locates the voice device based on its noise, solving the flexibility and accuracy issues of indoor sound source localization in existing technologies and achieving plug-and-play high-precision localization without the need for manual calibration.

CN116449292BActive Publication Date: 2026-06-02TIANJIN UNIV +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TIANJIN UNIV
Filing Date
2023-04-17
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies for indoor sound source localization require cumbersome manual calibration and prior knowledge, resulting in high costs and inflexibility, making it difficult to achieve plug-and-play functionality.

Method used

By collaborating with a robotic vacuum cleaner and a smart voice device, the robot constructs an indoor electronic map and locates the voice device based on its minimal noise. Combined with super-resolution sample offset measurement, inertial AoA estimation, and microphone array, the autonomous mapping and high-precision positioning of the voice device are achieved.

Benefits of technology

It requires no manual calibration, has flexible and high-precision indoor positioning capabilities, adapts to various scenario needs, reduces labor costs, and achieves plug-and-play functionality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116449292B_ABST
    Figure CN116449292B_ABST
Patent Text Reader

Abstract

The application discloses an acoustic device autonomous mapping method for voice positioning, and belongs to the technical field of wireless sensing; the application proposes an acoustic device autonomous mapping method for voice positioning, explores cooperation between a sweeping robot and a voice device, positions the spatial position of an indoor voice device in a map through a small noise generated during normal work of the sweeping robot based on an indoor electronic map constructed by the sweeping robot, thereby avoiding tedious manual calibration, and finally positions a sound source based on a voice map. Compared with previous partial acoustic positioning work, the application does not require any prior knowledge of an indoor space, has strong flexibility, and can better meet various requirements of users for indoor positioning scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless sensing technology, and more particularly to an acoustic device autonomous mapping method for voice localization. Background Technology

[0002] In recent years, with the rapid development and widespread adoption of the internet, the number and variety of smart devices have exploded. These diverse smart access devices have significantly improved our daily lives and provided a solid hardware foundation for applications such as smart homes and smart offices. The operation of smart devices requires the support of various input systems. Among them, voice input systems, as a low-cost and highly convenient remote command input control module, have been widely deployed in various devices. For example, smart appliances and smart voice assistants integrate voice modules, allowing us to remotely control these smart devices without additional input devices.

[0003] For smart service providers, context-aware services are a core competitive advantage in the current smart service market. Sound source location information is one of the most critical elements, and companies like Huawei, Xiaomi, Google, and Apple are all striving to achieve sound source localization to obtain effective contextual information. This is because location information can help resolve ambiguous issues. For example, in the smart home field, when a user says "turn on the lights," the system can determine which room's lights to turn on.

[0004] To achieve the aforementioned indoor sound source localization service, typical localization systems rely on prior knowledge of device location, device positioning, and an indoor floor plan. This requires not only the user to obtain an electronic map of the room but also to mark the spatial information of these devices on the map. This process incurs additional costs, especially when multiple smart voice devices are deployed indoors. To reduce these labor costs and enable plug-and-play applications, we envision an ideal smart home system: a robot that can automatically explore an indoor map and determine the spatial information of voice devices on the map, allowing the system to locate sound sources based on the aforementioned voice map.

[0005] Wireless sensing technology is a widely used technology in indoor positioning. Specifically, wireless sensing technologies include acoustic sensing, Wi-Fi sensing, Bluetooth, RFID, and computer vision. For the aforementioned construction of indoor scenes and indoor positioning blueprints based on smart devices, acoustic sensing is an ideal low-cost, high-quality solution. Acoustic sensing technology uses a series of acoustic processing methods to extract information such as the Angle of Arrival (AoA), Doppler Shift, Time of Flight (TOF), and signal attenuation contained in the sound signal, thereby sensing the surrounding environmental conditions and human activities. Depending on the type of sound signal used for sensing, traditional acoustic sensing technologies can be divided into modulation-based acoustic sensing and unknown signal-based acoustic sensing. For modulation-based acoustic sensing, the speaker device emits carefully modulated acoustic signals, such as continuous wave (CW) and frequency-modulated continuous wave (FMCW), allowing the microphone to extract channel state characteristics from the received known signals. Since the characteristics of the original signal are known, this method often achieves higher accuracy. For acoustic sensing based on unknown signals, microphone arrays receive random acoustic signals from sound sources, such as speech, music, footsteps, or even noise, and then extract useful information for pose perception. Because the system cannot control or predict such natural sounds whose frequencies and contents lack prior knowledge, it is difficult to extract environmental features using traditional channel estimation methods. The 2020 work VoLoc designed a time-domain-based method to estimate the AoA of a signal and combined it with the geometric parameters of nearby walls to estimate the user's position indoors, but this system has certain limitations on the shape of the room.

[0006] To address the aforementioned problems, this invention proposes an acoustic device autonomous mapping method for voice localization. Summary of the Invention

[0007] The purpose of this invention is to provide an acoustic device autonomous mapping method for voice positioning that can parse AoA information in unknown acoustic signals to complete the spatial location calibration of indoor voice devices and support high-precision indoor positioning, thereby solving the problems mentioned in the background art.

[0008] To achieve the above objectives, the present invention adopts the following technical solution:

[0009] An autonomous mapping method for acoustic devices for voice localization explores the cooperation between a robotic vacuum cleaner and a voice device. It utilizes an indoor electronic map constructed by the robotic vacuum cleaner and leverages the minimal noise generated during its normal operation to locate the spatial position of the voice device within the map, thus avoiding tedious manual calibration. Finally, the sound source is located based on the voice map. Specifically, the method includes the following:

[0010] S1. Super-resolution sample offset measurement: The frequency domain signal is acquired using the microphone of the intelligent voice device, interpolated in the time domain to complete the zero filling in the frequency domain, and then the sound signal is filtered using a filter. The filtered signal is then processed to calculate the fine-grained sample offset.

[0011] S2. Inertial-based AoA estimation: The actual trajectory sequence is collected using the microphone of the intelligent voice device, and the data is augmented to generate a training set. The AoA sequence is calculated based on the obtained training set. The super-resolution sample offset measurement method in S1 is combined to obtain a finer-grained sample offset, and then the AoA model is constructed.

[0012] S3, Microphone Array Localization: The robot with SLAM capability explores the layout of the room. Then, the microphone array in the smart voice device captures the sound of the robot during operation. Based on S2, the AoA of the sound signal is calculated. Complex numbers are used to represent the position and direction of the microphone array. The coordinate system of the SLAM map and AoA is synchronized, and then the spatial position of the microphone array is marked on the SLAM map.

[0013] S4. Room Structure and Positioning: Based on the content described in S3, mark the position of the microphone array on the electronic map to obtain accurate room information. Then, use the microphone array to locate the human voice, use wall reflections and multiple AoAs to determine the location of the target, and use geometric relationships to complete the human voice localization.

[0014] Preferably, S1 specifically includes the following:

[0015] S1.1. Obtain a received sequence of length N using a microphone, denoted as... h r ( t Perform a Fourier transform on the received sequence to obtain the frequency domain signal, specifically as follows:

[0016] H r ( t )= F ( h r ( t ))

[0017] in, F(·) represents the Fourier transform;

[0018] S1.2. Zero-padding is performed on the frequency domain obtained in S1.1, using a length of... K × N The inverse Fourier transform yields a new time-domain sequence, specifically represented as:

[0019] h’ r ( t )= F -1 ( H r ( t ), KN )

[0020] in, F -1 (·) indicates the inverse Fourier transform;

[0021] S1.3. Use a Hampel filter to filter the audio signal, eliminate high-frequency jumps, and then perform cross-correlation on the filtered signal to calculate the fine-grained sample offset.

[0022] Preferably, S2 specifically includes the following:

[0023] S2.1. Using a microphone to collect sound signals, rotate, scale, and translate a small sequence of actual trajectory signals to generate a new motion trajectory sequence without changing the shape of the original trajectory; represent the trajectory coordinates in a complex coordinate system, assuming the coordinates at time t are: l t = x t + you t

[0024] in, j The complex unit is represented by ; by transforming the above equation according to Euler's formula, the coordinates can be expressed as:

[0025] l t = r t e j t = r t cos ( t )+ r t sin ( t )= x t + you t

[0026] S2.2 Calculate AoA through geometry. In the calculation process, it is assumed that the microphone position is fixed and the candidate cases are obtained by simulating the robot's trajectory through rotation, translation, and scaling.

[0027] S2.3, Assuming the velocity of the sound signal is c, and the sound sampling rate is... f s Then the sample offset can be expressed as

[0028]

[0029] in, S off Indicates sample offset;

[0030] S2.4 Construct an unsupervised learning neural network AE embedding AoA estimation model, and use the results obtained in S2.3. S off Used as input values ​​for the model, and then trained to make predictions. Close to input value S off Using mathematical models Replacement decoder guides model input values S off Oriented rotation ensures that the latent features provided by the AE are AoAs; based on the above operations, the target encoder will approximate the inverse function of the decoder, thereby converting the sample offset into AoAs.

[0031] Preferably, S3 specifically includes the following:

[0032] S3.1 Layout Exploration and Sound Capture: The layout of the room is explored using a robot with SLAM capabilities. Then, the microphone array in the smart voice device captures the sound of the robot during operation, and the AoA of the sound signal is calculated based on S2.

[0033] S3.2 Array Representation: The position and orientation of the microphone array are represented by complex numbers, denoted as [formulas to be inserted here]. Q = x + you and The robot's position sequence is represented as ,in, l t = x t + you t, representing the position at time t; assuming the AoA sequence calculated in S3.1 is ,pass and calculate Q and ;

[0034] S3.3 Sequence Synchronization: The acoustic signal is sliced ​​with a fixed period t, then the average energy is detected to determine the start stage, human voice interference is eliminated by frequency detection, and finally the alignment of two sequence samples is completed by interpolation.

[0035] S3.4 Coordinate System Synchronization: Assume the robot reports the position as follows: p t The AoA estimate derived from the microphone array is i t Based on the content described in S3.1 to S3.3, it can be seen that ( Q , p t , It belongs to the robot coordinate system. i t It belongs to the microphone coordinate system; therefore, it can be obtained through a position-based representation. p t -Q , can be obtained through angle-based representation Considering unit vectors, then:

[0036]

[0037] Further transformation into:

[0038]

[0039] Iterate through the search space Q to maximize the average value of the equation within the T window:

[0040]

[0041] After determining the location of the microphone array, its direction is calculated using the following formula:

[0042]

[0043] Preferably, S4 specifically includes the following:

[0044] S4.1 For an array with N microphones, its original input signal is an N×M matrix, where M represents the number of samples;

[0045] S4.2. Zero-fill the signal and then use the short-time Fourier transform of the time-domain signal to obtain the time-frequency diagram;

[0046] S4.3. Use differential cancellation to eliminate interference at low frequencies to obtain a clearer waveform, and then traverse each time slot to determine the largest frequency component to obtain the channel frequency response of different microphones.

[0047] S4.4. Solve for AoA using the MUSIC algorithm; for the obtained AoA, use clustering methods to determine the direct path and reflection path, and use geometric relationships to complete the human voice localization.

[0048] Compared with the prior art, the present invention provides an acoustic device autonomous mapping method for voice localization, which has the following beneficial effects:

[0049] (1) This invention proposes an acoustic device autonomous mapping method for voice localization. It is based on the cooperation between intelligent robots and intelligent voice devices. The robot with SLAM function constructs an indoor electronic map and uses the small noise generated when it is working normally to locate the spatial position of the indoor voice device in the map, thereby avoiding tedious manual calibration. Finally, the voice map is used to locate the sound source. It is a voice map based on microphone array and intelligent robot. Compared with some previous acoustic localization work, this invention does not require any prior knowledge of the indoor space, has great flexibility, and can better meet the various needs of users for indoor positioning scenarios.

[0050] (2) This invention proposes an inertial-based super-resolution AoA estimation method that combines target motion and AoA estimation to improve the accuracy of moving target tracking. In addition, this method only requires the intelligent voice device to capture the small noises in the daily operation of the sweeping robot, without the need to send specially designed acoustic signals separately.

[0051] (3) The present invention develops an effective method to synchronize the coordinate systems of moving objects and acoustic devices, thereby enabling the system to achieve autonomous mapping of microphone arrays.

[0052] (4) The present invention designed a large number of experiments in multiple scenarios to verify the system's capabilities. The present invention used a commercial SLAM sweeping robot and microphone array to build a prototype and achieved high-precision microphone positioning effect. Attached Figure Description

[0053] Figure 1 This is a flowchart of an acoustic device autonomous mapping method for voice localization proposed in this invention;

[0054] Figure 2 This is a schematic diagram of the microphone array model in Embodiment 1 of the present invention;

[0055] Figure 3This is a schematic diagram of the neural network model in Embodiment 1 of the present invention;

[0056] Figure 4 This is a schematic diagram of coordinate transformation in Embodiment 1 of the present invention;

[0057] Figure 5 This is a schematic diagram of the human voice localization principle in Embodiment 1 of the present invention;

[0058] Figure 6 This is a flowchart of human voice localization in Embodiment 1 of the present invention. Detailed Implementation

[0059] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0060] First, the technical terms used in this invention will be explained:

[0061] Wireless sensing is a technology that utilizes ubiquitous wireless sensor signals in the environment to obtain channel state information, thereby sensing the surrounding environment and human activities. It is one of the most critical technologies in the Internet of Things (IoT) field. A wide variety of signals can be used for wireless sensing, such as Wi-Fi, sound, light signals, Bluetooth, and RFID.

[0062] Acoustic sensing technology is an important component of wireless sensing, and it is a technology that uses acoustic signals to perceive the environment. Specifically, it analyzes the changes in sound signals during propagation (such as direction, speed, delay, and energy) to determine the state and characteristics of the target being sensed.

[0063] Indoor positioning and tracking technology is one of the commonly used solutions in smart furniture. Because GPS signals are weak indoors, leading to significant positioning drift, wireless sensing technology is often used to meet indoor positioning needs.

[0064] The angle of arrival (AoA) is the direction from which a signal wave originates. In acoustic sensing, its purpose is to determine the direction from which a sound signal comes, and then, in turn, deduce the azimuth angle of the target sound source. In practice, we typically use microphone arrays to capture the signal angle of arrival; common arrays include 6-microphone circular arrays and 4-microphone linear arrays. Due to the far-field effect, the sound arriving at a microphone can be considered as a series of parallel lines. Therefore, the signals received by different microphones can be seen as versions of the original signal with different delays. The different delay levels of the received signals can help us determine the direction of arrival.

[0065] Coordinate system synchronization involves transforming the camera coordinate systems of different devices into a unified world coordinate system. In this invention, the camera coordinate systems of the voice device and the robot vacuum cleaner are synchronized and uniformly transformed into the world coordinate system of the map.

[0066] Angular resolution (AoA) refers to a system's ability to distinguish the direction of arrival of a signal. In a microphone array within a smart voice device, the speech signals received by each microphone have different delays. Since the received signal is a discrete sample of the sound wave from the microphone, we cannot obtain the accurate delay time in the time domain. Integer sample point offset means the system cannot distinguish delays smaller than one sampling time, further leading to multiple angles of arrival potentially corresponding to the same sample shift, causing AoA estimation errors.

[0067] The technical solution of the present invention will be described in detail below with reference to specific examples and accompanying drawings.

[0068] Example 1:

[0069] Without loss of generality, within a 6.4m × 6.4m open space, this invention constructed a prototype voice localization system using a commercially available SLAM robotic vacuum cleaner and a microphone array. The SLAM robotic vacuum cleaner can report its location and provide indoor electronic mapping services, functions already supported by some commercial robotic vacuum cleaners. To simulate common smart voice devices, we used a Seeed Studio 4-microphone linear array and a 6-microphone circular array to capture the operating noise of the SLAM robotic vacuum cleaner; their shapes are similar to Amazon's Echo and Alibaba's Tmall Genie. Each microphone array was connected to a Raspberry Pi 4B for control. For both microphone arrays, the distance between adjacent microphones was 4.75cm, with the 6 microphones in the circular array arranged in a regular hexagonal pattern. In our prototype system, the voice device collected sound at a sampling rate of 48 kHz. After data collection, we used a desktop computer to execute code written in Matlab. Based on the above voice localization system prototype, this invention proposes an acoustic device autonomous mapping method for voice localization, specifically including the following:

[0070] ① Super-resolution sample offset measurement section:

[0071] One sub-scenario of this invention is using the AoA of an audio signal to locate a continuously moving device. Under far-field effects, the signal received by the microphone array can be considered as a series of parallel lines. Ignoring multipath effects, the signal reaching the [missing information] will [missing information]. i The acoustic signals of each microphone are denoted as s i ( tAccording to geometric relationships, there is an additional flight distance for the signal to reach different microphone pairs. Assume the first... i The location of the microphone is , will the i The and the first j The spatial vector between the microphones is represented as follows: Therefore, the additional flight distance is:

[0072]

[0073] in, for i The unit vector in the direction, its principle is as follows: Figure 2 As shown. In the time domain, this distance is reflected as a time shift. Therefore, we can obtain ,in c Indicates the speed of sound.

[0074] However, due to the limitation of the sampling rate of the acoustic signal (48 kHz), the sample offset must be an integer, and the time-domain-based method is insufficient to distinguish different AoAs. Therefore, this invention proposes a new time-frequency analysis method.

[0075] This invention first performs zero-padding in the frequency domain of the signal, which is equivalent to interpolation in the time domain. Then, we use a special filter to filter the audio signal. Specifically, for a received sequence of length N... h r ( t First, its frequency domain signal is obtained, and then... H r ( t )= F ( h r ( t )) indicates that among them F (·) denotes the Fourier transform. Next, zero-padding is performed in the frequency domain, using a length of... K × N The inverse Fourier transform yields a new time-domain sequence, namely h’ r ( t )= F -1 ( H r ( t ), KN ),in, F -1(·) denotes the inverse Fourier transform. The above two steps are equivalent to interpolation in the time domain. Then, a cross-correlation operation is performed to calculate the sample offset. To obtain a finer-grained sample offset, this invention proposes a simple but more efficient solution: the Hampel filter. This will result in a new peak, which originates from redundant frequency components in the interpolation process, reflected as tiny ripples in the interpolation sequence. For this purpose, we use the Hampel filter to eliminate high-frequency jumps. Therefore, we can first use the Hampel and smoothing filters, and then cross-correlate the filtered signal to calculate the fine-grained sample offset.

[0076] ② Inertia-based AoA estimation part:

[0077] In practice, multipath signals (such as echoes) affect our estimation of AoA, but based on the fact that motion is continuous, the AoA of continuous signal segments does not change significantly. In other words, there is an implicit relationship between Ground Truth and the observed results. To address this, we designed a data-driven method based on neural networks.

[0078] First, to avoid the labor costs of manual data collection, we adopted an efficient method to generate the training set. Specifically, we generated new motion trajectory sequences by rotating, scaling, and translating a small number of actual trajectory sequences without changing the shape of the original trajectories, thus achieving data augmentation. We represent the trajectory coordinates in a complex coordinate system, assuming... t The coordinates of the time are l t = x t + you t ,in j To represent the complex unit, according to Euler's formula, we have:

[0079] l t = r t e j t = r t cos ( t )+ r t sin ( t )= x t + yout

[0080] By performing basic transformation operations, a new trajectory sequence can be obtained.

[0081] Next, the AoA sequence is calculated using geometry. Due to the relativity of positional relationships, even if we assume the microphone position is fixed, we can simulate all candidate cases by rotating, translating, and scaling the robot's trajectory.

[0082] For an AoA sequence, we need to calculate the sample offset. Assume the velocity of the sound signal is c, and the sound sampling rate is... f s Then the sample offset can be expressed as:

[0083]

[0084] in, S off Indicates sample offset.

[0085] Using the super-resolution sample offset measurement method described above, we can obtain sample offsets with finer granularity.

[0086] Finally, an unsupervised learning neural network AE is proposed to embed the AoA estimation model. For example... Figure 3 As shown, the predicted value is improved through training. Close to input value S off This allows us to obtain the encoder and decoder. To ensure that the latent features provided by AE are AoAs (... This invention uses a mathematical model. An alternative decoder is used to guide it to perform the directional transformation. Therefore, the target encoder will approximate the inverse function of this decoder, thus converting the sample's offset to AoAs.

[0087] ③ Microphone array positioning section:

[0088] This invention uses complex numbers to represent the position and orientation of the microphone array, namely and . Q = x + you and Similarly, the robot's position sequence is represented as... , where represents the position at time t. Assume the AoA sequence derived from the above steps is . ,pass and calculate Q and .

[0089] Sequence Synchronization: Due to the different sampling rates of the robot and the microphone, we need to synchronize the start of the sequence and each sample. We slice the acoustic signal in 0.01s windows, then detect the average energy to determine the initial stage and eliminate human speech interference through frequency detection. Next, we use an interpolation method to align the samples of the two sequences. In practice, the microphone array performs uniform sampling at 48 kHz, while the time intervals of the robot's reported position sequences are non-uniform. Therefore, based on the continuity of motion, we perform spline interpolation on the robot's position sequences to make the time interval between adjacent trajectory points 0.01 seconds.

[0090] Coordinate system synchronization: Since the robot and microphone use different coordinate systems, we need to synchronize them. Assume the robot reports a position of... p t The AoA estimate derived from the microphone array is i t ,So( Q , p t , It belongs to the robot coordinate system. i t It belongs to the microphone coordinate system. Therefore, it can be obtained through position-based representation. p t -Q This can be obtained through angle-based representation. Considering unit vectors, then:

[0091]

[0092] Further transformation into:

[0093]

[0094] Iterate through the search space Q to maximize the average value of the equation within the T window:

[0095]

[0096] After determining the location of the microphone array, its direction is calculated using the following formula:

[0097]

[0098] ④ Room structure and location:

[0099] After marking the microphone array's location on an electronic map, this array can be used to locate a person's voice. We use wall reflections and multiple AoAs to determine the target's location; the specific principle is as follows... Figure 5 .

[0100] For an array with N microphones, the original input signal is an N×M matrix, where M represents the number of samples. To improve frequency domain resolution, after zero-padding the signal, we use the Short Time Fourier Transform (STFT) of the time-domain signal to obtain a time-frequency map. Then, differential processing is used to eliminate low-frequency interference to obtain a clearer waveform. Next, we iterate through each time slot to determine the largest frequency component to obtain the Channel Frequency Response (CFR) for different microphones. Finally, the MUSIC algorithm is used to solve for the AoA (Aspect-Oriented Allocation). For the final AoA, we use clustering to determine the direct and reflection paths and utilize geometric relationships to complete voice localization. The workflow is detailed below. Figure 6 .

[0101] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. An acoustic device autonomous mapping method for speech localization, characterized in that, Specifically, it includes the following: S1. Super-resolution sample offset measurement: The frequency domain signal is acquired using the microphone array of the intelligent voice device, interpolated in the time domain to complete the zero filling in the frequency domain, and then the sound signal is filtered using a filter. The filtered signal is then processed to calculate the fine-grained sample offset. S2. Inertia-based AoA estimation: Obtain the actual trajectory sequence of the robot with SLAM function, generate a training set, calculate the AoA sequence based on the obtained training set, and combine the super-resolution sample offset measurement method in S1 to obtain a finer-grained sample offset, and then construct the AoA model. S3, Microphone Array Localization: The robot with SLAM capability explores the layout of the room. Then, the microphone array in the smart voice device captures the sound of the robot during operation. Based on S2, the AoA of the sound signal is calculated. Complex numbers are used to represent the position and direction of the microphone array. The coordinate system of the SLAM map and AoA is synchronized, and then the spatial position of the microphone array is marked on the SLAM map. S4. Room Structure and Positioning: Based on S3, the microphone array position is marked on the electronic map to obtain accurate room information. Then, the microphone array is used to locate the human voice. The wall reflection and multiple AoAs are used to determine the target's location, and geometric relationships are used to complete the human voice localization. S1 specifically includes the following: S1.

1. Obtain a received sequence of length N using a microphone, denoted as... h r ( t Perform a Fourier transform on the received sequence to obtain the frequency domain signal, specifically as follows: H r ( t )= F ( h r ( t )) in, F (·) represents the Fourier transform; S1.

2. Zero-padding is performed on the frequency domain obtained in S1.1, using a length of... K × N The inverse Fourier transform yields a new time-domain sequence, specifically represented as: h’ r ( t )= F -1 ( H r ( t ), KN ) in, F -1 (·) indicates the inverse Fourier transform; S1.

3. Use a Hampel filter to filter the audio signal, eliminate high-frequency jumps, and then perform cross-correlation on the filtered signal to calculate the fine-grained sample offset. S2 specifically includes the following: S2.

1. Using a microphone array to collect sound signals, rotate, scale, and translate the acquired actual trajectory sequence of the robot with SLAM functionality to generate a new motion trajectory sequence without changing the shape of the original trajectory; represent the trajectory coordinates in a complex coordinate system, assuming the coordinates at time t are: l t = x t + jy t in, j The complex unit is represented by ; by transforming the above equation according to Euler's formula, the coordinates can be expressed as: l t = ρ t e j t = ρ t cos ( t )+ ρ t sin ( t )= x t + jy t S2.2 Calculate AoA through geometry. In the calculation process, it is assumed that the microphone position is fixed and the candidate cases are obtained by simulating the robot's trajectory through rotation, translation, and scaling. S2.3, Assuming the velocity of the sound signal is c, and the sound sampling rate is... f s Then the sample offset can be expressed as in, S off Indicates sample offset; S2.4 Construct an unsupervised learning neural network AE embedding AoA estimation model, and use the results obtained in S2.

3. S off Used as input values ​​for the model, and then trained to make predictions. Close to input value S off Using mathematical models Replacement decoder guides model input values S off The targeted transformation ensures that the latent features provided by the AE are AoAs; based on the above operations, the target encoder will approximate the inverse function of the decoder, thereby converting the sample offset into AoAs; S3 specifically includes the following: S3.1 Layout Exploration and Sound Capture: The layout of the room is explored using a robot with SLAM capabilities. Then, the microphone array in the smart voice device captures the sound of the robot during operation, and the AoA of the sound signal is calculated based on S2. S3.2 Array Representation: The position and orientation of the microphone array are represented by complex numbers, denoted as [formulas to be inserted here]. Q = x + jy and The robot's position sequence is represented as ,in, l t = x t + jy t , l t Let represent the position at time t; assuming the AoA sequence calculated in S3.1 is . ,pass and calculate Q and ; S3.3 Sequence Synchronization: The acoustic signal is sliced ​​with a fixed period window, then the average energy is detected to determine the start stage, human voice interference is eliminated by frequency detection, and finally the alignment of two sequence samples is completed by interpolation. S3.4 Coordinate System Synchronization: Assume the robot reports the position as follows: p t The AoA estimate derived from the microphone array is θ t Combining S3.1 to S3.3, it can be seen that ( Q , p t , It belongs to the robot coordinate system. θ t It belongs to the microphone coordinate system; therefore, it can be obtained through a position-based representation. p t -Q This can be obtained through angle-based representation. Considering unit vectors, then: Further transformation into: Iterate through the search space Q to maximize the average value of the equation within the T window: After determining the location of the microphone array, its direction is calculated using the following formula: 。 2. The acoustic device autonomous mapping method for speech localization according to claim 1, characterized in that, S4 specifically includes the following: S4.1 For an array with N microphones, its original input signal is an N×M matrix, where M represents the number of samples; S4.

2. Zero-fill the signal and then use the short-time Fourier transform of the time-domain signal to obtain the time-frequency diagram; S4.

3. Use differential cancellation to eliminate interference at low frequencies to obtain a clearer waveform, and then traverse each time slot to determine the largest frequency component to obtain the channel frequency response of different microphones. S4.

4. Solve for AoA using the MUSIC algorithm; for the obtained AoA, use clustering methods to determine the direct path and reflection path, and use geometric relationships to complete the human voice localization.