An Internet of Things voice signal matching and recognition method

By constructing a wall acoustic characteristic database and dynamically adjusting the speech recognition model, the problems of speech feature changes and environmental noise interference in large indoor sports scenes are solved, and high-accuracy speech recognition and emergency rescue support are achieved.

CN119274544BActive Publication Date: 2025-07-08NANJING DINGSHAN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411326545.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-23
Publication Date
2025-07-08
Estimated Expiration
2044-09-23

AI Technical Summary

Technical Problem

In large indoor sports scenarios, changes in athlete's voice characteristics and environmental noise interference seriously affect the accuracy of speech recognition, and the prior art is difficult to effectively deal with.

Method used

Through acoustic modeling and simulation analysis, acoustic characteristics database of walls is constructed, speech data at different distances between the athlete and the wall is obtained, speech signal attenuation and interference models are established, recognition strategies are dynamically adjusted, and incremental learning optimization models are optimized.

Benefits of technology

It improves the accuracy of speech recognition, can quickly identify rescue instructions in complex environments, reduce accidental injuries, and provide emergency rescue support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119274544B_ABST
    Figure CN119274544B_ABST
Patent Text Reader

Abstract

The present application provides an Internet of Things voice signal matching and recognition method, including: analyzing voice features according to voice feature vectors in combination with the distance and angular position relationship between the mover and the wall, establishing an attenuation and interference model of the voice signal under wall occlusion, and obtaining the influence law of wall occlusion on the voice recognition accuracy rate; if the voice recognition result is that the mover issues a distress instruction, triggering an emergency rescue plan, sending an alarm message to the rescue center, and planning an optimal rescue route according to the position information of the mover to guide the rescue personnel to quickly reach the accident scene to minimize accidental injuries to the greatest extent; continuously optimizing the wall occlusion voice recognition model by using an incremental learning method, and dynamically updating the acoustic characteristic database and the voice attenuation and interference model according to the recognition accuracy rate and the false alarm rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to a method for matching and recognizing voice signals of an Internet of Things. Background Art

[0002] In large-scale indoor sports scenes, the speech characteristics of athletes will change significantly when they exercise vigorously. Rapid breathing and physical exertion cause changes in the fundamental frequency, resonance peak position, energy distribution, etc. of the speech pattern, making feature extraction and pattern matching of speech recognition more difficult. At the same time, the wall material of the sports venue also has complex reflection and interference effects on the sound, long reverberation time, and multi-path effects, which seriously affect the speech clarity and recognition accuracy. The dynamic range of speech signals varies greatly and the non-stationarity is strong. The coupling effect of the above factors makes robust speech recognition in noisy indoor sports environments face huge technical challenges. However, both the dynamic changes of speech features and the complexity of environmental noise make this process full of challenges. In other words, at present, in large-scale indoor sports scenes, how to deal with the changes in athlete's speech characteristics and the interference of environmental noise while ensuring the accuracy of speech recognition is a technical problem that needs to be solved urgently. Summary of the invention

[0003] The present invention provides a method for matching and recognizing voice signals of an Internet of Things, which mainly includes:

[0004] Acquire sound reflection data of different wall materials, thicknesses, shapes, surface roughness, internal structures, hole distributions, attachments, and temperatures. Through acoustic modeling and simulation analysis, determine the effects of various wall properties on speech signal frequency, amplitude, duration, energy, spectral characteristics, resonance frequency, echo characteristics, and signal-to-noise ratio, and form a database of wall acoustic characteristics.

[0005] According to the wall acoustic characteristics database, the speech data of the athlete at different distances from the wall during exercise is obtained as speech samples, and a speech feature data set is constructed to extract the frequency, amplitude, duration, energy, spectral characteristics, resonance frequency, echo characteristics and signal-to-noise ratio parameters of the speech signal to obtain the speech feature vectors at different distances from the wall.

[0006] According to the speech feature vector, combined with the distance and angle position relationship between the athlete and the wall, the speech features are analyzed, and the attenuation and interference model of the speech signal under the occlusion of the wall is established, and the influence of the wall occlusion on the accuracy of speech recognition is obtained;

[0007] Select corresponding acoustic parameters from the wall acoustic characteristic database according to the wall material, thickness, shape, surface roughness, internal structure, hole distribution, attachments, and temperature attributes, and combine them with the voice signal attenuation and interference model to construct a wall occlusion voice recognition model for accurately recognizing voice commands in a complex wall environment;

[0008] Preprocess the voice signal in the wall occlusion voice recognition model, extract the frequency, amplitude, duration, energy, spectral features, resonance frequency, echo characteristics, and signal-to-noise ratio characteristic parameters of the voice signal, and optimize the feature subset to reduce the interference of wall acoustic characteristics on voice recognition;

[0009] During the movement, real-time obtain the voice data and position information of the mover and input them into the wall occlusion voice recognition model. According to the distance and angular position relationship between the mover and the wall, dynamically adjust the acoustic parameters and voice feature weights of the wall occlusion voice recognition model, adaptively optimize the voice recognition strategy, and obtain the voice recognition result;

[0010] If the voice recognition result is that the mover issues a distress command, trigger the emergency rescue plan, send an alarm message to the rescue center, and plan the optimal rescue route according to the position information of the mover to guide the rescue personnel to quickly reach the accident scene to minimize accidental injuries to the greatest extent;

[0011] Adopt the incremental learning method to continuously optimize the wall occlusion voice recognition model, and dynamically update the acoustic characteristic database and the voice attenuation interference model according to the recognition accuracy rate and false alarm rate.

[0012] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:

[0013] The present invention discloses an Internet of Things voice signal matching and recognition method. This method establishes a wall acoustic characteristic database, collects voice samples at different distances, constructs a voice feature data set, analyzes the influence of wall occlusion on voice, and establishes a voice signal attenuation interference model. Combining wall attributes and acoustic parameters, constructs and trains a wall occlusion voice recognition model. During the movement, real-time obtain voice data and position information, dynamically adjust model parameters, and adaptively optimize the recognition strategy. When a distress command is recognized, trigger the rescue plan. The present invention also uses an incremental learning algorithm to continuously optimize the model, improve the voice recognition accuracy rate in a complex environment, provide reliable technical support for rapid rescue in an emergency, and effectively reduce accidental injuries. Brief Description of the Drawings

[0014] Figure 1 It is a flowchart of an Internet of Things voice signal matching and recognition method of the present invention.

[0015] Figure 2Schematic diagram of a method for matching and recognizing Internet of Things voice signals according to the present invention.

[0016] Figure 3 Another schematic diagram of a method for matching and recognizing Internet of Things voice signals according to the present invention. Detailed implementation manners

[0017] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this specification without creative efforts shall fall within the protection scope of this specification.

[0018] As Figures 1 - 3 , a method for matching and recognizing Internet of Things voice signals in this embodiment may specifically include:

[0019] S101. Obtain sound reflection data under different wall materials, thicknesses, shapes, surface roughnesses, internal structures, hole distributions, attachments, and temperatures. Through acoustic modeling and simulation analysis, determine the influence of various wall attributes on the frequency, amplitude, duration, energy, spectral characteristics, resonance frequency, echo characteristics, and signal-to-noise ratio of voice signals, and form a wall acoustic characteristic database.

[0020] Use a laser scanner to perform three-dimensional scanning and modeling on the wall surface to obtain wall surface roughness data; according to the wall surface roughness data, collect a wall temperature distribution map, and calculate the internal temperature gradient distribution of the wall; obtain the wall internal structure image from an X-ray tomography device, analyze the hole distribution and attachment conditions, and generate a three-dimensional model of the wall internal structure; according to the wall surface three-dimensional model, internal structure model, and temperature distribution data, construct an acoustic simulation model through finite element analysis software; use the ray tracing method to simulate the propagation path of sound waves in the wall within a specified frequency range, calculate the sound wave energy attenuation and reflection characteristics, and obtain the wall acoustic response function; use a microphone array to collect the reflected sound signals at different positions around the wall, extract the spectral characteristics of the signals; determine the arrival time and direction of each frequency component, and use the independent component analysis method for blind source separation to eliminate the influence of multiple reflections and reverberations, and construct a measured sound reflection characteristic data set.

[0021] Exemplarily, a laser scanner is used to perform three-dimensional scanning and modeling on the wall surface to obtain the wall surface roughness data. At the same time, an infrared thermal imaging technology is used to collect the wall temperature distribution map. Combining with the pre-established material thermophysical property parameter database, the internal temperature gradient distribution of the wall is calculated. An internal structure image of the wall is obtained from an X-ray tomography device, and the hole distribution and attachment conditions are analyzed to generate a three-dimensional model of the wall internal structure. According to the wall surface three-dimensional model, internal structure model and temperature distribution data, an acoustic simulation model is constructed by using finite element analysis software. The ray tracing method is used to simulate the propagation path of sound waves in the wall in the frequency range of 20 Hz to 20 kHz, and the finite difference time domain method is used to calculate the attenuation coefficient and reflection coefficient of sound waves propagating in the wall at different frequencies. The attenuation coefficient and reflection coefficient at each frequency are substituted into the acoustic response function model to obtain the acoustic frequency response function of the wall.

[0022]

[0023] This formula is the propagation equation of sound waves in a homogeneous medium, where p represents sound pressure, t represents time, and c represents the speed of sound. is the Laplace operator, representing the second derivative of space. The reflected sound signals are collected at different positions around the wall using a microphone array. The spectral features of the signals are extracted through Fourier transform. Combining the time delay estimation method and the direction estimation method, the arrival time and direction of each frequency component are determined. Independent component analysis is used for blind source separation to eliminate the influence of multiple reflections and reverberation, and a measured sound reflection characteristic data set is constructed, including obtaining the spectral features of the signal in different time periods by performing short-time Fourier transform on the sound signal. According to the phase difference of different frequency components in the spectral features, the generalized cross-correlation method is used to estimate the arrival time delay of each frequency component. According to the amplitude difference of different frequency components in the spectral features, the multiple signal classification algorithm is used to estimate the arrival direction of each frequency component. It is judged whether the time delay and direction of the frequency component satisfy the geometric relationship between the direct sound and the reflected sound. If satisfied, the frequency component is considered as the direct sound, otherwise it is considered as the reflected sound. For the frequency components judged as reflected sounds, the independent component analysis method is used for blind source separation to separate the reverberation signal into signals of different reflection sources. For the separated reflected sound signals, according to their time delay and direction information, the corresponding reflection surface position and reflection characteristic parameters are determined. The information such as the spectrum, time delay, direction, and reflection characteristics of the direct sound signal and each reflected sound signal is labeled to construct a measured data set of sound reflection characteristics. By comparing the prediction results of the simulation model and the measured data, multiple regression analysis is used to preliminarily evaluate the influence of each wall property on the acoustic characteristics and establish the initial parameter estimation values. The Markov chain Monte Carlo algorithm is used for Bayesian parameter estimation to optimize the acoustic model parameters and improve the prediction accuracy of the model for the wall acoustic characteristics. For different wall materials, thicknesses, shapes, surface roughnesses, internal structures, hole distributions, attachments, and temperature conditions, a wall acoustic characteristic database containing parameters such as the frequency, amplitude, duration, energy, spectral features, resonance frequency, echo characteristics, and signal-to-noise ratio of the voice signal is generated. The laser scanner scans a 1m×1m wall sample with an accuracy of 0.1mm to obtain the surface three-dimensional point cloud data, and the calculated surface roughness Ra value is 5μm. The infrared thermal imager collects the wall surface temperature distribution map with a resolution of 0.1°C. Combining the thermal conductivity of lime mortar of 0.93W / (m·K) in the pre-established material thermophysical parameter database, the temperature gradient inside the wall is calculated to be 0.5°C / cm through the heat conduction equation. The X-ray tomography equipment obtains the wall internal structure image with a resolution of 0.5mm, and identifies bubble holes with a diameter of 2 - 5mm, a porosity of 3%, and metal attachments of 2cm×2cm. Based on the obtained data, the finite element analysis software COMSOL Multiphysics is used to construct an acoustic simulation model, and the ray tracing method is used to simulate the sound wave propagation in the frequency range of 20Hz to 20kHz. Each 100Hz is a sampling point, and there are a total of 201 frequency points.Using the finite-difference time-domain method, with a grid size of 1 / 10 of the wavelength and a time step of 1 μs, calculate the acoustic wave energy attenuation and reflection characteristics to obtain the wall acoustic response function. Arrange 16 omnidirectional microphones to form a 4×4 array with a spacing of 20 cm. Collect the reflected sound signals at distances of 1 m, 2 m, and 3 m from the wall. The sampling rate is 48 kHz and the sampling duration is 5 s. Perform a 1024-point fast Fourier transform on the collected signals to extract spectral features. Use the Generalized Cross Correlation - Phase Transform (GCC-PHAT) algorithm to estimate the acoustic wave arrival time with an accuracy of 0.1 ms. Employ the Multiple Signal Classification (MUSIC) algorithm for sound source direction estimation with an azimuth resolution of 1°. Use the Fast Independent Component Analysis (FastICA) algorithm for independent component analysis to achieve blind source separation, eliminate the influence of multiple reflections and reverberation, and reconstruct the original sound source signal. Construct a measured acoustic reflection characteristic dataset containing parameters such as reflection coefficients and phase delays at various frequency points. Through multiple linear regression analysis, initially obtain relationships such as for every 1 cm increase in wall thickness, the reflection coefficient at 1 kHz in the middle frequency increases by 0.02, and for every 1 μm increase in surface roughness, the reflection coefficient at 10 kHz in the high frequency decreases by 0.01. Use the Metropolis-Hastings algorithm to implement Markov chain Monte Carlo sampling for Bayesian parameter estimation, iterate 10,000 times, and obtain the optimized acoustic model parameters. Finally, generate a wall acoustic characteristic database containing the frequency response curves, reverberation time RT60, early decay time EDT, speech transmission index STI, etc. of voice signals in the 100 Hz - 8 kHz frequency band under conditions of different materials (brick, concrete, gypsum board, etc.), thicknesses (5 - 30 cm), shapes (flat, concave-convex), surface roughnesses of 1 - 100 μm, internal structures (homogeneous, porous), hole distributions of 0 - 10%, attachments (metal, wood), and temperatures of 0 - 40 °C.

[0024] S102. According to the wall acoustic characteristic database, obtain the voice data of the exerciser at different positions at different distances from the wall during exercise as voice samples, construct a voice feature dataset, extract the frequency, amplitude, duration, energy, spectral features, resonance frequency, echo characteristics, and signal-to-noise ratio parameters of the voice signal, and obtain the voice feature vectors at different wall positions.

[0025] Use multiple omnidirectional microphone arrays to track the position of the mover in real time; record the change in the distance between the mover and the wall according to the mover's position, synchronously collect the voice data of the mover, and obtain the original voice sample set containing distance information; conduct a preliminary quality assessment on the original voice sample set, calculate the signal-to-noise ratio and conduct spectral analysis, and screen out valid voice samples according to the preset signal-to-noise ratio threshold; perform preprocessing on the valid voice samples, remove environmental low-frequency noise, eliminate echo interference, extract valid voice segments, and obtain clear voice signals; extract feature parameters from the clear voice signals, calculate the spectral features of the voice, estimate the formant frequency, extract the fundamental period, calculate the energy envelope, and obtain feature parameters such as the frequency, amplitude, duration, and energy of the voice; for the voice samples collected at different distances, calculate the signal-to-noise ratio and echo characteristic parameters, and combine with the voice feature parameters to construct a voice feature vector containing distance information.

[0026] Exemplarily, according to the wall acoustic characteristic database, a plurality of omnidirectional microphone arrays are arranged in a laboratory environment. The triangulation method is used to track the position of the mover in real time, record the change in the distance between the mover and the wall, synchronously collect the voice data of the mover, and form an original voice sample set containing distance information. A preliminary quality assessment is performed on the collected voice samples, the signal-to-noise ratio is calculated and spectral analysis is carried out. The signal-to-noise ratio threshold is set to 10 dB, and the effective samples are screened out. Preprocessing is performed on the screened voice samples. The environmental low-frequency noise is removed by setting a high-pass filter with a cut-off frequency of 100 Hz, the echo interference is eliminated by using the non-linear acoustic echo cancellation method, and the G.729B voice activity detection algorithm is used to extract the effective voice segments to obtain a clear voice signal. Feature parameters are extracted from the preprocessed voice signal. The spectral features of the voice are calculated using the short-time Fourier transform, the formant frequencies are estimated by the Levinson-Durbin recursive algorithm for linear prediction analysis, the pitch period is extracted by the FFT-based method for cepstrum analysis, the energy envelope of the voice signal is calculated, and the short-time zero-crossing rate and short-time energy are used simultaneously to handle the Doppler effect caused by movement, so as to obtain the feature parameters such as the frequency, amplitude, duration, and energy of the voice. For the voice samples collected at different distances, the signal-to-noise ratio and echo characteristic parameters are calculated, combined with the previously extracted voice features, a voice feature vector containing distance information is constructed, and non-linear dimensionality reduction is performed by kernel principal component analysis, and finally a low-dimensional representation reflecting the voice features at different distances from the wall position is obtained. Polynomial regression is used to analyze the change trend of the voice features at different distances, and the mapping relationship between the distance and the voice features is established. 16 omnidirectional microphones are arranged in a 10 m × 10 m laboratory to form a 4 × 4 array with a spacing of 2 m. The triangulation method is used to track the position of the mover, and the accuracy reaches ±5 cm. The mover maintains different distances from the wall in the range of 0.5 m to 5 m, and a position point is recorded every 0.5 m, with a total of 10 sampling positions. 5 seconds of voice is collected at each position, with a sampling rate of 44.1 kHz and 16-bit quantization. A preliminary quality assessment is performed on the collected voice samples, the power spectral density is calculated using the Welch method, with 1024-point FFT, 50% overlap, and Hanning window. The signal-to-noise ratio is calculated, and the threshold of 10 dB is set to screen out the effective samples. A fourth-order Butterworth high-pass filter with a cut-off frequency of 100 Hz is applied to the screened samples to remove the low-frequency noise. The frequency shift adaptive filtering algorithm is used for non-linear acoustic echo cancellation, and the convergence step size is set to 0.01. The G.729B algorithm is used for voice activity detection, with a frame length of 20 ms and a frame shift of 10 ms. The short-time Fourier transform is performed using a 512-point FFT, with a frame length of 25 ms and a frame shift of 10 ms to calculate the spectral features of the voice. Linear prediction analysis is carried out through a 10th-order Levinson-Durbin recursive algorithm to estimate the first 3 formant frequencies.Use the FFT-based method for cepstrum analysis to extract the fundamental period, with a search range of 50 - 500 Hz. Calculate the short-time energy and short-time zero-crossing rate, with a frame length of 20 ms, for dealing with the Doppler effect, and the thresholds are set to 0.015 and 0.1 respectively. For 10 different distance positions, extract 20-dimensional MFCC features, and combine with the previously obtained parameters such as frequency, amplitude, duration, energy, etc. to construct a 40-dimensional feature vector. Use kernel principal component analysis with a radial basis function kernel for non-linear dimensionality reduction, retain the principal components that explain 95% of the variance, and obtain a low-dimensional representation of about 15 dimensions. Finally, use third-order polynomial regression to analyze the changing trend of speech features at different distances, establish the mapping relationship between distance and speech features, and the coefficient of determination R of the regression model. 2 Reaches 0.85.

[0027] S103. According to the speech feature vector, combined with the distance and angular position relationship between the mover and the wall, analyze the speech features, establish an attenuation and interference model of the speech signal under wall occlusion, and obtain the influence law of wall occlusion on the speech recognition accuracy rate.

[0028] Obtain the mover's position, the wall's position, and the speech propagation path information, and map the information to a pre-constructed three-dimensional coordinate system; according to the mapping information in the three-dimensional coordinate system, calculate the propagation path length and the incident angle of the speech signal in space; use a material detector to measure the acoustic characteristic parameters of the wall material, and the acoustic characteristic parameters include the sound absorption coefficient and the scattering coefficient; according to the acoustic characteristic parameters, the propagation path length, and the incident angle, establish an attenuation model of the speech signal during spatial propagation; based on the multi-path propagation theory, analyze the time delay and phase difference between the direct sound and each reflected sound, and establish an acoustic interference model under wall occlusion conditions; use the attenuation model and the interference model to transform the original speech feature vector to obtain the speech feature vector after wall occlusion; input the transformed speech feature vector into a pre-trained speech recognizer to obtain the recognition result; compare the recognition result with the original speech content, calculate the word error rate; according to the word error rate, determine the influence law of wall occlusion on the speech recognition accuracy rate.

[0029] Exemplarily, according to the voice feature vector, the distance between the mover and the wall, and the angular position relationship, a three-dimensional coordinate system is constructed. The positions of the mover, the wall, and the voice propagation path are mapped into the coordinate system. The Monte Carlo ray tracing algorithm is used to simulate the propagation path of the voice signal in space, calculate the path lengths and incident angles of the direct sound, the reflected sound, and the diffracted sound. For complex wall structures, the triangular patch decomposition method is used to handle irregular reflections. The acoustic characteristic parameters of the wall material, including the sound absorption coefficient and the scattering coefficient, are obtained by using a material detector. Combining the incident angle and the frequency, the sound energy attenuation on each propagation path is calculated, considering air absorption, geometric divergence, and surface reflection loss. The frequency dependence of air absorption is calculated using the ISO9613-1 standard, and an attenuation model for the voice signal during spatial propagation is established. A sound pressure level meter is used to measure the actual sound pressure level at different positions and compared with the prediction results of the attenuation model to verify the accuracy of the model. Based on the multipath propagation theory, the time delay and phase difference between the direct sound and each reflected sound are analyzed, and the fast Fourier transform is used to calculate the frequency response function, considering the effects of the comb filtering effect and the Doppler effect on the voice signal, and an acoustic interference model under wall occlusion conditions is established. Using the established attenuation model and interference model, the original voice feature vector is transformed to obtain the voice features after being blocked by the wall. The transformed features are input into a pre-trained DeepSpeech speech recognizer. By comparing the recognition results with the original voice content, the word error rate is calculated as a quantitative index of the accuracy rate, and the recognition accuracy rates at different distances and angles are statistically analyzed. The differences in the recognition accuracy rates of vowels, consonants, and syllables are analyzed, the confusion matrix is used to evaluate the recognition performance, and multiple regression analysis is used to obtain the influence law of wall occlusion on the recognition accuracy rates of different types of voice content. In a laboratory of 10m×10m×3m, a three-dimensional coordinate system is established, the position of the mover is set as the origin (0,0,0), and the wall is located at 5m on the x-axis. Using the Monte Carlo ray tracing algorithm, 10,000 rays are emitted, and each ray is calculated for a maximum of 5 reflections. For a complex wall structure of 0.5m×0.5m, the triangular patch decomposition method is used to decompose it into 200 small triangles. The sound absorption coefficient of the wall in the frequency band of 125Hz - 4kHz is measured using a material detector, and the average value is 0.3, and the scattering coefficient is 0.2. According to the ISO9613-1 standard, under the conditions of 20°C and 50% relative humidity, the attenuation coefficient of the 500Hz sound wave in the air is calculated to be 2.8×10^-3 dB / m. After establishing the attenuation model, a sound pressure level meter is placed at 1m, 3m, and 5m away from the wall to measure the sound pressure level of the 1kHz pure tone and compared with the model prediction value, and the error is controlled within ±1dB. The 1024-point FFT is used to calculate the frequency response function, the sampling rate is 44.1kHz, and the comb filtering effect in the range of 0 - 22.05kHz is analyzed. Considering the Doppler effect generated when the mover moves at a speed of 1m / s, the maximum frequency shift is about 3Hz.After the original speech feature vectors are transformed by the attenuation and interference model, they are input into a pre-trained DeepSpeech speech recognizer. The recognizer is trained based on 5000 hours of Chinese speech data, with 5 hidden layers, each layer having 1024 neurons. The recognition results are compared with the original speech content, and the word error rate (WER) is calculated. At distances of 1m, 3m, and 5m, the average WERs are 10%, 15%, and 25% respectively. Through confusion matrix analysis, it is found that the recognition accuracy of vowels is 10% higher than that of consonants, and the error rate of monosyllabic words is 5% higher than that of polysyllabic words. Using multiple regression analysis, a linear relationship that the WER increases by 3% for every 1m increase in distance is obtained, and the R. 2 value is 0.92.

[0030] S104. According to the wall material, thickness, shape, surface roughness, internal structure, hole distribution, attachments, and temperature attributes, select the corresponding acoustic parameters from the wall acoustic property database, and combine them with the speech signal attenuation and interference model to construct a wall-obstructed speech recognition model for accurately recognizing speech commands in complex wall environments.

[0031] Obtain acoustic parameters from the wall acoustic property database. The acoustic parameters include sound absorption coefficient, scattering coefficient, transmission loss, and acoustic impedance. For the missing values in the acoustic parameters, Kriging interpolation method is used to complete them to obtain a complete set of wall acoustic property parameters. According to the wall acoustic property parameter set, simulate the propagation path of the speech signal in the complex wall environment. If multiple reflection and diffraction effects are detected, process the acoustic wave diffraction on the irregular surface. Through the propagation path model, calculate the frequency response function of the speech signal at the receiving point. Judge whether there are comb filtering effect and Doppler effect. If so, introduce a harmonic distortion model to process the nonlinear acoustic effect. Construct a time-frequency characteristic transformation model of the speech signal to transform the original speech features and obtain the speech features after being obstructed by the wall. Calculate the signal-to-noise ratio of the transformed speech features, determine the threshold for feature correlation analysis. Screen the effective features according to the threshold to generate speech data in the obstructed environment. Input the screened speech features into a pre-trained end-to-end speech recognition model based on Transformer. Judge the differences in recognition accuracy for different types of speech commands to obtain the speech recognition accuracy of the model in the complex wall environment.

[0032] Exemplarily, according to the wall material, thickness, shape, surface roughness, internal structure, hole distribution, attachments, and temperature properties, corresponding acoustic parameters are extracted from the wall acoustic property database, including the absorption coefficient, scattering coefficient, transmission loss, and acoustic impedance. The Kriging interpolation method is used to complement the missing parameters to construct a complete set of wall acoustic property parameters. A temperature sensor is used to monitor the wall temperature change in real time and dynamically update the acoustic parameters. The frequency-dependent ray tracing method is adopted to simulate the propagation path of the speech signal in the complex wall environment. Combining with the wall acoustic property parameter set, the energy attenuation and time delay on each propagation path are calculated. Considering the multiple reflection and diffraction effects, the Kirchhoff approximation method is used to handle the acoustic wave diffraction on the irregular surface, and a spatial propagation model of the speech signal is constructed. Based on the spatial propagation model, combined with the indoor sound field theory, the frequency response function of the speech signal at the receiving point is calculated. Considering the comb filtering effect and Doppler effect, and at the same time introducing a harmonic distortion model to handle the nonlinear acoustic effect, a time-frequency characteristic transformation model of the speech signal is constructed to transform the original speech features and obtain the speech features after being blocked by the wall. The signal-to-noise ratio calculation and feature correlation analysis are performed on the transformed speech features, and a threshold is set to screen out the effective features. Data augmentation techniques, such as adding simulated reflection and attenuation effects, are used to generate speech data in the occluded environment. The screened speech features are input into an end-to-end speech recognition model based on Transformer. The progressive transfer learning method is adopted. Based on the pre-trained model, the speech data in the occluded environment is used for fine-tuning to optimize the network parameters. The differences in the recognition accuracy of different types of speech commands (short commands, long sentences) are analyzed, and the perplexity is used to evaluate the model performance to improve the speech recognition accuracy of the model in the complex wall environment. In a laboratory of 10m×10m×3m, the wall acoustic parameters are extracted using the acoustic material database, such as the absorption coefficient of 0.1 - 0.9, scattering coefficient of 0.05 - 0.5, transmission loss of 20 - 60dB, and acoustic impedance of 2000 - 8000kg / m 2 s. The Kriging interpolation method is used with an interval of 0.1m 3The grid density is used to complete the missing parameters. 10 temperature sensors are arranged, sampling every 30 seconds with an accuracy of ±0.1℃, and the acoustic parameters are updated dynamically. A frequency-dependent ray tracing algorithm is used to emit 1000 rays in the range of 100Hz-8kHz, with a frequency point of every 100Hz, and a maximum of 5 reflections are calculated. The Kirchhoff approximation method is used to process the sound wave diffraction of the 0.5m×0.5m irregular wall. A 16-order FIR filter is constructed to simulate the comb filtering effect, considering the Doppler frequency shift of ±10Hz. A 3rd-order harmonic distortion model is introduced, and the distortion coefficient is set to 0.05. The signal-to-noise ratio is calculated for the transformed speech features, and a threshold of 6dB is set to retain features with a signal-to-noise ratio greater than the threshold. Feature correlation analysis is used to eliminate redundant features with a correlation coefficient greater than 0.95. 10,000 speech data in an occluded environment are generated by adding random reflections of 0-10dB and random attenuation of 0-20dB. An end-to-end speech recognition model with a 12-layer Transformer structure, a hidden layer dimension of 512, and 8 attention heads. Progressive transfer learning is used, and the learning rate is gradually increased from 1e-5 to 1e-4. The difference in recognition accuracy between short commands of 1-3 seconds and long sentences of 5-10 seconds is analyzed, and the perplexity is used to evaluate the model performance. The perplexity of short commands and long sentences is controlled within 50 and 100 respectively. The final model has a speech recognition accuracy of 95% in a complex wall environment.

[0033] S105. Preprocess the speech signal in the wall-occluded speech recognition model to extract the frequency, amplitude, duration, energy, spectral characteristics, resonance frequency, echo characteristics and signal-to-noise ratio characteristic parameters of the speech signal, and optimize the feature subset to reduce the interference of the wall acoustic characteristics on speech recognition.

[0034] After preprocessing the input speech signal, voice activity detection is performed to obtain valid speech segments. Feature parameters are extracted from the valid speech segments, and the feature parameters include: calculating the spectrum features of the valid speech segments; performing linear prediction analysis through the Levinson-Durbin recursive algorithm to obtain the formant frequency of the valid speech segments; extracting the pitch period of the valid speech segments using the real cepstrum method; and calculating the energy envelope of the valid speech segments. Modeling the feature parameters includes: using a Gaussian mixture model to model the sound features of different speakers; calculating the signal-to-noise ratio and echo characteristic parameters; combining the spectrum features, formant frequency, pitch period and energy envelope to construct a high-dimensional feature vector. The high-dimensional feature vector is optimized, and the optimization includes: normalizing the high-dimensional feature vector using the Z-score standardization method; performing dimensionality reduction and correlation analysis on the normalized feature vector using principal component analysis; and compensating for the influence of the acoustic characteristics of the wall. Feature selection is performed on the feature vector after optimization to obtain an optimized feature vector.

[0035] Exemplarily, preprocess the input speech signal. Remove ambient low-frequency noise by setting a high-pass filter with a cut-off frequency of 100 Hz, eliminate echo interference using a least mean square error adaptive filter, perform voice activity detection using the G.729B algorithm, extract the effective speech segments, and obtain a clear speech signal. Extract feature parameters from the preprocessed speech signal. Calculate the spectral features of the speech using the short-time Fourier transform, perform linear prediction analysis to estimate the formant frequencies through the Levinson-Durbin recursive algorithm, extract the pitch period using the real cepstrum method, and calculate the energy envelope of the speech signal to obtain feature parameters such as the frequency, amplitude, duration, and energy of the speech. Use a Gaussian mixture model to model the voice characteristics of different speakers to improve the robustness of feature extraction. For the extracted feature parameters, calculate the signal-to-noise ratio and echo characteristic parameters, and combine with the previously obtained features such as frequency, amplitude, duration, and energy to construct a high-dimensional feature vector. Use the Z-score normalization method to normalize the feature parameters of different scales to ensure the comparability of the data in each dimension of the feature vector. Perform preliminary dimensionality reduction and correlation analysis on the feature vector using principal component analysis to remove redundant information. Use the inverse filtering method of the acoustic transfer function to compensate for the influence of the wall acoustic characteristics and reduce the interference of wall occlusion on the speech features. Perform feature selection using Lasso regression based on L1 regularization, evaluate the performance of different feature subsets through cross-validation, set the cross-validation error increase not exceeding 1% as the stopping criterion, select the optimal feature subset, reduce the impact of redundant features on the recognition performance, and obtain the optimized feature vector. In a laboratory environment, preprocess the speech signal with a sampling rate of 16 kHz. Use a 4th-order Butterworth high-pass filter with a cut-off frequency set to 100 Hz to remove ambient low-frequency noise. Use a 32nd-order least mean square error adaptive filter with a step size of 0.01 to eliminate echo interference. Use an improved G.729B algorithm with a frame length of 20 ms and a frame shift of 10 ms to perform voice activity detection and extract the effective speech segments. Perform a 512-point short-time Fourier transform on the preprocessed signal with a frame length of 25 ms and a frame shift of 10 ms to calculate the spectral features of the speech. Use a 10th-order Levinson-Durbin recursive algorithm to perform linear prediction analysis and estimate the first 3 formant frequencies. Use the real cepstrum method with a window length of 30 ms to extract the pitch period, and the search range is 50 - 500 Hz. Calculate the short-time energy with a frame length of 25 ms to obtain the energy envelope. Use a Gaussian mixture model with 32 components to model the voice characteristics of 5 different speakers to improve the robustness of feature extraction. Calculate the signal-to-noise ratio, set a threshold of 6 dB, and retain the features with a signal-to-noise ratio greater than the threshold. Construct a high-dimensional vector containing 40-dimensional features and use the Z-score normalization method for normalization. Perform dimensionality reduction using principal component analysis and retain the principal components that explain 95% of the variance, reducing to 15 dimensions.Using the inverse filtering technology of the acoustic transfer function based on the all-pole model with an order of 20 to compensate for the influence of the wall acoustic characteristics. Lasso regression is used for feature selection, and the regularization parameter λ starts from 0.1 and increases by 0.1 in each iteration until the cross-validation error increases by more than 1%. Finally, a 10-dimensional optimal feature subset is selected to form an optimized feature vector.

[0036] S106. During the movement, the voice data and position information of the mover are obtained in real time and input into the wall occlusion speech recognition model. According to the distance and angular position relationship between the mover and the wall, the acoustic parameters and speech feature weights of the wall occlusion speech recognition model are dynamically adjusted, the speech recognition strategy is adaptively optimized, and the speech recognition result is obtained.

[0037] Obtain the three-dimensional position coordinates of the mover, which are calculated in real time by the multi-sensor fusion positioning technology using UWB, inertial sensors, and cameras; collect the voice signal of the mover according to the three-dimensional position coordinates, and the voice signal is collected by the microphone array and undergoes signal preprocessing and feature extraction to obtain the voice feature vector; calculate the distance and angular relationship between the mover and the wall, which are determined by the three-dimensional position coordinates; simulate the acoustic wave propagation path, which is calculated based on the distance and angular relationship; perform dynamic compensation on the voice feature vector, and the dynamic compensation is based on the acoustic propagation model updated by the acoustic wave propagation path, and the compensated voice feature vector is obtained by adjusting the frequency response, energy distribution, and delay characteristics; input the compensated voice feature vector into the adaptive speech recognizer, and the adaptive speech recognizer uses the recursive least squares algorithm for online learning, and the speech recognition result is obtained by updating the recognizer parameters.

[0038] Exemplarily, a multi-sensor fusion positioning technology that uses UWB, inertial sensors, and cameras is employed to obtain the three-dimensional position coordinates of a mover in real time. Extended Kalman filtering is used for data fusion. Meanwhile, a microphone array is used to collect the voice signals of the mover. The continuous voice signals are segmented and spliced through the sliding window method, and signal preprocessing and feature extraction are performed to obtain voice feature vectors and position information. According to the real-time position coordinates of the mover, the distance and angle relationship between the mover and the wall are calculated. The ray tracing algorithm for different frequency bands is used to simulate the propagation paths of sound waves with different frequencies, and the multi-path effect caused by complex wall structures is processed. Combining with the wall acoustic characteristic database, the attenuation coefficient and reflection coefficient in the acoustic propagation model are updated. A sound pressure level meter is used for real-time measurement at key positions to verify the accuracy of the updated acoustic propagation model. Based on the updated acoustic propagation model, dynamic compensation is performed on the voice feature vectors to adjust the frequency response, energy distribution, and time delay characteristics. At the same time, an exponential decay function based on distance and angle is used to dynamically adjust the weight coefficients of each dimension of the voice features. An adaptive filter is used to compensate for the Doppler effect caused by changes in different moving speeds and directions, improving the adaptability of the model to the motion state. The compensated voice feature vectors are input into an adaptive speech recognizer, and the recursive least squares algorithm is used for online learning to update the recognizer parameters in real time. The dynamic time warping algorithm is used to balance real-time performance and recognition accuracy, and the decision threshold is dynamically adjusted according to the confidence level of the recognition result, and finally the speech recognition result is output. In a 10m×10m×3m laboratory, 4 UWB base stations, 6 high-definition cameras, and 1 inertial measurement unit are deployed. The extended Kalman filtering algorithm is used to fuse multi-sensor data, and the positioning accuracy reaches ±5cm. 16 omnidirectional microphones are used to form a 4×4 array, with a sampling rate of 48kHz. A 200ms sliding window with an overlap rate of 50% is used to process continuous speech. The ray tracing algorithm is used to simulate the sound wave propagation in the range of 100Hz - 8kHz, with each 100Hz as a frequency band, and 1000 rays are emitted, and at most 5 reflections are calculated. A 15th-order all-pole model is used to construct the acoustic transfer function, and the model parameters are updated every 0.5s. A sound pressure level meter is placed at 1m, 2m, and 3m away from the wall to verify the model accuracy in real time, and the error is controlled within ±1dB. An exponential decay function w = exp(-0.1d - 0.05θ) based on distance d and angle θ is used to dynamically adjust the feature weights, where exp represents the natural exponential function, which is an exponential function with the mathematical constant e (approximately equal to 2.71828) as the base. A 10th-order adaptive FIR filter with a step size of 0.01 is used to compensate for the Doppler frequency shift of ±10Hz. The adaptive speech recognizer adopts a 12-layer Transformer structure, with a hidden layer dimension of 512 and 8 attention heads. The learning rate of the recursive least squares algorithm is set to 0.1, and the forgetting factor is 0.98. The dynamic time warping algorithm uses the Sakoe-Chiba band with a bandwidth of 10% of the sequence length.The initial value of the decision threshold is set to 0.8 and is dynamically adjusted according to the average confidence of the first 10 recognition results, with an adjustment step size of 0.02. Finally, within the range of the movement speed of 0 - 2 m / s and the steering angle of ±90°, the speech recognition accuracy remains above 95%.

[0039] S107. If the speech recognition result is that the mover issues a distress instruction, trigger the emergency rescue plan, send an alarm message to the rescue center, and plan the optimal rescue route according to the position information of the mover to guide the rescue personnel to quickly reach the accident scene to minimize accidental injuries to the greatest extent possible.

[0040] Use the BERT model to perform semantic analysis on the speech recognition result, extract keywords and match them with the preset distress instruction vocabulary to obtain the semantic matching degree score. According to the semantic matching degree score and the emotional urgency score, determine whether to trigger the distress instruction. If the distress instruction is triggered, construct and update the three-dimensional scene model in real time. According to the three-dimensional scene model, combined with the geographic information system, generate a detailed scene map containing information such as obstacles, channels, and entrances and exits. Package the distress instruction, the position of the mover, and the scene map information and send it to the rescue center. Receive the rescue instruction sent by the rescue center and perform multi-objective path planning using the A* algorithm. According to the multi-objective path planning result, calculate multiple candidate rescue routes and select the optimal rescue route from the multiple candidate rescue routes. Obtain the rescue effect data of the optimal rescue route and establish a rescue effect evaluation system. According to the rescue effect evaluation system, record the response time, arrival time, and rescue success rate indicators to obtain the analysis result of the key factors in the rescue process.

[0041] Exemplarily, the BERT model is used to perform semantic analysis on the speech recognition results, extract keywords and match them with a preset distress instruction word library. At the same time, speech emotion recognition technology is used to analyze the urgency of the speech. If the comprehensive score of semantic matching degree and emotional urgency exceeds the preset threshold, it is determined as a distress instruction, and the emergency rescue plan is triggered. According to the real-time position information of the mover and the surrounding environment data, the lidar SLAM technology is used to construct and update the three-dimensional scene model in real time. Combining with the geographic information system, a detailed scene map containing information such as obstacles, channels and entrances and exits is generated. Through the MQTT protocol and the 5G network, information such as distress instructions, mover positions, and scene maps are packaged and sent to the rescue center using an encrypted transmission protocol. At the same time, on-site emergency equipment such as sirens and emergency lighting is activated. After receiving the information, the rescue center immediately performs information parsing and verification, confirms the accident level and required resources, and dispatches the nearest rescue personnel and equipment. The improved A* algorithm is used for multi-target path planning, considering the rescue personnel's position, real-time traffic conditions and scene obstacles, calculating multiple candidate rescue routes, and comprehensively evaluating them according to factors such as time and safety to select the optimal rescue route. The rescue personnel are guided in real time through the navigation, and the road conditions information is continuously updated to dynamically adjust the route. The rescue effect is evaluated, and indicators such as response time, arrival time, and rescue success rate are recorded. Data mining technology is used to analyze the key factors in the rescue process, and the rescue strategy and resource allocation are continuously optimized. In practical applications, the pre-trained BERT-base-chinese model is used to perform semantic analysis on the speech recognition results, extract the top-5 keywords, and calculate the cosine similarity with a word library containing 200 distress-related words, with the threshold set to 0.8. At the same time, a speech emotion recognition model based on CNN-LSTM is used to classify the speech into three categories: "calm", "anxious", and "fear", and the urgency levels are assigned 0, 0.5, and 1 respectively. The rescue is triggered when the comprehensive score exceeds 0.9. The Velodyne VLP-16 lidar is used for SLAM mapping, with a scanning frequency of 10Hz, a measurement range of 100m, an accuracy of ±3cm, and the three-dimensional scene of a 10m×10m area is updated in real time. The MQTT protocol is used, and the QoS level is set to 2. Through the 5G network (uplink rate 100Mbps), data packets are transmitted, which are about 5MB and contain 1080p scene images, position coordinates, and voice files to the rescue center. The rescue center server with 64-core CPU and 256GB memory completes information parsing within 100ms, and determines the accident level from 1 to 5 through the decision tree algorithm. The improved A* algorithm is used for path planning, considering 5 goals, including time, distance, safety, road conditions, and resource consumption. A 100×100 grid map is set, and the road conditions information is updated every 500ms. The navigation system uses Kalman filtering to fuse GPS and IMU data, and the positioning accuracy reaches ±1m.The rescue effect evaluation system records 10 indicators, including the average response time target < 3 minutes, the arrival time target < 10 minutes, and the rescue success rate target > 95%. The random forest algorithm is used to analyze historical data, with a sample size of > 10,000 rescue records, to identify the key factors affecting the rescue effect, and the rescue strategy is updated monthly.

[0042] S108. Continuously optimize the wall occlusion speech recognition model using the incremental learning method, and dynamically update the acoustic feature database and the speech attenuation interference model according to the recognition accuracy rate and the false alarm rate.

[0043] Use the sliding window method to collect speech recognition results, calculate the F1 score according to the speech recognition results; according to the F1 score, use Gaussian process regression for Bayesian optimization to dynamically adjust the weights of the parameters in the acoustic feature database; based on the acoustic feature database, reconstruct the attenuation model and interference model of the speech signal, and use the recursive least squares method to adjust the model parameters in real time; perform incremental learning on the wall occlusion speech recognition model, combine the newly collected speech data with the historical data, and dynamically adjust the model structure and parameters; establish a model performance evaluation system, including accuracy rate, recall rate, and F1 score indicators, set dynamic thresholds, and trigger model retraining when the performance indicators are lower than the thresholds.

[0044] Exemplarily, the sliding window method is used to collect the speech recognition results in the recent period, and the F1 score is calculated to evaluate the recognition performance. At the same time, the sensor network is used to collect environmental parameters in real time, including attributes such as wall material, thickness, shape, surface roughness, internal structure, hole distribution, attachments, and temperature. According to the calculated F1 score, Gaussian process regression is used for Bayesian optimization to dynamically adjust the weights of the parameters in the acoustic characteristics database, focusing on optimizing the parameters that have a greater impact on the recognition effect, updating the acoustic characteristics database, and calculating the confidence interval of the parameter update. The k-fold cross-validation method is used to verify the effectiveness of the updated database to ensure the reasonable integration of new data and historical data. Based on the updated acoustic characteristics database, the attenuation model and interference model of the speech signal are reconstructed. The recursive least squares method is used to adjust the model parameters in real time, and the minimum description length criterion is used to select the optimal model structure to reduce the computational complexity and improve the adaptability of the model to environmental changes. The Learn++ algorithm is used for incremental learning of the wall occlusion speech recognition model, combining the newly collected speech data with the historical data, dynamically adjusting the model structure and parameters, adding an L2 regularization term to prevent overfitting, and continuously optimizing the model performance. A model performance evaluation system is established, including indicators such as accuracy, recall rate, and F1 score. A dynamic threshold is set, and when the performance indicator is lower than the threshold, the model is retrained to ensure continuous improvement of the recognition accuracy and reduction of the false alarm rate. In practical applications, a 30-minute sliding window is used, updated every 5 minutes, to collect speech recognition results and calculate the F1 score to evaluate the performance. 20 environmental sensors are deployed, including 5 material analyzers, 5 thickness gauges, 5 roughness sensors, and 5 temperature sensors, with a sampling frequency of 1 Hz. Gaussian process regression is used for Bayesian optimization, with 100 iterations and a learning rate of 0.01 to optimize 50 key parameters in the acoustic characteristics database. 5-fold cross-validation is used, and the validation threshold is set to 0.85. The reconstructed attenuation model and interference model use a 20th-order all-pole structure, and the forgetting factor of the recursive least squares method is set to 0.98, and the penalty term coefficient of the minimum description length criterion is 0.1. The Learn++ algorithm sets the number of base classifiers to 10, uses a decision tree as the base classifier, and the maximum depth is 5. An L2 regularization term of 0.001 is added to prevent overfitting. In the model performance evaluation system, the weights of accuracy, recall rate, and F1 score are 0.3, 0.3, and 0.4 respectively, and the initial value of the dynamic threshold is set to 0.9, adjusted in steps of 0.01 after each evaluation. When the comprehensive performance indicator is lower than the threshold for 3 consecutive times, the model is retrained. The speech recognition accuracy of the wall occlusion speech recognition model in a complex wall environment is improved from 85% initially to 95%, the false alarm rate is reduced from 10% to 3%, and the F1 score is improved from 0.87 to 0.96.

[0045] For those skilled in the art, it is obvious that the present application is not limited to the details of the above-described exemplary embodiments, and that the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. Therefore, in any respect, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present application. Any reference signs in the claims should not be construed as limiting the claims involved.

Claims

1. An Internet of Things voice signal matching and recognition method, characterized in that, The method includes: obtaining sound reflection data under different wall materials, thicknesses, shapes, surface roughnesses, internal structures, hole distributions, attachments, and temperatures, and determining the effects of various wall properties on the frequency, amplitude, duration, energy, spectral characteristics, resonance frequency, echo characteristics, and signal-to-noise ratio of voice signals through acoustic modeling and simulation analysis to form a wall acoustic characteristic database; obtaining voice data of a mover at different distances from the wall during movement as voice samples according to the wall acoustic characteristic database, constructing a voice feature dataset, extracting the frequency, amplitude, duration, energy, spectral characteristics, resonance frequency, echo characteristics, and signal-to-noise ratio parameters of the voice signal to obtain voice feature vectors at different wall positions; analyzing the voice features according to the voice feature vectors in combination with the distance and angular position relationship between the mover and the wall, establishing an attenuation and interference model of the voice signal under wall occlusion, and obtaining the influence law of wall occlusion on the voice recognition accuracy; selecting corresponding acoustic parameters from the wall acoustic characteristic database according to the wall material, thickness, shape, surface roughness, internal structure, hole distribution, attachment, and temperature attributes, and combining them with the voice signal attenuation and interference model to construct a wall occlusion voice recognition model for accurately recognizing voice commands in a complex wall environment; preprocessing the voice signal in the wall occlusion voice recognition model, extracting the frequency, amplitude, duration, energy, spectral characteristics, resonance frequency, echo characteristics, and signal-to-noise ratio characteristic parameters of the voice signal, and optimizing the feature subset to reduce the interference of wall acoustic characteristics on voice recognition; obtaining the voice data and position information of the mover in real time during movement and inputting them into the wall occlusion voice recognition model, dynamically adjusting the acoustic parameters and voice feature weights of the wall occlusion voice recognition model according to the distance and angular position relationship between the mover and the wall, adaptively optimizing the voice recognition strategy, and obtaining the voice recognition result; if the voice recognition result is that the mover issues a distress command, triggering an emergency rescue plan, sending an alarm message to the rescue center, and planning the optimal rescue route according to the position information of the mover to guide the rescue personnel to quickly reach the accident site to minimize accidental injuries to the greatest extent; continuously optimizing the wall occlusion voice recognition model using an incremental learning method, and dynamically updating the acoustic characteristic database and voice attenuation interference model according to the recognition accuracy and false alarm rate.

2. The method according to claim 1, wherein Obtain the sound reflection data under different wall materials, thicknesses, shapes, surface roughnesses, internal structures, hole distributions, attachments, and temperatures. Through acoustic modeling and simulation analysis, determine the effects of various wall properties on the frequency, amplitude, duration, energy, spectral characteristics, resonance frequency, echo characteristics, and signal-to-noise ratio of the voice signal, and form a wall acoustic characteristic database, including: performing three-dimensional scanning and modeling on the wall surface using a laser scanner to obtain the wall surface roughness data; collecting the wall temperature distribution map according to the wall surface roughness data, and calculating the internal temperature gradient distribution of the wall; obtaining the wall internal structure image from the X-ray tomography equipment, analyzing the hole distribution and attachment conditions, and generating a three-dimensional model of the wall internal structure; constructing an acoustic simulation model through finite element analysis software based on the wall surface three-dimensional model, internal structure model, and temperature distribution data; using the ray tracing method to simulate the propagation path of sound waves in the wall within a specified frequency range, calculating the sound wave energy attenuation and reflection characteristics, and obtaining the wall acoustic response function; using a microphone array to collect the reflected sound signals at different positions around the wall, and extracting the spectral characteristics of the signals; Determine the arrival time and direction of each frequency component, and use the independent component analysis method for blind source separation to eliminate the influence of multiple reflections and reverberation, and construct a measured sound reflection characteristic data set.

3. The method according to claim 1, wherein According to the wall acoustic characteristic database, obtain the voice data of the exerciser at different distances from the wall during exercise as voice samples, construct a voice feature data set, and extract the frequency, amplitude, duration, energy, spectral characteristics, resonance frequency, echo characteristics, and signal-to-noise ratio parameters of the voice signal to obtain the voice feature vector at different wall positions at different distances, including: using multiple omnidirectional microphone arrays to track the position of the exerciser in real time; recording the distance change between the exerciser and the wall according to the exerciser's position, and synchronously collecting the voice data of the exerciser to obtain the original voice sample set containing distance information; performing a preliminary quality assessment on the original voice sample set, calculating the signal-to-noise ratio and performing spectral analysis, and screening out effective voice samples according to the preset signal-to-noise ratio threshold; performing preprocessing on the effective voice samples, removing environmental low-frequency noise, eliminating echo interference, extracting effective voice segments, and obtaining clear voice signals; Extract the characteristic parameters from the clear voice signal, calculate the spectral characteristics of the voice, estimate the resonance peak frequency, extract the fundamental period, calculate the energy envelope, and obtain the characteristic parameters of the frequency, amplitude, duration, and energy of the voice; for the voice samples collected at different distances, calculate the signal-to-noise ratio and echo characteristic parameters, and combine the voice characteristic parameters to construct a voice feature vector containing distance information.

4. The method according to claim 1, wherein Based on the voice feature vectors, combined with the distance and angular position relationship between the mover and the wall, analyze the voice features, establish an attenuation and interference model of the voice signal under wall occlusion, and obtain the influence law of wall occlusion on the voice recognition accuracy rate, including: obtaining the mover position, wall position and voice propagation path information, and mapping the information to a pre-constructed three-dimensional coordinate system; according to the mapping information in the three-dimensional coordinate system, calculate the propagation path length and incident angle of the voice signal in space; use a material detector to measure the acoustic characteristic parameters of the wall material, and the acoustic characteristic parameters include the sound absorption coefficient and the scattering coefficient; according to the acoustic characteristic parameters, propagation path length and incident angle, establish an attenuation model of the voice signal in the space propagation process; based on the multipath propagation theory, analyze the time delay and phase difference between the direct sound and each reflected sound, and establish an acoustic interference model under wall occlusion conditions; use the attenuation model and the interference model to transform the original voice feature vectors to obtain the voice feature vectors after being occluded by the wall. Input the transformed voice feature vectors into a pre-trained voice recognizer to obtain the recognition results; compare the recognition results with the original voice content and calculate the word error rate; according to the word error rate, determine the influence law of wall occlusion on the voice recognition accuracy rate.

5. The method according to claim 1, characterized in that, According to the wall material, thickness, shape, surface roughness, internal structure, hole distribution, attachments and temperature attributes, select the corresponding acoustic parameters from the wall acoustic characteristic database, and combine them with the voice signal attenuation and interference models to construct a wall occlusion voice recognition model for accurately recognizing voice commands in a complex wall environment, including: obtaining acoustic parameters from the wall acoustic characteristic database, and the acoustic parameters include the sound absorption coefficient, scattering coefficient, transmission loss and acoustic impedance. For the missing values in the acoustic parameters, use Kriging interpolation method to complete them to obtain a complete set of wall acoustic characteristic parameters; according to the set of wall acoustic characteristic parameters, simulate the propagation path of the voice signal in a complex wall environment; if multiple reflection and diffraction effects are detected, process the acoustic wave diffraction on the irregular surface. Through the propagation path model, calculate the frequency response function of the voice signal at the receiving point; judge whether there are comb filtering effect and Doppler effect, and if so, introduce a harmonic distortion model to process the nonlinear acoustic effects; construct a time-frequency characteristic transformation model of the voice signal to transform the original voice features and obtain the voice features after being occluded by the wall. Calculate the signal-to-noise ratio of the transformed voice features to determine the threshold for feature correlation analysis; screen the effective features according to the threshold to generate voice data in the occlusion environment; input the screened voice features into a pre-trained end-to-end voice recognition model based on Transformer; judge the differences in the recognition accuracy rates of different types of voice commands to obtain the voice recognition accuracy rate of the model in a complex wall environment.

6. The method according to claim 1, characterized in that, Preprocess the speech signal in the wall-obstructed speech recognition model, extract the frequency, amplitude, duration, energy, spectral features, resonance frequency, echo characteristics, and signal-to-noise ratio characteristic parameters of the speech signal, and optimize the feature subset to reduce the interference of the wall acoustic characteristics on speech recognition, including: after preprocessing the input speech signal, perform speech activity detection to obtain effective speech segments; extract characteristic parameters from the effective speech segments, and the characteristic parameters include: calculate the spectral features of the effective speech segments; perform linear prediction analysis through the Levinson-Durbin recursive algorithm to obtain the formant frequencies of the effective speech segments; use the real cepstrum method to extract the fundamental period of the effective speech segments; calculate the energy envelope of the effective speech segments; model the characteristic parameters, and the modeling includes: use the Gaussian mixture model to model the voice characteristics of different speakers; calculate the signal-to-noise ratio and echo characteristic parameters; combine the spectral features, formant frequencies, fundamental period, and energy envelope to construct a high-dimensional feature vector; perform optimization processing on the high-dimensional feature vector, and the optimization processing includes: normalize the high-dimensional feature vector using the Z-score normalization method; use principal component analysis to perform dimensionality reduction and correlation analysis on the normalized feature vector; compensate for the influence of the wall acoustic characteristics; perform feature selection on the optimized feature vector to obtain the optimized feature vector.

7. The method according to claim 1, characterized in that, During the movement, real-time obtain the speech data and position information of the mover, and input them into the wall-obstructed speech recognition model. According to the distance and angular position relationship between the mover and the wall, dynamically adjust the acoustic parameters and speech feature weights of the wall-obstructed speech recognition model, adaptively optimize the speech recognition strategy, and obtain the speech recognition result, including: obtain the three-dimensional position coordinates of the mover, which are calculated in real time by the multi-sensor fusion positioning technology using UWB, inertial sensors, and cameras; collect the speech signal of the mover according to the three-dimensional position coordinates, and the speech signal is collected by the microphone array and undergoes signal preprocessing and feature extraction to obtain the speech feature vector; calculate the distance and angular relationship between the mover and the wall, which is determined by the three-dimensional position coordinates; simulate the acoustic wave propagation path, which is calculated based on the distance and angular relationship; perform dynamic compensation on the speech feature vector, and the dynamic compensation is performed based on the acoustic propagation model updated by the acoustic wave propagation path, and the compensated speech feature vector is obtained by adjusting the frequency response, energy distribution, and delay characteristics. Input the compensated speech feature vector into the adaptive speech recognizer, and the adaptive speech recognizer uses the recursive least squares algorithm for online learning to obtain the speech recognition result by updating the recognizer parameters.

8. The method according to claim 1, characterized in that If the voice recognition result is that the exerciser issues a distress instruction, an emergency rescue plan is triggered, an alarm message is sent to the rescue center, and the optimal rescue route is planned based on the position information of the exerciser to guide the rescue personnel to quickly reach the accident scene to minimize accidental injuries, including: using the BERT model to perform semantic analysis on the voice recognition result, extracting keywords and matching them with a preset distress instruction thesaurus to obtain a semantic matching degree score; judging whether to trigger a distress instruction according to the semantic matching degree score and the emotional emergency degree score; if a distress instruction is triggered, constructing and updating a three-dimensional scene model in real time; generating a detailed scene map containing obstacle, passage, and entrance / exit information based on the three-dimensional scene model in combination with the geographic information system; packing the distress instruction, the exerciser's position, and the scene map information and sending them to the rescue center; receiving the rescue instruction sent by the rescue center and performing multi-objective path planning using the A* algorithm; calculating multiple candidate rescue routes according to the multi-objective path planning result and selecting the optimal rescue route from the multiple candidate rescue routes; obtaining the rescue effect data of the optimal rescue route and establishing a rescue effect evaluation system; recording the response time, arrival time, and rescue success rate indicators according to the rescue effect evaluation system to obtain the key factor analysis result of the rescue process.

9. The method according to claim 1, wherein The incremental learning method is used to continuously optimize the wall occlusion voice recognition model, and the acoustic characteristic database and the voice attenuation interference model are dynamically updated according to the recognition accuracy and false alarm rate, including: using the sliding window method to collect voice recognition results and calculating the F1 score according to the voice recognition results; According to the F1 score, Bayesian optimization is performed using Gaussian process regression to dynamically adjust the weights of the parameters in the acoustic characteristic database; Based on the acoustic characteristic database, a new attenuation model and interference model of the voice signal are reconstructed, and the model parameters are adjusted in real time using the recursive least squares method; Incremental learning is performed on the wall occlusion voice recognition model, and the newly collected voice data is combined with the historical data to dynamically adjust the model structure and parameters; a model performance evaluation system is established, including accuracy, recall, and F1 score indicators, and a dynamic threshold is set. When the performance indicators are lower than the threshold, model retraining is triggered.

Citation Information

Patent Citations

  • Audio data processing method and apparatus

    CN107016996A

  • Speech recognition method based on artificial intelligence

    CN118155623A