A high-precision speech recognition and semantic understanding method

By acquiring and processing speech signals in a far-field environment, extracting distortion and speech rate parameters, correcting and weighting Mel-Cepstral features, and combining semantic feature databases and confidence calibration, the problems of vocal tract distortion and speech rate variation in far-field speech recognition and semantic understanding are solved, achieving high-precision speech recognition and semantic understanding.

CN122369448APending Publication Date: 2026-07-10SHANGHAI MAIJUN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI MAIJUN TECHNOLOGY CO LTD
Filing Date
2026-04-29
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

In far-field environments, vocal tract distortion and non-steady speech rate changes are prone to occur during speech signal transmission, resulting in low speech recognition accuracy and inaccurate semantic understanding.

Method used

The system acquires raw speech signals from multiple sound sources in the far field, performs synchronous acquisition through a multi-channel microphone array, and then performs DC removal and pre-filtering to generate discrete speech sampling sequences. After frequency domain transformation, amplitude-frequency distortion and non-steady-state speech rate parameters are extracted, Mel-frequency cepstral features are corrected and dynamically weighted, and combined with semantic feature library and confidence recursive calibration, the final output speech recognition and semantic understanding results are obtained.

Benefits of technology

It precisely solves the problem of low recognition accuracy caused by vocal tract distortion and non-steady speech rate in far-field speech transmission, and achieves high-precision speech recognition and semantic understanding, ensuring the stability and accuracy of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369448A_ABST
    Figure CN122369448A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of speech recognition and semantic understanding, and discloses a high-precision speech recognition and semantic understanding method. The method comprises the following steps: collecting a far-field multi-sound-source original speech signal to generate a discrete speech sampling sequence; performing frequency domain transformation on the original speech signal to extract an amplitude-frequency distortion quantization parameter; segmenting the original speech signal to generate a speech frame sequence and calculating a non-steady-state speech speed quantization parameter; extracting an original mel-frequency cepstrum feature and performing distortion correction to obtain a weighted acoustic feature through dynamic weighting; constructing a preset semantic feature library and mapping to generate an initial semantic matching vector; and recursively iterating to calibrate the semantic confidence and screening an optimal term to output a result. The application solves the problem of low far-field speech recognition accuracy in the prior art, realizes high-precision speech recognition and semantic understanding through multi-link collaborative optimization, and improves recognition stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of speech recognition and semantic understanding technology, and specifically relates to a high-precision speech recognition and semantic understanding method. Background Technology

[0002] The main problem with existing far-field speech recognition and semantic understanding technologies is that in far-field environments, vocal tract distortion is easily generated during speech signal transmission, and the speaker's speech rate is subject to non-steady-state changes. These two factors, whether alone or in combination, can lead to insufficient accuracy of extracted acoustic features, which in turn causes deviations in subsequent semantic matching and confidence calibration. Ultimately, this results in low speech recognition accuracy and inaccurate semantic understanding, failing to meet the actual needs for high-precision voice interaction in far-field scenarios.

[0003] Based on the above problems, there is an urgent need for a technical solution that can solve the problem of insufficient recognition accuracy caused by vocal tract distortion and non-steady speech rate. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and propose a high-precision speech recognition and semantic understanding method, comprising the following steps: S1: Acquire raw speech signals from multiple sound sources in the far field and generate discrete speech sampling sequences; S2: Perform frequency domain transformation on the discrete speech sampling sequence to extract the amplitude-frequency distortion parameters during the vocal tract transmission process; S3: Divide the discrete speech sampling sequence into speech frame sequences according to the preset frame length, traverse the speech frame sequences, count the inter-frame temporal intervals, and calculate the non-steady speech rate quantization parameters. S4: Extract the original Mel cepstral features, and use the vocal tract distortion acoustic feature correction formula based on the amplitude-frequency distortion variability parameter to correct the distortion of the original Mel cepstral features, thus obtaining the corrected Mel cepstral features. S5: Based on the non-steady speech rate quantization parameters, the speech rate linkage feature weight allocation formula is called to dynamically weight the modified Mel cepstral features to obtain the weighted acoustic features; S6: Construct a preset semantic feature library, map the weighted acoustic features to the semantic feature space corresponding to the preset semantic feature library, and generate an initial semantic matching vector; S7: Based on the weighted acoustic features and the initial semantic matching vector, call the semantic confidence recursive calibration formula to complete the recursive iterative calibration of semantic confidence; S8: Select the optimal semantic match based on the calibrated semantic confidence and output the speech recognition and semantic understanding results.

[0005] Preferably, the step of acquiring original speech signals from multiple sound sources in the far field involves using a multi-channel microphone array as the acquisition carrier. The processor controls the microphone array to synchronously acquire speech signals from multiple sound sources in the far field environment at a sampling frequency of 16kHz. The acquired original speech signals are mono-channel time-domain signals. After acquisition, the processor performs DC removal processing and pre-filtering on the original speech signals. The pre-filtering uses a Butterworth low-pass filter with a cutoff frequency of 8kHz. After processing, the processor performs discrete sampling on the signal to generate a discrete speech sampling sequence. The frame length of the discrete speech sampling sequence is set to 1024 sampling points.

[0006] In a further preferred embodiment, the step of performing frequency domain transformation on the discrete speech sampling sequence involves the processor performing a 1024-point Fast Fourier Transform on the discrete speech sampling sequence to obtain a frequency domain signal; the amplitude spectrum of the frequency domain signal is extracted, and the difference between the amplitude value of each frequency point and the standard channel transmission amplitude value is calculated. This difference is used as the amplitude-frequency distortion quantization parameter; the amplitude-frequency distortion quantization parameter includes the distortion amplitude difference and distortion phase difference corresponding to each frequency point. The distortion amplitude difference is expressed in decibels, and the distortion phase difference is expressed in radians; after extraction, the amplitude-frequency distortion quantization parameter is stored in the processor cache for subsequent steps of distortion correction of the original Mel-frequency cepstral features.

[0007] In a further preferred embodiment, the step of calculating the non-steady-state speech rate quantization parameters involves the processor traversing the speech frame sequence, statistically analyzing the time-domain interval between every two adjacent speech signals with a statistical precision of 1ms; calculating the physical speech rate of the current speech frame based on the time-domain interval between adjacent frames, where the physical speech rate is calculated as the number of speech frames per second; calculating the difference between the physical speech rates of two adjacent frames to obtain the physical speech rate change difference; and performing a ratio operation between the physical speech rate and the physical speech rate change difference with a standard reference speech rate of 10 frames per second to obtain the normalized speech rate quantization value and the normalized speech rate change difference; after the normalization operation is completed, the non-steady-state speech rate quantization parameters are obtained, which include the normalized speech rate quantization value and the normalized speech rate change difference.

[0008] Furthermore, when correcting the distortion of the original Mel cepstral features, a vocal tract distortion acoustic feature correction formula is used, which is as follows: ; in, , is the m-th corrected Mel cepstral feature, dimensionless, used to characterize acoustic features that can be directly used for subsequent weighted processing after channel distortion correction; The m-th dimension original Mel-Cepstral feature is dimensionless and is a fundamental acoustic feature directly extracted from discrete speech sampling sequences; is the k-th order amplitude-frequency distortion gain coefficient, which is dimensionless and used to adjust the correction strength of the corresponding order amplitude-frequency distortion on the original Mel cepstral features; is the k-th order distortion attenuation coefficient, which is dimensionless and used to control the attenuation rate during the amplitude-frequency distortion correction process; Let be the normalized frequency point of the k-th order speech feature, dimensionless, obtained from the ratio of the physical frequency point to the speech sampling rate, i.e. ,in The physical frequency point, measured in Hz, is the specific frequency value obtained after frequency domain transformation of the discrete speech sampling sequence. The voice sampling rate, measured in Hz, is a fixed sampling frequency set when acquiring the original voice signal. The total order of frequency decomposition is dimensionless and represents the total order after decomposing the frequency domain signal. This is a dimensionless global mean compensation for acoustic features, used to balance the overall amplitude of Mel cepstral features in different dimensions and avoid amplitude shifts in the corrected features. During the correction process, the processor calls this formula, substituting the original Mel cepstral features, amplitude-frequency distortion parameters, and normalized frequency points into the formula to calculate the corrected Mel cepstral features.

[0009] In a further preferred embodiment, when dynamically weighting the modified Mel-Cepstral features, the dynamic weight allocation of each dimension feature is determined based on the non-steady-state speech rate quantization parameters. The calculation logic of the dynamic weight allocation is related to the normalized speech rate quantization value and the difference in normalized speech rate change in the non-steady-state speech rate quantization parameters. Combined with the amplitude characteristics of the modified Mel-Cepstral features, the weight allocation result of each dimension feature is obtained through a preset operation relationship. Then, the modified Mel-Cepstral features are weighted according to the weight allocation result to obtain the weighted acoustic features. During the dynamic weighting process, the processor calls the preset operation logic, substitutes the modified Mel-Cepstral features and non-steady-state speech rate quantization parameters into the operation, and ensures that the weighting process matches the speech rate change trend, thereby improving the relevance of the acoustic features.

[0010] In a further preferred embodiment, when performing recursive iterative calibration of semantic confidence, based on weighted acoustic features and an initial semantic matching vector, the semantic confidence is iteratively updated through a preset recursive logic. During the recursive process, parameters such as confidence recursive forgetting factor, semantic feature matching degree, and semantic-acoustic feature correlation factor are introduced. Each parameter participates in the calculation according to preset values. After initializing the semantic confidence, the calibrated semantic confidence is obtained through multiple rounds of iterative updates. During the recursive iterative calibration process, the processor calls the recursive logic, takes the average value of the semantic feature matching degree of each dimension in the initial semantic matching vector as the initial semantic confidence, substitutes the weighted acoustic features and the initial semantic matching vector into the calculation, sets a fixed number of iterations, updates the semantic confidence after each iteration, and obtains the calibrated semantic confidence after the iteration ends.

[0011] Further optimized, the initial semantic matching degree is calculated by weighted acoustic features and the cosine similarity of each semantic feature in the preset semantic feature library; during the recursive iteration process, when the difference in semantic confidence between two adjacent iterations is less than 0.001, the iteration is terminated early, without needing to complete the preset 10 iterations; the confidence-based forgetting factor is used. The value is 0.3, which is the semantic-acoustic feature correlation factor. The value of is set according to the importance of semantic features. The association factor corresponding to core semantic features is 0.8, and the association factor corresponding to non-core semantic features is 0.2.

[0012] A further preferred step involves constructing a preset semantic feature library, which is stored in the processor's storage unit. The preset semantic feature library contains 10,000 semantic feature vectors, and the dimension of each semantic feature vector is consistent with the dimension of the weighted acoustic features. The next step involves mapping the weighted acoustic features to the semantic feature space corresponding to the preset semantic feature library. The processor executes a feature mapping algorithm to project the weighted acoustic features onto the preset semantic feature space, and calculates the matching degree between the weighted acoustic features and each semantic feature vector in the preset semantic feature library. The matching degree is calculated using cosine similarity. After calculation, the top 5 semantic feature vectors with the highest matching degrees are used as initial semantic matching vectors. The initial semantic matching vectors contain the semantic feature vectors and their corresponding matching degrees.

[0013] In a further optimized step, the processor compares the calibrated semantic confidence score with a preset confidence threshold of 0.7 in the step of outputting the speech recognition and semantic understanding results. When the calibrated semantic confidence score is greater than or equal to 0.7, the processor directly outputs the corresponding optimal semantic match as the final result. When the calibrated semantic confidence score is less than 0.7, the processor re-invokes the semantic confidence score recursive calibration formula and adjusts the confidence score recursive forgetting factor. The value is 0.2, and the recursive iterative calibration is performed again. If the confidence is still less than 0.7 after recalibration, the semantic matching item with the highest matching degree is output and marked as insufficient confidence. The final output format is text format, which includes speech recognition text and corresponding semantic understanding results. After output, it is stored in the storage unit and sent to the external display device through the communication interface.

[0014] Technical effects: The inventive technical point of this invention is that it corrects the original Mel-frequency cepstral features by using amplitude-frequency distortion parameters, dynamically weights acoustic features by combining non-steady-state speech rate quantization parameters, and recursively calibrates semantic confidence through a collaborative design. This accurately solves the main problems of low recognition accuracy caused by far-field speech transmission tract distortion and non-steady-state speech rate in the background technology, and achieves high-precision speech recognition and semantic understanding, ensuring recognition stability and accuracy. Attached Figure Description

[0015] Figure 1 The flowchart shows a far-field speech recognition and semantic understanding method based on amplitude-frequency distortion correction and speech rate weighting. Figure 2 This is a schematic diagram of the hardware architecture of a far-field speech recognition and semantic understanding system. Figure 3 Flowchart of the semantic feature matching and confidence recursive calibration method; Figure 4 This is a schematic diagram illustrating the entire process of far-field speech recognition and semantic understanding. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0017] Technical problems with existing technologies: In far-field environments, vocal tract distortion is prone to occur during the transmission of speech signals. At the same time, the speaker's speech rate is subject to non-steady-state changes. These two factors, whether alone or in combination, can lead to insufficient accuracy of extracted acoustic features, which in turn causes deviations in subsequent semantic matching and confidence calibration, ultimately resulting in low speech recognition accuracy and inaccurate semantic understanding.

[0018] Based on this, please refer to Figures 1-4This embodiment provides a high-precision speech recognition and semantic understanding method, including: acquiring original speech signals from multiple far-field sources to generate a discrete speech sampling sequence; performing frequency domain transformation on the discrete speech sampling sequence to extract amplitude-frequency distortion quantization parameters during vocal tract transmission; dividing the discrete speech sampling sequence into speech frame sequences according to a preset frame length, traversing the speech frame sequences, calculating the inter-frame time interval, and calculating non-steady-state speech rate quantization parameters; extracting the original Mel-frequency cepstral features, and applying the vocal tract distortion acoustic feature correction formula based on the amplitude-frequency distortion quantization parameters to correct the distortion of the original Mel-frequency cepstral features. The corrected Mel-Cepstral features are obtained; based on the non-steady-state speech rate quantization parameters, the speech rate linkage feature weight allocation formula is called to dynamically weight the corrected Mel-Cepstral features to obtain weighted acoustic features; a preset semantic feature library is constructed, and the weighted acoustic features are mapped to the semantic feature space corresponding to the preset semantic feature library to generate an initial semantic matching vector; based on the weighted acoustic features and the initial semantic matching vector, the semantic confidence recursive calibration formula is called to complete the recursive iterative calibration of semantic confidence; the optimal semantic matching item is selected according to the calibrated semantic confidence, and the speech recognition and semantic understanding results are output. The operation of acquiring far-field multi-source raw speech signals is completed by the hardware processor controlling the speech acquisition device. The acquired raw speech signals are continuous time-domain electrical signals. The processor discretizes the continuous time-domain electrical signals according to a fixed sampling rule to obtain discrete speech sampling sequences, which serve as the basic data carrier for all subsequent data processing. The frequency domain transformation of discrete speech sampling sequences is performed using a frequency domain analysis algorithm. Frequency domain transformation converts the time-domain speech signal into a frequency-domain signal distribution, thereby extracting the amplitude-frequency distortion parameters caused by the vocal tract transmission during speech signal transmission. These parameters directly reflect the degree of distortion in the amplitude and frequency of the speech signal caused by the vocal tract. The operation of dividing the discrete speech sampling sequence into speech frame sequences according to a preset frame length involves dividing continuous discrete sampled data into several fixed-length data units, each corresponding to one frame of speech signal. During the traversal of the speech frame sequence, the time-domain interval between adjacent frames is calculated. The value of the time-domain interval directly reflects the speaker's speaking pace. Based on the value of the time-domain interval, the non-steady-state speech rate quantization parameters are calculated, which fully characterize the real-time changes in the speaker's speech rate. The original Mel-Cepstral features are the core acoustic features extracted from speech signals, which can characterize the essential acoustic properties of speech signals. By correcting the distortion of the original Mel-Cepstral features based on amplitude-frequency distortion parameters, feature distortion caused by vocal tract transmission can be eliminated, restoring the true acoustic features of the speech signal. The corrected Mel-Cepstral features have eliminated distortion interference and have higher accuracy.A dynamic weighting operation is performed on the corrected Mel-Cepstral features. The weight ratios of different acoustic features are adjusted according to the non-steady-state speech rate quantization parameters, making the acoustic feature representation adaptable to the speaker's real-time speech rate. The resulting weighted acoustic features can more accurately match the actual pronunciation state of the speech signal. A pre-set semantic feature library stores standardized semantic feature data. Mapping the weighted acoustic features to the corresponding semantic feature space of the pre-set semantic feature library completes the conversion from acoustic features to semantic features. The generated initial semantic matching vector contains the preliminary matching results between acoustic and semantic features. A recursive iterative calibration operation of semantic confidence combines the weighted acoustic features and the initial semantic matching vector, iteratively calculating and optimizing the semantic matching confidence value through multiple iterations. Semantic data with large matching deviations are eliminated. The calibrated semantic confidence can truly reflect the reliability of semantic matching. Based on the calibrated semantic confidence, the optimal semantic matching item is selected, choosing the semantic data with the highest confidence value as the final matching result. The final output speech recognition and semantic understanding results completely restore the text content and semantic meaning of the far-field multi-source speech signal.

[0019] This technical solution addresses the recognition bias caused by far-field speech transmission distortion and non-steady-state changes in speech rate through a collaborative process encompassing speech signal acquisition, distortion extraction, speech rate calculation, feature correction, feature weighting, semantic mapping, confidence calibration, and result output. All execution steps are completed by the hardware processor in conjunction with the algorithm program. Those skilled in the art can fully implement the technical solution based on the above process without any creative effort.

[0020] The existing technology has the following technical problems: When collecting raw speech signals from multiple sound sources in the far field, there is a lack of specific acquisition carriers and signal preprocessing methods, which results in the raw signals containing interference such as clutter and DC components, affecting the accuracy of subsequent feature extraction.

[0021] Based on this, the steps for acquiring original speech signals from multiple sound sources in the far field involve using a multi-channel microphone array as the acquisition carrier. The processor controls the microphone array to synchronously acquire speech signals from multiple sound sources in the far field environment at a sampling frequency of 16kHz. The acquired original speech signals are mono-channel time-domain signals. After acquisition, the processor performs DC removal and pre-filtering on the original speech signals. The pre-filtering uses a Butterworth low-pass filter with a cutoff frequency of 8kHz. After processing, the processor performs discrete sampling on the signal to generate a discrete speech sampling sequence. The frame length of the discrete speech sampling sequence is set to 1024 sampling points.

[0022] The multi-channel microphone array consists of multiple pickup units arranged in a fixed spatial layout. This allows for simultaneous reception of sound signals from different directions in the far-field environment, avoiding the limitations of a single pickup unit. The processor sends a synchronous acquisition command to the multi-channel microphone array, enabling all pickup units to start signal acquisition at the same time. The 16kHz sampling frequency conforms to the audio signal acquisition standard, preserving the acoustic details of the audio signal. The acquired mono time-domain signal eliminates redundant data from the multi-channel signal, reducing the computational load of subsequent data processing. DC removal is performed using the processor's built-in DC component elimination algorithm, which directly filters out the constant DC voltage component in the original audio signal, eliminating signal offset issues caused by power supply interference. The Butterworth low-pass filter features a flat passband amplitude-frequency characteristic, and its 8kHz cutoff frequency accurately filters out high-frequency noise interference outside the audio signal frequency band, preserving the effective frequency components of the audio signal and preventing high-frequency noise from mixing into subsequent feature extraction processes. Discrete sampling converts the preprocessed continuous time-domain signal into discrete digital data. Each frame of discrete speech sampling sequence consists of 1024 sampling points. The fixed frame length setting ensures that subsequent operations such as frequency domain transformation and speech rate statistics have a unified data standard. The processor stores the generated discrete speech sampling sequence in a temporary storage area for subsequent steps to call.

[0023] This technical solution improves the quality of the original speech signal from the source by defining the acquisition carrier, sampling parameters, preprocessing procedures, and sampling rules. It eliminates various interferences in the acquisition process and provides clean basic data for subsequent acoustic feature extraction and semantic processing. Those skilled in the art can complete the acquisition and processing of the original speech signal by following the above acquisition and preprocessing procedures. The operation process is clear and there are no implementation obstacles.

[0024] The existing technology has the following technical problems: When performing frequency domain transformation on discrete speech sampling sequences, there is a lack of specific transformation methods and extraction logic for amplitude-frequency distortion parameters, which makes it impossible to accurately capture vocal tract distortion information and affects the acoustic feature correction effect.

[0025] Based on this, the step of frequency domain transformation of the discrete speech sampling sequence involves the processor performing a 1024-point Fast Fourier Transform on the discrete speech sampling sequence to obtain the frequency domain signal; the amplitude spectrum of the frequency domain signal is extracted, and the difference between the amplitude value of each frequency point and the standard channel transmission amplitude value is calculated. This difference is used as the amplitude-frequency distortion parameter. The amplitude-frequency distortion parameter includes the distortion amplitude difference and distortion phase difference corresponding to each frequency point. The distortion amplitude difference is expressed in decibels, and the distortion phase difference is expressed in radians. After extraction, the amplitude-frequency distortion parameter is stored in the processor cache for subsequent steps of distortion correction of the original Mel-frequency cepstral features.

[0026] The 1024-point Fast Fourier Transform (FFT) is the optimal frequency domain transformation method for discrete speech sampling sequences with fixed frame lengths. The transformation process completely maps the time-domain speech sampling data to the frequency domain, resulting in a frequency domain signal that clearly shows the amplitude and phase distribution of the speech signal at different frequency points. Amplitude spectrum extraction directly reads the amplitude data corresponding to each frequency point in the frequency domain signal. The standard channel transmission amplitude value is the standard amplitude data of the speech signal under distortion-free conditions. By subtracting the real-time amplitude from the standard amplitude, the obtained value can intuitively reflect the degree of amplitude distortion caused by channel transmission.

[0027] The amplitude difference of distortion is expressed in decibels, conforming to the amplitude quantization standard in the acoustic field, and can accurately characterize the magnitude of amplitude distortion. The phase difference of distortion is expressed in radians, completely preserving the phase distortion information of the speech signal. The amplitude-frequency distortion quantization parameter includes both amplitude and phase distortion data, comprehensively covering the distortion impact of vocal tract transmission on the speech signal. The processor cache has high-speed data read and write capabilities. Storing the amplitude-frequency distortion quantization parameter in the processor cache can shorten the data retrieval time of subsequent distortion correction steps and improve overall processing efficiency. This technical solution achieves accurate extraction of vocal tract distortion information through fixed-point fast Fourier transform, standardized amplitude difference calculation, and standardized parameter representation. All calculation steps are automatically executed by the processor according to a preset algorithm, with clear data processing logic and stable and reliable calculation results.

[0028] The existing technology has the following technical problems: When calculating the quantization parameters of non-steady speech rate, there is a lack of specific statistical methods and calculation logic, which makes it impossible to accurately capture the non-steady changes in speech rate and affects the pertinence of subsequent feature weighting.

[0029] Based on this, the steps for calculating the non-steady-state speech rate quantization parameters are as follows: The processor traverses the speech frame sequence, statistically analyzing the time interval between each two adjacent speech signals, with a statistical precision of 1ms. The physical speech rate of the current speech frame is calculated based on the time interval between adjacent frames, calculated as the number of speech frames per second. The difference between the physical speech rates of two adjacent frames is calculated to obtain the physical speech rate change difference. The physical speech rate and the physical speech rate change difference are then compared with a standard reference speech rate of 10 frames per second to obtain the normalized speech rate quantization value and the normalized speech rate change difference. After normalization, the non-steady-state speech rate quantization parameters are obtained, including the normalized speech rate quantization value and the normalized speech rate change difference. The processor's traversal of the speech frame sequence reads the data of each frame sequentially according to the generation order of the speech frames, statistically analyzing the time interval between adjacent frames. The 1ms statistical precision can capture subtle changes in the speaker's speech rate, avoiding speech rate data deviations caused by insufficient statistical precision. Physical speech rate is calculated using the number of speech frames per second (fps). The number of frames directly corresponds to the speed of speech; more frames mean a faster speech rate, and fewer frames mean a slower speech rate. This calculation method is intuitive and closely reflects the actual patterns of speech pronunciation. The physical speech rate variation difference is the numerical difference between the physical speech rates of two adjacent frames. The sign of the difference reflects the upward or downward trend of the speech rate, while the absolute value of the difference reflects the magnitude of the speech rate change, fully representing the non-steady-state fluctuations of speech rate. The standard reference speech rate is set to 10 fps, which is the average speech rate in everyday communication scenarios. Ratioing the physical speech rate and the physical speech rate variation difference to the standard reference speech rate eliminates the influence of the time dimension on subsequent calculations. The resulting normalized speech rate quantization value and normalized speech rate variation difference are both dimensionless parameters, adapting to the subsequent feature weighting calculation rules.

[0030] The non-steady-state speech rate quantization parameter includes both the real-time status and trend of speech rate, providing comprehensive and accurate reference data for the subsequent dynamic weighting of acoustic features. All statistical and calculation steps are automatically completed by the processor, and the data processing flow is coherent and without logical loopholes.

[0031] The existing technology has the following technical problems: when correcting distortion of the original Mel cepstral features, there is a lack of specific correction formulas and parameter definitions, which cannot effectively eliminate the feature deviation caused by vocal tract distortion, and the inconsistency of dimensions can easily lead to calculation errors.

[0032] Based on this, when correcting the distortion of the original Mel-Cepstral features, the acoustic feature correction formula for tract distortion is adopted. The acoustic feature correction formula for tract distortion is as follows: ; in , is the m-th corrected Mel cepstral feature, dimensionless, used to characterize acoustic features that can be directly used for subsequent weighted processing after channel distortion correction; The m-th dimension original Mel-Cepstral feature is dimensionless and is a fundamental acoustic feature directly extracted from discrete speech sampling sequences; is the k-th order amplitude-frequency distortion gain coefficient, which is dimensionless and used to adjust the correction strength of the corresponding order amplitude-frequency distortion on the original Mel cepstral features; is the k-th order distortion attenuation coefficient, which is dimensionless and used to control the attenuation rate during the amplitude-frequency distortion correction process; The normalized frequency point of the k-th order speech feature is dimensionless. , The physical frequency point, with the dimension Hz. The speech sampling rate is expressed in Hz. is the total order of frequency point decomposition, which is dimensionless; This is a dimensionless global mean compensation for acoustic features. During the correction process, the processor calls this formula, substituting the original Mel-Cepstral features, amplitude-frequency distortion parameters, and normalized frequency points into the formula to calculate the corrected Mel-Cepstral features. The design of the acoustic feature correction formula for duct distortion is based on the theory of frequency domain distortion compensation for speech signals. This theory posits that the influence of duct transmission on speech acoustic features exhibits an exponential decay law, and that constructing a distortion compensation term using an exponential function can accurately restore the true acoustic features.

[0033] The core operational logic of the formula is divided into three parts: the first part is the basic operation of the original Mel-Cepstral features; the second part is the exponential compensation operation for amplitude-frequency distortion; and the third part is the global amplitude balance compensation operation. The m-th dimension original Mel-Cepstral features are the essential acoustic representation of the speech signal. As the basic data for the correction operation, their dimensionless nature allows them to directly participate in multiplication operations. The k-th order amplitude-frequency distortion gain coefficient and the k-th order distortion attenuation coefficient work together. The gain coefficient amplifies the distortion compensation effect at the corresponding frequency point, while the attenuation coefficient controls the attenuation rate of the compensation. Both are dimensionless parameters, ensuring dimensional consistency in the operation process. The k-th order speech feature normalized frequency point is obtained by the ratio of the physical frequency point to the speech sampling rate. Both the physical frequency point and the speech sampling rate have dimensions in Hz. After the ratio operation, the dimensional influence is eliminated, resulting in a purely dimensionless normalized frequency point. This ensures that the exponential function operation conforms to mathematical rules and avoids logical errors caused by the intervention of physical dimensions in exponential operations. The total order of frequency decomposition corresponds to the number of frequency domain signal decompositions, determining the precision of distortion correction. The global mean compensation of acoustic features is used to balance the amplitudes of Mel-Cepstral features across all dimensions, preventing amplitude shifts in the corrected features and ensuring the reasonable distribution of the corrected features. The derivation of the formula starts from the physical laws of vocal tract distortion, first determining the core form of exponential distortion compensation, then introducing normalized frequency points to eliminate dimensional interference, and finally adding mean compensation to optimize feature distribution. The entire derivation process aligns with the technical principles of speech signal processing and is logically consistent. The processor substitutes the original Mel-Cepstral features, amplitude-frequency distortion parameters, and normalized frequency points into the formula, executing the operations sequentially: summation, exponential operation, multiplication, and addition. The resulting corrected Mel-Cepstral features completely eliminate interference from vocal tract distortion, restoring the true acoustic properties of the speech signal. The formula's computational flow is fixed and the parameters are clearly defined; those skilled in the art can directly call the formula to complete the distortion correction operation without needing to explore the computational logic.

[0034] The existing technology has the following technical problems: When dynamically weighting the modified Mel-Cepstral features, there is a lack of specific weight calculation logic, which makes it impossible to combine speech rate changes to achieve targeted weighting, resulting in insufficient targeting of acoustic features and affecting the accuracy of subsequent semantic matching.

[0035] Based on this, when dynamically weighting the modified Mel-Cepstral features, the dynamic weight allocation for each dimension of the features is determined based on the non-steady-state speech rate quantization parameters. The calculation logic of the dynamic weight allocation is related to the normalized speech rate quantization value and the difference in normalized speech rate change in the non-steady-state speech rate quantization parameters. Combining the amplitude characteristics of the modified Mel-Cepstral features, the weight allocation results for each dimension of the features are obtained through a preset operation relationship. Then, the modified Mel-Cepstral features are weighted according to the weight allocation results to obtain the weighted acoustic features. During the dynamic weighting process, the processor calls the preset operation logic, substituting the modified Mel-Cepstral features and non-steady-state speech rate quantization parameters into the operation to ensure that the weighting process matches the speech rate change trend and improves the relevance of the acoustic features. The core operation logic of the dynamic weight allocation is based on the speech rate linkage feature weight allocation formula, which is as follows: ; in, The ω-th dimension feature is dynamically assigned weights, which are dimensionless. The ω-th repaired positive post-Mel cepstral characteristic is dimensionless; This is the normalized speech rate quantization value for the current speech frame, dimensionless. , The physical speech rate is measured in frames per second. The standard reference speaking rate is measured in frames per second. The normalized speech rate variation difference between adjacent frames is dimensionless. , This represents the difference in physical speech rate, measured in frames per second. The total dimension of acoustic features is dimensionless. This is a dimensionless coefficient for smoothing speech rate changes. The formula is designed based on the correlation between speech acoustic features and speech rate. The speed of speech directly affects the effective dimensions of acoustic features. When the speech rate is fast, the weight of high-frequency acoustic features needs to be increased, and when the speech rate is slow, the weight of low-frequency acoustic features needs to be increased. By jointly operating the normalized speech rate parameter and the corrected acoustic features, dynamic adaptation of weights can be achieved.

[0036] The numerator of the formula is the product of the ω-th repaired Mel-Cepstral feature and the normalized speech rate quantization value of the current speech frame, directly reflecting the effectiveness of this dimension of the feature at the current speech rate. The denominator consists of two terms: the first term is the sum of the products of all repaired Mel-Cepstral features and the normalized speech rate quantization values, used to normalize the weights; the second term is the product of the speech rate change smoothing coefficient and the difference in normalized speech rate changes between adjacent frames, used to smooth weight fluctuations caused by sudden changes in speech rate and avoid abnormal weight values. The derivation of the formula starts from the relationship between speech rate and acoustic features, first constructing a correlation term between feature amplitude and speech rate, then designing a normalized denominator to ensure a reasonable range of weight values, and finally adding a smoothing term to optimize weight stability. The derivation process aligns with the common practice of technical improvements in the field of speech recognition. All parameters are normalized to eliminate the influence of dimensions. The ratio of physical speech rate to standard reference speech rate and the ratio of the difference in physical speech rate changes to standard reference speech rate are dimensionless parameters, ensuring the dimensional homogeneity of the overall formula operation. After substituting the corrected Mel-Cepstral features and non-steady-state speech rate quantization parameters into the formula, the processor first calculates the values ​​of the numerator and denominator, then obtains the dynamically assigned weights of the ω-th dimension features through division, and finally weights the corrected Mel-Cepstral features according to the weight values. The resulting weighted acoustic features are perfectly adapted to the current speaker's speech rate, highlighting the role of effective acoustic features and weakening the interference of ineffective features. The formula has a clear operational logic and complete parameter definitions. Those skilled in the art can directly perform dynamic weighting operations based on the formula. The weighting results are stable and closely match the actual pronunciation state.

[0037] The existing technology has the following technical problems: When completing the semantic confidence recursive iterative calibration, there is a lack of specific recursive logic and parameter settings, which leads to inaccurate confidence calibration and makes it impossible to accurately select the optimal semantic matching item.

[0038] Based on this, when performing recursive iterative calibration of semantic confidence, the semantic confidence is iteratively updated using a pre-defined recursive logic based on weighted acoustic features and an initial semantic matching vector. During the recursion process, parameters such as a confidence recursion forgetting factor, semantic feature matching degree, and semantic acoustic feature association factor are introduced. Each parameter participates in the calculation according to a pre-defined value. After initializing the semantic confidence, the calibrated semantic confidence is obtained through multiple rounds of iterative updates. During the recursive iterative calibration process, the processor calls this recursive logic, using the average of the semantic feature matching degrees of each dimension in the initial semantic matching vector as the initial semantic confidence. The weighted acoustic features and the initial semantic matching vector are substituted into the calculation, and a fixed number of iterations is set. After each iteration, the semantic confidence is updated, and the calibrated semantic confidence is obtained after the iteration ends. The recursive iterative calibration of semantic confidence relies on the semantic confidence recursive calibration formula, which is: ; in Let be the semantic confidence score after the (t+1)th iteration, which is dimensionless. Let be the initial semantic confidence level for the t-th iteration, which is dimensionless; The forgetting factor is a recursive confidence factor, which is dimensionless. Let be the semantic feature matching degree of the j-th dimension, which is dimensionless; The j-th dimension is a weighted acoustic feature, which is dimensionless; is a semantic acoustic feature association factor, dimensionless; The total dimension of semantic features is dimensionless. The formula is designed based on recursive iterative optimization theory, which posits that semantic confidence calibration needs to consider both historical matching results and current matching data. By balancing the weights of these two factors through a forgetting factor, stable optimization of confidence can be achieved. The formula consists of two core parts: the first part is the historical confidence retention term, which retains the effective confidence data from the previous iteration by multiplying the initial semantic confidence in the t-th iteration by the recursive forgetting factor; the second part is the current matching optimization term, which incorporates the current acoustic and semantic feature matching information by summing the products of the j-th dimension semantic feature matching degree, the j-th dimension weighted acoustic feature, and the semantic-acoustic feature correlation factor, and then multiplying this sum with the complement of the forgetting factor. The derivation of the formula starts from the basic principles of recursive optimization, first determining the fusion framework of historical and current data, then introducing the correlation parameters of semantic and acoustic features to ensure that the confidence update closely matches the actual matching effect, and finally setting iterative rules to achieve multiple optimizations. The derivation process conforms to the technical principles of data optimization processing. All parameters are dimensionless. The semantic feature matching degree characterizes the degree of matching between semantic and acoustic features. The semantic-acoustic feature correlation factor adjusts the influence weight of semantic features in different dimensions, with higher values ​​for the correlation factor corresponding to core semantic features to highlight the importance of core semantics. The processor substitutes the initial semantic confidence, weighted acoustic features, and initial semantic matching vector into the formula. In each iteration, the confidence value is updated by summing the historical and current terms. During the iteration process, the accuracy of the confidence is gradually optimized, and false matching data is eliminated. The final calibrated semantic confidence truly reflects the reliability of semantic matching. The formula has a clear recursive logic, well-defined parameter functions, and fixed iteration rules. Those skilled in the art can directly perform semantic confidence calibration based on the formula, and the calibration results are accurate and unbiased.

[0039] The existing technology has the following technical problems: In the process of semantic confidence recursive iterative calibration, the calculation method of the initial semantic matching degree is unclear, and there is a lack of iteration termination conditions and specific parameter values, which leads to low calibration efficiency and unstable results.

[0040] Based on this, the initial semantic matching degree is calculated by the cosine similarity between the weighted acoustic features and each semantic feature in the preset semantic feature library. During the recursive iteration process, when the difference in semantic confidence between two adjacent iterations is less than 0.001, the iteration is terminated early without completing the preset number of iterations. The confidence recursive forgetting factor is set to 0.3, and the semantic acoustic feature association factor is set according to the importance of the semantic features. The association factor corresponding to the core semantic features is set to 0.8, and the association factor corresponding to the non-core semantic features is set to 0.2. The cosine similarity calculation method can accurately quantify the matching degree between the weighted acoustic features and the semantic features in the preset semantic feature library. The calculation process is achieved by the ratio of the vector dot product to the vector magnitude. The calculation result is stable and unbiased. The value of the initial semantic matching degree is directly used as the core data of the initial semantic matching vector, providing a basis for confidence iteration. The iteration is terminated early when the difference in semantic confidence between two adjacent iterations is less than 0.001. This optimization rule improves computational efficiency while ensuring calibration accuracy. At this point, the confidence value has stabilized, and continuing the iteration will not produce a significant optimization effect, thus reducing unnecessary computation time. The confidence recursion forgetting factor is set to 0.3, meaning that the retention rate of historical confidence data is 30%, and the integration rate of current matching data is 70%. This value balances the reference value of historical data with the optimization effect of current data, adapting to the needs of far-field speech recognition scenarios. The semantic acoustic feature association factor is set according to the importance of semantic features. Core semantic features directly determine the core meaning of speech, and the association factor is set to 0.8. Non-core semantic features only play an auxiliary role, and the association factor is set to 0.2. This hierarchical setting highlights the matching weight of core semantics and improves the accuracy of semantic understanding.

[0041] During the recursive iteration process, the processor monitors the confidence difference between two adjacent iterations in real time. When the termination condition is met, the operation stops immediately and the calibrated confidence value is output. All parameter values ​​have been verified in actual scenarios, and the operation rules are fixed and unambiguous. Those skilled in the art can complete efficient and stable confidence calibration by following the above parameter values ​​and iteration rules.

[0042] The existing technology has the following technical problems: When constructing a preset semantic feature library and mapping weighted acoustic features, there is a lack of specific library structure and mapping logic, which leads to inaccurate semantic feature matching and unreasonable generation of initial semantic matching vectors.

[0043] Based on this, the process involves constructing a pre-defined semantic feature library, stored in the processor's memory. This library contains 10,000 semantic feature vectors, each with the same dimension as the weighted acoustic features. The next step is mapping the weighted acoustic features to the corresponding semantic feature space within the pre-defined semantic feature library. The processor executes a feature mapping algorithm to project the weighted acoustic features into the pre-defined semantic feature space, calculating the matching degree between the weighted acoustic features and each semantic feature vector in the pre-defined semantic feature library using cosine similarity. After calculation, the top 5 semantic feature vectors with the highest matching degrees are used as the initial semantic matching vectors. These initial matching vectors contain the semantic feature vectors and their corresponding matching degrees. The processor's memory has a large data storage capacity. The 10,000 semantic feature vectors in the pre-defined semantic feature library cover various speech scenarios, including daily communication, technical terms, and command control, meeting the semantic matching needs of far-field multi-source speech. The consistent dimension of each semantic feature vector with the weighted acoustic features ensures the validity of the vector matching operation. The core function of the feature mapping algorithm is to project data from the acoustic feature space to the semantic feature space, realizing the conversion of acoustic signals into semantic information. The projection process is performed according to a preset spatial mapping relationship, without data loss or distortion. The cosine similarity calculation method can accurately measure the similarity between weighted acoustic features and semantic feature vectors. The higher the matching degree value, the higher the fit between the two. The top 5 semantic feature vectors with the highest matching degree are selected as the initial semantic matching vectors. This can not only retain the effective semantic data with high matching degree, but also provide sufficient candidates for subsequent confidence calibration, avoiding the bias risk of a single matching result.

[0044] The initial semantic matching vector contains both semantic feature vectors and corresponding matching degree values, providing complete basic data for subsequent confidence recursive iterative calibration. The construction rules of the preset semantic feature library, the execution logic of the feature mapping algorithm, and the matching degree calculation and filtering rules are all clear and explicit. Those skilled in the art can directly complete the construction of the semantic feature library and the acoustic feature mapping operation, and the semantic matching results are accurate and comprehensive.

[0045] The existing technology has the following technical problems: when outputting speech recognition and semantic understanding results, it lacks specific confidence judgment logic and result processing methods, resulting in insufficient accuracy of the output results and inability to cope with situations with low confidence.

[0046] Based on this, the steps for outputting speech recognition and semantic understanding results involve the processor comparing the calibrated semantic confidence with a preset confidence threshold. When the calibrated semantic confidence is greater than or equal to the preset confidence threshold, the processor directly outputs the corresponding optimal semantic match as the final result. When the calibrated semantic confidence is less than the preset confidence threshold, the processor re-invokes the semantic confidence recursive calibration formula, adjusts the confidence recursive forgetting factor to 0.2, and performs recursive iterative calibration again. If the confidence is still less than the preset threshold after recalibration, the semantic match with the highest matching degree is output, and it is marked as having insufficient confidence. The final output format is text, containing the speech recognition text and the corresponding semantic understanding result. After output, it is stored in the storage unit and simultaneously sent to an external display device via the communication interface. The preset confidence threshold serves as the judgment standard for result output, effectively filtering semantic data with insufficient matching accuracy and ensuring the reliability of the output results. When the calibrated semantic confidence is greater than or equal to the preset confidence threshold, it indicates that the semantic matching result is accurate and reliable, and directly outputting the optimal semantic match can meet the usage requirements. If the semantic confidence score after calibration is less than the preset confidence threshold, the recursive forgetting factor is adjusted to 0.2 to reduce the retention rate of historical data and increase the optimization weight of the current matching data. The recursive iteration calibration is then performed again to further uncover more accurate matching results. If the confidence score is still less than the preset threshold after recalibration, the semantic match with the highest matching score is output and marked as having insufficient confidence. This provides users with reference results and clearly indicates the reliability level of the results, preventing misjudgment.

[0047] The final result is output in text format, including speech recognition text and semantic understanding results, intuitively presenting the complete information of far-field speech. The output data is simultaneously stored in the storage unit and transmitted to an external display device via a communication interface for convenient real-time viewing by users. The processor controls the entire process of result judgment, recalibration, data output, storage, and transmission. All judgment rules and processing methods are fixed and clear. Those skilled in the art can complete the final result output and display by following the above process. The result output is stable and adaptable to various application scenarios.

[0048] The theoretical design basis of the speech rate-linked feature weight allocation formula and the semantic confidence recursive calibration formula both originate from the fundamental theories of speech signal processing and data optimization, conforming to the common sense of technological improvement in the field of speech recognition. The specific derivation process is closely integrated with the implementation logic of the technical solution. The derivation of the speech rate-linked feature weight allocation formula starts from the intrinsic relationship between speech rate and acoustic features. Non-steady changes in speech rate directly change the effectiveness of acoustic features in different dimensions. A single fixed weight cannot adapt to real-time speech rate changes, so it is necessary to construct a dynamic weight calculation logic linked to the speech rate parameter. In the derivation process, the corrected Mel-Cepstral feature is first determined as the basis for weight calculation. Then, the normalized speech rate quantization value is introduced to represent the real-time speech rate state. The two are multiplied to obtain the effective score of a single-dimensional feature. Subsequently, the effective scores of all dimensions are summed to achieve normalization processing, ensuring that the weight values ​​are within a reasonable range. At the same time, the normalized speech rate change difference and smoothing coefficient are introduced to construct a smoothing term to avoid abnormal fluctuations in weight caused by sudden changes in speech rate. Finally, a complete weight allocation formula is formed. All calculation steps conform to the influence law of speech rate on acoustic features, without any design that violates technical common sense. The derivation of the semantic confidence recursive calibration formula is based on the core idea of ​​recursive iterative optimization. The calibration of semantic confidence needs to take into account both historical matching results and current acoustic semantic matching data. Relying solely on a single data source cannot guarantee the accuracy of the confidence. Therefore, the derivation process first constructs a dual-branch structure of historical confidence retention terms and current matching optimization terms. The weight ratio of the dual branches is adjusted by the confidence recursive forgetting factor. Then, the semantic feature matching degree, weighted acoustic features, and semantic-acoustic feature correlation factors are multiplied and summed to accurately represent the reliability of the current match. Finally, the values ​​of the dual branches are added together to obtain the iterative confidence. The confidence value is gradually optimized through multiple rounds of iteration. The derivation process conforms to the basic logic of data recursive optimization and can achieve accurate calibration of semantic confidence.

[0049] The core function of the vocal tract distortion acoustic feature correction formula is to eliminate the interference of vocal tract distortion on the original Mel-Cepstral features during far-field speech transmission, restoring the true acoustic properties of the speech signal. Each parameter and computational step in the formula has a clear technical meaning and function. The m-th corrected Mel-Cepstral feature is the output of the formula, serving as the core input data for subsequent dynamic weighting. Its dimensionless property ensures that it can directly participate in subsequent calculations, and the value of this parameter directly reflects the accuracy of the corrected acoustic features. The m-th original Mel-Cepstral feature is the basic input data of the formula, directly derived from feature extraction of discrete speech sampling sequences. It retains the original acoustic information of the speech signal and is the core object of distortion correction. The k-th order amplitude-frequency distortion gain coefficient is responsible for adjusting the compensation intensity of the corresponding order amplitude-frequency distortion. For frequencies with high distortion, the gain coefficient will automatically increase the compensation intensity; for frequencies with low distortion, the gain coefficient will decrease the compensation intensity, achieving differentiated and accurate correction. The k-th order distortion attenuation coefficient is used to control the attenuation rate of exponential compensation, preventing over-compensation from causing acoustic feature distortion and ensuring that the corrected features conform to the standard state without distortion. The k-th order normalized frequency point of the speech feature is the key parameter of the formula. It eliminates the dimension by using the ratio of the physical frequency point to the speech sampling rate, ensuring the exponential function operation conforms to basic mathematical rules. Simultaneously, it accurately represents the relative position of the frequency point within the entire speech band, providing a precise reference for distortion correction at different frequencies. The total order of frequency point decomposition determines the fineness of distortion correction; a higher order covers more comprehensive frequencies and results in more accurate correction. The global mean compensation of acoustic features is used to balance the amplitude distribution of all dimensions of Mel-Cepstral features, preventing some dimensions from having excessively high or low amplitudes after correction, thus ensuring the overall stability of the corrected features. The formula's calculation process starts from the original Mel-Cepstral features, first compensating for amplitude-frequency distortion through exponential operations, and then balancing the global amplitude through addition operations. The entire calculation process is dimensionally consistent and logically coherent, effectively solving the acoustic feature deviation problem caused by vocal tract distortion.

[0050] The core value of the speech rate-linked feature weighting formula is to achieve dynamic adaptation between acoustic feature weights and non-steady speech rate, thereby improving the targeting of weighted acoustic features. The parameters and operational logic of each part of the formula are designed around the correlation between speech rate and acoustic features. The dynamically assigned weight of the ω-th dimension feature is the output of the formula, directly determining the weight of the ω-th corrected Mel-Cepstral feature in subsequent semantic matching. Its value range is stable and closely matches practical needs. The ω-th corrected Mel-Cepstral feature is the basis for weight calculation. After correction of vocal tract distortion, it possesses high accuracy and can truly reflect the acoustic properties of speech. The normalized speech rate quantization value of the current speech frame is the core parameter representing the real-time speech rate. It is obtained by the ratio of the physical speech rate to the standard reference speech rate. After eliminating the influence of the time dimension, it can directly participate in the weighting calculation. The faster the speech rate, the higher the value of this parameter, and the corresponding feature weight will also increase accordingly. The difference in normalized speech rate variation between adjacent frames represents the fluctuation state of speech rate. The greater the speech rate fluctuation, the larger the value of this parameter. Combined with the speech rate variation smoothing coefficient, it can effectively smooth weight fluctuations and avoid feature distortion caused by abrupt changes in weight values. The total dimension of acoustic features determines the coverage of weight calculation, ensuring that Mel-spectrum features receive reasonable weight allocation after correction of all dimensions. The speech rate variation smoothing coefficient is used to adjust the degree of influence of speech rate fluctuations on weights, avoiding abnormal weights due to abrupt changes in speech rate and ensuring the stability of weight allocation. The formula uses the numerator to represent the correlation between single-dimensional features and speech rate, and the denominator to achieve weight normalization and smoothing. The calculation process is dimensionally homogeneous and logically rigorous, accurately adapting to the speaker's real-time speech rate state and solving the acoustic feature mismatch problem caused by non-steady-state changes in speech rate.

[0051] The core objective of the semantic confidence recursive calibration formula is to optimize the confidence value of semantic matching through recursive iteration, providing a reliable basis for selecting the optimal semantic match. All parameters and operational structures of the formula serve to accurately calibrate the confidence. The semantic confidence after the (t+1)th iteration is the output of the formula, representing the reliability of the semantic matching after one round of iteration optimization; a higher value indicates a more accurate matching result. The initial semantic confidence in the tth iteration is the result of the previous iteration, serving as the basis for this iteration and retaining valid information from historical matching. The confidence recursive forgetting factor is a core parameter balancing historical and current data; a reasonable value ensures the stability of confidence updates and avoids abrupt numerical changes. The j-th dimension semantic feature matching degree characterizes the degree of matching between the weighted acoustic features and the j-th dimension semantic features, representing the core data of the current matching information. The j-th dimension weighted acoustic feature is a precise acoustic feature adapted to speech rate, providing a reliable acoustic foundation for semantic matching. The semantic acoustic feature association factor adjusts the influence weights of different dimensions of semantic features, highlighting the importance of core semantics and improving the accuracy of semantic understanding. The total dimension of semantic features determines the coverage of confidence calibration, ensuring that semantic features of all dimensions can participate in confidence calculation. The formula integrates historical and current data through a dual-branch structure, and gradually optimizes the confidence value through multiple iterations. The calculation process is free of logical loopholes, and the parameters are fully defined. It can effectively solve the confidence deviation problem caused by the mismatch between acoustic and semantic features, and improve the overall accuracy of speech recognition and semantic understanding.

[0052] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A high-precision speech recognition and semantic understanding method, characterized in that, Includes the following steps: S1: Acquire raw speech signals from multiple sound sources in the far field and generate discrete speech sampling sequences; S2: Perform frequency domain transformation on the discrete speech sampling sequence to extract the amplitude-frequency distortion parameters during the vocal tract transmission process; S3: Divide the discrete speech sampling sequence into speech frame sequences according to the preset frame length, traverse the speech frame sequences, count the inter-frame temporal intervals, and calculate the non-steady speech rate quantization parameters. S4: Extract the original Mel cepstral features, and use the vocal tract distortion acoustic feature correction formula based on the amplitude-frequency distortion variability parameter to correct the distortion of the original Mel cepstral features, thus obtaining the corrected Mel cepstral features. S5: Based on the non-steady speech rate quantization parameters, the speech rate linkage feature weight allocation formula is called to dynamically weight the modified Mel cepstral features to obtain the weighted acoustic features; S6: Construct a preset semantic feature library, map the weighted acoustic features to the semantic feature space corresponding to the preset semantic feature library, and generate an initial semantic matching vector; S7: Based on the weighted acoustic features and the initial semantic matching vector, call the semantic confidence recursive calibration formula to complete the recursive iterative calibration of semantic confidence; S8: Select the optimal semantic match based on the calibrated semantic confidence and output the speech recognition and semantic understanding results.

2. The high-precision speech recognition and semantic understanding method according to claim 1, characterized in that, The steps for acquiring raw speech signals from multiple sound sources in the far field are as follows: A multi-channel microphone array is used as the acquisition carrier. The processor controls the microphone array to synchronously acquire speech signals from multiple sound sources in the far field environment at a sampling frequency of 16kHz. The acquired raw speech signals are mono-channel time-domain signals. After acquisition, the processor performs DC removal and pre-filtering on the raw speech signals. The pre-filtering uses a Butterworth low-pass filter with a cutoff frequency of 8kHz. After processing, the processor performs discrete sampling on the signal to generate a discrete speech sampling sequence. The frame length of the discrete speech sampling sequence is set to 1024 sampling points.

3. The high-precision speech recognition and semantic understanding method according to claim 1, characterized in that, The step of performing frequency domain transformation on the discrete speech sampling sequence involves the processor performing a 1024-point Fast Fourier Transform on the discrete speech sampling sequence to obtain the frequency domain signal. The amplitude spectrum of the frequency domain signal is then extracted, and the difference between the amplitude value at each frequency point and the standard channel transmission amplitude value is calculated. This difference is used as the amplitude-frequency distortion parameter. The amplitude-frequency distortion parameter includes the distortion amplitude difference and distortion phase difference corresponding to each frequency point. The distortion amplitude difference is expressed in decibels, and the distortion phase difference is expressed in radians. After extraction, the amplitude-frequency distortion parameter is stored in the processor cache for subsequent steps of distortion correction of the original Mel-frequency cepstral features.

4. The high-precision speech recognition and semantic understanding method according to claim 1, characterized in that, The steps for calculating the non-steady-state speech rate quantization parameters are as follows: The processor traverses the speech frame sequence, counts the time-domain interval between each two adjacent speech signals, and the statistical precision of the time-domain interval is 1ms; the physical speech rate of the current speech frame is calculated based on the time-domain interval between adjacent frames, and the physical speech rate is calculated as the number of speech frames per second; the difference between the physical speech rates of two adjacent frames is calculated to obtain the physical speech rate change difference; the physical speech rate and the physical speech rate change difference are respectively compared with the standard reference speech rate of 10 frames / second to obtain the normalized speech rate quantization value and the normalized speech rate change difference. After the normalization operation is completed, the non-steady-state speech rate quantization parameters are obtained, which include the normalized speech rate quantization value and the normalized speech rate change difference.

5. The high-precision speech recognition and semantic understanding method according to claim 1, characterized in that, When correcting distortion in the original Mel cepstral features, the acoustic feature correction formula for duct distortion is used. The acoustic feature correction formula for duct distortion is as follows: ; in, , is the m-th corrected Mel cepstral feature, dimensionless, used to characterize acoustic features that can be directly used for subsequent weighted processing after channel distortion correction; The m-th dimension original Mel-Cepstral feature is dimensionless and is a fundamental acoustic feature directly extracted from discrete speech sampling sequences; is the k-th order amplitude-frequency distortion gain coefficient, which is dimensionless and used to adjust the correction strength of the corresponding order amplitude-frequency distortion on the original Mel cepstral features; is the k-th order distortion attenuation coefficient, which is dimensionless and used to control the attenuation rate during the amplitude-frequency distortion correction process; Let be the normalized frequency point of the k-th order speech feature, dimensionless, obtained from the ratio of the physical frequency point to the speech sampling rate, i.e. ,in The physical frequency point, measured in Hz, is the specific frequency value obtained after frequency domain transformation of the discrete speech sampling sequence. The voice sampling rate, measured in Hz, is a fixed sampling frequency set when acquiring the original voice signal. The total order of frequency decomposition is dimensionless and represents the total order after decomposing the frequency domain signal. This is a dimensionless global mean compensation for acoustic features, used to balance the overall amplitude of Mel cepstral features in different dimensions and avoid amplitude shifts in the corrected features. During the correction process, the processor calls this formula, substituting the original Mel cepstral features, amplitude-frequency distortion parameters, and normalized frequency points into the formula to calculate the corrected Mel cepstral features.

6. The high-precision speech recognition and semantic understanding method according to claim 1, characterized in that, When dynamically weighting the modified Mel cepstral features, the dynamic weight allocation of each dimension feature is determined based on the non-steady-state speech rate quantization parameter. The calculation logic of the dynamic weight allocation is related to the normalized speech rate quantization value and the difference between the normalized speech rate change in the non-steady-state speech rate quantization parameter. Combined with the amplitude characteristics of the modified Mel cepstral features, the weight allocation result of each dimension feature is obtained through the preset operation relationship. Then, the modified Mel cepstral features are weighted according to the weight allocation result to obtain the weighted acoustic features. During the dynamic weighting process, the processor calls the preset operation logic, substitutes the corrected Mel-Cepstral features and non-steady-state speech rate quantization parameters into the operation, and ensures that the weighting process matches the speech rate change trend, thereby improving the relevance of the acoustic features.

7. The high-precision speech recognition and semantic understanding method according to claim 1, characterized in that, When performing recursive iterative calibration of semantic confidence, based on weighted acoustic features and an initial semantic matching vector, the semantic confidence is iteratively updated through a preset recursive logic. During the recursive process, parameters such as confidence recursive forgetting factor, semantic feature matching degree, and semantic-acoustic feature correlation factor are introduced. Each parameter participates in the calculation according to preset values. After initializing the semantic confidence, the calibrated semantic confidence is obtained through multiple rounds of iterative updates. During the recursive iterative calibration process, the processor calls the recursive logic, takes the average of the semantic feature matching degrees of each dimension in the initial semantic matching vector as the initial semantic confidence, substitutes the weighted acoustic features and the initial semantic matching vector into the calculation, sets a fixed number of iterations, updates the semantic confidence after each iteration, and obtains the calibrated semantic confidence after the iteration ends.

8. The high-precision speech recognition and semantic understanding method according to claim 7, characterized in that, The initial semantic matching degree is calculated by weighting the acoustic features and the cosine similarity between each semantic feature in the preset semantic feature library; During the recursive iteration process, if the difference in semantic confidence between two adjacent iterations is less than 0.001, the iteration is terminated early without having to complete the preset 10 iterations. Confidence-based recursive forgetting factor The value is 0.3, which is the semantic-acoustic feature correlation factor. The value of is set according to the importance of semantic features. The association factor corresponding to core semantic features is 0.8, and the association factor corresponding to non-core semantic features is 0.

2.

9. The high-precision speech recognition and semantic understanding method according to claim 1, characterized in that, The steps include: constructing a preset semantic feature library, which is stored in the processor's storage unit and contains 10,000 semantic feature vectors, each with the same dimension as the weighted acoustic features; mapping the weighted acoustic features to the corresponding semantic feature space of the preset semantic feature library, whereby the processor executes a feature mapping algorithm to project the weighted acoustic features onto the preset semantic feature space, and calculates the matching degree between the weighted acoustic features and each semantic feature vector in the preset semantic feature library using cosine similarity; after calculation, the top 5 semantic feature vectors with the highest matching degree are used as the initial semantic matching vectors, which contain the semantic feature vectors and their corresponding matching degrees.

10. The high-precision speech recognition and semantic understanding method according to claim 1, characterized in that, In the step of outputting speech recognition and semantic understanding results, the processor compares the calibrated semantic confidence with a preset confidence threshold of 0.7; when the calibrated semantic confidence is greater than or equal to 0.7, it directly outputs the corresponding optimal semantic match as the final result. When the calibrated semantic confidence is less than 0.7, the processor re-invokes the semantic confidence recursive calibration formula and adjusts the confidence recursive forgetting factor. The value is 0.2, and the recursive iterative calibration is performed again. If the confidence is still less than 0.7 after recalibration, the semantic matching item with the highest matching degree is output and marked as insufficient confidence. The final output format is text format, which includes speech recognition text and corresponding semantic understanding results. After output, it is stored in the storage unit and sent to the external display device through the communication interface.