An ai voice intent recognition system of a wearable device

By generating a theoretical time delay bias and a peak-finding search tolerance window in the wearable device, and combining it with binaural microphone signal processing, the problem of distinguishing between the wearer's and others' voices was solved, improving the accuracy of voice recognition in open environments and during movement.

CN122511236APending Publication Date: 2026-08-04宁波瀚时科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610646867.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing voice recognition technology has difficulty effectively distinguishing between the voice of the wearer and the voice of others in wearable devices, leading to false wake-ups and false executions, especially in open environments and during movement.

Method used

By acquiring the device structure constants and inertial measurement unit data of the wearable device, a theoretical time delay bias and peak search tolerance window are generated. Combined with the audio signals collected by the binaural microphones, a cross-power spectrum signal and spatial cross-correlation characteristic curve are generated. The actual speech time delay difference is extracted, sound source classification and judgment parameters are established, and the speaker's own speech is selected while the speech of others is blocked.

Benefits of technology

It improves the wearable device's ability to resist false triggering in open environments and during movement, reduces the probability of non-wearer voice entering the intent recognition process, and ensures that the recognition result comes from the wearer's voice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122511236A_ABST
    Figure CN122511236A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of voice recognition, in particular to an AI voice intention recognition system of a wearable device, which comprises a dynamic tolerance and bias extraction module, obtains device structure constants of the wearable device, generates a theoretical time delay bias, simultaneously extracts gyroscope data and accelerometer data according to an inertial measurement unit of the wearable device, generates a peak searching tolerance window and corresponding time domain tolerance limits. In the application, sound source attribution screening can be completed before semantic recognition, a bystander flag triggers an intention blocking instruction, the probability of non-wearer voice entering the intention recognition process is reduced, a self flag triggers a target voice segment extraction, a subsequent mel frequency acoustic feature matrix and an intention instruction triggering text are ensured to be derived from wearer voice, and therefore the anti-mis-triggering capability of the wearable device in an open environment, close-range multi-person communication and motion wearing scenarios can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech recognition technology, and more particularly to an AI speech intent recognition system for wearable devices. Background Technology

[0002] Speech recognition technology mainly involves the acquisition, processing, analysis, and semantic conversion of human speech signals. Its core lies in converting continuous acoustic signals into text, instructions, or control parameters that can be understood and executed by computing devices.

[0003] Existing technologies primarily focus on the acquisition, processing, analysis, and semantic conversion of human speech signals. In practice, continuous acoustic signals are typically treated as the object to be recognized, with a focus on converting speech content into text, commands, or control parameters. However, factors such as the identity of the sound source, the device's wearing structure, the spatial differences between the two microphones, and head movement often fail to establish stable constraints before semantic conversion. This leads to the easy inclusion of nearby speech, ambient sound from the same direction, and reflected sound fragments in the recognition process along with the wearer's speech. For example, when the wearer is traveling, attending a multi-person meeting, or engaging in outdoor activities, if someone nearby utters words similar to the device's commands, traditional speech processing can easily send the sound content directly to the subsequent recognition stage, causing false wake-ups, false executions, or incorrect text output. Furthermore, head movements and changes in wearing angle alter the time difference between the two microphones, making fixed thresholds or simple acoustic feature judgments unsuitable for dynamic wearing conditions. Therefore, improvements are needed. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing an AI voice intent recognition system for wearable devices.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: an AI voice intent recognition system for wearable devices includes: The dynamic tolerance and bias extraction module obtains the device structure constants of the wearable device and generates the theoretical time delay bias; at the same time, it extracts gyroscope data and accelerometer data from the inertial measurement unit of the wearable device and generates peak search tolerance window and corresponding time domain tolerance limits. The audio feature peak finding module collects two time-domain audio signals from the binaural microphones in the wearable device, performs complex frequency domain mapping, generates a cross-power spectrum signal, extracts the phase features of the cross-power spectrum signal, and generates a spatial cross-correlation feature curve; the peak finding search tolerance window is shifted to a position centered on the horizontal axis where the theoretical time delay offset is located, and the actual speech time delay difference is extracted. The spatial acoustic filtering module calculates the actual speech time delay difference with the theoretical time delay offset to generate an actual delay error value. When the actual delay error value is greater than the time domain tolerance limit, a bystander flag is established; when the actual delay error value is less than or equal to the time domain tolerance limit, a self flag is established, generating sound source classification parameters. Based on the sound source classification parameters and the actual speech time delay difference, an intent to block command and a target speech segment are generated.

[0006] Preferably, the system further includes: The voice intent feature module, in response to the intent blocking command, terminates the intent recognition process of the current audio signal; in response to the personal identifier in the sound source classification judgment parameters, segments the target speech segment into a multi-frame time-domain sequence according to the signal segmentation command; extracts the Mel frequency cepstral coefficient features in the multi-frame time-domain sequence to generate a Mel frequency acoustic feature matrix; and generates intent command trigger text based on the Mel frequency acoustic feature matrix.

[0007] Preferably, the steps for obtaining the peak-finding search tolerance window and the time-domain tolerance limit are as follows: Obtain the device structure constants of the wearable device, read the calibration coordinates of the mouth sound point, the calibration coordinates of the left ear microphone and the calibration coordinates of the right ear microphone, calculate the straight-line distance from the mouth sound point to the left ear microphone and the straight-line distance from the mouth sound point to the right ear microphone respectively, process the absolute value of the difference between the two straight-line distance values, determine the result as the straight-line distance difference, divide the straight-line distance difference by the sound speed constant, and generate the theoretical time delay offset. Synchronously, based on the inertial measurement unit of the wearable device, gyroscope data and accelerometer data are extracted. The axial component of angular velocity in the gyroscope data and the axial component of linear acceleration in the accelerometer data are time-stamped according to the same acquisition time. Axial component records with mismatched timestamps are removed. The amplitude of the retained axial component of angular velocity and the amplitude of the axial component of linear acceleration are jointly quantized time-by-time to generate the intensity of head movement. The numerical variation range of the intensity of head movement during continuous acquisition is read, and the peak-finding horizontal coordinate search boundary is adjusted according to the numerical variation range. The adjusted peak-finding horizontal coordinate search boundary is used as the peak-finding search tolerance window, and the allowable delay deviation range corresponding to the intensity of head movement is determined as the time-domain tolerance limit. The peak-finding search tolerance window and the corresponding time-domain tolerance limit are generated.

[0008] Preferably, the step of obtaining the spatial cross-correlation characteristic curve is as follows: Two time-domain audio signals are acquired using binaural microphones. The acquisition timestamps, sampling amplitude sequences, and sampling frequency identifiers of the two time-domain audio signals are read. The two time-domain audio signals are aligned according to the acquisition timestamps. The aligned two time-domain audio signals are divided into frames with the same frame length. Complex frequency domain mapping is performed on the sampling amplitude sequence in each frame. The complex amplitude corresponding to each frequency point is extracted. The two complex amplitudes are conjugately multiplied according to the same frequency point to generate a cross-power spectrum signal. Read the real part value, imaginary part value, and frequency point arrangement order of each frequency point in the cross power spectrum signal. Determine the phase angle value of each frequency point according to the real part value and imaginary part value. Perform continuous correction on the phase angle value of adjacent frequency points. Perform frequency-time domain transformation on the corrected phase angle value according to the frequency point arrangement order. Arrange the transformed horizontal axis delay value and vertical axis correlation strength value accordingly to generate a spatial cross-correlation characteristic curve.

[0009] Preferably, the step of obtaining the actual speech time delay difference is as follows: Adjust the center x-coordinate of the peak-finding search tolerance window to the x-coordinate position of the theoretical time delay offset, extract the spatial cross-correlation feature curve segment within the coverage area of ​​the peak-finding search tolerance window, compare the correlation intensity values ​​of adjacent y-coordinates in the spatial cross-correlation feature curve segment point by point, extract the extreme points of the y-coordinates where the correlation intensity value of the y-coordinate is greater than the correlation intensity value of the adjacent y-coordinates, and determine the x-coordinate value corresponding to the extreme points of the y-coordinates as the actual speech time delay difference.

[0010] Preferably, the step of obtaining the actual delay error value is as follows: The actual speech time delay difference and the theoretical time delay offset are adjusted to the same delay coordinate scale. The horizontal coordinate unit identifier of the actual speech time delay difference and the horizontal coordinate unit identifier of the theoretical time delay offset are verified. The delay values ​​with inconsistent unit identifiers are scaled. The absolute value of the difference between the scaled actual speech time delay difference and the theoretical time delay offset is processed, and the processed delay deviation is used as the actual delay error value.

[0011] Preferably, the step of acquiring the target speech segment is as follows: The actual delay error values ​​are written into the sound source determination comparison sequence one by one. The actual delay error values ​​and the time domain tolerance limit are compared one by one according to the sound source determination comparison sequence. If the actual delay error value is greater than the time domain tolerance limit, an outsider flag is established. If the actual delay error value is less than or equal to the time domain tolerance limit, a self flag is established. The outsider flag or the self flag is written into the same determination field to generate sound source classification determination parameters. The system reads the user identifier from the sound source classification parameters, retrieves the sampling amplitude sequence, acquisition timestamp, channel arrangement identifier, and actual speech time delay difference of the two time-domain audio signals, aligns the sampling points of the two time-domain audio signals according to the acquisition timestamp, and performs time-domain phase compensation on the two time-domain audio signals based on the actual speech time delay difference. It sets the channel weights of the left and right ear microphones according to the channel arrangement identifier, multiplies the left ear microphone channel weight by the corresponding sampling amplitude sequence of the left ear microphone, and multiplies the right ear microphone channel weight by the corresponding sampling amplitude sequence of the right ear microphone. It sums the weighted sampling amplitudes at the same sampling time, retains continuous sampling segments that satisfy the effective amplitude range of speech after summing, and obtains the target speech segment.

[0012] Preferably, the step of obtaining the Mel frequency acoustic feature matrix is ​​as follows: In response to the intent blocking command, the flow status identifier of the current audio signal is read, and the flow status identifier of the current audio signal is set to the termination state to stop the current audio signal from entering the subsequent intent recognition node. In response to the personal flag in the sound source classification and determination parameters, the target speech segment and signal segmentation command are invoked, and the sampling timestamp, sampling amplitude sequence, and start and end boundary identifier of the target speech segment are read. According to the frame length boundary and frame shift boundary specified by the signal segmentation command, the sampling amplitude sequence of the target speech segment is continuously segmented to generate a multi-frame time domain sequence. The sampled amplitude sequence, frame number identifier, and frame start and end timestamps within the multi-frame time-domain sequence are read frame by frame. Frequency component expansion is performed on the sampled amplitude sequence of each frame. The energy distribution value corresponding to each frequency component is extracted according to the Mel frequency scale. The energy distribution value is converted into cepstral component arrangement value. The cepstral component arrangement value is sequentially collected according to the frame number identifier to form Mel frequency cepstral coefficient features. Component normalization division is performed on each dimension of the cepstral component arrangement value in the Mel frequency cepstral coefficient features to generate the Mel frequency acoustic feature matrix.

[0013] Preferably, the step of obtaining the intent command trigger text is as follows: The feature coordinates, frame number identifiers, and cepstral dimension identifiers within the Mel frequency acoustic feature matrix are read line by line. The feature coordinates are then substituted into the dictionary state node probability set according to the frame number identifiers. The likelihood values ​​corresponding to the same candidate state node sequence are accumulated one by one. The likelihood values ​​of each candidate state node sequence are compared, and the candidate state node sequence with the largest likelihood value is selected. The state node sequence mapping vocabulary parameters corresponding to the candidate state node sequence are extracted. The text is concatenated according to the word order of the state node sequence mapping vocabulary parameters to generate the intent command trigger text.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, the theoretical time delay bias is extracted from the structural constant of the wearable device, and combined with gyroscope and accelerometer data to form a peak-finding search tolerance window and corresponding time-domain tolerance limits. This allows speech signal processing to no longer rely solely on the acoustic content itself, but instead incorporates the wearing structure, mouth articulation point, relative positions of the binaural microphones, and intensity of head movement into the pre-judgment process. Cross-power spectrum signals are generated from the two time-domain audio signals, and phase features are further extracted to form a spatial cross-correlation feature curve. The peak-finding search tolerance window is then shifted to the horizontal axis where the theoretical time delay bias is located to extract the actual speech time delay difference, transforming sound source localization from blind peak-finding across the entire domain. Peak finding is constrained around the wearer's vocal position to reduce interference from other people's voices, environmental reflections, and motion disturbances on peak value judgment. By comparing the actual speech time delay difference, theoretical time delay offset, and time domain tolerance limit, a personal marker or a bystander marker is established. Sound source attribution can be screened before semantic recognition. The bystander marker triggers an intent blocking command, reducing the probability of non-wearer voices entering the intent recognition process. The personal marker triggers target speech segment extraction, ensuring that the subsequent Mel frequency acoustic feature matrix and intent command trigger text originate from the wearer's voice. Therefore, it can improve the wearable device's anti-false triggering capability in open environments, close-range multi-person communication, and sports wearing scenarios. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the principle of the present invention; Figure 2 A schematic diagram showing the peak-finding search tolerance window width corresponding to the average intensity of head movements; Figure 3 This is a schematic diagram of the actual speech time delay difference in the spatial cross-correlation characteristic curve. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0017] Please see Figure 1-3 The present invention provides a technical solution: an AI voice intent recognition system for wearable devices includes: The dynamic tolerance and bias extraction module obtains the device structure constants of the wearable device and generates the theoretical time delay bias. At the same time, it extracts gyroscope and accelerometer data from the inertial measurement unit of the wearable device and generates peak search tolerance window and corresponding time domain tolerance limits. The audio feature peak finding module collects two time-domain audio signals from the binaural microphones in the wearable device, performs complex frequency domain mapping, generates a cross-power spectrum signal, extracts the phase features of the cross-power spectrum signal, and generates a spatial cross-correlation feature curve; the peak finding search tolerance window is shifted to the position centered on the horizontal axis where the theoretical time delay offset is located, and the actual speech time delay difference is extracted. The spatial acoustic filtering module calculates the difference between the actual speech time delay and the theoretical time delay bias to generate an actual delay error value. When the actual delay error value is greater than the time domain tolerance limit, a bystander flag is established; when the actual delay error value is less than or equal to the time domain tolerance limit, a self flag is established, generating sound source classification parameters. Based on the sound source classification parameters and the actual speech time delay difference, an intention blocking command and a target speech segment are generated. The voice intent feature module, in response to an intent blocking command, terminates the intent recognition process of the current audio signal; in response to the self-identity flag in the sound source classification judgment parameters, it segments the target speech segment into a multi-frame time-domain sequence according to the signal segmentation command; it extracts the Mel frequency cepstral coefficient features in the multi-frame time-domain sequence to generate a Mel frequency acoustic feature matrix; and it generates the intent command trigger text based on the Mel frequency acoustic feature matrix.

[0018] The steps for obtaining the peak-finding search tolerance window and time-domain tolerance limits are as follows: Obtain the device structure constants of the wearable device, read the calibration coordinates of the mouth sound point, the calibration coordinates of the left ear microphone and the calibration coordinates of the right ear microphone, calculate the straight-line distance from the mouth sound point to the left ear microphone and the straight-line distance from the mouth sound point to the right ear microphone respectively, process the absolute value of the difference between the two straight-line distance values, determine the result as the straight-line distance difference, divide the straight-line distance difference by the sound speed constant, and generate the theoretical time delay offset. Synchronously, based on the inertial measurement unit of the wearable device, gyroscope data and accelerometer data are extracted. The axial component of angular velocity in the gyroscope data and the axial component of linear acceleration in the accelerometer data are time-stamped according to the same acquisition time. Axial component records with mismatched timestamps are removed. The amplitude of the retained axial component of angular velocity and the amplitude of the axial component of linear acceleration are jointly quantized time-by-time to generate the intensity of head movement. The numerical variation range of head movement intensity during continuous acquisition is read, and the peak-finding horizontal coordinate search boundary is adjusted according to the numerical variation range. The adjusted peak-finding horizontal coordinate search boundary is used as the peak-finding search tolerance window, and the allowable delay deviation range corresponding to the intensity of head movement is determined as the time-domain tolerance limit. The peak-finding search tolerance window and the corresponding time-domain tolerance limit are generated.

[0019] Specifically, the device structure constants of the wearable device are obtained by reading the pre-calibrated coordinates of the mouth sound point, the left ear microphone, and the right ear microphone from the device's non-volatile storage. These three coordinates are all three-dimensional Cartesian coordinates. Then, the straight-line distance from the mouth sound point to the left ear microphone and the straight-line distance from the mouth sound point to the right ear microphone are calculated using the Euclidean distance formula. Next, the absolute value of the difference between these two calculated straight-line distance values ​​is processed, and the result is determined as the straight-line distance difference. Finally, the straight-line distance difference is divided by the sound speed constant at 20 degrees Celsius under standard atmospheric pressure (e.g., 343 m / s) to generate a reference theoretical time delay offset.

[0020] Synchronously, based on the inertial measurement unit of the wearable device, three-axis angular velocity data from the gyroscope and three-axis linear acceleration data from the accelerometer are continuously extracted at a sampling frequency of, for example, 100 times per second. To ensure the time synchronization of the data sources, the axial components of the angular velocity in the gyroscope data and the axial components of the linear acceleration in the accelerometer data are strictly time-stamped according to the acquisition timestamp of each data frame, and all axial component records whose timestamps do not match precisely are discarded. For each retained synchronous sampling time point, the Euclidean norm of the angular velocity vector and the linear acceleration vector, i.e., the magnitude of the angular velocity and the magnitude of the linear acceleration, are calculated respectively. Then, these two magnitudes are jointly quantized time-by-time. The quantization process is implemented through a weighted summation model, which multiplies the angular velocity magnitude by a weighting coefficient. (e.g., 0.6), multiply the linear acceleration magnitude by a weighting factor. (For example, 0.4), then multiply the two and add them together, where the weighting coefficient and It is an empirical value obtained based on statistical analysis of a large amount of experimental data, used to balance the influence of rotational and linear motion on changes in head posture, thereby obtaining a continuous time series of head movement intensity.

[0021] The system reads the changes in head movement intensity over the most recent consecutive acquisition time (e.g., the past 500 milliseconds) and calculates the average value over this period to quantify the user's current average head movement intensity. Based on the magnitude of this average value, the peak-finding horizontal axis search boundary is dynamically adjusted. Specifically, the width of the peak-finding search tolerance window consists of a base width and a dynamic width. For example, the base width is set to 2 milliseconds to cover the small measurement errors when the head is completely still. The dynamic width is obtained by multiplying the average head movement intensity by a width adjustment coefficient (e.g., 0.02 milliseconds / unit intensity). If the calculated average intensity is 50 units, then the total width is... The adjusted peak-finding x-axis search boundary is used as the peak-finding search tolerance window. Simultaneously, the allowable delay deviation range corresponding to the intensity of head movement is determined as the time-domain tolerance limit. Its value also consists of a base tolerance (e.g., 0.5 milliseconds) and a dynamic tolerance. The dynamic tolerance is calculated by multiplying the average intensity of head movement by a tolerance adjustment coefficient (e.g., 0.01 milliseconds / unit of intensity). In the example above, the time-domain tolerance limit is... Milliseconds are used to generate a peak-finding search tolerance window and corresponding temporal tolerance limits that are adapted to the current head movement state.

[0022] The steps for obtaining the spatial cross-correlation characteristic curve are as follows: Two time-domain audio signals are acquired using binaural microphones. The acquisition timestamps, sampling amplitude sequences, and sampling frequency identifiers of the two time-domain audio signals are read. The two time-domain audio signals are aligned according to the acquisition timestamps. The aligned two time-domain audio signals are divided into frames with the same frame length. Complex frequency domain mapping is performed on the sampling amplitude sequence in each frame. The complex amplitude corresponding to each frequency point is extracted. The two complex amplitudes are conjugately multiplied according to the same frequency point to generate a cross-power spectrum signal. Read the real part, imaginary part, and frequency order of each frequency point in the cross-power spectrum signal. Determine the phase angle value of each frequency point according to the real and imaginary part values. Perform continuous correction on the phase angle values ​​of adjacent frequency points. Perform frequency-time domain transformation on the corrected phase angle values ​​according to the frequency order. Arrange the transformed horizontal axis delay value and vertical axis correlation strength value accordingly to generate a spatial cross-correlation characteristic curve.

[0023] Specifically, based on two time-domain audio signals synchronously acquired by binaural microphones at a sampling rate of 16kHz and a depth of 16 bits, the acquisition timestamps, sampling amplitude sequences, and sampling frequency identifiers of these two signals are read. The starting points of the two signals are aligned based on the acquisition timestamps. The aligned two time-domain audio signals are then framed with a fixed frame length (e.g., 512 sampling points, corresponding to 32 milliseconds) and frame shift (e.g., 256 sampling points, corresponding to 16 milliseconds, i.e., 50% overlap). Before framing, a Hamming window function is applied to each frame to reduce spectral leakage. Then, a Fast Fourier Transform is performed on the sampling amplitude sequence within each frame to map it from the time domain to the complex frequency domain, thereby extracting the complex amplitude corresponding to each discrete frequency point. This complex amplitude contains the amplitude and phase information of that frequency point. Finally, the complex amplitudes of the left and right ear channels at the same frame and the same frequency point are conjugate-multiplied, specifically calculated as follows: ,in, Index representing discrete frequency points. The left ear channel signal frame at the frequency point Complex amplitude at the location, The right ear channel signal frame at the frequency point The conjugate of the complex amplitude at all frequencies Perform this operation to generate a cross-power spectrum signal.

[0024] Read each frequency point in the cross-power spectrum signal The corresponding complex value, and the real part value separated. and imaginary part values And the order of the frequency points, using the arctangent function. Calculate the phase angle value for each frequency point. Because the output range of the arctangent function is limited to arrive This can cause phase jumps, so it's necessary to continuously correct the phase angle values ​​of adjacent frequency points, i.e., phase unrolling. Specifically, this involves iterating from low to high frequencies, and if the absolute value of the phase difference between the current frequency point and the previous frequency point is greater than... Then, add or subtract the phase angle of the current and all subsequent frequency points. Multiples of an integer, until the difference falls into... Within the interval, a continuous phase spectrum is obtained. Subsequently, the corrected phase angle values ​​are converted to the time domain according to the frequency point arrangement. Here, the GCC-PHAT method is used, which performs phase weighting processing on the cross-power spectrum signal. The weighting function is as follows: Then, an inverse fast Fourier transform is performed on the weighted spectrum to obtain the cross-correlation function. Finally, the abscissa (time delay value) in the cross-correlation function is converted. ) and ordinate (related intensity values) Arrange them in a one-to-one correspondence to generate spatial cross-correlation characteristic curves.

[0025] The steps to obtain the actual voice time delay difference are as follows: Adjust the center x-coordinate of the peak-finding search tolerance window to the x-coordinate position of the theoretical time delay offset, extract the spatial cross-correlation feature curve segment within the coverage area of ​​the peak-finding search tolerance window, compare the correlation intensity values ​​of adjacent y-coordinates in the spatial cross-correlation feature curve segment point by point, extract the extreme points of the y-coordinates where the correlation intensity value of the y-coordinate is greater than the correlation intensity value of the adjacent y-coordinates, and determine the x-coordinate value corresponding to the extreme points of the y-coordinates as the actual speech time delay difference.

[0026] Specifically, the center x-coordinate of the peak-finding search tolerance window generated in the previous step is adjusted to the x-coordinate position of the theoretical time delay offset calculated based on the device structure constant. For example, if the theoretical time delay offset is 0.4 milliseconds and the width of the peak-finding search tolerance window is 3 milliseconds, the search interval is set to [0.4 - 1.5, 0.4 + 1.5], or [-1.1, 1.9] milliseconds. Then, from the spatial cross-correlation feature curve generated in the previous step, all curve segments with x-coordinate delay values ​​falling within this search interval are extracted. Next, a peak search is performed on the extracted spatial cross-correlation feature curve segments. Specifically, all data points within the segment are traversed to find the point with the largest y-coordinate correlation strength. The search for this point does not rely on gradients or derivatives but is completed by directly comparing the y-coordinate values ​​of all candidate points. The x-coordinate value of this point with the largest y-coordinate correlation strength is then determined, for example, within [-1.1, 1.9] milliseconds. [1.9] If the correlation intensity peak is found to occur at 0.45 milliseconds within the millisecond interval, then the value of 0.45 milliseconds is determined as the actual speech time delay difference.

[0027] The steps to obtain the actual delay error value are as follows: The actual speech time delay difference and the theoretical time delay offset are adjusted to the same delay coordinate scale. The horizontal coordinate unit of the actual speech time delay difference and the horizontal coordinate unit of the theoretical time delay offset are verified. The delay values ​​with inconsistent unit labels are scaled. The absolute value of the difference between the actual speech time delay difference and the theoretical time delay offset after scale conversion is processed. The processed delay deviation is used as the actual delay error value.

[0028] Specifically, to adjust the actual speech time delay difference and the theoretical time delay offset to the same delay coordinate scale, the first step is to verify the unit of the two values ​​on the horizontal axis. For example, the actual speech time delay difference is usually expressed in the number of sampling points, while the theoretical time delay offset is expressed in milliseconds. For instance, if the audio sampling frequency is 16kHz, the theoretical time delay offset is 0.4 milliseconds, but the actual speech time delay difference is 7 sampling points, then a scale conversion is needed for the delay values ​​with inconsistent unit labels. The actual speech time delay difference expressed in sampling points is converted to milliseconds. The conversion process is to divide the number of sampling points by the number of sampling points per millisecond. After unifying the units of the two values ​​to milliseconds, the absolute value of the difference between the actual speech time delay after scaling (0.4375 milliseconds) and the theoretical time delay offset (0.4 milliseconds) is taken. The calculation process is as follows: The processed delay deviation (0.0375 milliseconds) is taken as the actual delay error value.

[0029] The steps for obtaining the target speech segment are as follows: The actual delay error values ​​are written into the sound source judgment comparison sequence one by one. The actual delay error values ​​and the time domain tolerance limit are compared one by one according to the sound source judgment comparison sequence. If the actual delay error value is greater than the time domain tolerance limit, an outsider mark is established. If the actual delay error value is less than or equal to the time domain tolerance limit, a self mark is established. The outsider mark or self mark is written into the same judgment field to generate sound source classification judgment parameters. The system reads the user identifier from the sound source classification parameters, retrieves the sampling amplitude sequence, acquisition timestamp, channel arrangement identifier, and actual speech time delay difference of the two time-domain audio signals, aligns the sampling points of the two time-domain audio signals according to the acquisition timestamp, and performs time-domain phase compensation on the two time-domain audio signals based on the actual speech time delay difference. It sets the channel weights for the left and right ear microphones according to the channel arrangement identifier, multiplies the left ear microphone channel weight by the corresponding sampling amplitude sequence of the left ear microphone, and multiplies the right ear microphone channel weight by the corresponding sampling amplitude sequence of the right ear microphone. It sums the weighted sampling amplitudes at the same sampling time, retains continuous sampling segments that satisfy the effective amplitude range of the speech after summing, and obtains the target speech segment.

[0030] Specifically, the actual delay error value calculated for each frame is written into a first-in-first-out (FIFO) queue structure for sound source determination comparison. The length of this sequence is preset to 10 to smooth short-term determination fluctuations. Then, the latest actual delay error value is compared with the temporal tolerance limit dynamically generated based on the intensity of the current head movement according to the sound source determination comparison sequence. For example, if the currently calculated actual delay error value is 0.0375 milliseconds and the corresponding temporal tolerance limit is 1.0 millisecond, since 0.0375 milliseconds is less than 1.0 milliseconds, the preliminary determination result of the current frame is "the person". Subsequently, the 10 latest determination results in the sound source determination comparison sequence are statistically analyzed. If the number of frames determined as "the person" exceeds a preset majority voting threshold, such as 7, then the "person" flag is finally established. Conversely, if the actual delay error value is greater than the temporal tolerance limit, or if the number of frames of "the person" in the sequence statistics does not reach 7, then the "bystander" flag is established. Finally, this "bystander" flag or "person" flag obtained through time smoothing and majority voting is written into an independent determination field to generate sound source classification determination parameters.

[0031] The system reads the sound source classification parameters and performs subsequent operations only when the parameters indicate the user's identity. It retrieves the sampling amplitude sequence, acquisition timestamp, channel arrangement identifier, and actual speech time delay difference of the two time-domain audio signals corresponding to the current frame. It precisely aligns the sampling points of the two time-domain audio signals according to the acquisition timestamps and performs time-domain phase compensation on one of the signals (e.g., the right channel signal) based on the actual speech time delay difference. Specifically, it converts the actual speech time delay difference (e.g., 0.45 milliseconds) into a sampling point offset. For each sampling point, since the offset contains decimals, Sinc interpolation or linear interpolation methods are used to calculate the sampling amplitude at non-integer points. For example, in linear interpolation, the new signal at integer points... The amplitude of the original signal is and The amplitude-weighted average at each point is obtained, i.e. Subsequently, according to the channel arrangement identifier, the weights of the left and right ear microphone channels are set respectively, usually set to a fixed weight of 0.5 for equal gain merging. The weight of the left ear microphone channel is multiplied into the sampling amplitude sequence of the left ear microphone, and the weight of the right ear microphone channel is multiplied into the sampling amplitude sequence of the right ear microphone after phase compensation. The weighted sampling amplitudes at the same sampling time are summed to form an enhanced single-channel signal. Finally, the root mean square energy of the single-channel signal within a continuous 10-millisecond window is calculated and compared with a preset lower limit of the effective amplitude range of speech (e.g., -25dBFS, which is set after statistical analysis of the energy of a large number of quiet environment recordings). All continuous sampling segments with energy higher than this lower limit are retained to obtain the target speech segment.

[0032] The steps for obtaining the Mel frequency acoustic characteristic matrix are as follows: In response to the intent blocking command, the flow status identifier of the current audio signal is read and set to the termination state to stop the current audio signal from entering the subsequent intent recognition node. In response to the person flag in the sound source classification and judgment parameters, the target speech segment and signal segmentation command are called. The sampling timestamp, sampling amplitude sequence and start and end boundary identifier of the target speech segment are read. According to the frame length boundary and frame shift boundary specified by the signal segmentation command, the sampling amplitude sequence of the target speech segment is continuously segmented to generate a multi-frame time domain sequence. The sampled amplitude sequence, frame number identifier, and frame start and end timestamps of the multi-frame time-domain sequence are read frame by frame. Frequency component expansion is performed on the sampled amplitude sequence of each frame. The energy distribution value corresponding to each frequency component is extracted according to the Mel frequency scale. The energy distribution value is converted into cepstral component arrangement value. The cepstral component arrangement value is sequentially collected according to the frame number identifier to form Mel frequency cepstral coefficient features. Component normalization division operation is performed on each dimension of the cepstral component arrangement value in the Mel frequency cepstral coefficient features to generate Mel frequency acoustic feature matrix.

[0033] Specifically, in response to the intent blocking command, the process status identifier associated with the current audio signal is read, and its value in the process management table is changed from "running" (e.g., value 1) to "terminated" (e.g., value 0). This modification will cause the scheduler to skip the audio signal in subsequent processing cycles, thus preventing it from entering the subsequent intent recognition node. When responding to the person's flag in the sound source classification parameters, the target speech segment and a preset signal segmentation command are invoked. This command explicitly defines the frame length boundary as 400 sampling points (corresponding to 25 milliseconds at a 16kHz sampling rate) and the frame shift boundary. The process begins with 160 sampling points (corresponding to a duration of 10 milliseconds). Next, the starting sampling timestamp, complete sampling amplitude sequence, and start and end boundary markers of the target speech segment are read. Starting from the first sampling point of the sampling amplitude sequence, a segment of 400 sampling points is extracted as the first frame. Then, the extraction starting point is moved back 160 sampling points, and another segment of 400 sampling points is extracted as the second frame. This segmentation operation is repeated until the end position of the extraction window exceeds the start and end boundary markers of the target speech segment. All extracted segments are arranged in chronological order to generate a multi-frame time-domain sequence.

[0034] The sampling amplitude sequence, frame number identifier, and frame start and end timestamps of a multi-frame time-domain sequence are read frame by frame. For each frame's sampling amplitude sequence (e.g., 400 sampling points), a first-order high-pass pre-emphasis filter with a coefficient of 0.97 is first applied. Then, a Hamming window function is applied to the filtered sequence. Next, a 512-point Fast Fourier Transform is performed on the windowed sequence to expand its frequency components, and its power spectrum is calculated. Then, the power spectrum is passed through a Mel filter bank consisting of 26 triangular filters. The center frequency of this filter bank exhibits a linear distribution below 1 kHz and a logarithmic distribution above 1 kHz, covering a frequency range from 0 Hz to 8000 Hz. The power spectrum of each filter bank is then calculated. The logarithmic energy output from the filter yields a 26-dimensional energy distribution vector. This energy distribution vector is then processed by a discrete cosine transform, retaining the first 13 coefficients. These coefficients are the cepstral component arrangement values ​​for that frame. The cepstral component arrangement values ​​of all frames are sequentially collected according to the frame number to form the initial Mel frequency cepstral coefficient features. Finally, cepstral mean-variance normalization is performed on the cepstral component arrangement values ​​of each dimension (13 dimensions in total) of these Mel frequency cepstral coefficient features. This involves calculating the mean and standard deviation of the entire speech in each dimension, then subtracting the mean from the corresponding dimension of each frame and dividing by the standard deviation to generate the final Mel frequency acoustic feature matrix.

[0035] The steps to obtain the intent command trigger text are as follows: The feature coordinates, frame number identifiers, and cepstral dimension identifiers within the Mel frequency acoustic feature matrix are read line by line. The feature coordinates are then substituted into the dictionary state node probability set according to the frame number identifiers. The likelihood values ​​corresponding to the same candidate state node sequence are accumulated one by one. The likelihood values ​​of each candidate state node sequence are compared, and the candidate state node sequence with the largest likelihood value is selected. The state node sequence mapping vocabulary parameters corresponding to the candidate state node sequence are extracted. The text is then concatenated according to the word order of the state node sequence mapping vocabulary parameters to generate the intent command trigger text.

[0036] Specifically, the feature coordinates, frame number identifiers, and cepstral dimension identifiers within the Mel-frequency acoustic feature matrix are read line by line. Before decoding, an acoustic model (i.e., a dictionary of state node probabilities), a pronunciation dictionary, and a language model need to be constructed. The acoustic model adopts an HMM-GMM architecture, where each phoneme is represented by a three-state HMM, and the emission probability of each state is described by a Gaussian mixture model with 16 Gaussian components. This model is trained using a large-scale annotated speech database via the Baum-Welch algorithm. The pronunciation dictionary (a part of the state node sequence mapping to vocabulary parameters) is a pre-established mapping table of words and phoneme sequences. The language model adopts a ternary grammar model, trained on a large-scale text corpus, providing word order. The probability of a column; during decoding, each row of the Mel frequency acoustic feature matrix (i.e., the feature vector of each frame) is sequentially input into the decoder based on the Viterbi algorithm according to the frame number identifier. The decoder uses the acoustic model to calculate the emission probability of the current feature vector in each HMM state, and combines the transition probabilities provided by the language model and the constraints of the pronunciation dictionary to dynamically construct and expand, for example, paths in the search network. By accumulating the log-likelihood values, the globally optimal candidate state node sequence is found. After all frames have been processed, the final state is backtracked to obtain the candidate state node sequence with the largest likelihood value. Finally, this state sequence is converted back into a word sequence according to the pronunciation dictionary, and the text is concatenated according to the word order to generate the intent command trigger text.

Claims

1. An AI voice intent recognition system for wearable devices, characterized in that, The system includes: The dynamic tolerance and bias extraction module obtains the device structure constants of the wearable device and generates the theoretical time delay bias; at the same time, it extracts gyroscope data and accelerometer data from the inertial measurement unit of the wearable device and generates peak search tolerance window and corresponding time domain tolerance limits. The audio feature peak finding module collects two time-domain audio signals from the binaural microphones in the wearable device, performs complex frequency domain mapping, generates a cross-power spectrum signal, extracts the phase features of the cross-power spectrum signal, and generates a spatial cross-correlation feature curve; the peak finding search tolerance window is shifted to a position centered on the horizontal axis where the theoretical time delay offset is located, and the actual speech time delay difference is extracted. The spatial acoustic filtering module calculates the actual speech time delay difference with the theoretical time delay offset to generate an actual delay error value. When the actual delay error value is greater than the time domain tolerance limit, a bystander flag is established; when the actual delay error value is less than or equal to the time domain tolerance limit, a self flag is established, generating sound source classification parameters. Based on the sound source classification parameters and the actual speech time delay difference, an intent to block command and a target speech segment are generated.

2. The AI ​​voice intent recognition system for wearable devices according to claim 1, characterized in that, The system also includes: The voice intent feature module, in response to the intent blocking command, terminates the intent recognition process of the current audio signal; in response to the personal identifier in the sound source classification judgment parameters, segments the target speech segment into a multi-frame time-domain sequence according to the signal segmentation command; extracts the Mel frequency cepstral coefficient features in the multi-frame time-domain sequence to generate a Mel frequency acoustic feature matrix; and generates intent command trigger text based on the Mel frequency acoustic feature matrix.

3. The AI ​​voice intent recognition system for wearable devices according to claim 1, characterized in that, The steps for obtaining the peak-finding search tolerance window and time-domain tolerance limits are as follows: Obtain the device structure constants of the wearable device, read the calibration coordinates of the mouth sound point, the calibration coordinates of the left ear microphone and the calibration coordinates of the right ear microphone, calculate the straight-line distance from the mouth sound point to the left ear microphone and the straight-line distance from the mouth sound point to the right ear microphone respectively, process the absolute value of the difference between the two straight-line distance values, determine the result as the straight-line distance difference, divide the straight-line distance difference by the sound speed constant, and generate the theoretical time delay offset. Synchronously, based on the inertial measurement unit of the wearable device, gyroscope data and accelerometer data are extracted. The axial component of angular velocity in the gyroscope data and the axial component of linear acceleration in the accelerometer data are time-stamped according to the same acquisition time. Axial component records with mismatched timestamps are removed. The amplitude of the retained axial component of angular velocity and the amplitude of the axial component of linear acceleration are jointly quantized time-by-time to generate the intensity of head movement. The numerical variation range of the intensity of head movement during continuous acquisition is read, and the peak-finding horizontal coordinate search boundary is adjusted according to the numerical variation range. The adjusted peak-finding horizontal coordinate search boundary is used as the peak-finding search tolerance window, and the allowable delay deviation range corresponding to the intensity of head movement is determined as the time-domain tolerance limit. The peak-finding search tolerance window and the corresponding time-domain tolerance limit are generated.

4. The AI ​​voice intent recognition system for wearable devices according to claim 1, characterized in that, The steps for obtaining the spatial cross-correlation characteristic curve are as follows: Two time-domain audio signals are acquired using binaural microphones. The acquisition timestamps, sampling amplitude sequences, and sampling frequency identifiers of the two time-domain audio signals are read. The two time-domain audio signals are aligned according to the acquisition timestamps. The aligned two time-domain audio signals are divided into frames with the same frame length. Complex frequency domain mapping is performed on the sampling amplitude sequence in each frame. The complex amplitude corresponding to each frequency point is extracted. The two complex amplitudes are conjugately multiplied according to the same frequency point to generate a cross-power spectrum signal. Read the real part value, imaginary part value, and frequency point arrangement order of each frequency point in the cross power spectrum signal. Determine the phase angle value of each frequency point according to the real part value and imaginary part value. Perform continuous correction on the phase angle value of adjacent frequency points. Perform frequency-time domain transformation on the corrected phase angle value according to the frequency point arrangement order. Arrange the transformed horizontal axis delay value and vertical axis correlation strength value accordingly to generate a spatial cross-correlation characteristic curve.

5. The AI ​​voice intent recognition system for wearable devices according to claim 1, characterized in that, The steps for obtaining the actual voice time delay difference are as follows: Adjust the center x-coordinate of the peak-finding search tolerance window to the x-coordinate position of the theoretical time delay offset, extract the spatial cross-correlation feature curve segment within the coverage area of ​​the peak-finding search tolerance window, compare the correlation intensity values ​​of adjacent y-coordinates in the spatial cross-correlation feature curve segment point by point, extract the extreme points of the y-coordinates where the correlation intensity value of the y-coordinate is greater than the correlation intensity value of the adjacent y-coordinates, and determine the x-coordinate value corresponding to the extreme points of the y-coordinates as the actual speech time delay difference.

6. The AI ​​voice intent recognition system for wearable devices according to claim 1, characterized in that, The steps for obtaining the actual delay error value are as follows: The actual speech time delay difference and the theoretical time delay offset are adjusted to the same delay coordinate scale. The horizontal coordinate unit identifier of the actual speech time delay difference and the horizontal coordinate unit identifier of the theoretical time delay offset are verified. The delay values ​​with inconsistent unit identifiers are scaled. The absolute value of the difference between the scaled actual speech time delay difference and the theoretical time delay offset is processed, and the processed delay deviation is used as the actual delay error value.

7. The AI ​​voice intent recognition system for wearable devices according to claim 1, characterized in that, The steps for obtaining the target speech segment are as follows: The actual delay error values ​​are written into the sound source determination comparison sequence one by one. The actual delay error values ​​and the time domain tolerance limit are compared one by one according to the sound source determination comparison sequence. If the actual delay error value is greater than the time domain tolerance limit, an outsider flag is established. If the actual delay error value is less than or equal to the time domain tolerance limit, a self flag is established. The outsider flag or the self flag is written into the same determination field to generate sound source classification determination parameters. The system reads the user identifier from the sound source classification parameters, retrieves the sampling amplitude sequence, acquisition timestamp, channel arrangement identifier, and actual speech time delay difference of the two time-domain audio signals, aligns the sampling points of the two time-domain audio signals according to the acquisition timestamp, and performs time-domain phase compensation on the two time-domain audio signals based on the actual speech time delay difference. It sets the channel weights of the left and right ear microphones according to the channel arrangement identifier, multiplies the left ear microphone channel weight by the corresponding sampling amplitude sequence of the left ear microphone, and multiplies the right ear microphone channel weight by the corresponding sampling amplitude sequence of the right ear microphone. It sums the weighted sampling amplitudes at the same sampling time, retains continuous sampling segments that satisfy the effective amplitude range of speech after summing, and obtains the target speech segment.

8. The AI ​​voice intent recognition system for wearable devices according to claim 2, characterized in that, The steps for obtaining the Mel frequency acoustic feature matrix are as follows: In response to the intent blocking command, the flow status identifier of the current audio signal is read, and the flow status identifier of the current audio signal is set to the termination state to stop the current audio signal from entering the subsequent intent recognition node. In response to the personal flag in the sound source classification and determination parameters, the target speech segment and signal segmentation command are invoked, and the sampling timestamp, sampling amplitude sequence, and start and end boundary identifier of the target speech segment are read. According to the frame length boundary and frame shift boundary specified by the signal segmentation command, the sampling amplitude sequence of the target speech segment is continuously segmented to generate a multi-frame time domain sequence. The sampled amplitude sequence, frame number identifier, and frame start and end timestamps within the multi-frame time-domain sequence are read frame by frame. Frequency component expansion is performed on the sampled amplitude sequence of each frame. The energy distribution value corresponding to each frequency component is extracted according to the Mel frequency scale. The energy distribution value is converted into cepstral component arrangement value. The cepstral component arrangement value is sequentially collected according to the frame number identifier to form Mel frequency cepstral coefficient features. Component normalization division is performed on each dimension of the cepstral component arrangement value in the Mel frequency cepstral coefficient features to generate the Mel frequency acoustic feature matrix.

9. The AI ​​voice intent recognition system for wearable devices according to claim 8, characterized in that, The steps for obtaining the intent command trigger text are as follows: The feature coordinates, frame number identifiers, and cepstral dimension identifiers within the Mel frequency acoustic feature matrix are read line by line. The feature coordinates are then substituted into the dictionary state node probability set according to the frame number identifiers. The likelihood values ​​corresponding to the same candidate state node sequence are accumulated one by one. The likelihood values ​​of each candidate state node sequence are compared, and the candidate state node sequence with the largest likelihood value is selected. The state node sequence mapping vocabulary parameters corresponding to the candidate state node sequence are extracted. The text is concatenated according to the word order of the state node sequence mapping vocabulary parameters to generate the intent command trigger text.