A voice data recognition method and system based on an AI voice algorithm

By employing multimodal collaborative triggering, adaptive noise reduction, and dynamic beamforming optimization based on AI speech algorithms, the robustness and adaptability issues of speech data recognition technology in complex scenarios have been addressed. This has enabled highly accurate and coherent speech data recognition, thereby improving the performance and user experience of the voice interaction system.

CN121237092BActive Publication Date: 2026-04-10HUAQIAO UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAQIAO UNIVERSITY
Filing Date
2025-11-14
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing speech data recognition technologies suffer from problems such as insufficient robustness of triggering mechanisms, poor adaptability of noise reduction processing, lack of dynamic beamforming, lack of data storage continuity, and insufficient long-term adaptability in complex scenarios, which affect the performance and user experience of speech interaction systems.

Method used

A method based on AI speech algorithms, including multimodal collaborative triggering acquisition, AI adaptive noise reduction processing, dynamic beam optimization and adjustment, and semantic association caching enhancement, is adopted to achieve high accuracy and robust recognition of speech data through microphone arrays, multimodal sensors, and AI processing units.

Benefits of technology

It significantly improves the accuracy of speech data recognition and system adaptability in complex environments, reduces the false trigger rate, maintains speech clarity, ensures improved signal-to-noise ratio of target sound sources, and establishes coherent storage of cross-scene data, supporting context-coherent intelligent interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237092B_ABST
    Figure CN121237092B_ABST
Patent Text Reader

Abstract

The application discloses a voice data recognition method and system based on an AI voice algorithm, relates to the technical field of AI voice recognition, and solves the problem of low voice data recognition capability, and comprises the following steps: S1, multi-modal cooperative triggering collection: lip muscle electrical signals and voiceprint features are synchronously collected through a multi-modal sensor, an activation instruction is generated through feature fusion, and voice collection is started by triggering; S2, AI adaptive noise reduction processing: a generative adversarial network is used to separate noise from an original audio signal, an environmental noise feature is separated to generate a dynamic noise reduction mask, and the integrity of a human voice feature is preserved; S3, beam dynamic optimization adjustment: based on a reinforcement learning algorithm, real-time audio quality is analyzed, the beam pointing and gain parameters of a microphone array are dynamically adjusted, and a target sound source is focused; and S4, semantic association cache enhancement: real-time semantic analysis is performed on collected voice data, and the application greatly improves the voice data recognition capability of the AI voice algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech recognition, and more particularly to a speech data recognition method based on an AI speech algorithm and a system thereof. BACKGROUND

[0002] With the rapid development of artificial intelligence and Internet of Things technology, voice interaction has become one of the core ways of human-computer interaction, and is widely used in intelligent terminals, vehicle-mounted systems, smart homes, medical assistance and other fields. As the front link of voice interaction, the quality of voice data collection directly determines the accuracy of subsequent speech recognition and semantic understanding, and is the key foundation for ensuring the performance of voice interaction systems.

[0003] However, the existing speech data recognition technology still has many technical bottlenecks in complex scenarios:

[0004] 1. Insufficient robustness of triggering mechanism: Traditional voice collection relies on a single audio energy threshold for triggering, which is easily disturbed by environmental noise and non-target speech in noisy environments (such as public places and industrial workshops), leading to frequent false triggering or missing collection of key speech, seriously affecting system response efficiency and user experience.

[0005] 2. Poor adaptability of noise reduction processing: Existing noise reduction algorithms are mostly based on fixed noise models or single filtering strategies, making it difficult to cope with complex noise scenarios that change dynamically (such as mixed steady-state noise and transient noise). Excessive noise reduction can distort the voice features, while insufficient noise reduction can leave noise interference, both of which can reduce the effective utilization of voice data.

[0006] 3. Lack of dynamicity of beamforming: Traditional microphone array beamforming parameters are mostly fixed, and cannot track the changes in the direction of the target sound source in real time. In scenarios where multiple sound sources are superimposed or the sound source moves, it is difficult to focus on the target speech, resulting in insufficient gain of the target signal and poor interference suppression effect.

[0007] 4. Lack of coherence in data storage: Existing voice data storage mostly uses isolated files, lacks semantic association analysis of voice content, and cannot establish logical associations between data collected across scenarios and time periods, resulting in low utilization of historical data and difficulty in supporting contextually coherent intelligent interaction needs.

[0008] 5. Insufficient long-term adaptability of the system: Environmental noise characteristics and user speech features change dynamically over time, and existing collection systems lack a closed-loop optimization mechanism, so algorithm parameters cannot be updated adaptively, and the collection quality can degrade over time.

[0009] Therefore, the present application discloses a voice data recognition method that can achieve high accuracy, high robustness and high coherence in complex dynamic environments, which is an urgent need to improve the performance of voice interaction systems. SUMMARY

[0010] In order to solve the above technical problems, the application discloses a voice data recognition method and system based on an AI voice algorithm, which can improve the voice data recognition ability through artificial intelligence and effectively solve the above technical problems.

[0011] In order to achieve the above technical effects, the application adopts the following technical solutions:

[0012] A voice data recognition method based on an AI voice algorithm is applied to a collection terminal integrated with a microphone array, a multi-modal sensor and an AI processing unit. The method collects original audio signals through the microphone array, transmits the signals to the AI processing unit after filtering processing, and supports the basic format conversion and storage of single-channel voice data. The method includes the following steps:

[0013] S1, multi-modal cooperative triggering collection:

[0014] The multi-modal sensor synchronously collects lip myoelectric signals and voiceprint features, generates an activation instruction through feature fusion, and triggers voice collection start;

[0015] S2, AI adaptive noise reduction processing:

[0016] The generative adversarial network is used to separate noise from the original audio signal, separate environmental noise features, and generate beam dynamic optimization adjustment;

[0017] S3, beam dynamic optimization adjustment:

[0018] Based on the reinforcement learning algorithm, the real-time audio quality is analyzed, the beam pointing and gain parameters of the microphone array are dynamically adjusted, and the target sound source is focused;

[0019] S4, semantic association cache enhancement:

[0020] The collected voice data is subjected to real-time semantic analysis, an association index with historical interaction data is established, and coherent storage of cross-scene collected data is realized.

[0021] As a further technical solution of the application, the S1 multi-modal cooperative triggering collection method includes:

[0022] S11, lip myoelectric signal collection:

[0023] The flexible patch sensor integrates a high-density myoelectric electrode array and a miniature impedance detection circuit, and the lip skin is attached with a medical-grade conductive gel. The electrode signal is output after pre-processing by a preamplifier differential amplification module and a 50Hz notch filter;

[0024] S12, voiceprint feature extraction:

[0025] Each unit of the microphone array is connected to an independent preamplifier circuit. The initial audio segment is converted into a spectrogram by a short-time Fourier transform and then input into a voiceprint encoder to extract frame-level voiceprint features.

[0026] S13, Feature Fusion:

[0027] A trigger fusion model is constructed, which includes an electromyography feature branch and a voiceprint feature branch. Feature interaction is achieved through a cross-attention layer. The generated activation vector needs to be verified by a dynamic threshold module. The threshold is adjusted in real time according to the environmental noise level, and data acquisition is started when both the duration and feature matching degree are met.

[0028] As a further technical solution of the present invention, the S2AI adaptive noise reduction processing includes:

[0029] S21. Noise Feature Modeling:

[0030] Background noise collected by environmental sensors is decomposed into low-frequency, mid-frequency, and high-frequency noise components by a multi-band frequency divider, and then input into the dual discriminator of the generative adversarial network. The first discriminator focuses on noise or speech differentiation, while the second discriminator evaluates the integrity of human voice features.

[0031] S22, Dynamic Mask Generation:

[0032] The application generator adopts the U-Net architecture. The application encoder outputs a noise feature map, and the application decoder combines an attention gating mechanism to focus on the human voice frequency band, generate a time-frequency domain noise reduction mask, and adaptively attenuate the noise-dominant frequency band.

[0033] S23. Voice Feature Restoration:

[0034] The residual connection network is integrated into multi-scale residual blocks. By using skip connections to fuse the signal features before and after masking, the harmonic structure and prosodic features of the human voice frequency band are enhanced. The restoration process incorporates speech activity detection results to filter out non-speech segments.

[0035] As a further technical solution of the present invention, the S3 beam dynamic optimization adjustment includes:

[0036] S31. Quality Assessment Feedback:

[0037] The application's real-time evaluation module calculates a comprehensive quality score using the Speech Proficiency Index (STOI) and noise attenuation, and generates a state vector for reinforcement learning by combining the azimuth sensor data from the microphone array.

[0038] S32, Strategy Iteration and Optimization:

[0039] The deep deterministic policy gradient algorithm is adopted by the intelligent agent through reinforcement learning, the state space includes noise type classification results, target sound source azimuth and current beam parameters, the action space outputs beam angle adjustment and gain adjustment through continuous control quantity, and the reward function fuses signal-to-noise ratio improvement value and power consumption coefficient to dynamically calculate and optimize output;

[0040] S33, hardware parameter updating:

[0041] The optimized parameters are written into the digital signal processor register of the microphone array through the real-time control bus, the control signal output by the register drives the beam forming circuit through the digital-to-analog conversion module, and sub-second level adjustment of the beam pointing and gain is realized.

[0042] As a further technical solution of the application, the S4 semantic correlation cache enhancement includes:

[0043] S41, real-time semantic analysis:

[0044] The collected voice data are classified and entity extracted through the distilled BERT architecture of the lightweight semantic model, and core semantic labels including actions, objects and scenes are generated;

[0045] S42, historical correlation retrieval:

[0046] The current semantic labels and historical semantic data are converted into low-dimensional vectors through the local sensitive hashing algorithm by using the distributed index structure of the cache database, and the time decay factor is introduced when similarity matching, so that higher weight is given to the high correlation data in the near period;

[0047] S43, dynamic cache strategy:

[0048] Hierarchical storage is realized based on semantic importance score; core semantic data is stored in the cache area, and associated auxiliary data is archived to the large capacity storage area through the algorithm based on semantic compression, and low value data is cleaned up regularly through heat ranking.

[0049] As a further technical solution of the application, the fusion model working method in S13 further includes:

[0050] The outputs of the myoelectric feature branch and the voiceprint feature branch are time-synchronized through the feature alignment module, the cross-attention layer calculates the feature interaction weight through the multi-head attention mechanism, and the weight value is filtered through the gating function; the feature matching degree in the trigger condition is measured by cosine similarity and edit distance, and the dynamic threshold module is associated with the environmental noise sensor data.

[0051] As a further technical solution of the application, the working principle of the double discriminator structure of the generative adversarial network is:

[0052] The first discriminator outputs a noise probability distribution, the second discriminator outputs a human voice feature integrity score, and the generator loss function is a weighted sum of the losses of the two discriminators; a cyclic consistency constraint is introduced in the training process to ensure the semantic consistency of the denoised speech and the original clean speech, and the cycle loss is calculated by the edit distance.

[0053] As a further technical solution of the application, the reinforcement learning agent in S32 further comprises:

[0054] The state encoding module focuses on key environmental features using an attention mechanism; the experience replay pool uses a priority sampling strategy to assign higher sampling probabilities to high-reward samples; and the output layer of the policy network is processed by batch normalization to stabilize parameter updates, and the adjustment amount exceeding the hardware physical limit is filtered by the feasibility verification module before the action is executed;

[0055] The similarity matching method in S42 further comprises:

[0056] The semantic vector is compressed to a low dimension by a knowledge distillation technique; the similarity calculation introduces a domain knowledge graph to assist in correlation reasoning: when the direct similarity is below a threshold, the matching range is expanded through the entity relationship path in the graph; the correlation index adopts an incremental update mechanism, and only the relevant index branches are updated for newly collected data, reducing the computational overhead.

[0057] As a further technical solution of the application, the method further comprises a S5 quality closed-loop optimization step, comprising:

[0058] Periodically analyze the types of recognition errors of the collected data through a confusion matrix, and select difficult example samples to form a fine-tuning data set based on an uncertainty sampling strategy; during transfer learning fine-tuning, use a parameter freezing and layer-by-layer unfreezing strategy to only update the last three layers of the denoising model and the output layer of the reinforcement learning strategy network; the optimization effect is quantitatively evaluated by the collection quality score, and when the score decreases for three consecutive periods, the model is retrained, wherein the quality score at least includes the fusion recognition accuracy, the denoising signal-to-noise ratio, or the collection response speed.

[0059] The application also uses the following technical solutions:

[0060] A voice data recognition system based on an AI voice algorithm, comprising:

[0061] A multi-modal triggering module connected to a multi-modal sensor and an AI processing unit, the multi-modal sensor integrates a flexible lip muscle electrode patch and a voiceprint collection subunit, and the multi-modal triggering module generates an activation instruction through a cross-attention fusion model after receiving the muscle electrical signal and the voiceprint feature, triggering the voice collection to start;

[0062] The AI adaptive noise reduction module is arranged in the AI processing unit, and includes a generative adversarial network and a residual repair network.

[0063] The beam dynamic optimization module is connected with the microphone array and the AI processing unit, and includes a quality evaluation submodule, a reinforcement learning agent and a hardware control interface.

[0064] The semantic association cache module communicates with the AI processing unit and the storage unit, and includes a lightweight semantic analysis submodule, a distributed index submodule and a dynamic storage controller.

[0065] The quality closed-loop optimization module is connected with the AI processing unit, identifies error types through a confusion matrix analysis, fine-tunes the algorithm parameters of the AI adaptive noise reduction module and the beam dynamic optimization module by using transfer learning, and maintains long-term adaptability of the system.

[0066] The beneficial technical effects of the present application are as follows:

[0067] 1. The multi-modal sensor synchronously collects the lip muscle electrical signal and the voiceprint feature, generates an activation instruction through a cross-attention fusion model, and combines a dynamic threshold verification mechanism (which adjusts in real time according to the environmental noise level) to effectively distinguish target speech from environmental noise and non-target sound, avoid missing collection of key speech, and significantly improve the robustness of the triggering mechanism in a complex environment.

[0068] 2. The double-discriminator architecture of the generative adversarial network is combined with the residual repair network to realize dynamic noise separation and human voice feature repair.

[0069] 3. The beam dynamic optimization mechanism is constructed based on the reinforcement learning algorithm, and through real-time quality evaluation feedback and policy iteration, the beam angle and gain of the microphone array are dynamically adjusted at a sub-second level.

[0070] 4. Extract core semantic tags through lightweight semantic parsing, combine distributed index with time decay factor association retrieval mechanism, and establish semantic association of cross-scene collected data. Hierarchical dynamic storage strategy realizes high-speed access of core data and efficient archiving of auxiliary data, providing data support for contextually coherent intelligent interaction.

[0071] 5. Introduce quality closed-loop optimization module, identify error types through confusion matrix analysis, and realize adaptive updating of algorithm parameters through transfer learning and difficult example sample fine-tuning strategy. In the scene of dynamic changes of environmental noise characteristics and user voice features, the system keeps the quality score (fusion of recognition accuracy, signal-to-noise ratio, response speed) stable for a long time, solving the problem of performance degradation over time of traditional systems.

[0072] 6. Through lightweight semantic model (distilled BERT architecture), priority sampling experience replay pool, incremental index update and other optimization designs, the algorithm performance is guaranteed while the computational complexity is reduced. Reinforcement learning reward function integrates power consumption coefficient, dynamically balances data collection quality and hardware energy consumption, and is suitable for mobile terminals, embedded terminals and other resource-constrained terminals. BRIEF DESCRIPTION OF DRAWINGS

[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, a brief introduction will be given below to the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings described below are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings, wherein:

[0074] Figure 1 Flowchart of the speech data recognition method based on AI speech algorithm in the present application;

[0075] Figure 2 Flowchart of the S1 multi-modal cooperative triggering collection described in the present application;

[0076] Figure 3 Flowchart of AI self-adaptive noise reduction processing in the present application;

[0077] Figure 4 Flowchart of beam dynamic optimization adjustment in the present application;

[0078] Figure 5 Flowchart of semantic association cache enhancement in the present application;

[0079] Figure 6 Application embodiment diagram in the present application. DETAILED DESCRIPTION

[0080] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0081] like Figures 1-5 As shown, a speech data recognition method based on AI speech algorithms is applied to a data acquisition terminal integrating a microphone array, a multimodal sensor, and an AI processing unit. The method acquires raw audio signals through the microphone array, processes them through filtering, and then transmits them to the AI ​​processing unit, supporting basic format conversion and storage of single-channel speech data. The method includes the following steps:

[0082] S1. Multimodal collaborative triggering acquisition:

[0083] The system synchronously collects lip electromyography signals and voiceprint features using a multimodal sensor, and generates activation commands through feature fusion to trigger the start of voice acquisition.

[0084] S2, AI adaptive noise reduction processing:

[0085] Generative adversarial networks are used to separate noise from the original audio signal, extract environmental noise features to generate a dynamic noise reduction mask, and preserve the integrity of human voice features.

[0086] S3, Beam dynamic optimization adjustment;

[0087] Based on reinforcement learning algorithms, real-time audio quality is analyzed, and the beam pointing and gain parameters of the microphone array are dynamically adjusted to focus on the target sound source.

[0088] S4, enhanced semantic association caching;

[0089] Real-time semantic analysis is performed on the collected voice data, and an index is established to associate it with historical interaction data, enabling continuous storage of data collected across different scenarios.

[0090] In the above embodiments, the existing technology relies on a single audio energy threshold trigger, which is susceptible to environmental noise interference. This invention constructs a dual verification mechanism through S1 multimodal collaborative triggering acquisition:

[0091] Simultaneously acquire lip electromyography signals (through a high-density electrode array of flexible patch sensors) and voiceprint features (extracted by short-time Fourier transform of microphone array units) to form a dual basis for judgment based on physiological signals and acoustic features;

[0092] A cross-attention fusion model is used to realize feature interaction. The triggering conditions are adjusted in real time through a dynamic threshold module (the threshold is related to the environmental noise level), and a dual verification standard of duration and feature matching degree (cosine similarity + edit distance) is set.

[0093] The mechanism eliminates the limitation of single audio trigger from the signal source dimension, significantly reduces the false trigger rate in noisy environments, and ensures that key speech is not missed through feature fusion.

[0094] In view of the problem that the fixed noise model cannot cope with dynamic noise, the application realizes dynamic noise suppression through S2AI adaptive noise reduction processing:

[0095] The double discriminator architecture of the generative adversarial network (GAN) is adopted, the first discriminator focuses on noise / speech distinction, and the second discriminator evaluates the integrity of the human voice feature, and the noise component is decomposed by a multi-band frequency divider to realize fine modeling;

[0096] The generator generates a time-frequency domain noise reduction mask based on the U-Net architecture and the attention gate mechanism, and adaptively attenuates the noise-dominant frequency band; the residual connection network strengthens the human voice harmonic structure through multi-scale residual blocks, and combines the speech activity detection to filter non-speech segments.

[0097] The scheme breaks through the limitation of the fixed filtering strategy, and realizes "noise reduction without feature loss" in the mixed steady-state and transient noise scene, that is, even if the signal-to-noise ratio is as low as 5dB, the speech intelligibility index can still be improved.

[0098] In view of the solution of "beamforming dynamicity deficiency", the application realizes real-time adaptive focusing through S3 beam dynamic optimization adjustment:

[0099] Based on the reinforcement learning algorithm, a closed-loop optimization mechanism is constructed, the speech intelligibility index (STOI) and the noise attenuation amount are calculated through the real-time evaluation module, and the state vector is generated combined with the azimuth sensor data;

[0100] The reinforcement learning agent using the deep deterministic policy gradient algorithm includes the noise type, the sound source azimuth and the current beam parameter in the state space, and outputs continuous beam angle and gain adjustment amount; the reward function is used to fuse the signal-to-noise ratio improvement value and the power consumption coefficient to dynamically optimize the strategy;

[0101] Subsecond parameter update is realized through the real-time control bus to drive the beamforming circuit to dynamically adjust the pointing and gain.

[0102] The mechanism enables the microphone array to always focus on the target sound source in the sound source movement and multi-sound source superposition scene, and the target signal-to-noise ratio is improved by 15-20dB, effectively suppressing the sidelobe interference.

[0103] In view of the "lack of data storage continuity", the prior art stores speech data in isolation, and lacks semantic association. The application enhances the construction of cross-scene data logic through S4 semantic association cache:

[0104] A distilled BERT lightweight semantic model is used to extract core semantic tags (actions, objects, and scenes) of voice data, and an index basis in the content dimension is established;

[0105] A distributed index structure is combined with a local sensitive hashing algorithm to convert semantic tags into low-dimensional vectors, a time decay factor is introduced to weight highly relevant data in the near time period, and the matching range is expanded through knowledge graph assisted correlation reasoning;

[0106] Based on the semantic importance score, hierarchical storage is realized, core data is stored in the cache area, auxiliary data is archived through semantic compression, and the calculation overhead is reduced through incremental index updating.

[0107] This scheme breaks through the isolated file storage mode, improves the reuse rate of historical data by more than 40%, and provides data support for coherent context interaction.

[0108] In order to solve the problem of performance degradation caused by changes in environment and user characteristics, the application constructs an adaptive updating mechanism through a quality closed-loop optimization step (S5):

[0109] Periodically identify error types through confusion matrix analysis, and use an uncertainty sampling strategy to select difficult example samples to form a fine-tuning data set;

[0110] Based on transfer learning, the algorithm parameters are dynamically updated, and the parameter freezing and layer-by-layer unfreezing strategies are used to update only the last three layers of the denoising model and the output layer of the reinforcement learning strategy network, thereby reducing the calculation cost.

[0111] Set the collection quality score (fusion of recognition accuracy, signal-to-noise ratio, and response speed) as a monitoring index, and trigger model retraining when it decreases for 3 consecutive periods.

[0112] This mechanism ensures that the system can maintain stable collection performance in the scenario of long-term changes in environmental noise characteristics and user voice characteristics, and avoids the performance degradation problem of traditional systems.

[0113] In summary, through the cooperative design of multi-modal triggering, adaptive noise reduction, dynamic beam optimization, semantic association storage, and closed-loop optimization, the present application comprehensively solves the core bottlenecks of existing voice collection technology in complex dynamic environments, and significantly improves the collection quality and system adaptability.

[0114] In further embodiments, the S1 multi-modal cooperative triggering collection method comprises:

[0115] S11, lip myoelectric signal collection;

[0116] A high-density myoelectric electrode array and a miniature impedance detection circuit are integrated through a flexible patch sensor, a medical-grade conductive gel is attached to the lips, and the electrode signal is preprocessed through a preamplifier module and a 50Hz notch filter and then output;

[0117] S12, voiceprint feature extraction;

[0118] Each unit of the microphone array is connected to an independent preamplification circuit, and the collected initial voice segment is converted into a frequency spectrum graph through short-time Fourier transform and input into a voiceprint encoder to extract frame-level voiceprint features;

[0119] S13, feature fusion;

[0120] A trigger fusion model is constructed, the trigger fusion model includes a myoelectric feature branch and a voiceprint feature branch, feature interaction is realized through a cross-attention layer, the generated activation vector needs to be verified through a dynamic threshold module, the threshold is adjusted in real time according to the environmental noise level, and the threshold is started to collect when meeting the double conditions of continuous duration and feature matching degree.

[0121] As a further technical solution of the application, the S2AI adaptive noise reduction processing includes:

[0122] S21, noise feature modeling:

[0123] The background noise collected by the environmental sensor is decomposed into low-frequency, medium-frequency and high-frequency noise components through a multi-band frequency divider, and is input into the two discriminators of the generative adversarial network; wherein the first discriminator focuses on noise or speech distinction, and the second discriminator evaluates the completeness of the human voice feature;

[0124] S22, dynamic mask generation:

[0125] The application encoder outputs a noise feature map, and the application decoder focuses on the human voice frequency band combined with the attention gate mechanism to generate a time-frequency domain noise reduction mask for adaptive attenuation of noise-dominant frequency bands;

[0126] S23, human voice feature repair:

[0127] The residual connection network integrates a multi-scale residual block, fuses the signal features before and after the mask processing through a jump connection, strengthens the harmonic structure and prosody features of the human voice frequency band, and introduces a speech activity detection result to filter non-speech segments in the repair process.

[0128] Lip articulation is inevitably accompanied by muscle activity, and the voiceprint features are unique to individuals. The spatiotemporal correlation between the two can effectively distinguish real speech from environmental noise / non-target sound, and combined with dynamic threshold adjustment, it realizes robust triggering. In implementation, a flexible patch sensor (integrating a 32-channel high-density electromyography electrode array and a miniature impedance detection circuit) is used, which is tightly attached to the skin around the lips through a medical-grade conductive gel, ensuring stable impedance (the impedance detection circuit monitors the contact state in real time). The weak electromyography signals (μV level) collected by the electrode are amplified by a pre-differential amplification module (gain 1000 times), and the power frequency interference is filtered out by a 50Hz notch filter. Finally, the preprocessed electromyography signals with a bandwidth of 0-5kHz are output, providing a clean data source for feature extraction.

[0129] Each unit of the microphone array is independently connected to a low-noise preamplifier circuit (signal-to-noise ratio ≥85dB) to adjust the gain of the initial speech segment (gain range 20-60dB adjustable) to avoid signal saturation or distortion. Short-time Fourier transform (window length 20ms, step length 10ms) is used to convert the time-domain speech signal into a two-dimensional frequency spectrum (containing time-frequency-amplitude information), which is input into the voiceprint encoder based on the deep residual network to extract the frame-level voiceprint feature vector (dimension 128). The output of the electromyography feature branch (extracting the time-domain peak value and frequency center feature of the electromyography signal) and the voiceprint feature branch is calibrated by a time synchronization module (based on hardware timestamp alignment, error ≤1ms) to ensure temporal and spatial consistency. Through 8 cross-attention layers, the feature interaction weight is calculated (the weight value is filtered by the sigmoid gating function to eliminate invalid associations), and a fusion activation vector is generated. The dynamic threshold module receives environmental noise sensor data (noise level is divided into 5 levels), and each level corresponds to a different threshold.

[0130] The technical principle and implementation process of S2AI adaptive noise reduction processing are as follows:

[0131] AI adaptive noise reduction processing is based on generative adversarial learning and residual repair principle, through a three-level processing mechanism of dynamic noise modeling, time-frequency mask generation and human voice feature enhancement, to realize the goal of "precise noise suppression-human voice complete reservation". The core logic is: using GAN double discriminator game training to realize the precise distinction between noise and speech, combining attention mechanism to focus on the key frequency band of human voice, and through residual connection compensation feature loss, to adapt to the complex noise scene of dynamic change.

[0132] The background noise collected by the environmental sensor is decomposed into three components of low frequency (≤200 Hz), medium frequency (200 Hz-3 kHz), and high frequency (≥3 kHz) by a multi-band frequency divider (cutoff frequency 200 Hz / 3 kHz), which correspond to typical noise types such as mechanical noise, speech interference, and electromagnetic noise. The first discriminator of the generative adversarial network (3-layer convolutional neural network) inputs the noise / speech segment and outputs the noise probability distribution (0-1); the second discriminator (convolutional network with attention mechanism) inputs the denoised speech and outputs the completeness score of the human voice features (0-100 points). The double discriminators optimize the parameters through a cooperative training mechanism to provide accurate discrimination basis for subsequent noise reduction. The encoder of the generator (4 layers of down-sampling convolution) extracts features from the original audio signal and outputs a 128x128 noise feature map, which locates the area with high noise energy. The decoder (4 layers of up-sampling convolution) combines the attention gate mechanism (focusing on the 200Hz-3.4kHz core frequency band of human voice) to generate a time-frequency domain denoising mask with the same dimension as the input signal (mask value 0-1, 0 means complete preservation, and 1 means complete attenuation). High mask values are assigned to noise-dominant frequency bands (such as low-frequency mechanical noise bands) to achieve adaptive attenuation, and low mask values are assigned to human voice frequency bands to preserve details. The residual connection network contains three multi-scale residual blocks (using 3x3, 5x5, and 7x7 convolution kernels, respectively), which fuse the original signal features before mask processing with the processed signal features through a skip connection to compensate for the loss of human voice details during noise reduction. The harmonic structure (fundamental frequency + 2-5 harmonics) and prosody features (speech rate, pitch variation) of the human voice frequency band (200Hz-3.4kHz) are enhanced through reinforcement learning.

[0133] Through the above step-by-step implementation, S1 realizes accurate triggering in complex environments, and S2 completes adaptive suppression of dynamic noise and preservation of human voice, providing a high-quality data source for subsequent speech processing.

[0134] In order to further verify the above technical solution, the above technical solution is tested and verified,

[0135] The test environment is a factory workshop (steady-state noise 85dB), a busy street (non-steady-state noise 78dB), and a high-speed train (moving noise 92dB). The control group of the test is the traditional speech trigger (single microphone) + spectral subtraction noise reduction, and the experimental group is the proposed scheme (electromyography-voiceprint dual-mode trigger + GAN-VAE noise reduction). The evaluation index table is shown in Table 1.

[0136] Table 1 Evaluation index table

[0137]

[0138] The results of the multi-modal trigger performance comparison test are shown in Table 2.

[0139] Table 2 Multimodal trigger performance comparison test results table

[0140]

[0141] As can be seen from the above, the myoelectric-acoustic dual-mode trigger controls the false trigger rate to be <1% under extreme noise (more than 30 times lower than the traditional scheme), and the weak voice detection rate is increased to >95%, proving that the biological signal fusion mechanism effectively solves the trigger failure problem in a noisy environment. The performance comparison table of the noise reduction algorithm is shown in Table 3.

[0142] Table 3 Performance comparison table of noise reduction algorithm

[0143]

[0144] As can be seen from the above test, the STOI of GAN-VAE collaborative noise reduction is 0.94 (close to 0.98 of pure speech), proving that the double discriminator design and harmonic repair mechanism significantly improve speech intelligibility; MOS is 4.5 (better than the industry standard of 4.0), verifying the effect of complete protection of human voice features; the delay is 22ms, which meets the real-time interaction requirement (<50ms), and reflects the architecture advantage of U-Net+residual network. The dynamic mask generation efficiency is shown in Table 4.

[0145] Table 4 Dynamic mask generation efficiency

[0146]

[0147] As can be seen from the above test, the dynamic mask completes noise suppression in <25ms, and the human voice frequency band retention rate is >89%, proving that the attention gate mechanism and multi-frequency divider can effectively adapt to complex acoustic environments.

[0148] The myoelectric-acoustic dual-mode trigger still maintains a false trigger rate of ≤0.8% under 90dB noise, solving the long-standing problem of "voice wake-up failure in high-noise environments" in the industry.

[0149] GAN-VAE collaborative noise reduction makes the MOS score reach 4.5 (traditional scheme ≤3.8), which is the first time to realize "near-lossless" speech acquisition in strong noise. The end-to-end processing delay is ≤25ms (traditional AI noise reduction ≥48ms), which meets the stringent requirements of real-time interaction. In the mobile scene (speed >5m / s), the speech intelligibility fluctuation range is only ±1.3%, proving that the adversarial disturbance enhanced training effectively improves the model robustness. The competitive analysis of the technical scheme is shown in Table 5.

[0150] Table 5 Competitive analysis of technical scheme

[0151]

[0152] Further, the S3 beam dynamic optimization adjustment comprises:

[0153] S31, quality evaluation feedback:

[0154] The application real-time evaluation module calculates the comprehensive quality score through the speech intelligibility index STOI and the noise attenuation amount, and generates a state vector of reinforcement learning in combination with the azimuth sensor data of the microphone array;

[0155] S32, policy iteration optimization:

[0156] The reinforcement learning agent adopts a deep deterministic policy gradient algorithm, the state space includes noise type classification results, target sound source azimuth and current beam parameters, the action space outputs beam angle adjustment and gain adjustment through continuous control variables, and the reward function dynamically calculates and optimizes the output by fusing the signal-to-noise ratio improvement value and the power consumption coefficient;

[0157] S33, hardware parameter update:

[0158] The optimized parameters are written into the digital signal processor register of the microphone array through the real-time control bus, the control signal output by the register drives the beam forming circuit through the digital-to-analog conversion module, and sub-second level adjustment of the beam pointing and gain is realized.

[0159] In the above embodiment, the beam forming is regarded as a continuous control problem, the audio quality improvement is taken as the target, the reinforcement learning agent continuously learns the optimal adjustment strategy in the dynamic environment, the microphone array always focuses on the target sound source, and the noise reduction effect and hardware energy consumption are balanced, solving the problem that the traditional fixed beam parameters cannot adapt to the sound source movement and multiple interference scenes.

[0160] The real-time evaluation module performs a double-channel analysis on the noise-reduced audio signal, calculates the speech intelligibility (value range 0-1, 1 representing complete clarity) through the speech transmission index (STOI), and at the same time, counts the noise attenuation (unit dB, calculation formula: "noise energy before processing-noise energy after processing"). A weighted sum formula is used to generate a quality score (quality score = 0.6 x STOI + 0.4 x normalized value of noise attenuation), where the noise attenuation is mapped to the 0-1 interval through min-max normalization. By fusing the quality score, the microphone array azimuth sensor data (accuracy ± 0.5°, sampling frequency 10 Hz), and the environmental noise type label (such as mechanical noise, human voice interference, etc., output by the S2 module), a 128-dimensional reinforcement learning state vector is generated, fully reflecting the current acoustic environment and device state. The deep deterministic policy gradient (DDPG) algorithm is used to construct the agent, which includes the actor network (policy network) and the critic network (value network). The actor network is responsible for outputting specific adjustment actions, and the critic network evaluates the action value and guides policy optimization. The state space dimension is 20, including the noise type classification result (one-hot encoding 8 dimensions), the target sound source azimuth (2 dimensions: horizontal angle / tilt angle), and the current beam parameter (10 dimensions: including the phase and gain of 5 microphone units); the action space is a 2-dimensional continuous control quantity, corresponding to the beam angle adjustment amount (range -15° to +15°) and the gain adjustment amount (range -10 dB to +10 dB). The reward function is designed as dynamic reward value = 0.7 x SNR improvement value + 0.2 x (1-power consumption coefficient) + 0.1 x STOI change rate. The SNR improvement value is the difference between the SNR before and after parameter adjustment, the power consumption coefficient is positively related to the beam adjustment amplitude (to avoid excessive energy consumption caused by frequent large adjustments), and the STOI change rate ensures continuous improvement of speech intelligibility. The experience replay pool uses a priority sampling strategy (capacity 100,000 samples), giving higher sampling probability to high-reward samples; the network parameters are updated every 100 time steps, and the actor network is optimized through the gradient feedback of the critic network, while using soft update of the target network (soft update coefficient τ = 0.001) to ensure training stability.

[0161] The parameters are written into the digital signal processor (DSP) registers of the microphone array through the SPI real-time control bus (transmission rate 1 Mbps), and the double-buffering mechanism is used to avoid signal jitter during parameter updating. The digital control signal output by the DSP register is converted into an analog voltage signal by a 16-bit digital-to-analog conversion module (DAC, conversion accuracy 0.001 V / LSB), which drives the variable gain amplifier and phase adjuster in the beamforming circuit.

[0162] Through the above closed-loop process, the S3 beam dynamic optimization adjustment can continuously optimize the spatial filtering characteristics of the microphone array in complex scenarios such as multi-source interference and sound source movement.

[0163] In specific embodiments, the S4 semantic association cache enhancement includes:

[0164] S41, real-time semantic analysis:

[0165] Through the distilled BERT architecture of the lightweight semantic model, the collected voice data is classified and extracted, and core semantic labels containing actions, objects, and scenes are generated;

[0166] S42, historical association retrieval:

[0167] Through the distributed index structure of the cache database, the current semantic label and historical semantic data are converted into low-dimensional vectors through the local sensitive hashing algorithm, and a time decay factor is introduced when similarity matching is performed, giving higher weight to high correlation data in the near period;

[0168] S43, dynamic cache strategy: based on semantic importance score to realize hierarchical storage; core semantic data is stored in the cache area, and associated auxiliary data is archived to the large capacity storage area through the semantic compression algorithm, and low value data is cleaned up regularly through the heat ranking.

[0169] In specific embodiments, the core logic of the semantic association cache enhancement is: converting voice data from raw audio to structured semantic labels, establishing cross-scene data association through feature vector similarity matching, dynamically allocating storage resources combined with semantic importance, and achieving the cache optimization goal of "fast access to high-value data and efficient archiving of low-value data", providing data support for contextually coherent intelligent interaction.

[0170] Implementation process:

[0171] 1. S41 real-time semantic analysis:

[0172] Model architecture design: adopt distilled BERT lightweight architecture (parameter reduction by 60%), retain 12-layer Transformer encoder core structure, migrate the semantic understanding ability of the pre-training model to the lightweight model through the knowledge distillation technology, adapt to the algorithm power limit of terminal equipment (inference delay ≤200ms).

[0173] Semantic label generation: Two-stage parsing of the directed speech data output by S3: the first stage identifies the user's core intent (such as "query", "control", "chat", etc., with a classification accuracy of ≥92%) through an intent classification head (with a softmax activation function); the second stage extracts key entities such as actions (such as "turn on", "adjust"), objects (such as "light", "temperature"), and scenes (such as "living room", "bedroom") through an entity extraction head (based on CRF conditional random fields), ultimately generating a 15-20 dimensional core semantic label vector.

[0174] Real-time optimization: Quantitative reasoning (INT8 precision) is used to reduce computational load, combined with a sliding window mechanism (window size 5s) for long speech block processing, ensuring continuous semantic analysis of real-time speech streams.

[0175] 2. S42 historical correlation retrieval:

[0176] Distributed index construction: The cache database uses a distributed index structure based on Elasticsearch to convert semantic labels into a hybrid index mode of inverted index and vector index. The inverted index is used for fast matching of entity keywords, and the vector index stores semantic feature vectors to support similarity retrieval.

[0177] Feature vector conversion: High-dimensional semantic labels (15-20 dimensions) are compressed into 64-dimensional low-dimensional vectors through the Local Sensitivity Hashing (LSH) algorithm, reducing storage and computational overhead. The hash function uses a multi-bucket mapping strategy to ensure that semantically similar labels fall into the same hash bucket, improving retrieval efficiency.

[0178] Similarity matching mechanism: Calculate the cosine similarity (threshold set to 0.7) between the current semantic vector and the historical vector, and introduce a time decay factor to give 1.2 times weight to highly relevant data (similarity ≥0.8) within the past 30 days, giving priority to recent interaction data. When the direct similarity is below the threshold, expand the matching path through the domain knowledge graph (containing 500+ entity relationships), such as indirectly associating "light" with "lighting" through "synonymous relationship".

[0179] 3. S43 dynamic caching strategy:

[0180] Semantic importance scoring: Build a scoring model to quantify the value of semantic data, with the scoring formula being "importance score = 0.5 x intent priority + 0.3 x entity scarcity + 0.2 x historical access frequency". The intent priority is assigned according to the degree of interaction importance (such as "emergency control" assigned a value of 1.0, "chitchat" assigned a value of 0.3), and the entity scarcity reflects the frequency of the entity appearing in historical data (low-frequency entities have higher weights).

[0181] Hierarchical storage implementation: based on the score results, the storage resources are allocated: core semantic data with importance score ≥ 0.7 (such as high-frequency control instructions, key entity information) is stored in the cache area (DDR4 memory is used, read and write speed ≥ 20GB / s), and the complete semantic label and the original voice segment are retained; the associated auxiliary data with a score of 0.3-0.7 (such as context chatting content) is processed by a semantic compression algorithm (based on Transformer sequence compression, compression ratio 3:1), and then archived to a large-capacity storage area (such as eMMC flash memory, capacity ≥ 64GB); low-value data with a score < 0.3 is marked as to be cleaned up.

[0182] Cache update mechanism: an LRU (Least Recently Used) and hotness sorting combined cleaning strategy is used, and the hotness of the data in the large-capacity storage area is calculated every morning (hotness = access frequency x time decay coefficient), and the low-value data ranked after 20% is automatically deleted; the cache area uses a timing refresh mechanism (once every hour), and the core data that has not been accessed for more than 72 hours is downgraded to the large-capacity storage area, ensuring efficient use of cache resources.

[0183] Through the above process, the S4 semantic association cache enhancement realizes the semantic storage and associated retrieval of voice data.

[0184] In further technical embodiments, the fusion model working method in S13 further includes:

[0185] The outputs of the myoelectric feature branch and the voiceprint feature branch are time-synchronized through a feature alignment module, a cross-attention layer uses a multi-head attention mechanism to calculate feature interaction weights, and the weight values are filtered for invalid associations through a gating function; the feature matching degree in the trigger condition is measured by cosine similarity and edit distance, and the dynamic threshold module is associated with environmental noise sensor data.

[0186] In the above specific implementation process,

[0187] The technical principle and specific implementation process of the S13 fusion model are:

[0188] Technical principle: the S13 fusion model first ensures the consistency of the two modal features in time sequence through time synchronization, then captures deep associations between features and filters noise interactions using a multi-head attention mechanism, and finally realizes robust trigger judgment in complex scenarios through double-metric feature matching and environment-adaptive dynamic threshold, avoiding the defect that single-modal features are easily disturbed.

[0189] Specific implementation process:

[0190] The electromyographic signal (from the flexible patch sensor) and the voiceprint feature (from the microphone array) carry hardware timestamps (accuracy ±0.1 ms) respectively, and the feature alignment module calculates the delay of the two signals (typical delay range 5-20 ms) based on the system clock through the timestamp difference. Linear interpolation is used to fill in or downsample the feature sequence with large delay, so that the time steps of the electromyographic feature sequence (sampling rate 100 Hz) and the voiceprint feature sequence (sampling rate 16 kHz downsampled to 100 Hz) are completely matched (one feature frame every 10 ms), ensuring temporal consistency. The time correlation coefficient of the synchronized features is calculated through a sliding window (window size 100 ms), and when the coefficient is ≥0.85, it is determined that the synchronization is valid, otherwise a re-calibration is triggered to avoid fusion errors caused by timing misalignment.

[0191] Cross-attention layer: feature interaction and weight filtering;

[0192] The multi-head attention configuration adopts 8-head attention mechanism, and the electromyographic feature vector (64 dimensions) and the voiceprint feature vector (128 dimensions) are respectively mapped to 8 parallel subspaces (each subspace dimension is 8 and 16 respectively) through linear transformation. When calculating the interaction weight, in each subspace, the attention weight is calculated through matrix operation of Query (electromyographic feature), Key (voiceprint feature), and Value (voiceprint feature) (formula: ), which captures the correlation of different modal features in the fine-grained dimension. Gating function filtering: after concatenating the 8-head attention weights, the invalid correlation is filtered through the sigmoid gating function ( , where ) The interaction with a weight value <0.2 is considered as noise correlation and is set to zero, and the high-confidence feature interaction (weight value ≥0.2) is preserved.

[0193] The fusion vector generation is to weight and sum the filtered weights and Value vectors, and output a 128-dimensional fusion feature vector through linear transformation, which integrates the key information of the two modalities.

[0194] The feature matching degree adopts a double measurement mechanism, and the cosine similarity calculation is to calculate the cosine similarity (formula: ) between the current fusion feature vector and the preset target feature template (generated through the user enrollment process), which measures the overall similarity of the vector space, and the threshold is set to 0.75. The edit distance measurement is to calculate the dynamic time warping (DTW) edit distance of the feature sequence (arranged by time steps), which measures the matching degree of the sequence in the time sequence change, and the normalized threshold is set to 0.3 (the smaller the value, the more similar the sequence).

[0195] The double judgment rule is that when the cosine similarity is greater than or equal to 0.75 and the edit distance is less than or equal to 0.3, it is determined that the feature matching is effective; if either condition is not met, it is determined that there is no matching, thereby avoiding the limitation of a single measurement (for example, the cosine similarity is not sensitive to the timing change, and the edit distance is not sensitive to the overall vector distribution).

[0196] The dynamic threshold module receives environmental noise sensor data in real time, and divides the noise level into five levels (level 0: quiet environment, noise ≤ 30 dB; level 4: extremely noisy, noise ≥ 70 dB). The basic trigger threshold (fusion feature activation value) is dynamically increased with the noise level: the threshold of level 0 is set to 0.5, and the threshold is increased by 0.1 for each level (the threshold of level 4 is set to 0.9), while the duration requirement is relaxed (level 0 needs to last for 300 ms, and level 4 can be shortened to 150 ms), balancing the anti-interference ability and response sensitivity. When the fusion feature activation value is greater than or equal to the dynamic threshold, the double conditions of feature matching degree are met, and the duration meets the standard, the dynamic threshold module outputs the acquisition start instruction, otherwise it remains in standby state, effectively reducing the risk of false triggering in noisy environment. Through the above implementation process, the S13 fusion model realizes the accurate fusion and dynamic verification of the electromyography and voiceprint features.

[0197] In the above embodiment, the working principle of the double-discriminator structure of the generative adversarial network adopts a cooperative training mechanism, which is as follows:

[0198] The first discriminator outputs a noise probability distribution, the second discriminator outputs a human voice feature integrity score, and the generator loss function is the weighted sum of the losses of the two discriminators; a cycle consistency constraint is introduced in the training process to ensure the semantic consistency of the denoised speech and the original clean speech, and the cycle loss is calculated by the edit distance.

[0199] The technical principle and specific implementation process of the double-discriminator cooperative training mechanism of the generative adversarial network are as follows:

[0200] In the double-discriminator cooperative training process of the generative adversarial network (GAN), the first discriminator focuses on the binary classification of noise and speech, guiding the generator to accurately separate the noise; the second discriminator focuses on the evaluation of the integrity of human voice features, constraining the generator to retain the key features of the speech; both are weighted and cooperatively optimized by the loss function to generate the generator, while introducing a cycle consistency constraint to ensure the semantic consistency of the denoised speech and the original clean speech.

[0201] In specific implementation, the first discriminator (noise classification discriminator): adopts a 3-layer convolutional neural network (CNN) structure, the input is a 256x256 time-frequency spectrogram (noise-containing speech or pure noise), noise features are extracted through a convolutional layer (3x3 convolution kernel, step 2), and after global average pooling, a fully connected layer is connected, and a noise probability distribution (dimension 1, sigmoid activation, the value closer to 1 indicates that it is more likely to be noise) is output. The network focuses on learning the spectral features of steady-state noise (such as fan noise) and transient noise (such as collision noise) to achieve accurate differentiation between noise and speech.

[0202] The second discriminator (feature integrity discriminator): adopts a hybrid architecture of "CNN + attention mechanism", the input is the noise-reduced speech time-frequency spectrogram, speech features (such as harmonic structure, formant) are extracted through 4 residual convolutional blocks, the channel attention module is introduced to focus on the 200Hz-3.4kHz human voice core frequency band, and finally the human voice feature integrity score (dimension 1, range 0-100, the higher the value, the more complete the feature preservation) is output through a fully connected layer. The network focuses on learning the integrity criteria of key features such as speech prosody, intonation, and harmonic distribution.

[0203] The generator (U-Net architecture) receives noise-containing speech and outputs noise-reduced speech; the first discriminator classifies "generator output + real noise", and the second discriminator scores "generator output + real clean speech", forming a closed-loop confrontation of "generator optimization of noise reduction effect → discriminator evaluation of defects → generator reverse correction".

[0204] Loss function weighting design:

[0205] First discriminator loss : adopts binary cross-entropy loss, the formula is ; where x is pure noise, y is clean speech, n is superimposed noise, G is the generator, , and the first discriminator. Second discriminator loss : adopts mean square error loss, the formula is ; where is the second discriminator, and the goal is to make the feature score of the noise-reduced speech close to the clean speech. Generator loss : is the weighted sum of the double discriminator loss, the formula is ; where (against the first discriminator) (4), (against the second discriminator) (5), and the weight (prior to guarantee noise suppression).

[0206] The cycle loss calculation uses dynamic time warping (DTW) edit distance to quantify semantic differences, and the noise-reduced speech (G(x)) ) Align the mel-spectrogram sequence of the original clean speech (y) and calculate the minimum edit distance (total cost of insertion, deletion, and replacement operations) between the sequences, which is normalized as the cycle loss , the formula is , wherein is the sequence length.

[0207] Total loss function optimization: the final loss of the generator is , wherein (balance the adversarial loss and semantic constraints), guiding the generator to preserve semantic features while reducing noise.

[0208] Co-training to build a training data set, including noisy speech (mixed steady-state / transient noise) with different signal-to-noise ratios (0dB-20dB), corresponding clean speech and pure noise samples, divided into training set and validation set according to 9:1, and each batch input batchsize=16.

[0209] Iterative training: adopt an alternating training strategy, first fix the generator parameters, update the double discriminators (learning rate , Adam optimizer, ); then fix the discriminator parameters and update the generator (learning rate , Adam optimizer, ). Verify every 100 rounds to calculate the signal-to-noise ratio (SNR) and speech intelligibility index (STOI) of the denoised speech.

[0210] Convergence criterion: when the SNR improves by ≥15dB, STOI≥0.9 and there is no significant improvement for 50 consecutive rounds on the validation set, stop training and save the generator and discriminator parameters. Gradient clipping (maximum gradient norm=1.0) is used to prevent gradient explosion during training.

[0211] Through the above co-training mechanism, the generator can accurately separate noise (noise suppression increased by 20%) in a complex noise environment, while preserving more than 95% of the human voice feature integrity through the second discriminator and cycle constraint, solving the technical problem of "noise suppression and feature preservation cannot be achieved simultaneously" in traditional GAN noise reduction

[0212] As a further embodiment of the application, the reinforcement learning agent in S32 further comprises:

[0213] The state encoding module focuses on key environmental features through an attention mechanism; the experience replay pool uses a priority sampling strategy to give higher sampling probability to high-reward samples; the output layer of the policy network is processed by batch normalization to stabilize parameter updates, and the adjustment amount exceeding the hardware physical limit is filtered by the feasibility verification module before the action is executed.

[0214] In the above embodiments, the specific implementation process includes:

[0215] 1. State encoding module: attention mechanism focuses on key features:

[0216] Feature input composition: the original state features received by the state encoding module include multi-dimensional information, covering environmental noise types (8-dimensional one-hot encoding), target sound source azimuth angle (2-dimensional: horizontal angle / tilt angle), current beam parameters (10-dimensional: phase and gain), real-time quality score (1-dimensional), etc., totaling 21-dimensional original feature vector.

[0217] Attention mechanism design: a scaled dot-product attention structure is used, and the original feature vector is transformed linearly to generate Query, Key, and Value matrices (all with dimensions of 21x32). The similarity score of Query and Key is calculated (formula: ), and after normalization by the softmax function, the attention weight is obtained, which emphasizes the weight of features that are strongly related to beam optimization (such as sound source azimuth angle and quality score) and weakens the weight of secondary features (such as low-impact categories in environmental noise types). Encoding output: the attention weight and the Value matrix are weighted and summed to output a 32-dimensional enhanced state feature vector. Compared with the original feature vector, the weight of key features is increased by 30%-50%, making the agent more focused on environmental information that plays a decisive role in beam adjustment. Experience replay pool: priority sampling strategy improves sample efficiency

[0218] Priority calculation: the priority value (p) of each experience is dynamically calculated based on the TD error (time difference error), with the formula (where is the TD error, to avoid a priority of 0). The larger the TD error, the higher the potential value of the experience to policy optimization, and the larger the priority value. Sampling implementation: a priority-based random sampling strategy is used, and the probability of each experience being selected is positively related to the priority value (probability formula: balancing priority and randomness). After sampling, the importance sampling weight corrects the gradient update, avoiding the excessive influence of high-priority samples on training and improving sample utilization efficiency by more than 20%.

[0219] In specific embodiments, the policy network (actor network) uses a 3-layer fully connected neural network, with the first two layers outputting 64-dimensional and 32-dimensional features, respectively, which are processed by the ReLU activation function and then input into the output layer. The output layer needs to output continuous beam angle adjustment and gain adjustment, which are sensitive to parameter fluctuations.

[0220] Batch normalization deployment: a batch normalization layer is added before the output layer to standardize the input 32-dimensional features (formula: wherein is the batch mean, is the batch variance, ), and then adjusting the distribution (x) by a scaling parameter (s) and a shift parameter (b) to have a mean of 0 and a variance of 1 for each batch of input features.

[0221] Stabilization effect: Batch normalization reduces the standard deviation of parameter updates in the output layer by 40-60%, avoids strategy shocks caused by fluctuations in the distribution of input features, and improves the convergence speed of the loss function in the training process by 25%, and the stability of the final converged strategy is significantly enhanced.

[0222] In specific embodiments, according to the hardware parameters of the microphone array, the preset action feasibility constraint is: the beam angle adjustment range is [-12°, +12°] (exceeding the range may cause mechanical structure damage), and the gain adjustment range is [-10dB, +10dB] (exceeding the range may cause signal saturation or insufficient gain).

[0223] In specific embodiments, after the strategy network outputs the original action vector (including angle adjustment and gain adjustment), the feasibility verification module checks the action components one by one. If the angle adjustment exceeds the range, it is truncated to the nearest boundary value (such as 12° when the original output is 15°); if the gain adjustment exceeds the range, it is also truncated to the [-10dB, +10dB] interval.

[0224] In specific embodiments, an action change rate constraint is set, the angle change amount of adjacent two adjustments does not exceed 5° / time, and the gain change amount does not exceed 3dB / time, to avoid signal instability caused by sudden changes in hardware parameters.

[0225] The similarity matching method in S42 further includes:

[0226] The semantic vector is compressed to low dimension by knowledge distillation technology; the similarity calculation introduces a domain knowledge graph to assist in correlation reasoning: when the direct similarity is below a threshold, the matching range is expanded through the entity relationship path in the graph; the correlation index adopts an incremental update mechanism, and newly collected data only updates the relevant index branch, reducing the computational overhead.

[0227] ​​​The further technical solution of the similarity matching in S42 is based on the optimization principle of "efficient compression-association expansion-lightweight update", and the core logic is: using the knowledge distillation technology to compress the vector dimension under the premise of preserving the core semantic information, reducing the storage and calculation cost; introducing the domain knowledge graph to construct the entity relationship network, and expanding the matching range through the relationship path when the direct similarity is insufficient, and strengthening the integrity of the semantic association; using the incremental update mechanism to update only the local index branch, avoiding the high overhead of full index reconstruction, and realizing efficient and comprehensive semantic matching.

[0228] In specific embodiments, a "teacher-student" distillation framework is used, the teacher model is a pre-trained high-dimensional semantic encoder (outputting a 128-dimensional vector, preserving rich semantic details), and the student model is a lightweight network (3 layers of full connection + ReLU activation), and the goal is to compress the 128-dimensional vector to 32-dimensional while preserving the core semantic features.

[0229] In specific embodiments, during the training process, the student model is jointly optimized by minimizing the KL divergence loss ( ) output by the teacher model and the classification loss ( ) of the original semantic label, and the total loss , ensuring the semantic fidelity of the compressed vector.

[0230] In specific embodiments, the matching degree of the compressed vector and the original label is compared by cosine similarity before and after compression, and the average cosine similarity of the compressed vector and the teacher model vector is required to be ≥0.92, ensuring that there is no significant loss of core semantic information. The final 32-dimensional compressed vector has a storage overhead reduced by 75% compared to the 128-dimensional original vector, and the similarity calculation speed is improved by 3 times.

[0231] In specific embodiments, the domain knowledge graph construction includes a voice interaction domain knowledge graph containing 500+ entities and 800+ relationships, the entities cover actions (such as "adjust" and "query"), objects (such as "light" and "temperature"), and scenes (such as "living room" and "bedroom"), and the relationship types include synonymous relationships (such as "light-illumination"), hierarchical relationships (such as "air conditioner-home appliance"), and association relationships (such as "temperature-air conditioner"), etc., and a graph database (Neo4j) is used for storage.

[0232] In specific embodiments, the cosine similarity of the current 32-dimensional semantic vector and the historical vector is calculated (the threshold is set to 0.7), if ≥0.7, it is directly determined as a high correlation match, and there is no need to expand; if <0.7, the knowledge graph assisted reasoning is triggered.

[0233] In specific embodiments, for the core entity in the current semantic label (such as "air conditioner"), 1-2 jump relationship paths (path length ≤ 2, to avoid excessive expansion leading to noise) are retrieved in the knowledge graph to obtain associated entities (such as "temperature" and "cooling"). The historical semantic vectors corresponding to the associated entities are included in the matching range, the similarity is recalculated, and the entities on the path are assigned a decay weight (the first jump weight is 0.8, and the second jump weight is 0.5).

[0234] In specific embodiments, the associated index of the cache database adopts a "main index-branch index" hierarchical structure, the main index stores the hash value of the core semantic label, the branch index is divided according to the entity type (action / object / scene), and each branch contains the semantic vector and the associated pointer of the entity of this type.

[0235] In specific embodiments, after the newly collected data generates a semantic vector, the entity type branch to which it belongs is located through the main index (such as "object = light" corresponding to "object branch index"), and only the update of this branch index is triggered, and the other irrelevant branches remain unchanged.

[0236] In specific embodiments, S42 similarity matching reduces the calculation overhead by compressing the vector dimension, expands the associated range with the help of the knowledge graph, and improves the index efficiency through incremental updating. In actual tests, the average response time of semantic matching is shortened to 80ms, the potential association detection rate is increased by 25%, and the index update overhead is reduced by 80%, providing efficient support for the coherent storage and retrieval of cross-scene voice data.

[0237] As a further technical solution of the application, the method further comprises an S5 quality closed-loop optimization step, which includes:

[0238] Periodically analyze the types of recognition errors of the collected data through a confusion matrix, and select difficult example samples to form a fine-tuning data set based on an uncertainty sampling strategy; adopt a parameter freezing and layer-by-layer unfreezing strategy during transfer learning fine-tuning, and only update the last three layers of the noise reduction model and the output layer of the reinforcement learning strategy network; the optimization effect is quantitatively evaluated by the collection quality score, and when the score decreases for three consecutive periods, the model is retrained, wherein the quality score at least includes the fusion recognition accuracy, the noise reduction signal-to-noise ratio, or the collection response speed.

[0239] The core logic of the S5 quality closed-loop optimization step is: analyze the types of recognition errors through a confusion matrix quantitative analysis, locate the weak links of the model; focus on difficult example samples by using an uncertainty sampling, improve the pertinence of fine-tuning; reduce the fine-tuning cost by using the parameter freezing strategy of transfer learning, realize lightweight and efficient model updating; monitor the optimization effect in real time through multi-dimensional quality score, dynamically trigger the retraining mechanism, and ensure that the system maintains high collection quality for a long time, solving the defect that the traditional static model cannot adapt to the dynamic environment.

[0240] Specific implementation process:

[0241] 1. Error type analysis: Confusion matrix quantification diagnosis:

[0242] Data collection range: Collect the collected data and corresponding annotation results (manual annotation of error types) in the past period regularly (default every week), with a sample size of no less than 500 and covering different scenarios (quiet / noisy) and different user characteristics.

[0243] Confusion matrix construction: Construct a multi-dimensional confusion matrix, with rows representing actual labels (such as "correct trigger - correct noise reduction - correct beam" "false trigger" "excessive noise reduction" "beam offset" and other 8 categories) and columns representing model output results. The matrix elements are the number of samples for the corresponding error types. Calculate the error rate (number of error samples / total number of samples) and the proportion of each type of error through the matrix.

[0244] Key error positioning: Focus on high-proportion error types (proportion ≥ 10%), such as "excessive noise reduction under low signal-to-noise ratio" "beam tracking delay for moving sound sources", etc. Analyze the root cause in combination with error scenario characteristics (such as noise type, sound source moving speed) to provide a clear direction for subsequent fine-tuning.

[0245] 2. Difficult instance sample selection: Uncertainty sampling strategy:

[0246] Uncertainty measurement indicators: For unannotated samples in the collected data set, quantify uncertainty through model prediction confidence, using "prediction entropy value" and "maximum probability difference" dual indicators for evaluation: the higher the prediction entropy value (indicating that the model is less certain about the result), the smaller the maximum probability difference (indicating low distinction between classes), the higher the probability of being judged as a difficult example.

[0247] Sampling implementation: Set entropy value threshold (≥ 0.8) and maximum probability difference threshold (≤ 0.2), and select samples that meet either condition from the candidate samples; use stratified sampling to ensure that difficult examples cover main error types (such as 40% of noise reduction error samples, 30% of beam error samples, etc.), forming a fine-tuning data set (10%-15% of the size of the base training set).

[0248] Sample verification: Manually review the difficult example samples selected, remove annotated errors or worthless samples (such as irrecoverable data caused by extreme noise), and ensure the effectiveness of the fine-tuning data set.

[0249] 3. Transfer learning fine-tuning: Parameter freezing and layer-by-layer unfreezing:

[0250] Model parameter freezing strategy: For the core model, a hierarchical freezing mechanism is adopted: the first 5 layers of the denoising model (generative adversarial network) are frozen (to retain the basic noise feature learning ability), and only the last 3 layers (convolutional layers and residual blocks responsible for fine denoising and feature repair) are unfrozen; the first 2 layers of the reinforcement learning strategy network are frozen (to retain the environment feature encoding ability), and only the output layer (responsible for beam parameter adjustment output) is unfrozen.

[0251] Layer-by-layer unfreezing process: In the early stage of fine-tuning (first 5 epochs), only the output layer is trained; in the middle stage (5-10 epochs), the second-to-last layer is unfrozen and trained jointly with the output layer; in the later stage (10-15 epochs), all target layers are trained, and a small learning rate (1 / 10 of the base learning rate, i.e., 1e-5) is used to avoid parameter oscillation.

[0252] Loss function design: The denoising model fine-tuning loss uses the double discriminator loss + cycle consistency loss (weight adjustment to 0.5 of the original weight); the strategy network fine-tuning loss uses the TD error loss of reinforcement learning, combined with the high reward weight of difficult example samples (difficult example sample reward coefficient x 1.5).

[0253] 4. Quality evaluation and retraining trigger: Quantitative monitoring mechanism:

[0254] Quality score calculation: A multi-dimensional acquisition quality score is constructed, with the formula "Quality score = 0.4 x fusion recognition accuracy + 0.3 x denoising signal-to-noise ratio + 0.3 x acquisition response speed (normalized)". Where fusion recognition accuracy ≥ 95% is 1 point, each 1% decrease is 0.02 points; denoising signal-to-noise ratio ≥ 20 dB is 1 point, each 2 dB decrease is 0.1 points; response speed ≤ 500 ms is 1 point, each 100 ms increase is 0.1 points, total score range 0-1 points.

[0255] Periodic evaluation: Calculate the quality score every optimization period (default 1 week), compare with historical scores to draw trend curve, evaluate fine-tuning effect. If the score improves by ≥0.05, it is determined that the optimization is effective, and the fine-tuned model parameters are retained.

[0256] Re-training trigger condition: When the quality score decreases for 3 consecutive periods (cumulative decrease ≥0.1), or a single indicator (such as fusion recognition accuracy) falls below the threshold (≤85%), trigger model retraining: reconstruct the training set based on the original training set + historical difficult example samples, retrain the full model with the initial learning rate, with 60% of the base training rounds to ensure that the model quickly recovers performance.

[0257] Through the above implementation process, the S5 quality closed loop optimization step realizes the continuous iteration of system performance: the difficult example sample fine-tuning reduces the error rate of the main error type by 30%-40%, the parameter freezing strategy reduces the fine-tuning calculation amount by 60%, and the quality score monitoring ensures that the system can still maintain stable performance when the environment changes (the quality score fluctuation is less than or equal to 0.05 in long-term use), effectively solving the problem of insufficient long-term adaptability of traditional systems.

[0258] An AI voice algorithm-based voice data recognition system, comprising:

[0259] A multi-modal triggering module connected with a multi-modal sensor and an AI processing unit, the multi-modal sensor integrating a flexible lip muscle eletrode patch and a voiceprint collection subunit, the multi-modal triggering module generating an activation instruction through a cross-attention fusion model after receiving an eletrode signal and a voiceprint feature, and triggering voice collection start;

[0260] An AI adaptive noise reduction module deployed in the AI processing unit, containing a generative adversarial network and a residual repair network, the generative adversarial network training a double discriminator model based on noise information of an environment sensor, dynamically generating a time-frequency noise reduction mask, and the residual repair network strengthening a human voice feature through a multi-scale residual block;

[0261] A beam dynamic optimization module connected with a microphone array and the AI processing unit, including a quality evaluation sub-module, a reinforcement learning agent, and a hardware control interface, the quality evaluation sub-module outputting an audio quality score, the reinforcement learning agent generating beam angle and gain adjustment parameters based on the score, and the hardware control interface driving a beamforming circuit of the microphone array;

[0262] A semantic association cache module in communication with the AI processing unit and a storage unit, containing a lightweight semantic analysis sub-module, a distributed index sub-module, and a dynamic storage controller, the semantic analysis sub-module extracting core semantic labels of voice data, the distributed index sub-module establishing an associated index with historical data, and the dynamic storage controller realizing hierarchical storage and cache update;

[0263] A quality closed loop optimization module connected with the AI processing unit, analyzing error types through a confusion matrix, and fine-tuning algorithm parameters of the AI adaptive noise reduction module and the beam dynamic optimization module through transfer learning, to maintain long-term adaptability of the system.

[0264] In the above embodiment, the voice data recognition system solves the problems of poor robustness, weak adaptability, and low data utilization of traditional voice collection systems in complex environments through the organic linkage of five modules. The core logic is: precise starting is achieved by a multi-modal trigger module to exclude invalid interference; noise and human voice are separated by an AI adaptive noise reduction module to preserve voice features; real-time focusing on target sound sources is achieved by a beam dynamic optimization module to improve signal gain; cross-scene data association is established by a semantic association cache module to enhance historical reuse; and algorithm parameters are continuously iterated by a quality closed-loop optimization module to maintain long-term performance stability.

[0265] The specific implementation process is:

[0266] 1. Multi-modal trigger module: precise triggering of dual modalities:

[0267] Hardware deployment: multi-modal sensor integrates two types of core perception units - flexible lip muscle electrode patch (contains 32-channel high-density myoelectric electrode array and micro impedance detection circuit, adheres to the lip skin through medical-grade conductive gel) and voiceprint collection subunit (6 units of microphone array, each connected to an independent preamplification circuit), both of which are connected to the AI processing unit (equipped with NVIDIA Jetson Xavier NX) through a high-speed data bus (USB 3.0).

[0268] Signal preprocessing: myoelectric signals are output after pre-differential amplification (gain 1000 times) and 50Hz notch filtering; voiceprint signals are converted into frequency spectrum by short-time Fourier transform, extracting frame-level voiceprint features (128-dimensional vector).

[0269] Fusion trigger logic: the module has a cross-attention fusion model that synchronizes myoelectric features (64-dimensional) and voiceprint features (128-dimensional) in time (based on hardware timestamp, error ≤1ms), calculates feature interaction weights (filters invalid associations with weights <0.2) through 8 attention mechanisms, and generates a 128-dimensional activation vector. The dynamic threshold module adjusts the trigger condition in real time (such as setting the threshold to 0.9 under 4-level noise, with a duration of ≥150ms), and outputs the collection start instruction after double verification, with a false trigger rate of less than 5%.

[0270] 2. AI adaptive noise reduction module: dynamic noise suppression and human voice preservation:

[0271] Hardware deployment: the module is deployed on the GPU core (21 TOPS of computing power) of the AI processing unit, communicates with the environment sensor (noise sensor + temperature and humidity sensor), and obtains background noise data (sampling rate 16kHz) in real time.

[0272] Dual-network collaborative processing:

[0273] GAN training: Train a dual discriminator model based on noise information from environmental sensors (decomposed into low / mid / high frequency components) - the first discriminator (3-layer CNN) outputs a noise probability distribution, and the second discriminator (CNN + channel attention) outputs a score for the completeness of the vocal features; the generator uses a U-Net architecture to output a noise feature map through an encoder, and a decoder generates a time-frequency noise reduction mask (attenuation ≥ 20 dB for noise-dominant frequency bands) with attention gate.

[0274] Residual repair enhancement: The residual connection network contains 3 multi-scale residual blocks (3x3 / 5x5 / 7x7 convolution kernels), which fuse the signal features before and after mask processing through a jump connection, strengthen the vocal harmonic structure (200Hz-3.4kHz), and improve the speech intelligibility index (STOI) by ≥25% after repair.

[0275] Real-time processing flow: After the original audio signal is input, noise classification, mask generation, and feature repair are completed within 10ms, and the pure speech signal is output to the beam dynamic optimization module.

[0276] 3. Beam dynamic optimization module: Real-time focused gain adjustment:

[0277] Hardware connection: The module is connected to the microphone array (6-unit linear array, 2cm spacing) and the AI processing unit through the SPI control bus, including a quality evaluation submodule (embedded DSP chip), a reinforcement learning agent (deployed on the AI processing unit CPU), and a hardware control interface (digital-to-analog conversion module, 16-bit precision).

[0278] Dynamic optimization flow:

[0279] Quality evaluation: Real-time calculation of speech intelligibility index (STOI) and noise attenuation, combined with azimuth sensor data (accuracy ±0.5°) to generate a 21-dimensional state vector.

[0280] Policy iteration: The reinforcement learning agent uses the deep deterministic policy gradient algorithm, and the state encoding module focuses on the sound source azimuth and quality score (weight increased by 30-50%) through the attention mechanism; the experience replay pool (capacity 100,000) uses priority sampling (high reward sample sampling probability x 1.5); the output layer of the policy network is processed by batch normalization, and the output beam angle (-12° to +12°) and gain (-10dB to +10dB) adjustment amount.

[0281] Hardware driving: After the optimized parameters are verified for feasibility (filtering adjustment amounts that exceed physical limits), they are written to the DSP register, and the digital-to-analog conversion module drives the beamforming circuit to achieve sub-second (≤500ms) parameter updates, with a target sound source signal-to-noise ratio improvement of 15-20dB.

[0282] 4. Semantic association cache module: coherent storage of data across scenarios:

[0283] Hardware architecture: modules communicate with AI processing unit (PCIe 3.0 interface) and storage unit (high-speed DDR4 memory + large-capacity eMMC flash), semantic parsing sub-module is deployed in NPU core (4 TOPS of computing power) of AI processing unit, distributed index sub-module is built based on Elasticsearch.

[0284] Semantic processing and storage:

[0285] Semantic parsing: lightweight semantic model (distilled BERT, parameter reduction by 60%) classifies intent (accuracy ≥ 92%) and extracts entities from voice data, generating 15-20 dimensional core semantic labels containing actions / objects / scenarios.

[0286] Association retrieval: compress semantic labels into 64-dimensional low-dimensional vectors through local sensitive hashing, combine with domain knowledge graph (500+ entities) for association reasoning (1-2 hop relationship path), and weight highly relevant data in the last 30 days by 1.2.

[0287] Hierarchical storage: dynamic storage controller based on semantic importance score (≥ 0.7 stored in DDR4, 0.3-0.7 stored in eMMC after semantic compression), clean 20% low-value data daily, core data access delay ≤ 50ms.

[0288] 5. Quality closed-loop optimization module: long-term stable performance iteration:

[0289] Module deployment: communicate with CPU and storage unit of AI processing unit, trigger optimization process through scheduled tasks (weekly), and rely on human annotation interface to obtain error type annotations.

[0290] Closed-loop optimization process:

[0291] Error analysis: construct a confusion matrix of 8 types of errors, calculate error rate and proportion, and locate high-proportion errors (such as "excessive noise reduction under low signal-to-noise ratio").

[0292] Difficult example fine-tuning: use uncertainty sampling (entropy ≥ 0.8 or maximum probability difference ≤ 0.2) to select difficult example samples (accounting for 40% of fine-tuning set), freeze the first 5 layers of the denoising model and the first 2 layers of the strategy network during transfer learning, and only update the last 3 layers / output layer.

[0293] Effect monitoring: quality score = 0.4 x fusion recognition accuracy + 0.3 x noise reduction signal-to-noise ratio + 0.3 x response speed (normalized), trigger retraining (based on original + historical difficult example data, training rounds are 60% of the basic rounds) when the quality score decreases for 3 consecutive periods (cumulative amplitude >= 0.1), and the quality score fluctuation is <= 0.05 in long-term use.

[0294] System synergy effect: each module is closely coordinated through data flow (voice signal / semantic label) and control flow (trigger instruction / optimization parameter), and can still maintain a trigger accuracy of more than 90% in a noisy environment (noise >= 70dB), the voice signal-to-noise ratio after noise reduction is improved by >= 15dB, the beam tracking delay is <= 500ms, the historical data reuse rate is improved by 40%, and high-quality, high-robustness and high-coherence voice data acquisition in complex dynamic scenes is realized, providing core support for subsequent voice interaction applications.

[0295] Although the specific embodiments of the present application are described above, those skilled in the art should understand that these specific embodiments are only illustrative, and those skilled in the art can make various omissions, substitutions and changes to the details of the above method and system without departing from the principles and essence of the present application. For example, the above method steps are combined, so as to perform substantially the same function to achieve substantially the same result according to the substantially same method, which belongs to the scope of the present application. Therefore, the scope of the present application is only limited by the appended claims.

Claims

1. A speech data recognition method based on AI speech algorithms, applied to a data acquisition terminal integrating a microphone array, a multimodal sensor, and an AI processing unit, wherein the method acquires raw audio signals through a microphone array, transmits them to the AI ​​processing unit after filtering, and supports basic format conversion and storage of single-channel speech data; characterized in that: Includes the following steps: S1. Multimodal collaborative triggering acquisition: The system synchronously collects lip electromyography signals and voiceprint features using a multimodal sensor, and generates activation commands through feature fusion to trigger the start of voice acquisition. S2, AI adaptive noise reduction processing: Generative adversarial networks are used to separate noise from the original audio signal, extracting environmental noise features to generate a dynamic noise reduction mask while preserving the integrity of human voice features. S3, Beam dynamic optimization adjustment; Based on reinforcement learning algorithms, real-time audio quality is analyzed, and the beam pointing and gain parameters of the microphone array are dynamically adjusted to focus on the target sound source. S4, enhanced semantic association caching; Real-time semantic analysis of the collected voice data is performed, and an index is established to associate the data with historical interaction data, so as to achieve continuous storage of data collected across different scenarios. The S1 multimodal collaborative triggering acquisition method includes: S11. Acquisition of lip electromyography signals; A high-density electromyography electrode array and a micro impedance detection circuit are integrated through a flexible patch sensor. The electrode signal is then output after being pre-processed by a pre-amplifier module and a 50Hz notch filter through a medical-grade conductive gel attached to the skin of the lips. S12. Voiceprint feature extraction; Each unit of the microphone array is connected to an independent preamplifier circuit. The initial audio segment is converted into a spectrogram by a short-time Fourier transform and then input into a voiceprint encoder to extract frame-level voiceprint features. S13, Feature fusion; A trigger fusion model is constructed, which includes an electromyography feature branch and a voiceprint feature branch. Feature interaction is achieved through a cross-attention layer. The generated activation vector needs to be verified by a dynamic threshold module. The threshold is adjusted in real time according to the environmental noise level, and data acquisition is started when both the duration and feature matching degree are met.

2. The identification method according to claim 1, characterized in that: The S2AI adaptive noise reduction process includes: S21. Noise Feature Modeling: Background noise collected by environmental sensors is decomposed into low-frequency, mid-frequency, and high-frequency noise components by a multi-band frequency divider, and then input into the dual discriminator of the generative adversarial network. The first discriminator focuses on noise or speech differentiation, while the second discriminator evaluates the integrity of human voice features. S22, Dynamic Mask Generation: The application generator adopts the U-Net architecture. The application encoder outputs a noise feature map, and the application decoder, combined with an attention gating mechanism, focuses on the human voice frequency band to generate a time-frequency domain noise reduction mask and adaptively attenuates the noise-dominant frequency band. S23. Voice Feature Restoration: By integrating residual connection networks into multi-scale residual blocks, and fusing signal features before and after masking through skip connections, the harmonic structure and prosodic features of the human voice frequency band are enhanced. The restoration process incorporates speech activity detection results to filter out non-speech segments.

3. The identification method according to claim 1, characterized in that: The S3 beam dynamic optimization adjustment includes: S31. Quality Assessment Feedback: The application's real-time evaluation module calculates a comprehensive quality score using the Speech Proficiency Index (STOI) and noise attenuation, and generates a state vector for reinforcement learning by combining the azimuth sensor data from the microphone array. S32, Strategy Iteration and Optimization: By employing a deep deterministic policy gradient algorithm through a reinforcement learning agent, the state space includes noise type classification results, target sound source azimuth angle, and current beam parameters. The action space outputs beam angle adjustment and gain adjustment through continuous control. The output is dynamically calculated and optimized by fusing the signal-to-noise ratio improvement value and power consumption coefficient through a reward function. S33, Hard motion space is updated via continuous control input parameters: The optimized parameters are written into the digital signal processor register of the microphone array via a real-time control bus. The control signal output from the register drives the beamforming circuit through the digital-to-analog converter module, realizing sub-second adjustment of beam pointing and gain. The reinforcement learning agent in S32 also includes: The state coding module uses an attention mechanism to focus on key environmental features; The experience replay pool employs a priority sampling strategy, assigning higher sampling probabilities to high-reward samples; The output layer of the policy network stabilizes parameter updates through batch normalization, and before the action is executed, the feasibility verification module filters out adjustment amounts that exceed the physical limits of the hardware.

4. The identification method according to claim 1, characterized in that: The S4 semantic association cache enhancements include: S41. Real-time semantic parsing: By using a lightweight semantic model with a distilled BERT architecture, intent classification and entity extraction are performed on the collected speech data to generate core semantic tags containing actions, objects, and scenes. S42. Historical Link Search: The cache database adopts a distributed index structure, and the locality-sensitive hash algorithm is used to convert the current semantic tags and historical semantic data into low-dimensional vectors. A time decay factor is introduced when similarity matching, and higher weights are given to highly relevant data in the recent period. S43, Dynamic Caching Strategy: Layered storage is implemented based on semantic importance scoring; core semantic data is stored in a high-speed cache, while related auxiliary data is archived to a large-capacity storage area using a semantic compression-based algorithm, and low-value data is periodically cleaned up by sorting by popularity.

5. The identification method according to claim 1, characterized in that: The fusion model working method in S13 also includes: The outputs of the electromyography feature branch and the voiceprint feature branch are synchronized in time through the feature alignment module. The cross-attention layer uses a multi-head attention mechanism to calculate the feature interaction weights, and the weight values ​​are filtered out by a gating function to remove invalid associations. The feature matching degree in the triggering condition is measured by both cosine similarity and edit distance. The dynamic threshold module associates environmental noise sensor data.

6. The identification method according to claim 2, characterized in that: The working principle of the dual discriminator structure of the generative adversarial network using the collaborative training mechanism is as follows: The first discriminator outputs the noise probability distribution, the second discriminator outputs the human voice feature integrity score, and the generator loss function is the weighted sum of the losses of the two discriminators. The training process introduces a cycle consistency constraint to ensure the semantic consistency between the denoised speech and the original clean speech. The cycle loss is calculated by the edit distance.

7. The identification method according to claim 4, characterized in that: In step S42, the similarity matching method further includes: Semantic vector computation is compressed to a low dimension using knowledge distillation techniques. Similarity calculation, incorporating domain knowledge graphs to assist in association reasoning: When the direct similarity is below the threshold, the matching range is expanded through entity relationship paths in the graph; The associated index uses an incremental update mechanism, where newly collected data only updates the associated index branches, reducing computational overhead.

8. The identification method according to claim 1, characterized in that: The method also includes an S5 quality closed-loop optimization step, comprising: Regularly analyze the types of identification errors in the collected data using confusion matrix analysis, and select difficult sample samples to form a fine-tuning dataset based on an uncertainty sampling strategy; During fine-tuning of transfer learning, a parameter freezing and layer-by-layer unfreezing strategy is adopted, updating only the last three layers of the denoising model and the output layer of the reinforcement learning policy network; The optimization effect is quantitatively evaluated by collecting quality scores. When the score drops for three consecutive cycles, the model is retrained. The quality scores include at least the fusion recognition accuracy, noise reduction signal-to-noise ratio, or acquisition response speed.

9. A speech data recognition system based on an AI speech algorithm, characterized in that, The method described by any one of claims 1-8 includes: A multimodal triggering module is connected to a multimodal sensor and an AI processing unit. The multimodal sensor integrates a flexible lip electromyography patch and a voiceprint acquisition subunit. After receiving electromyography signals and voiceprint features, the multimodal triggering module generates an activation command through a cross-attention fusion model to trigger the start of voice acquisition. The AI ​​adaptive noise reduction module is deployed within the AI ​​processing unit and includes a generative adversarial network and a residual repair network. The generative adversarial network trains a dual discriminator model based on noise information from environmental sensors and dynamically generates a time-frequency noise reduction mask. The residual repair network enhances human voice features through multi-scale residual blocks. The beam dynamic optimization module connects the microphone array and the AI ​​processing unit. It includes a quality assessment submodule, a reinforcement learning agent, and a hardware control interface. The quality assessment submodule outputs an audio quality score. The reinforcement learning agent generates beam angle and gain adjustment parameters based on the score and drives the beamforming circuit of the microphone array through the hardware control interface. The semantic association caching module communicates with the AI ​​processing unit and the storage unit. It includes a lightweight semantic parsing submodule, a distributed indexing submodule, and a dynamic storage controller. The semantic parsing submodule extracts the core semantic tags of the speech data, the distributed indexing submodule establishes an association index with historical data, and the dynamic storage controller implements hierarchical storage and cache updates. The quality closed-loop optimization module, connected to the AI ​​processing unit, identifies error types through confusion matrix analysis and uses transfer learning to fine-tune the algorithm parameters of the AI ​​adaptive noise reduction module and the beam dynamic optimization module to maintain the long-term adaptability of the system.

Citation Information

Patent Citations

  • Voice acquiring method, device and system, and facility

    CN110767228A

  • Dynamic sound filtering and classifying system and method in real-time microphone connection environment and storage medium

    CN117636890A

  • Intelligent real-time interactive question-answering system based on virtual digital human

    CN120318388A