Voice interaction method and system of AI intelligent robot
By building a dynamic noise feature library and differentiated noise reduction processing, combined with multimodal verification, the recognition accuracy and safety issues of AI intelligent robots in complex noise environments are solved, and efficient voice interaction is achieved in factory environments.
Patent Information
- Application Number
- CN202511156872.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-09-19
AI Technical Summary
The existing AI intelligent robot voice interaction system has low recognition accuracy in complex noisy environments and lacks a multimodal verification mechanism, posing a security risk.
Build a dynamic noise feature library that includes steady-state noise, impact noise, and human voice interference features. Combine real-time audio acquisition with differentiated noise reduction processing, switch the speech recognition model through Mel-frequency cepstral coefficient feature matching analysis, calculate the confidence value and perform multimodal verification, and output the permission control signal.
Accurate, reliable and secure voice interaction of AI intelligent robots is achieved in complex noisy environments, improving the robustness, accuracy and security of the interaction.
Smart Images

Figure CN120673768A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of voice interaction technology, and in particular to an AI intelligent robot voice interaction method and system. Background Art
[0002] With the rapid development of artificial intelligence (AI) technology, the voice interaction function of AI intelligent robots has become a research hotspot in the field of human-computer interaction. Traditional voice interaction systems can achieve high recognition accuracy in quiet environments, but their performance degrades significantly in complex and noisy environments.
[0003] Existing voice interaction technologies face the following challenges. Most systems use a single noise reduction algorithm, making it difficult to address diverse noise types. For example, steady-state noise (such as the continuous operation of air conditioners and fans), impact noise (such as keyboard tapping and objects colliding), and human voice interference (such as the sound of people chatting nearby) all have different impacts on voice signals, making a single algorithm incapable of addressing them effectively.
[0004] In terms of interaction reliability, most existing voice interaction systems rely solely on voice commands and lack multimodal verification mechanisms. This makes it easy for the system to misjudge user commands in noisy environments or fail to effectively identify unauthorized users, posing a security risk.
[0005] Therefore, there is an urgent need for an AI intelligent robot voice interaction method to solve the above problems. Summary of the Invention
[0006] The purpose of the present invention is to provide an AI intelligent robot voice interaction method, comprising the following steps: Acquire a noise sample set, construct a dynamic noise feature library including steady-state noise, impact noise, and human voice interference features based on the noise sample set, and acquire a pre-stored gesture command library; Collect audio signals in real time, obtain the current environmental signal-to-noise ratio based on the audio signals, perform differentiated noise reduction processing on audio signals in the low-frequency band, mid-frequency band, and high-frequency band, and generate a noise-reduced speech signal; Acquire Mel-frequency cepstral coefficient features according to the noise reduction speech signal, perform matching analysis on the Mel-frequency cepstral coefficient features and a dynamic noise feature library, output a noise scene classification result, and switch a corresponding speech recognition model; Outputting a voice interaction signal through the voice recognition model, obtaining a voice command according to the voice interaction signal, and calculating a confidence value according to the voice interaction signal and a current environment signal-to-noise ratio; Obtain a dynamic confidence threshold, compare the confidence value with the dynamic confidence threshold, and output multimodal verification data; Analyze and judge the multimodal verification data and output the permission control signal.
[0007] Furthermore, the steps of collecting audio signals in real time, obtaining the current environmental signal-to-noise ratio according to the audio signals, performing differentiated noise reduction processing on the audio signals in the low-frequency band, the mid-frequency band, and the high-frequency band, and generating the noise-reduced speech signal include: Sampling an audio signal at a sampling rate of 16 kHz, dividing the audio signal into frames and performing Hanning window processing on the audio signal, and obtaining a processed audio signal; Calculating an energy ratio of an active speech segment to a silent segment in the processed audio signal, and obtaining a current environmental signal-to-noise ratio; Decompose the processed audio signal into three frequency bands: low frequency band, mid frequency band and high frequency band; The noise reduction process is performed on the processed audio signals in the low frequency band, mid frequency band and high frequency band respectively. A time domain waveform is obtained, and a noise reduction speech signal is obtained according to the time domain waveform.
[0008] Furthermore, the steps of obtaining Mel-frequency cepstral coefficient features according to the noise reduction speech signal, performing matching analysis on the Mel-frequency cepstral coefficient features and a dynamic noise feature library, outputting a noise scene classification result, and switching a corresponding speech recognition model include: Performing frame processing on the noise reduction speech signal to obtain a frame signal; The framed signal is processed by a 26-channel Mel filter bank to obtain logarithmic energy, and the first 12-dimensional Mel-frequency cepstral coefficient features are extracted by discrete cosine transform; Perform cosine similarity matching on the Mel-frequency cepstral coefficient feature and the dynamic noise feature library, and output a noise scene classification result, wherein the noise scene classification result includes a steady-state noise class and an impact noise class; For steady-state noise, the DNN-HMM hybrid model is used; For impact noise scenarios, the temporal attention Transformer model is used.
[0009] Furthermore, the steps of outputting a voice interaction signal through the voice recognition model, obtaining a voice instruction according to the voice interaction signal, and calculating a confidence value according to the voice interaction signal and a current environment signal-to-noise ratio include: Acquiring a voice command according to the voice interaction signal; Obtain the word probability and the total energy value of the entire frequency band based on the voice interaction signal; If the word probability is lower than 0.4, the voice command is deemed invalid; If the term probability is not less than 0.4, extracting the energy value of the 500-4000 Hz human voice main frequency band in the voice interaction signal ratio, and obtaining the voice activity detection energy ratio based on the 500-4000 Hz human voice main frequency band energy value and the total energy value of the entire frequency band; The confidence value is calculated according to the term probability, the voice activity detection energy ratio, and the current environment signal-to-noise ratio.
[0010] Furthermore, the steps of obtaining a dynamic confidence threshold, comparing the confidence value with the dynamic confidence threshold, and outputting multimodal verification data include: Obtaining a dynamic confidence threshold, where the confidence threshold increases linearly within a preset interval; Compare the confidence value to the dynamic confidence threshold: When the confidence value is not lower than the dynamic confidence qualification threshold, the voice command is output; When the confidence value is lower than the dynamic confidence qualification threshold, the operator's gesture is acquired, analyzed and matched with the pre-stored gesture instruction library, and the voice instruction and gesture instruction are output simultaneously; Based on the comparison result of the confidence value and the dynamic confidence qualification threshold, multimodal verification data including voice commands or voice commands and gesture commands is output.
[0011] Furthermore, the step of analyzing and judging the multimodal verification data and outputting the authority control signal includes: Get the voice permission level including preset voiceprint features; Extracting input voiceprint features of voice commands in multimodal verification data; Determine whether the input voiceprint features match the preset voiceprint features and voice permission level; If there is a match, and when the multimodal verification data only contains voice commands, the voice commands are executed; If there is no match, and when the multimodal verification data only contains voice commands, the voice commands will not be executed and a voice reminder will be issued; If they match, and when the multimodal verification data includes voice commands and gesture commands, then determine whether the voice commands and gesture commands match; If the two match, the voice command and gesture command in the multimodal verification data are executed; If the two do not match, the program will not be executed and a warning will be triggered; If there is no match, and when the multimodal verification data includes voice commands and gesture commands, it will not be executed and a warning will be triggered.
[0012] The present invention also discloses an AI intelligent robot voice interaction system, comprising: An acquisition module is used to acquire a noise sample set, construct a dynamic noise feature library including steady-state noise, impact noise and human voice interference features based on the noise sample set, and acquire a pre-stored gesture command library; An acquisition module is used to acquire audio signals in real time, obtain the current environmental signal-to-noise ratio based on the audio signals, perform differentiated noise reduction processing on the audio signals in the low-frequency band, mid-frequency band, and high-frequency band, and generate a noise-reduced speech signal; A switching module is used to obtain Mel-frequency cepstral coefficient features based on the noise reduction speech signal, match and analyze the Mel-frequency cepstral coefficient features with a dynamic noise feature library, output a noise scene classification result, and switch the corresponding speech recognition model; a calculation module, configured to output a voice interaction signal through the voice recognition model, obtain a voice instruction according to the voice interaction signal, and calculate a confidence value according to the voice interaction signal and a current environment signal-to-noise ratio; An output module is used to obtain a dynamic confidence threshold, compare the confidence value with the dynamic confidence threshold, and output multimodal verification data; The verification execution module is used to analyze and judge the multimodal verification data and output the permission control signal.
[0013] Furthermore, the acquisition module includes: Sampling an audio signal at a sampling rate of 16 kHz, dividing the audio signal into frames and performing Hanning window processing on the audio signal, and obtaining a processed audio signal; Calculating an energy ratio of an active speech segment to a silent segment in the processed audio signal, and obtaining a current environmental signal-to-noise ratio; Decompose the processed audio signal into three frequency bands: low frequency band, mid frequency band and high frequency band; The noise reduction process is performed on the processed audio signals in the low frequency band, mid frequency band and high frequency band respectively. A time domain waveform is obtained, and a noise reduction speech signal is obtained according to the time domain waveform.
[0014] The present application also provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0015] The present application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the above method when executed by a processor.
[0016] The beneficial effects of this application are: The present invention obtains noise samples to build a dynamic noise feature library and a pre-stored gesture command library. It combines real-time audio acquisition with differentiated noise reduction, switching of speech recognition models adapted to noise scenarios, confidence calculation, multimodal verification under dynamic threshold comparison, and authority control analysis to synergistically achieve accurate, reliable, and secure voice interaction for AI intelligent robots in complex noisy environments, thereby improving the robustness, accuracy, and security of the interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a flowchart of a method proposed in one embodiment of the present application.
[0018] Figure 2 This is a schematic diagram of the system structure proposed in one embodiment of the present application.
[0019] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0020] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0021] like Figure 1 As shown, the present application provides an AI intelligent robot voice interaction method, comprising the following steps: S1, obtaining a noise sample set, constructing a dynamic noise feature library including steady-state noise, impact noise, and human voice interference features based on the noise sample set, and obtaining a pre-stored gesture command library; S2, collecting audio signals in real time, obtaining a current environmental signal-to-noise ratio based on the audio signals, performing differentiated noise reduction processing on audio signals in the low-frequency band, the mid-frequency band, and the high-frequency band, and generating a noise-reduced speech signal; S3, obtaining Mel-frequency cepstral coefficient features according to the noise reduction speech signal, performing matching analysis on the Mel-frequency cepstral coefficient features and a dynamic noise feature library, outputting a noise scene classification result, and switching a corresponding speech recognition model; S4, outputting a voice interaction signal through the voice recognition model, obtaining a voice command according to the voice interaction signal, and calculating a confidence value according to the voice interaction signal and a current environment signal-to-noise ratio; S5, obtaining a dynamic confidence threshold, comparing the confidence value with the dynamic confidence threshold, and outputting multimodal verification data; S6: Analyze and judge the multimodal verification data and output the permission control signal.
[0022] As described in the above steps S1-S6, the present invention obtains noise samples to construct a dynamic noise feature library and a pre-stored gesture command library, and combines real-time audio acquisition with differentiated noise reduction, switching of speech recognition models adapted to noise scenarios, confidence calculation, multimodal verification under dynamic threshold comparison, and authority control analysis to synergistically achieve accurate, reliable, and secure voice interaction of AI intelligent robots in complex noise environments, thereby improving the robustness, accuracy, and security of the interaction.
[0023] This invention is particularly suitable for intelligent inspection robots with voice interaction capabilities within factories. Due to the production and manufacturing environment, factories are subject to a large variety of noises, including steady-state noise from production line machinery (such as the continuous roar of motors, with a frequency range of 200-800Hz), impact noise from material handling or equipment collisions (such as metal clashing, with a burst frequency range of 1-6kHz), and the chatter of workers (human voice interference, mainly distributed between 500-4000Hz). These noises can severely distort the voice commands issued by inspection personnel (such as "stop operation" and "record data"). Furthermore, different personnel within the factory (such as operators, technicians, and administrators) have different access rights to the robots. Without strict access control, misoperation can lead to production accidents. Therefore, it is necessary to address issues such as the effective extraction of voice signals in complex noise environments, the dynamic adaptation of recognition models to noise scenarios, the accurate assessment of command credibility, the collaborative verification of multimodal information, and the strict management of operational permissions to ensure that robots can accurately respond to legitimate commands in factory environments.
[0024] Traditional solutions have significant limitations in factory scenarios: For noise processing, a single noise reduction algorithm (such as global spectral subtraction) cannot simultaneously suppress steady-state and impact noise. For example, when processing workshop motor noise, it will over-attenuate mid-frequency signals containing human voice commands. The speech recognition model uses a fixed DNN-HMM, which reduces command recognition accuracy when interfered with by high-frequency impact noise (such as the sound of falling parts). Furthermore, the lack of a multimodal verification mechanism makes it easy for commands to be misjudged due to noise when relying solely on voice. Sudden noise changes can lead to a large number of missed or misjudgments. Permission management relies solely on password input, which is disconnected from voice interaction and does not meet the factory's requirements for fast operation. This invention specifically addresses these issues, enabling precise interaction and secure control in noisy environments.
[0025] Specifically, a noise sample set is obtained and a dynamic noise signature library, encompassing steady-state noise, impact noise, and human voice interference, is constructed from this noise sample set. A pre-existing gesture command library is also acquired to provide the robot with a noise reference and a basis for multimodal verification. Construction of the dynamic noise signature library requires the collection of typical factory noise samples: One hour of continuous recording of production line motor operation as steady-state noise samples, transient sounds such as material impact and tool drops as impact noise samples, and the conversations of different workers as human voice interference samples. These samples are then framed (25ms / frame) and Fourier transformed to extract spectral features (such as center frequency, bandwidth, and amplitude variance). This creates a signature library that supports dynamic updates (with new samples added every two weeks to account for new noise sources). The pre-existing gesture command library contains factory-specific interactive gestures, such as a "clenched fist" to confirm an action and an "open palm" to cancel. Gesture images are captured using a camera and their contour features are extracted and stored, providing a comparison standard for subsequent multimodal verification. For example, if a new high-frequency fan is introduced, its steady-state noise signature can be added to the signature library to ensure the robot can recognize this new noise type.
[0026] Audio signals are collected in real time, the current environmental signal-to-noise ratio is obtained based on the audio signals, and differentiated noise reduction processing is performed on the audio signals in the low-frequency band, mid-frequency band, and high-frequency band. A noise-reduced voice signal is generated to achieve targeted suppression of factory noise, and finally the noise-reduced voice signal is synthesized. Mel-frequency cepstral coefficient features are obtained based on the noise-reduced voice signal, and the Mel-frequency cepstral coefficient features are matched and analyzed with the dynamic noise feature library. The noise scene classification result is output, and the corresponding speech recognition model is switched to achieve adaptation of the noise scene and the recognition model. The speech recognition model outputs a voice interaction signal, obtains a voice instruction based on the voice interaction signal, and calculates a confidence value based on the voice interaction signal and the current environmental signal-to-noise ratio, so that the reliability of the instruction can be quantified.
[0027] Obtaining a dynamic confidence threshold, comparing the confidence value with the dynamic confidence threshold, and outputting multimodal verification data can ensure the validity of instructions through dynamic judgment. Analyzing and judging multimodal verification data and outputting permission control signals can ensure operational security.
[0028] Through the synergistic effect of the above steps, this method enables the factory's intelligent inspection robot with voice interaction function to improve the accuracy of voice command recognition in complex noisy environments, thereby reducing the error rate of operation. At the same time, it ensures operational safety through permission control, significantly improving the practical value of the robot in industrial scenarios.
[0029] In one embodiment, the steps of collecting audio signals in real time, obtaining a current environmental signal-to-noise ratio based on the audio signals, performing differentiated noise reduction processing on audio signals in the low-frequency band, the mid-frequency band, and the high-frequency band, and generating a noise-reduced speech signal include: S21, collecting an audio signal at a sampling rate of 16 kHz, dividing the audio signal into frames and performing Hanning window processing on the audio signal, and obtaining a processed audio signal; S22, calculating the energy ratio of the active speech segment to the silent segment in the processed audio signal, and obtaining the current environment signal-to-noise ratio; S23, decomposing the processed audio signal into three frequency bands: a low frequency band, a mid frequency band, and a high frequency band; S24, performing noise reduction processing on the processed audio signals in the low frequency band, the mid frequency band, and the high frequency band, respectively, obtaining time domain waveforms, and obtaining a noise-reduced speech signal according to the time domain waveforms.
[0030] As described in steps S21-S24 above, by preprocessing the audio signals collected in real time in the factory environment, quantifying the signal-to-noise ratio, and performing differentiated noise reduction in different frequency bands, accurate suppression of different types of industrial noise is achieved, and ultimately a high-quality noise-reduced voice signal is generated, providing a clear and reliable signal basis for the subsequent voice recognition of the intelligent inspection robot with voice interaction function, ensuring the effective extraction of voice commands in complex factory noise environments.
[0031] Factory environments are plagued by a variety of noises with distinct characteristics. Steady-state noise generated by production line motors, fans, and other equipment is primarily concentrated in the low-frequency band (0-1kHz), characterized by persistence and periodicity. Mid-frequency noise (1-4kHz) from workers' conversations and equipment operation overlaps with human voice commands, easily masking critical speech information. Impact noise, such as material collisions and tool drops, is primarily distributed in the high-frequency band (4-8kHz), characterized by bursts and a wide spectrum. These noises directly distort speech signals. Without targeted treatment, inspectors' commands (such as "check equipment temperature" or "record abnormal data") will be difficult for robots to accurately recognize. Therefore, real-time noise analysis and adaptive noise reduction strategies tailored to frequency bands are essential to minimize noise while preserving the key characteristics of speech signals. This is a key requirement for addressing complex factory noise interference.
[0032] Specifically, audio signals are sampled at a 16kHz sampling rate, framed, and Hanning windowed to provide stable data for subsequent signal analysis and noise reduction. The intelligent inspection robot, equipped with voice interaction capabilities, uses an omnidirectional microphone to collect real-time factory ambient sound at a 16kHz sampling rate. This sampling rate fully covers the primary frequency range of speech signals (0-8kHz) and satisfies the Nyquist sampling theorem to prevent frequency aliasing. The continuous audio signal is then framed, using a 25ms frame length and a 10ms frame shift (60% overlap) to render the non-stationary audio signal approximately stationary within the short timeframe, facilitating spectral analysis. A Hanning window is applied to each frame to smooth out signal edges, reduce spectral leakage caused by framing, and ensure that the spectral characteristics of each frame accurately reflect those of the original signal. In factory scenarios, this preprocessing effectively reduces the interference of signal abrupt changes caused by mechanical vibration on subsequent analysis. For example, when collecting speech near an assembly line in a workshop, frame segmentation and windowing stabilize the spectrum of each frame, providing a reliable basis for noise type determination.
[0033] The energy ratio of active speech segments to silent segments in the processed audio signal is calculated, and the current ambient signal-to-noise ratio (SNR) is obtained to quantify the noise intensity in the factory environment. Active speech segments are distinguished from silent segments by setting energy thresholds: active speech segments are defined as segments with energy above -35dBFS (corresponding to the signals of patrol personnel speaking), and silent segments are defined as segments with energy below -45dBFS (corresponding to signals containing only ambient noise). The energy ratio between these two segments is the SNR. For example, in production line areas where equipment is operating, the SNR is typically 5-10dB, indicating high noise energy. In downtime and maintenance areas, the SNR is higher (15-20dB), indicating relatively low noise interference. The quantified SNR directly guides subsequent adjustments to the noise reduction strength: environments with low SNRs require stronger noise reduction parameters, while environments with high SNRs can appropriately reduce the noise reduction strength to minimize speech distortion.
[0034] Decomposing the processed audio signal into three frequency bands—low, mid, and high—is the prerequisite for achieving differentiated noise reduction. Based on the frequency distribution characteristics of factory noise and speech, the time-domain processed audio signal is converted to the frequency domain using a Fast Fourier Transform (FFT). The frequency bands are then divided into the following ranges: low frequency (0-1kHz), primarily containing steady-state noise from motors and fans; mid frequency (1-4kHz), which concentrates the core frequency components of vocal commands (such as the main energy of vowels and consonants); and high frequency (4-8kHz), which contains speech details (such as fricative sounds like "sh" and "s") and impulsive noise. This division clearly identifies the noise type within each frequency band, providing clear targets for targeted noise reduction. For example, in an automotive welding workshop, the low frequency band is primarily noise from the welding robot's motor, while the high frequency band includes the impulsive noise of welding torch sparks. These bands can be treated separately.
[0035] The system performs noise reduction on processed audio signals in the low, mid, and high frequency bands, extracting time-domain waveforms. Based on these waveforms, it then generates a noise-reduced speech signal, enabling precise suppression of factory noise through algorithm adaptation. For low-frequency steady-state noise, an adaptive noise cancellation algorithm is employed: using noise extracted during silent periods as a reference signal, the filter coefficients are adjusted in real time using the LMS (least mean square) algorithm to maximize the offset between the noise estimate output by the filter and the noise components in the input signal. This effectively suppresses low-frequency steady-state factory noise and dynamically tracks noise changes (such as frequency offsets caused by changes in motor load). An improved spectral subtraction algorithm is used in the mid-frequency band: a noise model is first constructed using noise power spectrum estimation (based on silent segment data), which is then subtracted from the noisy speech spectrum. A speech presence probability model is also introduced (retaining 70% of the spectral energy when the speech presence probability is greater than 0.6). This avoids the "musical noise" caused by traditional spectral subtraction and ensures that the inspector's mid-frequency commands are clearly intelligible. For example, when processing conversations between workers and robots, this algorithm effectively removes background equipment noise while preserving key vowel features in the "Start Inspection Procedure" command. A wavelet threshold denoising algorithm is used in the high-frequency band to address impulsive noise. The signal is decomposed into four layers using the db4 wavelet basis. High-frequency wavelet coefficients containing impulsive noise are hard-thresholded (with the threshold set to 0.015 times the maximum coefficient). Only significant coefficients with amplitudes exceeding the threshold are retained. The signal is then reconstructed using an inverse wavelet transform. This removes sudden impulsive noise (such as the sound of falling parts) while preserving high-frequency speech details. Finally, the frequency domain signals processed in the three frequency bands are converted back to the time domain using an inverse FFT and concatenated to form the complete de-noised speech signal. For example, in an assembly workshop containing low-frequency motor noise (600Hz), medium-frequency conversation noise (2kHz) and high-frequency stamping noise (5kHz), after differentiated noise reduction, the signal-to-noise ratio of the voice signal is improved by 10-15dB, and the recognition accuracy of the command "Check workstation No. 3" is improved by more than 30% compared with traditional methods.
[0036] Through the synergistic effect of the above steps, complex noise can be accurately suppressed in a factory environment. Compared with traditional global noise reduction methods, frequency band processing improves the noise suppression rate of each frequency band, while reducing the average energy loss of the voice signal. This lays a high-quality signal foundation for the subsequent extraction of Mel-frequency cepstral coefficient features and the accurate operation of the speech recognition model of the intelligent inspection robot with voice interaction function, significantly improving the reliability of the robot's voice interaction in industrial scenarios.
[0037] In one embodiment, the steps of obtaining Mel-frequency cepstral coefficient features based on the noise reduction speech signal, performing matching analysis on the Mel-frequency cepstral coefficient features and a dynamic noise feature library, outputting a noise scene classification result, and switching a corresponding speech recognition model include: S31, performing frame processing on the noise reduction speech signal to obtain a frame signal; S32, processing the framed signal through a 26-channel Mel filter bank to obtain logarithmic energy, and extracting the first 12-dimensional Mel-frequency cepstral coefficient features through discrete cosine transform; S33, performing cosine similarity matching on the Mel-frequency cepstral coefficient feature and a dynamic noise feature library, and outputting a noise scene classification result, wherein the noise scene classification result includes a steady-state noise class and an impact noise class; S34, for the steady-state noise class, a DNN-HMM hybrid model is used; S35, for the impact noise scene, switches to the temporal attention Transformer model.
[0038] As described in steps S31-S35 above, the Mel-frequency cepstral coefficient features are extracted from the noise-reduced voice signal obtained by the factory intelligent inspection robot, matched with the dynamic noise feature library to identify the noise scene type, and the adapted voice recognition model is switched accordingly to achieve accurate recognition of voice commands in different factory noise environments, ensuring that the inspection robot can accurately understand the operator's instructions (such as "start equipment detection" and "record abnormal parameters").
[0039] Steady-state noise (such as the continuous sound of motors running on assembly lines) and impulsive noise (such as the sound of parts colliding or tools falling) in factory environments have significantly different interference patterns on speech signals. Steady-state noise has a regular spectral distribution and primarily affects the stationary characteristics of speech. Impulsive noise, with its bursty and wide spectrum characteristics, easily disrupts the temporal continuity of speech. Different speech recognition models have different adaptability to these two types of noise. The DNN-HMM hybrid model excels at modeling stationary signals and performs stably under steady-state noise. The temporal attention Transformer model, by focusing on key speech segments, is more resistant to the sudden interference of impulsive noise. Therefore, identifying the current noise scene type through feature matching and dynamically switching models is crucial to addressing the problem of reduced recognition accuracy caused by multiple noise types in factories. This step, through a coherent process of feature extraction, scene classification, and model switching, enables targeted adaptation to different noise scenarios, significantly improving recognition stability. This is particularly applicable to complex factory environments with frequently changing noise types.
[0040] The noise-reduced speech signal is framed to obtain a framed signal, which serves as the foundation for feature extraction. The intelligent inspection robot frames the noise-reduced speech signal using the same parameters as preprocessing: a 25ms frame length and a 10ms frame shift (60% overlap). This ensures that each frame maintains short-term stability. This processing is compatible with the short-term stability of factory noise. For example, when processing speech containing intermittent stamping noise, framing can confine the impact noise fragments to specific frames, facilitating targeted analysis during subsequent feature extraction.
[0041] The framed signal is processed using a 26-channel Mel filter bank to obtain logarithmic energy. This is then processed through a discrete cosine transform (DCT) to extract the first 12 Mel-frequency cepstral coefficients (MFCCs), which are used to extract speech features that effectively distinguish noise types. The center frequencies of the 26 channels of the Mel filter bank are distributed on a Mel scale (covering 100-8000Hz), consistent with human auditory perception. This allows the system to highlight key frequency components of voice commands (e.g., 500-4000Hz) while suppressing non-critical factory noise. The spectral energy obtained after filter bank processing is then logarithmized (enhancing the discrimination of low-energy features). A discrete cosine transform (DCT) is then used to remove inter-feature correlations, extracting the first 12 Mel-frequency cepstral coefficients (MFCCs). These 12-dimensional features focus on the speech's spectral envelope and improve the discrimination between steady-state and impulsive factory noise. For example, when extracting features from speech containing steady-state motor noise, the low-frequency MFCC coefficients show minimal fluctuations. However, for speech containing impulsive noise, the high-frequency coefficients exhibit significant abrupt changes, providing a clear basis for scene classification.
[0042] The Mel-frequency cepstral coefficient features are matched against a dynamic noise feature library using cosine similarity, outputting a noise scene classification result. This classification includes both steady-state and impact noise categories, enabling accurate identification of factory noise scenarios. The dynamic noise feature library contains pre-stored MFCC feature templates for typical factory noises: The steady-state noise template (e.g., motors and fans) is constructed by continuously collecting 30 minutes of steady-state noise and extracting the average MFCC features; the impact noise template (e.g., parts impact and tool drops) is constructed by collecting 500 impact events and averaging the MFCC features. Dynamic updates are supported (new noise samples are added monthly). The cosine similarity (ranging from 0 to 1) between the MFCC features of the current speech and the two templates is calculated. If the similarity with the steady-state template exceeds 0.75, the scene is classified as steady-state noise (e.g., an assembly line operation area); if the similarity with the impact template exceeds 0.75, the scene is classified as impact noise (e.g., a material loading and unloading area). For example, in a welding workshop, if the similarity between the current speech MFCC feature and the impact noise template reaches 0.82, it is determined to be an impact noise scene, providing a basis for model switching.
[0043] For steady-state noise, a DNN-HMM hybrid model is used; for impulsive noise scenarios, a temporal attention Transformer model is used, improving recognition accuracy through model adaptation. The structural parameters of the DNN-HMM hybrid model are optimized for steady-state factory noise: the DNN contains three fully connected hidden layers (128 neurons per layer, with a Reluctant Unit (ReLU) activation function) to model the nonlinear mapping of speech features. The HMM has five states corresponding to the onset, transition, steady-state, decay, and end phases of speech. The state transition probabilities are set based on speech patterns under steady-state factory noise, effectively capturing continuous and stable speech patterns. In steady-state noise scenarios (such as assembly lines), the temporal attention Transformer model is optimized for impulsive noise. The encoder incorporates a six-layer self-attention mechanism (8 attention heads, 512 hidden layer dimensions). By calculating inter-frame attention weights, it focuses on key segments of the human voice command (such as "abnormal" and "stop") and suppresses interference from sudden impulsive noise. The decoder uses a greedy search algorithm to generate recognition results, achieving 30% better immunity to impulsive noise than the DNN-HMM. For example, when there is high-frequency impact noise in the material area, the temporal attention Transformer model can ignore noise frames through the attention mechanism. For example, it can accurately identify the "pause conveyor belt" instruction, which improves the recognition accuracy compared to the fixed model.
[0044] Through the synergistic effect of the above steps, the factory intelligent inspection robot can reduce the fluctuation range of the voice command recognition accuracy in scenarios where noise types dynamically switch. In workshop areas where steady-state and impact noise alternate (such as assembly-material linkage areas), the average recognition accuracy is improved, better meeting the factory inspection requirements for voice interaction reliability.
[0045] In one embodiment, the steps of outputting a voice interaction signal through the voice recognition model, obtaining a voice instruction according to the voice interaction signal, and calculating a confidence value according to the voice interaction signal and a current environment signal-to-noise ratio include: S41, obtaining a voice command according to the voice interaction signal; S42, obtaining the word probability and the total energy value of the entire frequency band (0-8000 Hz) according to the voice interaction signal; S43, if the term probability is lower than 0.4, the voice command is determined to be invalid; S44: If the term probability is not less than 0.4, extracting the energy value of the 500-4000 Hz human voice main frequency band in the voice interaction signal ratio, and obtaining a voice activity detection energy ratio based on the energy value of the 500-4000 Hz human voice main frequency band and the total energy value of the entire frequency band; S45 , calculating a confidence value according to the term probability, the voice activity detection energy ratio, and the current environment signal-to-noise ratio.
[0046] As described in steps S41-S45 above, by performing a multi-dimensional analysis of the voice interaction signals output by the factory intelligent inspection robot and calculating the confidence value in combination with the ambient noise level, a quantitative evaluation of the effectiveness of the voice commands is achieved, providing a reliable basis for subsequent multimodal verification and ensuring that the inspection robot only responds to high-quality and effective commands in a complex factory environment.
[0047] The quality of voice command recognition in factory environments is influenced by multiple factors: the word probability output by the speech recognition model directly reflects the degree of match, but may be affected by similar pronunciations; the energy percentage of the main vocal frequency band (500-4000Hz) reflects the purity of the valid speech components; a high percentage indicates less noise interference; and the ambient signal-to-noise ratio (SNR) reflects the overall acoustic conditions; low SNR environments reduce the reliability of recognition results. These three factors collectively determine the credibility of voice commands, and a single metric cannot fully assess this. For example, a high word probability may lead to misjudgment due to noise similar to the command pronunciation, while a high energy percentage may lose its meaning under strong noise. Therefore, calculating confidence values through multi-parameter fusion is crucial to addressing the issue of inaccurate judgment of voice command validity in factory environments.
[0048] Traditional solutions have limitations in factory scenarios: most use the probability output by the recognition model as the sole judgment criterion (e.g., setting a threshold of 0.5), failing to consider the impact of signal energy distribution and ambient noise. For example, in a noisy area of an assembly line, even a word probability of 0.6 could lead to misidentification of actual commands due to noise interference. In a quiet control room, however, a probability of 0.5 could be valid due to clear signals. This single criterion results in high rates of false positives (falsely accepting invalid commands) and false negatives (falsely rejecting valid commands), failing to meet the reliability requirements of factory inspections. This step, however, integrates word probability, voice activity detection energy ratio, and ambient signal-to-noise ratio to construct a multi-dimensional confidence assessment system, improving judgment accuracy and making it particularly suitable for factory environments with high noise fluctuations.
[0049] The voice command, derived from the voice interaction signal, forms the basis for evaluation. The output of the recognition model (DNN-HMM or Transformer) adapted to the noise scenario for the voice interaction signal includes the decoded text command (e.g., "Check Unit 3") and the corresponding acoustic feature parameters. For example, in the steady-state noise environment of an assembly workshop, the DNN-HMM model outputs the voice command "Start vibration detection," which serves as the raw data for subsequent evaluation.
[0050] The key evaluation parameters are extracted from the voice interaction signal by obtaining the term probability and total energy across the entire frequency band (0-8000Hz). The term probability is a quantitative indicator (ranging from 0 to 1) of the recognition model's ability to match the output text with the input speech. Calculated by the softmax function in the model's output layer, it reflects the similarity between the speech and the corresponding command in the term library. For example, the recognition result for "stop the conveyor belt" corresponds to a term probability of 0.85, indicating that the model is highly confident that the input speech is this command. The total energy across the entire frequency band is calculated by calculating the sum of squares of the voice interaction signal's time domain waveform (in dBFS). It reflects the overall signal strength and typically ranges from -50 to -20 dBFS in factory scenarios. Higher energy values indicate a more pronounced speech signal (e.g., an operator speaking at close range).
[0051] If the term probability is lower than 0.4, the voice command is deemed invalid. This is a mechanism for quickly filtering out low-quality results. The threshold of 0.4 is based on factory scenario statistics: when the term probability is lower than 0.4, the error rate between the command and the actual voice is too high (for example, misidentifying "record data" as "ignore data"), making further evaluation meaningless. For example, in a material area with high impact noise, the recognition model might output "turn off power" even though the term probability is only 35%. In this case, the command is immediately deemed invalid, preventing the erroneous execution of high-risk commands.
[0052] If the term probability is at least 0.4, a further evaluation is conducted on potentially valid commands. The energy of the 500-4000Hz mainband of human voice is extracted from the voice interaction signal, and its ratio to the total energy of the entire frequency band (i.e., the voice activity detection energy ratio) is calculated. This ratio reflects the proportion of valid speech components in the total signal (ranging from 0 to 1). A bandpass filter (passband 500-4000Hz, stopband attenuation 40dB / decade) is used to extract the mainband signal. Its energy is calculated and then divided by the total energy of the entire frequency band to obtain the ratio. For example, if the mainband energy is -30dBFS and the total energy of the entire frequency band is -25dBFS, the energy ratio is 0.316 (i.e., 10^[(-30+25) / 10]), indicating that the valid speech component accounts for 31.6%. In factory scenarios, if this ratio is below 0.2, it indicates that noise energy is dominant, and even high term probabilities may be unreliable.
[0053] The confidence value is calculated based on the word probability, voice activity detection energy ratio and current environment signal-to-noise ratio to achieve multi-dimensional quantitative evaluation. The confidence value calculation formula is: ; in, represents the confidence value, represents the term probability, represents the voice activity detection energy ratio, Indicates the current environment signal-to-noise ratio, represents the weight of the term probability, represents the weight of the voice activity detection energy ratio, The weight of the current environment's signal-to-noise ratio is obtained through the above steps (ranging from 0 to 20 dB, normalized to 0 to 1). The weight distribution is determined based on the statistical impact of various parameters on the effectiveness of instructions in the factory scenario. For example, in an assembly area with a signal-to-noise ratio of 10 dB, the probability of an instruction entry is 0.7 and the energy ratio is 0.6, then the confidence =0.5×0.7+0.3×0.6+0.2×(10 / 20)=0.35+0.18+0.1=0.63, indicating that the command has moderate reliability. This calculation method comprehensively considers model judgment, signal purity, and environmental noise, and improves the accuracy of command effectiveness judgment in factory scenarios compared to single-metric evaluation.
[0054] Through the synergistic effect of the above steps, factory intelligent inspection robots can quantitatively evaluate the reliability of voice commands. In a typical factory environment with a signal-to-noise ratio of 5-15dB, the probability of misjudging invalid commands as valid and the probability of misjudging valid commands as invalid are both reduced. This provides an accurate quantitative basis for subsequent dynamic threshold comparison and multimodal verification, significantly improving the reliability and security of voice interaction.
[0055] In one embodiment, the steps of obtaining a dynamic confidence threshold, comparing the confidence value with the dynamic confidence threshold, and outputting multimodal verification data include: S51, obtaining a dynamic confidence threshold, wherein the confidence threshold increases linearly within a preset interval; S52, comparing the confidence value with the dynamic confidence qualification threshold: S53, when the confidence value is not lower than the dynamic confidence qualification threshold, outputting a voice command; S54, when the confidence value is lower than the dynamic confidence qualification threshold, obtaining the operator's operation gesture, analyzing the operator's operation gesture and matching it with the pre-stored gesture instruction library, and outputting the voice instruction and gesture instruction simultaneously; S55 , outputting multimodal verification data including voice instructions or voice instructions and gesture instructions based on a comparison result of the confidence value and the dynamic confidence qualification threshold.
[0056] As described in the above steps S51-S55, by obtaining a dynamic confidence qualification threshold adapted to the factory environment noise level, the confidence value of the voice command is compared with the threshold, and multimodal verification data containing voice commands or a combination of voice and gesture commands is output according to the comparison result, so as to realize dynamic judgment and multi-dimensional verification of the validity of the command by the factory intelligent inspection robot, ensuring that in a factory environment with large noise fluctuations, it can not only quickly respond to high-reliability commands, but also reduce the risk of erroneous execution of low-reliability commands through multimodal verification.
[0057] The noise intensity in a factory environment fluctuates significantly: when a production line is operating at full capacity, the signal-to-noise ratio may be as low as 5dB (high noise), while when equipment is shut down for maintenance, the signal-to-noise ratio can rise to 20dB (low noise). In a high-noise environment, voice commands are susceptible to interference, resulting in a low confidence value. If a fixed threshold (such as 0.6) is used, valid commands may be mistakenly judged as invalid. In a low-noise environment, even if the confidence value is slightly lower, the command may be valid. A fixed threshold may lead to excessive and unnecessary verification steps. Therefore, the qualified threshold needs to be dynamically adjusted according to the environmental noise. That is, the greater the noise, the higher the threshold to strictly screen commands; the lower the noise, the lower the threshold to improve interaction efficiency. At the same time, gesture verification is introduced for low-confidence commands. This is a key requirement for balancing command reliability and interaction efficiency in factory scenarios.
[0058] Obtaining a dynamic confidence threshold, which increases linearly within a preset range, is key to adapting the threshold to the environment. The preset range is set to 0.5-0.8 based on factory noise characteristics, where 0.5 corresponds to a high signal-to-noise ratio (20dB, low noise) and 0.8 corresponds to a low signal-to-noise ratio (5dB, high noise). The threshold scales linearly with the signal-to-noise ratio, calculated as: dynamic threshold = 0.8 - 0.3 × (P-5) / 15 (P represents the signal-to-noise ratio, which is in the range of 5-20dB). For example, when the inspection robot is in an assembly line area with a signal-to-noise ratio of 10dB, the dynamic threshold is 0.8 - 0.3 × (10-5) / 15 = 0.7, meaning a confidence value of at least 0.7 is required to qualify. In a control room with a signal-to-noise ratio of 18dB, the dynamic threshold is 0.8 - 0.3 × (18-5) / 15 = 0.54. Lowering the threshold reduces unnecessary verification. Dynamic adjustment ensures that the threshold is positively correlated with the noise level, strictly controlling in high-noise environments and improving efficiency in low-noise environments.
[0059] The confidence value is compared with the dynamic confidence threshold, and different processing strategies are adopted based on the results. When the confidence value is at least the dynamic threshold, the command is highly reliable, and the voice command can be directly output. For example, in a warehouse with a signal-to-noise ratio of 15dB, a command has a confidence value of 0.68 and a calculated dynamic threshold of 0.62. Since 0.68 ≥ 0.62, the voice command "Count Shelf Quantity" is output. If the confidence value is below the dynamic threshold, the command is significantly affected by noise, and gesture verification is required to enhance reliability. The robot uses a front-facing camera to capture the operator's gestures in real time, extracting gesture contour features (such as the number of fingers and palm orientation), and then performs a cosine similarity match (threshold 0.85) with a pre-stored gesture command library (including factory-specific gestures such as "making a fist" to confirm and "waving" to cancel). If a match is successful, both the voice command and the gesture command are output simultaneously. For example, in a welding workshop with a signal-to-noise ratio of 6dB, the confidence value of a certain command is 0.65, and the dynamic threshold is 0.78 (because 6dB is close to the low signal-to-noise ratio of 5dB). Because 0.65 is less than 0.78, the operator makes a "fist" gesture (which has a similarity of 0.92 with the "confirm" gesture in the library), and then the voice command "stop the welding machine" and the "confirm" gesture command are output, ensuring the effectiveness of the command through dual modality.
[0060] Based on the comparison results of the confidence value and the dynamic confidence qualification threshold, multimodal verification data containing voice commands or voice commands and gesture commands is output, providing a comprehensive basis for subsequent permission control. In factory scenarios, this output method takes into account both efficiency and safety: high-reliability commands are output in a single mode to speed up response (such as routine inspection commands in the control room), and low-reliability commands are output in a dual mode to reduce risks (such as equipment operation commands in the workshop). For example, in the assembly line area, the command confidence of "record operating parameters" is 0.75 ≥ the dynamic threshold of 0.7, and a single voice command is output; in the material handling area, the command confidence of "lift the robotic arm" is 0.6 < the dynamic threshold of 0.75, and a voice command and a matching "raise your hand up" gesture command are output. In both cases, clear data support is provided for subsequent permission verification.
[0061] Through the synergistic effect of the above steps, the average time it takes for factory intelligent inspection robots to respond to commands in noise fluctuation scenarios can be shortened, while the error execution rate can be reduced. This significantly improves operational safety while ensuring interactive efficiency. It is especially suitable for complex scenarios in factories that involve both routine inspections and high-risk operations.
[0062] In one embodiment, the step of analyzing and judging the multimodal verification data and outputting the authority control signal includes: S61, obtaining a voice permission level including a preset voiceprint feature; S63, extracting input voiceprint features of the voice command in the multimodal verification data; S63, determining whether the input voiceprint feature matches the preset voiceprint feature and the voice permission level; S64, if there is a match, and when the multimodal verification data only includes the voice command, executing the voice command; S65, if there is no match, and when the multimodal verification data only includes a voice command, the voice command is not executed, and a voice reminder is issued; S66, if there is a match, and when the multimodal verification data includes a voice instruction and a gesture instruction, determining whether the voice instruction and the gesture instruction match; S661, if the two match, executing the voice command and gesture command in the multimodal verification data; S662, if the two do not match, then the execution is not performed and a warning is triggered; S67: If there is no match, and when the multimodal verification data includes voice commands and gesture commands, it is not executed and a warning is triggered.
[0063] As described in the above steps S61-S67, by performing voiceprint matching, authority verification and instruction consistency judgment on the multimodal verification data obtained by the factory intelligent inspection robot, accurate authority control signals are output to ensure that only valid instructions from authorized personnel can be executed. At the same time, multimodal instruction comparison is used to avoid misoperation, providing a strict authority control mechanism for factory production safety.
[0064] Factory environments often have operators with varying levels of authority (e.g., administrators can execute shutdown commands, while operators can only record data), and the execution of these commands may impact equipment safety. Without strict permission verification, unauthorized personnel may issue high-risk commands, or commands may be misidentified and executed (e.g., "pause" being identified as "start"), potentially leading to production accidents. Therefore, user identity verification through voiceprint features, command legitimacy based on permission levels, and consistency verification of multimodal commands are crucial to balancing operational flexibility and security in factories.
[0065] Traditional solutions have limitations in factory scenarios: most rely solely on password or card authentication, disconnected from voice interaction and cumbersome operation. Alternatively, they only verify the content of voice commands without verifying user permissions, leading to uncontrolled access. For example, a traditional system might execute an operator's "stop" command. However, this new solution, by binding voiceprints to permissions, only allows administrators to execute that command, effectively solving the problem of separating permissions from interaction.
[0066] Obtaining a voice permission level that includes preset voiceprint features serves as the basis for permission control. Preset voiceprint features are pre-entered by factory personnel: voice samples are collected from administrators, technicians, operators, and other personnel with different permissions (for example, 10 commonly used factory commands per person). Voiceprint features are extracted (using 12th-order Mel-frequency cepstral coefficients + 3rd-order difference coefficients) and associated with permission levels (e.g., administrators correspond to "emergency operation, parameter modification" permissions, while operators correspond to "data recording, status query" permissions). These permissions are stored in the robot's local database and support regular updates (voiceprint templates are updated quarterly to adapt to voiceprint changes). For example, administrator Zhang's voiceprint features correspond to permissions for commands such as "stop" and "adjust parameters," while operator Li's voiceprint features correspond only to permissions for commands such as "record temperature" and "upload data."
[0067] Extracting voiceprint features from voice commands in multimodal verification data enables real-time verification of user identity. Voice command segments (1-3 seconds in length) from the multimodal data are preprocessed (framing and windowing, 20ms frame length, 10ms window shift). 12th-order voiceprint features are extracted using linear predictive coding (LPC). Cosine similarity (0-1) is calculated with pre-set voiceprint features. A match is determined if the similarity is ≥ 0.9. For example, when an operator issues the "Start Inspection" command, their voiceprint is extracted and compared with the pre-set operator voiceprint. A match is determined if the similarity is 0.92.
[0068] The system determines whether the input voiceprint features match the preset voiceprint features and voice permission level, and performs different operations based on the situation. When the multimodal verification data only contains voice commands: if the voiceprint matches and the command is within the corresponding permission range (such as an administrator issuing a "stop" command), the command is executed to ensure the efficiency of the authorized operation; if the voiceprint does not match (such as an outsider issuing "adjust pressure") or the command exceeds the permission (such as an operator issuing "stop"), the command is not executed and a voice reminder (such as "Insufficient permission, please contact the administrator") is issued to prevent unauthorized operations. For example, in an assembly workshop, if an operator issues the command "Record conveyor speed", the voiceprint matches and the permission is met, the robot will execute the recording operation; if the operator issues the command "Stop conveyor", the robot will not execute the operation due to insufficient permission and will issue a reminder.
[0069] When multimodal verification data includes voice commands and gesture commands: If the voiceprint matches, further semantic consistency between the voice and gesture commands is determined (e.g., the voice command "Stop" matches the "make a fist" gesture, or the voice command "record" matches the "wave" gesture). The matching criteria are based on a pre-stored command mapping table (e.g., "Stop - Make a fist," "Start - Thumbs Up"), and the semantic similarity of the commands is calculated (a threshold of 0.85 or above indicates a match). If the two commands match (e.g., "Check device" and the "point to device" gesture), the command is executed to ensure the validity of the multimodal command. If the two commands do not match (e.g., the voice command "Stop" and the "wave" gesture), the command is not executed and an audible and visual warning (e.g., a flashing red warning light and a beeping sound) is triggered to prevent misuse. If the voiceprint does not match (e.g., an unauthorized person issues the "Start device" command accompanied by a gesture), the command is not executed and a warning is triggered to prevent malicious operation. For example, in a welding workshop, if the administrator speaks "pause welding" and makes an "open palm" gesture, and the two match and the voiceprint permissions are met, the robot will pause; if the voice says "pause" but the gesture is "clenched fist" (corresponding to "confirm start"), a warning will be triggered.
[0070] Through the synergistic effect of the above steps, the factory intelligent inspection robot's unauthorized command interception rate can be improved and the multi-modal command misexecution rate can be reduced. While ensuring operational flexibility, it provides a strict permission barrier for the safe operation of the production line. It is especially suitable for factory scenarios involving high-risk operations (such as chemical workshops and heavy machinery production lines).
[0071] The present invention also discloses an AI intelligent robot voice interaction system, comprising: Acquisition module 1 is used to acquire a noise sample set, construct a dynamic noise feature library containing steady-state noise, impact noise and human voice interference features based on the noise sample set, and acquire a pre-stored gesture command library; Acquisition module 2 is used to collect audio signals in real time, obtain the current environmental signal-to-noise ratio based on the audio signals, perform differentiated noise reduction processing on the audio signals in the low-frequency band, mid-frequency band, and high-frequency band, and generate a noise-reduced speech signal; Switching module 3, used to obtain Mel-frequency cepstral coefficient features based on the noise reduction speech signal, match and analyze the Mel-frequency cepstral coefficient features with the dynamic noise feature library, output the noise scene classification result, and switch the corresponding speech recognition model; Calculation module 4, configured to output a voice interaction signal through the voice recognition model, obtain a voice command according to the voice interaction signal, and calculate a confidence value according to the voice interaction signal and a current environment signal-to-noise ratio; Output module 5, used to obtain the dynamic confidence qualified threshold, compare the confidence value with the dynamic confidence qualified threshold, and output multimodal verification data; The verification execution module 6 is used to analyze and judge the multimodal verification data and output an authority control signal.
[0072] In one embodiment, the acquisition module includes: Sampling an audio signal at a sampling rate of 16 kHz, dividing the audio signal into frames and performing Hanning window processing on the audio signal, and obtaining a processed audio signal; Calculating an energy ratio of an active speech segment to a silent segment in the processed audio signal, and obtaining a current environmental signal-to-noise ratio; Decompose the processed audio signal into three frequency bands: low frequency band, mid frequency band and high frequency band; The noise reduction process is performed on the processed audio signals in the low frequency band, mid frequency band and high frequency band respectively. A time domain waveform is obtained, and a noise reduction speech signal is obtained according to the time domain waveform.
[0073] The present application also provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0074] The present application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the above method when executed by a processor.
[0075] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, value library or other media provided in this application and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct RAM bus dynamic RAM (DRDRAM), and RAM bus dynamic RAM (RDRAM).
[0076] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, apparatus, article, or method comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, apparatus, article, or method. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, apparatus, article, or method comprising the element.
[0077] The above description is only a preferred embodiment of the present invention and does not limit the scope of this application. Any equivalent results or equivalent process transformations made using the contents of the description and drawings of the present invention, or directly or indirectly applied in other related technical fields, are also included in the scope of protection of this application.
Claims
1. An AI intelligent robot voice interaction method, characterized in that: The following steps are involved: Acquire a noise sample set, construct a dynamic noise feature library including steady-state noise, impact noise, and human voice interference features based on the noise sample set, and acquire a pre-stored gesture command library; Collect audio signals in real time, obtain the current environmental signal-to-noise ratio based on the audio signals, perform differentiated noise reduction processing on audio signals in the low-frequency band, mid-frequency band, and high-frequency band, and generate a noise-reduced speech signal; Acquire Mel-frequency cepstral coefficient features according to the noise reduction speech signal, perform matching analysis on the Mel-frequency cepstral coefficient features and a dynamic noise feature library, output a noise scene classification result, and switch a corresponding speech recognition model; Outputting a voice interaction signal through the voice recognition model, obtaining a voice command according to the voice interaction signal, and calculating a confidence value according to the voice interaction signal and a current environment signal-to-noise ratio; Obtain a dynamic confidence threshold, compare the confidence value with the dynamic confidence threshold, and output multimodal verification data; Analyze and judge the multimodal verification data and output the permission control signal.
2. The AI intelligent robot voice interaction method according to claim 1, characterized in that: The steps of collecting audio signals in real time, obtaining the current environmental signal-to-noise ratio according to the audio signals, performing differentiated noise reduction processing on audio signals in the low-frequency band, the mid-frequency band, and the high-frequency band, and generating a noise-reduced speech signal include: Sampling an audio signal at a sampling rate of 16 kHz, dividing the audio signal into frames and performing Hanning window processing on the audio signal, and obtaining a processed audio signal; Calculating an energy ratio of an active speech segment to a silent segment in the processed audio signal, and obtaining a current environmental signal-to-noise ratio; Decompose the processed audio signal into three frequency bands: low frequency band, mid frequency band and high frequency band; The noise reduction process is performed on the processed audio signals in the low frequency band, mid frequency band and high frequency band respectively. A time domain waveform is obtained, and a noise reduction speech signal is obtained according to the time domain waveform.
3. The AI intelligent robot voice interaction method according to claim 1, characterized in that: The steps of obtaining Mel-frequency cepstral coefficient features according to the noise reduction speech signal, performing matching analysis on the Mel-frequency cepstral coefficient features and a dynamic noise feature library, outputting a noise scene classification result, and switching a corresponding speech recognition model include: Performing frame processing on the noise reduction speech signal to obtain a frame signal; The framed signal is processed by a 26-channel Mel filter bank to obtain logarithmic energy, and the first 12-dimensional Mel-frequency cepstral coefficient features are extracted by discrete cosine transform; Perform cosine similarity matching on the Mel-frequency cepstral coefficient feature and the dynamic noise feature library, and output a noise scene classification result, wherein the noise scene classification result includes a steady-state noise class and an impact noise class; For steady-state noise, the DNN-HMM hybrid model is used; For impact noise scenarios, the temporal attention Transformer model is used.
4. The AI intelligent robot voice interaction method according to claim 1, characterized in that: The steps of outputting a voice interaction signal through the voice recognition model, acquiring a voice instruction according to the voice interaction signal, and calculating a confidence value according to the voice interaction signal and a current environment signal-to-noise ratio include: Acquiring a voice command according to the voice interaction signal; Obtain the word probability and the total energy value of the entire frequency band based on the voice interaction signal; If the word probability is lower than 0.4, the voice command is deemed invalid; If the term probability is not less than 0.4, extracting the energy value of the 500-4000 Hz human voice main frequency band in the voice interaction signal ratio, and obtaining the voice activity detection energy ratio based on the 500-4000 Hz human voice main frequency band energy value and the total energy value of the entire frequency band; The confidence value is calculated according to the term probability, the voice activity detection energy ratio, and the current environment signal-to-noise ratio.
5. The AI intelligent robot voice interaction method according to claim 4, characterized in that: The steps of obtaining a dynamic confidence threshold, comparing the confidence value with the dynamic confidence threshold, and outputting multimodal verification data include: Obtaining a dynamic confidence threshold, where the confidence threshold increases linearly within a preset interval; Compare the confidence value to the dynamic confidence threshold: When the confidence value is not lower than the dynamic confidence qualification threshold, the voice command is output; When the confidence value is lower than the dynamic confidence qualification threshold, the operator's gesture is acquired, analyzed and matched with the pre-stored gesture instruction library, and the voice instruction and gesture instruction are output simultaneously; Based on the comparison result of the confidence value and the dynamic confidence qualification threshold, multimodal verification data including voice commands or voice commands and gesture commands is output.
6. The AI intelligent robot voice interaction method according to claim 5, characterized in that: The step of analyzing and judging the multimodal verification data and outputting the authority control signal includes: Get the voice permission level including preset voiceprint features; Extracting input voiceprint features of voice commands in multimodal verification data; Determine whether the input voiceprint features match the preset voiceprint features and voice permission level; If there is a match, and when the multimodal verification data only contains voice commands, the voice commands are executed; If there is no match, and when the multimodal verification data only contains voice commands, the voice commands will not be executed and a voice reminder will be issued; If they match, and when the multimodal verification data includes voice commands and gesture commands, then determine whether the voice commands and gesture commands match; If the two match, the voice command and gesture command in the multimodal verification data are executed; If the two do not match, the program will not be executed and a warning will be triggered; If there is no match, and when the multimodal verification data includes voice commands and gesture commands, it will not be executed and a warning will be triggered.
7. An AI intelligent robot voice interaction system, characterized in that: include: An acquisition module is used to acquire a noise sample set, construct a dynamic noise feature library including steady-state noise, impact noise and human voice interference features based on the noise sample set, and acquire a pre-stored gesture command library; An acquisition module is used to acquire audio signals in real time, obtain the current environmental signal-to-noise ratio based on the audio signals, perform differentiated noise reduction processing on the audio signals in the low-frequency band, mid-frequency band, and high-frequency band, and generate a noise-reduced speech signal; A switching module is used to obtain Mel-frequency cepstral coefficient features based on the noise reduction speech signal, match and analyze the Mel-frequency cepstral coefficient features with a dynamic noise feature library, output a noise scene classification result, and switch the corresponding speech recognition model; a calculation module, configured to output a voice interaction signal through the voice recognition model, obtain a voice instruction according to the voice interaction signal, and calculate a confidence value according to the voice interaction signal and a current environment signal-to-noise ratio; An output module is used to obtain a dynamic confidence threshold, compare the confidence value with the dynamic confidence threshold, and output multimodal verification data; The verification execution module is used to analyze and judge the multimodal verification data and output the permission control signal.
8. The AI intelligent robot voice interaction system according to claim 7, characterized in that: The acquisition module includes: Sampling an audio signal at a sampling rate of 16 kHz, dividing the audio signal into frames and performing Hanning window processing on the audio signal, and obtaining a processed audio signal; Calculating an energy ratio of an active speech segment to a silent segment in the processed audio signal, and obtaining a current environmental signal-to-noise ratio; Decompose the processed audio signal into three frequency bands: low frequency band, mid frequency band and high frequency band; The noise reduction process is performed on the processed audio signals in the low frequency band, mid frequency band and high frequency band respectively. A time domain waveform is obtained, and a noise reduction speech signal is obtained according to the time domain waveform.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to claims 1 to 6 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to claims 1 to 6 are implemented.
Citation Information
Patent Citations
Method and device for automatically identifying cough
CN101894551A
Method for adjusting confidence coefficient threshold of voice recognition and electronic device
CN103578468A
Voice navigation method and device and terminal equipment
CN110085217A
Medical equipment and control method thereof
CN110368097A
Speech recognition model training method and apparatus, storage medium and electronic device
CN110544469A
Cited By
Voice interaction method and system of artificial intelligence safety assistant
CN121034307A
Digital human interaction method and system based on multi-mode sensing intelligent action switching
CN121050590A
Instrument control method and system based on voice interaction
CN121214930A
Instrument control method and system based on voice interaction
CN121214930B
Patrol robot voice interaction system with environment perception capability
CN121260164A