A fume hood window lifting voice analysis system based on semantic enhancement
By using harmonic notch suppression and glottal excitation confidence-weighted feature extraction, combined with grammatical constraints and multi-layer verification, the problem of low accuracy in recognizing fume hood window lifting commands was solved, and safe and reliable voice control was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PAI LAB EQUIP CO LTD
- Filing Date
- 2026-03-19
- Publication Date
- 2026-07-31
AI Technical Summary
Existing voice analysis systems suffer from low signal-to-noise ratios and a lack of reverse interaction in semantic understanding in fume hood laboratory scenarios, resulting in low accuracy in recognizing fume hood window lifting commands and posing a risk of miscontrol.
Harmonic notch suppression, glottal excitation confidence-weighted feature extraction, and grammatical constraint recognition are employed, combined with multi-layer verification of motion direction, position rationality, and noise consistency, to achieve hierarchical safe execution of lifting commands.
It significantly improves the detection accuracy and robustness of fume hood window lifting commands, reduces the risk of false triggering, and enhances the safety of laboratory operations.
Smart Images

Figure CN121862096B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech recognition technology, specifically to a speech analysis system for lifting and lowering a fume hood window based on semantic enhancement. Background Technology
[0002] Speech analysis technology is an important means of realizing human-computer interaction. Its core technical path usually includes three stages: acoustic front-end processing, speech recognition, and semantic understanding. Existing speech analysis systems generally adopt a unidirectional linear architecture in which acoustic processing and semantic understanding are independent. That is, after the acoustic model converts the collected speech signal into text, the natural language understanding module performs intent recognition on the text. The knowledge of the semantic layer is not visible to the acoustic layer processing process, and there is no reverse interaction path between the two.
[0003] In laboratory applications using fume hoods, the aforementioned architecture exhibits significant shortcomings. Firstly, the continuous broadband mechanical noise generated by the fume hood exhaust fan highly overlaps with the human voice spectrum, resulting in an extremely low signal-to-noise ratio. Consequently, the recognized text output by the acoustic model inevitably contains numerous erroneous words. Since semantic knowledge cannot be used to correct the acoustic layer, this erroneous text is directly transmitted to the semantic understanding module, amplifying the cumulative recognition errors. The system lacks the ability to semantically correct noisy recognized text. Secondly, the speech generated by laboratory operators during work is a continuous stream of experimental operation language, with equipment control commands embedded within the experimental statements. These two are indistinguishable acoustically. Due to a lack of specialized semantic knowledge in the laboratory domain, the existing system cannot accurately pinpoint the semantic boundaries of control commands, easily misinterpreting experimental statements as control intentions or overlooking genuine lifting commands. These two shortcomings combine to result in low accuracy in recognizing and extracting fume hood window lifting commands, posing a risk of miscontrol in safety-sensitive laboratory operating scenarios.
[0004] To address this, a voice analysis system for the lifting and lowering of fume hood windows based on semantic enhancement is proposed. Summary of the Invention
[0005] The purpose of this invention is to provide a voice analysis system for lifting and lowering fume hood windows based on semantic enhancement. Through rotational speed feedback adaptive harmonic suppression, glottal excitation confidence weighted feature extraction and grammatical constraint recognition, combined with multi-layer verification of motion direction, position rationality and noise consistency, the system can achieve hierarchical and safe execution of lifting and lowering commands.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A voice analysis system for lifting and lowering fume hood windows based on semantic enhancement, comprising:
[0008] The harmonic notch suppression module determines the blade passing frequency and harmonic frequencies in real time based on the speed feedback signal, performs harmonic notch suppression on the collected voice signal, and obtains the harmonic suppressed voice signal and the residual energy vector of each harmonic frequency band.
[0009] The time-frequency feature extraction module calculates the glottal excitation autocorrelation coefficient frame by frame for the harmonic-suppressed speech signal, cross-verifies the glottal excitation confidence with the residual energy vector, constructs a time-frequency reliability mask, and extracts acoustic features by weighting the Mel band to obtain a reliability weighted feature matrix and a frame-by-frame glottal excitation confidence sequence.
[0010] The steps for determining the reliability weights for each Mel frequency band are as follows: the overall glottal excitation confidence of the current frame is applied to all frequency bands; the residual harmonic energy ratio corresponding to the frequency range of the residual energy vector is used to allocate frequency band reliability weights based on the residual harmonic energy ratio; and the reliability weights for each Mel frequency band in the current frame are obtained based on the glottal excitation confidence and the residual harmonic energy ratio.
[0011] The semantic enhancement instruction detection module explicitly weights the frame-level acoustic posterior probability with the glottal excitation confidence sequence, suppresses the contribution of low-confidence frames to the detection score of rising and falling keywords, and applies subject grammatical constraints to candidate rising and falling verbs to distinguish between rising and falling instructions and experimental statements, thereby obtaining candidate texts of rising and falling instructions and comprehensive detection confidence.
[0012] The safety execution decision module sequentially performs motion direction feasibility verification, target position quantification rationality verification, and broadband turbulent noise energy and opening feedback value consistency verification on candidate texts. The three-layer verification results, together with the comprehensive detection confidence, drive the hierarchical safety execution decision.
[0013] Preferably, the process of obtaining the harmonic-suppressed speech signal and the residual energy vector of each harmonic frequency band is as follows: the rotational speed feedback signal is collected and converted into an instantaneous rotational speed value; the fundamental frequency through which the current blade passes is determined based on the number of blades and the instantaneous rotational speed; and the harmonic frequency values of each order are calculated in sequence to form a set of harmonic frequencies; taking each frequency in the set of harmonic frequencies as the center, the suppression bandwidth of each notch unit is determined by combining the short-time fluctuation amplitude of the rotational speed signal within the time period; and the notch units are cascaded and applied to the speech signal to output a harmonic-suppressed speech signal.
[0014] After notch suppression is completed, the signal energy before and after suppression in the neighborhood of each harmonic center frequency is counted to obtain the degree of harmonic suppression completion in each frequency band, and the residual energy vector is formed by arranging them according to the harmonic order.
[0015] Preferably, the construction process of the time-frequency reliability mask is as follows: the harmonic-suppressed speech signal is divided into frames according to a preset frame length and frame shift, and the autocorrelation peak is searched in the delay interval within the corresponding Chinese pronunciation fundamental frequency range for each frame signal. The glottal excitation confidence of the current frame is defined based on the peak amplitude and the zero-delay autocorrelation amplitude.
[0016] The reliability weights are determined by combining two factors for each Mel frequency band, and then arranged in frame order and frequency band order to form a time-frequency reliability mask.
[0017] Preferably, the process of obtaining the reliability weighted feature matrix and the frame-by-frame glottic excitation confidence sequence is as follows: for the harmonic-suppressed speech signal, the filter output energy is calculated in each Mel frequency band according to the framing results; the frequency band energy output is scaled by the reliability weight of the corresponding frame and frequency band in the time-frequency reliability mask; for frequency bands with reliability weights lower than the preset reliability judgment threshold, the long-term average value of historical frame energy accumulated after the frequency band is started is used to replace the current frame observation value; for frequency bands with reliability weights not lower than the judgment threshold, the current frame's true observed energy value is retained.
[0018] Logarithms are taken for the energy values after scaling and / or substitution of each frequency band, and discrete cosine transform is performed to extract cepstral coefficients. These coefficients are then arranged in frame order to form a reliability-weighted feature matrix. Simultaneously, the ratio of the peak value of glottal excitation autocorrelation to the zero-delay autocorrelation value in each frame is arranged in frame order to form a frame-by-frame glottal excitation confidence sequence.
[0019] Preferably, the process of obtaining the candidate text for the elevation and descent instructions and the comprehensive detection confidence is as follows: taking the reliability-weighted feature matrix as input, the posterior probability of each phoneme is calculated through the acoustic model, and the posterior probability of the current frame is multiplicatively scaled by the confidence value corresponding to each frame in the frame-by-frame glottal excitation confidence sequence; when performing the elevation and descent keyword sequence search on the posterior probability sequence, the total keyword score is mainly driven by high-confidence frames, and the influence of random activation of low-confidence frames on the total score is systematically suppressed, forming a reliability-weighted keyword detection score;
[0020] For the detected candidate ascending / descending verbs, search for subject nouns within the preset word position range: if a subject noun belonging to the preset experimental quantifier vocabulary exists, the verb is determined to be in the experimental statement syntax structure, and a penalty reduction is applied to the instruction confidence; if no subject noun exists and / or the subject belongs to the window component vocabulary, the original detection confidence is maintained, and it is confirmed as an ascending / descending instruction structure.
[0021] The reliability-weighted keyword detection score is combined with the grammatical constraint verification result to form a comprehensive detection confidence score, which together with the corresponding candidate ascending / descending verbs and quantification modifiers constitute the candidate text for ascending / descending instructions.
[0022] Preferably, the three-layer verification process is as follows:
[0023] Motion direction feasibility verification: Read the opening feedback value to determine the current position of the window, extract the motion direction indicated by the rising and falling verbs in the candidate text, and if the current position is already at the upper limit of the physical travel and the instruction direction is to continue to increase the opening, and / or is already at the lower limit of the travel and the instruction direction is to continue to decrease the opening, the instruction is determined to be unexecutable, the direction verification fails to pass the mark is output, the voice prompt is triggered and the subsequent verification is terminated.
[0024] Target position quantization rationality verification: Map proportional positions in candidate text to target opening percentages, add a preset step amount to incremental positions based on the current opening, and map endpoint positions to the upper and / or lower limits of the travel range. Calculate the displacement between the target position and the current position. If the displacement does not exceed the minimum effective travel limit, the target position and the current position are determined to be substantially the same, and a rationality verification failure flag is output; otherwise, a pass flag and the target position value are output.
[0025] Broadband turbulent noise energy and opening consistency verification: Extract the total broadband turbulent noise energy outside the harmonic frequency band in the non-speech segment, compare it with the opening-broadband turbulent noise energy comparison table established by offline calibration, and determine whether the current broadband noise energy falls into the calibration confidence interval corresponding to the sensor reported opening value. If it falls into the interval, output a consistency verification pass mark; otherwise, output an abnormal mark.
[0026] Preferably, the method for obtaining hierarchical safety execution decisions is as follows: combining the results of three-layer verification with the overall detection confidence level, and judging the execution conditions according to serial logic: when any layer verification output fails and / or an anomaly mark is displayed, the corresponding voice prompt is triggered and the current execution process is terminated, and no drive commands are output to the motion lifting.
[0027] After all three layers of verification pass, a tiered response is executed based on the overall detection confidence interval: when the confidence level is in the high confidence interval, the target position value is output to the motion lift, driving the window to move to the target position; when the confidence level is in the medium confidence interval, the speech synthesis module broadcasts the parsed target action description, waits for the operator to press a button for confirmation, and then outputs the direction motion lift drive command; when the confidence level is in the low confidence interval, the current candidate text is discarded and the operator is asked to reissue the command via voice prompt.
[0028] The motion lifting and lowering forms a closed loop based on the target position value and the opening feedback value. After the window reaches the target position, the motion stops, and the speech synthesis broadcasts the execution completion status, completing the complete semantic recognition to safe execution closed loop.
[0029] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0030] 1. This invention dynamically calculates the blade passage frequency and its multiple harmonic frequencies by real-time reading of the impeller speed feedback signal. It then constructs a cascaded notch filter unit that adaptively adjusts its bandwidth according to operating conditions to specifically suppress structured periodic noise in the speech signal. Simultaneously, it statistically analyzes the energy differences before and after harmonic suppression to form a residual energy vector, providing a quantitative basis for subsequent reliability modeling. This scheme can adjust the suppression parameters in real-time according to speed fluctuations, avoiding speech distortion caused by over-suppression or misidentification caused by insufficient suppression. It maintains high speech intelligibility and keyword detection stability even under high wind speed and strong turbulence conditions.
[0031] 2. This invention constructs glottal excitation confidence by calculating the autocorrelation peak of the glottal excitation frame by frame, and generates a time-frequency reliability mask by combining the proportion of harmonic residual energy, thus differentially weighting the Mel band features. Simultaneously, it applies explicit multiplicative suppression to low-confidence frames at the acoustic posterior probability level, ensuring that keyword scores are primarily driven by high-reliability speech segments. This scheme achieves a two-layer reliability control from acoustic feature construction to posterior probability fusion, systematically reducing the probability of random activation induced by fan noise, significantly improving the accuracy and robustness of rising and falling keyword detection, and reducing the risk of false triggering.
[0032] 3. After obtaining candidate instruction text, this invention introduces subject grammatical constraints to distinguish between operational instructions and experimental statements. Furthermore, it employs three layers of physical logic verification—feasibility of motion direction verification, rationality verification of target position quantification, and consistency verification of broadband turbulence noise and opening feedback—to progressively screen instruction executability. Finally, it implements a high, medium, and low confidence-based graded response strategy based on comprehensive detection confidence. This invention deeply couples semantic understanding, real-time equipment status, and safety boundary conditions, effectively avoiding safety hazards caused by out-of-bounds movement, malfunctions, and sensor anomalies, significantly improving the inherent safety of fume hood operation. Attached Figure Description
[0033] Figure 1 A schematic diagram of the voice analysis logic flow for lifting and lowering the window of a fume hood based on semantic enhancement, provided for this invention;
[0034] Figure 2 A schematic diagram of a voice analysis system for lifting and lowering a fume hood window based on semantic enhancement, provided for this invention;
[0035] Figure 3 This is a schematic diagram of the three-layer security verification process provided by the present invention. Detailed Implementation
[0036] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention.
[0037] Example 1:
[0038] This invention addresses the need for voice-controlled window raising and lowering in fume hood laboratory settings by proposing a speech analysis system based on semantic enhancement. The core technical path of this system consists of four sequentially connected processing stages: harmonic noise suppression, time-frequency reliability feature extraction, semantic enhancement command detection, and physical verification for secure execution. The output data of each stage serves as the direct input to the next stage, forming a complete data flow chain.
[0039] Please see Figure 1 and Figure 2 The present invention provides a voice analysis system for lifting and lowering a fume hood window based on semantic enhancement. The technical solution is as follows: a harmonic notch suppression module determines the passing frequency of the blades and the harmonic frequencies of each order in real time based on the rotational speed feedback signal, performs harmonic notch suppression on the collected voice signal, and obtains the harmonic suppressed voice signal and the residual energy vector of each harmonic frequency band.
[0040] The time-frequency feature extraction module calculates the glottal excitation autocorrelation coefficient frame by frame for the harmonic-suppressed speech signal, cross-verifies the glottal excitation confidence with the residual energy vector, constructs a time-frequency reliability mask, and extracts acoustic features by weighting the Mel band to obtain a reliability weighted feature matrix and a frame-by-frame glottal excitation confidence sequence.
[0041] The semantic enhancement instruction detection module explicitly weights the frame-level acoustic posterior probability with the glottal excitation confidence sequence, suppresses the contribution of low-confidence frames to the detection score of rising and falling keywords, and applies subject grammatical constraints to candidate rising and falling verbs to distinguish between rising and falling instructions and experimental statements, thereby obtaining candidate texts of rising and falling instructions and comprehensive detection confidence.
[0042] The safety execution decision module sequentially performs motion direction feasibility verification, target position quantification rationality verification, and broadband turbulent noise energy and opening feedback value consistency verification on candidate texts. The three-layer verification results, together with the comprehensive detection confidence, drive the hierarchical safety execution decision.
[0043] Furthermore, the process of obtaining the harmonic-suppressed speech signal and the residual energy vectors of each harmonic frequency band is as follows:
[0044] The speed feedback signal is continuously collected and converted into an instantaneous speed value. The fundamental frequency of the current blade is determined based on the product of the number of exhaust fan blades and the instantaneous speed. The harmonic frequencies of each order are calculated in turn to form a set of harmonic frequencies that are updated in real time.
[0045] Centered on each frequency in the harmonic frequency set, the suppression bandwidth of each notch unit is dynamically determined by combining the short-term fluctuation amplitude of the speed signal within that time period. When the speed fluctuation amplitude is large, the bandwidth is appropriately widened to cover the frequency jitter range. When the speed is stable, the bandwidth is narrowed to maximize the retention of speech components outside the notch. The above notch units are cascaded and applied to the speech signal to output a harmonic suppressed speech signal.
[0046] After harmonic suppression is completed, the signal energy before and after suppression in the neighborhood of each harmonic center frequency is counted. The ratio of the two values represents the degree of harmonic suppression in each frequency band. The residual energy vector is formed by arranging the two values according to the harmonic order and is transmitted to the next frequency along with the harmonic-suppressed speech signal. The higher the residual proportion in the residual energy vector, the less sufficient the harmonic component suppression in that frequency band is. The subsequent processing will reduce the reliability weight of the corresponding frequency band accordingly.
[0047] Specifically, during operation, the fume hood exhaust fan generates harmonic noise centered on the blade passage frequency and its harmonics. The blade passage frequency is determined by the number of fan blades and the instantaneous rotational speed. The number of blades is an inherent parameter of the exhaust fan and is entered during the configuration phase. The instantaneous rotational speed is continuously acquired through a speed feedback signal. This speed feedback signal is generated by a Hall effect sensor installed in the exhaust fan drive circuit. This sensor outputs a pulse every time the blades rotate through a certain angle, and the instantaneous rotational speed value is calculated by counting the number of pulses per unit time. The instantaneous rotational speed value is continuously updated with a fixed sampling period (this sampling period is determined during system debugging based on the exhaust fan's rotational speed response characteristics, requiring a sufficiently short sampling period to capture instantaneous changes in rotational speed, typically on the order of milliseconds). Based on this, the fundamental frequency of the blade passage and its second to fifth harmonics are calculated in real time, forming the set of harmonic frequencies at the current moment.
[0048] For each frequency in the harmonic frequency set, a narrowband notch filter unit is constructed to suppress harmonic components in the speech signal, centered on that frequency. The suppression bandwidth of each notch filter unit is dynamically adjusted according to the short-time fluctuation amplitude of the current speed signal. The short-time fluctuation amplitude is characterized by calculating the dispersion of speed samples within a preset time window. Specifically, the absolute value of the difference between each speed sample value and its mean within the time window is averaged to obtain the mean absolute deviation, which serves as a quantitative representation of the fluctuation amplitude.
[0049] The method for obtaining the speed fluctuation range judgment threshold is as follows. Upon initial power-on, a two-stage automatic calibration process is initiated. In the first stage, with the exhaust fan running at rated speed under no-load conditions (no voice signal input, window fixed in a certain position), speed feedback signals are continuously collected for at least two minutes, and the sample-by-sample average absolute deviation sequence of this signal is recorded. In the second stage, the window is driven to perform at least one complete lifting and lowering operation, simultaneously collecting the full-range speed feedback signal and calculating the corresponding sample-by-sample average absolute deviation sequence. All average absolute deviation samples collected in the two stages are merged, and the median value of the merged distribution is statistically analyzed. This median value is used as the judgment threshold between "large fluctuation range" and "stable speed." When the real-time average absolute deviation is higher than this threshold, the current speed fluctuation range is determined to be large, and the notch filter bandwidth is adjusted to be wider; when it is lower than or equal to this threshold, the current speed is determined to be stable, and the notch filter bandwidth is adjusted to be narrower. By introducing speed fluctuation data under window movement conditions, it is ensured that the threshold value can effectively distinguish between the two actual operating states: no-load stability and window movement.
[0050] Each notch filter unit adopts a second-order infinite impulse response structure, which is updated in real time according to the current center frequency and bandwidth. Each notch filter unit is cascaded in sequence to act on the speech signal and output a harmonic-suppressed speech signal.
[0051] The residual energy vector is calculated as follows: After notch filtering, for each harmonic order, the energy of the speech signal before and after notch filtering is calculated within a frequency band centered on the order frequency and encompassing the corresponding notch bandwidth. The ratio of the processed energy to the unprocessed energy is taken as the residual energy percentage for that harmonic order, with a value range of 0 to 1. The residual energy percentages for each harmonic order are arranged in ascending order to form the residual energy vector. Each element of this vector corresponds to the degree of suppression for a given harmonic order; the closer the value is to zero, the more sufficient the harmonic suppression for that order; the closer the value is to 1, the less sufficient the harmonic suppression for that order. The harmonic-suppressed speech signal and the residual energy vector are then passed together to the next stage.
[0052] Furthermore, the construction process of the time-frequency reliability mask is as follows:
[0053] The harmonic-suppressed speech signal is divided into frames according to a preset frame length and frame shift. For each frame, the autocorrelation peak is searched within the delay interval of the corresponding Chinese pronunciation fundamental frequency range. The ratio of the peak amplitude to the zero-delay autocorrelation amplitude is defined as the glottal excitation confidence of the current frame. A high confidence indicates that the frame is dominated by periodic glottal excitation and the acoustic information is reliable; a low confidence indicates that the frame is dominated by noise or unvoiced sounds and the acoustic information is unreliable.
[0054] For each Mel frequency band, two factors are combined to determine the reliability weight: the first is the overall glottal excitation confidence of the current frame, which applies to all frequency bands; the second is the proportion of residual harmonic energy in the residual energy vector corresponding to the frequency range of that frequency band. A higher residual proportion reduces the reliability weight of that frequency band, while a lower residual proportion does not result in any additional penalty for the reliability weight of that frequency band. The two factors are multiplied and normalized to obtain the reliability weight of each Mel frequency band in the current frame, which is then arranged in frame order and frequency band order to form a time-frequency reliability mask. Each element in the time-frequency reliability mask corresponds to the reliability weight of a specific frame and a specific Mel frequency band, and as a whole, it reflects the acoustic reliability distribution of the speech signal in both time and frequency dimensions, which is used for subsequent feature extraction weighting.
[0055] Specifically, the harmonic-suppressed speech signal is processed by framing according to a preset frame length and frame shift. The frame length and frame shift are determined during the system design phase based on the short-term stationarity requirements of the speech signal. The frame length value must satisfy the requirement that the statistical characteristics of the speech signal are approximately stationary within that duration, and the frame shift value must satisfy the requirement that there is sufficient overlap between adjacent frames to ensure temporal continuity.
[0056] For each speech frame, a normalized autocorrelation sequence is calculated within the pitch period delay interval corresponding to the fundamental frequency range of Mandarin pronunciation (the fundamental frequency of Mandarin pronunciation is usually distributed within the normal vocal frequency range of adults, and the corresponding pitch period interval is approximately several milliseconds to more than ten milliseconds; the specific interval is determined during the system configuration phase based on the vocal characteristics of the target user group). The maximum peak value within this delay interval is searched, and the ratio of this maximum peak value to the autocorrelation value at zero delay is defined as the glottal excitation confidence of the current frame. When this ratio is close to one, it indicates that the current frame signal has significant quasi-periodicity, dominated by glottal excitation; when this ratio is close to zero, it indicates that the current frame signal has no periodicity, dominated by noise or unvoiced sounds.
[0057] The time-frequency reliability mask is constructed as follows: For each Mel band in the current frame, the glottal excitation confidence is first used as a global factor; then, the proportion of residual energy corresponding to the harmonic order covered by the frequency range of the Mel band is extracted from the residual energy vector as the frequency band specific penalty factor of the band.
[0058] During the calibration phase, based on the distribution of the residual energy proportion of each harmonic order after notch filtering, the mean of the residual proportion plus one standard deviation is used as the criterion for "high residual proportion." That is, when the residual proportion of a frequency band exceeds this threshold, it is considered that harmonic suppression in that band is insufficient, and the reliability weight of the corresponding frequency band is penalized with a reduction. Conversely, the mean minus one standard deviation is used as the criterion for "low residual proportion." That is, when the residual proportion of a frequency band is below this threshold, it is considered that harmonic suppression in that band is sufficient, and the reliability weight of the corresponding frequency band is not subject to additional penalty. For residual proportions between the two thresholds, the penalty intensity is calculated using linear interpolation to achieve a smooth transition from no penalty to full penalty.
[0059] After multiplying the glottal excitation confidence level by the band-specific penalty factor, the result is normalized: the product of the glottal excitation confidence level and the band-specific penalty factor at their maximum possible values (both being one) (i.e., the theoretical maximum value of one) is used as the divisor. The product of each band is then divided by this theoretical maximum value to obtain the reliability weight of each Mel band in the current frame. Since the theoretical maximum value is a fixed constant, this normalization method does not depend on the actual data of the current frame and can still maintain numerical stability in frames with extremely low glottal excitation confidence levels, with the weight values remaining stable between 0 and 1.
[0060] Furthermore, the process of obtaining the reliability-weighted feature matrix and the frame-by-frame glottal excitation confidence sequence is as follows:
[0061] The filter output energy of the harmonic-suppressed speech signal is calculated in each Mel frequency band according to the framing results, and the energy output of the frequency band is scaled by the reliability weight of the corresponding frame and frequency band in the time-frequency reliability mask.
[0062] For frequency bands with a reliability weight lower than the preset reliability judgment threshold, the long-term average value of historical frame energy accumulated in the frequency band after system startup is used to replace the current frame observation value to avoid introducing the distortion characteristics of noise-dominated frames into subsequent identification; for frequency bands with a reliability weight not lower than the threshold, the current frame's true observed energy value is retained.
[0063] Logarithms are taken for the energy values after scaling or substitution of each frequency band, and discrete cosine transform is performed to extract cepstral coefficients. These coefficients are then arranged in frame order to form a reliability-weighted feature matrix. Simultaneously, the ratio of the peak value of the glottal excitation autocorrelation to the zero-delay autocorrelation in each frame is arranged in frame order to form a frame-by-frame glottal excitation confidence sequence. The reliability-weighted feature matrix and the frame-by-frame glottal excitation confidence sequence are passed forward together. The former is used for acoustic posterior probability calculation, and the latter is used for frame-level posterior probability weighting.
[0064] For each frame of the harmonic-suppressed speech signal, a corresponding triangular filter bank is applied to each Mel frequency band, and the filter output energy of each band is calculated. The energy value of that frequency band is scaled by the reliability weight of the corresponding frame and frequency band in the time-frequency reliability mask, that is, the filter output energy is multiplied by the reliability weight of the corresponding position.
[0065] The reliability threshold is determined during the system calibration phase as follows: Under the calibration environment, at least fifty frames are collected for each of the following types of frames: voiced frames (voiced frames, where the operator reads words containing the main vowels) and unvoiced frames (voiceless frames and pure noise frames). The glottal excitation confidence distribution of the two types of frames is statistically analyzed. The 10th percentile of the confidence distribution of the voiced frames is compared with the 90th percentile of the confidence distribution of the voiceless and noise frames. The median value of the two is taken as the reliability threshold to make the false positive rate tend to be balanced in both directions. This threshold value is stored in the system configuration file.
[0066] For frequency bands where the reliability weight is lower than the reliability judgment threshold, the filter output energy value of the current frame is replaced by the average of the sliding window of all historical frame energy values of the frequency band since the system started (the sliding window length is determined during the calibration stage based on the typical duration of the speech signal to cover enough historical frames so that the average value has statistical stability); for frequency bands where the reliability weight is not lower than the threshold, the actual filter output energy value of the current frame is retained.
[0067] The energy values of each frequency band, after scaling or substitution, are logarithmically calculated to base 10 to form a logarithmic filter bank energy sequence. A discrete cosine transform is then performed on this sequence, and the first few order coefficients of the transform result are taken as the cepstral coefficients of the current frame. These coefficients are arranged in frame order to form a reliability-weighted feature matrix. Simultaneously, the glottal excitation confidence scores (i.e., the ratio of peak autocorrelation to zero-delay autocorrelation) calculated for each frame are arranged in frame order to form a frame-by-frame glottal excitation confidence score sequence. The reliability-weighted feature matrix and the frame-by-frame glottal excitation confidence score sequence are then propagated forward together.
[0068] Furthermore, the process of obtaining the candidate texts for the elevation / decrease commands and the overall detection confidence is as follows:
[0069] Using the reliability-weighted feature matrix as input, the posterior probability of each phoneme is calculated frame by frame through the acoustic model. The posterior probability of each frame is multiplicatively scaled using the confidence value corresponding to each frame in the frame-by-frame glottal excitation confidence sequence. The posterior probability contribution of low-confidence frames is suppressed, while the posterior probability contribution of high-confidence frames is preserved. When performing ascending and descending keyword sequence search on the posterior probability sequence, the total keyword score is mainly driven by high-confidence frames, and the impact of random activation of low-confidence frames on the total score is systematically suppressed, forming a reliability-weighted keyword detection score.
[0070] For the detected candidate ascending / descending verbs, search for subject nouns within the preset word positions before them: if there is a subject noun belonging to the preset experimental quantifier vocabulary, the verb is determined to be in the experimental statement syntax structure, and its instruction confidence is penalized by a reduction in weight; if there is no subject noun before it or the subject belongs to the window component vocabulary, the original detection confidence is maintained, and it is confirmed as an ascending / descending instruction structure.
[0071] The reliability-weighted keyword detection score is combined with the grammatical constraint verification result to form a comprehensive detection confidence score. This score, together with the corresponding candidate ascending / descending verbs and their quantification modifiers, constitutes the ascending / descending instruction candidate text, which is then passed on.
[0072] The acoustic model receives a reliability-weighted feature matrix and outputs the frame-level posterior probability of each phoneme for each frame. Taking the reliability-weighted feature matrix as input, the acoustic model processes the cepstral coefficient vector of each frame in the matrix through three processing layers, ultimately outputting the posterior probability vector of the frame belonging to each phoneme category.
[0073] The input projection layer performs a linear transformation on the cepstral coefficient vector of the current frame, mapping it to a higher-dimensional intermediate representation space. The output real-valued vector is then fed into the hidden representation layer group. The hidden representation layer group consists of fully connected layers connected in series. Each layer performs a linear affine transformation and element-wise nonlinear activation on the output of the previous layer. The nonlinear activation truncates negative values to zero and preserves positive values. After multiple transformations, a high-level acoustic feature vector is output. The output probability layer performs a linear classification transformation on this vector to obtain a logical score vector of the same dimension as the phoneme set. Then, it performs an exponential operation on all elements and normalizes the vector so that the sum of all elements is one, outputting the posterior probability vector of the phonemes for the current frame.
[0074] The phoneme set consists of basic Mandarin pronunciation units and is fixed after being determined during the system design phase. The output layer dimension is consistent with the size of this set. Frame-level posterior probability vectors are accumulated sequentially by frame, and then multiplicatively scaled using a frame-by-frame glottal excitation confidence sequence. The posterior probabilities of all phonemes in low-confidence frames are proportionally reduced, while the posterior probabilities of high-confidence frames remain unchanged. The scaled sequence is used to calculate the cumulative score of keyword phoneme paths, ensuring that keyword detection scores are primarily driven by acoustically reliable frames. Before system deployment, the acoustic model is adaptively trained or fine-tuned using speech data recorded by target users in a fume hood operating environment to ensure that the model's output posterior probability distribution adapts to the acoustic conditions of the scenario.
[0075] Multiplicative scaling is applied to the posterior probabilities of each frame using a frame-by-frame glottic excitation confidence sequence: the posterior probabilities of all phonemes in each frame are multiplied by the glottic excitation confidence value of that frame, so that the posterior probabilities of low-confidence frames (frames dominated by noise or unvoiced sounds) are generally reduced, while the posterior probabilities of high-confidence frames (frames dominated by glottic excitations) are completely preserved.
[0076] Perform a search for ascending and descending keyword sequences on the frame-level posterior probability sequence after multiplicative scaling: Maintain a list of ascending and descending keywords that contains all possible ascending and descending verbs and their common variant pronunciations (such as "ascend", "descend", "open", "close", "stop", "adjust", "reach", etc. and their combined variants) that may appear in the scenario of the lifting and lowering operation of the fume hood window in Chinese. Search for the phoneme sequence corresponding to the above keywords in the frame-level posterior probability sequence, calculate the cumulative posterior probability score for each keyword, and divide this score by the length of the keyword phoneme sequence (normalization processing) to obtain the reliability weighted detection score for the keyword.
[0077] For the detected candidate ascending and descending verbs, scan and identify the text within a continuous preset number of word positions in front of them (the number of word positions is determined based on the position range where the subject usually appears in the Chinese subject-predicate structure during the system design phase, generally two to three word positions), and query whether there are words belonging to the experimental quantifier list (the experimental quantifier list contains common non-equipment quantifiers in the laboratory operation scenario, such as "temperature", "concentration", "pressure", "volume", "mass", etc.). If a word in the experimental quantifier list is scanned, it is determined that the candidate ascending and descending verb is in the syntactic structure of an experimental statement, and a weight reduction process is applied to its reliability weighted detection score (multiply by a preset penalty coefficient, which is determined through mixed corpus testing during the calibration phase, and the value principle is to be low enough to reduce the mis-trigger rate of experimental statement instructions to an acceptable level, while not being too low to completely shield edge cases); if no word in the experimental quantifier list is scanned or a word in the window component list is scanned (the window component list contains equipment component words such as "window", "pane", "opening", etc.), the original reliability weighted detection score is maintained.
[0078] The comprehensive detection confidence is generated in the following way: Based on the reliability weighted detection score after grammar constraint processing, perform a normalization mapping of it to the value range from zero to one to obtain the comprehensive detection confidence. The candidate ascending and descending verb and its subsequent quantitative modifiers (including proportional expressions, incremental expressions, and endpoint expressions) together form the candidate text for the lifting and lowering instruction, which is passed backward together with the comprehensive detection confidence.
[0079] Further, the verification process of the three-layer verification refers to Figure 3 , specifically:
[0080] Feasibility verification of the movement direction: Read the opening feedback value to determine the current position of the window, extract the movement direction indicated by the ascending and descending verb in the candidate text. If the current position is already at the physical travel upper limit and the instruction direction is to continue increasing the opening, or is already at the travel lower limit and the instruction direction is to continue decreasing the opening, it is determined that the instruction is physically not executable, output a direction verification failure flag, trigger a voice prompt, and terminate the subsequent verification;
[0081] Target position quantification rationality verification: Map proportional position expressions in candidate text to target opening percentages, add a preset step amount to incremental position expressions based on the current opening, and map endpoint position expressions to the upper or lower limit of the travel range. Calculate the displacement between the target position and the current position. If the displacement does not exceed the minimum effective travel limit, the target position and the current position are determined to be substantially the same, and a rationality verification failure flag is output; if it exceeds the limit, a pass flag and the target position value are output.
[0082] Broadband turbulent noise energy and opening consistency verification: Extract the total broadband turbulent noise energy outside the harmonic frequency band in the most recent non-speech segment, compare it with the opening-broadband turbulent noise energy comparison table established by offline calibration, and determine whether the current broadband noise energy falls within the calibration confidence interval corresponding to the sensor reported opening value. If it falls within the interval, output a consistency verification pass mark; if it does not fall within the interval, output an anomaly mark and request operator confirmation.
[0083] Specifically, the feasibility verification of the movement direction involves reading the current opening feedback value output by the window position sensor and determining the relationship between the current position of the window and the upper and lower limits of the physical travel. The upper and lower limits of the travel are determined during system installation and debugging by manually driving the window to the mechanical limit point and recording the sensor value, which is then stored in the system configuration file. The movement direction indicated by rising / falling verbs in the candidate text is extracted ("rise" and "open" correspond to increasing opening, "fall" and "close" correspond to decreasing opening, and "stop" corresponds to maintaining the current state). If the current opening feedback value is at the upper limit of the travel (allowing for deviations within the sensor's measurement error range, which is determined by sensor repeatability testing during calibration) and the command direction is increasing opening, or if the current opening feedback value is at the lower limit of the travel and the command direction is decreasing opening, the output direction verification fails, triggering the speech synthesis to broadcast the prompt, terminating the subsequent two layers of verification, and ending the current command process; otherwise, the output direction verification passes, and the second layer of verification begins.
[0084] Target position quantification rationality verification: Semantic normalization processing is performed on the position quantification expressions in the candidate text. Proportional expressions (such as "half", "one-third", "seventy percent", etc.) are mapped to a target opening percentage value based on the maximum travel of the window; incremental expressions (such as "raise a little", "lower a little", etc.) are based on the current opening feedback value, superimposed with the preset step amount mapped by the corresponding fuzzy adverb (the step amount classification table is determined during the calibration phase by conducting a semantic survey of target users, based on their subjective evaluation average of the displacement amplitude expected by words such as "a little", "some", "slight", etc., divided into three levels: small, medium, and large); endpoint expressions (such as "highest", "fully open", "completely closed", etc.) are directly mapped to the upper or lower limit of the travel. After mapping, the absolute value of the difference between the target position value and the current opening feedback value (i.e., the required displacement) is calculated. The minimum effective travel limit is determined during the system debugging phase based on the starting characteristics and mechanical clearance of the window lifting motor, taking the minimum driving amount that can cause visible displacement of the window as the lower limit. If the required displacement is lower than the minimum effective travel limit, a reasonableness check failure flag is output, triggering the speech synthesis to broadcast the prompt; if it is not lower than the limit, a reasonableness check pass flag and the target position value are output, and the process proceeds to the third layer of verification.
[0085] Verification of consistency between broadband turbulent noise energy and opening: Within the most recently detected inactive speech segment (speech activity detection uses the glottal excitation confidence level processed in frames as an auxiliary basis; if the confidence level is lower than the reliability judgment threshold for several consecutive frames, it is determined to be a non-speech segment), extract frequency components other than harmonic bands from the harmonic-suppressed speech signal, calculate the total energy value of this component within a duration not shorter than a preset stable estimation window (the length of the stable estimation window is determined based on the short-term stability of noise during the calibration stage, and the value must ensure that the statistical standard deviation of broadband turbulent noise energy estimation is lower than the preset proportion of the overall mean within this duration; this proportion is determined by repeated measurement results under different window lengths during the calibration stage), query the "window opening - broadband turbulent noise energy" lookup table established in offline calibration, determine the acoustically reasonable opening range corresponding to this energy value, and judge whether the current opening feedback value reported by the sensor falls within this range. The lookup table is established as follows: With the exhaust fan operating at its actual rated speed, the viewing window is fixed at each calibrated opening position. A sufficient duration of pure noise signal is collected at each position (ideally at least 60 seconds). The mean and standard deviation of the broadband turbulent noise energy at each position are calculated. The mean plus or minus twice the standard deviation is used as the energy confidence interval for each position. After confirming that the confidence intervals of adjacent calibrated positions do not overlap (if there is overlap, the density of calibrated positions is increased until there is no overlap), the data is stored in the lookup table. If the current energy value falls within the confidence interval corresponding to the current opening, a consistency verification pass flag is output; otherwise, an anomaly flag is output, triggering a voice synthesis broadcast requesting operator confirmation.
[0086] The broadband turbulent noise energy and opening consistency verification is performed only when the window is stationary. The window is determined to be stationary if, within the most recent stable estimation window duration, the change in the opening feedback value is less than the measurement error range determined by the sensor repeatability test. If the window is in motion (the opening feedback value is constantly changing), the third-level verification is skipped, a conditional pass flag is directly output, and in subsequent graded safety execution decisions, the current comprehensive detection confidence interval is lowered by one level: instructions originally in the high confidence interval are downgraded to the waiting confirmation process corresponding to the medium confidence interval, and instructions originally in the medium confidence interval are downgraded to the request retransmission process corresponding to the low confidence interval. This maintains a necessary safety margin when acoustic verification cannot be performed, preventing misjudgments due to temporary inconsistencies between the acoustic and sensor states caused by window movement.
[0087] Furthermore, the method for obtaining graded security execution decisions is as follows:
[0088] The three-layer verification results and the comprehensive detection confidence are input, and the execution conditions are judged according to the serial logic: when any layer verification output fails or is marked as abnormal, the corresponding voice prompt is triggered and the current execution process is terminated without outputting any driving instructions;
[0089] After all three layers of verification pass, a tiered response is executed based on the overall detection confidence interval: when the confidence level is in the high confidence interval, the target position value is directly output to the motion control, driving the window to move to the target position; when the confidence level is in the medium confidence interval, the speech synthesis module broadcasts the parsed target action description, waits for the operator to press a button for confirmation, and then outputs the driving command; when the confidence level is in the low confidence interval, the current candidate text is discarded and the operator is asked to reissue the command through a voice prompt.
[0090] A closed loop is formed based on the target position value and the opening feedback value. After the window reaches the target position, it stops moving. The speech synthesis announces the execution completion status, thus completing the complete semantic recognition to safe execution closed loop.
[0091] Specifically, the three-layer verification results and the comprehensive detection confidence are fed into the execution decision module and processed in a serial logic: the first, second and third layer verification marks are checked in sequence. If any layer outputs a failure or an abnormal mark, the decision process is stopped, the corresponding prompt text is output to the speech synthesis, and no driving command output is generated.
[0092] The high confidence threshold is determined during the system calibration phase using the following method: At least one hundred real control command speech samples and at least one hundred experimental statement speech samples are collected. The overall detection confidence level is calculated for all samples. The 25th percentile of the confidence level in the real control command samples is used as a candidate high confidence threshold, and the 75th percentile of the confidence level in the experimental statement samples is used as a candidate low confidence threshold. If the candidate high confidence threshold is higher than the candidate low confidence threshold (i.e., there is a clear separation in the confidence level distributions of the two types of samples), both are directly used as the high and low confidence thresholds. If there is overlap, the intersection of the confidence level distributions of the two types of samples is used as the low confidence threshold, and the value at a preset range above the low confidence threshold (determined based on an engineering trade-off between acceptable false trigger rate and false negative rate) is used as the high confidence threshold. These two thresholds divide the overall detection confidence level into three intervals: above the high confidence threshold is the high confidence interval, between the two thresholds is the medium confidence interval, and below the low confidence threshold is the low confidence interval.
[0093] After all three layers of verification pass, the comprehensive detection confidence interval is read: when it is in the high confidence interval, the target position value output by the target position quantization rationality verification is directly transmitted to the motion module, driving the window to move towards the target position; when it is in the medium confidence interval, the speech synthesis module converts the target position value into a natural language description and broadcasts it, waiting for the operator to press the button to confirm the signal, and only after receiving the confirmation signal will the target position value be transmitted to the motion module; when it is in the low confidence interval, the current rise and fall command candidate text is discarded, and the speech synthesis module broadcasts a prompt requesting a re-issue of the command.
[0094] After receiving the target position value, the motion module uses the opening feedback value as the real-time feedback quantity and continuously transmits the deviation between the current opening and the target position value to the motor driver, driving the window to move towards the target position until the deviation is lower than the sensor measurement error range (this range is determined by the sensor repeatability test during the calibration phase). Once the target position is determined, the driving stops, and the voice synthesis announces the completion status.
[0095] Example 2:
[0096] Based on Example 1, this embodiment adds a feedforward pre-compensation branch in the harmonic suppression stage, cross-frame consistency verification in the reliability mask construction stage, and a feedback dynamic correction mechanism in the execution decision stage. These are explained below.
[0097] During the window motion transition phase, feedforward pre-compensation is performed on the notch center frequency using the drive direction signal of the motion module:
[0098] While the motion module sends a drive command to the motor driver, the motion direction information of the command is synchronously transmitted to the harmonic suppression processing module. The harmonic suppression processing module determines the expected offset direction of the blade passing frequency corresponding to the current motion direction based on the preset speed-load response characteristics. Before the speed feedback signal reflects the actual speed change, the expected offset direction is applied to the current notch center frequency.
[0099] During the system debugging phase, the speed-load response characteristics are recorded through multiple window raising and lowering operations. The time delay and corresponding frequency offset between the moment the drive command is issued and the moment the speed feedback signal detects the corresponding speed change are recorded. The average value of multiple measurements is used as the reference value of the feedforward offset and stored in the system configuration file. When the difference between the actual instantaneous speed of the speed feedback signal and the feedforward predicted value exceeds the speed fluctuation amplitude judgment limit, the notch center frequency is corrected with the actual feedback value, and the update mode is restored to feedback-dominated mode.
[0100] Specifically, when the viewing window rises, the cross-sectional area of the air inlet decreases, the airflow velocity inside the cabinet increases, the load on the exhaust fan increases, the speed drops briefly, and the blade passing frequency shifts towards lower frequencies; when the viewing window falls, the cross-sectional area increases, the airflow velocity decreases, the load on the exhaust fan decreases, the speed rises briefly, and the blade passing frequency shifts towards higher frequencies. The direction of these physical relationships can be predicted given that the characteristics of the exhaust fan driver are determined.
[0101] During the system debugging phase, the drive window was executed at least five complete upward and five complete downward strokes, respectively. The time interval between the moment the drive command was issued and the moment the Hall sensor detected the start of the corresponding speed change was recorded for each operation, as well as the maximum deviation of the blade passing frequency relative to the stable value before the movement within that time interval. The average time interval and the average frequency deviation were calculated for both upward and downward operations. The average frequency deviation corresponding to the upward operation was defined as the upward feedforward offset (negative sign, indicating pre-adjustment towards lower frequencies), and the average frequency deviation corresponding to the downward operation was defined as the downward feedforward offset (positive sign, indicating pre-adjustment towards higher frequencies). The average time intervals for the upward and downward operations were used as the feedforward durations for the upward and downward directions, respectively.
[0102] For cases where the window performs short-distance movement (the target displacement is less than the preset proportion of the total stroke), the actual wind resistance change is less than the change during the full stroke. The feedforward offset is used after being linearly reduced according to the ratio of the target displacement to the total stroke. This processing is derived based on the inverse relationship between the air intake cross-sectional area and the airflow velocity, without introducing additional calibration parameters.
[0103] During system operation, the motion module sends drive commands to the motor driver while simultaneously transmitting motion direction information to the harmonic suppression processing module. Upon receiving this information, the harmonic suppression processing module immediately corrects the center frequency of all notch filters by the corresponding feedforward offset. The feedforward correction continues for the duration of the corresponding feedforward duration from the moment the drive command is issued, and then switches to the conventional update mode dominated by sensor feedback. The switching is time-based, ensuring that feedforward protection covers the complete period of sensor response delay, while avoiding misjudgment of the switching timing due to unstable frequency difference judgment.
[0104] By introducing the motion drive direction signal as feedforward, the center frequency of the notch is actively pre-compensated during the physical sampling delay of the Hall sensor, eliminating the harmonic suppression hysteresis window during the transition phase of the window motion. This ensures that harmonic suppression remains effective throughout the entire window motion process, preventing residual harmonics from interfering with the voice signal during the transition phase, and further improving the recognition stability of lifting commands in motion.
[0105] Furthermore, when constructing the time-frequency reliability mask, a temporal consistency verification based on forward neighboring frames is introduced for the single-frame glottal excitation confidence:
[0106] The pitch period delay corresponding to the glottal excitation autocorrelation peak of the current frame is compared with the pitch period delay of the previous frame and the two previous frames to calculate the deviation of the pitch period delay between forward adjacent frames. During the calibration stage, the maximum value of the pitch period delay deviation between forward adjacent voiced frames when the operator reads the calibration words is used as the pitch period stability judgment limit.
[0107] For the current frame, the original reliability weight calculated by cross-validation of glottal excitation confidence and residual energy vector is maintained only if the following three conditions are met simultaneously: First, the glottal excitation confidence of the current frame exceeds the reliability judgment threshold; Second, the pitch period delay deviation between the current frame and the previous frame, and between the previous two frames, is lower than the pitch period stability judgment threshold; Third, the glottal excitation confidence of the previous frame and between the previous two frames exceeds the lowest coherence confidence threshold determined by the 90th percentile of the noise frame confidence distribution during the calibration stage.
[0108] If any of the above three conditions are not met, the glottal excitation confidence of the current frame will be reset to zero, the reliability weight of each Mel band in the frame will be reduced to zero accordingly, and the corresponding features will be filled with the long-term average of historical frames.
[0109] Specifically, during the system calibration phase, operators are instructed to read aloud words containing rising and falling vowels at a normal speaking speed, and at least thirty consecutive voiced segments are collected. The absolute deviation of the pitch period delay between adjacent frames within each voiced segment is statistically analyzed, and the 95th percentile of the deviation values of all adjacent frames is taken as the threshold for determining pitch period stability.
[0110] During the calibration phase, pure noise signals under normal operating conditions of the fume hood without personnel activity are collected for no less than two minutes. The glottal excitation confidence level is calculated for all frames, and the 90th percentile value of the distribution is taken as the lowest coherent confidence limit.
[0111] During runtime, the following three conditions are simultaneously verified for the current frame: the glottal excitation confidence of the current frame exceeds the reliability threshold; the deviation of the pitch period delay of the current frame from the previous frame and the two frames before that is lower than the pitch period stability threshold; and the glottal excitation confidence of the previous frame and the two frames before that exceeds the minimum coherence confidence threshold. If all three conditions are met, the original reliability weights are maintained; if any condition is not met, the glottal excitation confidence of the current frame is reset to zero, the reliability weights of each Mel band are reduced to zero, and the corresponding features are filled with the long-term average of historical frames.
[0112] The above verification only uses data from the preceding adjacent frames and does not depend on subsequent frames that have not yet arrived, which meets the unidirectional causality requirement of streaming processing. The processing delay introduced is only two frames in length, which is within the acceptable range of the overall system delay.
[0113] By verifying cross-frame temporal consistency, isolated high-confidence frames (sources of noise impulses) are effectively distinguished from multiple consecutive high-confidence frames (sources of real speech), reducing the contamination of the reliability mask by fume hood pneumatic impulse noise. This allows the time-frequency reliability mask to more accurately reflect the distribution of real speech frames, improving the acoustic quality of the subsequent reliability-weighted feature matrix and the noise impulse resistance of keyword detection.
[0114] Furthermore, a dynamic correction mechanism based on execution result feedback is introduced for the comprehensive detection confidence level:
[0115] After each window movement (ascent or descent) is completed, the system enters a short feedback confirmation window period. The duration of this window period is determined during the system design phase based on the typical reaction time required for the operator to provide voice confirmation of the execution result and is stored in the system configuration file. If a new instruction with the same target position value as the completed position value is detected during the window period, the operator is deemed satisfied, no new drive output is generated, and the confidence sample is added to the calibration sample library. If a new instruction with a different target position value is detected, the system directly enters the three-layer verification process. If no new instruction is detected during the window period, the system ends with an empty window as the default satisfaction signal.
[0116] The dynamic correction mechanism uses accumulated confidence samples to periodically update the high-confidence and low-confidence limits. Once a sufficient number of samples is accumulated as determined in the system design phase, the limits are recalculated, enabling the decision limits for graded security execution to continuously adapt to changes in actual usage conditions.
[0117] Specifically, the window duration is obtained as follows: During the system debugging phase, record the time interval between the operator's voice response to the execution result in no less than twenty operations. Take the smaller value between the 85th percentile of all recorded values and the maximum allowable window duration determined in the system design phase based on the operation safety specifications, and use it as the feedback confirmation window duration, which is then stored in the system configuration file.
[0118] During the window period, the system continuously executes the keyword detection process, comparing the detected candidate instruction target position value with the position value of the currently executed instruction. If they are the same, it is considered a satisfactory confirmation, and the comprehensive detection confidence score corresponding to the candidate text of the up / down instruction that triggered the current execution decision is included as a positive sample in the calibration sample library. If they are different, it is considered a correction intention, and the new instruction directly enters the three-layer verification process. If no new instruction is detected at the end of the window period, the end of the window is considered a default satisfactory signal, and the aforementioned comprehensive detection confidence score is also included in the positive sample library.
[0119] Negative sample collection is explicitly limited to the following: under normal operating conditions, only after the feedback confirmation window ends and before the next movement is completed, the comprehensive detection confidence scores corresponding to candidate words that have undergone grammatical constraint weight reduction in the keyword detection process are collected. All candidate words within the window period are not included in the negative sample library, ensuring that the timing of positive and negative sample collection is independent of each other.
[0120] When both the number of positive and negative samples reach the minimum recalibrated sample size determined during the system design phase, a boundary recalculation is triggered: the high-confidence and low-confidence boundaries are redefined based on the relationship between the 25th percentile of the positive sample confidence distribution and the 75th percentile of the negative sample confidence distribution, the system configuration file is updated, and subsequent decision-making adopts the updated boundary values. After each update, the sample database is cleared, and the accumulation of samples required for the next round of updates begins anew.
[0121] By using the execution result feedback window and the dynamic update mechanism of the boundary, the system's hierarchical safety execution decision-making boundary can be continuously and adaptively updated according to the actual use environment (changes in operators' vocal habits, noise characteristic drift caused by equipment aging, etc.). This avoids the problem of increased misjudgment rate due to changes in environmental conditions during long-term use of fixed calibration boundaries. At the same time, positive samples are implicitly collected by using execution result feedback, eliminating the need for operators to actively perform recalibration operations.
[0122] In summary, this invention achieves high-precision extraction of lifting control commands in the context of strong noise in fume hoods by deeply coupling three layers of logic verification: reliability modeling at the feature level, syntactic constraints at the semantic level, and physical level. By systematically suppressing noise-induced false triggering, it enhances the anti-interference capability of the voice control system and the inherent safety of equipment operation.
[0123] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A semantic-enhanced fume hood sash lift voice analysis system based on, characterized by, include: The harmonic notch suppression module determines the blade passing frequency and harmonic frequencies in real time based on the speed feedback signal, performs harmonic notch suppression on the collected voice signal, and obtains the harmonic suppressed voice signal and the residual energy vector of each harmonic frequency band. The time-frequency feature extraction module calculates the glottal excitation autocorrelation coefficient frame by frame for the harmonic-suppressed speech signal, cross-verifies the glottal excitation confidence with the residual energy vector, constructs a time-frequency reliability mask, and extracts acoustic features by weighting the Mel band to obtain a reliability weighted feature matrix and a frame-by-frame glottal excitation confidence sequence. The steps for determining the reliability weights for each Mel frequency band are as follows: the overall glottal excitation confidence of the current frame is applied to all frequency bands; the residual harmonic energy ratio corresponding to the frequency range of the residual energy vector is used to allocate frequency band reliability weights based on the residual harmonic energy ratio; and the reliability weights for each Mel frequency band in the current frame are obtained based on the glottal excitation confidence and the residual harmonic energy ratio. The semantic enhancement instruction detection module explicitly weights the frame-level acoustic posterior probability with the glottal excitation confidence sequence, suppresses the contribution of low-confidence frames to the detection score of rising and falling keywords, and applies subject grammatical constraints to candidate rising and falling verbs to distinguish between rising and falling instructions and experimental statements, thereby obtaining candidate texts of rising and falling instructions and comprehensive detection confidence. The safety execution decision module sequentially performs motion direction feasibility verification, target position quantification rationality verification, and broadband turbulent noise energy and opening feedback value consistency verification on candidate texts. The three-layer verification results, together with the comprehensive detection confidence, drive the hierarchical safety execution decision.
2. A semantic-enhanced fume hood sash lift voice analysis system according to claim 1, wherein, The process of obtaining the harmonic-suppressed speech signal and the residual energy vector of each harmonic frequency band is as follows: the rotational speed feedback signal is collected and converted into an instantaneous rotational speed value. The fundamental frequency of the current blade is determined based on the number of blades and the instantaneous rotational speed, and the harmonic frequency values of each order are calculated in sequence to form a set of harmonic frequencies. Taking each frequency in the set of harmonic frequencies as the center, the suppression bandwidth of each notch unit is determined by combining the short-time fluctuation amplitude of the rotational speed signal within the time period. The notch units are cascaded and applied to the speech signal to output the harmonic-suppressed speech signal. After notch suppression is completed, the signal energy before and after suppression in the neighborhood of each harmonic center frequency is counted to obtain the degree of harmonic suppression completion in each frequency band, and the residual energy vector is formed by arranging them according to the harmonic order.
3. The voice analysis system for lifting and lowering a fume hood window based on semantic enhancement according to claim 2, characterized in that, The process of constructing the time-frequency reliability mask is as follows: the harmonic-suppressed speech signal is divided into frames according to the preset frame length and frame shift. For each frame signal, the autocorrelation peak is searched in the delay interval within the corresponding Chinese pronunciation fundamental frequency range. The glottal excitation confidence of the current frame is defined based on the peak amplitude and the zero-delay autocorrelation amplitude. The reliability weights are determined by combining two factors for each Mel frequency band, and then arranged in frame order and frequency band order to form a time-frequency reliability mask.
4. The voice analysis system for lifting and lowering a fume hood window based on semantic enhancement according to claim 3, characterized in that, The process of obtaining the reliability weighted feature matrix and the frame-by-frame glottal excitation confidence sequence is as follows: the filter output energy of the harmonic-suppressed speech signal is calculated in each Mel frequency band according to the framing results, and the frequency band energy output is scaled by the reliability weight of the corresponding frame and frequency band in the time-frequency reliability mask. For frequency bands with a reliability weight lower than the preset reliability judgment threshold, the long-term average value of historical frame energy accumulated after the frequency band is started is used to replace the current frame observation value; for frequency bands with a reliability weight not lower than the judgment threshold, the current frame's true observed energy value is retained. Logarithms are taken for the energy values after scaling and / or substitution of each frequency band, and discrete cosine transform is performed to extract cepstral coefficients. These coefficients are then arranged in frame order to form a reliability-weighted feature matrix. Simultaneously, the ratio of the peak value of glottal excitation autocorrelation to the zero-delay autocorrelation value in each frame is arranged in frame order to form a frame-by-frame glottal excitation confidence sequence.
5. The voice analysis system for lifting and lowering a fume hood window based on semantic enhancement according to claim 4, characterized in that, The process of obtaining candidate texts for rising and falling commands and the comprehensive detection confidence is as follows: taking the reliability-weighted feature matrix as input, the posterior probability of each phoneme is calculated through the acoustic model, and the posterior probability of the current frame is multiplicatively scaled by the confidence value corresponding to each frame in the frame-by-frame glottal excitation confidence sequence; when performing rising and falling keyword sequence search on the posterior probability sequence, the total keyword score is driven by high-confidence frames, and the impact of random activation of low-confidence frames on the total score is systematically suppressed, forming a reliability-weighted keyword detection score; For the detected candidate ascending / descending verbs, search for subject nouns within the preset word position range: if a subject noun belonging to the preset experimental quantifier vocabulary exists, the verb is determined to be in the experimental statement syntax structure, and a penalty reduction is applied to the instruction confidence; if no subject noun exists and / or the subject belongs to the window component vocabulary, the original detection confidence is maintained, and it is confirmed as an ascending / descending instruction structure. The reliability-weighted keyword detection score is combined with the grammatical constraint verification result to form a comprehensive detection confidence score, which together with the corresponding candidate ascending / descending verbs and quantification modifiers constitute the candidate text for ascending / descending instructions.
6. The voice analysis system for lifting and lowering a fume hood window based on semantic enhancement according to claim 5, characterized in that, The verification process of the three-layer check is as follows: Motion direction feasibility verification: Read the opening feedback value to determine the current position of the window, extract the motion direction indicated by the rising and falling verbs in the candidate text, and if the current position is already at the upper limit of the physical travel and the instruction direction is to continue to increase the opening, and / or is already at the lower limit of the travel and the instruction direction is to continue to decrease the opening, the instruction is determined to be unexecutable, the direction verification fails to pass the mark is output, the voice prompt is triggered and the subsequent verification is terminated. Target position quantization rationality verification: Map proportional positions in candidate text to target opening percentages, add a preset step amount to incremental positions based on the current opening, and map endpoint positions to the upper and / or lower limits of the travel range. Calculate the displacement between the target position and the current position. If the displacement does not exceed the minimum effective travel limit, the target position and the current position are determined to be substantially the same, and a rationality verification failure flag is output; otherwise, a pass flag and the target position value are output. Broadband turbulent noise energy and opening consistency verification: Extract the total broadband turbulent noise energy outside the harmonic frequency band in the non-speech segment, compare it with the opening-broadband turbulent noise energy comparison table established by offline calibration, and determine whether the current broadband noise energy falls into the calibration confidence interval corresponding to the sensor reported opening value. If it falls into the interval, output a consistency verification pass mark; otherwise, output an abnormal mark.
7. A voice analysis system for lifting and lowering a fume hood window based on semantic enhancement according to claim 6, characterized in that, The method for obtaining graded safety execution decisions is as follows: combining the results of three-layer verification with the overall detection confidence level, and judging the execution conditions according to serial logic: when any layer verification output fails and / or an anomaly mark is displayed, the corresponding voice prompt is triggered and the current execution process is terminated, and no drive commands are output to the motion lifting. After all three layers of verification pass, a graded response is executed based on the comprehensive detection confidence interval: when the confidence is in the high confidence interval, the target position value is output to the motion lift, driving the window to move to the target position; when the confidence is in the medium confidence interval, the speech synthesis module broadcasts the parsed target action description, and waits for the operator to press the button to confirm before outputting the direction motion lift drive command. When the confidence level is in the low confidence interval, discard the candidate text and request the operator to reissue the command via voice prompt; The motion lifting and lowering forms a closed loop based on the target position value and the opening feedback value. After the window reaches the target position, the motion stops, and the speech synthesis broadcasts the execution completion status, completing the complete semantic recognition to safe execution closed loop.