Intelligent manufacturing workshop safety interaction control method and system based on large model

By combining energy concentration and adaptive Kalman filtering, the Fourier transform window length is dynamically adjusted, solving the problem of difficulty in balancing noise suppression and voice fidelity in traditional methods, and realizing more efficient voice-safe interactive control in smart manufacturing workshops.

CN122024731AActive Publication Date: 2026-05-12SHANDONG BLUEBIRD IND INTERNET CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG BLUEBIRD IND INTERNET CO LTD
Filing Date
2026-04-13
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Traditional short-time Fourier transform cannot simultaneously achieve noise suppression and voice fidelity in smart manufacturing workshops, resulting in insufficient signal clarity and accuracy, which affects production safety and efficiency.

Method used

A method combining energy concentration and adaptive Kalman filtering is adopted to dynamically adjust the Fourier transform window length. Local impulsive noise is identified by energy concentration, and noise smoothing and impulse suppression are performed by combining adaptive Kalman filtering. The window length is dynamically adjusted to adapt to different noise environments.

Benefits of technology

It achieves the avoidance of noise trailing and speech feature masking in impact noise scenarios, improves frequency resolution in stable speech scenarios, provides clearer and more stable speech signal input, and improves the reliability and accuracy of interactive control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024731A_ABST
    Figure CN122024731A_ABST
Patent Text Reader

Abstract

The invention relates to the field of industrial control, in particular to an intelligent manufacturing workshop safety interaction control method and system based on a large model, and the method comprises the steps: collecting a sound signal of a workshop environment, calculating the instantaneous energy of a sliding window, constructing an energy concentration ratio, and discriminating local impact noise; introducing adaptive Kalman filtering by taking an energy concentration ratio as a prior factor, tracking a smooth energy change rate, calculating a window length parameter, and adaptively determining a dynamic window length; and after short-time Fourier transform noise reduction, multiplexing or reconstructing window length according to frame similarity, and inputting pure voice into the large voice recognition model to realize safe interaction control. According to the invention, the problems of noise trailing, insufficient frequency resolution and voice distortion easily caused by a traditional fixed window length are solved, the equipment impact noise and the artificial voice instruction can be effectively distinguished, the signal processing precision and the anti-interference capability are improved, and the method is suitable for high-reliability voice safety interaction in a complex industrial environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial control, and more particularly to a method and system for safe interactive control in intelligent manufacturing workshops based on large-scale models. Background Technology

[0002] The workshop production model centered on digitalization, networking, and intelligence is gradually replacing the traditional model. Human-machine collaboration, equipment interconnection, and intelligent scheduling have become typical characteristics of intelligent manufacturing workshops. Workshop safety interaction, as a guarantee for the stable operation of intelligent manufacturing workshops, rapidly transmits production instructions to key links. The accuracy, real-time nature, and anti-interference capability of this interaction directly affect workshop production safety, operational efficiency, and the level of intelligence.

[0003] However, smart manufacturing workshops are filled with various noise sources, including equipment operation, machining, and airflow. Voice interaction signals are easily interfered with by background noise, leading to distortion and blurred features. This can cause errors in production instruction transmission, affecting workshop efficiency, or even equipment malfunctions and human-machine collaboration accidents, threatening workshop safety. Existing technologies often use Fourier transform to process voice signals to smooth noise. Short-time Fourier transform (SFT) is one of the most widely used signal processing methods. It achieves time-frequency domain analysis and noise filtering by dividing the voice signal into multiple short-time frames and performing Fourier transforms on each frame.

[0004] However, the traditional short-time Fourier transform uses a fixed window length for signal framing, which presents an irreconcilable technical contradiction and has become the core bottleneck for voice safety interaction in smart manufacturing workshops: although a long window can improve frequency resolution and accurately identify speech frequency features, it can also cause instantaneous noise energy to spread to adjacent speech frames due to the time-domain smoothing effect, forming a trailing phenomenon that masks speech formant features and contaminates the effective speech signal; although a short window can accurately locate instantaneous noise and reduce energy spread across frames, it can significantly reduce frequency resolution, making it impossible to distinguish between the speech fundamental frequency and background noise, and easily losing the core frequency features of the speech, resulting in speech distortion after denoising. Neither of these can meet the requirements of voice safety interaction in workshops for signal clarity and accuracy. Summary of the Invention

[0005] To address the problem that in short-time Fourier transform with a fixed window length, long windows are prone to noise trailing and masking speech features due to temporal smoothing effects, while short windows reduce frequency resolution and cause speech distortion after denoising, neither of which can meet the requirements for signal clarity and accuracy in safe voice interaction in workshops, this invention provides solutions in the following aspects.

[0006] In the first aspect, the intelligent manufacturing workshop safety interactive control method based on a large model includes: collecting sound signals from the intelligent manufacturing workshop; defining a sliding window for signal analysis and calculating the instantaneous energy of each sliding window; constructing the energy concentration of the current frame sound signal based on the instantaneous energy of all sliding windows; determining whether there is local impulsive noise in the current frame sound signal through the energy concentration to obtain a preliminary judgment result; combining the preliminary judgment result with the energy concentration as a priori feedback factor introduced into an adaptive Kalman filter; using the difference in instantaneous energy between adjacent windows as the energy change rate of the sliding window; performing noise smoothing and impact suppression tracking processing based on the energy change rate to obtain the optimal window after removing interference. The estimated energy change rate is fused with the energy concentration to obtain the Fourier transform window length adjustment parameter, which filters out noise misjudgments caused by operators issuing loud commands. Based on the window length adjustment parameter, the dynamic window length of the current frame is calculated between the preset maximum and minimum window lengths. The dynamic window length is used to perform Fourier transform on the sound signal and perform noise reduction processing. The denoised sound signal is transmitted as a safety interaction input signal to the backend speech recognition model. The similarity evaluation value of the temporal features of the current frame sound signal and the next frame sound signal is calculated. Based on the similarity evaluation value, it is determined whether to reuse or reconstruct the dynamic window length of the next frame sound signal to complete the safety interaction control of the intelligent manufacturing workshop.

[0007] Preferably, the steps for calculating the instantaneous energy of each sliding window are as follows: The original sound signals within each sliding window are summed sequentially, and the summation result is divided by the length of the sliding window to obtain the instantaneous energy of the sliding window.

[0008] Preferably, the energy concentration is calculated as follows: Calculate the ratio between the maximum instantaneous energy of all sliding windows and the mean instantaneous energy of all sliding windows. Use an exponential function to perform an exponential decay mapping on the standard deviation of the instantaneous energy of all sliding windows. The product of the ratio and the result of the exponential decay mapping is taken as the energy concentration of the sound signal in the current frame.

[0009] Preferably, the step of determining whether there is local impulsive noise in the current frame audio signal by energy concentration is as follows: If the energy concentration value is greater than or equal to the preset concentration threshold, it is determined that there is local impulsive noise in the current frame audio signal; if the energy concentration value is less than the preset concentration threshold, it is determined that there is no local impulsive noise in the current frame audio signal, and it is a steady-state background noise or a smooth and continuous voice command signal from the operator, thus obtaining a preliminary judgment result.

[0010] Preferably, the window length adjustment parameter is obtained in the following way: The system state vector is constructed by the difference between the instantaneous energy of each window in the current frame and the instantaneous energy of the adjacent windows. The state transition matrix and observation matrix are established based on the energy time-series change characteristics, and the instantaneous energy of the sliding window is used as the observation value. The observation noise covariance is dynamically adjusted based on the energy concentration of the current frame, and the Kalman gain is adaptively adjusted based on the observation noise covariance to complete the tracking and smoothing of the energy change rate, and obtain the optimal energy change rate estimate after removing interference. The ratio between the estimated optimal energy change rate and the mean instantaneous energy of all sliding windows in the current frame is used as the relative mutation rate. The average of the relative mutation rates of all windows is then multiplied by the energy concentration. The result of the multiplication is then subjected to a negative exponential mapping to obtain the window length adjustment parameter.

[0011] Preferably, the dynamic window length is calculated as follows: Multiply the window length adjustment parameter by the preset maximum window length to obtain the maximum window length weighted term; multiply the difference between 1 and the window length adjustment parameter by the preset minimum window length to obtain the minimum window length weighted term; sum the maximum window length weighted term and the minimum window length weighted term, round the sum down, and use the processed result as the dynamic window length corresponding to the current frame audio signal.

[0012] Preferably, the step of determining whether to perform dynamic window length reuse or reconstruction of the next frame's audio signal based on the similarity evaluation value includes: The similarity evaluation value of the waveforms between the current frame audio signal and the next frame audio signal after denoising is calculated. If the similarity evaluation value is greater than or equal to the preset similarity threshold, it is determined that the sound field state is highly similar. The dynamic window length of the current frame is reused as the window length in the Fourier transform of the next frame to perform denoising. Conversely, if the similarity evaluation value is less than the preset similarity threshold, it is determined that the sound field state has changed abruptly. The dynamic window length is reconstructed by performing instantaneous energy calculation, energy meter accuracy construction, window length adjustment parameter quantization, and dynamic window length calculation on the next frame audio signal.

[0013] Secondly, a large-scale model-based intelligent manufacturing workshop safety interactive control system includes a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the aforementioned large-scale model-based intelligent manufacturing workshop safety interactive control method is implemented.

[0014] The present invention has the following effects: 1. This invention determines the window length adjustment parameter by combining energy concentration and adaptive Kalman filtering, thereby achieving adaptive dynamic adjustment of the short-time Fourier transform window length. In impulsive noise scenarios, the window length is automatically reduced to avoid noise trailing and speech feature masking, while in stable speech scenarios, the window length is automatically expanded to improve frequency resolution. This fundamentally overcomes the shortcomings of traditional fixed window lengths that cannot simultaneously achieve noise suppression and speech fidelity.

[0015] 2. This invention introduces energy concentration as a priori factor into adaptive Kalman filtering to track and smooth the rate of energy change. This effectively suppresses electrical noise and sudden equipment impact interference, accurately distinguishes between equipment impact noise and loud voice commands from operators, avoids misjudgments and abnormal window length jitter caused by amplitude increases, provides clearer and more stable input signals for the backend speech recognition model, and improves the reliability of interactive control.

[0016] 3. This invention achieves the reuse or reconstruction of dynamic window length by evaluating the temporal waveform similarity of adjacent frames, making full use of the short-term stability of the acoustic environment in the workshop. When there are no sudden changes in the sound field, the historical optimal window length can be directly reused to reduce redundant calculations. When there are sudden changes in the sound field, the calculation is recalculated to ensure accuracy. This achieves cost-saving of computing power without reducing the noise reduction effect, and is more suitable for the continuous, low-latency voice security interaction requirements of intelligent manufacturing workshops. Attached Figure Description

[0017] Figure 1 This is a flowchart of steps S1-S5 in the intelligent manufacturing workshop safety interactive control method based on a large model according to an embodiment of the present invention.

[0018] Figure 2 This is a structural block diagram of the intelligent manufacturing workshop safety interactive control system based on a large model, according to an embodiment of the present invention. Detailed Implementation

[0019] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.

[0020] Reference Figure 1 The intelligent manufacturing workshop safety interaction control method based on a large model includes steps S1-S5, as follows: S1: Collect sound signals from the smart manufacturing workshop, define the sliding window for signal analysis, and calculate the instantaneous energy of each sliding window.

[0021] Since voice safety interaction in smart manufacturing workshops requires the extraction of effective features from the original sound signal, and the original signal is a continuous time-domain signal that cannot be directly used to distinguish between impact noise and voice commands, it is necessary to first complete signal acquisition and frame window delineation, convert the continuous signal into discrete window-level analysis units, and then extract the first core time-domain feature through instantaneous energy calculation, laying the data foundation for subsequent multi-dimensional feature analysis and noise discrimination.

[0022] MEMS microphone arrays are deployed in the production areas of the smart manufacturing workshop (such as around machine tools, robotic arm operating areas, and material conveyor lines—areas with high-frequency voice interaction). The deployment location and number are determined based on the workshop equipment layout and operating range to acquire sound signals from various areas within the workshop, ensuring that the sound signals include both operator commands and ambient background sound. The acquisition frequency is set to 40kHz, and a continuous acquisition length of 1 second is used as one frame of sound signal to be processed. A fixed-length sliding window is used as the basic unit for signal analysis. In this embodiment, the length of the sliding window is set as follows: The sliding step size is 0.5ms, so that the sliding window can perform segment-by-segment coverage analysis of the sound signal in a continuous and seamless manner, ensuring that no sound signal is missed.

[0023] In intelligent manufacturing workshops for machining, sudden impact noises such as CNC lathe start-up and shutdown, workpiece clamping impacts, and tool changing impacts are the main sources of noise interference in speech recognition. The core characteristics of this type of noise are instantaneous high energy and energy highly concentrated in a local time window, which is significantly different from the steady-state background noise of continuous equipment operation and the voice commands issued smoothly by operators (energy evenly distributed along the time axis). Traditional technology uses fixed-window-length Fourier transform to process signals, without dedicated noise recognition logic, and cannot distinguish between impact noise, steady-state background noise, and effective speech in mixed signals. This leads to either long window lengths causing impact noise energy to spread and cause trailing, masking speech formants, or short window lengths reducing frequency resolution and causing speech distortion, ultimately resulting in low speech recognition accuracy and equipment control malfunctions. Instantaneous energy, as a basic time-domain feature characterizing the strength of sound signals, can intuitively reflect the energy level of the signal within a window and is the core basis for distinguishing between instantaneous high-energy impact noise and steady-state background noise. Therefore, it is necessary to calculate the instantaneous energy of the entire frame of sound signal in units of a defined sliding window.

[0024] Using a defined sliding window as a unit, the instantaneous energy of the entire frame of sound signal is calculated window by window. The values ​​of all sound signals within a single sliding window are summed sequentially, and the summation result is divided by the length of the sliding window to obtain the instantaneous energy of the sliding window.

[0025] Specifically, the instantaneous energy satisfies the following relationship: ; In the formula, Indicates the first Each sliding window corresponds to the instantaneous energy of the sound signal. This represents the length of the sliding window for a single audio signal. Indicates the first The starting signal point of a sliding window, This represents the original sound signals from the smart manufacturing workshop. Indicates the first The original sound signal.

[0026] Based on the instantaneous energy corresponding to all sliding windows of the entire frame of sound signal, an instantaneous energy dataset for each frame of sound signal is generated.

[0027] S2: Construct the energy concentration of the current frame audio signal based on the instantaneous energy of all sliding windows, and use the energy concentration to determine whether there is local impulsive noise in the current frame audio signal, and obtain a preliminary judgment result.

[0028] The instantaneous energy of a single window alone cannot characterize the energy distribution of the entire frame signal. It is also impossible to accurately distinguish between impulse noise, which has a high concentration of energy only in a local time, and speech signals or steady-state background noise with a stable overall energy distribution and no local abrupt changes. Therefore, it is necessary to analyze the energy concentration based on the instantaneous energy dataset of the entire frame, integrate the discrete window energy features into a noise discrimination index for the entire frame signal, and achieve preliminary discrimination of impulse noise through threshold comparison. This provides a screening basis for the accurate calculation of subsequent window length adjustment parameters and reduces the interference of invalid features on subsequent calculations.

[0029] First, the ratio of the maximum to the average instantaneous energy of all sliding windows within the current frame's audio signal is calculated. This ratio reflects the degree of local concentration and amplitude difference of the instantaneous energy. Simultaneously, using the standard deviation of the instantaneous energy of all sliding windows as the independent variable, an exponential decay mapping is performed using an exponential function. This ensures that the mapping result monotonically decreases as the energy fluctuation increases, thereby suppressing interference from drastic energy fluctuations in subsequent discrimination. Multiplying the ratio by the exponential decay mapping result to construct the energy concentration is to comprehensively consider both the degree of local energy concentration and the overall stability of the energy, avoiding misclassification of signals with large overall energy fluctuations but no local concentration as impulse noise based solely on a single ratio, thus improving the accuracy of noise discrimination.

[0030] Multiplying the above ratio by the result of the exponential decay mapping yields the energy concentration of the current frame's audio signal, which is used to comprehensively characterize the degree of concentration and stability of the audio signal energy on the time axis.

[0031] The energy concentration of the current frame's audio signal is compared with a preset concentration threshold to achieve a preliminary determination of the signal type. If the energy concentration is greater than or equal to the preset concentration threshold, it indicates that the energy of the current frame audio signal has obvious accumulation in the local time range and the signal has significant abrupt change characteristics. Based on this, it is determined that there is local impact noise in the current frame audio signal.

[0032] If the energy concentration is less than the preset concentration threshold, it indicates that the energy distribution of the current frame sound signal is relatively uniform and smooth. Based on this, it is determined that there is no local impact noise in the current frame sound signal, and the signal is the steady-state background noise generated by the operation of the equipment or the smooth and continuous voice command signal issued by the operator, thus completing the preliminary judgment of the current frame sound signal.

[0033] For example, the preset concentration threshold is 2.5. The energy concentration of three types of signals, namely steady-state background noise, operator's steady voice commands, and equipment impact noise, is collected and statistically analyzed. The energy concentration of steady-state background noise and steady voice commands is generally lower than 2.5, while the energy concentration of impact noise is generally higher than 2.5.

[0034] After initial discrimination, signal frames containing local impulsive noise and normal signal frames can be quickly screened out, providing clear signal feature guidance for subsequent steps: for frames containing impulsive noise, more accurate window length adjustment parameters need to be calculated to shrink the window length and suppress noise; for normal signal frames, a larger window length can be retained to ensure speech frequency resolution. The initial discrimination step is a key preliminary step for achieving adaptive window length adjustment, which directly determines the targeting of subsequent feature calculation and window length adjustment.

[0035] Through the aforementioned preliminary discrimination steps, it is possible to quickly distinguish whether the current frame's audio signal is localized impulsive noise, steady-state background noise, or a stable voice command signal from the operator. This provides a preliminary discrimination basis for the subsequent adaptive adjustment of the Fourier transform window length. Preliminary discrimination can eliminate obvious impulsive noise interference in advance during signal preprocessing, avoiding misjudging abnormal equipment noises, sudden impacts, and other noises as valid speech. Simultaneously, it preserves steady-state background noise and clear voice commands under normal operating conditions. This not only improves the targeting of signal feature analysis but also lays the foundation for subsequent dynamic window length denoising and accurate recognition by large-scale speech recognition models, effectively improving the reliability and anti-interference capability of workshop safety interactive control.

[0036] S3: Based on the preliminary discrimination results, the energy concentration is introduced as a prior feedback factor into the adaptive Kalman filter. The difference in instantaneous energy between adjacent windows is used as the energy change rate of the sliding window. Based on the energy change rate, noise smoothing and impact suppression tracking processing are performed to obtain the energy change rate estimate of the optimal window after removing interference. The energy change rate estimate is fused with the energy concentration to obtain the Fourier transform window length adjustment parameter to filter out noise misjudgments caused by operators issuing loud instructions.

[0037] To accurately obtain the window length adjustment parameters that characterize the stability of the signal, an adaptive Kalman filter model is constructed based on the instantaneous energy of each sliding window: The instantaneous energy of each sliding window in the current frame, and the difference in instantaneous energy between the sliding window and the adjacent previous window, are used together as state components to construct the system state vector.

[0038] Based on the physical characteristic that the energy of sound signals changes continuously and smoothly in time, a corresponding state transition matrix and observation matrix are established, and the instantaneous energy of the sliding window is used as the observation value of the filtering system.

[0039] Using the energy concentration of the current frame as the basis for dynamic adjustment, the observation noise covariance is constructed and updated. The Kalman gain is then adaptively adjusted using this observation noise covariance to reduce the impact of interference signals such as impulse noise on the observation results, thereby achieving stable tracking of the energy change rate and noise smoothing, and thus obtaining the optimal energy change rate estimate after removing electrical noise and impulse interference.

[0040] The specific steps for constructing and updating the observation noise covariance based on the energy concentration of the current frame are as follows: First, a baseline noise covariance constant is preset for the workshop under a stable, steady-state background noise environment. The baseline noise covariance constant is set to 0.01. Through extensive field measurements and statistical analysis of the intelligent manufacturing workshop under a pure steady-state background noise environment without equipment impact or operator noise, the energy fluctuation range of the microphone acquisition circuit's inherent electrical noise and the environmental background noise is obtained. The baseline covariance constant is used to characterize the basic observation noise level of the system under pure, stable operating conditions without interference or impact. It can provide a stable benchmark for the adaptive adjustment of the subsequent dynamic observation noise covariance, ensuring that reasonable and reliable dynamic amplification can be achieved based on energy concentration when encountering impact noise.

[0041] The dynamic observation noise covariance corresponding to the current window is obtained by multiplying the baseline noise covariance constant with the natural exponential function value with the energy concentration of the current frame as the index. As the energy concentration of the current frame changes in real time, the dynamic observation noise covariance is updated synchronously, so that the dynamic observation noise covariance can be adaptively adjusted with the degree of signal impact, thereby adapting to the complex and ever-changing noise environment of the workshop.

[0042] Furthermore, the estimated optimal energy change rate is compared with the mean instantaneous energy of all sliding windows in the current frame to obtain the normalized relative mutation rate. The average of the relative mutation rates of all windows in the current frame is then calculated, multiplied by the energy concentration, and subjected to negative exponential mapping to normalize the result. The interval is used to obtain the window length adjustment parameters suitable for Fourier transform.

[0043] Specifically, the window length adjustment parameter satisfies the following relationship: ; In the formula, This indicates the parameter for adjusting the Fourier transform window length. This indicates the energy concentration of the sound signal in the current frame. Represented by constants An exponential function with base 0. This indicates the total number of sliding windows used to divide the current frame's audio signal. The first result obtained after adaptive Kalman filtering is... The optimal energy change rate estimate for each window. This represents the average instantaneous energy of all sliding windows within the current frame.

[0044] In the complex noise environment of a workshop, equipment impact noise typically exhibits a dual characteristic of highly concentrated local energy and drastic energy abrupt changes in adjacent windows. In contrast, loud voices from operators only show an overall increase in amplitude, with energy changes in adjacent windows remaining relatively gradual. To accurately distinguish between the two types of signals and achieve adaptive smooth adjustment of the window length, the formula introduces energy concentration to characterize the local impact degree of the signal. Simultaneously, the estimated optimal energy change rate is compared with the instantaneous average energy to obtain a normalized relative mutation rate, eliminating the interference of the overall signal amplitude on the judgment result and avoiding misjudging loud voices from operators as impact noise. By averaging the relative mutation rates of all windows within a frame, the global average mutation degree of the sound signal in the current frame can be reflected, improving parameter stability. Finally, a negative exponential mapping is used to normalize the comprehensive feature values ​​to... The interval allows the window length adjustment parameter to decrease smoothly as the signal impact increases and increase as the signal stabilizes, which is highly compatible with the subsequent dynamic window length weighted calculation rules.

[0045] By adopting the above settings, while ensuring continuous and seamless window length adjustment, it is possible to accurately identify impact noise and filter false triggers caused by increased voice amplitude, while suppressing electrical noise and sudden transient interference in the acquisition circuit, thus significantly improving the robustness and reliability of the window length adjustment parameters.

[0046] S4: Based on the window length adjustment parameter, the dynamic window length of the current frame is calculated between the preset maximum and minimum window lengths. The dynamic window length is used to perform Fourier transform on the audio signal and perform noise reduction processing. The noise-reduced audio signal is then transmitted to the backend speech recognition model as a secure interactive input signal.

[0047] Traditional fixed window lengths cannot simultaneously achieve both time-domain positioning accuracy and frequency-domain resolution. However, window length adjustment parameters can accurately reflect the stability of the current frame signal. Therefore, based on further analysis of the window length adjustment parameters, we can achieve adaptive adjustment of the dynamic window length within a preset range, so that the window length is accurately matched with the signal characteristics. The specific steps are as follows: Using the window length adjustment parameter as a weighting coefficient, the preset maximum and minimum window lengths are calculated separately: the window length adjustment parameter is multiplied by the preset maximum window length to obtain the maximum window length weighting term; the difference between 1 and the window length adjustment parameter is multiplied by the preset minimum window length to obtain the minimum window length weighting term. The maximum and minimum window length weighting terms are summed, and the result is rounded down to the nearest integer. The resulting integer is the dynamic window length corresponding to the current frame's audio signal. This dynamic window length can adaptively adjust within a preset range according to the stability of the audio signal. The more stable the signal, the closer the window length is to the maximum window length; the stronger the signal impact, the closer the window length is to the minimum window length. This balances frequency domain resolution and temporal domain positioning accuracy, adapting to the audio signal analysis needs in complex noise environments like workshops.

[0048] Specifically, the dynamic window length satisfies the following relationship: ; In the formula, This indicates the dynamic window length corresponding to the audio signal of the current frame. This indicates the parameter for adjusting the Fourier transform window length. This represents the preset maximum window length for the Fourier transform. This represents the preset minimum window length for the Fourier transform. This indicates a floor operation. Since the window length adjustment parameter has a range of values ​​within... If there is local noise in the current frame, the window length adjustment parameter is set to a smaller value, and the window length shrinks; if the current frame contains a stable voice command or steady-state background noise, the window length adjustment parameter is set to a larger value, and the window length is longer. Weighted interpolation ensures that the window length is continuous and smooth, without jumps or spikes, avoiding Fourier transform spectrum jitter and distortion, and realizing adaptive continuous change of the window length within a reasonable range. It avoids the window length being too large or too small, and uses the window length adjustment parameter to directly reflect the stability of the signal, making the dynamic window length strongly correlated with the signal characteristics.

[0049] By performing a short-time Fourier transform on the sound signal using a dynamic window, impact noise can be suppressed while retaining effective speech features. After denoising the transformed signal, the clean speech signal is transmitted to the back-end speech recognition model, which can improve the model's recognition accuracy. Ultimately, this enables safe and accurate interactive control in the intelligent manufacturing workshop, completing the entire process from window length adaptation to signal processing and then to safe interaction.

[0050] The weighted interpolation method ensures a smooth and continuous change in window length, free from jumps and glitches, thus avoiding Fourier transform spectrum jitter and distortion. It tends towards a larger window length when the signal is stable to maintain frequency domain resolution, and towards a smaller window length when the signal changes abruptly to improve time domain positioning accuracy. This setting achieves adaptive matching of window length to signal characteristics while effectively suppressing workshop impact noise interference and preserving complete voice features, thereby improving the stability and accuracy of voice interaction control in intelligent manufacturing workshops.

[0051] S5: Calculate the similarity evaluation value of the temporal features between the current frame audio signal and the next frame audio signal. Based on the similarity evaluation value, determine whether to perform dynamic window length reuse or reconstruction on the next frame audio signal to complete the safe interactive control of the intelligent manufacturing workshop.

[0052] To fully utilize the short-term stationary characteristics of the acoustic environment in intelligent manufacturing workshops and reduce redundant computational consumption during continuous signal processing, after denoising the current frame's audio signal, a waveform similarity assessment value is further calculated between the current frame's audio signal and the newly acquired next frame's audio signal. Based on the comparison result of this similarity assessment value and a preset similarity threshold, a dynamic window length reuse or reconstruction decision is executed. If the similarity evaluation value is greater than or equal to the preset similarity threshold, it is determined that the sound field environment and signal features corresponding to the sound signal of the current frame and the next frame are highly similar, and no sudden interference or voice state switching has occurred. The dynamic window length calculated in the current frame is directly reused as the analysis window length of the sound signal in the Fourier transform of the next frame, and denoising processing is directly performed based on the window length. If the similarity evaluation value is less than the preset similarity threshold, it is determined that the sound field environment or signal characteristics have undergone a sudden change, and the original dynamic window length is no longer suitable for the current signal. The complete steps of instantaneous energy calculation, energy concentration construction, window length adjustment parameter quantization, and dynamic window length calculation are re-executed for the next frame of sound signal to complete the reconstruction of the dynamic window length, so as to ensure the noise reduction accuracy and signal analysis reliability.

[0053] This invention also provides a safety interactive control system for intelligent manufacturing workshops based on a large model. For example... Figure 2 As shown, the system includes a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement the intelligent manufacturing workshop safety interactive control method based on a large model according to the first aspect of the present invention. The system also includes other components well known to those skilled in the art, such as a communication bus and a communication interface, the settings and functions of which are known in the art and therefore will not be described in detail here.

[0054] It should be noted that those skilled in the art can make various modifications and improvements without departing from the inventive concept, and these all fall within the scope of protection of this invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A safety interactive control method for intelligent manufacturing workshops based on a large model, characterized in that, include: Collect sound signals from the smart manufacturing workshop, define the sliding window for signal analysis, and calculate the instantaneous energy of each sliding window; The energy concentration of the current frame audio signal is constructed based on the instantaneous energy of all sliding windows. The presence of local impulsive noise in the current frame audio signal is determined by the energy concentration, and a preliminary judgment result is obtained. Based on the preliminary judgment results, the energy concentration is introduced as a prior feedback factor into the adaptive Kalman filter. The difference in instantaneous energy between adjacent windows is used as the energy change rate of the sliding window. Based on the energy change rate, noise smoothing and impact suppression tracking processing are performed to obtain the energy change rate estimate of the optimal window after removing interference. The energy change rate estimate is fused with the energy concentration to obtain the Fourier transform window length adjustment parameter to filter out noise misjudgments caused by operators giving loud instructions. The dynamic window length of the current frame is calculated between the preset maximum and minimum window lengths based on the window length adjustment parameters. The dynamic window length is used to perform Fourier transform on the audio signal and perform noise reduction. The noise-reduced audio signal is then transmitted to the backend speech recognition model as a secure interactive input signal. Calculate the similarity evaluation value of the temporal features between the current frame audio signal and the next frame audio signal. Based on the similarity evaluation value, determine whether to perform dynamic window length reuse or reconstruction on the next frame audio signal to complete the safe interactive control of the intelligent manufacturing workshop.

2. The intelligent manufacturing workshop safety interactive control method based on a large model according to claim 1, characterized in that, The steps for calculating the instantaneous energy of each sliding window are as follows: The original sound signals within each sliding window are summed sequentially, and the summation result is divided by the length of the sliding window to obtain the instantaneous energy of the sliding window.

3. The intelligent manufacturing workshop safety interactive control method based on a large model according to claim 1, characterized in that, The energy concentration is calculated as follows: Calculate the ratio between the maximum instantaneous energy of all sliding windows and the mean instantaneous energy of all sliding windows. Use an exponential function to perform an exponential decay mapping on the standard deviation of the instantaneous energy of all sliding windows. The product of the ratio and the result of the exponential decay mapping is taken as the energy concentration of the sound signal in the current frame.

4. The intelligent manufacturing workshop safety interactive control method based on a large model according to claim 1, characterized in that, The steps for determining whether there is localized impulsive noise in the current frame's audio signal based on energy concentration are as follows: If the energy concentration value is greater than or equal to the preset concentration threshold, it is determined that there is local impulsive noise in the current frame audio signal; if the energy concentration value is less than the preset concentration threshold, it is determined that there is no local impulsive noise in the current frame audio signal, and it is a steady-state background noise or a smooth and continuous voice command signal from the operator, thus obtaining a preliminary judgment result.

5. The intelligent manufacturing workshop safety interactive control method based on a large model according to claim 1, characterized in that, The window length adjustment parameter is obtained as follows: The system state vector is constructed by the difference between the instantaneous energy of each window in the current frame and the instantaneous energy of the adjacent windows. The state transition matrix and observation matrix are established based on the energy time-series change characteristics, and the instantaneous energy of the sliding window is used as the observation value. The observation noise covariance is dynamically adjusted based on the energy concentration of the current frame, and the Kalman gain is adaptively adjusted based on the observation noise covariance to complete the tracking and smoothing of the energy change rate, and obtain the optimal energy change rate estimate after removing interference. The ratio between the estimated optimal energy change rate and the mean instantaneous energy of all sliding windows in the current frame is used as the relative mutation rate. The average of the relative mutation rates of all windows is then multiplied by the energy concentration. The result of the multiplication is then subjected to a negative exponential mapping to obtain the window length adjustment parameter.

6. The intelligent manufacturing workshop safety interactive control method based on a large model according to claim 1, characterized in that, The dynamic window length is calculated as follows: Multiply the window length adjustment parameter by the preset maximum window length to obtain the maximum window length weighted term; multiply the difference between 1 and the window length adjustment parameter by the preset minimum window length to obtain the minimum window length weighted term. The maximum window length weighted term and the minimum window length weighted term are summed, the sum is rounded down, and the result is used as the dynamic window length corresponding to the audio signal of the current frame.

7. The intelligent manufacturing workshop safety interactive control method based on a large model according to claim 1, characterized in that, Based on the similarity evaluation value, the steps for determining whether to perform dynamic window length reuse or reconstruction on the next frame of audio signal include: The similarity evaluation value of the waveforms between the current frame audio signal and the next frame audio signal after denoising is calculated. If the similarity evaluation value is greater than or equal to the preset similarity threshold, it is determined that the sound field state is highly similar. The dynamic window length of the current frame is reused as the window length in the Fourier transform of the next frame to perform denoising. Conversely, if the similarity evaluation value is less than the preset similarity threshold, it is determined that the sound field state has changed abruptly. The dynamic window length is reconstructed by performing instantaneous energy calculation, energy meter accuracy construction, window length adjustment parameter quantization, and dynamic window length calculation on the next frame audio signal.

8. A safety interactive control system for intelligent manufacturing workshops based on a large model, characterized in that: include: A processor and a memory, wherein the memory stores computer program instructions that, when executed by the processor, implement the intelligent manufacturing workshop safety interactive control method based on a large model according to any one of claims 1-7.