Solvent-free compound machine capable of realizing voice interaction

Through multimodal signal fusion and adaptive learning technology, the problem of speech recognition errors in solvent-free composite machines in complex sound fields is solved, and high-accuracy voice control and process parameter optimization are achieved to meet the needs of intelligent factories.

CN120472899APending Publication Date: 2025-08-12SHANDONG QINGYANG INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510715126.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing voice interaction system of solvent-free composite machines cannot effectively deal with the complex sound field of high-frequency motor noise and low-frequency vibration of the composite machine, resulting in speech control recognition errors, risk of misoperation and low intelligence.

Method used

Multimodal signal fusion technology is adopted to fuse audio and vibration signals through an anti-noise audio acquisition module and a vibration sensor, combined with adaptive learning technology, adapt to the operator's accent and optimize process parameters, including the design of the hardware layer, software layer and data communication layer.

Benefits of technology

It improves the accuracy of voice command recognition, adapts to complex sound fields, adapts to the accent of new operators, optimizes process parameters, and meets the usage needs of intelligent factories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472899A_ABST
    Figure CN120472899A_ABST
Patent Text Reader

Abstract

The invention discloses a solvent-free compound machine capable of voice interaction, which comprises a solvent-free compound machine body, and a voice interaction system is arranged in the solvent-free compound machine body; the voice interaction system comprises a hardware layer, a software layer and a data communication layer, the software layer is connected with the hardware layer and the data communication layer, the data communication layer is connected with the hardware layer, and the data communication layer is connected with a cloud. Multi-mode signal fusion is adopted, audio and vibration are fused, the low-frequency voice signal-to-noise ratio is increased, voice instruction recognition accuracy can be improved, a complex sound field mixed by high-frequency motor noise and low-frequency vibration of a compound machine can be effectively dealt with, the adaptive learning technology is combined, and the recognition accuracy of the compound machine is improved. And accent of a new operator can be adapted and process parameters can be optimized, so that the adaptation range is enlarged, and the use requirements are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of solvent-free laminating machines, and in particular to a solvent-free laminating machine capable of voice interaction. Background Art

[0002] A solvent-free laminating machine is an environmentally friendly laminating equipment used for the production of flexible packaging materials. It is mainly used to bond two or more substrates together through solvent-free adhesives. Compared with traditional solvent-based laminating machines, its biggest feature is that it does not use organic solvents, so it is more environmentally friendly, energy-saving and safe. Solvent-free laminating machines in the existing technology are usually manually operated. The manual operation method has a low degree of intelligence and cannot adapt to the trend of smart factories. Although there are some solvent-free laminating machines for voice interactive control in the existing technology, their voice systems usually rely on a single noise reduction algorithm and cannot cope with the complex sound field of the compounding machine mixed with high-frequency motor noise and low-frequency vibration. As a result, there is a risk of recognition errors during voice control, leading to misoperation. Based on the above situation, this application proposes a solvent-free laminating machine with voice interaction. Summary of the Invention

[0003] Based on the technical problems existing in the background technology, the present invention proposes a solvent-free compounding machine with voice interaction.

[0004] The present invention provides a voice-interactive solventless compounding machine, comprising a solventless compounding machine body, wherein the solventless compounding machine body has a built-in voice interaction system;

[0005] The voice interaction system includes a hardware layer, a software layer and a data communication layer. The software layer is connected to the hardware layer and the data communication layer. The data communication layer is connected to the hardware layer, and the data communication layer is connected to the cloud.

[0006] The hardware layer includes an anti-noise audio acquisition module, a main control calculation module, an execution feedback module and an environment perception module;

[0007] The software layer includes a speech preprocessing module, a speech recognition and command parsing module, an acoustic scene adaptation module, a device control module, and a self-learning optimization module;

[0008] The data communication layer includes a real-time control communication module, a cloud collaboration module and a local data cache module.

[0009] Preferably, the anti-noise audio acquisition module includes four industrial-grade MEMS microphones arranged linearly with a spacing of 8 cm, and the microphones are equipped with silicone shock-proof brackets and low-frequency vibration sensors. The four microphones directionally pick up the operator's voice, suppress the equipment operation noise, and compensate for the low-frequency voice signal through the vibration sensor;

[0010] The main control computing module is used to run noise reduction and speech recognition algorithms and coordinate the work of various modules;

[0011] The execution feedback module uses voice broadcast and light to prompt the command execution status. The voice broadcast uses a speaker with a frequency response range of 100Hz-16kHz. The three-color light is defined as three flashes of red light for "insufficient adhesive" and two flashes of yellow light for "abnormal tension".

[0012] The environmental perception module includes a temperature sensor, a humidity sensor, and a Hall-effect speed sensor, which are used to monitor the device's temperature, humidity, and speed, respectively, providing context for noise suppression. The temperature sensor and humidity sensor are mounted on the motor housing, and the Hall-effect speed sensor is aligned with the main drive shaft. The sensor zero point is automatically calibrated upon daily power-up.

[0013] Preferably, the microphone adopts the TDOA algorithm when performing source positioning, and its formula is:

[0014]

[0015] where Δt ij is the time delay difference between microphones i and j, C is the speed of sound (340 m / s), θ is the incident angle of the sound source, r i is the spatial coordinate of microphone i, d ij is the path difference from the sound source to the two microphones;

[0016] The vibration sensor uses a vibration signal weighted fusion technique to compensate for low-frequency speech signals, and the formula used is:

[0017] S fuscd (f) = w audio (f)·S audio (f)+w vid (f)·S vid (f);

[0018] Among them S vid (f) The weight of the low frequency band (<500Hz) is set to 0.7, and the high frequency band is close to 0. fuscd (f) is the fused spectrum, w audio (f) is the audio signal weight, low frequency is 0.3, high frequency is 1.0, w vid (f) is the vibration signal weight, low frequency 0.7, high frequency 0.0, S audio (f) is the microphone audio spectrum.

[0019] Preferably, the speech preprocessing module is used for real-time noise reduction and enhanced speech clarity, and its specific logical steps are as follows:

[0020] S201: Perform 5-500 Hz bandpass filtering on the vibration signal collected by the vibration sensor to eliminate high-frequency mechanical noise interference. During the filtering process, the vibration energy needs to be calculated using the following formula: Among them E vid is the short-time energy of the vibration signal, x vid [n] is the nth sampling point of the vibration signal, N is the calculation frame length, and its value is 160;

[0021] S202: Audio vibration weighted fusion, the microphone signal is processed in frequency bands, and the low frequency band is calculated using weighted fusion technology: S fuscd (f)=0.7·S vid (f)+0.3S audio (f), the high frequency band uses the microphone signal directly;

[0022] S203: Noise modeling based on the device speed is performed. Pure noise samples are recorded at different speeds when the device is unloaded. The average spectrum is calculated and the reference noise spectrum is obtained by looking up the table based on the speed sensor signal N(t). The formula used is:

[0023]

[0024] in is the current noise power spectrum estimate, N max The maximum permissible speed of the equipment;

[0025] And perform real-time noise update. When a speechless segment is detected, the noise power spectrum is updated. The formula used is:

[0026]

[0027] is the updated noise power spectrum estimate, is the noise power spectrum estimate of the previous frame, f is the frequency, and Y(f) is the power spectrum of the noisy speech in the current frame;

[0028] S204: Perform frequency domain processing, perform FFT on each frame of audio (20ms length, 50% overlap), obtain the spectrum Y(f), and calculate the spectrum after noise reduction:

[0029]

[0030] Where α(f) is the frequency-related over-subtraction factor, which is set to 1.8 in the low frequency band and 1.2 in the high frequency band; β is the spectrum floor coefficient, which is set to 0.002;

[0031] And retain the original phase φ(f), synthesize the time domain signal, the formula used is; in is the time domain signal, IFFT(·) is the inverse fast Fourier transform, φ(f) is the phase spectrum of the original noisy speech, is the phase information expressed in the form of complex exponential, t is time, and f is frequency;

[0032] S205: Automatically adjust the gain according to the output voice amplitude, and enable the transient suppression filter when a burst noise is detected. The formula used to adjust the gain is:

[0033] Where G is the gain coefficient, x ^ [m] is the mth sampling point of the time domain signal after denoising, and M is the calculation frame length.

[0034] Preferably, the speech recognition and instruction parsing module is used to convert speech into text instructions and parse them into device control parameters;

[0035] The acoustic scene adaptation module is used to identify whether the current noise scene is steady or transient, and dynamically adjust the noise reduction strategy. The specific logical steps are as follows:

[0036] S301: Feature extraction: Extract audio features and equipment operating condition features. The audio features are Mel-Frequency Cepstral Coefficients (MFCCs). Extract the first 13 dimensions (including energy terms) to characterize the noise spectrum characteristics and detect transient noise (sudden high-frequency components of valve sounds). The formula used is:

[0037] Where x[n] is the audio sampling point, N is the frame length, n is the index of the current sampling point, ranging from 1 to N, x[n] is the amplitude of the discrete signal at index n, x[n-1] is the amplitude of the discrete signal at index n-1, and sgn(·) is the sign function;

[0038] The equipment operating characteristics include speed value and vibration energy mean value;

[0039] S302: Scene classification is performed using a lightweight neural network. The feature data extracted in S201 is input to the lightweight neural network, and the output is 0, 1, or 2, where 0 represents steady-state noise, 1 represents transient noise, and 2 represents speech-dominated noise. The decision logic is also set. The spectrum of the steady-state noise is stable, the ZCR is low (<0.15), the speed fluctuation is small (ΔN < 50 RPM), and the ZCR of the transient noise increases sharply (>0.3).

[0040] S303: Dynamically adjust the noise reduction strategy: enhance low-frequency suppression for steady-state noise, enable transient filters for transient noise, disable vibration fusion for voice-dominated noise, and protect high-frequency voice.

[0041] S304: Perform real-time calibration: When a classification error is detected (speech is misclassified as transient noise), the frame features are recorded and incremental learning is triggered. The formula used for incremental learning is: Where λ = 0.01, only the weight of the last layer is fine-tuned to avoid catastrophic forgetting. is the total loss function, is the cross entropy loss, θ is the current model parameter, and θ0 is the initial model parameter.

[0042] Preferably, the device control module is used to safely execute instructions. For high-risk instructions (emergency stop), voiceprint matching + PLC status confirmation is required. 20 high-frequency instructions ("pause" and "slowdown") are pre-stored and activated when the network is interrupted.

[0043] The self-learning optimization module is used to adapt to the operator's accent and optimize the process parameter recommendation model, which is a baseline model trained based on a standard process library or historical data provided by the manufacturer;

[0044] The specific logical steps for its accent adaptation optimization are as follows:

[0045] S6011: When a new operator uses it for the first time, the system guides the recording of 10 core commands ("accelerate", "stop", "increase coating amount"). Regarding the environmental requirements, the recording is performed under 85dB background noise to ensure consistency with actual working conditions;

[0046] S6012: Perform feature extraction, extracting Mel-frequency cepstral coefficients and new operator voiceprint features, and perform normalization processing to eliminate volume differences and retain pronunciation characteristics;

[0047] S6013: Perform incremental learning to update the industrial speech recognition model: Calculate the loss function and fine-tune the parameters. The loss function calculation formula is the same as that used in S204. When fine-tuning the parameters, retain the speech feature extraction layer of the pre-trained model and only update the weights of the fully connected classification layer to adapt to the new accent characteristics. The industrial speech recognition model is the initial deployment of the system.

[0048] S6014: Push the new model parameters to the edge device through the encrypted channel;

[0049] The specific logical steps for optimizing the process parameters are as follows:

[0050] S6021: Extract key features from historical data, including process parameters such as speed, temperature, and coating weight, as well as result indicators such as composite strength and appearance defect rate. K-means is used to group similar process parameters and mark high-quality combinations.

[0051] S6022: Filter historical high-quality parameter groups and prioritize them based on the current substrate type and working conditions;

[0052] S6023: If the sensor detects an anomaly, it recommends solution parameters for similar historical scenarios;

[0053] S6024: Record the parameters and production results actually used by the operator, retrain the recommendation model every week, and weight the new data.

[0054] Preferably, the real-time control communication module is used to interact with the PLC, transmit control instructions and status data, read and write PLC registers via Modbus TCP during protocol configuration, and synchronize device status via EtherCAT;

[0055] The cloud collaboration module is used for remote monitoring, model updates, and fault diagnosis. It pushes alarm information via MQTT, transmits encrypted model files via HTTPS, and automatically synchronizes logs that were not uploaded during the interruption after the network is restored.

[0056] The local data cache module is used to store voice logs and device operation data.

[0057] The present invention also proposes a method for using a solvent-free laminating machine capable of voice interaction, comprising the following steps:

[0058] S1: Synchronously collects audio and vibration signals through the microphone array and vibration sensor in the anti-noise audio acquisition module:

[0059] S2: The audio and vibration signals collected in S1 are weighted and fused using a speech preprocessing module to improve the signal-to-noise ratio of low-frequency speech. The collected signals are also preprocessed for noise reduction to suppress ambient noise (motor and valve sounds) and retain pure speech.

[0060] S3: The acoustic scene adaptation module identifies whether the current noise scene is steady or transient and dynamically adjusts the noise reduction strategy;

[0061] S4: The speech recognition and command parsing module converts the pre-processed speech in S2 into text commands, understands the command intent through the text commands, and converts them into device executable parameters;

[0062] S5: Write the device executable parameters in S4 into the PLC register in the device control module to control the device execution. At the same time, the execution feedback module prompts the instruction execution status through voice broadcast and light;

[0063] S6: When a new operator issues a voice command, the self-learning optimization module is triggered to use incremental learning technology to update the industrial speech recognition model and regularly optimize the process parameter recommendation model;

[0064] S7: Stores voice logs and device operation data through the local data cache module and uploads key data to the cloud.

[0065] Compared with the existing technology, the beneficial effects of the present invention are:

[0066] The present invention adopts multimodal signal fusion to fuse audio and vibration, thereby improving the low-frequency voice signal-to-noise ratio, thereby improving the accuracy of voice command recognition, and can effectively cope with the complex sound field of the compound machine where high-frequency motor noise and low-frequency vibration are mixed. In combination with adaptive learning technology, it can adapt to the accent of new operators and optimize process parameters, thereby increasing the adaptation range and meeting usage requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 This is a system block diagram of a voice-interactive solvent-free compounding machine proposed by the present invention;

[0068] Figure 2 This is a flow chart of a solvent-free laminating machine with voice interaction proposed by the present invention. DETAILED DESCRIPTION

[0069] The present invention will be further explained below with reference to specific embodiments.

[0070] Example

[0071] Reference Figure 1-2 ,This embodiment proposes a voice-interactive solventless compounding machine, including a solventless compounding machine body, and the solventless compounding machine body has a built-in voice interaction system;

[0072] The voice interaction system includes a hardware layer, a software layer, and a data communication layer. The software layer is connected to the hardware layer and the data communication layer. The data communication layer is connected to the hardware layer, and the data communication layer is connected to the cloud.

[0073] The hardware layer includes the noise-resistant audio acquisition module, the main control computing module, the execution feedback module and the environmental perception module;

[0074] The anti-noise audio acquisition module includes four industrial-grade MEMS microphones arranged linearly with an 8cm spacing. Silicone shock-absorbing brackets and low-frequency vibration sensors are installed on the microphones. These four microphones directionally pick up the operator's voice, suppressing equipment operating noise, and use the vibration sensor to compensate for low-frequency voice signals.

[0075] The microphone uses the TDOA algorithm for source positioning, and its formula is:

[0076]

[0077] where Δt ij is the time delay difference between microphones i and j, C is the speed of sound (340 m / s), θ is the incident angle of the sound source, ri is the spatial coordinate of microphone i, d ij is the path difference from the sound source to the two microphones;

[0078] The vibration sensor uses vibration signal weighted fusion technology to compensate for low-frequency voice signals. The formula used is:

[0079] S fuscd (f) = w audio (f)·S audio (f)+w vid (f)·S vid (f);

[0080] Among them S vid (f) The weight of the low frequency band (<500Hz) is set to 0.7, and the high frequency band is close to 0. fuscd (f) is the fused spectrum, w audio (f) is the audio signal weight, low frequency is 0.3, high frequency is 1.0, w vid (f) is the vibration signal weight, low frequency 0.7, high frequency 0.0, S audio (f) is the microphone audio spectrum;

[0081] The main control computing module is used to run noise reduction and speech recognition algorithms and coordinate the work of various modules;

[0082] The execution feedback module uses voice broadcast and light to prompt the command execution status. The voice broadcast uses a speaker with a frequency response range of 100Hz-16kHz. The three-color light is defined as three flashes of red light for "insufficient adhesive" and two flashes of yellow light for "abnormal tension".

[0083] The environmental sensing module includes a temperature sensor, a humidity sensor, and a Hall-effect speed sensor, which monitor device temperature, humidity, and speed, respectively, providing context for noise suppression. The temperature and humidity sensors are mounted on the motor housing, while the Hall-effect speed sensor is aligned with the main drive shaft. The sensor zero point is automatically calibrated upon daily power-up.

[0084] The software layer includes a speech preprocessing module, a speech recognition and command parsing module, an acoustic scene adaptation module, a device control module, and a self-learning optimization module;

[0085] The speech preprocessing module is used for real-time noise reduction and enhanced speech clarity. Its specific logical steps are as follows:

[0086] S201: Perform 5-500 Hz bandpass filtering on the vibration signal collected by the vibration sensor to eliminate high-frequency mechanical noise interference. During the filtering process, the vibration energy needs to be calculated using the following formula: Among them E vid is the short-time energy of the vibration signal, xvid [n] is the nth sampling point of the vibration signal, N is the calculation frame length, and its value is 160;

[0087] S202: Audio vibration weighted fusion, the microphone signal is processed in frequency bands, and the low frequency band is calculated using weighted fusion technology: S fuscd (f)=0.7·S vid (f)+0.3S audio (f), the high frequency band uses the microphone signal directly;

[0088] S203: Noise modeling based on the device speed is performed. Pure noise samples are recorded at different speeds when the device is unloaded. The average spectrum is calculated and the reference noise spectrum is obtained by looking up the table based on the speed sensor signal N(t). The formula used is:

[0089]

[0090] in is the current noise power spectrum estimate, N max The maximum permissible speed of the equipment;

[0091] And perform real-time noise update. When a speechless segment is detected, the noise power spectrum is updated. The formula used is:

[0092]

[0093] is the updated noise power spectrum estimate, is the noise power spectrum estimate of the previous frame, f is the frequency, and Y(f) is the power spectrum of the noisy speech in the current frame;

[0094] S204: Perform frequency domain processing, perform FFT on each frame of audio (20ms length, 50% overlap), obtain the spectrum Y(f), and calculate the spectrum after noise reduction:

[0095]

[0096] Where α(f) is the frequency-related over-subtraction factor, which is set to 1.8 in the low frequency band and 1.2 in the high frequency band; β is the spectrum floor coefficient, which is set to 0.002;

[0097] And retain the original phase φ(f), synthesize the time domain signal, the formula used is; in is the time domain signal, IFFT(·) is the inverse fast Fourier transform, φ(f) is the phase spectrum of the original noisy speech, is the phase information expressed in the form of complex exponential, t is time, and f is frequency;

[0098] S205: Automatically adjust the gain according to the output voice amplitude, and enable the transient suppression filter when a burst noise is detected. The formula used to adjust the gain is:

[0099] Where G is the gain coefficient, x ^ [m] is the mth sampling point of the time domain signal after noise reduction, and M is the calculation frame length;

[0100] The speech recognition and command parsing module is used to convert speech into text commands and parse them into device control parameters;

[0101] The acoustic scene adaptation module is used to identify whether the current noise scene is steady-state or transient and dynamically adjust the noise reduction strategy. Its specific logical steps are as follows:

[0102] S301: Feature extraction: Extract audio features and equipment operating condition features. The audio features are Mel-Frequency Cepstral Coefficients (MFCCs). Extract the first 13 dimensions (including energy terms) to characterize the noise spectrum characteristics and detect transient noise (sudden high-frequency components of valve sounds). The formula used is:

[0103] Where x[n] is the audio sampling point, N is the frame length, n is the index of the current sampling point, ranging from 1 to N, x[n] is the amplitude of the discrete signal at index n, x[n-1] is the amplitude of the discrete signal at index n-1, and sgn(·) is the sign function;

[0104] The equipment operating characteristics include speed value and vibration energy mean value;

[0105] S302: Scene classification is performed using a lightweight neural network. The feature data extracted in S201 is input to the lightweight neural network, and the output is 0, 1, or 2, where 0 represents steady-state noise, 1 represents transient noise, and 2 represents speech-dominated noise. The decision logic is also set. The spectrum of the steady-state noise is stable, the ZCR is low (<0.15), the speed fluctuation is small (ΔN < 50 RPM), and the ZCR of the transient noise increases sharply (>0.3).

[0106] S303: Dynamically adjust the noise reduction strategy: enhance low-frequency suppression for steady-state noise, enable transient filters for transient noise, disable vibration fusion for voice-dominated noise, and protect high-frequency voice.

[0107] S304: Perform real-time calibration: When a classification error is detected (speech is misclassified as transient noise), the frame features are recorded and incremental learning is triggered. The formula used for incremental learning is: Where λ = 0.01, only the weight of the last layer is fine-tuned to avoid catastrophic forgetting. is the total loss function, is the cross entropy loss, θ is the current model parameter, and θ0 is the initial model parameter;

[0108] The equipment control module is used to safely execute instructions. For high-risk instructions (emergency stop), voiceprint matching + PLC status confirmation is required. 20 high-frequency instructions ("pause" and "slowdown") are pre-stored and activated when the network is interrupted.

[0109] The self-learning optimization module is used to adapt to the operator's accent and optimize the process parameter recommendation model. The process parameter recommendation model is a baseline model trained based on the manufacturer's standard process library or historical data.

[0110] The specific logical steps for its accent adaptation optimization are as follows:

[0111] S6011: When a new operator uses it for the first time, the system guides the recording of 10 core commands ("accelerate", "stop", "increase coating amount"). Regarding the environmental requirements, the recording is performed under 85dB background noise to ensure consistency with actual working conditions;

[0112] S6012: Perform feature extraction, extracting Mel-frequency cepstral coefficients and new operator voiceprint features, and perform normalization processing to eliminate volume differences and retain pronunciation characteristics;

[0113] S6013: Perform incremental learning to update the industrial speech recognition model: Calculate the loss function and fine-tune the parameters. The loss function calculation formula is the same as that used in S204. When fine-tuning the parameters, retain the speech feature extraction layer of the pre-trained model and only update the weights of the fully connected classification layer to adapt to the new accent characteristics. The industrial speech recognition model is the initial deployment of the system.

[0114] S6014: Push the new model parameters to the edge device through the encrypted channel;

[0115] The specific logical steps for optimizing the process parameters are as follows:

[0116] S6021: Extract key features from historical data, including process parameters such as speed, temperature, and coating weight, as well as result indicators such as composite strength and appearance defect rate. K-means is used to group similar process parameters and mark high-quality combinations.

[0117] S6022: Filter historical high-quality parameter groups and prioritize them based on the current substrate type and working conditions;

[0118] S6023: If the sensor detects an anomaly, it recommends solution parameters for similar historical scenarios;

[0119] S6024: Record the parameters and production results actually used by operators, retrain the recommendation model weekly, and weight the new data;

[0120] The data communication layer includes a real-time control communication module, a cloud collaboration module, and a local data cache module;

[0121] The real-time control communication module is used to interact with the PLC, transmit control instructions and status data, read and write PLC registers via Modbus TCP during protocol configuration, and synchronize device status via EtherCAT.

[0122] The cloud collaboration module is used for remote monitoring, model updates, and fault diagnosis. It pushes alarm information via MQTT, transmits encrypted model files via HTTPS, and automatically synchronizes logs that were not uploaded during the interruption after the network is restored.

[0123] The local data cache module is used to store voice logs and device operation data.

[0124] This embodiment also provides a method for using a solvent-free laminating machine capable of voice interaction, comprising the following steps:

[0125] S1: Synchronously collects audio and vibration signals through the microphone array and vibration sensor in the anti-noise audio acquisition module:

[0126] S2: The audio and vibration signals collected in S1 are weighted and fused using a speech preprocessing module to improve the signal-to-noise ratio of low-frequency speech. The collected signals are also preprocessed for noise reduction to suppress ambient noise (motor and valve sounds) and retain pure speech.

[0127] S3: The acoustic scene adaptation module identifies whether the current noise scene is steady or transient and dynamically adjusts the noise reduction strategy;

[0128] S4: The speech recognition and command parsing module converts the pre-processed speech in S2 into text commands, understands the command intent through the text commands, and converts them into device executable parameters;

[0129] S5: Write the device executable parameters in S4 into the PLC register in the device control module to control the device execution. At the same time, the execution feedback module prompts the instruction execution status through voice broadcast and light.

[0130] S6: When a new operator issues a voice command, the self-learning optimization module is triggered to use incremental learning technology to update the industrial speech recognition model and regularly optimize the process parameter recommendation model;

[0131] S7: Stores voice logs and device operation data through the local data cache module and uploads key data to the cloud;

[0132] This embodiment uses multimodal signal fusion to fuse audio and vibration, thereby improving the low-frequency voice signal-to-noise ratio, thereby improving the accuracy of voice command recognition, and can effectively cope with the complex sound field of the compound machine where high-frequency motor noise and low-frequency vibration are mixed. In combination with adaptive learning technology, it can adapt to the accent of new operators and optimize process parameters, thereby increasing the adaptation range and meeting usage requirements.

[0133] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A solvent-free compounding machine capable of voice interaction, comprising a solvent-free compounding machine body, characterized in that: The solvent-free compounding machine body is equipped with a voice interaction system; The voice interaction system includes a hardware layer, a software layer and a data communication layer. The software layer is connected to the hardware layer and the data communication layer. The data communication layer is connected to the hardware layer, and the data communication layer is connected to the cloud. The hardware layer includes an anti-noise audio acquisition module, a main control calculation module, an execution feedback module and an environment perception module; The software layer includes a speech preprocessing module, a speech recognition and command parsing module, an acoustic scene adaptation module, a device control module, and a self-learning optimization module; The data communication layer includes a real-time control communication module, a cloud collaboration module and a local data cache module.

2. The voice-interactive solvent-free compounding machine according to claim 1, characterized in that: The anti-noise audio acquisition module includes four industrial-grade MEMS microphones arranged linearly with an 8cm spacing. Silicone shock-proof brackets and low-frequency vibration sensors are installed on the microphones. The four microphones directionally pick up the operator's voice, suppressing equipment operating noise, and the vibration sensor compensates for low-frequency voice signals. The main control computing module is used to run noise reduction and speech recognition algorithms and coordinate the work of various modules; The execution feedback module uses voice broadcast and light to prompt the command execution status. The voice broadcast uses a speaker with a frequency response range of 100Hz-16kHz. The three-color light is defined as three flashes of red light for "insufficient adhesive" and two flashes of yellow light for "abnormal tension". The environmental perception module includes a temperature sensor, a humidity sensor, and a Hall-effect speed sensor, which are used to monitor the device's temperature, humidity, and speed, respectively, providing context for noise suppression. The temperature sensor and humidity sensor are mounted on the motor housing, and the Hall-effect speed sensor is aligned with the main drive shaft. The sensor zero point is automatically calibrated upon daily power-up.

3. The voice-interactive solvent-free compounding machine according to claim 2, characterized in that: The microphone uses the TDOA algorithm when performing source positioning, and its formula is: where Δt ij is the time delay difference between microphones i and j, C is the speed of sound, θ is the incident angle of the sound source, r i is the spatial coordinate of microphone i, d ij is the path difference from the sound source to the two microphones; The vibration sensor uses a vibration signal weighted fusion technique to compensate for low-frequency speech signals, and the formula used is: S fuscd (f)=w audio (f)·S audio (f)+w vid (f)·S vid (f); Among them S vid (f) The weight of the low frequency band (<500Hz) is set to 0.7, and the high frequency band is close to 0. fuscd (f) is the fused spectrum, w audio (f) is the audio signal weight, low frequency is 0.3, high frequency is 1.0, w vid (f) is the vibration signal weight, low frequency 0.7, high frequency 0.0, S audio (f) is the microphone audio spectrum.

4. The voice-interactive solvent-free compounding machine according to claim 1, characterized in that: The speech preprocessing module is used for real-time noise reduction and enhanced speech clarity. Its specific logical steps are as follows: S201: Perform 5-500 Hz bandpass filtering on the vibration signal collected by the vibration sensor to eliminate high-frequency mechanical noise interference. During the filtering process, the vibration energy needs to be calculated using the following formula: Among them E vid is the short-time energy of the vibration signal, x vid [n] is the nth sampling point of the vibration signal, N is the calculation frame length, and its value is 160; S202: Audio vibration weighted fusion, the microphone signal is processed in frequency bands, and the low frequency band is calculated using weighted fusion technology: S fuscd (f)=0.7·S vid (f)+0.3S audio (f), the high frequency band uses the microphone signal directly; S203: Noise modeling based on the device speed is performed. Pure noise samples are recorded at different speeds when the device is unloaded. The average spectrum is calculated and the reference noise spectrum is obtained by looking up the table based on the speed sensor signal N(t). The formula used is: in is the current noise power spectrum estimate, N max The maximum permissible speed of the equipment; And perform real-time noise update. When a speechless segment is detected, the noise power spectrum is updated. The formula used is: Where γ=0.95, is the updated noise power spectrum estimate, is the noise power spectrum estimate of the previous frame, f is the frequency, and Y(f) is the power spectrum of the noisy speech in the current frame; S204: Perform frequency domain processing, perform FFT on each frame of audio (20ms length, 50% overlap), obtain the spectrum Y(f), and calculate the spectrum after noise reduction: Where α(f) is the frequency-related over-subtraction factor, which is set to 1.8 in the low frequency band and 1.2 in the high frequency band; β is the spectrum floor coefficient, which is set to 0.002; And retain the original phase φ(f), synthesize the time domain signal, the formula used is; in is the time domain signal, IFFT(·) is the inverse fast Fourier transform, φ(f) is the phase spectrum of the original noisy speech, is the phase information expressed in the form of complex exponential, t is time, and f is frequency; S205: Automatically adjust the gain according to the output voice amplitude, and enable the transient suppression filter when a burst noise is detected. The formula used to adjust the gain is: Where G is the gain coefficient, x ^ [m] is the mth sampling point of the time domain signal after denoising, and M is the calculation frame length.

5. The solvent-free compounding machine capable of voice interaction according to claim 1, characterized in that: The speech recognition and instruction parsing module is used to convert speech into text instructions and parse them into device control parameters; The acoustic scene adaptation module is used to identify whether the current noise scene is steady or transient, and dynamically adjust the noise reduction strategy. The specific logical steps are as follows: S301: Feature extraction: extract audio features and equipment operating condition features. The audio features are Mel-frequency cepstral coefficients. Extract the first 13 dimensional coefficients to characterize the noise spectrum characteristics and detect transient noise. The formula used is: Where x[n] is the audio sampling point, N is the frame length, n is the index of the current sampling point, ranging from 1 to N, x[n] is the amplitude of the discrete signal at index n, x[n-1] is the amplitude of the discrete signal at index n-1, and sgn(·) is the sign function; The equipment operating characteristics include speed value and vibration energy mean value; S302: Scene classification is performed using a lightweight neural network. The feature data extracted in S201 is input to the lightweight neural network, and the output is 0, 1, or 2, where 0 represents steady-state noise, 1 represents transient noise, and 2 represents speech-dominated noise. The decision logic is also set. The spectrum of the steady-state noise is stable, the ZCR is low (<0.15), the speed fluctuation is small (ΔN < 50 RPM), and the ZCR of the transient noise increases sharply (>0.3). S303: Dynamically adjust the noise reduction strategy: enhance low-frequency suppression for steady-state noise, enable transient filters for transient noise, disable vibration fusion for voice-dominated noise, and protect high-frequency voice. S304: Perform real-time calibration: When a classification error is detected, the frame features are recorded and incremental learning is triggered. The formula used for incremental learning is: Where λ = 0.01, only the weight of the last layer is fine-tuned to avoid catastrophic forgetting. is the total loss function, is the cross entropy loss, θ is the current model parameter, and θ0 is the initial model parameter.

6. The voice-interactive solvent-free compounding machine according to claim 1, characterized in that: The device control module is used to safely execute instructions. For high-risk instructions (emergency stop), voiceprint matching + PLC status confirmation is required. 20 high-frequency instructions ("pause" and "slowdown") are pre-stored and activated when the network is interrupted; The self-learning optimization module is used to adapt to the operator's accent and optimize the process parameter recommendation model, which is a baseline model trained based on a standard process library or historical data provided by the manufacturer; The specific logical steps for its accent adaptation optimization are as follows: S6011: When a new operator uses it for the first time, the system guides them to record 10 core instructions. The recording environment is required to be under 85dB background noise to ensure consistency with actual working conditions. S6012: Perform feature extraction, extracting Mel-frequency cepstral coefficients and new operator voiceprint features, and perform normalization processing to eliminate volume differences and retain pronunciation characteristics; S6013: Perform incremental learning to update the industrial speech recognition model: Calculate the loss function and fine-tune the parameters. The loss function calculation formula is the same as that used in S204. When fine-tuning the parameters, retain the speech feature extraction layer of the pre-trained model and only update the weights of the fully connected classification layer to adapt to the new accent characteristics. The industrial speech recognition model is the initial deployment of the system. S6014: Push the new model parameters to the edge device through the encrypted channel; The specific logical steps for optimizing the process parameters are as follows: S6021: Extract key features from historical data, including process parameters such as speed, temperature, and coating weight, as well as result indicators such as composite strength and appearance defect rate. K-means is used to group similar process parameters and mark high-quality combinations. S6022: Filter historical high-quality parameter groups and prioritize them based on the current substrate type and working conditions; S6023: If the sensor detects an anomaly, it recommends solution parameters for similar historical scenarios; S6024: Record the parameters and production results actually used by the operator, retrain the recommendation model every week, and weight the new data.

7. The voice-interactive solvent-free compounding machine according to claim 1, characterized in that: The real-time control communication module is used to interact with the PLC, transmit control instructions and status data, read and write PLC registers through Modbus TCP during protocol configuration, and synchronize device status through EtherCAT; The cloud collaboration module is used for remote monitoring, model updates, and fault diagnosis. It pushes alarm information via MQTT, transmits encrypted model files via HTTPS, and automatically synchronizes logs that were not uploaded during the interruption after the network is restored. The local data cache module is used to store voice logs and device operation data.

8. The method for using a solvent-free compounding machine capable of voice interaction according to any one of claims 1 to 7, characterized in that: The following steps are involved: S1: Synchronously collects audio and vibration signals through the microphone array and vibration sensor in the anti-noise audio acquisition module: S2: The audio and vibration signals collected in S1 are weighted and fused using a speech preprocessing module to improve the signal-to-noise ratio of low-frequency speech. The collected signals are also preprocessed for noise reduction to suppress ambient noise and preserve pure speech. S3: The acoustic scene adaptation module identifies whether the current noise scene is steady or transient and dynamically adjusts the noise reduction strategy; S4: The speech recognition and command parsing module converts the pre-processed speech in S2 into text commands, understands the command intent through the text commands, and converts them into device executable parameters; S5: Write the device executable parameters in S4 into the PLC register in the device control module to control the device execution. At the same time, the execution feedback module prompts the instruction execution status through voice broadcast and light; S6: When a new operator issues a voice command, the self-learning optimization module is triggered to use incremental learning technology to update the industrial speech recognition model and regularly optimize the process parameter recommendation model; S7: Stores voice logs and device operation data through the local data cache module and uploads key data to the cloud.