An intelligent voice interaction system and method applied to a laser cutting machine

By introducing an acoustic front-end module, a dual-domain verification module, and a secure communication module into the laser cutting machine, the problems of low voice recognition rate and misoperation in industrial noise environments have been solved, and a high-accuracy and secure voice interaction system has been achieved.

CN122369438APending Publication Date: 2026-07-10JINAN BODOR LASER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JINAN BODOR LASER CO LTD
Filing Date
2026-03-27
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

The existing voice interaction system of laser cutting machine has a low recognition rate in industrial noise environment and fails to effectively combine the equipment status for instruction verification, resulting in a high risk of misoperation.

Method used

An acoustic front-end module is used for noise suppression and voice enhancement. Signals are collected by a ring microphone array and an accelerometer. A dual-domain verification module performs dual verification at both the physical signal and semantic levels. A secure communication module performs dynamic token verification and handshake protocol interaction to ensure the reliability and security of commands.

Benefits of technology

It improves speech recognition accuracy in noisy environments, reduces the risk of misoperation, and achieves adaptive and secure device status and voice interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369438A_ABST
    Figure CN122369438A_ABST
Patent Text Reader

Abstract

This application relates to the field of industrial automation control technology, specifically to an intelligent voice interaction system and method applied to a laser cutting machine. The system includes: an acoustic front-end module that collects ambient sound and performs noise suppression, outputting a differential feature spectrum; a voice recognition module that outputs recognized text; a dual-domain verification module that receives the differential feature spectrum and the recognized text, determines valid human voice based on the differential feature spectrum at the physical signal level, performs semantic legality verification on the recognized text, and generates an instruction execution decision; a secure communication module that interacts with a PLC via a bus for dynamic token verification and handshake protocol interaction to ensure secure instruction transmission; and an execution control module located within the PLC that controls the operation of the laser cutting machine based on the decision. Through the acoustic front-end's noise adaptive processing, the multi-layered protection mechanism of dual-domain verification, and the secure communication protocol, the system solves problems such as low voice recognition rate, conflict between instructions and equipment status, and insecure communication in high-noise industrial environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of industrial automation control technology, specifically to an intelligent voice interaction system and method for laser cutting machines, and in particular to a voice interaction technology that can adapt to changes in environmental noise and be context-aware based on equipment status. Background Technology

[0002] Currently, most laser cutting machines are operated via touchscreens or physical buttons, which have the following technical drawbacks: First, operators cannot operate the machine when their hands are occupied; second, the high noise levels in industrial environments lead to low voice recognition accuracy; third, voice commands are disconnected from the equipment status, making it easy to make mistakes; and fourth, voice interaction only focuses on voice recognition itself and does not take into account the specific working environment and status of the laser cutting machine.

[0003] Existing speech recognition technologies suffer from a significant drop in recognition rate in high-noise industrial environments and lack a linkage mechanism with equipment operating status. Current intelligent voice interaction systems fail to address speech recognition issues in noisy environments, do not consider noise spikes caused by changes in equipment power, and do not implement dynamic instruction set verification based on equipment status.

[0004] Therefore, there is an urgent need for an intelligent voice interaction system that can adapt to changes in environmental noise and be context-aware based on equipment status, in order to solve the voice control problem of laser cutting machines in industrial noise environments and improve the ease of operation and safety. Summary of the Invention

[0005] To address the above problems, this invention provides an intelligent voice interaction system and method for use in laser cutting machines.

[0006] In a first aspect, the present invention provides an intelligent voice interaction system for use in laser cutting machines, comprising: The acoustic front-end module is used to acquire sound signals in the working environment of the laser cutting machine and perform noise suppression and speech enhancement processing, outputting a differential feature spectrum; A speech recognition module, connected to the acoustic front-end module, is used to perform speech recognition on the signal output by the acoustic front-end module and output the speech-recognized text. A dual-domain verification module, connected to the acoustic front-end module and the speech recognition module, is used to receive the differential feature spectrum and the speech recognition text, and after determining the valid human voice based on the differential feature spectrum at the physical signal level, perform semantic legality verification on the speech recognition text in combination with the current operating status of the laser cutting machine, and generate instruction execution decision; A secure communication module, connected to the dual-domain verification module, is used to receive the instruction execution decision and interact with the programmable logic controller via a bus for dynamic token verification and handshake protocol interaction. The execution control module, located within the programmable logic controller and connected to the security communication module, is used to execute decision-making control of the laser cutting machine according to the instructions, and to feed back the execution results to the dual-domain verification module to update the subsequent verification basis.

[0007] A complete industrial voice control closed loop is constructed by organically combining acoustic front-end modules, speech recognition modules, dual-domain verification modules, secure communication modules, and execution control modules.

[0008] As a preferred embodiment of the technical solution of the present invention, the acoustic front-end module includes: A circular microphone array is positioned at the top of the control panel, offset from the center axis of the cutting head's movement trajectory by a predetermined distance; it is used to acquire multi-channel sound signals. An accelerometer, integrated at the center of a ring microphone array, is used to collect structural noise signals caused by high-frequency vibrations of the laser cutting machine. The vibration coupling compensation unit is used to adaptively filter and cancel the sound signal collected by the microphone using the signal collected by the accelerometer, remove the mechanical noise introduced by vibration, and output the compensated sound signal. The beamforming unit is used to perform time-delay summation on the compensated sound signal and dynamically adjust the beam pointing angle according to the real-time position of the laser cutting head, outputting a beamforming signal. ; The dynamic noise baseline update unit dynamically adjusts the forgetting factor of the noise baseline according to the real-time output power change rate of the laser cutting machine, so as to adapt to the sudden changes in environmental noise caused by the transient changes in laser power, and outputs a dynamic noise baseline. ; The differential characteristic spectrum calculation unit is used to calculate the beamforming signal based on the beamforming signal. and the dynamic noise baseline The differential feature spectrum is calculated and output to the dual-domain verification module.

[0009] The eccentrically positioned ring microphone array is optimized to address the obstruction caused by the gantry structure of the laser cutting machine, avoiding the blind spots in specific directions inherent in traditional circular arrays and ensuring that commands issued by the operator from different positions are effectively captured. An accelerometer integrated at the center of the array collects structural noise signals caused by the high-frequency vibration of the laser cutting head in real time, providing a reference source for vibration coupling compensation. The vibration coupling compensation unit uses the accelerometer signal to adaptively filter and cancel the microphone signal, improving the signal-to-noise ratio. The beamforming unit dynamically adjusts the beam pointing angle based on the real-time position of the laser cutting head, overcoming the interference of the high-speed movement of the cutting head on sound source localization and ensuring that the main lobe of the beam is always aligned with the operator's location. The dynamic noise baseline update unit is linked to the laser power change rate; when a power transient causes a sudden change in environmental noise, it adaptively adjusts the baseline update rate to avoid misinterpreting sudden noise as human voice. The differential characteristic spectrum calculation unit takes the logarithm of the ratio of the beamforming signal to the dynamic noise baseline, outputting a normalized characteristic spectrum that eliminates the influence of absolute volume.

[0010] As a preferred embodiment of the technical solution of the present invention, the dynamic noise baseline update unit updates the noise baseline according to the following formula:

[0011] in, As a dynamic forgetting factor, For the current moment ,frequency The noise baseline value, The noise baseline value is the value from the previous time step, and the dynamic forgetting factor is determined by the probability of speech presence. Real-time output power change rate of laser cutting machine Joint decision:

[0012] in, Basic forgetting rate, Let the probability of the speech existing at the current moment be . This is an adjustment amount for the forgetting factor when speech is present, used to reduce the forgetting factor when speech is detected. Freeze baseline updates; This represents the current percentage of the laser cutting machine's output power. This is a forgetting factor adjustment coefficient used to temporarily reduce power during sudden power changes. To accelerate baseline updates to adapt to abrupt changes in environmental noise; When the laser power changes transiently Adaptive reduction accelerates noise baseline updates to adapt to abrupt changes in environmental noise.

[0013] When speech is detected Xiang Shi Reduce and freeze baseline updates to prevent human voices from being learned as noise; When the laser power changes transiently Xiang Shi Temporarily reduce the noise level to accelerate baseline updates, enabling the noise baseline to quickly track environmental changes. Basal forgetting rate =0.98 ensures smooth baseline updates under steady-state noise, while dynamic adjustment ensures rapid response in special scenarios.

[0014] As a preferred embodiment of the technical solution of the present invention, the beamforming unit dynamically calculates the sound source arrival angle based on the real-time position coordinates of the laser cutting head and adjusts the beam pointing angle so that the main lobe of the beam is always aligned with the area where the operator is located, thereby overcoming the interference of the cutting head movement on the sound source positioning.

[0015] During the operation of a laser cutting machine, the cutting head moves at high speed on a gantry, while the operator is usually in a fixed position. Traditional fixed beam pointing schemes cannot adapt to changes in the relative position of the sound source and the cutting head. When the cutting head moves to certain positions, it may block or reflect sound wave propagation, affecting the beamforming effect.

[0016] This technical solution acquires the cutting head's position coordinates in real time, dynamically calculates the sound source's angle of arrival, and adjusts the beam direction to ensure that the main lobe of the beam is always aligned with the operator's area. This design fully utilizes the laser cutting machine's own position feedback system, achieving adaptive beam control without the need for additional sensors.

[0017] As a preferred embodiment of the technical solution of the present invention, the dual-domain verification module includes: A time-frequency differential mapping unit, connected to the acoustic front-end module, is used to receive the differential feature spectrum and determine whether the current signal is a valid human voice from the physical signal level based on the differential feature spectrum. The semantic vector drift detection unit, connected to the speech recognition module, is used to receive speech recognition text and determine from a semantic perspective whether the speech recognition text matches the current device state.

[0018] The time-frequency differential mapping unit only focuses on whether it is human voice, not the content; the semantic vector drift detection unit only focuses on whether the command is legal, not the clarity of the signal. The two units perform orthogonal verification, without interfering with each other, ensuring command reliability from different dimensions. Even if speech recognition fails in a noisy environment, the physical layer may intercept non-human voice signals; even if the physical layer passes the test, the semantic layer can still intercept illegal commands based on the device status. This multi-layered protection reduces the risk of misoperation.

[0019] The two units can be optimized and upgraded independently. The time-frequency differential mapping unit can adjust the algorithm for different noise environments, and the semantic vector drift detection unit can adjust the instruction set for different device types, which has good scalability.

[0020] As a preferred embodiment of the technical solution of the present invention, the time-frequency differential mapping unit is based on the differential feature spectrum. Calculate the score for valid human voice recognition:

[0021] Among them, the lower limit frequency of the main frequency band of human voice Upper limit frequency of the main frequency band of human voice The main frequency band for human voice, when When the threshold is exceeded, the current frame is determined to be a valid human voice; the differential feature spectrum The acoustic front-end module calculates the result using the following formula:

[0022] In the formula, It is a very small positive number, used to prevent the denominator from being zero.

[0023] The characteristic spectrum itself has been normalized to the noise baseline, representing the signal-to-noise ratio of the signal relative to the noise. It is insensitive to absolute volume and is more suitable for industrial noise environments. The main frequency band of human voice is retained for integration, while high-frequency noise from laser cutting and low-frequency mechanical vibrations are eliminated, further reducing the probability of false triggering. The sum of differential energy across the entire frequency band is used as the judgment score, which has clear physical meaning, simple threshold setting, and low computational cost, making it suitable for real-time processing.

[0024] As a preferred embodiment of the technical solution of the present invention, the semantic vector drift detection unit includes: The expected instruction pool generation submodule is used to dynamically generate the allowed instruction set based on the current process stage of the laser cutting machine. The process stage includes at least one of perforation, cutting, retraction, and standby. The semantic encoding submodule is used to encode speech recognition text into semantic vectors. ; The consistency scoring submodule is used to calculate the state consistency score according to the following formula. ,according to The value generation instruction execution decision;

[0025] This is the semantic vector of the current speech recognition text, generated by the semantic encoding module; For the expected instruction set generated based on the current process stage, a semantic vector containing several expected instructions is provided. ; for and cosine similarity, The state weighting factor is determined based on the current safety status of the laser cutting machine: 1.5 in emergency situations and 1.0 in normal operating situations. The set of allowed instructions is dynamically generated based on the current process stage of the laser cutting machine (piercing, cutting, retraction, standby, etc.), and the allowed instructions differ in different states. For example, start instructions are prohibited in the cutting state, and pause instructions are prohibited in the standby state, fundamentally preventing conflicts between instructions and states.

[0026] A lightweight BERT model is used to encode the identified text into semantic vectors. Compared with traditional keyword matching, it can handle synonyms, near-synonyms, and colloquial expressions, improving user experience. Semantic-level fuzzy matching is achieved by calculating the cosine similarity between the current vector and the expected instruction vector. A state weight factor is introduced to increase sensitivity in emergency situations, ensuring timely execution of critical instructions. Based on the similarity score, three decisions are triggered: direct execution, secondary confirmation, and discard, balancing operational convenience and security.

[0027] As a preferred embodiment of the technical solution of the present invention, the secure communication module includes: A dynamic token generator is used to generate dynamic tokens that increment with the instruction sequence. The generation rules of the tokens are bound to the current security level of the laser cutting machine. The security level includes at least the protective door status and the laser ready status. The handshake protocol unit is used to execute a four-step handshake protocol with the programmable logic controller via an industrial fieldbus. The host computer writes the instruction and dynamic token, and sets the request flag. After verifying the token and device status, the programmable logic controller resets the request flag. After executing an instruction, the programmable logic controller sets the acknowledge flag. After the host computer confirms the response, it resets the response flag and updates the token.

[0028] The token increments sequentially with the instruction sequence and is updated after each instruction interaction. Even if an attacker intercepts the communication messages, they cannot reuse the token, effectively preventing replay attacks. The token generation rule is bound to the current security level of the laser cutting machine. When the equipment is in a high-risk state, the token verification rule automatically adds an extra security bit to further enhance security. A four-step interaction—request sending → programmable logic controller (PLC) verification → instruction execution → host computer confirmation—ensures are not lost or repeatedly executed. The PLC resets the request flag during the verification phase to prevent duplicate triggering; the host computer resets the response flag during the confirmation phase to ensure each instruction has a clear result. Automatic retries are triggered when communication is abnormal, improving system robustness.

[0029] As a preferred embodiment of the technical solution of the present invention, the execution control module feeds back the instruction execution result to the dual-domain verification module, and the dual-domain verification module updates the expected instruction pool according to the execution result to realize context-adaptive voice interaction; When bus communication interruption exceeds a preset threshold, the execution control module controls the laser cutting machine to enter a safety holding mode, prohibiting voice-initiated start commands but allowing voice-initiated emergency stop commands.

[0030] The execution control module feeds back the command execution results to the dual-domain verification module. The dual-domain verification module updates the expected command pool based on the execution results, achieving context-adaptive voice interaction. This mechanism enables the voice interaction system to have memory and learning capabilities, dynamically adjusting the interaction strategy as the device status changes. When bus communication interruption exceeds a preset threshold, the programmable logic controller (PLC) automatically enters a safety hold mode, prohibiting all voice-initiated start commands but allowing voice-initiated emergency stop commands. This design ensures that the device will not go out of control due to loss of voice control during communication failures, while retaining emergency shutdown capabilities, complying with industrial safety standards. The PLC returns execution status codes to the voice system, which then broadcasts different prompts based on the codes, such as "command executed," "device is working," "execution failed," and "please close the safety door first," achieving both visual and voice feedback to enhance the user experience.

[0031] Secondly, the present invention also provides an intelligent voice interaction method for laser cutting machines, applied to the system described in the first aspect, comprising the following steps: S1. Collect ambient sound signals, perform vibration coupling compensation and beamforming through a ring microphone array and accelerometer, output beamforming signals, and dynamically update the noise baseline according to the real-time output power of the laser cutting machine, thereby calculating the differential characteristic spectrum; S2. Based on the differential feature spectrum, calculate the total energy within the main frequency band of human voice. When the total energy exceeds a preset threshold, determine that the current signal is a valid human voice. Perform speech recognition on the beamforming signal and output the speech recognition text. S3. Only when the speech is determined to be valid human voice, perform semantic validity verification on the output speech recognition text: The expected instruction set is dynamically generated based on the current process stage of the laser cutting machine. Encode the speech-recognized text into a semantic vector; Calculate the similarity between the semantic vector and each expected instruction vector in the expected instruction set, and combine it with the current security state weight factor to obtain the state consistency score; The command execution decision is generated based on the state consistency score: when the score is higher than the first threshold, the command is executed directly; when the score is between the second threshold and the first threshold, a second voice confirmation is triggered; when the score is lower than the second threshold, the command is discarded. S4. The instruction execution decision is securely transmitted to the programmable logic controller via dynamic token verification and handshake protocol interaction with the programmable logic controller through the industrial fieldbus. S5. Execute the decision control of the laser cutting machine according to the instructions, and feed back the execution result to step S3 to update the expected instruction set for the next moment, so as to realize context-adaptive voice interaction.

[0032] As can be seen from the above technical solutions, this application has the following advantages: By setting up independent acoustic front-end modules and speech recognition modules, signal processing and content recognition are decoupled. The acoustic front-end focuses on extracting effective signals and outputting differential feature spectra in noisy environments, providing high-quality input for subsequent processing. The dual-domain verification module simultaneously receives differential feature spectra and speech recognition text, performing dual verification at both the physical signal and semantic levels. This fundamentally solves the misidentification problem caused by traditional solutions relying solely on a single confidence level, effectively preventing security risks caused by conflicts between commands and device states. The introduction of dynamic token verification and handshake protocol interaction through the secure communication module ensures the reliability, uniqueness, and security of voice commands during industrial bus transmission, preventing replay attacks and command loss.

[0033] By placing the execution control module inside the programmable logic controller (PLC), the PLC's inherent safety mechanisms and logic interlocking functions are fully utilized, achieving a deep integration of voice control and equipment safety protection. Attached Figure Description

[0034] To more clearly illustrate the technical solution of this application, the accompanying drawings used in the description will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0035] Figure 1A block diagram of a system provided in one embodiment of the present invention.

[0036] Figure 2 A block diagram of a system provided for another embodiment of the present invention.

[0037] Figure 3 This is a flowchart illustrating the method provided in an embodiment of the present invention. Detailed Implementation

[0038] To make the purpose, features, and advantages of this application more apparent and understandable, specific embodiments and accompanying drawings will be used to clearly and completely describe the technical solution protected by this application. Obviously, the embodiments described below are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0039] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this application and in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.

[0040] like Figure 1 As shown, this embodiment of the invention provides an intelligent voice interaction system for use in laser cutting machines, comprising: The acoustic front-end module is used to acquire sound signals in the working environment of the laser cutting machine and perform noise suppression and speech enhancement processing, outputting a differential feature spectrum; A speech recognition module, connected to the acoustic front-end module, is used to perform speech recognition on the signal output by the acoustic front-end module and output the speech-recognized text. A dual-domain verification module, connected to the acoustic front-end module and the speech recognition module, is used to receive the differential feature spectrum and the speech recognition text, and after determining the valid human voice based on the differential feature spectrum at the physical signal level, perform semantic legality verification on the speech recognition text in combination with the current operating status of the laser cutting machine, and generate instruction execution decision; A secure communication module, connected to the dual-domain verification module, is used to receive the instruction execution decision and interact with the programmable logic controller via a bus for dynamic token verification and handshake protocol interaction. The execution control module, located within the programmable logic controller and connected to the security communication module, is used to execute decision-making control of the laser cutting machine according to the instructions, and to feed back the execution results to the dual-domain verification module to update the subsequent verification basis.

[0041] The acoustic front-end module is installed on top of the laser cutting machine's control panel. It collects ambient sound signals, performs noise suppression and speech enhancement processing, and outputs a differential feature spectrum. The speech recognition module connects to the acoustic front-end module and performs speech recognition on the signal output by the acoustic front-end module, outputting the recognized speech text. The dual-domain verification module connects to both the acoustic front-end module and the speech recognition module. It receives the differential feature spectrum and the recognized speech text, determines the validity of the speech at the physical signal level, performs semantic validity verification on the recognized speech text, and generates an instruction execution decision. The secure communication module connects to the dual-domain verification module and receives the instruction execution decision. It interacts with the programmable logic controller (PLC) via an industrial fieldbus for dynamic token verification and handshake protocol communication. The execution control module, located within the PLC, communicates with the secure communication module and controls the laser cutting machine's operating status according to the instruction execution decision.

[0042] In some embodiments, such as Figure 2 As shown, the acoustic front-end module includes: a ring microphone array, an accelerometer, a vibration coupling compensation unit, a beamforming unit, a dynamic noise baseline update unit, and a differential characteristic spectrum calculation unit.

[0043] The circular microphone array adopts an eccentric six-microphone ring structure, containing six MEMS digital microphones (PDM interface), with an array radius of R=45mm, optimized for the 500Hz-4kHz human voice frequency range. This array is positioned at the top of the laser cutting machine's control panel, offset 150mm from the center axis of the cutting head's movement trajectory. This is to reduce the direct airflow impact during cutting head movement and to avoid sound wave obstruction by the laser cutting machine's gantry structure.

[0044] An accelerometer sensor is integrated at the center of a circular microphone array to collect structural noise signals caused by high-frequency vibrations of the laser cutting head. This accelerometer sensor is a triaxial accelerometer, capable of comprehensively sensing vibrations in all directions of the machine body.

[0045] The vibration coupling compensation unit is used to adaptively filter and cancel the sound signal collected by the microphone using the signal acquired by the accelerometer, thus removing mechanical noise introduced by vibration. The specific calculation formula is as follows:

[0046] in, For the first The raw sound signal collected by each microphone The vibration signal is acquired by an accelerometer, and LMS is the least mean square adaptive filtering algorithm. The coupling coefficient, obtained through offline calibration, ranges from 0.1 to 0.3. After vibration coupling compensation, approximately 30% of low-frequency mechanical vibration noise components can be removed, resulting in a compensated audio signal. .

[0047] The beamforming unit performs a short-time Fourier transform (STFT) on the compensated audio signal, performs a delay summation in the frequency domain, and outputs a beamforming signal. The specific calculation formula is as follows:

[0048] in, To compensate for the sound signal The frequency domain representation, The angle of arrival of the sound source is estimated in real time using the GCC-PHAT algorithm. For the first The time delay of each microphone relative to the reference point; As a super-directional weight, the main lobe width is set to 60°.

[0049] Specifically, the beamforming unit is based on the real-time position coordinates of the laser cutting head ( Dynamically calculate the angle of arrival of the sound source. And adjust the beam pointing angle so that the main lobe of the beam is always aligned with the area where the operator is located, thus overcoming the interference of the cutting head movement on the sound source localization.

[0050] The dynamic noise baseline update unit dynamically adjusts the forgetting factor of the noise baseline according to the real-time output power change rate of the laser cutting machine, in order to adapt to the sudden changes in environmental noise caused by laser power transients, and outputs a dynamic noise baseline. ; The updated formula is as follows:

[0051] in, As a dynamic forgetting factor, For the current moment ,frequency The noise baseline value, The noise baseline value is the value from the previous time step, and the dynamic forgetting factor is determined by the probability of speech presence. Real-time output power change rate of laser cutting machine Joint decision:

[0052] in, Basic forgetting rate, The probability of speech presence at the current moment ranges from [0,1] and is provided by the Voice Activity Detection (VAD) algorithm. This is the adjustment amount for the forgetting factor when speech is present, with a value of 0.08, used to reduce the forgetting factor when speech is detected. Freeze baseline updates to prevent human voices from being learned as noise; This represents the current percentage of the laser cutting machine's output power. This is a forgetting factor adjustment coefficient used to temporarily reduce power during sudden power changes. To accelerate baseline updates to adapt to abrupt changes in environmental noise; The value range is [0.90, 0.99].

[0053] When the laser power changes transiently Adaptive reduction accelerates noise baseline updates to adapt to abrupt changes in environmental noise.

[0054] To prevent baseline divergence (such as excessively low baseline after prolonged silence), an energy lower limit clamping mechanism is set:

[0055] in η =0.5, ensuring that the baseline energy is not less than half of the median energy of the past 10 frames.

[0056] The differential characteristic spectrum calculation unit is used to calculate the beamforming signal based on the beamforming signal. and the dynamic noise baseline The differential feature spectrum is calculated and output to the dual-domain verification module.

[0057] The differential feature spectrum The following formula is used for calculation:

[0058] In the formula, It is a very small positive number, used to prevent the denominator from being zero. This difference characteristic spectrum It characterizes the signal-to-noise ratio of the current signal relative to the noise baseline, and serves as the input to the subsequent dual-domain verification module.

[0059] In this embodiment of the invention, the beamforming unit dynamically calculates the sound source arrival angle based on the real-time position coordinates of the laser cutting head and adjusts the beam pointing angle so that the main lobe of the beam is always aligned with the area where the operator is located, thus overcoming the interference of the cutting head movement on the sound source positioning.

[0060] In some embodiments, the speech recognition module is connected to the acoustic front-end module and is used to perform speech recognition on the signal output by the acoustic front-end module, outputting the recognized text. In this embodiment, the speech recognition module uses a deep learning model for speech recognition. The speech signal model can be represented as:

[0061] in y ( t () is the received signal. s( t () represents the original speech signal. n ( t () represents noise.

[0062] The basic formula for speech recognition is:

[0063] in For word strings, Given the input speech feature sequence, For acoustic models, For language models.

[0064] In this embodiment, the CTC loss function is used to solve the problem of inconsistent input and output lengths:

[0065] Where B is the compression function, which compresses the path Mapped to label sequence , For possible state sequences, Given the input speech feature sequence, Indicates the first t Speech features of frames T It is the total number of frames, which is usually much greater than the length of the output text. For all that can be compressed into The set of paths Is the model in the first t Frame prediction is The probability, This represents the probability of a single path.

[0066] In addition, the speech recognition module also includes dereverberation preprocessing, which uses reverberation estimation based on acoustic models and a dereverberation network based on deep learning to reduce signal distortion caused by sound reflections.

[0067] In some embodiments, the dual-domain verification module includes a time-frequency differential mapping unit and a semantic vector drift detection unit.

[0068] The time-frequency differential mapping unit is connected to the acoustic front-end module to receive the differential feature spectrum and determine whether the current signal is valid human voice at the physical signal level based on the differential feature spectrum. The semantic vector drift detection unit is connected to the speech recognition module to receive the speech recognition text and determine whether the speech recognition text matches the current device state at the semantic level.

[0069] The input to the time-frequency difference mapping unit includes: the difference feature spectrum from the acoustic front-end module. In this embodiment, the frame length is set to 25ms, the frame shift to 10ms, and the FFT points to 512. Frequency band filtering retains only the 300Hz-3400Hz main human voice frequency band, eliminating high-frequency noise from laser cutting (>4kHz) and low-frequency mechanical vibrations (<200Hz). The time-frequency differential mapping unit calculates the differential characteristic spectrum based on this... Calculate the score for valid human voice recognition:

[0070] Among them, the lower limit frequency of the main frequency band of human voice (Value taken as 300Hz) up to the upper limit of the main frequency band of human voice (Value taken as 3400Hz) is the main frequency band for human voice, when When the threshold is exceeded, the current frame is determined to be valid human voice; when Exceeding the preset threshold If the current frame is deemed to contain valid human voice, it is discarded as ambient noise. This threshold can be obtained through experimental calibration; in this embodiment, it is set to 0.5.

[0071] The semantic vector drift detection unit includes: an expected instruction pool generation submodule, a semantic encoding submodule, and a consistency scoring submodule.

[0072] The expected instruction pool generation submodule is used to dynamically generate the allowed instruction set based on the current process stage of the laser cutting machine. The laser cutting machine's process stages include, but are not limited to: piercing, cutting, retraction, and standby. For example: When the device status is cutting, the expected instruction pool ={Stop, Pause, Reduce Power}; When the device is in standby mode, the expected instruction pool ={Startup, Return to Origin, Parameter Settings}; When the device status is perforated, the expected instruction pool ={Pause, Cancel Piercing}.

[0073] The semantic encoding submodule is used to encode speech recognition text into semantic vectors. In this embodiment, a lightweight BERT model is used to encode the recognized text into a 128-dimensional semantic vector.

[0074] The consistency scoring submodule is used to calculate the state consistency score according to the following formula. ,according to The value generation instruction execution decision;

[0075] in, This is the semantic vector of the current speech recognition text, generated by the semantic encoding module; For the expected instruction set generated based on the current process stage, a semantic vector containing several expected instructions is provided. ; for and The cosine similarity, with values ​​ranging from [ [1,1], the closer the value is to 1, the closer the semantics are; The state weighting factor is determined based on the current safety status of the laser cutting machine: 1.5 in emergency situations and 1.0 in normal operating situations. The consistency scoring submodule is based on Value generation instruction execution decision: when When the value is ≥0.85, the instruction is executed directly; When 0.60≤ When the value is less than 0.85, a second voice confirmation is triggered; when If the value is less than 0.60, it is determined to be a misidentified or irrelevant instruction, discarded, and logged.

[0076] In some embodiments, the secure communication module includes a dynamic token generator and a handshake protocol unit.

[0077] This system adopts a two-layer architecture of upper-level semantic parsing and lower-level logic verification. The speech recognition system (upper computer) is responsible for intent understanding, and the PLC (lower computer) is responsible for safe execution and status feedback. The two exchange data via an industrial fieldbus (such as ModbusTCP, Profinet, or EtherCAT). This embodiment uses the register address allocation table shown in Table 1. Table 1

[0078] A dynamic token generator is used to generate dynamic tokens that increment sequentially with the command sequence. In this embodiment, the token is a 32-bit unsigned integer, automatically incremented by 1 after each successful command interaction. Specifically, the token generation rules are tied to the current security level of the laser cutting machine, which includes at least the status of the protective door and the laser's ready state. For example, when the protective door is open, the token verification rules automatically add an extra security check bit to prevent accidental operation.

[0079] The handshake protocol unit is used to execute a four-step handshake protocol with the PLC via an industrial fieldbus. (Define time window) and timeout threshold .

[0080] Step 1: Send Request The host computer writes CMD_ID, CMD_PARAM, and CMD_TOKEN; The host computer sets REQ_TRIGGER from 0 to 1 (triggered on rising edge); Record sending time .

[0081] Step 2: PLC Receiver Verification The PLC detected the rising edge of REQ_TRIGGER; PLC checks whether CMD_TOKEN matches (to prevent replay attacks); PLC verifies internal safety status (such as interlock conditions and protective door status). The PLC resets REQ_TRIGGER to 0 (to prevent repeated triggering).

[0082] Step 3: Command Execution The PLC executes specific actions (such as starting the laser and adjusting the power). The PLC writes the execution result to STS_CODE; The PLC sets ACK_TRIGGER from 0 to 1.

[0083] Step 4: Host computer confirmation The host computer polls ACK_TRIGGER; If a rising edge is detected, verify the returned STS_CODE and ACK_TOKEN. The host computer resets ACK_TRIGGER to 0; Update local Token_next; process complete.

[0084] If the host computer is If no ACK_TRIGGER transition is detected, a timeout retry is triggered. The maximum number of retries is 3. Exceeding this limit will be considered a communication failure and trigger an alarm.

[0085] In some embodiments, the execution control module is located within the PLC and communicates with the safety communication module. It is used to execute decisions and control the operating state of the laser cutting machine according to instructions. The execution control module includes a state machine and a logic interlock unit.

[0086] The state machine is used to manage the operating status of the laser cutting machine, including but not limited to: standby, piercing, cutting, retraction, emergency stop, and fault states. The state machine makes decisions based on received instructions to switch states and feeds back the current state in real time to the expected instruction pool generation module of the dual-domain verification module. This updates the expected instruction set for the next moment, enabling context-adaptive voice interaction.

[0087] Logical interlock units are used to ensure the safety of instruction execution. For example: When the protective door is open, the command to start the laser must not be executed; When the laser temperature is too high, the power increase command is prohibited from being executed; Rapid feed commands are prohibited when the cutting head is too close to the workpiece.

[0088] If bus communication interruption exceeds a set threshold (e.g., 500ms), the PLC automatically enters safety hold mode, disabling all voice-initiated start commands, but allowing voice-initiated emergency stop commands (hard-wire backup required). The host computer interface will display "communication offline"; please switch to manual mode.

[0089] The PLC returns the STS_CODE to the speech system, and the speech system reads different prompts according to the code, as shown in Table 2: Table 2

[0090] The execution control module feeds back the command execution results to the dual-domain verification module. The dual-domain verification module updates the expected command pool based on the execution results, enabling context-adaptive voice interaction. For example, after executing the start command, the device status switches from standby to cutting. At this time, the expected command pool is automatically updated to {stop, pause, reduce power}, preventing the start command from being accidentally triggered again during the cutting process.

[0091] By employing an eccentric ring microphone array, accelerometer vibration compensation, dynamic noise baseline updates, and dynamic beamforming direction adjustment, the speech recognition accuracy is improved from <65% to >90% in strong steady-state industrial noise environments of 80-100dB. A dual-domain verification mechanism verifies command legitimacy at both the physical signal and semantic levels, effectively preventing security risks caused by conflicts between commands and equipment states. Dynamic token verification and a four-step handshake protocol ensure the reliability and security of command transmission; combined with the PLC's internal logic interlocking unit, intrinsic safety is achieved. By feeding the execution results back to the expected command pool, dynamic adaptation between equipment status and voice interaction is achieved, enhancing the user experience.

[0092] like Figure 3 As shown, an intelligent voice interaction method for laser cutting machines, applied to the system described in the above embodiments, includes the following steps: S1. Collect ambient sound signals, perform vibration coupling compensation and beamforming through a ring microphone array and accelerometer, output beamforming signals, and dynamically update the noise baseline according to the real-time output power of the laser cutting machine, thereby calculating the differential characteristic spectrum; Ambient sound signals are acquired and preprocessed through an acoustic front-end module. The acoustic front-end module includes an off-center ring microphone array and an accelerometer integrated at the center of the array.

[0093] The ring microphone array adopts an eccentric six-microphone ring structure, containing six MEMS digital microphones with an array radius of R=45mm. It is positioned at the top of the laser cutting machine's operation panel, 150mm off-center from the central axis of the cutting head's movement trajectory. This layout is optimized to mitigate the obstruction caused by the laser cutting machine's gantry structure and reduce direct airflow impact during cutting head movement. An accelerometer is integrated at the center of the array to collect structural noise signals caused by the high-frequency vibration of the laser cutting head.

[0094] S2. Based on the differential feature spectrum, calculate the total energy within the main frequency band of human voice. When the total energy exceeds a preset threshold, determine that the current signal is a valid human voice. Perform speech recognition on the beamforming signal and output the speech recognition text. S3. Only when the speech is determined to be valid human voice, perform semantic validity verification on the output speech recognition text: The expected instruction set is dynamically generated based on the current process stage of the laser cutting machine. Encode the speech-recognized text into a semantic vector; Calculate the similarity between the semantic vector and each expected instruction vector in the expected instruction set, and combine it with the current security state weight factor to obtain the state consistency score; The command execution decision is generated based on the state consistency score: when the score is higher than the first threshold, the command is executed directly; when the score is between the second threshold and the first threshold, a second voice confirmation is triggered; when the score is lower than the second threshold, the command is discarded. S4. The instruction execution decision is securely transmitted to the programmable logic controller via dynamic token verification and handshake protocol interaction with the programmable logic controller through the industrial fieldbus. S5. Execute the decision control of the laser cutting machine according to the instructions, and feed back the execution result to step S3 to update the expected instruction set for the next moment, so as to realize context-adaptive voice interaction.

[0095] The process of dynamically generating the expected instruction set based on the current process stage of the laser cutting machine specifically includes: When the laser cutting machine is in the piercing stage, the expected instruction set includes pause and cancel piercing; When the laser cutting machine is in the cutting stage, the expected instruction set includes stop, pause, and power reduction; When the laser cutting machine is in the retraction phase, the expected instruction set includes stop and pause; When the laser cutting machine is in standby mode, the expected instruction set includes start, return to origin, and parameter settings.

[0096] The handshake protocol interaction specifically includes: The host computer writes the instruction and dynamic token, and sets the request flag. After verifying the token and device status, the programmable logic controller resets the request flag. After executing an instruction, the programmable logic controller sets the acknowledge flag. After the host computer confirms the response, it resets the response flag and updates the token. The dynamic token increments with the instruction sequence, and its generation rule is bound to the current security level of the laser cutting machine. The security level includes at least the protective door status and the laser ready status.

[0097] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, mobile hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, or other media capable of storing program code. It includes several instructions to cause a computer terminal (which may be a personal computer, server, or a second terminal, network terminal, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0098] In the embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0099] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0100] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0101] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An intelligent voice interaction system for use in laser cutting machines, characterized in that, include: The acoustic front-end module is used to acquire sound signals in the working environment of the laser cutting machine and perform noise suppression and speech enhancement processing, outputting a differential feature spectrum; A speech recognition module, connected to the acoustic front-end module, is used to perform speech recognition on the signal output by the acoustic front-end module and output the speech-recognized text. A dual-domain verification module, connected to the acoustic front-end module and the speech recognition module, is used to receive the differential feature spectrum and the speech recognition text, and after determining the valid human voice based on the differential feature spectrum at the physical signal level, perform semantic legality verification on the speech recognition text in combination with the current operating status of the laser cutting machine, and generate instruction execution decision; A secure communication module, connected to the dual-domain verification module, is used to receive the instruction execution decision and interact with the programmable logic controller via a bus for dynamic token verification and handshake protocol interaction. An execution control module, located within the programmable logic controller and connected to the security communication module, is used to execute decision-making control of the laser cutting machine's operating status according to the instructions.

2. The intelligent voice interaction system for laser cutting machines according to claim 1, characterized in that, The acoustic front-end module includes: A circular microphone array is positioned at the top of the control panel, offset from the center axis of the cutting head's movement trajectory by a predetermined distance; it is used to acquire multi-channel sound signals. An accelerometer, integrated at the center of a ring microphone array, is used to collect structural noise signals caused by high-frequency vibrations of the laser cutting machine. The vibration coupling compensation unit is used to adaptively filter and cancel the sound signal collected by the microphone using the signal collected by the accelerometer, remove the mechanical noise introduced by vibration, and output the compensated sound signal. The beamforming unit is used to perform time-delay summation on the compensated sound signal and dynamically adjust the beam pointing angle according to the real-time position of the laser cutting head, outputting a beamforming signal. ; The dynamic noise baseline update unit dynamically adjusts the forgetting factor of the noise baseline according to the real-time output power change rate of the laser cutting machine, so as to adapt to the sudden changes in environmental noise caused by the transient changes in laser power, and outputs a dynamic noise baseline. ; The differential characteristic spectrum calculation unit is used to calculate the beamforming signal based on the beamforming signal. and the dynamic noise baseline The differential feature spectrum is calculated and output to the dual-domain verification module.

3. The intelligent voice interaction system for laser cutting machines according to claim 2, characterized in that, The dynamic noise baseline update unit updates the noise baseline according to the following formula: in, As a dynamic forgetting factor, For the current moment ,frequency The noise baseline value, The noise baseline value at the previous moment is used as the dynamic forgetting factor, which is determined by the probability of speech presence. Real-time output power change rate of laser cutting machine Joint decision: in, Basic forgetting rate, Let the probability of the speech existing at the current moment be . This is an adjustment amount for the forgetting factor when speech is present, used to reduce the forgetting factor when speech is detected. Freeze baseline updates; This represents the current percentage of the laser cutting machine's output power. This is a forgetting factor adjustment coefficient used to temporarily reduce power during sudden power changes. To accelerate baseline updates to adapt to abrupt changes in environmental noise; When the laser power changes transiently Adaptive reduction accelerates noise baseline updates to adapt to abrupt changes in environmental noise.

4. The intelligent voice interaction system for laser cutting machines according to claim 2, characterized in that, The beamforming unit dynamically calculates the sound source arrival angle based on the real-time position coordinates of the laser cutting head and adjusts the beam pointing angle so that the main lobe of the beam is always aligned with the area where the operator is located, thus overcoming the interference of the cutting head movement on the sound source positioning.

5. The intelligent voice interaction system for laser cutting machines according to claim 3, characterized in that, The two-domain verification module includes: A time-frequency differential mapping unit, connected to the acoustic front-end module, is used to receive the differential feature spectrum and determine whether the current signal is a valid human voice from the physical signal level based on the differential feature spectrum. The semantic vector drift detection unit, connected to the speech recognition module, is used to receive speech recognition text and determine from a semantic perspective whether the speech recognition text matches the current device state.

6. The intelligent voice interaction system for laser cutting machines according to claim 5, characterized in that, The time-frequency difference mapping unit is based on the difference feature spectrum. Calculate the score for valid human voice recognition: Among them, the lower limit frequency of the main frequency band of human voice Upper limit frequency of the main frequency band of human voice The main frequency band for human voice, when When the threshold is exceeded, the current frame is determined to be a valid human voice; the differential feature spectrum The acoustic front-end module calculates the result using the following formula: In the formula, It is a very small positive number, used to prevent the denominator from being zero.

7. The intelligent voice interaction system for laser cutting machines according to claim 5, characterized in that, The semantic vector drift detection unit includes: The expected instruction pool generation submodule is used to dynamically generate the allowed instruction set based on the current process stage of the laser cutting machine. The process stage includes at least one of perforation, cutting, retraction, and standby. The semantic encoding submodule is used to encode speech recognition text into semantic vectors. ; The consistency scoring submodule is used to calculate the state consistency score according to the following formula. ,according to The value generation instruction execution decision; This is the semantic vector of the current speech recognition text. For the expected instruction set generated based on the current process stage, a semantic vector containing several expected instructions is provided. ; for and cosine similarity, This is the state weighting factor, which is determined based on the current safety status of the laser cutting machine.

8. The intelligent voice interaction system for laser cutting machines according to claim 1, characterized in that, The secure communication module includes: A dynamic token generator is used to generate dynamic tokens that increment with the instruction sequence. The generation rules of the tokens are bound to the current security level of the laser cutting machine. The security level includes at least the protective door status and the laser ready status. The handshake protocol unit is used to execute the handshake protocol with the programmable logic controller via the industrial fieldbus. The host computer writes the instruction and dynamic token, and sets the request flag. After verifying the token and device status, the programmable logic controller resets the request flag. After executing an instruction, the programmable logic controller sets the acknowledge flag. After the host computer confirms the response, it resets the response flag and updates the token.

9. The intelligent voice interaction system for laser cutting machines according to claim 1, characterized in that, The execution control module feeds back the instruction execution result to the dual-domain verification module, which updates the expected instruction pool based on the execution result to achieve context-adaptive voice interaction. When bus communication interruption exceeds a preset threshold, the execution control module controls the laser cutting machine to enter a safety holding mode, prohibiting voice-initiated start commands but allowing voice-initiated emergency stop commands.

10. An intelligent voice interaction method for use in a laser cutting machine, applied to the system described in any one of claims 1 to 9, characterized in that, Includes the following steps: S1. Collect ambient sound signals, perform vibration coupling compensation and beamforming through a ring microphone array and accelerometer, output beamforming signals, and dynamically update the noise baseline according to the real-time output power of the laser cutting machine, thereby calculating the differential characteristic spectrum; S2. Based on the differential feature spectrum, calculate the total energy within the main frequency band of human voice. When the total energy exceeds a preset threshold, determine that the current signal is a valid human voice. Perform speech recognition on the beamforming signal and output the speech recognition text. S3. Only when the speech is determined to be valid human voice, perform semantic validity verification on the output speech recognition text: The expected instruction set is dynamically generated based on the current process stage of the laser cutting machine. Encode the speech-recognized text into a semantic vector; Calculate the similarity between the semantic vector and each expected instruction vector in the expected instruction set, and combine it with the current security state weight factor to obtain the state consistency score; The command execution decision is generated based on the state consistency score: when the score is higher than the first threshold, the command is executed directly; when the score is between the second threshold and the first threshold, a second voice confirmation is triggered; when the score is lower than the second threshold, the command is discarded. S4. The instruction execution decision is securely transmitted to the programmable logic controller via dynamic token verification and handshake protocol interaction with the programmable logic controller through the industrial fieldbus. S5. Execute the decision control of the laser cutting machine according to the instructions, and feed back the execution result to step S3 to update the expected instruction set for the next moment, so as to realize context-adaptive voice interaction.