Electromechanical equipment remote control device based on artificial intelligence voice interaction
Through voice interaction and multimodal perception technology based on artificial intelligence, combined with predictive risk assessment, the problems of inconvenient user interaction, insufficient state perception and insufficient risk prediction of traditional electromechanical equipment remote control systems are solved, and higher intelligence and safe remote control are achieved.
Patent Information
- Application Number
- CN202510747387.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-08-26
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional remote control systems for electromechanical equipment have shortcomings in terms of user interaction convenience, real-time and depth of device status perception, potential risk prediction of operation instructions, and system adaptive optimization capabilities, resulting in poor security and user experience.
The remote control device of electromechanical equipment based on artificial intelligence voice interaction is adopted, combined with user interaction devices, edge cognition bodies and predictive active safety interaction engines, and through multimodal sensor data perception, automatic speech recognition and natural language understanding, a deep understanding of user intentions and device status is achieved, and predictive risk assessment and adaptive optimization are carried out.
It improves the intelligence level and safety of remote control of electromechanical equipment, realizes active prediction and avoidance of potential risks, optimizes equipment maintenance strategies, and improves the convenience of user interaction and the adaptability of the system.
Smart Images

Figure CN120544548A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a remote control device for electromechanical equipment based on artificial intelligence voice interaction. Background Art
[0002] With the continuous improvement of industrial automation and intelligence, the demand for remote control of electromechanical equipment is growing. Traditional remote control systems for electromechanical equipment still have certain limitations in terms of user interaction convenience, real-time and in-depth perception of equipment status, and safety assurance during operation. Existing remote control systems often rely on preset graphical user interfaces or fixed control logic, which may require users to perform complex or non-intuitive operations when issuing commands. At the same time, these systems often lack a comprehensive and in-depth understanding of the operating status of electromechanical equipment itself. In most cases, they can only obtain basic operating parameters and lack real-time insight and early warning capabilities for potential equipment anomalies and health decline trends.
[0003] When executing user commands, traditional systems often lack a proactive assessment mechanism for the risks that the commands themselves might pose given the current state of the device. This means that even if a user command could cause equipment overload, damage, or pose a safety hazard, the system may still proceed without further delay. Safety assurance relies heavily on the operator's experience and standards. Furthermore, these systems typically lack the ability to self-learn and optimize based on actual usage and user feedback, making it difficult to continuously improve performance and user experience after deployment.
[0004] Therefore, the present invention proposes a remote control device for electromechanical equipment based on artificial intelligence voice interaction to solve the shortcomings of the existing technology. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention provides a remote control device for electromechanical equipment based on artificial intelligence voice interaction, which solves the shortcomings of the remote control system of electromechanical equipment in terms of the convenience of voice interaction, real-time deep perception of equipment status, active prediction and avoidance of potential risks of operating instructions, and system adaptive optimization capabilities, thereby improving the intelligence level, safety and user experience of remote control.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a remote control device for electromechanical equipment based on artificial intelligence voice interaction, including: a user interaction device, used to collect user voice control instructions and present information to the user; an electromechanical device, used to perform predetermined operations according to the control instructions; the electromechanical equipment includes a receiving terminal, a control unit and an actuator, the receiving terminal is electrically connected to the control unit, and the control unit is electrically connected to the actuator; the user interaction device is communicatively connected to the receiving terminal of the electromechanical equipment; an edge cognition is coupled to the electromechanical equipment, and the edge cognition is used to collect multimodal sensor data of the electromechanical equipment and generate equipment insight information based on the multimodal sensor data; a predictive active safety interaction engine is communicatively connected to the user interaction device and the edge cognition, used to process user voice instructions to determine user intentions, and generate control instructions for the electromechanical equipment in combination with the user intentions and the equipment insight information.
[0007] Preferably, the user interaction device includes: a microphone for collecting the user's voice instructions; a display unit for presenting a graphical user interface and visual information related to the operation of the electromechanical device to the user; and a speaker for broadcasting voice feedback and system prompts to the user.
[0008] Preferably, the receiving terminal of the electromechanical device is used to receive the control instructions from the user interaction device; the control unit of the electromechanical device includes a programmable logic controller or a microcontroller unit, and the control unit is used to parse the control instructions and drive the actuator of the electromechanical device accordingly; the actuator of the electromechanical device includes a motor, a valve, a heater or a robotic arm, and the actuator is used to perform physical operations corresponding to the control instructions; the communication connection between the user interaction device and the receiving terminal of the electromechanical device is a wireless communication connection.
[0009] Preferably, the edge cognitive body includes: a multimodal sensor array coupled to the electromechanical device, for collecting the multimodal sensor data, wherein the multimodal sensor data includes vibration signals, acoustic signals or thermal imaging data; an edge intelligent processor connected to the multimodal sensor array, for learning the digital representation of the state of the electromechanical device from the multimodal sensor data, detecting the operating anomaly of the electromechanical device based on the digital representation, and evaluating the health index of the electromechanical device or predicting its remaining effective life, and using the operating anomaly, the health index or the remaining effective life as part of the device insight information.
[0010] Preferably, the predictive active safety interaction engine includes: a voice instruction processing module, which is used to convert the user voice instruction into text through automatic speech recognition, and analyze the text through natural language understanding to determine the user intention; a risk assessment module, which is connected to the voice instruction processing module and the edge cognitive body, and the risk assessment module is used to perform a predictive risk assessment on the operation corresponding to the user intention based on the user intention and the device insight information; a decision and instruction generation module is connected to the risk assessment module, and is used to generate the control instruction when the result of the predictive risk assessment indicates that the risk is lower than a preset threshold; an alarm prompt module, which generates a safety warning for user perception, or adjusts the control instruction to reduce the risk when the result of the predictive risk assessment indicates that the risk exceeds the preset threshold.
[0011] The present invention also provides a method for remotely controlling electromechanical equipment based on artificial intelligence voice interaction, the method comprising the following steps:
[0012] S1. Collecting user voice commands through a user interaction device, and processing the user voice commands by the predictive active safety interaction engine using an automatic speech recognition model and a natural language understanding model based on machine learning to determine user intent;
[0013] S2. Collecting multimodal sensor data from electromechanical equipment through an edge cognito, and processing the multimodal sensor data using a deep learning-based digital representation learning algorithm and anomaly detection algorithm, combined with an equipment health status prediction model, to generate equipment insight information.
[0014] S3. The predictive active safety interaction engine combines the user intent with the device insight information, performs a risk assessment using a predictive risk assessment model, and generates an operation instruction based on the risk assessment result, where the operation instruction is a control instruction for the electromechanical device or a safety warning perceived by the user;
[0015] S4. Perform subsequent processing according to the operation instruction: if the operation instruction is the control instruction, transmit it to the control unit of the electromechanical device to drive the actuator to perform a predetermined operation; if the operation instruction is the safety warning, present the safety warning to the user through the user interaction device.
[0016] Preferably, the step S1 includes: the user voice command collected by the user interaction device is represented by the original audio signal V audio , is input into the automatic speech recognition model; the automatic speech recognition model uses the acoustic and language knowledge obtained through machine learning training to transform the original audio signal V audio Convert to the corresponding text sequence Tasr , the conversion process is: T asr =M ASR (V audio θ ASR );
[0017] Among them, M ASR represents the function of the automatic speech recognition model, and θ ASR represents a set of learned parameters of the automatic speech recognition model; a text sequence T output by the automatic speech recognition model asr As input, passed to the natural language understanding model;
[0018] The natural language understanding model processes the text sequence T asr Conduct in-depth semantic analysis and key information extraction to identify the user's specific intention; User Intention I u is further structured to include a clear intent category C I and a set of related parameters or slot information P slots The user intends to u The determination process is: I u ={C I ,P slots}=M NLU (T asr θ NLU );
[0019] Among them, M NLU represents the function of the natural language understanding model, θ NLU Represents the learned parameter set of the natural language understanding model, and the final determined structured user intention I u .
[0020] Preferably, the step S2 includes: preprocessing the collected multimodal sensor data to obtain the preprocessed multimodal sensor data X proc_msa (t), where t is a timestamp; the pre-processed multimodal sensor data X proc_msa (t) Input to the encoder f of the digital representation learning algorithm DRL_enc , the encoder is used to learn and extract the digital representation Z of the state of the electromechanical device at time t state (t), which is calculated as: Z state (t) = f DRL_enc (X proc_msa (t);θ DRL_enc );
[0021] Among them, θ DRL_enc For the encoder f DRL_enc The learned model parameters of
[0022] The abnormality detection algorithm is used to calculate the abnormality score S of the electromechanical device at time t. anomaly (t), thereby detecting the abnormal operation of the electromechanical equipment; the abnormality score S anomaly The calculation of (t) is based on either of the following methods: directly based on the digital representation Z state (t) performing evaluation, or based on the pre-processed multimodal sensor data X proc_msa (t) and the decoder f by the digital representation learning algorithm DRL_dec From the digital representation Z state (t) Reconstructed data The deviation between is evaluated, where the reconstructed data is:
[0023] Among them, θ DRL_dec For the decoder f DRL_dec Model parameters; and, using the equipment health status prediction model, based on the digital representation Z state (t), evaluate the health index H(t) of the electromechanical equipment or predict its remaining useful life RUL(t); output the operating abnormality, the health index H(t) or the remaining useful life RUL(t) as the equipment insight information.
[0024] Preferably, the step S3 includes: the predictive active safety interaction engine is integrated into the determined user intention, denoted as I u , and the generated device insight information, denoted as D insight ; The user intention I u and the device insight information D insight , and combined with the preset electromechanical equipment operation safety regulations and knowledge base, recorded as K safety , are provided as input to the predictive risk assessment model; the core function of the predictive risk assessment model is denoted as M PRA , used to evaluate the execution and the user intention I u The potential risk level associated with the corresponding operation is denoted as R pred ; The risk assessment process is: R pred =M PRA (I u ,D insight ,K safety θ PRA );
[0025] Among them, θ PRA Represents the predictive risk assessment model M PRAThe learned or configured parameter set; through the intelligent decision-making process, the calculated potential risk level M PRA With a preset security risk threshold τ risk Compare to determine the subsequent operation instructions: If the potential risk level R pred Below the security risk threshold τ risk , the intelligent decision-making process determines that the current operation is safe or has an acceptable risk, and generates a u Directly corresponding control instruction C cmd As the operation instruction; if the potential risk level R pred Equal to or exceeding the security risk threshold τ risk , the intelligent decision-making process determines that the current operation has a high risk and generates a safety warning A for the user to perceive. sfty As the operation instruction, or the intelligent decision-making process attempts to adjust the original user intention I u The operating parameters in the ts are used to reduce the risk and generate the adjusted, risk-reduced control instructions C′ cmd As the operation instruction; the control instruction C that is finally generated cmd 、Safety Warning A sfty , or the adjusted control instruction C′ cmd is the operation instruction.
[0026] Preferably, the step S4 includes: collecting user feedback provided by the user for the generated device insight information or the generated operation instruction through the user interaction device, and the user feedback is recorded as F u , the user feedback F u It can be the user's explicit evaluation of the system output, correction information, or implicit preference reflected by the user's subsequent behavior; based on the collected user feedback F u , update or adjust the parameters of the artificial intelligence model in the device to optimize its future performance; the updating process specifically includes updating a model in the edge cognitive body, or updating a model in the predictive active safety interaction engine; when updating the model in the edge cognitive body, it involves adjusting the encoder f of the digital representation learning algorithm DRL_enc The model parameters θ DRL_enc , or its decoder f DRL_dec The model parameters θ DRL_dec , or the relevant parameters of the anomaly detection algorithm, or the parameters of the equipment health status prediction model, this update is expressed as: θ′ EC =Φ EC (θ EC ,F u );
[0027] '
[0028] Among them, θ EC represents the edge cognitive model parameter set to be updated, θ EC represents the updated parameter set, θ EC Represents the update function for the edge cognition model;
[0029] When updating the model in the predictive active safety interaction engine, it involves adjusting the model parameters θ of the automatic speech recognition model. ASR , or the model parameters θ of the natural language understanding model NLU , or the model parameter θ of the predictive risk assessment model PRA , this update is expressed as: θ′ PPSIE =Φ PPSIE (θ PPSIE ,F u );
[0030] Among them, θ PPSIE Represents the parameter set of the predictive active safety interaction engine model to be updated, θ′ PPSIE represents the updated parameter set, Φ PPSIE Represents the update function for the predictive active safety interaction engine model.
[0031] The present invention provides a remote control device for electromechanical equipment based on artificial intelligence voice interaction. It has the following beneficial effects:
[0032] 1. This invention achieves deep intelligence in the control of electromechanical equipment by integrating AI-based voice interaction, real-time device insights provided by edge cognition, and a predictive proactive safety interaction engine. This improves operational safety. Compared to traditional remote control systems that typically only passively execute commands and lack comprehensive assessment of the real-time status of the equipment and potential operational risks, this invention proactively identifies and avoids high-risk operations, effectively preventing safety incidents caused by misoperation or poor equipment condition.
[0033] 2. This invention utilizes edge cognition to perform multimodal data perception and intelligent analysis on electromechanical equipment, generating accurate, real-time insights into equipment, including operational anomalies, health indicators, and remaining useful life predictions. This optimizes equipment maintenance strategies and management efficiency. While existing technologies often rely on fixed maintenance cycles or post-failure responses, this invention facilitates predictive maintenance through forward-looking health assessments, reducing unplanned downtime, extending equipment lifespan, and lowering operational costs.
[0034] 3. This invention utilizes automatic speech recognition and natural language understanding technologies, tightly integrated with a predictive active safety interaction engine, enabling users to interact with electromechanical devices through natural and convenient voice commands. This interactive method significantly improves the system's usability and operational efficiency. Compared to the complex graphical user interfaces or cumbersome keystrokes commonly used in existing technologies, this invention simplifies the human-computer interaction process and lowers the operational threshold, making it particularly suitable for emergency situations or scenarios where manual operation is inconvenient.
[0035] 4. The present invention's iterative optimization mechanism for artificial intelligence models based on user feedback allows the system to continuously improve the performance of its internal AI models by learning from users' explicit evaluations or implicit preferences in actual operations. This closed-loop adaptive learning capability ensures the continuous evolution of the system's intelligence level. Compared to most existing AI systems, whose model parameters are fixed after deployment and have difficulty adapting to changes in working conditions or user habits, the present invention solves the problems of insufficient adaptability and potential long-term performance degradation, enabling the system to better serve users and specific application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 A perspective view of the device of the present invention;
[0037] Figure 2 A schematic diagram of the electromechanical device of the present invention;
[0038] Figure 3 Schematic diagram of the edge cognitive body of the present invention;
[0039] Figure 4 Schematic diagram of the predictive active safety interaction engine of the present invention;
[0040] Figure 5 is a schematic diagram of the device architecture of the present invention;
[0041] Figure 6 Schematic diagram of the method of the present invention.
[0042] Among them, 1. User interaction device; 11. Microphone; 12. Display unit; 13. Speaker; 2. Electromechanical equipment; 21. Receiving terminal; 22. Control unit; 23. Actuator; 3. Edge cognitive body; 31. Multimodal sensor array; 32. Edge intelligent processor; 4. Predictive active safety interaction engine; 41. Voice command processing module; 42. Risk assessment module; 43. Decision and command generation module; 44. Alarm prompt module. DETAILED DESCRIPTION
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0044] See also Figure 1-5 , an embodiment of the present invention provides a remote control device for electromechanical equipment based on artificial intelligence voice interaction, comprising: a user interaction device 1, for collecting user voice control instructions and presenting information to the user; the user interaction device 1 comprises: a microphone 11, for collecting the user voice instructions, such as a high-fidelity microphone array to capture clear voice signals; a display unit 12, for presenting a graphical user interface and visual information related to the operation of the electromechanical equipment 2 to the user; such as a touch screen, a liquid crystal display or a projection device, for presenting a graphical user interface (GUI), the real-time operating status of the electromechanical equipment 2, visual information related to the operation, and possible risk warnings to the user; a speaker 13, for broadcasting voice feedback and system prompts to the user; broadcasting voice feedback to the user, such as operation confirmation, system prompt sound or detailed voice alarm information, constitutes a complete audio-visual interactive experience;
[0045] The electromechanical device 2 is configured to perform a predetermined operation according to the control instruction; the electromechanical device 2 includes a receiving terminal 21, a control unit 22, and an actuator 23, wherein the receiving terminal 21 is electrically connected to the control unit 22, and the control unit 22 is electrically connected to the actuator 23; the receiving terminal 21 of the electromechanical device 2 is configured to receive the control instruction from the user interaction device 1; the control unit 22 of the electromechanical device 2 includes a programmable logic controller or a microcontroller unit, and the control unit 22 is configured to parse the control instruction and drive the actuator 23 of the electromechanical device 2 accordingly; the actuator 23 of the electromechanical device 2 includes a motor, a valve, a heater, or a robotic arm, and the actuator 23 is configured to perform a physical operation corresponding to the control instruction; the communication connection between the user interaction device 1 and the receiving terminal 21 of the electromechanical device 2 is a wireless communication connection; the user interaction device 1 is communicatively connected to the receiving terminal 21 of the electromechanical device 2;
[0046] Specifically, the receiving terminal 21 is responsible for receiving control instructions from the user interaction device 1 or other designated sources. The communication connection between the receiving terminal 21 and the user interaction device 1 can be a wired connection, such as industrial Ethernet, or a more flexible wireless communication connection, such as Wi-Fi, Bluetooth, LoRa, or 5G and other industrial wireless networks. The control unit 22 is electrically connected to the receiving terminal 21, and usually includes a programmable logic controller (PLC) or a microcontroller unit (MCU). Its core function is to parse the received control instructions and drive the actuator 23 to act according to the instruction logic. The actuator 23 is electrically connected to the control unit 22 and is the component that actually performs physical operations. Specifically, it can include at least one of a motor (driving rotation or linear motion), a valve (controlling fluid on / off or flow), a heater (controlling temperature), or a multi-degree-of-freedom robotic arm (performing complex operation tasks).
[0047] An edge cognito 3 is coupled to the electromechanical device 2 , and the edge cognito 3 is used to collect multimodal sensor data of the electromechanical device 2 and generate device insight information based on the multimodal sensor data;
[0048] The edge cognitive body 3 includes: a multimodal sensor array 31 coupled to the electromechanical device 2, for collecting the multimodal sensor data, wherein the multimodal sensor data includes vibration signals, acoustic signals, or thermal imaging data; an edge intelligent processor 32 connected to the multimodal sensor array 31, for learning a digital representation of the state of the electromechanical device 2 from the multimodal sensor data, detecting an operational anomaly of the electromechanical device 2 based on the digital representation, and evaluating the health index of the electromechanical device 2 or predicting its remaining useful life, and using the operational anomaly, the health index, or the remaining useful life as part of the device insight information;
[0049] Specifically, the multimodal sensor array 31 is directly or indirectly installed at the key parts of the electromechanical equipment 2, and is used to comprehensively collect multimodal sensor data that reflects the status of the equipment. These data types may include but are not limited to vibration signals, acoustic signals, or thermal imaging data. The edge intelligent processor 32 is connected to the multimodal sensor array 31 and has the ability to perform intelligent processing near the source of data generation. The processor is used to learn a compact and information-rich digital representation of the state of the electromechanical equipment 2 from the massive, original multimodal sensor data through advanced algorithms. Based on these digital representations, the edge intelligent processor 32 can detect operational anomalies of the electromechanical equipment 2 in real time, such as sudden signs of failure or subtle changes from normal operating conditions. Furthermore, it can also evaluate the current health index of the electromechanical equipment 2, or predict its remaining useful life (RUL) using a predictive model.
[0050] The predictive active safety interaction engine 4 is in communication with the user interaction device 1 and the edge cognitive entity 3, and is configured to process user voice commands to determine user intent, and generate control commands for the electromechanical device 2 based on the user intent and the device insight information;
[0051] The predictive active safety interaction engine 4 includes: a voice instruction processing module 41, which is used to convert the user voice instruction into text through automatic speech recognition, and analyze the text through natural language understanding to determine the user intention; a risk assessment module 42, which is connected to the voice instruction processing module 41 and the edge cognitive body 3, and the risk assessment module 42 is used to perform a predictive risk assessment on the operation corresponding to the user intention based on the user intention and the device insight information; a decision and instruction generation module 43 is connected to the risk assessment module 42, and is used to generate the control instruction when the result of the predictive risk assessment indicates that the risk is lower than a preset threshold; an alarm prompt module 44, which generates a safety warning for the user to perceive, or adjusts the control instruction to reduce the risk when the result of the predictive risk assessment indicates that the risk exceeds the preset threshold;
[0052] Specifically, the voice command processing module 41 receives voice commands from the user interaction device 1, converts them into text using automatic speech recognition (ASR) technology, and then analyzes the text using natural language understanding (NLU) technology to accurately grasp the user's true intent. The risk assessment module 42 is key to achieving "active safety." It receives user intent from the voice command processing module 41 and device insights from the edge cognition 3. Based on this information, combined with pre-set safety rules and a knowledge base, it performs a predictive assessment of the potential risk of executing the operation corresponding to the user's intent under the current device state. The decision and instruction generation module 43 is connected to the risk assessment module 42. When the predictive risk assessment results indicate that the potential risk is below a preset safety threshold, this module generates control instructions for the electromechanical device 2. The alarm prompt module 44 intervenes when the predictive risk assessment results indicate that the risk exceeds the preset threshold. It can generate a safety warning for the user to perceive or, in some cases, proactively adjust the parameters of the original control instruction to generate a new one with lower risk.
[0053] See also Figure 6 The present invention also provides a method for remotely controlling electromechanical equipment based on artificial intelligence voice interaction, the method comprising the following steps:
[0054] S1. Collecting user voice commands through the user interaction device 1, and processing the user voice commands by the predictive active safety interaction engine 4 using an automatic speech recognition model and a natural language understanding model based on machine learning to determine the user's intention;
[0055] In this embodiment, the core purpose of the steps of user voice command processing and intention determination is to accurately convert the operation instructions issued by the user through natural voice into machine-understandable, structured user intentions, laying the foundation for subsequent risk assessment and instruction execution.
[0056] Generally, this step starts with the user issuing a voice command through the microphone 11 of the user interaction device 1. The microphone 11 is responsible for collecting the original acoustic waveform issued by the user, which is the physical carrier of the user's voice command and is recorded as the original audio signal V in this article. audio The user interaction device 1 collects the original audio signal V audio After preliminary digital processing, it is safely and reliably transmitted to the voice command processing module 41 inside the predictive active safety interaction engine 4 for subsequent in-depth analysis.
[0057] In the voice instruction processing module 41, the automatic speech recognition (ASR) model first processes the received original audio signal V audio Specifically, the automatic speech recognition model M ASR The purpose is to map the time domain audio signal sequence into the corresponding text character sequence. ASR It is usually built based on machine learning, especially deep learning technology, for example, it can adopt a network structure including recurrent neural network (RNN), long short-term memory network (LSTM), gated recurrent unit (GRU) or more advanced Transformer, and combined with a sequence-to-sequence conversion framework such as connectionist temporal classification (CTC) loss function or attention mechanism. ASR The performance of depends on a set of learned parameters within it, denoted as θ ASR , these parameters θ ASR It is obtained by supervised learning training on a dataset containing a large number of diverse speech samples and their corresponding text annotations. The function of the automatic speech recognition model can be conceptually expressed as: asr =M ASR (V audio θ ASR );
[0058] Among them, M ASR represents the function of the automatic speech recognition model, and θ ASR represents the learned parameter set of the automatic speech recognition model; and T asr It represents the text sequence output by the model that corresponds to the input speech content. This text sequence T asr It is the basis for subsequent natural language understanding processing.
[0059] Get the text sequence T asr Afterwards, the text sequence will be passed as input to the natural language understanding (NLU) model in the voice instruction processing module 41. NLU The core task is to asr Perform in-depth semantic analysis and key information extraction to accurately identify and parse the user's true intention in the voice command. Specifically, the natural language understanding model M NLU A series of processes such as lexical analysis, syntactic analysis, named entity recognition, relationship extraction and intent classification will be performed. Similar to the automatic speech recognition model, the natural language understanding model M NLU It is usually based on machine learning technology, such as a model based on the Transformer architecture, or a combination of traditional rule-based systems and statistical models. NLU The performance of also depends on a set of learned parameters within it, denoted as θ NLU , these parameters θ NLU It is obtained by training on a dataset built for specific application scenarios, which contains a large number of text instructions and their corresponding structured intent annotations.
[0060] Through the processing of natural language understanding model, the user's intention I u is further structured into a category C containing explicit intentions I and a set of related parameters or slot information P slots The intention category C I Indicates the type of operation the user wants to perform, such as "start device", "stop device", "adjust parameters", "query status", etc. The parameter or slot information P slots The specific details required or relevant to execute the intent category are provided. For example, if the intent category is "adjust parameters", the slot information may include "device number", "parameter name" and "target parameter value"; if the intent category is "query status", the slot information may include "device number" and "queried indicator name". u The determination process is expressed as: u ={C I ,P slots}=M NLU (T asr θ NLU );
[0061] Among them, M NLU represents the function of the natural language understanding model, θ NLURepresents the learned parameter set of the natural language understanding model, and the final determined structured user intention I u ; and I u Represents the finalized, including intent category C I and parameter or slot information P slots Structured user intent.
[0062] In a possible implementation, the parameter or slot information P slots It can be a set of key-value pairs, for example: slots ={(slot1,value1),(slot2,value2),…,(slot n ,value n )};
[0063] Among them, slot i Represents the name of the i-th parameter, value i Represents its corresponding value.
[0064] Finally, the structured user intention I determined by this step u , in its clear and unambiguous form, provides the subsequent modules of the predictive active safety interaction engine 4 with the necessary input information for analysis, judgment and decision-making, thereby ensuring that the entire remote control process can accurately respond to user requests.
[0065] S2. Collecting multimodal sensor data from the electromechanical device 2 through the edge cognito 3, and processing the multimodal sensor data using a deep learning-based digital representation learning algorithm and anomaly detection algorithm, combined with a device health status prediction model, to generate device insight information.
[0066] In this embodiment, the core purpose of the steps of multimodal data perception and insight generation of electromechanical equipment is to conduct real-time and comprehensive monitoring and analysis of the operating status of the electromechanical equipment 2 through the edge cognitive body 3, and to extract equipment insight information that can characterize the health status of the equipment, potential anomalies and future trends, thereby providing key data support for subsequent predictive risk assessment and intelligent decision-making.
[0067] Typically, this step is performed by an edge cognito 3 tightly coupled to the electromechanical device 2. The edge cognito 3 continuously collects multiple types of sensor data from different components or aspects of the electromechanical device 2 through its configured multimodal sensor array 31. This multimodal sensor data may include, but is not limited to, vibration signals to reflect the smoothness, imbalance, looseness, or wear of mechanical components; acoustic signals to capture various sounds generated by the device during operation, thereby identifying auditory features such as unusual noises and changes in noise patterns that may indicate abnormalities; and thermal imaging data to monitor the temperature distribution in key areas of the device, promptly identifying hot spots or abnormal temperature rises, indicating potential electrical failures or excessive friction.
[0068] After collecting the raw multimodal sensor data, the edge intelligent processor 32 inside the edge cognitive body 3 performs a series of preprocessing operations on the data. The purpose of preprocessing is to improve the quality of the data and make it more suitable for subsequent complex model analysis. Specifically, preprocessing may include filtering the signal to remove noise interference, aligning the timestamps of data collected by different sensors to ensure synchronization, normalizing or standardizing the data to eliminate the influence of different dimensions, and extracting preliminary statistical features or time-frequency domain features. After preprocessing, the data obtained is recorded as preprocessed multimodal sensor data X in this article. proc_msa (t), where t represents the timestamp of data acquisition or the discrete time sampling point, which represents the state of the data at a specific moment.
[0069] Next, the pre-processed multimodal sensor data X proc_msa (t) is input to the core module of the deep learning-based digital representation learning algorithm deployed in the edge intelligent processor 32. This module is usually embodied as a pre-trained or online learned encoder function, denoted as f DRL_enc The encoder f DRL_enc The design goal is to extract the preprocessed multimodal sensor data X from the preprocessed multimodal sensor data X which may have high dimensions, complex nonlinear relationships and redundant information. proc_msa (t), autonomously learn and extract a low-dimensional, information-intensive digital representation that can effectively and comprehensively reflect the key operating state of the electromechanical device 2 at time t, which is denoted as Z in this article. state (t). This digital representation learning process aims to map the raw sensor data into a latent feature space that is easier to analyze and understand. This learning and transformation process is represented as: state (t) = f DRL_enc (X proc_msa (t);θ DRL_enc );
[0070] Among them, θDRL_enc For the encoder f DRL_enc The learned model parameters of f DRL_enc Represents the encoder function in the digital representation learning algorithm, which can be implemented as a deep neural network structure, such as a convolutional autoencoder, a variational autoencoder, or an autoencoder based on a recurrent neural network (RNN); θ DRL_enc represents the encoder f DRL_enc The set of learned parameters of the model is usually learned by optimizing a specific objective function (such as minimizing the reconstruction error, maximizing the mutual information between data and representation, etc.) on a large amount of historical normal operation data or data containing specific failure modes; state (t) is the model output, which is a digital representation of the state of the electromechanical device 2 at time t.
[0071] After obtaining the digital representation Z capable of representing the device state state (t) Afterwards, the edge intelligent processor 32 uses an anomaly detection algorithm to evaluate whether the current operating state of the electromechanical device 2 deviates from its normal operating mode, and calculates a quantitative anomaly score S accordingly. anomaly (t). The abnormal score S anomaly The calculation of (t) can be based on various methods:
[0072] In one possible implementation, the anomaly detection algorithm can act directly on the digital representation Z state (t). For example, statistical methods (such as determining the probability density of the new representation based on a Gaussian mixture model) or machine learning methods (such as a single-class support vector machine or isolation forest) can be used to represent the digital representation Z state (t) The boundary of the normal operating mode is constructed in the multi-dimensional feature space formed. When the new state representation falls outside this boundary or deviates too much from the center of the normal mode, it is judged as abnormal and a corresponding abnormality score is given.
[0073] As another option, especially when the digital representation learning algorithm adopts a structure including an encoder and a decoder such as an autoencoder, anomaly detection can be based on the reconstruction error. Specifically, there is a DRL_enc The corresponding decoder function is denoted as f DRL_dec , whose model parameter is θ DRL_dec The decoder f DRL_dec The function is to try to represent Z from the compressed digital state (t) reconstructs an approximate version of the original input data, denoted as This reconstruction process is expressed as:
[0074] Among them, θ DRL_dec For the decoder f DRL_dec Model parameters; Subsequently, the preprocessed multimodal sensor data X can be calculated proc_msa (t) Instead of reconstructing data The difference or distance between them is used to obtain the quantitative anomaly score S anomaly (t). Generally speaking, when the device is operating in a normal state, its data can be well reconstructed by the model, and the reconstruction error is small; however, when the device is abnormal, its data pattern deviates from the normal state, and the model is difficult to accurately reconstruct, resulting in an increase in the reconstruction error, thus the anomaly score S anomaly By comparing this anomaly score with the preset threshold, it can be determined whether the device has an operational anomaly.
[0075] In addition, in order to more comprehensively evaluate the long-term health status of the equipment and predict its future trends, the edge intelligent processor 32 also uses an equipment health status prediction model. This model is usually based on the digital representation Z accumulated and formed over a period of time. state (t) time series data. For example, a deep learning model based on a recurrent neural network (RNN), a long short-term memory network (LSTM), a gated recurrent unit (GRU), or a temporal convolutional network (TCN) that can process sequence data can be used to model the equipment degradation trend reflected in the digital representation sequence. Through this model, a quantitative health index H(t) of the electromechanical equipment 2 at the current time t can be evaluated. The health index H(t) is usually a normalized value (for example, between 0 and 1, 1 indicates complete health and 0 indicates complete failure), or its remaining useful life RUL(t) under the current working conditions and wear state can be predicted, that is, the time the equipment can continue to operate reliably before requiring major maintenance or reaching the end of its service life.
[0076] Finally, through the above series of processing procedures, the edge cognitive body 3 detects the abnormal operation (for example, by comparing the abnormality score S anomaly (t) and the abnormal event or state obtained by the threshold), the evaluated equipment health index H(t), or the predicted equipment remaining useful life RUL(t), or any combination of these information, are integrated into structured equipment insight information. This equipment insight information D insight It will then be output as the key basis for subsequent predictive active safety interactive engine 4 to conduct risk assessment and decision-making.
[0077] S3, the predictive active safety interaction engine 4 combines the user intention with the device insight information, performs a risk assessment using a predictive risk assessment model, and generates an operation instruction based on the result of the risk assessment, the operation instruction being a control instruction for the electromechanical device 2 or a safety warning perceived by the user;
[0078] In this embodiment, the steps of combining user intention with device insight information to perform predictive risk assessment and generate operation instructions accordingly, the core purpose of which is to make a forward-looking judgment on the feasibility and safety of executing user intention under the current state of electromechanical equipment 2 through the predictive active safety interaction engine 4, and generate a final operation instruction based on this judgment that can both meet user needs and ensure operational safety.
[0079] Generally, this step is completed by the risk assessment module 42, the decision and instruction generation module 43, and the alarm prompt module 44 within the predictive active safety interaction engine 4. The risk assessment module 42 first integrates two key information as the basis for evaluation:
[0080] The first one is the structured user intention output from the aforementioned user voice command processing and intention determination step, denoted as I u , the user's intention I u It specifies the type of operation the user wishes to perform and its related parameters;
[0081] The second is the real-time device insight information output from the aforementioned electromechanical equipment multimodal data perception and insight generation steps, denoted as D insight , the device insights information D insight It reflects the current operating status, health status or potential abnormality of the electromechanical equipment (2).
[0082] After obtaining and integrating the user intention I u and the device insight information D insight Afterwards, the risk assessment module 42 will further combine a pre-built or dynamically learned electromechanical equipment operation safety regulations and knowledge base, denoted as K safety The knowledge base K safety The data store contains information related to specific electromechanical equipment2, including safety operating specifications, known operating contraindications, equipment design limit parameters, risk patterns derived from historical failure cases, industry safety standards, and safe operating guidelines for specific operating conditions. This information provides important context and constraints for risk assessment.
[0083] Then, the risk assessment module 42 uses the above integrated user intention I u , device insight information D insight and safety regulations and knowledge base K safety, through its internally deployed predictive risk assessment model, denoted as M PRA , to execute the current user intention I u The potential risk level that may be associated with the corresponding operation is quantitatively or qualitatively evaluated, and the evaluation result is recorded as R pred The predictive risk assessment model M PRA The specific implementation methods can be varied. For example, a system based on expert rules can be used to convert K safety The security rules in the code are encoded as logical judgment conditions; machine learning-based methods can also be used, such as training a classifier (such as a decision tree, support vector machine or deep neural network) or regressor to enable it to u and D insight Predict the corresponding risk level R pred In more complex implementations, models based on Bayesian networks or causal reasoning can also be used to analyze the interactions between various factors and assess risks. This risk assessment process can be conceptually expressed as: pred =M PRA (I u ,D insight ,K safety θ PRA );
[0084] Among them, θ PRA Represents the predictive risk assessment model M PRA The potential risk level R is the set of learned or configured parameters (for machine learning models) or the preset rule parameters and logical structure (for rule-based systems). pred The output can be a continuous risk score or a discrete risk level, such as "low risk," "medium risk," or "high risk." The purpose of this predictive risk assessment is to proactively identify and quantify potential operational risks before actual operational instructions are issued to the electromechanical equipment 2, thereby avoiding safety incidents or equipment damage caused by poor equipment status or improper user instructions.
[0085] After obtaining the predictive risk assessment model M PRA Output potential risk level R pred Afterwards, the decision and instruction generation module 43 and the alarm prompt module 44 in the predictive active safety interaction engine 4 will make intelligent decisions based on the risk level to determine the final generated operation instructions. Specifically, the calculated potential risk level R pred and a pre-set security risk threshold τ that can be adjusted according to specific application scenarios risk Make a comparison.
[0086] If the comparison results show that the potential risk level R pred Below the security risk threshold τ risk (ie R pred <τ risk ), it indicates that the risk of executing the operation corresponding to the user intention in the current device state is within an acceptable range, or the operation is safe. In this case, the decision and instruction generation module 43 will convert the original user intention I u Directly converted into a control instruction for the electromechanical device 2, denoted as C cmd The control instruction C cmd The control unit 22 of the electromechanical device 2 will be clearly instructed on how to drive its actuator 23 to complete the operation requested by the user. cmd This constitutes the final output operation instruction of this step.
[0087] On the contrary, if the comparison result shows that the potential risk level R pred Equal to or exceeding the security risk threshold τ risk (ie R pred ≥τ risk ), it indicates that executing the operation corresponding to the user's intention under the current device state has a high risk and may cause damage to the device, threaten the safety of the operator, or affect the normal operation of production. In this case, the alarm prompt module 44 or the decision and instruction generation module 43 will take one or more of the following measures:
[0088] As an option, the system will generate a security warning for the user to perceive, denoted as A sfty . This safety warning A sfty The user interaction device 1 will clearly point out the risks of the current operation to the user and may explain the reasons for the risks (for example, the device temperature is too high, the load exceeds the safety range, or it matches a known dangerous operation mode, etc.). sfty This constitutes the final output of the operation instructions in this step, which aims to prevent the execution of high-risk operations and enhance the user's security awareness.
[0089] In another possible implementation, especially when the system has a higher level of intelligent decision-making capabilities, the decision and instruction generation module 43 or the alarm prompt module 44 may try to proactively respond to the original user intention without completely violating the user's core intention. uThe system adjusts the operating parameters in the system to find an alternative operation plan with lower risk. For example, if the user requests to start the device at a higher speed, but the system assesses that the speed risk is too high under the current device state, the system may automatically adjust the startup speed to a relatively lower value that is still within the safe range. After such adjustments, a new control instruction with effectively reduced risk is generated, which is recorded as C′. cmd At this time, the adjusted control instruction C' cmd The final output of this step is the operation instruction. This proactive adjustment capability reflects the system's "active safety" feature, which means that it not only passively rejects dangerous instructions, but also attempts to guide users to perform safer operations.
[0090] In summary, this step integrates user intent, real-time device insights, and security knowledge to conduct a forward-looking risk assessment. Based on the assessment results, it intelligently decides whether to directly execute the user instruction, issue a security warning, or optimize the instruction for security. Finally, it outputs a security-considered operation instruction (the operation instruction may be C cmd 、A sfty or C′ cmd The operation instruction will then be passed to the subsequent steps for processing to ensure the safety and reliability of the entire electromechanical equipment remote control process.
[0091] S4. Perform subsequent processing according to the operation instruction: if the operation instruction is the control instruction, transmit it to the control unit 22 of the electromechanical device 2 to drive the actuator 23 to perform a predetermined operation; if the operation instruction is the safety warning, present the safety warning to the user through the user interaction device 1;
[0092] In this embodiment, step S4, i.e., the step of collecting user feedback and updating or adjusting the parameters of the artificial intelligence model, has the core purpose of building a closed-loop learning and optimization mechanism, so that the device described in the present invention can continuously improve the performance of its internal artificial intelligence model based on the actual usage feedback of the user, thereby enhancing the adaptability, accuracy and user experience of the system.
[0093] Generally, this step begins with collecting user feedback on specific information generated by the system in the previous step through the user interaction device 1. Specifically, the object of user feedback can be the device insight information generated by the edge cognitive agent 3 in step S2, or the feedback provided on the operation instruction generated by the predictive active safety interaction engine 4 in step S3. This user feedback is denoted as F in this article. u .
[0094] The user feedback F uThe form of F can be various. In one possible implementation, it can be an explicit evaluation actively provided by the user through the interface of the user interaction device 1 (such as a touch screen, button) or voice channel, such as a score on the accuracy of a certain voice recognition, a judgment on the accuracy of a certain device insight information, or direct feedback on the rationality of a certain risk assessment result. As another option, the user feedback F u Alternatively, implicit preferences or evaluations can be indirectly inferred by analyzing the user's subsequent behavioral sequences during interaction with the system. For example, if a user immediately stops the relevant operation after the system issues a security warning, this could be seen as an implicit recognition of the warning's effectiveness. Conversely, if a user repeatedly attempts an operation that the system blocks, this could indicate that the user disagrees with the risk assessment results or that the system fails to accurately understand the user's true and safe needs.
[0095] After collecting the user feedback F u Based on this feedback, the system then updates parameters or adjusts the structure of one or more core AI models deployed within the device, aiming to optimize their performance when performing similar tasks in the future. This update process is a key step in enabling the autonomous evolution of the system's intelligence level. Specifically, this update process can include optimizing relevant models deployed within the edge cognito 3, or optimizing relevant models deployed within the predictive active safety interaction engine 4, or both.
[0096] When the update object is the model inside the edge cognitive body 3, the collected user feedback F u It can be used to guide the adjustment of relevant model parameters. Specifically, it may involve:
[0097] Adjust the encoder f in the digital representation learning algorithm DRL_enc The model parameters θ DRL_enc , or its corresponding decoder f DRL_dec The model parameters θ DRL_dec , so that the learned digital representation of the device status can more accurately reflect the actual condition of the device, or improve the distinguishability between normal and abnormal states.
[0098] Adjust the relevant parameters of the anomaly detection algorithm, such as adjusting the judgment threshold of the anomaly score, or updating the classifier parameters used to distinguish normal from abnormal patterns, to improve the precision and recall rate of anomaly detection.
[0099] Adjust the parameters of the equipment health status prediction model, for example, by using the actual equipment failure time or maintenance records provided by the user to modify the parameters of the health index assessment model or the remaining useful life RUL(t) prediction model to improve its prediction accuracy. The update process of the internal model of the edge cognitive body 3 is expressed as: θ′EC =Φ EC (θ EC ,F u );
[0100] Among them, θ EC Represents the edge cognitive model parameter set to be updated, θ′ EC represents the updated parameter set, Φ EC Represents a preset update function or learning algorithm, which is responsible for updating the model according to the current model parameters θ EC and received user feedback F u , calculate the updated model parameters. For example, Φ EC It can be an online learning algorithm based on gradient descent, a policy update rule based on reinforcement learning, or a method based on Bayesian update. EC Representatives after user feedback F u After optimization, the edge cognitive body 3 has a new set of parameters for the internal correlation model.
[0101] When the update object is the model inside the predictive active safety interaction engine 4, the collected user feedback F u It can also be used to guide the adjustment of relevant model parameters. Specifically, it may involve: adjusting the model parameters θ of the automatic speech recognition (ASR) model ASR , for example, using the correct text transcription provided by the user to fine-tune the acoustic model or language model to improve the recognition accuracy under specific user accents, specific environmental noise or specific domain terms. Adjust the model parameters θ of the natural language understanding (NLU) model NLU For example, based on the user's confirmation or correction of the intention understanding result, the performance of the intention classifier or slot filling model is optimized so that it can more accurately grasp the user's true operation intention. PRA The model parameters θ PRA For example, if user feedback indicates that the system issues unnecessary warnings too frequently (too conservative) or fails to identify certain actual risks (too aggressive), this feedback can be used to adjust the weights, thresholds, or decision logic in the risk assessment model to make its risk judgment more in line with actual needs and safety standards.
[0102] The updating process of the internal model of the predictive active safety interaction engine 4 is expressed as: θ′ PPSIE =Φ PPSIE (θ PPSIE ,F u );
[0103] Among them, θ PPSIE Represents the parameter set of the predictive active safety interaction engine model to be updated, θ′PPSIE represents the updated parameter set, Φ PPSIE Represents the update function for the predictive active safety interaction engine model.
[0104] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A remote control device for electromechanical equipment based on artificial intelligence voice interaction, characterized in that: include: A user interaction device (1) for collecting user voice control instructions and presenting information to the user; an electromechanical device (2) for performing a predetermined operation according to the control instruction; The electromechanical device (2) comprises a receiving terminal (21), a control unit (22) and an actuator (23), wherein the receiving terminal (21) is electrically connected to the control unit (22), and the control unit (22) is electrically connected to the actuator (23); The user interaction device (1) is communicatively connected to a receiving terminal (21) of the electromechanical device (2); An edge cognitive body (3) is coupled to the electromechanical device (2), and the edge cognitive body (3) is used to collect multimodal sensor data of the electromechanical device (2) and generate device insight information based on the multimodal sensor data; A predictive active safety interaction engine (4) is in communication with the user interaction device (1) and the edge cognitive body (3), and is used to process user voice instructions to determine user intentions, and to generate control instructions for the electromechanical device (2) in combination with the user intentions and the device insight information.
2. The electromechanical equipment remote control device based on artificial intelligence voice interaction according to claim 1 is characterized in that: The user interaction device (1) comprises: A microphone (11) for collecting user voice commands; a display unit (12) for presenting a graphical user interface and visual information related to the operation of the electromechanical device (2) to a user; The speaker (13) is used to broadcast voice feedback and system prompts to the user.
3. The electromechanical equipment remote control device based on artificial intelligence voice interaction according to claim 1 is characterized in that: The receiving terminal (21) of the electromechanical device (2) is used to receive the control instruction from the user interaction device (1); The control unit (22) of the electromechanical device (2) includes a programmable logic controller or a microcontroller unit, and the control unit (22) is used to parse the control instructions and drive the actuator (23) of the electromechanical device (2) accordingly; The actuator (23) of the electromechanical device (2) includes a motor, a valve, a heater or a robotic arm, and the actuator (23) is used to perform a physical operation corresponding to the control instruction; The communication connection between the user interaction device (1) and the receiving terminal (21) of the electromechanical device (2) is a wireless communication connection.
4. The electromechanical equipment remote control device based on artificial intelligence voice interaction according to claim 1 is characterized in that: The edge cognitive body (3) includes: A multimodal sensor array (31) is coupled to the electromechanical device (2) for collecting the multimodal sensor data, wherein the multimodal sensor data includes vibration signals, acoustic signals or thermal imaging data; An edge intelligent processor (32) is connected to the multimodal sensor array (31) and is used to learn a digital representation of the state of the electromechanical device (2) from the multimodal sensor data, detect an operational anomaly of the electromechanical device (2) based on the digital representation, and evaluate the health index of the electromechanical device (2) or predict its remaining effective life, and use the operational anomaly, the health index or the remaining effective life as a component of the device insight information.
5. The electromechanical equipment remote control device based on artificial intelligence voice interaction according to claim 1 is characterized in that: The predictive active safety interaction engine (4) includes: A voice instruction processing module (41) is used to convert the user's voice instruction into text through automatic speech recognition, and analyze the text through natural language understanding to determine the user's intention; a risk assessment module (42) connected to the voice instruction processing module (41) and the edge cognitive body (3), the risk assessment module (42) being configured to perform a predictive risk assessment on an operation corresponding to the user intention based on the user intention and the device insight information; A decision and instruction generation module (43) is connected to the risk assessment module (42) and is used to generate the control instruction when the result of the predictive risk assessment indicates that the risk is lower than a preset threshold; An alarm prompt module (44) generates a safety warning for the user to perceive, or adjusts the control instruction to reduce the risk when the result of the predictive risk assessment indicates that the risk exceeds the preset threshold.
6. A method for remotely controlling electromechanical equipment based on artificial intelligence voice interaction, applied to the device according to any one of claims 1 to 5, characterized in that: The method comprises the following steps: S1. Collecting user voice commands through the user interaction device (1), and processing the user voice commands by the predictive active safety interaction engine (4) using an automatic speech recognition model and a natural language understanding model based on machine learning, thereby determining the user's intention; S2, collecting multimodal sensor data of the electromechanical device (2) through the edge cognitive body (3), and processing the multimodal sensor data by the edge cognitive body (3) using a digital representation learning algorithm and anomaly detection algorithm based on deep learning, and combining the multimodal sensor data with a device health status prediction model, thereby generating device insight information; S3, the predictive active safety interaction engine (4) combines the user intention with the device insight information, performs risk assessment through a predictive risk assessment model, and generates an operation instruction based on the result of the risk assessment, the operation instruction being a control instruction for the electromechanical device (2) or a safety warning for the user to perceive; S4. Perform subsequent processing according to the operation instruction: If the operation instruction is the control instruction, transmitting it to the control unit (22) of the electromechanical device (2) to drive the actuator (23) to perform a predetermined operation; If the operation instruction is the safety warning, the safety warning is presented to the user via the user interaction device (1).
7. The method for remotely controlling electromechanical equipment based on artificial intelligence voice interaction according to claim 6, characterized in that: The step S1 comprises: The user voice command collected by the user interaction device (1) is represented by the original audio signal V audio , is input into the automatic speech recognition model; the automatic speech recognition model uses the acoustic and language knowledge obtained through machine learning training to transform the original audio signal V audio Convert to the corresponding text sequence T asr , the conversion process is: T asr =M ASR (V audio ;θ ASR ); Among them, M ASR represents the function of the automatic speech recognition model, and θ ASR represents a set of learned parameters of the automatic speech recognition model; a text sequence T output by the automatic speech recognition model asr As input, passed to the natural language understanding model; The natural language understanding model processes the text sequence T asr Conduct in-depth semantic analysis and key information extraction to identify the user's specific intention; User Intention I u is further structured to include a clear intent category C I and a set of related parameters or slot information P slots The user intends to u The determination process is: I u ={C I ,P slots }=M NLU (T asr ;θ NLU ); Among them, M NLU represents the function of the natural language understanding model, θ NLU Represents the learned parameter set of the natural language understanding model, and the final determined structured user intention I u .
8. The method for remotely controlling electromechanical equipment based on artificial intelligence voice interaction according to claim 6, characterized in that: The step S2 comprises: The collected multimodal sensor data is preprocessed to obtain the preprocessed multimodal sensor data X proc_msa (t), where t is a timestamp; the pre-processed multimodal sensor data X proc_msa (t) Input to the encoder f of the digital representation learning algorithm DRL_enc , the encoder is used to learn and extract the digital representation Z of the state of the electromechanical device (2) at time t state (t), which is calculated as: Z state (t)=f DRL_enc (X proc_msa (t);θ DRL_enc ); Among them, θ DRL_enc For the encoder f DRL_enc The learned model parameters of The abnormality detection algorithm is used to calculate the abnormality score S of the electromechanical device (2) at time t. anomaly (t), thereby detecting the abnormal operation of the electromechanical device (2); the abnormality score S anomaly (t) is calculated based on either: Directly based on the digital representation Z state (t) performing evaluation, or based on the pre-processed multimodal sensor data X proc_msa (t) and the decoder f by the digital representation learning algorithm DRL_dec From the digital representation Z state (t) Reconstructed data The deviation between is evaluated, where the reconstructed data is: Among them, θ DRL_dec For the decoder f DRL_dec Model parameters; and, using the equipment health status prediction model, based on the digital representation Z state (t), evaluating the health index H(t) of the electromechanical equipment (2) or predicting its remaining useful life RUL(t); The operation abnormality, the health index H(t) or the remaining useful life RUL(t) is output as the device insight information.
9. The method for remotely controlling electromechanical equipment based on artificial intelligence voice interaction according to claim 6, characterized in that: The step S3 comprises: The predictive active safety interaction engine (4) is integrated into the determined user intention, denoted as I u , and the generated device insight information, denoted as D insight ; The user intention I u and the device insight information D insight , and combined with the preset electromechanical equipment operation safety regulations and knowledge base, recorded as K safety , are provided as input to the predictive risk assessment model; the core function of the predictive risk assessment model is denoted as M PRA , used to evaluate the execution and the user intention I u The potential risk level associated with the corresponding operation is denoted as R pred ; The risk assessment process is: R pred =M PRA (I u ,D insight ,K safety ;θ PRA ); Among them, θ PRA Represents the predictive risk assessment model M PRA The set of learned or configured parameters; Through the intelligent decision-making process, the calculated potential risk level M PRA With a preset security risk threshold τ risk Compare to determine the subsequent operation instructions: If the potential risk level R pred Below the security risk threshold τ risk , the intelligent decision-making process determines that the current operation is safe or has an acceptable risk, and generates a u Directly corresponding control instruction C cmd As the operating instruction; If the potential risk level R pred Equal to or exceeding the security risk threshold τ risk , the intelligent decision-making process determines that the current operation has a high risk and generates a safety warning A for the user to perceive. sfty As the operation instruction, or the intelligent decision-making process attempts to adjust the original user intention I u The operating parameters in the ts are used to reduce the risk and generate the adjusted, risk-reduced control instructions C′ cmd As the operation instruction; the control instruction C that is finally generated cmd 、Safety Warning A sfty , or the adjusted control instruction C′ cmd is the operation instruction.
10. The method for remotely controlling electromechanical equipment based on artificial intelligence voice interaction according to claim 6, characterized in that: The step S4 comprises: The user interaction device (1) collects user insight information generated by the device, or user feedback provided by the user for the generated operation instruction, and the user feedback is recorded as F u , the user feedback F u It can be the user's explicit evaluation of the system output, correction information, or implicit preferences reflected through the user's subsequent behavior; Based on the collected user feedback F u , updating or adjusting parameters of the artificial intelligence model in the device to optimize its future performance; the updating process specifically includes updating a model in the edge cognitive body (3), or updating a model in the predictive active safety interaction engine (4); When updating the model in the edge cognizer (3), it involves adjusting the encoder f of the digital representation learning algorithm. DRL_enc The model parameters θ DRL_enc , or its decoder f DRL_dec The model parameters θ DRL_dec , or the relevant parameters of the anomaly detection algorithm, or the parameters of the equipment health status prediction model, this update is expressed as: θ′ EC =Φ EC (i EC ,F u ); Among them, θ EC Represents the edge cognitive model parameter set to be updated, θ′ EC represents the updated parameter set, θ EC Represents the update function for the edge cognition model; When updating the model in the predictive active safety interaction engine (4), it involves adjusting the model parameters θ of the automatic speech recognition model. ASR , or the model parameters θ of the natural language understanding model NLU , or the model parameter θ of the predictive risk assessment model PRA , this update is expressed as: θ′ PPSIE =Φ PPSIE (i PPSIE ,F u ); Among them, θ PPSIE Represents the parameter set of the predictive active safety interaction engine model to be updated, θ′ PPSIE represents the updated parameter set, Φ PPSIE Represents the update function for the predictive active safety interaction engine model.
Citation Information
Cited By
Generative rainstorm and flood intelligent risk-avoiding dialogue method, system and equipment and medium
CN120708622A
Human-computer interaction method and system based on AI large model
CN121742373A