Fuzzy instruction analysis method and system based on dual-channel noise reduction and dynamic semantic map
By adopting multimodal industrial speech recognition and fuzzy instruction analysis methods with dual-channel noise reduction and dynamic semantic maps in industrial environments, the problem of insufficient high-noise and fuzzy instruction analysis in industrial environments is solved, high-accurate speech recognition and accurate fuzzy instruction analysis are achieved, and the operation efficiency and intelligence level of industrial scenarios are improved.
Patent Information
- Application Number
- CN202411927885.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-27
AI Technical Summary
The lack of high noise, complex working conditions and fuzzy command analysis capabilities in industrial environments have led to a significant decline in the recognition accuracy and resolution capabilities of existing voice control technologies in industrial scenarios, which cannot meet actual needs.
Multimodal industrial speech recognition and fuzzy instruction analysis methods based on dual-channel noise reduction and dynamic semantic maps are adopted. Through the fusion of dual-channel speech noise reduction technology and multimodal data, combined with deep learning models and graph neural networks, speech recognition and fuzzy instruction analysis in high-noise environments are realized.
Significantly improve speech recognition accuracy in high-noise environments, realize accurate analysis of fuzzy instructions, and enhance operational efficiency and intelligence level in industrial scenarios.
Smart Images

Figure CN120048269A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent speech parsing and multimodal data fusion, and particularly to a fuzzy instruction parsing method and system based on dual-channel noise reduction and dynamic semantic maps. Background Art
[0002] With the rapid development of Industry 4.0 and intelligent manufacturing technologies, the requirements for the intelligence and efficiency of operations in industrial scenarios are increasing day by day. Operations in traditional industrial environments mostly rely on manual methods such as touch screens, buttons, or keyboards, which are not only less efficient but also prone to operation errors and employee fatigue under high-load or complex working conditions. Especially in large industrial enterprises (such as steel mills, chemical plants), employees often need to perform corresponding operations on multiple devices. Facing heavy workloads, the limitations of traditional methods become more obvious.
[0003] In recent years, voice control, as a natural interaction method, has gradually received attention and demonstrated significant advantages in liberating hands and simplifying operation processes. However, the application of voice control technology in industrial scenarios still faces the following main challenges:
[0004] (1) Insufficient recognition accuracy in high-noise environments: Industrial environments are filled with various background noises, which can seriously affect the accuracy of speech recognition. Existing speech recognition technologies are mostly used in relatively quiet environments, and their recognition accuracy significantly decreases in complex industrial scenarios, unable to meet actual needs.
[0005] (2) Insufficient ability to parse fuzzy instructions: During actual use, employees often use fuzzy language to describe device positions to speed up the operation process, such as fuzzy instructions based on relative positions like "Switch to the device screen behind the rough rolling mill.", descriptive fuzzy instructions like "Check if there are any devices with abnormal temperatures currently.", and fuzzy instructions dependent on dynamic environments like "Adjust the humidity to the appropriate range." Existing voice control systems mostly rely on predefined instruction sets and cannot effectively parse such fuzzy commands, easily leading to operation delays or errors.
[0006] (3) Insufficient adaptability to multimodal fusion: Industrial environments are complex and variable, including dynamic factors such as light, temperature, humidity, and noise. Existing multimodal fusion technologies usually adopt fixed weight strategies and lack real-time dynamic adjustment capabilities, making it difficult to fully utilize the comprehensive advantages of voice, vision, and sensor data, resulting in unstable parsing effects.
[0007] To address the above problems, there is an urgent need for a voice parsing method that can adapt to complex industrial environments, which can not only achieve accurate speech recognition under high-noise conditions, but also parse user fuzzy instructions and optimize the multimodal fusion strategy according to the dynamic environment to improve the operation efficiency and intelligence level in industrial scenarios. Summary of the Invention
[0008] Based on the above background, the present invention proposes a multi-modal industrial speech recognition and fuzzy instruction parsing method and system based on dual-channel noise reduction and dynamic semantic maps. By combining noise reduction speech recognition, multi-modal data fusion, and fuzzy instruction parsing technologies, it can accurately parse industrial fuzzy instructions under high-noise and complex working conditions, and is applicable to various industrial scenarios such as production monitoring, equipment control, and operation navigation, providing support for the intelligent upgrade of industry.
[0009] The technical solution adopted by the present invention is as follows:
[0010] A fuzzy instruction parsing method based on dual-channel noise reduction and dynamic semantic maps, the method comprising the following steps:
[0011] (1) Noise reduction speech recognition: Combining dual-channel speech noise reduction and speech recognition, extract the target speech signal in a high-noise environment and complete speech-to-text conversion, and output the text data corresponding to the target speech signal;
[0012] (2) Multi-modal fusion: Based on real-time environment perception, dynamically adjust the weights of speech data, visual data, and sensor data;
[0013] (3) Fuzzy instruction parsing: Receive the text data output in step (1), and use the weights of the optimized speech data, visual data, and sensor data output in step (2) as the key inputs for updating the semantic map and parsing fuzzy instructions to achieve semantic map update; Use the dynamically updated semantic map combined with the device layout, environmental data, and user instruction history to locate the target device and parse the fuzzy instructions.
[0014] Further, step (1) specifically includes the following steps:
[0015] (1.1) Establish a voiceprint recognition channel: Extract the voiceprint features of the target user by using a deep learning model, combine a real-time comparison mechanism, output the voice instruction signal of the target user, and block the voice interference of non-target users;
[0016] (1.2) Establish a noise separation channel: Model the frequency characteristics of industrial noise through an autoencoder and a convolutional neural network CNN to obtain a noise reduction model; Use the noise reduction model to perform preliminary noise reduction on the input noisy audio, extract the speech signal and generate a noise mask;
[0017] (1.3) Dual-channel collaborative optimization: The recognition result output by the voiceprint recognition channel provides a reference signal for the noise separation channel; The clear speech signal extracted by the noise separation channel is fed back as input to the voiceprint recognition channel to optimize the accuracy of voiceprint matching;
[0018] Through the windowing and overlapping technique, combine the audio signals of the voiceprint recognition channel and the noise separation channel, and output a clear audio signal that conforms to the voiceprint of the target user, that is, output the target voice signal;
[0019] (1.4) Transmit the target voice signal output in step (1.3) to the speech recognition model to complete speech recognition, and output the text data corresponding to the target voice signal.
[0020] Further, step (1.1) to establish a voiceprint recognition channel specifically includes:
[0021] (1.11) Voiceprint feature extraction: Use spectrum analysis and Mel frequency cepstral coefficients to extract the voice features of all users, and combine convolutional neural networks and multi-layer residual networks for deep learning to construct a voiceprint feature library;
[0022] (1.12) Real-time voiceprint matching: When the user issues a voice command, collect the voice command signal in real time and extract the corresponding voiceprint features, and compare the voiceprint features corresponding to the voice command signal with the voiceprints in the voiceprint feature library to identify the target user;
[0023] (1.13) Mask non-target voice signals: According to the comparison result in step (1.12), mask the non-target voice signals below the preset threshold in the comparison result, and output the optimized voice command signal.
[0024] Further, step (2) specifically includes the following steps:
[0025] (2.1) Multimodal data input: Obtain three types of modal data: voice data, visual data, and environmental sensor data;
[0026] (2.2) Weight calculation: Use the Transformer multi-head attention mechanism to calculate the initial weights ω audio 、ω vision 、ω sensor ;
[0027] (2.3) Weight optimization: Through the DQN network combined with environmental data and reward mechanism, optimize the weight allocation strategy of the three types of modal data of voice data, visual data, and environmental sensor data for weight optimization.
[0028] Further, step (2.3) specifically includes:
[0029] Take the sensor data of the environment and the initial weights obtained in step (2.2) as the current environmental state vector S of the DQN network. The form of the current environmental state vector S is:
[0030] S = [Senv , S fusion = [Noise Level, Temperature, Light Intensity, …,
[0031] ω audio , ω vision , ω sensor
[0032] Among them, S env represents the sensor data of the current environment, including noise data, temperature data, and light intensity data; S fusion is the initial weights ω audio , ω vision , ω sensor of the three-modal data of speech data, visual data, and environmental sensor data calculated through step (2.2);
[0033] According to the current environmental state vector S and historical experience, the DQN network selects the historical strategy A based on past similar environmental states;
[0034] Then, the DQN network optimizes the weight combination strategy of the three-modal data of speech data, visual data, and environmental sensor data according to the current environmental state vector S and the historical strategy through the reward mechanism.
[0035] Furthermore, the weight selection formula of the DQN network is:
[0036]
[0037] Among them, Q(S, A) represents the value of taking the historical strategy under the current environmental state vector S, r is the reward, α is the learning rate, and γ is the discount factor; represents the maximum Q value among all decisions A′ in the next state S′;
[0038] According to the current environmental state vector S, select the strategy with the highest Q value, that is, the optimal strategy A*:
[0039]
[0040] According to the optimal strategy A * , adjust the weights of the three-modal data of speech data, visual data, and environmental sensor data to obtain the optimized modal weights
[0041] Furthermore, step (3) includes the following steps:
[0042] (3.1) Build a semantic map using a graph neural network (GNN), dynamically adjust the weights of device nodes by combining real-time sensor data, and optimize node associations based on the environmental state;
[0043] (3.2) Update the priority of device areas based on the historical frequency of user instructions, and combine the optimized modal weights output in step (2) to achieve semantic map update;
[0044] (3.3) Use the all-MiniLM-L6-v2 model to convert the text data output in step (1) into high-dimensional semantic vectors, parse the instructions through a graph propagation mechanism, complete the positioning of the target device or target area, and output the parsing result.
[0045] Further, the method for semantic map update in step (3.2) is as follows:
[0046] Represent each device area as a node of the graph, and the edges between nodes represent the spatial relationship or operation dependency relationship of device areas; the update process is described as:
[0047]
[0048] where, represents the weight of the i-th device at time t + 1; j represents the j-th device within the neighborhood of the i-th device, represents the weight of the j-th device at time t; represents the cumulative weight of the nodes within the neighborhood of node i, reflecting the cooperation relationship between the corresponding device and the neighboring areas; η is the learning rate, used to control the step size or speed of node weight update; λ is the neighborhood influence coefficient, used to control the influence degree of neighborhood nodes on the weight of the target node; Freq(i) is the historical frequency of the instructions of the i-th device, and Sensor(i) is the contribution value of the sensor state.
[0049] Further, the calculation formula of Sensor(i) is:
[0050] Sensor(i) = Sensor audio (i) + Sensor vision (i) + Sensor sensor (i);
[0051]
[0052] where, Sensor raw_audio (i), Sensor raw_vision (i) and Sensor raw_sensor (i) are the original sensor data from three modalities of voice data, visual data, and environmental sensor data, which have been normalized and standardized; is the optimized modal weight obtained from (2).
[0053] A fuzzy instruction parsing system based on dual-channel noise reduction and dynamic semantic map, the system includes the following modules:
[0054] (1) Noise reduction speech recognition module: Combining dual-channel speech noise reduction and speech recognition, extract the target speech signal in a high-noise environment and complete speech-to-text conversion, outputting the text data corresponding to the target speech signal;
[0055] (2) Multimodal fusion module: Based on real-time environment perception, dynamically adjust the weights of speech data, visual data, and sensor data;
[0056] (3) Fuzzy instruction parsing module: Receive the text data output by the noise reduction speech recognition module, use the optimized weights of speech data, visual data, and sensor data output by the multimodal fusion module as the key inputs for updating the semantic map and parsing fuzzy instructions to achieve semantic map update; Use the dynamically updated semantic map combined with device layout, environmental data, and user instruction history to locate the target device and parse fuzzy instructions.
[0057] The present invention has the following beneficial effects:
[0058] (1) Improve speech recognition accuracy: The method provided by the present invention can achieve accurate extraction of speech signals in an industrial high-noise environment through the dual-channel noise reduction technology, combined with the collaborative optimization mechanism of the voiceprint recognition channel and the noise separation channel; It not only effectively shields non-target user speech and environmental noise, but also can dynamically adapt to changing industrial noise scenarios, greatly improving the speech recognition accuracy in complex environments, providing technical support for the reliable transmission of production instructions.
[0059] (2) Realize dynamic multimodal fusion optimization: The present invention dynamically adjusts the weights of speech, visual, and sensor data based on the real-time perception of the industrial environment (such as noise level, light intensity, temperature and humidity, etc.), based on the Transformer adaptive weighting mechanism and the DQN optimization strategy, to ensure the parsing accuracy in complex industrial scenarios. Compared with the traditional fixed-weight fusion method, the present invention can automatically optimize decisions according to environmental changes, making the instruction parsing always maintain high robustness and efficiency under dynamic conditions.
[0060] (3)Enhance the ability to parse ambiguous instructions: The present invention utilizes a dynamic semantic map and GNN technology to parse the relative position or ambiguous orientation information contained in ambiguous language according to the real-time device layout, sensor data, and user instruction history. For example, when the user only describes "the device near the heating furnace", the system can automatically match the target device and generate corresponding control instructions. This ability to parse ambiguous instructions can significantly improve the flexibility and accuracy of instruction issuance, and is particularly suitable for intelligent control in multi-device intensive areas.
[0061] (4)Improve industrial operation efficiency: The intelligent voice parsing system of the present invention liberates the operator's hands, enabling them to complete tasks such as device operation, parameter adjustment, or production scheduling through voice, significantly simplifying the operation steps in complex process flows, and improving production efficiency and operation convenience. Description of the Drawings
[0062] Figure 1 It is a flowchart of the dual-channel noise reduction speech recognition steps in an embodiment of the present invention;
[0063] Figure 2 It is a flowchart of the multi-modal fusion steps in an embodiment of the present invention;
[0064] Figure 3 It is a flowchart of the ambiguous instruction parsing steps in an embodiment of the present invention;
[0065] Figure 4 It is a structural diagram of the ambiguous instruction parsing system in an embodiment of the present invention;
[0066] Figure 5 It is a working flowchart of the ambiguous instruction parsing system in an embodiment of the present invention;
[0067] Figure 6 It is a flowchart of the ambiguous instruction parsing method based on dual-channel noise reduction and dynamic semantic map in an embodiment of the present invention. Detailed Embodiments
[0068] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0069] On the contrary, the present invention covers any alternatives, modifications, equivalent methods, and solutions made within the spirit and scope of the present invention defined by the claims. Further, in order to enable the public to have a better understanding of the present invention, some specific details are described in detail in the following detailed description of the present invention. Those skilled in the art can fully understand the present invention without the description of these details.
[0070] Embodiment 1: A fuzzy instruction parsing method based on dual-channel noise reduction and dynamic semantic map, as follows Figure 6 The method includes the following steps:
[0071] (1) Noise-reduced speech recognition: Combining dual-channel speech noise reduction and speech recognition, extract the target speech signal in a high-noise environment and complete speech-to-text conversion, and output the text data corresponding to the target speech signal;
[0072] (2) Multimodal fusion: Based on real-time environment perception, dynamically adjust the weights of speech data, visual data, and sensor data;
[0073] (3) Fuzzy instruction parsing: Receive the text data output in step (1), and use the weights of the optimized speech data, visual data, and sensor data output in step (2) as the key inputs for updating the semantic map and parsing fuzzy instructions to achieve semantic map update; Use the dynamically updated semantic map to combine the device layout, environmental data, and user instruction history to locate the target device and parse the fuzzy instructions.
[0074] In step (1) of this embodiment, as follows Figure 1 By synergistically fusing the voiceprint recognition channel and the noise separation channel, and introducing a two-way optimization mechanism between the two channels, improve the speech recognition efficiency and robustness in an industrial high-noise environment; Step (1) specifically includes the following steps:
[0075] (1.1) Establish a voiceprint recognition channel: Extract the voiceprint features of the target user by using a deep learning model, combine with a real-time comparison mechanism, output the voice instruction signal of the target user, and block the voice interference of non-target users;
[0076] (1.2) Establish a noise separation channel: Model the frequency characteristics of industrial noise through an autoencoder and a convolutional neural network CNN to obtain a noise reduction model; Use the noise reduction model to perform preliminary noise reduction on the input noisy audio, extract the speech signal and generate a noise mask;
[0077] (1.3) Dual-channel collaborative optimization: The recognition result output by the voiceprint recognition channel provides a reference signal for the noise separation channel; The clear speech signal extracted by the noise separation channel is fed back to the voiceprint recognition channel as input to optimize the accuracy of voiceprint matching;
[0078] Through the windowing and overlapping technique, combine the audio signals of the voiceprint recognition channel and the noise separation channel to output a clear audio signal that conforms to the voiceprint of the target user, that is, output the target speech signal;
[0079] (1.4) Transmit the target speech signal output in step (1.3) to the speech recognition model to complete speech recognition, and output the text data corresponding to the target speech signal.
[0080] Among them, step (1.1) establishes a voiceprint recognition channel, specifically including:
[0081] (1.11) Voiceprint feature extraction: Use spectral analysis and Mel Frequency Cepstral Coefficients (MFCC) to extract the voice features of all users, and combine Convolutional Neural Network (CNN) and Multi-Layer Residual Network (ResNet) for deep learning to construct a voiceprint feature library;
[0082] (1.12) Real-time voiceprint matching: When the user issues a voice command, the voice command signal is collected in real time and the corresponding voiceprint features are extracted. The voiceprint features corresponding to the voice command signal are compared with the voiceprints in the voiceprint feature library to identify the target user;
[0083] (1.13) Mask non-target voice signals: According to the comparison result of step (1.12), mask the non-target voice signals below the preset threshold in the comparison result and output the optimized voice command signal.
[0084] Specifically, step (1.3) specifically includes:
[0085] (1.31) Voiceprint recognition to noise separation feedback: Calculate the similarity between the input voice of the voiceprint recognition channel and the target voiceprint features. When the similarity is below the preset threshold, feedback the frequency band information of the unmatched voice to the noise separation channel to adjust the noise mask generation strategy;
[0086] (1.32) Noise separation to voiceprint recognition feedback: Output the clear voice signal extracted by the noise separation channel to the voiceprint recognition channel. The voiceprint recognition channel optimizes the voiceprint feature extraction process in step (1.11) according to the received clear voice signal. According to the clear voice signal output by the noise separation channel, adjust the feature sensitivity in the feature extraction process of step (1.11) to more accurately identify the target voice features and reduce noise interference; At the same time, in the process of real-time voiceprint matching in step (1.12), the target voiceprint template will be dynamically updated according to the clear audio signal after noise separation. Through the feedback mechanism, learn and update the voiceprint features of the target user to improve the matching accuracy in a noisy environment.
[0087] Finally, through the windowing and overlapping technique, combine the audio signals of the voiceprint recognition channel and the noise separation channel to output a clear audio signal that conforms to the voiceprint of the target user, that is, output the target voice signal.
[0088] In the present invention, in step (1), a dual-channel noise reduction model is adopted. Identity verification is performed through the voiceprint recognition channel, and industrial noise and background interference are removed through the noise separation channel. The two channels work together to ensure that a pure voice signal is extracted and only the voice of the target employee is processed, and then it is input into a voice recognition model trained with industrial professional voice data for further recognition.
[0089] In step (1.1), in the initialization stage, each employee on the post needs to record their voice samples, extract features from the voice samples, establish voiceprint features, and perform feature learning through a deep learning network. These features will be stored in the voiceprint feature library as the unique identifier of the employee's identity. The first step of the voiceprint recognition channel is to extract features from the original voice signal, and use spectrograms and MFCCs to extract the unique audio features of each employee. Subsequently, the extracted biometric features are subjected to feature learning and comparison through a deep learning network, and the structure of ResNet is introduced to enhance the extraction accuracy of the employee's voiceprint features. The multi-layer convolutional network of ResNet can capture the subtle differences in the audio signal, thereby improving the recognition reliability.
[0090] In daily use, when an employee issues an instruction, the real-time collected voice is compared with the features in the voiceprint library to ensure that only the voice of the employee is subjected to voice parsing, thereby blocking other background sounds and the voice signals of non-post employees, and the recognition result is output.
[0091] In step (1.2) of this embodiment, the noise separation channel is based on a deep noise reduction model combining CNN and an autoencoder. It preliminarily reduces the noise of the input noisy audio, extracts a relatively pure voice signal and generates a noise mask by analyzing the frequency characteristics of common noises in the industrial environment. The generated noise mask will be used as a reference input for the voiceprint recognition channel in the subsequent dynamic feedback process to help it enhance the extraction ability of the target voice signal in a high-noise scenario.
[0092] In step (1.3) of this embodiment, the recognition result output by the voiceprint recognition channel provides a reference signal for the noise separation channel, thereby guiding the noise separation channel to further optimize the noise reduction process, especially for enhancing the voice signal of the target employee; the clear audio generated by the noise separation channel will be fed back as input to the voiceprint recognition channel to further optimize the accuracy of voiceprint matching. When the output result of the noise separation channel or the voiceprint recognition channel does not reach the preset matching threshold, the next round of noise reduction and voiceprint comparison is triggered until the similarity score of voiceprint recognition meets the preset threshold (0.8) or reaches the preset maximum number of iterations (3 times), completing the deep voice noise reduction and extracting the clear voice signal. Subsequently, through the windowing and overlapping technique, the audio signals of the two channels are combined to output a clear audio signal that conforms to the voiceprint of the target user as the final audio output. Finally, the clear audio signal after collaborative optimization will be transmitted to the speech recognition model trained with industrial professional speech data to complete speech recognition and output the corresponding text data.
[0093] In this embodiment, step (2) specifically includes the following steps:
[0094] (2.1) Multi-modal data input: Obtain three types of modal data: voice data, visual data, and environmental sensor data (such as noise level, light intensity, etc.).
[0095] (2.2) Weight calculation: Use the Transformer multi-head attention mechanism to calculate the initial weights ω audio , ω vision , ω sensor of the three types of modal data: voice data, visual data, and environmental sensor data; specifically, for the input features of each modality, use the multi-head self-attention mechanism to assign weights to each modality, and generate the attention scores of each modality through the multi-head attention mechanism according to the changes in environmental information such as noise and light, and normalize the attention scores to the initial weights.
[0096] (2.3) Weight optimization: Optimize the weight allocation strategy of the three types of modal data: voice data, visual data, and environmental sensor data through the DQN network (Deep Q-Network) combined with environmental data and the reward mechanism for weight optimization.
[0097] Take the environmental sensor data and the initial weights obtained in step (2.2) as the current environmental state vector S of the DQN network. The form of the current environmental state vector S is:
[0098] S = [S env , S fusion = [Noise Level, Temperature, Light Intensity, …, ω audio , ω vision , ωsensor
[0099] Among them, S env represents the sensor data of the current environment, including noise data, temperature data, and light data; S fusion is the initial weight ω of the three-modal data of speech data, visual data, and environmental sensor data calculated through step (2.2) audio , ω vision , ω sensor ;
[0100] According to the current environmental state vector S and historical experience, the DQN network selects the historical policy A based on past similar environmental states; the historical policy A is an optimal or approximately optimal modal weight allocation policy based on the actions taken in past similar states;
[0101] Then, the DQN network optimizes the weight combination policy of the three-modal data of speech data, visual data, and environmental sensor data according to the current environmental state vector S and the historical policy through a reward mechanism.
[0102] The weight selection formula of the DQN network is:
[0103]
[0104] Among them, Q(S,A) represents the value of taking the historical policy under the current environmental state vector S, r is the reward, α is the learning rate, and γ is the discount factor; represents the maximum Q value among all decisions A' in the next state S';
[0105] According to the current environmental state vector S, select the policy with the highest Q value, that is, the optimal policy A*:
[0106]
[0107] According to the optimal policy A * , adjust the weights of the three-modal data of speech data, visual data, and environmental sensor data to obtain the optimized modal weights
[0108] In the present invention, the multi-modal fusion step is achieved by equipping multiple environmental sensors, including temperature, humidity, noise, light, etc. sensors, which are installed around key devices. The sensor data is transmitted to the edge computing device in real time through industrial Ethernet or a wireless communication module, enabling the adjustment of modal weights according to environmental changes. This step combines speech, vision, and sensor data, and dynamically adjusts the weights of each modality to improve the accuracy of instruction parsing. The GPU is used to execute the adaptive weighting mechanism in the Transformer model, and the DQN is used to optimize the fusion strategy of different modality data.
[0109] The present invention monitors the environmental conditions in real time through sensor data, and uses the multi-head attention mechanism of Transformer to automatically assign weights to each modality to ensure the parsing stability in different environments. The modality weight strategy is optimized based on the DQN module according to environmental changes to achieve the globally optimal weight combination. When parsing instructions, this method combines the adaptive weighting of Transformer and the strategy optimization of DQN to realize the fusion of multi-modal data and ensure high robustness and accuracy.
[0110] Among them, the multi-head attention mechanism of the Transformer is used to calculate the initial weights of speech, visual, and sensor data according to the environmental state, specifically including: (1) extracting each modality data and generating embedded data; (2) generating the attention scores of each modality through the multi-head attention mechanism; (3) normalizing the attention scores into initial weights. The optimization strategy based on DQN includes: (1) inputting the environmental state as the state variable of DQN; (2) dynamically adjusting the weight allocation of the initial weights according to the reward mechanism, where the reward function is defined based on the instruction parsing accuracy and execution efficiency; (3) using a deep learning model to iteratively update the weight decision parameters to ensure global optimality.
[0111] In this embodiment, as Figure 3 shown, step (3) includes the following steps:
[0112] (3.1) Use the graph neural network GNN to construct a semantic map, dynamically adjust the device node weights in combination with real-time sensor data, and optimize the node associations according to the environmental state;
[0113] (3.2) Update the device area priority based on the historical frequency of user instructions, and combine the optimized modality weights output in step (2) to achieve semantic map update;
[0114] (3.3) Use the all-MiniLM-L6-v2 model to convert the text data output in step (1) into high-dimensional semantic vectors, parse the instructions through the graph propagation mechanism, complete the positioning of the target device or target area, and output the parsing result. This step combines the semantic map updated by GNN to more accurately compare the instruction embedding and the high-dimensional embedding vectors of the device area. Through the propagation mechanism of the graph structure, the correlation of the instruction embedding in the spatial structure is maximized, so as to accurately extract the intention contained in the fuzzy instruction and accurately output the parsing result.
[0115] Step (3) of the present invention is used to recognize and understand the user's ambiguous instructions; accept the text instructions output by step (2), and when the user uses the orientation description of non-specific device names, enhance the recognition effect through the dynamically updated semantic map and voice embedding, combine with GNN to capture the complex relationships between device areas, and improve the parsing accuracy of ambiguous instructions.
[0116] Specifically, the method for updating the semantic map in step (3.2) is as follows:
[0117] Each device area is represented as a node of the graph, and the edges between the nodes represent the spatial relationship or operation dependency relationship of the device areas; the update process is described as:
[0118]
[0119] Among them, represents the weight of the i-th device at time t + 1; j represents the j-th device in the neighborhood of the i-th device, represents the weight of the j-th device at time t; represents the cumulative weight of the nodes within the neighborhood of node i, reflecting the collaborative relationship between the corresponding device and the neighboring areas; η is the learning rate, used to control the step size or speed of node weight update; λ is the neighborhood influence coefficient, used to control the influence degree of neighborhood nodes on the weight of the target node; Freq(i) is the historical frequency of the i-th device instruction, and Sensor(i) is the contribution value of the sensor state.
[0120] Among them, the calculation formula of Sensor(i) is:
[0121] Sensor(i) = Sensor audio (i) + Sensor vision (i) + Sensor sensor (i);
[0122]
[0123] Among them, Sensor raw_audio (i), Sensor raw_vision (i) and Sensor raw_sensor (i) are the original sensor data that have been normalized and standardized from three modalities of voice data, visual data, and environmental sensor data; is the optimized modality weight obtained from (2).
[0124] Through the node update mechanism of the GNN, when a user frequently issues commands for a certain area, the GNN will automatically increase the weight and priority of that area to optimize the response speed in the specific area. By integrating real-time sensor data and the optimized modal weights output by the multimodal fusion module, the device location and environmental characteristics are reflected in the semantic map. Therefore, the method provided by the present invention can identify and process environmental information such as "high-noise area" or "high-temperature area".
[0125] Embodiment 2: A fuzzy instruction parsing system based on dual-channel noise reduction and dynamic semantic map, adopting the method described in Embodiment 1, as Figure 4 shown, the system includes the following modules:
[0126] (1) Noise reduction speech recognition module: Combining dual-channel speech noise reduction and speech recognition, extracting the target speech signal in a high-noise environment and completing speech-to-text conversion, and outputting the text data corresponding to the target speech signal;
[0127] Specifically, in this embodiment, a circular microphone array including 8 to 10 MEMS microphones is adopted, the spacing between each microphone is 20 ± 5 cm, and the overall diameter of the array is about 40 ± 10 cm. This microphone array is located above the console and covers the entire working area of the operator, and this design can capture sounds from different directions. The microphone array locates and enhances the speech signal from the specified direction through the beamforming algorithm, and then inputs the collected speech into the noise reduction speech recognition module.
[0128] (2) Multimodal fusion module: Based on real-time environment perception, dynamically adjusting the weights of speech data, visual data, and sensor data;
[0129] (3) Fuzzy instruction parsing module: Receiving the text data output by the noise reduction speech recognition module, using the optimized weights of speech data, visual data, and sensor data output by the multimodal fusion module as the key inputs for updating the semantic map and parsing fuzzy instructions to achieve semantic map update; Using the dynamically updated semantic map, combining the device layout, environmental data, and user instruction history to locate the target device and parse the fuzzy instructions.
[0130] The fuzzy instruction parsing module uses the GNN to manage and dynamically update the semantic map. The GNN can adjust the weights and label priorities of the device locations in real time according to the user's historical inputs and instruction frequencies, enabling more effective identification and parsing of fuzzy instructions in complex environments.
[0131] Through the node update mechanism of the GNN, when the user frequently issues commands for a certain area, the GNN will automatically increase the weight and priority of that area, thereby optimizing the response speed in the specific area. The core functions of this module include adaptive semantic map update, semantic embedding matching of fuzzy commands, and environmental information integration based on sensor data. By integrating the optimized weights of the multi-modal fusion module and utilizing the GNN's real-time graph structure update ability, the device location and environmental characteristics are reflected in the semantic map, dynamically adjusting the regional label weights to achieve the fusion of multi-modal data and ensure high robustness and accuracy.
[0132] Use the all-MiniLM-L6-v2 model to convert the received text commands in Step 1 into high-dimensional semantic vectors, capture the semantic information of the commands, complete the text command embedding, and represent the semantic features of the fuzzy commands. Combining with the semantic map updated by the GNN, the command embedding and the high-dimensional embedding vectors of the device area can be compared more accurately.
[0133] During fuzzy command matching, the GNN helps to further optimize the command parsing process. Through the propagation mechanism of the graph structure, the correlation of the command embedding in the spatial structure is maximized. Even if the user uses fuzzy descriptions, the command can be automatically parsed to the corresponding area and the accurate execution command can be output.
[0134] As Figure 2 shown, through the dual-channel voice noise reduction and speech recognition of the noise reduction speech recognition module, the speech extraction and recognition in a high-noise environment are completed, and the recognized text result is output to the fuzzy command parsing module; at the same time, the multi-modal fusion module dynamically adjusts the weights of voice, vision, and sensor data according to environmental changes and outputs them to the fuzzy command parsing module; the fuzzy command parsing module is used to locate the target area and complete the fuzzy command parsing by combining the dynamically updated semantic map, the device layout, and the environmental information.
[0135] The present invention has excellent adaptability in different industrial scenarios. Through the dynamic semantic map technology, the system can flexibly update the parsing logic according to the requirements of different industrial scenarios (such as device monitoring, production scheduling, abnormal alarm, etc.) to achieve accurate response to commands. At the same time, the dual-channel noise reduction and multi-modal fusion strategy provide good modular expansion capabilities, enabling this technology to be extended to fields such as automated production lines and equipment maintenance management, providing comprehensive support for the intelligent upgrade of industry.
[0136] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A fuzzy instruction parsing method based on dual-channel noise reduction and dynamic semantic map, characterized in that: The method comprises the following steps: (1) Noise reduction speech recognition: Combine dual-channel speech noise reduction and speech recognition to extract the target speech signal in a high-noise environment and complete speech-to-text conversion, outputting the text data corresponding to the target speech signal; (2) Multimodal fusion: Dynamically adjust the weights of voice data, visual data, and sensor data based on real-time environmental perception; (3) Fuzzy command parsing: receiving the text data output from step (1), using the weights of the optimized voice data, visual data, and sensor data output from step (2) as key inputs for updating the semantic map and parsing the fuzzy command, thereby updating the semantic map; using the dynamically updated semantic map in combination with the device layout, environmental data, and user command history to locate the target device and parse the fuzzy command.
2. The fuzzy instruction parsing method based on dual-channel noise reduction and dynamic semantic map according to claim 1 is characterized in that: Step (1) specifically includes the following steps: (1.1) Establish a voiceprint recognition channel: By using a deep learning model to extract the voiceprint features of the target user, combined with a real-time comparison mechanism, the target user's voice command signal is output and the voice interference of non-target users is blocked; (1.2) Establishing a noise separation channel: Modeling the frequency characteristics of industrial noise through an autoencoder and a convolutional neural network (CNN) to obtain a noise reduction model; using the noise reduction model to perform preliminary noise reduction on the input noisy audio, extract the speech signal and generate a noise mask; (1.3) Dual-channel collaborative optimization: The recognition results output by the voiceprint recognition channel provide a reference signal for the noise separation channel; the clear speech signal extracted by the noise separation channel is fed back to the voiceprint recognition channel as input to optimize the accuracy of voiceprint matching; Through the windowing and overlapping technology, the audio signals of the voiceprint recognition channel and the noise separation channel are combined to output a clear audio signal that conforms to the target user's voiceprint, that is, the target voice signal; (1.4) The target speech signal outputted from step (1.3) is transmitted to a speech recognition model to complete speech recognition and output text data corresponding to the target speech signal.
3. The fuzzy instruction parsing method based on dual-channel noise reduction and dynamic semantic map according to claim 2 is characterized in that: Step (1.1) establishes a voiceprint recognition channel, specifically including: (1.11) Voiceprint feature extraction: Spectral analysis and Mel frequency cepstral coefficients are used to extract the voice features of all users. Deep learning is performed by combining convolutional neural networks and multi-layer residual networks to build a voiceprint feature library. (1.12) Real-time voiceprint matching: When a user issues a voice command, the voice command signal is collected in real time and the corresponding voiceprint features are extracted. The voiceprint features corresponding to the voice command signal are compared with the voiceprints in the voiceprint feature library to identify the target user; (1.13) Shielding non-target voice signals: According to the comparison result of step (1.12), shielding the non-target voice signals below a preset threshold in the comparison result, and outputting an optimized voice command signal.
4. The fuzzy instruction parsing method based on dual-channel noise reduction and dynamic semantic map according to claim 1 is characterized in that: Step (2) specifically includes the following steps: (2.1) Multimodal data input: Acquire three modal data: voice data, visual data, and environmental sensor data; (2.2) Weight calculation: Use the Transformer multi-head attention mechanism to calculate the initial weights ω of the three modal data: speech data, visual data, and environmental sensor data. audio ,ω vision ,ω sensor ; (2.3) Weight optimization: By combining environmental data with the reward mechanism through the DQN network, the weight allocation strategy of the three modal data, namely speech data, visual data and environmental sensor data, is optimized to perform weight optimization.
5. The fuzzy instruction parsing method based on dual-channel noise reduction and dynamic semantic map according to claim 4 is characterized in that: Step (2.3) specifically includes: The sensor data of the environment and the initial weights obtained in step (2.2) are used as the current environment state vector S of the DQN network. The current environment state vector S is in the form of: S=[S env ,S fusion ]=[Noise Level,Temperature,Light Intensity,…, oh audio Oh, oh vision Oh, oh sensor ]; Among them, S env Represents the sensor data of the current environment, including noise data, temperature data, and light data; S fusion is the initial weight ω of the three modal data, namely, speech data, visual data and environmental sensor data, calculated by step (2.2) audio ,ω vision ,ω sensor ; Based on the current environment state vector S and historical experience, the DQN network selects a historical strategy A based on similar environment states in the past; Then the DQN network optimizes the weight combination strategy of the three modal data, namely speech data, visual data and environmental sensor data, through the reward mechanism according to the current environment state vector S and historical strategy.
6. The fuzzy instruction parsing method based on dual-channel noise reduction and dynamic semantic map according to claim 5 is characterized in that: The weight selection formula of the DQN network is: Among them, Q(S,A) represents the value of adopting the historical strategy under the current environment state vector S, r is the reward, α is the learning rate, and γ is the discount factor; Indicates that in the next state S ′ Under this condition, the maximum Q value among all decisions A′; According to the current environment state vector S, select the strategy with the highest Q value, that is, the optimal strategy A*: According to the optimal strategy A * , adjust the weights of the three modal data: voice data, visual data, and environmental sensor data, and obtain the optimized modal weights 7. The fuzzy instruction parsing method based on dual-channel noise reduction and dynamic semantic map according to claim 6 is characterized in that: Step (3) comprises the following steps: (3.1) Use graph neural network (GNN) to build semantic maps, dynamically adjust device node weights based on real-time sensor data, and optimize node associations based on environmental conditions; (3.2) Update the device area priority based on the historical frequency of user instructions, combined with the optimized modal weight output in step (2) Implement semantic map updates; (3.3) The text data outputted from step (1) is converted into a high-dimensional semantic vector using the all-MiniLM-L6-v2 model, instructions are parsed through a graph propagation mechanism, the target device or target area is located, and the parsing result is outputted.
8. The fuzzy instruction parsing method based on dual-channel noise reduction and dynamic semantic map according to claim 7 is characterized in that: The method for updating the semantic map in step (3.2) is: Each device region is represented as a node of the graph, and the edges between nodes represent the spatial relationship or operation dependency of the device regions; the update process is described as: in, represents the weight of the i-th device at time t+1; j represents the j-th device in the neighborhood of the i-th device, represents the weight of the jth device at time t; represents the cumulative weight of nodes in the neighborhood of node i, reflecting the collaborative relationship between the corresponding device and the adjacent area; η is the learning rate, which is used to control the step size or speed of node weight update; λ is the neighborhood influence coefficient, which is used to control the degree of influence of neighborhood nodes on the target node weight; Freq(i) is the historical frequency of the i-th device instruction, and Sensor(i) is the sensor state contribution value.
9. The fuzzy instruction parsing method based on dual-channel noise reduction and dynamic semantic map according to claim 7 is characterized in that: The calculation formula for Sensor(i) is: Sensor(i)=Sensor audio (i)+Sensor vision (i)+Sensor sensor (i); Among them, Sensor raw_audio (i) Sensor raw_vision (i) and Sensor raw_sensor (i) is the normalized and standardized raw sensor data from three modalities: speech data, visual data, and environmental sensor data; is the optimized modal weight obtained from (2).
10. A fuzzy instruction parsing system based on dual-channel noise reduction and dynamic semantic map, characterized in that: The system includes the following modules: (1) Noise reduction speech recognition module: Combines dual-channel speech noise reduction and speech recognition to extract the target speech signal in a high-noise environment and complete speech-to-text conversion, outputting the text data corresponding to the target speech signal; (2) Multimodal fusion module: Dynamically adjust the weights of voice data, visual data, and sensor data based on real-time environmental perception; (3) Fuzzy command parsing module: Receives the text data output by the noise reduction speech recognition module, uses the weights of the optimized speech data, visual data, and sensor data output by the multimodal fusion module as the key input for updating the semantic map and parsing fuzzy commands, and realizes the update of the semantic map; uses the dynamically updated semantic map combined with the device layout, environmental data, and user command history to locate the target device and parse the fuzzy commands.
Citation Information
Cited By
Voice song requesting interaction data processing method and system based on multi-mode fusion
CN120279870A
Voice-on-demand interaction data processing method and system based on multimodal fusion
CN120279870B
Multi-information fusion voice interaction method and system for range hood
CN120544569A
Live broadcast book-speaking noise processing method, system and device, and medium
CN120636436A