Job ar intelligent auxiliary management system based on voice ai interaction driving

By combining environmentally-aware voice interaction and AR multimodal interaction compensation modules with composite sensor arrays and deep neural networks, the problems of low efficiency, poor interaction, and insufficient real-time performance in existing job management systems have been solved, achieving efficient, accurate, and real-time job management.

CN121096334BActive Publication Date: 2026-03-24GUANGZHOU TODIAN NEW ENERGY TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-01
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing job management systems are inefficient, have poor user experience, and lack real-time capabilities. Voice AI interaction has low recognition accuracy in industrial environments, and AR technology has limited applications and cannot provide real-time and accurate guidance information.

Method used

It employs an environment-aware voice interaction module, an AR multimodal interaction compensation module, a dynamic noise robust data processing module, and a job management module, combined with a composite sensor array, adaptive beamforming processing, deep neural networks, and long short-term memory networks, to achieve real-time noise reduction and semantic understanding of voice signals, provide visualization and tactile interaction compensation, and monitor and adjust the job process in real time.

Benefits of technology

It improved operational efficiency and accuracy, optimized the interactive experience, enhanced real-time performance, reduced human error, and achieved automation of work processes and real-time data processing, thereby improving the efficiency and quality of production and warehouse management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121096334B_ABST
    Figure CN121096334B_ABST
Patent Text Reader

Abstract

The application discloses a kind of job AR intelligent auxiliary management systems based on voice AI interaction drive, applied to intelligent storage field, comprising: environmental perception type voice interaction module, AR multimodal interaction compensation module, dynamic noise robustness data processing module, job management module and database module;The environmental perception type voice interaction module is used to collect job personnel voice signal and suppress environmental noise;The dynamic noise robustness data processing module is used to carry out noise reduction processing, semantic understanding and closed loop optimization to voice signal, and generates feedback information and control instruction;The AR multimodal interaction compensation module is used to receive the feedback information of the dynamic noise robustness data processing module and provide visual and tactile interaction compensation, the job management module is used to dispatch management equipment according to control instruction;The application can effectively improve the efficiency of job, accuracy is high, real-time is strong and can realize interactive experience optimization.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent warehousing, and particularly to a job AR intelligent auxiliary management system based on voice AI interaction driving. BACKGROUND

[0002] In the field of job management today, traditional job management systems occupy a considerable market share. Among them, the job management mode relying on manual recording and paper document circulation has a long history and is widely applied. Under this mode, job personnel need to manually fill in various paper forms to record relevant information of job tasks, such as task content, completion progress, equipment usage, etc. For example, in the warehouse scenario, the record of goods in and out of the warehouse is usually recorded by the staff on the paper account book, including the information of goods name, quantity, in and out time, and the person in charge. In the workshop production scenario, workers record the processing steps, processing time and raw materials used on the paper work order. Then, these paper documents are circulated among different departments or personnel to realize the transmission of information and the promotion of job flow.

[0003] With the preliminary development of information technology, simple digital systems have emerged. Such systems transfer part of the job management process from offline to online, realizing a certain degree of informatization. Job personnel can access the system through computer terminals to perform data entry and query operations. For example, some enterprises use simple spreadsheets or database management systems to record job data. Job personnel input job-related information in the system, and the system can perform simple storage and statistical analysis of these data. In terms of job scheduling, the system can generate a preliminary job plan according to preset rules and display it to management personnel in the form of an electronic spreadsheet or a simple report. However, this simple digital system has relatively limited functions and is not capable of dealing with complex job scenarios and diverse job requirements. The existing management system has the following deficiencies:

[0004] (1) Low efficiency: Manual operation is prone to various errors in job management processes, such as data entry errors, information omissions, etc. When filling out paper documents, staff may mistakenly write the wrong quantity of goods or job time due to negligence, which not only requires additional time for checking and correcting, but also may cause deviations in subsequent job processes, affecting the overall job progress. Moreover, manual operation is time-consuming, especially when dealing with a large number of job tasks, the speed of manual recording and organizing data cannot meet the actual demand. The paper process is cumbersome, from filling out paper documents, passing to storage, each link needs to spend a lot of manpower and time cost. In terms of data aggregation and analysis, it is difficult to quickly and accurately aggregate and analyze the data of paper documents, and managers need to spend a lot of time manually organizing and calculating data to obtain key information such as job progress and equipment utilization, which seriously affects the timeliness and accuracy of decision-making, and thus affects job progress.

[0005] (2) Poor interactive experience: The interface interaction design of traditional systems is often not intuitive, and the operation process is complex. Job personnel need to spend a lot of time learning how to use the system and familiarize themselves with various operation instructions and interface layouts. For some older or less familiar with information technology staff, the learning cost is higher, which may lead to their resistance to using the system and reduce their work enthusiasm. When obtaining key information, job personnel need to switch between multiple interfaces and perform tedious operations to find the required information, making it difficult to quickly obtain key information, which affects work efficiency and job quality to some extent.

[0006] (3) Lack of real-time: In traditional job management systems, data updates often have delays. After the actual situation of the job site changes, the relevant information cannot be timely reflected in the system, because data entry needs manual operation, and the speed and frequency of manual operation are limited, which cannot achieve real-time updating. For example, in workshop production, if the equipment fails or the production progress is completed ahead of schedule, etc., the situation cannot be fed back to the system in time, and the management personnel are difficult to adjust the job plan in time, which may cause waste of production resources or production delay. This delay in data updating makes the system unable to reflect the real situation of the job site in real time, reducing the guiding role of the system in job management.

[0007] (4) Voice AI interaction application problems in job scene: Although voice AI interaction has great potential in job scene application, it still faces many problems. In industrial environment, noise interference is one of the main reasons for low voice recognition accuracy. The huge noise generated by workshop equipment running, such as the roar of machine tools, the running sound of conveyors, and the noise of warehouse machinery such as forklift engines, will superimpose with the voice signal of the job personnel, making the signal-to-noise ratio of the voice signal collected by the microphone array low, thereby increasing the difficulty of voice recognition. When multiple job personnel work at the same time, the voice signals will interfere with each other, and the traditional beamforming algorithm is difficult to accurately locate the target sound source, resulting in the voice recognition system unable to accurately recognize the instructions of each job personnel. In addition, inaccurate voice instruction understanding is also a common problem. Due to the complexity and diversity of natural language, after the voice recognition system converts the voice signal into text, it may not accurately understand the semantics of the text and the intention of the job personnel, resulting in incorrect instruction execution.

[0008] (5) Limitations of AR technology application in job management: There are still some limitations in the application of AR technology in job management. The combination of AR display content and actual job flow is not close enough, and it cannot provide real-time and accurate guidance information for job personnel. In the warehouse scene, the goods storage location information displayed by the AR glasses may deviate from the actual location, or cannot be dynamically updated according to the real-time operation of the job personnel, which makes the job personnel still need to spend time to judge and confirm when using AR technology, and cannot fully play the advantages of AR technology. The single interaction mode is also a problem, and most of the current AR applications only support simple gesture or touch interaction, lacking intelligent guidance function. Job personnel need to manually operate AR devices to obtain information or input instructions during operation, which to some extent distracts the attention of job personnel and affects the job efficiency. Moreover, the lack of intelligent guidance function makes job personnel unable to get effective guidance from the system when facing complex job tasks, increasing the difficulty of job and the possibility of errors.

[0009] Therefore, there is an urgent need for a voice AI interaction driven job AR intelligent auxiliary management system that can effectively improve the efficiency of job, has high accuracy, strong real-time performance and can optimize the interactive experience, SUMMARY

[0010] The purpose of the present application is to overcome the shortcomings of the prior art and provide a voice AI interaction driven job AR intelligent auxiliary management system that can effectively improve the efficiency of job, has high accuracy, strong real-time performance and can optimize the interactive experience.

[0011] To achieve the above invention purpose, the technical solution adopted by the present application is as follows:

[0012] The voice AI interaction-driven AR intelligent auxiliary management system is applied in the field of intelligent warehousing and includes: an environmentally aware voice interaction module, an AR multimodal interaction compensation module, a dynamic noise robust data processing module, an operation management module, and a database module.

[0013] The environment-aware voice interaction module is used to collect the voice signals of the operators and suppress environmental noise. It is connected to the dynamic noise robust data processing module. The environment-aware voice interaction module includes a composite sensor array unit and an adaptive beamforming processing unit.

[0014] The AR multimodal interaction compensation module is used to receive feedback information from the dynamic noise robustness data processing module and provide visualization and tactile interaction compensation. It is bidirectionally connected to the dynamic noise robustness data processing module.

[0015] The dynamic noise robust data processing module is connected to the environment-aware voice interaction module, the AR multimodal interaction compensation module, the job management module, and the database module, respectively, and is used to perform noise reduction processing, semantic understanding, and closed-loop optimization on the voice signal, and generate feedback information and control commands.

[0016] The operation management module is connected to the dynamic noise robust data processing module and the production equipment or storage equipment, and is used to schedule and manage the equipment according to control commands; the operation management module includes a status monitoring unit, which is used to monitor the operating status of the storage equipment and collect the operating data of the storage equipment;

[0017] The database module is connected to the dynamic noise robust data processing module to store work processes, equipment information, product information, and contextualized noise data;

[0018] The environment-aware voice interaction module synchronously acquires voice signals, environmental noise time-domain waveforms, frequency-domain features, and spatial distribution data through a composite sensor array. After processing by the adaptive beamforming processing unit, it outputs an anti-noise voice signal to the dynamic noise robustness data processing module. The dynamic noise robustness data processing module performs real-time noise modeling, noise reduction processing, and confidence-based feedback-based recognition model optimization on the input anti-noise voice signal, outputting semantically understood work instructions and feedback information. The AR multimodal interaction compensation module receives the work instructions and feedback information and displays them visually. Simultaneously, it acquires correction instructions confirmed by the operator's gestures or tactile feedback and sends them back to the dynamic noise robustness data processing module. The dynamic noise robustness data processing module generates control instructions based on the correction instructions and sends them to the work management module. Meanwhile, the operator... The AR multimodal interaction compensation module prompts the corresponding warehousing operations; the operation management module controls the warehousing equipment to perform warehousing operations according to the control commands sent by the dynamic noise robustness data processing module; the status monitoring unit monitors the operating status of the warehousing equipment in real time, collects the operating data of the warehousing equipment, and transmits the data to the dynamic noise robustness data processing module; the dynamic noise robustness data processing module monitors and adjusts the operation process in real time based on the real-time collected data to ensure the smooth progress of the operation; when the warehousing operation is completed, the AR multimodal interaction compensation module confirms the completion of the warehousing operation by obtaining the operator's voice commands or tactile feedback and sends the confirmation information to the dynamic noise robustness data processing module; after receiving the confirmation information, the dynamic noise robustness data processing module updates the warehousing information and operation process records in the database module.

[0019] Preferably, the environmentally aware voice interaction module includes a processor; the processor is a Rockchip RK3588; the composite sensor array unit includes a microphone array, an environmental noise sensor, and an orientation sensor connected to the processor via I2C and SPI interfaces; the environmental noise sensor includes a MEMS accelerometer and an ultrasonic sensor; the adaptive beamforming processing unit is connected to the composite sensor array unit via a parallel data bus to transmit noise spectrum and spatial distribution data in real time, and outputs the noise-resistant voice signal to the dynamic noise robust data processing module via a USB 3.0 interface.

[0020] Preferably, the microphone array adopts a multi-channel distributed layout, and achieves directional enhancement of the voice signal based on the delay summation beamforming principle. It strengthens the operator's voice and suppresses surrounding interference through spatial filtering. Its core model is as follows:

[0021]

[0022] Where y(t) represents the output signal, x m (t) represents the signal received by the m-th microphone, τ m The time delay of this channel is M, and the number of microphones is M; the MEMS accelerometer in the environmental noise sensor extracts the spectral characteristics of the vibration signal through Fourier transform, as shown in the formula:

[0023]

[0024] Where X(f) represents the frequency domain signal, x(t) is the time domain vibration acceleration signal, f is the frequency, and j is the imaginary unit; the MEMS signal reflects the amplitude-frequency relationship of the mechanical structure noise source; the ultrasonic sensor obtains the direction information of the sound source through ranging and sound intensity localization methods, and its basic localization model is the sound wave propagation time difference method:

[0025]

[0026] Where θ represents the angle between the sound source and the front of the array, c is the speed of sound, Δt is the time difference between the left and right channels, and d is the distance between the two sensors.

[0027] Preferably, the adaptive beamforming processing unit works closely with the composite sensor array unit to construct a dynamic noise field modeling and target sound source localization mechanism by receiving the noise spectrum characteristics and spatial distribution information output in real time from the environmental noise sensor and the orientation sensor. It also dynamically calculates the weighting coefficients of the microphone array based on the minimum variance distortionless response algorithm, thereby forming a high-gain pickup beam pointing towards the operator's mouth. The core optimization objective is:

[0028]

[0029] Where w represents the microphone array weighting vector, R is the noise covariance matrix, and d is the steering vector in the target direction. H Let w be the conjugate transpose; the analytical solution to this optimization problem is:

[0030]

[0031] Among them, w optThe optimal weight vector determines the contribution weight of each microphone channel to the final output signal. The oral cavity spatial coordinates provided by the orientation sensor are used to construct the steering vector d in real time, while the spectral data sensed by the environmental noise sensor participates in constructing the covariance matrix R, thereby enabling the beamforming process to respond to the current sound field state. When the environmental noise sensor detects a strong noise event from a specific direction, the adaptive beamforming processing unit adjusts the array weighting strategy to form a main lobe between 30° and 60° in front, while simultaneously forming a strongly suppressed side lobe in the rear, thereby enhancing the target speech and suppressing the noise direction signal. Finally, the processed high signal-to-noise ratio speech signal is input in real time to the dynamic noise robust data processing module for subsequent semantic analysis and task instruction generation.

[0032] Preferably, the dynamic noise robustness data processing module includes a real-time noise modeling submodule, a recognition-denoising joint optimization submodule, an analysis and decision-making submodule, an instruction generation submodule, a semantic understanding unit, and a speech recognition unit. The real-time noise modeling submodule works in conjunction with the environment-aware voice interaction module to dynamically model and suppress complex background noise in industrial settings. Its core mechanism is based on an improved Hidden Markov Model (HMM) for noise state transition modeling and parameter estimation, and integrates deep neural networks and long short-term memory (LSTM) network structures. The HMM models the characteristics of background noise changing over time through a state transition probability matrix and an observation probability function. The model form is as follows:

[0033]

[0034] Where P(O|λ) represents the observation sequence probability under model parameter λ, Q is the hidden state sequence, O is the observed noise feature sequence, P(Q|λ) is the state transition probability, and P(O|Q,λ) is the observation probability under a given state sequence. Combined with the Baum-Welch algorithm, the transition and output probabilities of the noise state can be updated online, achieving dynamic estimation of the noise power spectral density. This estimation result, along with the noisy speech signal acquired by the current microphone array, is input into a DNN-LSTM hybrid model. This DNN-LSTM hybrid model first extracts static spectral features using a deep neural network, then models the contextual relevance in the time series using a long short-term memory network. Its output is the estimated clean speech spectrum. The overall prediction process can be represented as:

[0035]

[0036] Where Y(f,t) represents the noisy speech spectrum, and N(f,t) represents the real-time noise spectrum characteristics. The model represents the predicted clean speech spectrum; during training, the goal is to minimize the reconstruction error, and the mean squared error loss function is used for optimization.

[0037]

[0038] Where S(f,t) is the actual speech spectrum. For the model output, T and F represent the number of time frames and the number of frequency points, respectively. Preferably, the recognition-denoising joint optimization submodule, as a core component of the dynamic noise robustness data processing module, forms a closed-loop feedback path with the real-time noise modeling submodule and the semantic understanding unit. It drives the adaptive optimization of the denoising model and the recognition model through a confidence evaluation mechanism of the speech recognition results. The recognition-denoising joint optimization submodule generates a confidence score for the recognition results based on the connection-time classification loss function, which quantifies the degree of matching between the output text of the speech recognition unit and the target label. Its core loss function is:

[0039]

[0040] in, Let represent the connection-time classification loss, where x is the input speech feature sequence, y is the target text label sequence, and π represents all possible paths. B -1 ( y ) Let P(π|x) represent the set of paths that can be mapped from the target sequence, and let P(π|x) be the probability of a path given the input. The recognition-denoising joint optimization submodule calculates the confidence score in real time and compares it with a set threshold. When the score is lower than 80%, the parameter adjustment strategy of the denoising algorithm is triggered. The spectrum adjustment process is expressed as follows:

[0041]

[0042] in, Y(f,t) represents the corrected speech spectrum, and Y(f,t) represents the original noisy spectrum. To estimate the noise spectrum, α is the attenuation coefficient, and ∈ is the minimum threshold to avoid negative values. After completing error correction, the recognition-denoising joint optimization submodule combines the corrected text with the original speech to form training samples. It then updates the parameters of the recognition model and denoising network in real time through backpropagation to minimize prediction errors and enhance adaptability to similar noise scenarios. The overall AR intelligent auxiliary management system forms an online learning closed loop of "collection-denoising-recognition-feedback-optimization." The optimization process aims to minimize the joint loss function, expressed as:

[0043] L total =λ1L CTC +λ2L denoise

[0044] Among them, L total For the joint loss function, L total To identify the loss, L denoise For noise reduction loss, λ1 and λ2 are adjustable weighting coefficients;

[0045] The semantic understanding unit is connected to the recognition-denoising joint optimization submodule to perform semantic analysis on the optimized text and determine the specific content and intent of the work instructions. The data integration submodule retrieves cargo-related information from the database module, including the cargo's storage location and inbound or outbound process. The analysis and decision-making submodule analyzes and processes the integrated data and determines the optimal warehousing path and operation steps based on a preset work optimization algorithm. The instruction generation submodule generates corresponding feedback information and control instructions based on the results of the analysis and decision-making submodule.

[0046] Preferably, the AR multimodal interaction compensation module includes an AR display auxiliary confirmation unit and a tactile feedback error correction unit; the AR display auxiliary confirmation unit is embedded in the AR glasses worn by the operator and maintains a real-time communication connection with the dynamic noise robust data processing module; the processor of the AR display auxiliary confirmation unit is an ARM Cortex-A53 processor, which interacts with the dynamic noise robust data processing module through the MIPI_DSI interface to display data, and the gesture correction signal is transmitted back through the SPI interface; when the voice recognition unit recognizes the user's command text, a semi-transparent floating window will be generated in the AR glasses' field of view to display the current recognition result and multiple candidate words with high confidence, avoiding misjudgment of the work command due to a single voice recognition error; the floating window adopts a side-by-side arrangement or a vertical sliding list form to quickly complete the confirmation.

[0047] Preferably, the tactile feedback error correction unit is integrated into the AR glasses or smart gloves worn by the operator. It is connected in real-time to the dynamic noise robust data processing module to construct an instant perception and intervention mechanism for industrial work sites. The tactile feedback error correction unit uses an STM32H7 series processor, receives device status conflict signals through the GPIO interface, and provides early warning through vibrations at different frequencies. Simultaneously, it interacts with the dynamic noise robust data processing module via the UART interface to complete control command interaction. When the tactile feedback error correction unit detects that a voice command may cause risk or misoperation, it sends an early warning signal to the operator through physical vibration, guiding them to actively check whether their current operational intention matches the actual state. When the data processing module receives the voice recognition result and performs semantic understanding, it automatically retrieves the device status information from the database module and compares it with the command. If a conflict is found, the tactile feedback mechanism is immediately triggered.

[0048] Preferably, the operation management module further includes a task scheduling unit, an equipment control unit, and a processor. The task scheduling unit controls the forklift to assist operators in handling goods according to control instructions sent by the dynamic noise robustness data processing module. The equipment control unit generates specific equipment control signals to control the forklift to travel along a specified path and accurately transport goods to the inbound or outbound location. The processor of the operation management module uses an industrial-grade ARM Cortex-A72 and interacts with the dynamic noise robustness data processing module via Ethernet using the Modbus TCP protocol to exchange control instructions. When the processor of the operation management module is connected to production equipment or warehousing equipment, it adapts to industrial sensors through an RS-485 interface and controls the warehousing equipment using the CANopen protocol. The status monitoring data of the warehousing equipment is transmitted to the dynamic noise robustness data processing module via WiFi through the MQTT protocol, forming a closed-loop management of the operation process.

[0049] Preferably, the database module contains a scenario-based noise database that collects noise from typical industrial scenarios, including various noise types and covering a sound pressure level range of 75-115 dB; the typical industrial scenario noise includes at least lathe noise, warehouse rack handling noise, and pneumatic tool noise.

[0050] Beneficial effects

[0051] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0052] (1) Efficiency Improvement: This invention is driven by voice AI interaction. Operators only need to speak the work instructions to quickly identify and execute the corresponding operations, greatly reducing the time and workload of manual input. In warehousing scenarios, operators do not need to manually search for the location of goods; they can quickly obtain the location of goods and perform operations through voice commands, shortening the time for goods to enter and leave the warehouse, realizing automated scheduling and management of the work process, reducing manual intervention, and improving the continuity and smoothness of operations. In workshop production scenarios, the system can automatically arrange the processing tasks of production equipment according to production tasks and equipment status, avoiding errors and delays that may occur during manual scheduling, and improving production efficiency. In addition, this invention can also collect and process on-site data in real time, promptly identify problems and make adjustments, further improving operational efficiency.

[0053] (2) Improved Accuracy: The application of speech recognition and AR technology in this invention effectively reduces human error. The speech recognition unit uses advanced algorithms to accurately recognize the voice commands of operators, avoiding errors that may occur with manual input. The AR multimodal interaction compensation module intuitively displays the work information in the operator's field of vision, enabling the operator to clearly understand the work requirements and operating steps, reducing operational errors caused by misunderstandings of information. This invention has real-time verification and feedback functions, which can monitor and verify the operator's operations in real time, and promptly detect and correct errors. During the process of goods entering and leaving the warehouse, the quantity, model, and other information of the goods are verified in real time. If an error is found, a prompt will be issued immediately, requiring the operator to check and correct it, ensuring the accuracy of the operation.

[0054] (3) Optimized Interactive Experience: The natural voice interaction and immersive AR display of this invention make operation more convenient and intuitive, reducing the learning cost for operators. Operators do not need to learn complex operation interfaces and instructions; they can complete tasks simply by interacting with the system using natural language, thus improving their work enthusiasm and satisfaction. When using AR glasses for work, operators can experience an immersive interactive experience, as if virtual information is integrated with the real environment, enhancing the fun and immersion of the work. In addition, this invention also supports multimodal interaction, allowing operators to choose from various interaction methods such as voice, gestures, and touch according to their needs, improving the flexibility and convenience of interaction.

[0055] (4) Enhanced Real-Time Performance: This invention can collect various data from the work site in real time, including equipment status and work progress, and process and analyze this data in real time. Based on the real-time data, the system can provide timely decision-making support to managers, helping them to make quick decisions and adjust work plans. In workshop production, when equipment malfunctions, this invention can immediately detect it and feed the fault information back to managers. Managers can then promptly arrange for maintenance personnel to carry out repairs based on the information provided by the system, avoiding the impact of malfunctions on production. Furthermore, it can optimize the work process based on real-time data, improving work efficiency and quality. Attached Figure Description

[0056] Figure 1 This is a system architecture diagram of the present invention;

[0057] Figure 2 This is the system signaling diagram of the present invention;

[0058] Figure 3 This is a schematic diagram showing the connection between the dynamic noise robust data processing module, the AR multimodal interaction compensation module, and the job management module of the present invention.

[0059] Figure 4 This is a schematic diagram showing the connection between the environmentally aware voice interaction module and the dynamic noise robust data processing module of the present invention. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of the present invention clearer and more complete, the present invention will be further described in detail below with reference to the embodiments. Obviously, the embodiments described below are some embodiments of the present invention, but the scope of protection claimed by the present invention is not limited to the specific embodiments below.

[0061] like Figures 1 to 4 As shown, the AR intelligent auxiliary management system for operations driven by voice AI interaction is applied in the field of intelligent warehousing. It includes: an environmentally perceptive voice interaction module, an AR multimodal interaction compensation module, a dynamic noise robust data processing module, an operation management module, and a database module.

[0062] like Figure 1 and Figure 4 As shown, the environment-aware voice interaction module is used to collect the voice signals of operators and suppress environmental noise. It is connected to the dynamic noise robustness data processing module. The environment-aware voice interaction module can effectively solve the problem of decreased voice command recognition accuracy caused by industrial environmental noise interference. The environment-aware voice interaction module includes a composite sensor array unit, an adaptive beamforming processing unit, and a processor. The processor uses Rockchip RK3588, an ARM architecture core processor, and is equipped with RAM memory, ROM memory, and a power management IC to construct the information processing unit. The composite sensor array unit includes a microphone array, an environmental noise sensor, and a position sensor connected to the processor via I2C and SPI interfaces. The environmental noise sensor includes a MEMS accelerometer and an ultrasonic sensor. The adaptive beamforming processing unit is connected to the composite sensor array unit through a parallel data bus to transmit noise spectrum and spatial distribution data in real time, and outputs the noise-resistant voice signal to the dynamic noise robustness data processing module through a USB 3.0 interface.

[0063] The microphone array adopts a multi-channel distributed layout and achieves directional enhancement of the voice signal based on the delay-summing beamforming principle. It enhances the operator's voice and suppresses surrounding interference through spatial filtering. Its core model is as follows:

[0064]

[0065] Where y(t) represents the output signal, x m (t) represents the signal received by the m-th microphone, τ mThe time delay of this channel is M, and the number of microphones is M; the MEMS accelerometer in the environmental noise sensor extracts the spectral characteristics of the vibration signal through Fourier transform, as shown in the formula:

[0066]

[0067] Where X(f) represents the frequency domain signal, x(t) is the time domain vibration acceleration signal, f is the frequency, and j is the imaginary unit; the MEMS signal reflects the amplitude-frequency relationship of the mechanical structure noise source; the ultrasonic sensor obtains the direction information of the sound source through ranging and sound intensity localization methods, and its basic localization model is the sound wave propagation time difference method:

[0068]

[0069] Where θ represents the angle between the sound source and the front of the array, c is the speed of sound, Δt is the time difference between the left and right channels, and d is the distance between the two sensors. The orientation sensor, in conjunction with the head or face tracking system, acquires the relative position of the worker's mouth and its spatial geometric relationship with the noise source in real time. This provides spatial constraints for speech signal extraction and supports the spatial filtering input for sound source separation and echo suppression algorithms, thereby improving the stability and accuracy of speech signal processing in dynamic industrial environments.

[0070] The adaptive beamforming processing unit works closely with the composite sensor array unit. By receiving the noise spectrum characteristics and spatial distribution information output in real time from the environmental noise sensor and the orientation sensor, it constructs a dynamic noise field modeling and target sound source localization mechanism. Based on the minimum variance distortionless response algorithm, it dynamically calculates the weighting coefficients of the microphone array, thereby forming a high-gain pickup beam pointing towards the operator's mouth. This method minimizes output power while maintaining the amplitude of the target speech signal without distortion, achieving the effect of suppressing noise from non-target directions. The core optimization objective is:

[0071]

[0072] Where w represents the microphone array weighting vector, R is the noise covariance matrix, and d is the steering vector in the target direction. H Let w be the conjugate transpose; the analytical solution to this optimization problem is:

[0073]

[0074] Among them, w optThe optimal weight vector determines the contribution weight of each microphone channel to the final output signal. In actual operation, the oral cavity spatial coordinates provided by the orientation sensor are used to construct the guide vector d in real time, while the spectral data sensed by the environmental noise sensor participates in constructing the covariance matrix R, thereby enabling the beamforming process to respond to the current sound field state. When the environmental noise sensor detects a strong noise event from a specific direction (such as a forklift approaching), the adaptive beamforming processing unit adjusts the array weighting strategy to form a main lobe between 30° and 60° in front, while simultaneously forming a strongly suppressed side lobe behind, thereby enhancing the target speech and suppressing the noise direction signal. Finally, the processed high signal-to-noise ratio speech signal is input in real time to the dynamic noise robust data processing module for subsequent semantic analysis and operation instruction generation.

[0075] The AR multimodal interaction compensation module receives feedback information from the dynamic noise robustness data processing module and provides visualization and tactile interaction compensation. It is bidirectionally connected to the dynamic noise robustness data processing module. The AR multimodal interaction compensation module includes an AR display auxiliary confirmation unit and a tactile feedback error correction unit. The AR display auxiliary confirmation unit is embedded in the AR glasses worn by the operator and maintains a real-time communication connection with the dynamic noise robustness data processing module. The processor of the AR display auxiliary confirmation unit is an ARM Cortex-A53 processor. It interacts with the dynamic noise robustness data processing module through the MIPI_DSI interface to display data, and the gesture correction signal is returned through the SPI interface. When the voice recognition unit recognizes the user's command text, it will generate a semi-transparent floating window in the AR glasses' field of view, displaying the current recognition result and multiple candidate words with high confidence, avoiding misjudgment of the work command due to a single voice recognition error. The floating window adopts a side-by-side arrangement or a vertical sliding list form to quickly complete the confirmation. If "pick up goods" is mistakenly identified as "place goods," the floating window will display relevant operation options such as "place goods," "pick up goods," and "return goods." Operators can simply tap the target word with their index finger or swipe left or right to make the selection, and the system will instantly receive the selected content and update the recognition result. This entire process forms an interactive closed loop of "voice input—AR confirmation—gesture correction," significantly improving recognition reliability and enhancing the system's usability under noise interference. This mechanism allows for user intervention to correct the recognition result before actual execution, ensuring the accuracy of task operations and effectively reducing the risk of mistransmission of instructions due to voice misrecognition. It is particularly suitable for industrial scenarios with high operational precision requirements, such as warehousing and equipment control.

[0076] The tactile feedback error correction unit is integrated into the AR glasses or smart gloves worn by the operator. It connects in real-time with the dynamic noise robust data processing module to construct an instant perception and intervention mechanism for industrial work sites. The tactile feedback error correction unit uses an STM32H7 series processor, receives device status conflict signals via the GPIO interface, and provides early warnings through vibrations at different frequencies. Simultaneously, it interacts with the dynamic noise robust data processing module via the UART interface to complete control command interactions. When the system detects that a voice command may cause a risk or misoperation, the tactile feedback error correction unit sends an early warning signal to the operator through physical vibration, guiding them to actively check whether their current operational intention matches the actual state. When the data processing module receives the voice recognition result and performs semantic understanding, it automatically retrieves device status information from the database module and compares it with the command. If a conflict is found, the tactile feedback mechanism is immediately triggered. For example, when an operator issues the command "Start machine tool No. 3," the system detects that machine tool No. 3 is currently in a fault-prone stop state before issuing the instruction. The tactile feedback unit then alerts the operator with two consecutive high-frequency, short vibrations, indicating a risk in executing the command. After receiving the vibration alert, the operator can choose to re-enter the command via voice or further review the equipment status information through the AR display auxiliary confirmation unit before making a judgment. The tactile feedback signal uses different combinations of vibration frequencies, durations, and intervals to encode various warning levels. For example, a single short vibration indicates a minor conflict, continuous vibration indicates a serious conflict, and intermittent vibration indicates an ambiguous state requiring manual confirmation. This achieves a more refined operational safety alert mechanism, improving the system's intelligent fault tolerance and operational robustness during human-machine interaction. This mechanism is particularly suitable for high-noise, high-complexity workshop or warehouse environments. In situations where audiovisual channels are limited or voice recognition errors increase, it forms a multimodal interactive closed loop that complements voice and vision, effectively avoiding misoperation, equipment damage, or safety hazards caused by misjudgment.

[0077] like Figure 3As shown, the dynamic noise robust data processing module is connected to the environment-aware voice interaction module, the AR multimodal interaction compensation module, the job management module, and the database module, respectively. It is used for noise reduction, semantic understanding, and closed-loop optimization of voice signals, and to generate feedback information and control commands. The dynamic noise robust data processing module includes a real-time noise modeling submodule, a recognition-noise reduction joint optimization submodule, a data integration submodule, an analysis and decision-making submodule, a command generation submodule, a semantic understanding unit, a speech recognition unit, and a speech preprocessing unit. The real-time noise modeling submodule works collaboratively with the environment-aware voice interaction module to dynamically model and suppress complex background noise in industrial environments. Its core mechanism is based on an improved Hidden Markov Model (HMM) for noise state transition modeling and parameter estimation, and integrates deep neural networks and long short-term memory (LSTM) network structures. The HMM models the characteristics of background noise changing over time through a state transition probability matrix and an observation probability function. The model form is as follows:

[0078]

[0079] Where P(O|λ) represents the observation sequence probability under model parameter λ, Q is the hidden state sequence, O is the observed noise feature sequence, P(Q|λ) is the state transition probability, and P(O|Q,λ) is the observation probability under a given state sequence. Combined with the Baum-Welch algorithm, the transition and output probabilities of the noise state can be updated online, achieving dynamic estimation of the noise power spectral density. This estimation result, along with the noisy speech signal acquired by the current microphone array, is input into a DNN-LSTM hybrid model. This DNN-LSTM hybrid model first extracts static spectral features using a deep neural network, then models the contextual relevance in the time series using a long short-term memory network. Its output is the estimated clean speech spectrum. The overall prediction process can be represented as:

[0080]

[0081] Where Y(f,t) represents the noisy speech spectrum, and N(f,t) represents the real-time noise spectrum characteristics. The model represents the predicted clean speech spectrum; during training, the goal is to minimize the reconstruction error, and the mean squared error loss function is used for optimization.

[0082]

[0083] Where S(f,t) is the actual speech spectrum. For the model output, T and F represent the number of time frames and the number of frequency points, respectively. Through this structure, the real-time noise modeling submodule can not only accurately adapt to the non-stationary noise characteristics in the workshop, but also effectively improve the quality of speech signal reconstruction, providing high-confidence speech input for the dynamic noise robustness data processing module.

[0084] The recognition-denoising joint optimization submodule, as a core component of the dynamic noise robustness data processing module, forms a closed-loop feedback path with the real-time noise modeling submodule and the semantic understanding unit. It drives the adaptive optimization of the denoising and recognition models through a confidence evaluation mechanism of the speech recognition results. The recognition-denoising joint optimization submodule generates confidence scores for the recognition results based on a connection-time classification loss function, which quantifies the degree of matching between the output text of the speech recognition unit and the target label. Its core loss function is:

[0085]

[0086] in, Let x represent the connection-time classification loss, y represent the input speech feature sequence, π represent the target text label sequence, and B represent all possible paths. -1 ( y ) Let P(π|x) represent the set of paths that can be mapped from the target sequence, and let P(π|x) be the probability of a path given the input. The recognition-denoising joint optimization submodule calculates the confidence score in real time and compares it with a set threshold. When the score is below 80%, the parameter adjustment strategy of the denoising algorithm is triggered, wherein the attenuation coefficient α of the spectral subtraction is adaptively increased to enhance the noise suppression capability. The spectral adjustment process is expressed as follows:

[0087]

[0088] in, Y(f,t) represents the corrected speech spectrum, and Y(f,t) represents the original noisy spectrum. To estimate the noise spectrum, α is the attenuation coefficient, and ∈ is the minimum threshold to avoid negative values. After completing error correction, the recognition-denoising joint optimization submodule combines the corrected text with the original speech to form training samples. It then updates the parameters of the recognition model and denoising network in real time through backpropagation to minimize prediction errors and enhance adaptability to similar noise scenarios. The overall AR intelligent auxiliary management system forms an online learning closed loop of "collection-denoising-recognition-feedback-optimization." The optimization process aims to minimize the joint loss function, expressed as:

[0089] L total =λ1L CTC +λ2L denoise

[0090] Among them, Ltotal For the joint loss function, L total To identify the loss, L denoise For noise reduction loss, λ1 and λ2 are adjustable weighting coefficients. Through the above mechanism, this submodule achieves robust processing and continuous evolution of unstructured speech data, significantly improving the recognition accuracy and interaction stability of the system in complex industrial environments.

[0091] The semantic understanding unit is connected to the recognition-denoising joint optimization submodule to perform semantic analysis on the optimized text and determine the specific content and intent of the work instructions. The data integration submodule retrieves cargo-related information from the database module, including the cargo's storage location and inbound / outbound processes. The analysis and decision-making submodule analyzes and processes the integrated data, determining the optimal inbound path and operation steps based on a preset work optimization algorithm. The instruction generation submodule generates corresponding feedback information and control instructions based on the results from the analysis and decision-making submodule.

[0092] Before leaving the factory, the system of this invention is pre-trained for different noise scenarios to generate scenario-specific noise reduction parameter configuration files (such as "workshop processing mode" and "warehouse handling mode"). Operators can switch the adaptation mode with one click through the AR interface. The dynamic noise robustness data processing module retrieves the corresponding noise reduction parameters according to the selected mode to process the noise.

[0093] The operation management module is connected to the dynamic noise robust data processing module and the production or storage equipment, and is used to schedule and manage the equipment according to control commands. The operation management module includes a status monitoring unit, a task scheduling unit, an equipment control unit, and a processor. The status monitoring unit is used to monitor the operating status of the storage equipment and collect its operating data. The task scheduling unit is used to arrange forklifts to assist operators in handling goods according to the control commands sent by the dynamic noise robust data processing module. The equipment control unit is used to generate specific equipment control signals to control the forklifts to travel along a specified path and accurately transport the goods to the storage location. The processor of the operation management module adopts an industrial-grade ARM Cortex-A72 and interacts with the dynamic noise robust data processing module via Ethernet using the Modbus TCP protocol to exchange control commands. When the processor of the operation management module is connected to the production or storage equipment, it adapts to industrial sensors through an RS-485 interface and controls the storage equipment using the CANopen protocol. The status monitoring data of the storage equipment is transmitted to the dynamic noise robust data processing module via WiFi through the MQTT protocol, forming a closed-loop management of the operation process.

[0094] The database module is connected to the dynamic noise robust data processing module and is used to store work processes, equipment information, product information, and scenario-based noise data. The database module contains a scenario-based noise database that collects typical industrial scenario noise, including various noise types and covering a sound pressure level range of 75-115dB. The typical industrial scenario noise includes at least lathe noise, warehouse rack handling noise, and pneumatic tool noise.

[0095] The environment-aware voice interaction module synchronously acquires voice signals, environmental noise time-domain waveforms, frequency-domain features, and spatial distribution data through a composite sensor array. After processing by the adaptive beamforming processing unit, it outputs an anti-noise voice signal to the dynamic noise robustness data processing module. The dynamic noise robustness data processing module performs real-time noise modeling, noise reduction processing, and confidence-based feedback-based recognition model optimization on the input anti-noise voice signal, outputting semantically understood work instructions and feedback information. The AR multimodal interaction compensation module receives the work instructions and feedback information and displays them visually. Simultaneously, it acquires correction instructions confirmed by the operator's gestures or tactile feedback and sends them back to the dynamic noise robustness data processing module. The dynamic noise robustness data processing module generates control instructions based on the correction instructions and sends them to the work management module. Meanwhile, the operator... The AR multimodal interaction compensation module prompts the corresponding warehousing operations; the operation management module controls the warehousing equipment to perform warehousing operations according to the control commands sent by the dynamic noise robustness data processing module; the status monitoring unit monitors the operating status of the warehousing equipment in real time, collects the operating data of the warehousing equipment, and transmits the data to the dynamic noise robustness data processing module; the dynamic noise robustness data processing module monitors and adjusts the operation process in real time based on the real-time collected data to ensure the smooth progress of the operation; when the warehousing operation is completed, the AR multimodal interaction compensation module confirms the completion of the warehousing operation by obtaining the operator's voice commands or tactile feedback and sends the confirmation information to the dynamic noise robustness data processing module; after receiving the confirmation information, the dynamic noise robustness data processing module updates the warehousing information and operation process records in the database module.

[0096] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0097] (1) Efficiency Improvement: This invention is driven by voice AI interaction. Operators only need to speak the work instructions to quickly identify and execute the corresponding operations, greatly reducing the time and workload of manual input. In warehousing scenarios, operators do not need to manually search for the location of goods; they can quickly obtain the location of goods and perform operations through voice commands, shortening the time for goods to enter and leave the warehouse, realizing automated scheduling and management of the work process, reducing manual intervention, and improving the continuity and smoothness of operations. In workshop production scenarios, the system can automatically arrange the processing tasks of production equipment according to production tasks and equipment status, avoiding errors and delays that may occur during manual scheduling, and improving production efficiency. In addition, this invention can also collect and process on-site data in real time, promptly identify problems and make adjustments, further improving operational efficiency.

[0098] (2) Improved Accuracy: The application of speech recognition and AR technology in this invention effectively reduces human error. The speech recognition unit uses advanced algorithms to accurately recognize the voice commands of operators, avoiding errors that may occur with manual input. The AR multimodal interaction compensation module intuitively displays the work information in the operator's field of vision, enabling the operator to clearly understand the work requirements and operating steps, reducing operational errors caused by misunderstandings of information. This invention has real-time verification and feedback functions, which can monitor and verify the operator's operations in real time, and promptly detect and correct errors. During the process of goods entering and leaving the warehouse, the quantity, model, and other information of the goods are verified in real time. If an error is found, a prompt will be issued immediately, requiring the operator to check and correct it, ensuring the accuracy of the operation.

[0099] (3) Optimized Interactive Experience: The natural voice interaction and immersive AR display of this invention make operation more convenient and intuitive, reducing the learning cost for operators. Operators do not need to learn complex operation interfaces and instructions; they can complete tasks simply by interacting with the system using natural language, thus improving their work enthusiasm and satisfaction. When using AR glasses for work, operators can experience an immersive interactive experience, as if virtual information is integrated with the real environment, enhancing the fun and immersion of the work. In addition, this invention also supports multimodal interaction, allowing operators to choose from various interaction methods such as voice, gestures, and touch according to their needs, improving the flexibility and convenience of interaction.

[0100] (4) Enhanced Real-Time Performance: This invention can collect various data from the work site in real time, including equipment status and work progress, and process and analyze this data in real time. Based on the real-time data, the system can provide timely decision-making support to managers, helping them to make quick decisions and adjust work plans. In workshop production, when equipment malfunctions, this invention can immediately detect it and feed the fault information back to managers. Managers can then promptly arrange for maintenance personnel to carry out repairs based on the information provided by the system, avoiding the impact of malfunctions on production. Furthermore, it can optimize the work process based on real-time data, improving work efficiency and quality.

[0101] The workflow of this invention will be explained in detail below using warehousing operations as an example:

[0102] like Figure 2As shown, the operator wears AR glasses and ensures the microphone array is functioning properly. When goods need to be stored, the operator issues a voice command through the AR multimodal interaction compensation module, such as "I want to store goods XX." The microphone array collects the operator's voice signal and transmits it to the adaptive beamforming processing unit for noise reduction. The signal is then input to the voice preprocessing unit for noise reduction, gain adjustment, and other preprocessing operations to improve the quality of the voice signal. Next, the voice recognition unit converts the preprocessed voice signal into a digital signal and uses a voice recognition algorithm to recognize the digital signal, obtaining the text-based operation command "I want to store goods XX," which is then transmitted to the semantic understanding unit of the data processing module. Upon receiving the operation command, the semantic understanding unit uses a natural language processing algorithm to perform semantic analysis on the text, determining that the specific content and intent of the operation command is a goods storage operation. Then, the data integration submodule retrieves information related to the XX goods from the database module, such as the storage location and warehousing process. Simultaneously, it receives real-time data collected by sensors at the work site, such as warehouse temperature, humidity, and occupancy of the storage area. The analysis and decision-making submodule analyzes and processes the integrated data, determining the optimal warehousing path and operation steps based on a preset operation optimization algorithm. The instruction generation submodule generates corresponding feedback information and control instructions based on the results of the analysis and decision-making submodule. The feedback information is communicated to the operators via the AR multimodal interaction compensation module, such as "Please move XX goods to shelf X on shelf X," and the storage location and warehousing path of the goods are displayed in AR format within the AR glasses' field of view. Following the prompts from the AR multimodal interaction compensation module, the operators move the goods to the designated location. Simultaneously, the task scheduling unit of the operation management module, based on control instructions sent by the dynamic noise robustness data processing module, controls forklifts and other equipment to assist the operators in moving the goods. The equipment control unit generates specific equipment control signals, controlling the forklifts to travel along the designated path and accurately move the goods to the warehousing location. The status monitoring unit monitors the operating status of equipment such as forklifts in real time, collects equipment operating data such as travel speed and cargo load weight, and transmits the data to the dynamic noise robustness data processing module. Based on this real-time data, the dynamic noise robustness data processing module monitors and adjusts the operation process in real time to ensure smooth operation. Once the goods are received into the warehouse, the operator confirms the completion of the warehousing operation through voice commands or tactile feedback from the AR multimodal interaction compensation module. Upon receiving the confirmation, the dynamic noise robustness data processing module updates the cargo inventory information and operation process records in the database module, completing the warehousing operation.

[0103] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Based on the disclosure and teachings of the above specification, those skilled in the art can also make changes and modifications to the above embodiments. Therefore, this invention is not limited to the specific embodiments disclosed and described above, and some modifications and changes to the invention should also fall within the scope of protection of the claims of this invention. Furthermore, although some specific terms are used in this specification, these terms are only for convenience of explanation and do not constitute any limitation on this invention.

Claims

1. A voice-AI interaction-driven AR-based intelligent auxiliary management system for operations, applied in intelligent warehousing or production fields, characterized in that: include: The system includes an environment-aware voice interaction module, an AR multimodal interaction compensation module, a dynamic noise robust data processing module, a job management module, and a database module. The environment-aware voice interaction module is used to collect the voice signals of the operators and suppress environmental noise. It is connected to the dynamic noise robust data processing module. The environment-aware voice interaction module includes a composite sensor array unit and an adaptive beamforming processing unit. The AR multimodal interaction compensation module is used to receive feedback information from the dynamic noise robustness data processing module and provide visualization and tactile interaction compensation. It is bidirectionally connected to the dynamic noise robustness data processing module. The dynamic noise robust data processing module is connected to the environment-aware voice interaction module, the AR multimodal interaction compensation module, the job management module, and the database module, respectively, and is used to perform noise reduction processing, semantic understanding, and closed-loop optimization on the voice signal, and generate feedback information and control commands. The operation management module is connected to the dynamic noise robust data processing module and the production equipment or storage equipment, and is used to schedule and manage the equipment according to control commands; the operation management module includes a status monitoring unit, which is used to monitor the operating status of the storage equipment and collect the operating data of the storage equipment; The database module is connected to the dynamic noise robust data processing module to store work processes, equipment information, product information, and contextualized noise data; The environment-aware voice interaction module synchronously collects voice signals, environmental noise time-domain waveforms, frequency-domain characteristics, and spatial distribution data through a composite sensor array. After processing by the adaptive beamforming processing unit, it outputs noise-resistant voice signals to the dynamic noise robust data processing module. The dynamic noise robustness data processing module performs real-time noise modeling, noise reduction processing, and confidence-based recognition model optimization on the input noise-resistant speech signal, and outputs semantically understood work instructions and feedback information; the AR multimodal interaction compensation module receives work instructions and feedback information and displays them visually, and at the same time obtains correction instructions confirmed by the operator's gestures or tactile feedback and sends them back to the dynamic noise robustness data processing module. The dynamic noise robustness data processing module generates control commands based on the correction instructions and sends them to the operation management module; simultaneously, operators perform corresponding storage operations based on prompts from the AR multimodal interaction compensation module; the operation management module controls the storage equipment to perform storage operations based on the control commands sent by the dynamic noise robustness data processing module; the status monitoring unit monitors the operating status of the storage equipment in real time, collects the operating data of the storage equipment, and transmits the data to the dynamic noise robustness data processing module; The dynamic noise robustness data processing module monitors and adjusts the operation process in real time based on the real-time collected data to ensure the smooth progress of the operation. When the warehousing operation is completed, the AR multimodal interaction compensation module confirms the completion of the warehousing operation by obtaining the voice command or tactile feedback of the operator and sends the confirmation information to the dynamic noise robustness data processing module. After receiving the confirmation information, the dynamic noise robustness data processing module updates the warehousing information and operation process records in the database module. The dynamic noise robustness data processing module includes a real-time noise modeling submodule, a recognition-noise reduction joint optimization submodule, a data integration submodule, an analysis and decision-making submodule, an instruction generation submodule, a semantic understanding unit, and a speech recognition unit; The real-time noise modeling submodule works in conjunction with the environment-aware voice interaction module to dynamically model and suppress complex background noise in industrial settings. Its core mechanism is based on an improved Hidden Markov Model (HMM) for noise state transition modeling and parameter estimation, integrating deep neural networks and long short-term memory (LSTM) network structures. The HMM models the characteristics of background noise changing over time using a state transition probability matrix and an observation probability function. The model form is as follows: in, Let Q represent the probability of the observed sequence under the model parameters λ, where Q is the hidden state sequence and O is the observed noise feature sequence. Let be the state transition probability. To estimate the observation probability under a given state sequence, the Baum-Welch algorithm is used to update the transition and output probabilities of noisy states online, enabling dynamic estimation of the noise power spectral density. This estimation result, along with the noisy speech signal acquired by the microphone array, is input into a DNN-LSTM hybrid model. This model first extracts static spectral features using a deep neural network, then models the contextual relevance in the time series using a long short-term memory network. Its output is the estimated clean speech spectrum. The overall prediction process can be represented as: in, This represents the spectrum of noisy speech. Indicates the real-time noise spectrum characteristics. The model is the predicted clean speech spectrum; during training, the goal is to minimize the reconstruction error, and the mean squared error loss function is used for optimization. in, The actual speech spectrum is represented by T and F, which are the number of time frames and the number of frequency points, respectively.

2. The AR-assisted management system for assignments based on voice AI interaction as described in claim 1, characterized in that: The environmentally aware voice interaction module includes a processor; the processor is a Rockchip RK3588; the composite sensor array unit includes a microphone array, an environmental noise sensor, and an orientation sensor connected to the processor via I2C and SPI interfaces; the environmental noise sensor includes a MEMS accelerometer and an ultrasonic sensor; the adaptive beamforming processing unit is connected to the composite sensor array unit via a parallel data bus to transmit noise spectrum and spatial distribution data in real time, and outputs the noise-resistant voice signal to the dynamic noise robustness data processing module via a USB 3.0 interface.

3. The AR-assisted management system for assignments based on voice AI interaction as described in claim 2, characterized in that: The microphone array adopts a multi-channel distributed layout, which enhances the operator's voice and suppresses surrounding interference through spatial filtering. Its core model is as follows: in, Indicates the output signal. The signal received by the m-th microphone is... The time delay of this channel is M, and the number of microphones is M; the MEMS accelerometer in the environmental noise sensor extracts the spectral characteristics of the vibration signal through Fourier transform, as shown in the formula: in, Represents frequency domain signals, The signal represents the vibration acceleration in the time domain, where f is the frequency and j is the imaginary unit. MEMS signals reflect the amplitude-frequency relationship of noise sources in mechanical structures. Ultrasonic sensors acquire sound source direction information through ranging and sound intensity localization methods; their basic localization model is the time difference of sound wave propagation method. Where θ represents the angle between the sound source and the front of the array, c is the speed of sound, Δt is the time difference between the left and right channels, and d is the distance between the two sensors.

4. The AR-assisted management system for assignments based on voice AI interaction as described in claim 3, characterized in that: The adaptive beamforming processing unit collaborates with the composite sensor array unit to construct a dynamic noise field modeling and target sound source localization mechanism by receiving the noise spectrum characteristics and spatial distribution information output in real time from the environmental noise sensor and the orientation sensor. It also dynamically calculates the weighting coefficients of the microphone array based on the minimum variance distortionless response algorithm, thereby forming a high-gain pickup beam pointing towards the operator's mouth. The core optimization objective is: Where w represents the microphone array weighting vector, R is the noise covariance matrix, and d is the steering vector in the target direction. Let w be the conjugate transpose; the analytical solution to this optimization problem is: in, The optimal weight vector determines the contribution weight of each microphone channel to the final output signal. The oral cavity spatial coordinates provided by the orientation sensor are used to construct the steering vector d in real time, while the spectral data sensed by the environmental noise sensor participates in constructing the covariance matrix R, thereby enabling the beamforming process to respond to the current sound field state. When the environmental noise sensor detects a strong noise event from a specific direction, the adaptive beamforming processing unit adjusts the array weighting strategy to form a main lobe between 30° and 60° in front, while simultaneously forming a strongly suppressed side lobe in the rear, thereby enhancing the target speech and suppressing the noise direction signal. Finally, the processed high signal-to-noise ratio speech signal is input in real time to the dynamic noise robust data processing module for subsequent semantic analysis and task instruction generation.

5. The AR-assisted management system for assignments based on voice AI interaction as described in claim 1, characterized in that: The recognition-denoising joint optimization submodule is a core component of the dynamic noise robust data processing module. Together with the real-time noise modeling submodule and the semantic understanding unit, it forms a closed-loop feedback path. It drives the adaptive optimization of the denoising and recognition models through a confidence evaluation mechanism of the speech recognition results. The recognition-denoising joint optimization submodule generates confidence scores for the recognition results based on a connection-time classification loss function. These scores quantify the degree of matching between the output text of the speech recognition unit and the target label. Its core loss function is: in, Let represent the connection-time classification loss, where x is the input speech feature sequence and y is the target text label sequence. For all possible paths, This represents the set of paths that can be mapped from the target sequence. Given the probability of a path under a given input; the recognition-denoising joint optimization submodule calculates the confidence score in real time and compares it with a set threshold. When the score is below 80%, the parameter adjustment strategy of the denoising algorithm is triggered. The spectrum adjustment process is represented as follows: in, This is the corrected speech spectrum. The original noisy spectrum, To estimate the noise spectrum, α is the attenuation coefficient. To avoid negative values, a minimum threshold is required. After error correction, the joint optimization submodule for recognition and noise reduction combines the corrected text with the original speech to form training samples, and updates the parameters of the recognition model and noise reduction network in real time through backpropagation. The overall AR intelligent auxiliary management system forms an online learning closed loop of "collection-noise reduction-recognition-feedback-optimization," with the optimization process aiming to minimize the joint loss function, expressed as: in, For the joint loss function, To identify the loss, For noise reduction loss, λ1 and λ2 are adjustable weighting coefficients; The semantic understanding unit is connected to the recognition-denoising joint optimization submodule to perform semantic analysis on the optimized text and determine the specific content and intent of the work instructions. The data integration submodule retrieves cargo-related information from the database module, including the cargo's storage location and inbound or outbound process. The analysis and decision-making submodule analyzes and processes the integrated data and determines the optimal warehousing path and operation steps based on a preset work optimization algorithm. The instruction generation submodule generates corresponding feedback information and control instructions based on the results of the analysis and decision-making submodule.

6. The AR-assisted management system for assignments based on voice AI interaction as described in claim 5, characterized in that: The AR multimodal interaction compensation module includes an AR display auxiliary confirmation unit and a tactile feedback error correction unit. The AR display auxiliary confirmation unit is embedded in the AR glasses worn by the operator and maintains a real-time communication connection with the dynamic noise robust data processing module. The processor of the AR display auxiliary confirmation unit is an ARM Cortex-A53 processor, which interacts with the dynamic noise robust data processing module through the MIPI_DSI interface to exchange display data, and the gesture correction signal is transmitted back through the SPI interface. Once the speech recognition unit recognizes the user's command text, it will generate a semi-transparent floating window in the AR glasses' field of view to display the current recognition result and multiple candidate words with high confidence, thus avoiding misjudgment of the operation command due to a single speech recognition error. The floating windows are arranged side-by-side or in a sliding list format, allowing for quick confirmation.

7. The AR-assisted management system for assignments based on voice AI interaction as described in claim 6, characterized in that: The tactile feedback error correction unit is integrated into the AR glasses or smart gloves worn by the operator. It connects in real-time with the dynamic noise robust data processing module to construct an instant perception and intervention mechanism for industrial work sites. The tactile feedback error correction unit uses an STM32H7 series processor, receives device status conflict signals via the GPIO interface, and provides early warnings through vibrations at different frequencies. Simultaneously, it interacts with the dynamic noise robust data processing module via the UART interface to complete control command interactions. When the tactile feedback error correction unit detects that a voice command may cause risk or misoperation, it sends an early warning signal to the operator through physical vibration, guiding them to actively check whether their current operational intention matches the actual state. When the data processing module receives the voice recognition result and performs semantic understanding, it automatically retrieves device status information from the database module and compares it with the command. If a conflict is found, the tactile feedback mechanism is immediately triggered.

8. The AR-assisted management system for assignments based on voice AI interaction as described in claim 1, characterized in that: The operation management module also includes a task scheduling unit, an equipment control unit, and a processor. The task scheduling unit controls the forklift to assist operators in handling goods according to control commands sent by the dynamic noise robustness data processing module. The equipment control unit generates specific equipment control signals to control the forklift to travel along a designated path and accurately transport goods to the inbound or outbound location. The processor of the operation management module uses an industrial-grade ARM Cortex-A72 and interacts with the dynamic noise robustness data processing module via Ethernet using the Modbus TCP protocol to exchange control commands. When the processor of the operation management module is connected to production equipment or warehousing equipment, it adapts to industrial sensors through an RS-485 interface and controls the warehousing equipment using the CANopen protocol. The status monitoring data of the warehousing equipment is transmitted to the dynamic noise robustness data processing module via WiFi through the MQTT protocol, forming a closed-loop management of the operation process.

9. The AR-assisted management system for assignments based on voice AI interaction as described in claim 1, characterized in that: The database module contains a scenario-based noise database that collects noise from typical industrial scenarios, including various noise types and covering a sound pressure level range of 75-115 dB. The typical industrial scenario noise includes at least lathe noise, warehouse rack handling noise, and pneumatic tool noise.

Citation Information

Patent Citations

  • Audio-visual voice noise reduction method based on multi-mode gating lifting model

    CN116013297A

  • Emotional speech synthesis method and device based on AI large model

    CN117174073A