A patrol robot voice interaction system with environmental perception capability

By using environmental voiceprint modeling and acoustic interference recognition technologies, combined with dual confirmation of voice commands and semantic coupling verification, the problem of patrol robots being accidentally triggered in high-noise environments has been solved, improving the robustness and control security of the voice recognition system.

CN121260164BActive Publication Date: 2026-08-25PUYUAN (ZHEJIANG) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511516760.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-08-25
Estimated Expiration
2045-10-23

AI Technical Summary

Technical Problem

In high-noise environments, the voice recognition system of patrol robots is prone to false triggering due to noise interference, leading to task interruption, incorrect response, and reduced security.

Method used

The system employs an environmental acoustic modeling module, an acoustic interference identification module, a command double confirmation module, a semantic coupling verification module, and a confidence dynamic adjustment module. It identifies noise risks through frequency domain redundancy analysis, sound pressure change gradient, and instantaneous energy mutation characteristics. Combined with the double confirmation and semantic coupling verification of voice commands, it dynamically adjusts the confidence threshold to ensure the validity and consistency of commands.

Benefits of technology

It significantly improves the robustness and false trigger defense capabilities of the speech recognition system in high-noise scenarios, enhances the safety of intelligent decision-making and control of robots in complex environments, and avoids untimely command execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121260164B_ABST
    Figure CN121260164B_ABST
Patent Text Reader

Abstract

The application discloses a patrol robot voice interaction system with environment sensing capability, and has the characteristics that it comprises an environment voiceprint modeling module, an acoustic interference identification module, an instruction double confirmation module, a semantic coupling verification module, a confidence dynamic adjustment module and an instruction behavior consistency filtering module; the environment voiceprint modeling module constructs an environment voiceprint real-time extraction module with a frequency domain redundancy analysis mechanism, continuously collects environment acoustic signals on a patrol path, and constructs a frequency band mapping vector based on the collected acoustic signals; the application has the advantages that it can accurately identify the voice mis-triggering risk in a high-noise environment, and improve the identification robustness through the semantic verification and double confirmation mechanism; meanwhile, the confidence dynamic adjustment and behavior consistency filtering are combined to realize the context rationality judgment of the control instruction, and ensure the safe and reliable execution of the instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and robot control technology, and in particular to a voice interaction system for a patrol robot with environmental perception capabilities. Background Technology

[0002] A patrol robot voice interaction system with environmental awareness refers to an intelligent robot system that integrates multimodal perception and natural language interaction. It can perceive surrounding environmental information (such as temperature, humidity, light, sound, and moving objects) in real time during patrol missions and intelligently adjust its voice interaction strategy based on environmental changes. This system typically integrates perception modules such as visual recognition, sound localization, LiDAR, and infrared sensing, combined with speech recognition and semantic understanding technologies. This enables the robot not only to understand human commands but also to provide semantic feedback, adjust tasks, and proactively communicate based on the current environmental state. It achieves intelligent response to abnormal events and human-machine collaborative decision-making, enhancing the intelligence, safety, and interactivity of patrols.

[0003] In high-noise environments such as industrial plants, underground utility tunnels, or areas with large machinery operations, patrol robots often encounter errors in interpreting command semantics when performing voice interaction functions. This is often due to background noise in external sound sources with frequency bands similar to the voice commands (e.g., the sound of impact drills, metallic clanging, high-frequency alarm tones). Particularly when the noise spectrum overlaps acoustically with specific high-authority commands (such as "stop," "execute emergency plan," "activate isolation procedure," etc.), the recognition system is prone to misinterpreting them as valid control commands, leading to erroneous responses. For example, the robot might suddenly interrupt its patrol process, forcibly switch to emergency mode, activate temporary lockdown procedures, or send false alarms to the control center without human intervention, severely impacting task continuity, area operational safety, and the stability of collaborative control. Existing voice recognition modules generally lack effective mechanisms to detect semantic false triggers in high-noise backgrounds, making it difficult to guarantee robustness in complex acoustic environments. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a voice interaction system for patrol robots with environmental perception capabilities.

[0005] To address the aforementioned technical problems, the technical solution of this invention is: a voice interaction system for a patrol robot with environmental awareness capabilities, comprising an environmental voiceprint modeling module, an acoustic interference recognition module, a command dual confirmation module, a semantic coupling verification module, a confidence dynamic adjustment module, and a command behavior consistency filtering module.

[0006] The environmental acoustic modeling module constructs a real-time environmental acoustic extraction module with a frequency domain redundancy analysis mechanism, continuously collects environmental acoustic signals along the patrol path, and constructs a frequency band mapping vector based on the collected acoustic signals.

[0007] The acoustic interference identification module constructs an acoustic interference discrimination model based on the frequency band mapping vector, extracts the sound pressure change gradient and instantaneous energy change features, and identifies areas with the risk of voice mis-triggering based on the features.

[0008] The dual confirmation module for voice commands is activated, limiting each key voice command to two consecutive recognitions, and only when the two recognition results are completely consistent is it determined to be a valid control command.

[0009] The semantic coupling verification module performs reverse matching verification of voice semantics for each control command to be executed, calculates the semantic coupling degree between the control command and the current environmental state features, and only when the coupling degree meets the set threshold condition, the command enters the next processing flow.

[0010] The confidence dynamic adjustment module dynamically calculates the speech recognition confidence threshold based on the interference intensity of the current recognition environment and the historical recognition accuracy, and judges the recognition validity of control commands based on the confidence threshold.

[0011] The command behavior consistency filtering module inputs the identified valid and semantically verified control commands into the voice behavior confirmation logic module. It performs logical filtering based on the continuity of the current patrol path and the preset behavior expectation state, and triggers the execution response only when the control command is consistent with the behavior expectation.

[0012] Preferably, the specific steps for collecting environmental acoustic signals along the patrol route and constructing a frequency band mapping vector based on the collected acoustic signals are as follows:

[0013] Deploy a multi-point acoustic acquisition sensor array and install multiple high-sensitivity wideband microphone modules around the main body of the patrol robot to continuously collect environmental acoustic signals;

[0014] The acquired acoustic signals are processed by short-time Fourier transform to extract the spectral statistical features of the frequency band sub-intervals and to calibrate the background frequency band weighting coefficients.

[0015] Construct a frequency band mapping algorithm model based on a graph structure to generate frequency band mapping vectors;

[0016] The frequency band mapping vector is stored in the environmental voiceprint cache unit and a voiceprint status database is constructed.

[0017] Preferably, the specific steps for accurately identifying speech-triggered risk areas through sound pressure change gradient and instantaneous energy change feature analysis are as follows:

[0018] Construct sound pressure time-series curves based on frequency band mapping vectors and calculate multi-scale sound pressure change gradient feature vectors;

[0019] Extract the instantaneous energy mutation features within a unit time period and construct an acoustic energy time series diagram by combining them with a time sliding window;

[0020] An acoustic interference discrimination model is constructed using a discrimination mechanism that combines rule-driven and learning-driven approaches, and acoustic interference risk level values ​​are generated.

[0021] Locate areas with moderate to high risk on the geographic path map and generate labels for areas with high false trigger risk.

[0022] Preferably, by performing multiple rounds of semantic comparison and consistency judgment on key control commands, the following steps are taken to ensure the reliability and consistency of each control action executed in a noisy background, and to prevent false commands from being issued due to brief sound interference or spectrum overlap:

[0023] Receives the high false trigger risk area label output by the acoustic interference identification module and switches to high sensitivity control mode;

[0024] The key control commands initially recognized are then recognized a second time within the speech buffer, and semantic consistency is compared.

[0025] The consistency confidence check of the recognition results is performed by combining the robot's current position, task status, and behavioral logic.

[0026] Based on the verification results, the pre-delay mechanism will be triggered or an invalid instruction will be output and a failure log will be recorded.

[0027] Preferably, the following steps are taken to calculate the "semantic coupling degree" index as the basis for judging the execution of control instructions, ensuring that control behavior can only continue under the conditions of logical rationality and contextual matching:

[0028] The control commands that are double-confirmed will be semantically vectorized.

[0029] Extract the current environment state parameters and construct the environment semantic vector;

[0030] A semantic similarity scoring mechanism is used to calculate the semantic coupling score between control commands and environmental states;

[0031] The coupling score is compared with a preset threshold to determine whether the control command is allowed to enter the subsequent processing flow.

[0032] Preferably, the specific steps for dynamically generating a speech recognition confidence threshold by quantifying the interference intensity of the current acoustic environment and combining it with the recognition performance data accumulated by the robot in different acoustic scenarios are as follows:

[0033] Extract the spectral distribution, sound pressure fluctuation characteristics, and instantaneous sound energy change frequency of the current area, and calculate the environmental interference intensity index;

[0034] Analyze historical speech recognition results and construct a statistical model for recognition accuracy;

[0035] The current confidence threshold is generated by integrating the environmental interference intensity index with historical recognition accuracy.

[0036] The identification confidence score is compared with the confidence threshold to determine the effectiveness of the control command identification.

[0037] Preferably, after the voice control command is recognized as valid and passes semantic verification, it is determined whether the current robot patrol path state has logical continuity to avoid the command interrupting the original path behavior. A path behavior vector is constructed based on the robot's historical path data and the current patrol trajectory, represented as the historical path behavior vector and the current path segment vector, respectively. The path state similarity parameter is obtained by calculating the cosine similarity between the two, and the calculation expression is as follows:

[0038] In the formula, It is a historical path behavior vector. It is the vector of the current path segment. It is the modulus of the historical path behavior vector. It is the modulus of the current path segment vector. It is path state similarity;

[0039] Based on the path behavior judgment, we analyze whether the currently identified control command conforms to the behavior expectation set in the task planning stage. Through semantic modeling, we represent the current control command as a control command semantic vector and the expected behavior corresponding to the current task stage as a behavior semantic vector. We calculate the degree of semantic deviation between the two to obtain the behavior expectation deviation parameter. The calculation expression is as follows:

[0040] In the formula, It is a behavioral semantic vector. It is a control instruction semantic vector. It is the magnitude of the behavioral semantic vector. It controls the magnitude of the semantic vector of the instruction. It is the deviation from expected behavior;

[0041] Path state similarity Deviation from behavioral expectations The inputs are used together as inputs for logical consistency evaluation, and the control logic consistency score is calculated through a weighted fusion method. The specific calculation formula is as follows: In the formula, It is a score for logical consistency. It is the path weight coefficient. It is the behavior matching weight coefficient. It is a control response flag. It is the threshold for logical consistency judgment.

[0042] The beneficial effects of this invention are:

[0043] This invention introduces environmental acoustic modeling and acoustic interference recognition modules to achieve dynamic perception and noise risk warning of the acoustic environment in which the patrol robot operates. The system can construct an acoustic model based on frequency band mapping vectors, accurately identify dangerous frequency bands in the environment with potential for false triggering, and dynamically locate high-risk areas for false triggering by combining features such as sound pressure gradient changes and instantaneous energy mutations. Furthermore, the system initiates a dual confirmation and semantic coupling verification mechanism for voice commands, ensuring that all critical control commands are only allowed to proceed to the next processing stage after passing triple verification of voice recognition, semantic understanding, and scene rationality. This significantly improves the robustness and false triggering prevention capabilities of the voice recognition system in high-noise environments.

[0044] This invention also achieves the self-consistency determination of the behavioral logic of voice control commands in complex patrol mission scenarios by constructing a dynamic confidence adjustment mechanism and a command behavior consistency filtering module. The system dynamically adjusts the confidence threshold based on the interference intensity in the recognition environment and the historical voice recognition accuracy. It performs a two-way semantic and path comparison between valid commands and the current patrol path state and expected task behavior, calculating a consistency score for command execution. Response commands are only allowed to be executed when the comprehensive score meets a preset safety threshold. This design effectively avoids the problem of robots executing inappropriate commands when behavioral logic is discontinuous or task states are mismatched, enhances the system's context awareness and task execution stability, and comprehensively improves the intelligent decision-making level and control safety of voice interaction control in high-risk industrial environments. Attached Figure Description

[0045] Figure 1 A schematic diagram of a patrol robot voice interaction system with environmental awareness capabilities; Detailed Implementation

[0046] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings. It should be noted that these descriptions are for the purpose of aiding understanding the present invention, but do not constitute a limitation thereof. Furthermore, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0047] The working principle of this invention: A voice interaction system for a patrol robot with environmental perception capabilities, including an environmental voiceprint modeling module, an acoustic interference recognition module, a command double confirmation module, a semantic coupling verification module, a confidence dynamic adjustment module, and a command behavior consistency filtering module.

[0048] The environmental acoustic modeling module constructs a real-time environmental acoustic extraction module with a frequency domain redundancy analysis mechanism, continuously collects environmental acoustic signals along the patrol path, and constructs a frequency band mapping vector based on the collected acoustic signals.

[0049] To effectively address the issue of accidental triggering when robots execute voice commands in complex acoustic environments, a real-time environmental acoustic signature extraction module with a frequency domain redundancy analysis mechanism is first constructed. This module continuously collects and analyzes background sound signals during patrol missions, generating frequency band mapping vectors that can be used for subsequent acoustic interference detection and dynamic adjustment of voice recognition. The specific steps are as follows:

[0050] First, a multi-point environmental acoustic acquisition sensor array is deployed, including multiple high-sensitivity broadband microphone modules, installed around the patrol robot in a spatially distributed structure. These microphone modules continuously acquire raw acoustic signals from the surrounding environment at preset time intervals (e.g., every 0.5 seconds) during the robot's patrol, simultaneously recording metadata such as sampling timestamps, spatial azimuth information, and microphone numbers. The raw acoustic data is processed by an analog-to-digital converter to form a time-stamped sequence of raw sound waves, ensuring a complete and traceable record of global acoustic changes along the patrol path.

[0051] Secondly, the acquired raw acoustic data undergoes frequency domain feature transformation processing. The Short Time Fourier Transform (STFT) algorithm is used to transform each segment of the raw signal data in the time and frequency domains, obtaining the spectral distribution of the acoustic signal within different time windows. To improve the robustness of the frequency domain representation, a frequency domain redundancy feature enhancement mechanism is introduced, specifically including steps such as frequency band sub-interval division, frequency amplitude statistics, and background frequency band weight coefficient calibration. Specifically, the frequency band sub-interval division step divides the complete spectrum into multiple frequency blocks according to a predefined frequency band range (e.g., every 100Hz as a group), and calculates the total energy and dominant frequency drift trend in each band. The frequency amplitude statistics step extracts statistical features such as peak amplitude, mean amplitude, and range for each frequency band, forming a multi-dimensional feature vector describing the dynamics of the spectrum. The background frequency band weight calibration step evaluates the probability and stability of frequency band occurrence based on a certain time sliding window (e.g., data within the last 5 minutes), assigning corresponding noise background weight indices to the frequency intervals, thereby forming a frequency domain feature distribution model expressing spatial acoustic redundancy.

[0052] Third, the frequency-domain processed acoustic feature vectors are input into a set of frequency band mapping algorithm models based on graph structures. These models analyze the temporal correlation and frequency coupling between features of different frequency bands, construct a multi-node frequency correlation graph, and further transform it into a "frequency band mapping vector." This vector is used to express the co-occurrence relationships and trends of typical frequency components in the current ambient sound field. For example, in continuous acoustic sampling, if a certain high-frequency band appears stably in multiple samplings and is accompanied by an increase in energy in a certain low-frequency band, it can be recorded as a "high-frequency-low-frequency coupling mode" in the frequency band mapping vector. This vector can also embed spatial directionality labels to indicate which sensor direction a specific frequency band originates from, providing spatial awareness support for subsequent acoustic interference source localization and regional risk assessment.

[0053] Finally, the frequency band mapping vector is output and temporarily stored in the environmental voiceprint cache unit, while a time-series-based voiceprint state database is constructed. Each frequency band mapping vector serves as a voiceprint sampling record, carrying information such as time tags, path node tags, and spectral feature tags, and can be used as input for subsequent processing units such as the acoustic interference discrimination model, semantic command coupling analysis module, and dynamic confidence adjustment module. This cache unit supports a real-time update mechanism, that is, as the robot moves, historical data from older time periods is continuously removed and the latest data is introduced, ensuring that the system maintains high timeliness and consistency in its understanding of the acoustic state of the current environment. Through the implementation of the above steps, not only can real-time structured perception of complex sound fields in the patrol path be achieved, but a solid data foundation and modeling support are also laid for subsequent high-risk area identification and voice command misrecognition prevention.

[0054] The acoustic interference identification module constructs an acoustic interference discrimination model based on the frequency band mapping vector, extracts the sound pressure change gradient and instantaneous energy change features, and identifies areas with the risk of voice mis-triggering based on the features.

[0055] To improve the stability and anti-interference capability of speech recognition in complex acoustic environments and avoid the problem of false speech triggering by robots in high-noise areas, this implementation method, based on obtaining the frequency band mapping vector, further constructs an acoustic interference discrimination model. Through the analysis of sound pressure change gradient and instantaneous energy mutation characteristics, it accurately identifies areas at risk of false speech triggering. The construction and recognition process includes the following steps:

[0056] First, a sound pressure time-series curve is constructed based on the frequency band mapping vector, and a multi-scale sound pressure change analysis mechanism is introduced. This mechanism continuously monitors the sound pressure value in each frequency band using fixed time windows (e.g., 1 second, 3 seconds, and 10 seconds), and calculates its first and second derivatives, representing the sound pressure change rate and acceleration, respectively. Through these derivative calculations, a "sound pressure change gradient feature vector" for each frequency band over a continuous time period is obtained, which can be used to describe whether there are abnormally drastic fluctuations or abnormal trends in the sound pressure of a specific frequency band. If a frequency band exhibits continuous steep increases, steep decreases, or unstable oscillations across multiple time window scales, it is initially marked as a potential interference frequency band and proceeds to the next step of judgment.

[0057] Secondly, a transient energy mutation feature extraction mechanism is introduced to assist in identifying sudden interference sources. This step, based on a short-time energy model, calculates the total acoustic energy value of the current ambient sound wave within a unit time period and constructs an "acoustic energy time sequence map" using a time sliding window. In the acoustic energy map, the system sets an adaptive energy mutation threshold by statistically analyzing the average energy level and standard deviation. Once the acoustic energy value is significantly higher than the background mean by more than two standard deviations within a certain time period, and the rate of change exceeds a preset range (e.g., a change of more than 50% within 0.5 seconds), it is considered a "transient energy mutation event." Combining this with previously marked potential interference frequency bands, the system further locates the specific time point and frequency range where the energy mutation occurs and identifies it as a "high-disturbance audio frequency block."

[0058] Third, an acoustic interference discrimination model is constructed. This model takes a frequency band mapping vector as input and uses sound pressure change gradient and energy mutation features as core criteria, employing a hybrid discrimination mechanism combining rule-driven and learning-driven approaches. The rule-driven part constructs a multi-condition matching logic judgment tree through a set of threshold rules (such as sound pressure variation rate threshold, frequency band stability coefficient, and energy mutation duration). The learning-driven part introduces a lightweight neural network (such as a convolutional neural network or a graph neural network) to train interference pattern classifiers under different sound field structures. Through this discrimination model, the system outputs in real time whether there is interference risk in the sensitive frequency bands for speech recognition in the current environment and generates an "acoustic interference risk level value" (e.g., divided into levels 0-3). Level 0 represents no risk, level 1 is a slight risk, level 2 is a moderate risk, and level 3 is a high-risk area. The discrimination model can be updated online by incorporating historical false triggering events, improving model adaptability and accuracy.

[0059] Finally, an acoustic risk spatial distribution map is constructed on the geographical path map of the patrol mission, and areas identified as having moderate or higher risks are spatially located and marked. The system generates "high-risk false triggering area labels" by combining acoustic interference risk level values ​​with the robot's current movement path, spatial coordinates, and task status. These labels provide environmental basis for the subsequent activation of the dual confirmation mechanism and coupling degree verification mechanism of the voice command module. These risk area labels have time-limited and location-specific attributes, automatically revoking after environmental changes stabilize or extending their validity period if acoustic characteristics continue to meet interference requirements. Furthermore, these labels are also synchronously uploaded to the environmental modeling database of the control center for multiple robots to share acoustic field risk information, thereby achieving coordinated patrol behavior and linked risk warnings. Through these steps, not only is high-precision identification of interference characteristics in complex acoustic environments achieved, but a context-driven active defense mechanism is also provided for the entire voice interaction system, effectively reducing the probability of false triggering of critical commands.

[0060] The dual-confirmation module for voice commands activates in areas where there is a risk of accidental voice triggering. It limits each key voice command to two consecutive recognitions, and only when the two recognition results are completely consistent is it determined to be a valid control command.

[0061] To prevent the robot from erroneously executing high-authority control commands due to misidentification in complex noisy environments, the system automatically activates a dual voice command confirmation module when it detects an area with a risk of accidental voice triggering. This module performs multiple rounds of semantic comparison and consistency checks on key control commands to ensure the reliability and consistency of every control action executed in a noisy environment, preventing erroneous commands due to brief sound interference or spectral overlap. Its specific implementation steps include the following four aspects:

[0062] First, based on the high-risk false triggering area label output by the preceding acoustic interference discrimination module, the system automatically switches the robot's voice interaction system currently in that area to "high-sensitivity control mode." In this mode, all recognized control commands must undergo multiple verification processes and are no longer allowed to be executed immediately. The system has a pre-set set of key command trigger words, such as "stop," "execute contingency plan," "start isolation," and "switch mode." This word list is set by the system's security policy management module and supports dynamic expansion. Any voice command belonging to this word list, once initially recognized in the current acoustic environment, enters a double confirmation process and cannot directly respond to control logic.

[0063] Secondly, the system initiates a continuous recognition mechanism, immediately reconstructing the key control commands recognized in the initial recognition within the speech buffer. To ensure consistency and stability, this step employs a time-locked speech signal caching mechanism, locking the audio signal segment relied upon by the initial recognition into a short-term buffer channel (e.g., 1 second long), and re-invoking the speech recognition engine for a second round of recognition without accepting new input. The recognition parameters here are set completely independently, and different threshold strategies can be used with the initial recognition model to enhance model stability. The dual recognition output results are rigorously verified by a semantic comparison algorithm. Only when the two recognition contents are completely consistent in word structure, semantic logic, and command keywords will the system proceed to the next judgment step. If the recognition results show any inconsistency in characters, word order, or command meaning, the control request is immediately marked as "not confirmed," and the system logs this, refusing to execute the response command.

[0064] Third, a context-based behavioral relevance weighting mechanism is introduced to verify the consistency and confidence of dual recognition results. Even if the text content of the two recognitions is completely identical, the system will still jointly judge the result with contextual information such as the current robot task status, location information, and task stage. For example, when the robot is patrolling a moving path and receives the instruction "stop," the system will assess whether there are necessary contextual factors such as path anomalies, obstacle warnings, or human intervention requests. If no relevant matching factors exist, the instruction can be judged as having insufficient confidence even if the voice is identical. This step introduces a semantic environment verification model, which evaluates the rationality of the instruction through logistic regression or deep matching algorithms and gives a confidence score. Only when the confidence score is higher than a set threshold (e.g., 0.8) is the voice instruction marked as a "highly consistent valid instruction"; otherwise, it is still considered an invalid request.

[0065] Finally, based on the final confirmation result, the system decides whether to write the voice command into the command buffer queue and pass it to the control execution layer module for execution. If the command is confirmed to be valid, the system will still add an "execution pre-delay" mechanism before execution, such as not accepting new commands for 0.3 seconds, and monitoring for sudden changes in sound signals or human interruption during this period to further improve control security. At the same time, the system will bind the confirmed command with a risk area label as a reference sample for recognition in similar scenarios in the future, and upload it to the global environment database to enhance the context recognition capability of the machine learning model. If the confirmation process fails, the system will trigger voice prompt feedback, such as "Command not recognized successfully, please repeat," and allow the user to reissue the command in the next time period. At the same time, the failure record will be synchronized to the backend analysis system for accidental triggering optimization training.

[0066] Through the above four steps, the voice command dual confirmation module can achieve multiple defenses against critical commands in high-interference scenarios through structured judgment and semantic comparison mechanisms, significantly improving the system's robustness and the security level of control decisions, and providing reliable voice interaction guarantees for the entire patrol mission.

[0067] The semantic coupling verification module performs reverse matching verification of voice semantics for each control command to be executed, calculates the semantic coupling degree between the control command and the current environmental state features, and only when the coupling degree meets the set threshold condition, the command enters the next processing flow.

[0068] To address the issue of control commands being misrecognized and erroneously executed in unexpected environments, the system constructs a speech-semantic reverse matching verification mechanism before critical control commands are triggered. This mechanism determines whether the currently recognized command and the environmental state have sufficient semantic consistency and behavioral rationality. The system calculates a "semantic coupling degree" index as the basis for judging the execution of control commands, ensuring that control behavior can only proceed under logically reasonable and context-matching conditions. The specific implementation steps are as follows:

[0069] First, upon receiving a control command verified through a dual voice confirmation mechanism, the system immediately extracts the semantic vector representation of the command. To this end, the system uses a natural language processing model (such as a semantic embedding model based on the Transformer architecture) to encode keywords, verb phrases, and contextual expressions from the speech recognition result, mapping the speech content into a context-aware semantic vector. This semantic vector not only reflects the semantic structure of the language text but also preserves the directionality of the command's intent. For example, "stop immediately" and "slow down" are grammatically similar but have different semantic strengths, and their vector distances will be distinguishable. The core of this step is that the embedding dimension of the semantic expression should be able to support subsequent comparison operations with multimodal environmental information.

[0070] Secondly, the system extracts multimodal environmental features related to voice commands from the current environmental state perception module, including but not limited to patrol path location information, surrounding obstacle detection status, current task execution stage, historical behavior records, noise level, and the presence of other voice sources. These state features are then transformed into structured semantic environment description vectors. This vector can be constructed using scene rule template extraction combined with knowledge graph modeling. For example, if the system detects that the robot is at the end of a closed path, there is high-frequency sound interference in the surrounding area, and the task log shows that a command has just been completed, then the scene will be modeled as a semantic state of "task completed + restricted environment + high interference," forming the corresponding environmental semantic vector.

[0071] Third, the system inputs the aforementioned instruction semantic vector and environmental state semantic vector into the semantic coupling degree calculation module, and performs semantic matching analysis using a semantic similarity scoring mechanism (such as cosine similarity or bidirectional matching neural network). The calculation result is a coupling degree score within the range [0,1]. The higher the value, the stronger the semantic consistency and the more the instruction conforms to the current environmental context. For example, if the "stop moving forward" instruction is issued when the robot is approaching an unknown obstacle and the environment in front is abnormal, the coupling degree score will usually be higher than 0.8; conversely, if the instruction occurs when the robot is autonomously patrolling in a safe passage without any warnings or emergencies, the coupling degree score will be lower than 0.3. The system pre-sets a dynamic coupling degree threshold as the boundary for determining whether to enter the next process. This threshold can be dynamically adjusted based on the task type, environmental complexity, and safety level. For example, it can be set to 0.6 in high-risk areas and relaxed to 0.4 in regular patrol phases.

[0072] Finally, the system determines whether the instruction can be executed based on the semantic coupling score and the current task strategy. If the score is higher than the threshold set for the current stage, the system will allow the control instruction to enter the subsequent confidence adjustment and behavior logic filtering process; if the score is lower than the threshold, the system will refuse to execute the instruction and trigger the voice feedback mechanism to prompt the user "The instruction is not applicable to the current state, please confirm and try again," and record the coupling failure log between the instruction and the environment state for subsequent model optimization and learning correction. In addition, the system also supports an instruction "pending confirmation queue." For instructions with critical scores (such as a difference of less than 0.05), they can be stored in the cache for re-verification in the next cycle, and the coupling degree can be re-evaluated when the environment changes. This design can avoid erroneously rejecting control behaviors that are actually reasonable after the environment changes, and enhance the system's context adaptation capability. Through the above four steps, the semantic reverse matching verification mechanism not only realizes the leap from speech recognition to semantic understanding, but also introduces "understanding the rationality of execution" into the voice control process, which greatly improves the intelligent response level and control safety boundary of the robot voice interaction system in complex environments.

[0073] The confidence dynamic adjustment module dynamically calculates the speech recognition confidence threshold based on the interference intensity of the current recognition environment and the historical recognition accuracy, and judges the recognition validity of control commands based on the confidence threshold.

[0074] To enhance the stability and robustness of the speech recognition system in complex and dynamic environments and further improve the security of control command execution, this implementation proposes a dynamic adjustment mechanism for speech recognition confidence based on environmental interference intensity and historical recognition accuracy. This mechanism dynamically generates a speech recognition confidence threshold by quantifying the interference intensity of the current acoustic environment and combining it with recognition performance data accumulated by the robot under different acoustic scenarios, thereby achieving intelligent judgment of the validity of the recognition results. The specific steps are as follows:

[0075] First, the system acquires the interference intensity parameters of the current environment. It extracts information such as the spectral distribution, sound pressure fluctuation characteristics, and frequency of instantaneous acoustic energy surges in the area where the patrol robot is located from the preceding acoustic interference discrimination module. Based on a multi-parameter fusion algorithm, it calculates an environmental interference intensity index, which expresses the degree of interference that the current acoustic environment may cause to speech recognition on a percentage basis. For example, in scenarios with high spectral overlap, drastic sound pressure changes, and frequent short-term energy surges, the interference intensity index can reach over 80 points; while in relatively quiet areas or areas with stable structural noise, the index may be below 30 points. This interference intensity index not only has a snapshot attribute of the current moment but also incorporates a time window statistically calculated trend slope to characterize the dynamic evolution trend of environmental interference, thus making the confidence calculation time-sensitive.

[0076] Secondly, historical recognition accuracy data is extracted, and a statistical model of speech recognition performance is constructed. The system archives and analyzes the robot's speech recognition results over a period of time, including core indicators such as recognition success rate, false recognition rate, and command execution failure rate. It also performs multi-dimensional statistical classification based on the recognition context, speech source characteristics, and command type, forming a historical accuracy model structured as "scene-command-recognition status." For example, the accuracy rate of the "stop" command is 62% in a high-noise construction site scenario, while it is 94% in an office environment. The system will retrieve the corresponding accuracy data from the statistical model based on the current environment and the type of recognition command, and evaluate its reliability range. Simultaneously, the model has a time-weighted decay mechanism, prioritizing data accumulated within the most recent period (e.g., 24 hours or 50 consecutive recognitions) to avoid the influence of changes in environmental conditions on historical data.

[0077] Third, a dynamic confidence threshold generation function is constructed by integrating the environmental interference intensity index and historical recognition accuracy. This function is a non-linear weighted function, and its core objective is to set an adaptive confidence threshold for the speech recognition results, preventing erroneous triggering of control logic when noise is severe or the recognition model lacks confidence. The output of this function is the "speech recognition confidence threshold at the current moment." For example, when the environmental interference intensity is 70 and the historical recognition accuracy is 60%, the system will calculate a higher confidence threshold (e.g., 0.85), requiring the speech recognition engine to achieve a confidence score of at least 85% for the instruction to be considered reliable. However, when the interference index is below 30 and the historical accuracy is above 90%, the confidence threshold may decrease to 0.65, thereby improving response sensitivity. This dynamic calculation mechanism can also be further modified based on factors such as the robot's current task priority and the risk level of the control instruction. For example, for instructions involving security lockouts, the threshold can be forcibly increased to enhance protection strength.

[0078] Finally, the validity of the current recognition result is determined based on the dynamically calculated confidence threshold. After recognizing a control command, the system obtains its corresponding confidence score from the speech recognition engine. This score is typically generated based on dimensions such as model consistency with the recognized content, semantic completeness, and acoustic matching. This score is compared with the current confidence threshold. If the recognition confidence score is equal to or higher than the threshold, the command is deemed valid and enters the behavior consistency filtering process. If the score is lower than the threshold, even if the recognized content is complete and error-free, it will be marked as a low-confidence command, will not enter the execution process, and will trigger a voice feedback mechanism to prompt the user, "The current environment has strong interference; please repeat the command." Simultaneously, this event will be recorded in the recognition performance log, updating the model's learning data as a reference for future optimization of the confidence function. Through this mechanism, the system not only improves the robustness of the speech recognition system in dynamic and complex environments but also significantly reduces the false recognition rate caused by environmental interference, further enhancing the autonomous control security and interactive intelligence level of the patrol robot.

[0079] The command behavior consistency filtering module inputs the control commands that are recognized as valid and have passed semantic verification into the voice behavior confirmation logic module. It performs logical filtering by combining the continuity of the current patrol path with the preset behavior expectation state, and triggers the execution response only when the control command is consistent with the behavior expectation.

[0080] After the voice control command is recognized as valid and passes semantic verification, it is first necessary to determine whether the current robot patrol path state has logical continuity to avoid the command interrupting the original path behavior. Based on the robot's historical path data and the current patrol trajectory, a path behavior vector is constructed, represented as the historical path behavior vector and the current path segment vector, respectively. By calculating the cosine similarity between the two, the path state similarity parameter is obtained. The calculation expression is as follows:

[0081] In the formula, It is a historical path behavior vector, representing the set of patrol path features executed by the robot within a certain period of time (such as the past 3 minutes or N path points). It is stored in vector form and can include encoded combinations of behavioral data such as path point coordinate sequences, travel direction, movement speed, rotation angle, and dwell time. It is the current path segment vector, representing the set of real-time behavioral data within the path segment where the robot is currently located. It is the modulus of the historical path behavior vector. It is the modulus of the current path segment vector. It is the path state similarity, which indicates the consistency of the current path behavior with the historical path trend in terms of direction, and the range is [0,1].

[0082] The vector modulus (also known as the norm of a vector) is a measure of the "length" of a vector in space, reflecting the amount of information, strength, or scale it carries. For an n-dimensional vector, its modulus is defined as the square root of the sum of the squares of its components. The role of the vector modulus is to provide a numerical scale, allowing for comparison, normalization, or judgment of "size" between different vectors. In speech behavior path judgment scenarios, the modulus plays a particularly important role, separating the directionality from the magnitude of a vector: by dividing the dot product of two vectors by the product of their respective moduli, the influence of their "absolute size" can be eliminated, retaining only the directional consistency, thus obtaining "cosine similarity." Simply put, in path or semantic vector processing, the vector modulus is a normalization benchmark for calculating similarity, ensuring that we are comparing "whether the behavioral trend is consistent," rather than "how big the steps are," which is crucial for the stability and accuracy of logical judgments.

[0083] The closer the value is to 1, the higher the consistency between the current patrol path and existing behavioral trends, and the stronger the path coherence; conversely, it indicates that the current path behavior deviates from the historical patrol logic, which may be a signal of miscontrol. This parameter serves as the first core input for behavioral logic judgment and is used in subsequent consistency score calculations.

[0084] Based on the path behavior judgment, we analyze whether the currently identified control command conforms to the behavior expectation set in the task planning stage. Through semantic modeling, we represent the current control command as a control command semantic vector and the expected behavior corresponding to the current task stage as a behavior semantic vector. We calculate the degree of semantic deviation between the two to obtain the behavior expectation deviation parameter. The calculation expression is as follows:

[0085] In the formula, This is a behavioral semantic vector, which represents the semantic representation of the "behavior to be performed" output by the task planning module at the current time point, patrol phase, or task context. It is generated by the patrol robot task scheduling system based on the task plan, path nodes, time period, and context state. For example, if the current phase is "turning inspection", It can correspond to behavioral targets such as "slow forward movement" and "turn detection". It is a semantic vector of control commands. This vector represents the semantic representation of user control commands identified by the speech recognition system. These are natural language commands issued by the user via voice and confirmed by the recognition system, such as "stop patrol," "activate isolation," and "advance five meters." It is the magnitude of the behavioral semantic vector. It controls the magnitude of the semantic vector of the instruction. It is the behavioral expectation deviation, which represents the degree of difference between the current voice command and the expected behavior. The value range is [0,1]. The closer the value is to 0, the higher the semantic consistency between the current control command and the expected task behavior; the closer the value is to 1, the more the command content deviates from the current task context.

[0086] This parameter is used to measure whether a voice command is "reasonable in terms of timing and task status," thereby further eliminating pseudo-reasonable commands that are semantically correct but "inappropriate in timing."

[0087] Path state similarity Deviation from behavioral expectations These factors, together, serve as inputs for the logical consistency assessment. A weighted fusion method is used to calculate the control logic consistency score. This comprehensive score represents the rationality of the control instruction within the current path and task context. The specific calculation formula is as follows: In the formula, The logical consistency score is the "overall judgment value" before the entire voice command can be executed. It comprehensively reflects whether the path state and semantic state match. The value range is 0-1. The higher the score, the more consistent it is with the current task context. These are path weight coefficients, used to control the impact of path state similarity on the total score. The influence ratio, ranging from 0 to 1, with a default recommended value of [value missing]. , This is the behavior matching weight coefficient, which controls the degree of influence of the expected behavior semantic matching degree on the total logical score. It is set to... ,satisfy The constraint relationship, It is a control response flag, a binary value, where 1 indicates that the instruction can be executed and 0 indicates that execution is prohibited. It serves as the final judgment signal for the instruction execution module, allowing only... When a command enters the control execution queue, When prompted, voice feedback may be given, such as "The command cannot be executed in the current state. Please confirm and try again." It is the threshold for judging logical consistency, the lower limit of the control system's tolerance for "logical rationality", with a value between 0.7 and 0.8, depending on the task's fault tolerance rate.

[0088] If the calculation yields If the value is above this threshold, the final response is marked. A value of 1 indicates that the control instruction can be allowed to enter the execution phase; otherwise, it is judged as logically inconsistent and execution is refused.

[0089] This judgment mechanism organically integrates multi-dimensional features such as the recognition result of control commands, the degree of semantic matching, and the path behavior logic, realizing a highly intelligent and adaptive voice command execution authorization mechanism. While ensuring the continuity of patrol paths and the consistency of task context, it effectively improves the accuracy and security of control responses.

[0090] This invention introduces environmental acoustic modeling and acoustic interference recognition modules to achieve dynamic perception and noise risk warning of the acoustic environment in which the patrol robot operates. The system can construct an acoustic model based on frequency band mapping vectors, accurately identify dangerous frequency bands in the environment with potential for false triggering, and dynamically locate high-risk areas for false triggering by combining features such as sound pressure gradient changes and instantaneous energy mutations. Furthermore, the system initiates a dual confirmation and semantic coupling verification mechanism for voice commands, ensuring that all critical control commands are only allowed to proceed to the next processing stage after passing triple verification of voice recognition, semantic understanding, and scene rationality. This significantly improves the robustness and false triggering prevention capabilities of the voice recognition system in high-noise environments.

[0091] This invention also achieves the self-consistency determination of the behavioral logic of voice control commands in complex patrol mission scenarios by constructing a dynamic confidence adjustment mechanism and a command behavior consistency filtering module. The system dynamically adjusts the confidence threshold based on the interference intensity in the recognition environment and the historical voice recognition accuracy. It performs a two-way semantic and path comparison between valid commands and the current patrol path state and expected task behavior, calculating a consistency score for command execution. Response commands are only allowed to be executed when the comprehensive score meets a preset safety threshold. This design effectively avoids the problem of robots executing inappropriate commands when behavioral logic is discontinuous or task states are mismatched, enhances the system's context awareness and task execution stability, and comprehensively improves the intelligent decision-making level and control safety of voice interaction control in high-risk industrial environments.

[0092] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and these variations still fall within the protection scope of the present invention.

Claims

1. A voice interaction system for a patrol robot with environmental perception capabilities, characterized in that, It includes an environmental acoustic signature modeling module, an acoustic interference recognition module, a command double confirmation module, a semantic coupling verification module, a confidence dynamic adjustment module, and a command behavior consistency filtering module. The environmental acoustic modeling module constructs a real-time environmental acoustic extraction module with a frequency domain redundancy analysis mechanism, continuously collects environmental acoustic signals along the patrol path, and constructs a frequency band mapping vector based on the collected acoustic signals. The acoustic interference identification module constructs an acoustic interference discrimination model based on the frequency band mapping vector, extracts the sound pressure change gradient and instantaneous energy change features, and identifies areas with the risk of voice mis-triggering based on the features. The dual confirmation module for voice commands is activated, limiting each key voice command to two consecutive recognitions, and only when the two recognition results are completely consistent is it determined to be a valid control command. The semantic coupling verification module performs reverse matching verification of voice semantics for each control command to be executed, calculates the semantic coupling degree between the control command and the current environmental state features, and only when the coupling degree meets the set threshold condition, the command enters the next processing flow. The confidence level dynamic adjustment module dynamically calculates the speech recognition confidence threshold based on the interference intensity of the current recognition environment and the historical recognition accuracy, and judges the recognition validity of control commands based on the confidence threshold. The command behavior consistency filtering module inputs the control commands that are recognized as valid and have passed semantic verification into the voice behavior confirmation logic module. It performs logical filtering by combining the continuity of the current patrol path with the preset behavior expectation state, and triggers the execution response only when the control command is consistent with the behavior expectation. After the voice control command is recognized as valid and passes semantic verification, it is determined whether the current robot patrol path state has logical continuity to avoid the command interrupting the original path behavior. Based on the robot's historical path data and the current patrol trajectory, a path behavior vector is constructed, represented as the historical path behavior vector and the current path segment vector, respectively. The path state similarity parameter is obtained by calculating the cosine similarity between the two, and the calculation expression is as follows: In the formula, It is a historical path behavior vector. It is the vector of the current path segment. It is the modulus of the historical path behavior vector. It is the modulus of the current path segment vector. It is path state similarity; Based on the path behavior judgment, we analyze whether the currently identified control command conforms to the behavior expectation set in the task planning stage. Through semantic modeling, we represent the current control command as a control command semantic vector and the expected behavior corresponding to the current task stage as a behavior semantic vector. We calculate the degree of semantic deviation between the two to obtain the behavior expectation deviation parameter. The calculation expression is as follows: In the formula, It is a behavioral semantic vector. It is a control instruction semantic vector. It is the magnitude of the behavioral semantic vector. It controls the magnitude of the semantic vector of the instruction. It is the deviation from expected behavior; Path state similarity Deviation from behavioral expectations The inputs are used together as inputs for logical consistency evaluation, and the control logic consistency score is calculated through a weighted fusion method. The specific calculation formula is as follows: In the formula, It is a score for logical consistency. It is the path weight coefficient. It is the behavior matching weight coefficient. It is a control response flag. It is the threshold for logical consistency judgment.

2. The voice interaction system for a patrol robot with environmental perception capability according to claim 1, characterized in that, The specific steps for collecting environmental acoustic signals along the patrol route and constructing a frequency band mapping vector based on the collected acoustic signals are as follows: Deploy a multi-point acoustic acquisition sensor array and install multiple high-sensitivity wideband microphone modules around the main body of the patrol robot to continuously collect environmental acoustic signals; The acquired acoustic signals are processed by short-time Fourier transform to extract the spectral statistical features of the frequency band sub-intervals and to calibrate the background frequency band weighting coefficients. Construct a frequency band mapping algorithm model based on a graph structure to generate frequency band mapping vectors; The frequency band mapping vector is stored in the environmental voiceprint cache unit and a voiceprint status database is constructed.

3. The voice interaction system for a patrol robot with environmental perception capability according to claim 1, characterized in that, The specific steps for accurately identifying speech-triggered risk areas through sound pressure change gradient and instantaneous energy mutation feature analysis are as follows: Construct sound pressure time-series curves based on frequency band mapping vectors and calculate multi-scale sound pressure change gradient feature vectors; Extract the instantaneous energy mutation features within a unit time period and construct an acoustic energy time series diagram by combining them with a time sliding window; An acoustic interference discrimination model is constructed using a discrimination mechanism that combines rule-driven and learning-driven approaches, and acoustic interference risk level values ​​are generated. Locate areas with moderate to high risk on the geographic path map and generate labels for areas with high false trigger risk.

4. The voice interaction system for a patrol robot with environmental perception capability according to claim 1, characterized in that, By performing multiple rounds of semantic comparison and consistency judgment on key control commands, the reliability and consistency of each control action executed in the background of noise interference are ensured, and the false commands are prevented due to brief sound interference or spectrum overlap. The specific steps are as follows: Receives the high false trigger risk area label output by the acoustic interference identification module and switches to high sensitivity control mode; The key control commands initially recognized are then recognized a second time within the speech buffer, and semantic consistency is compared. The consistency confidence check of the recognition results is performed by combining the robot's current position, task status, and behavioral logic. Based on the verification results, the pre-delay mechanism will be triggered or an invalid instruction will be output and a failure log will be recorded.

5. The voice interaction system for a patrol robot with environmental perception capability according to claim 1, characterized in that, By calculating the "semantic coupling degree" index as the basis for judging the execution of control instructions, the specific steps to ensure that control behavior can only continue to proceed under the conditions of logical rationality and contextual matching are as follows: The control commands that are double-confirmed will be semantically vectorized. Extract the current environment state parameters and construct the environment semantic vector; A semantic similarity scoring mechanism is used to calculate the semantic coupling score between control commands and environmental states; The coupling score is compared with a preset threshold to determine whether the control command is allowed to enter the subsequent processing flow.

6. The voice interaction system for a patrol robot with environmental perception capability according to claim 1, characterized in that, The specific steps for dynamically generating a speech recognition confidence threshold by quantifying the interference intensity of the current acoustic environment and combining it with the recognition performance data accumulated by the robot under different acoustic scenarios are as follows: Extract the spectral distribution, sound pressure fluctuation characteristics, and instantaneous sound energy change frequency of the current area, and calculate the environmental interference intensity index; Analyze historical speech recognition results and construct a statistical model for recognition accuracy; The current confidence threshold is generated by integrating the environmental interference intensity index with historical recognition accuracy. The identification confidence score is compared with the confidence threshold to determine the effectiveness of the control command identification.

Citation Information

Patent Citations

  • Vehicle-mounted microphone with speech recognition and control functions

    CN106218557A

  • Voice interaction method and system of AI intelligent robot

    CN120673768A