Multi-modal embodied interaction method for outdoor cleaning robot
Patent Information
- Application Number
- CN202610646574.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-12
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-05-12
AI Technical Summary
[0007]1、机器人交互方式单一,缺乏自然对话能力与上下文感知能力;
[0070] 1. Improve the naturalness of interaction: By synergistically integrating voice wake-up, speech recognition, knowledge base question answering, and edge language generation, the robot can complete natural language interaction in outdoor scenarios, thereby improving the user experience;
Smart Images

Figure CN122172977B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of robot interaction and autonomous control technology, specifically relating to a multimodal embodied interaction method for an outdoor cleaning robot. Background Technology
[0002] With the continuous development of artificial intelligence, embedded computing, and mobile robotics technologies, outdoor cleaning robots have been gradually applied in open environments such as parks, residential areas, squares, and campuses to perform tasks such as road sweeping, garbage collection, and inspection reminders. Compared with indoor service robots, outdoor cleaning robots face challenges such as complex environmental noise, high pedestrian mobility, high continuity of work tasks, and unstable network conditions. Therefore, higher requirements are placed on the naturalness, real-time performance, and reliability of human-computer interaction methods.
[0003] Existing outdoor cleaning robots typically employ preset announcements, remote terminal control, or simple voice command triggering for human-machine interaction. These methods suffer from limited interaction modes, insufficient contextual understanding, and weak behavioral coordination capabilities. On one hand, existing robots mostly use fixed-template voice announcements, only outputting standardized prompts under preset conditions. They lack comprehensive perception of user identity, environmental status, and interaction scenarios, making it difficult to generate natural, coherent, and targeted responses based on actual situations. On the other hand, in dynamic operating environments, while robots can avoid obstacles, they usually only passively adjust their movements, unable to proactively explain their intentions, detour directions, or next actions to pedestrians. This makes it difficult for pedestrians to predict the robot's trajectory, reducing human-machine collaboration efficiency and even causing path conflicts.
[0004] Furthermore, some existing methods rely on remote servers or large cloud-based models to perform speech understanding and text generation for certain intelligent interaction capabilities. While these methods have certain advantages in semantic expression, they are easily affected by factors such as network latency, network interruptions, privacy protection, and operating costs in outdoor scenarios, making it difficult to meet the robot's requirements for low latency, high reliability, and local autonomy during continuous operation. If control is entirely based on local fixed rules, it will result in insufficient interaction capabilities, making it difficult to handle open-domain question answering, colloquial expressions, or complex management commands.
[0005] In terms of operation and maintenance, robots currently rely on back-end terminals, mobile controllers, or dedicated host computers for status queries and task management. On-site management personnel find it difficult to operate the robots directly through natural language, resulting in insufficient maintenance efficiency and ease of interaction.
[0006] In summary, the existing technologies have the following problems:
[0007] 1. The robot's interaction methods are limited, lacking natural dialogue capabilities and context awareness;
[0008] 2. The robot lacks a proactive behavior prediction mechanism in dynamic work scenarios, resulting in poor human-robot collaboration;
[0009] 3. The intelligent interaction capabilities rely excessively on cloud computing, making it difficult to meet the requirements of real-time performance, reliability, and privacy in outdoor scenarios;
[0010] 4. Management personnel have limited on-site maintenance capabilities and cannot efficiently control the robot and query its status through natural voice.
[0011] 5. Multi-source audio broadcasts are prone to conflicts, and critical security information cannot be prioritized for delivery.
[0012] 6. In outdoor environments with noise interference, multiple pedestrians, and dynamic open spaces, robots are prone to making incorrect responses or erroneous actions due to false wake-up, voice misrecognition, or misunderstanding of the scene, which affects the reliability of interaction and the safety of execution. Summary of the Invention
[0013] To address the shortcomings of existing technologies and achieve the goal of enabling robots to maintain stable and efficient operation in complex outdoor environments, as well as to conduct natural, proactive, and reliable on-site interactions with pedestrians and management personnel, this invention adopts the following technical solution:
[0014] A multimodal embodied interaction method for outdoor cleaning robots, applied to outdoor cleaning robots, includes the following process:
[0015] The interaction mode is activated through multimodal perception; voice detection is used to judge human voices, and public interaction mode and management interaction mode are triggered according to different wake words; pedestrians are perceived by the collected environmental image sequence, and active safety prompt mode is triggered according to the spatial relationship between pedestrians and robots.
[0016] By using a speech recognition link with speech activity detection and speech endpoint detection, speech content is recognized after the wake word to obtain text recognition results;
[0017] The interaction intent and execution conditions are verified through multimodal fusion confirmation; a fusion confirmation score is performed based on the confidence level of the text recognition result, the spatial relationship, and the consistency between the text recognition result and the interaction mode and the robot's operating state, respectively.
[0018] When the fusion confirmation score meets the requirements, intelligent path decision-making generates response content or robot action instructions; if it is the public interaction mode, it responds according to the question; if it is the management interaction mode, it generates robot actions or outputs clarification prompts according to management commands.
[0019] Furthermore, the robot uses a voice wake-up channel to perform frame segmentation, windowing, and frequency domain analysis on the acquired raw audio stream. Based on the background noise estimation results, it suppresses steady-state wind noise, mechanical noise, and broadband interference in the environment to obtain an enhanced voice signal. Subsequently, it performs voice activity detection and wake-up word detection based on the enhanced voice signal.
[0020] Furthermore, the speech activity detection employs a method based on a joint decision of frame-level statistical features and temporal continuity to determine whether each audio frame contains human voices; the frame-level statistical features include one or more of short-time energy, spectral entropy, zero-crossing rate, and dominant frequency band energy distribution.
[0021] Furthermore, the robot detects, associates, and continuously tracks pedestrian targets. Based on the positional changes of the target detection box in consecutive image frames, combined with camera calibration parameters and the robot's own pose information, it estimates the relative distance, relative orientation, and movement trend between the pedestrian and the robot.
[0022] Let the position of the pedestrian target in the robot's coordinate system at a certain moment be... The relative distance between the pedestrian and the robot It can be represented as:
[0023]
[0024] Corresponding relative azimuth angle It can be represented as:
[0025]
[0026] The motion trend is obtained by acquiring the position of the pedestrian target in the robot coordinate system in two adjacent frames to obtain the relative motion vector of the pedestrian. The robot combines the relative motion vector of the pedestrian with the robot's planned motion vector to obtain the angle between the pedestrian's motion direction and the robot's planned motion direction, so as to generate an active safety warning risk score. When the score is greater than the threshold, the active safety warning mode is triggered.
[0027] Let the positions of the pedestrian target in the robot coordinate system in two adjacent frames be respectively and Then the pedestrian's relative motion vector for:
[0028]
[0029] Robot combines pedestrian relative motion vectors Planning motion vectors for robots Calculate the risk score for proactive safety alerts. :
[0030]
[0031] in, This represents the angle between the pedestrian's direction of movement and the robot's planned direction of movement. and These are the weighting coefficients. To prevent tiny constants with a denominator of zero.
[0032]
[0033] When the above formula is satisfied, the system triggers an active safety alert mode, where, The threshold for triggering safety prompts;
[0034] When a pedestrian is detected entering a preset safe area and the proactive safety warning risk score is greater than a threshold, the robot triggers the proactive safety warning mode; and generates a warning message containing detour direction, work intention or avoidance reminder based on the pedestrian's relative position, movement direction, the robot's current planned path and preset safe distance threshold.
[0035] The safe zone can be set as a fan-shaped area, a circular area, or a forward risk area that is dynamically adjusted according to the current direction of movement, centered on the robot body.
[0036] Furthermore, the voice activity detection process divides the audio signal collected after the wake-up word into frames with a fixed frame length, and obtains the short-time energy and zero-crossing rate of each frame. A voice activity determination function is constructed based on the short-time energy and zero-crossing rate. When the short-time energy is greater than the energy threshold and the zero-crossing rate is less than the zero-rate threshold, the current frame is determined to be a voice frame; otherwise, it is a non-voice frame.
[0037]
[0038] in, Indicates the first Frame audio signal, M represents the number of frames. This indicates the short-term energy of the frame;
[0039] To enhance the ability to distinguish between non-speech noise and silent segments, the calculation of the first... Frame zero-crossing rate :
[0040]
[0041] in, It is a symbolic function;
[0042] Based on short-time energy and zero-crossing rate, the system constructs a speech activity determination function:
[0043]
[0044] in, Energy threshold The zero-crossing rate threshold, This indicates that the current frame is a speech frame. This indicates that the current frame is a non-speech frame.
[0045] Furthermore, the speech endpoint detection determines the speech start point when multiple consecutive frames are identified as speech frames, and determines the speech end point when multiple consecutive frames are not speech frames, in order to extract complete and valid speech segments; pre-emphasis, framing, and windowing processing are performed on the speech segments, and filter bank energy features are extracted.
[0046] Based on the frequency domain of a single frame audio signal and the response coefficient of a single filter at the frequency index, the filter bank energy is obtained by accumulating the frequency points and taking the logarithm. The resulting logarithmic filter bank energy is input into a streaming speech recognition network for frame-by-frame decoding, and the text recognition result is output.
[0047] Furthermore, speech with high confidence is identified as reliable speech through speech confidence, and a speech recognition confidence score is generated.
[0048] The spatial relationships are used to confirm the visual presence of pedestrians within the interaction area, so as to generate a visual presence consistency score.
[0049] Pedestrian interaction intent is determined by identifying pedestrian orientation information and movement trends from the environmental images, thereby generating an orientation and interaction intent score.
[0050] When the public interaction mode is triggered, the text recognition result is verified according to the public Q&A rules. When the management interaction mode is triggered, it is further confirmed whether the current voice input meets the management control conditions and an permission or pattern matching score is generated. If the interaction mode and the text recognition result are inconsistent, a pattern mismatch prompt or a request for reconfirmation is output.
[0051] In the management interaction mode, the recognized management control commands are verified to be consistent with the action intent involved in the text recognition results and the current robot operating state, and a consistency score between the command and the scene state is generated. When the recognition results conflict with the robot's current state, task stage, safety conditions or execution conditions, the system outputs a clarification prompt or a refusal to execute prompt to prevent miscontrol.
[0052] The fusion confirmation score is a weighted sum of the speech recognition credibility score, visual presence consistency score, orientation and interaction intent score, permission or pattern matching score, and command and scene state consistency score.
[0053] When the identified text corresponds to a high-risk management action, motion control command, or task switching command, the fusion confirmation threshold is increased to further reduce the risk of erroneous actions caused by misidentification.
[0054] Furthermore, in the public interaction mode, the user question is obtained from the text recognition result and input into the semantic decision module. Local public knowledge base matching is performed, the encoding distance between the user question text and the preset question text is calculated to obtain a literal similarity score, and combined with the keyword overlap rate to obtain an enhanced literal matching score. Semantic vectors of the user question text and the preset question text are generated using a semantic encoding model, and a semantic similarity score is obtained through similarity. The enhanced literal matching score and the semantic similarity score are weighted and fused to obtain a comprehensive matching score. If the score reaches a threshold, it is determined that the user question matches a candidate question in the local public knowledge base, and the standard answer is retrieved according to the candidate question for response. Otherwise, the process switches to the terminal-side language generation path, and the user question text and historical dialogue context are input into the language generation model to generate natural language that conforms to the current context for response.
[0055] The lightweight language generation model employs one or more of the following techniques: parameter compression, low-bit quantization, distillation transfer, efficient parameter adaptation, and computation graph optimization, in order to adapt to the limited computing resources on the robot's edge and reduce inference latency.
[0056] The system introduces a content constraint mechanism during the language generation process to screen the generated content for security, friendliness, and task relevance in order to avoid outputting inappropriate content.
[0057] When the overall matching score of the knowledge base falls within a preset fuzzy range, the system prioritizes generating a clarifying response rather than directly providing a definitive answer.
[0058] The response text is generated by the edge speech synthesis module. The prosodic structure of the output speech is adjusted by combining the pause position of the response text, keyword stress and speech rate control parameters to improve the naturalness and clarity of the broadcast. In order to adapt to the high noise environment outdoors, the broadcast volume and broadcast speed are adaptively adjusted according to the ambient noise intensity. When the length of the response text exceeds the preset threshold, the speech synthesis module adopts a streaming output method of generating and playing at the same time to shorten the first packet broadcast delay.
[0059] Furthermore, in the management interaction mode, the management command text is obtained from the text recognition result and input into the management command parsing module. Local preset action library matching is performed, and the encoding distance between the management command text and the preset command text is calculated to obtain a literal similarity score. Combined with the keyword overlap rate, an enhanced literal matching score is obtained. A semantic encoding model is used to generate semantic vectors for the management command text and the preset command text respectively, and a semantic similarity score is obtained through similarity. The enhanced literal matching score and the semantic similarity score are weighted and fused to obtain a comprehensive matching score. If the score reaches a threshold, it is determined that the management command matches a candidate command in the local preset action library, and a standard action sequence is retrieved and executed according to the candidate command. Otherwise, a prompt message indicating that the command cannot be recognized is output, and the robot's current operating state remains unchanged to prevent miscontrol.
[0060] The action sequence includes at least one or more of the following: start cleaning, pause task, return to charging, standby at a fixed point, status announcement, and task switching.
[0061] The action sequence is parsed by the robot's underlying control module according to the preset control protocol, and drives the mobile chassis, cleaning mechanism, obstacle avoidance mechanism or status display module to perform the corresponding tasks. Before the action is executed, the current safety status of the robot is checked. If low battery, obstacle avoidance conflict, execution channel occupation or permission mismatch is detected, the corresponding action is delayed or refused to be executed.
[0062] For voice prompts, the local speech synthesis module is invoked to generate corresponding audio, and the volume and speed of the audio are adaptively adjusted based on the current ambient noise level.
[0063] Furthermore, unified intelligent scheduling and broadcasting of multi-source audio are implemented;
[0064] A unified audio playback management module is adopted to unify the multi-source audio formats of interactive response audio, proactive safety prompt audio, and robot status broadcast audio. The formatting process includes at least one or more of the following: sampling rate unification, quantization bit width conversion, channel mapping, buffer management, and target playback format encapsulation, to ensure the compatibility of multi-source audio in the same playback link.
[0065] The broadcast requests are categorized and scheduled according to preset priority rules, with the priority order being: proactive safety prompt audio and emergency alarm audio, public interaction response audio, and robot status broadcast audio.
[0066] The scheduling priority score is calculated based on the audio event type, urgency, waiting time, and repetition rate, and the audio request with the highest score is played first. If the current playback channel is idle, the newly arrived audio request is played immediately. If the current playback channel is busy, the newly arrived audio request enters the waiting queue, and the queuing order is determined according to the priority score. A preemption threshold is set to adjust the sorting. After the high-priority audio finishes playing, playback is resumed based on the remaining duration and original priority of the interrupted audio.
[0067] The system also includes a repetitive broadcast suppression mechanism. For prompt audios triggered repeatedly from the same event source within a preset time window, the system performs processes such as merging and deduplication, delaying broadcast, or suppressing broadcast to reduce the interference of invalid repetitive reminders on users.
[0068] The system also includes an audio segment integrity protection mechanism, which supports resuming playback from the interrupted position or replaying from the current semantic integrity boundary for low-priority audio that has been interrupted due to preemption.
[0069] The advantages and beneficial effects of this invention are as follows:
[0070] 1. Improve the naturalness of interaction: By synergistically integrating voice wake-up, speech recognition, knowledge base question answering, and edge language generation, the robot can complete natural language interaction in outdoor scenarios, thereby improving the user experience;
[0071] 2. Enhance proactive collaboration capabilities: By introducing visual perception and proactive safety prompt mechanisms, the robot can proactively explain its movement intentions and detour directions when pedestrians approach, reducing the risk of human-robot path conflicts;
[0072] 3. Reduced cloud dependency: By deploying capabilities such as speech recognition, semantic matching, language generation, and speech synthesis on the edge computing platform, this invention can run stably in offline or weak network environments, improving real-time performance, reliability, and privacy security.
[0073] 4. Balancing efficiency and flexibility: By setting up a dual-channel decision-making mechanism of knowledge base response path and edge language generation path, it can quickly respond to high-frequency questions and handle open-ended questions and context-related dialogues.
[0074] 5. Improve management convenience: By setting up a second wake-up word for management personnel and a management command parsing mechanism, on-site maintenance personnel can directly control robot tasks using natural voice, thereby improving management efficiency;
[0075] 6. Ensure priority delivery of safety information: Through unified management of multi-source audio, priority scheduling, and preemption recovery mechanisms, it is possible to ensure that safety prompt audio is output first in complex broadcasting scenarios, thereby enhancing the safety of robot operation.
[0076] 7. Reduce the risk of false recognition and improve execution reliability: By introducing a multimodal fusion confirmation mechanism between speech recognition and decision execution, consistency verification is performed on speech content, visual targets, interaction orientation, permission conditions and scene status, which can effectively reduce false wake-up, false recognition and false control, and improve the reliability of robot interaction and execution safety. Attached Figure Description
[0077] Figure 1 This is a flowchart of a multimodal embodied interaction method for an outdoor cleaning robot according to an embodiment of the present invention.
[0078] Figure 2 This is a flowchart illustrating the intelligent path decision-making process for generating response content or robot actions in an embodiment of the present invention.
[0079] Figure 3 This is a schematic diagram illustrating the scheduling principle of the audio playback management module in this embodiment of the invention. Detailed Implementation
[0080] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0081] This embodiment provides a multimodal embodied interaction method for outdoor cleaning robots. By collaboratively integrating voice wake-up, visual perception, speech recognition, semantic matching, edge language generation, motion control, and multi-source audio scheduling, the robot can maintain stable and efficient operation in complex outdoor environments, and also engage in natural, proactive, and reliable on-site interaction with pedestrians and management personnel. Specifically, it includes the following steps:
[0082] Step 1: Activate the interaction mode through multimodal perception;
[0083] The robot operates in parallel with the voice wake-up channel and the visual perception channel to perceive voice signals, pedestrian targets and spatial relationships in the environment in real time, and trigger public interaction mode, management interaction mode or proactive safety prompt mode based on the perception results.
[0084] 1. Voice wake-up channel: The acquired raw audio stream is framed, windowed, and analyzed in the frequency domain. Based on the background noise estimation results, steady-state wind noise, mechanical noise, and broadband interference in the environment are suppressed to obtain an enhanced speech signal. Subsequently, speech activity detection and wake-up word detection are performed based on the enhanced speech signal.
[0085] (1) Speech activity detection adopts a method based on the joint decision of frame-level statistical features and temporal continuity to determine whether each audio frame contains human voice; the frame-level statistical features include one or more of short-time energy, spectral entropy, zero-crossing rate, and main frequency band energy distribution.
[0086] (2) The wake word detection includes at least a first wake word and a second wake word, wherein:
[0087] The first wake word is used to trigger the public interaction mode;
[0088] The second wake word is used to trigger the management interaction mode.
[0089] When the first wake word or its approximate expression is detected and reaches the preset wake-up threshold, the system enters the public interaction mode; when the second wake word or its approximate expression is detected and reaches the preset management threshold, the system enters the management interaction mode and outputs a management mode entry prompt.
[0090] 2. Visual perception channel: The system acquires environmental image sequences through vehicle-mounted cameras and uses a lightweight target detection and tracking model deployed on the edge computing platform to detect, associate, and continuously track pedestrian targets.
[0091] The system estimates the relative distance, relative orientation, and movement trend between the pedestrian and the robot based on the position changes of the target detection box in consecutive image frames, combined with camera calibration parameters and the robot's own pose information.
[0092] Let the position of the pedestrian target in the robot's coordinate system at a certain moment be... The relative distance between the pedestrian and the robot It can be represented as:
[0093]
[0094] Corresponding relative azimuth angle It can be represented as:
[0095]
[0096] Let the positions of the pedestrian target in the robot coordinate system in two adjacent frames be respectively and Then the pedestrian's relative motion vector for:
[0097]
[0098] The system combines pedestrian relative motion vectors Planning motion vectors for robots Calculate the risk score for proactive safety alerts. :
[0099]
[0100] in, This represents the angle between the pedestrian's direction of movement and the robot's planned direction of movement. and These are the weighting coefficients. To prevent tiny constants with a denominator of zero.
[0101]
[0102] When equation (5) is satisfied, the system triggers the active safety prompt mode, where, The threshold for triggering safety prompts;
[0103] When a pedestrian is detected to have entered the preset safety zone and meets the conditions of equation (5), the system will automatically trigger the active safety prompt mode without requiring the user to wake it up. Based on the pedestrian's relative position, direction of movement, the robot's current planned path and the preset safety distance threshold, the system will generate prompts that include detour direction, work intention or avoidance reminders.
[0104] Preferably, the safety zone can be set as a fan-shaped area, a circular area, or a forward risk area that is dynamically adjusted according to the current direction of movement, centered on the robot body.
[0105] In this embodiment of the invention, the robot continuously collects environmental audio and images in front of it while in standby cleaning mode.
[0106] On the voice side, the system performs frame segmentation and spectrum analysis on the acquired audio stream, suppresses wind noise and mechanical noise based on background noise estimation results, and uses voice activity detection to determine whether there are human voice segments in the input audio. When the first wake word is recognized, it enters the public interaction mode; when the second wake word is recognized, it enters the management interaction mode and broadcasts the administrator mode entry prompt.
[0107] On the visual side, the system uses a camera to detect and track pedestrians in real time, calculates the relative distance between the pedestrian and the robot according to equation (1), calculates the relative azimuth angle according to equation (2), and estimates the relative motion vector of the pedestrian according to equation (3). Furthermore, the system calculates the proactive safety warning risk score according to equation (4) and automatically triggers the proactive safety warning process when equation (5) is satisfied. The system generates a warning message based on the robot's current path planning results, such as indicating that the robot will detour around the pedestrian from the left or right.
[0108] Step 2: Identify speech content through a speech recognition link with speech activity detection and endpoint detection;
[0109] After the interactive mode is activated, the effective speech segments input after wake-up are enhanced, endpoints are located, acoustic features are extracted, and streaming recognition is performed to obtain the user's text input.
[0110] Specifically, after the public interaction mode or management interaction mode is activated, the system processes the voice stream collected after the wake word in real time.
[0111] First, the system divides the input audio signal into frames with a fixed frame length. Let the first frame be... The frame audio signal is ,in Then the short-time energy of that frame Represented as:
[0112]
[0113] To enhance the ability to distinguish between non-speech noise and silent segments, the system can also calculate the... Frame zero-crossing rate :
[0114]
[0115] in, It is a symbolic function.
[0116] Based on short-time energy and zero-crossing rate, the system constructs a speech activity determination function:
[0117]
[0118] in, Energy threshold The zero-crossing rate threshold, This indicates that the current frame is a speech frame. This indicates that the current frame is a non-speech frame.
[0119] During endpoint detection, when continuous Frame satisfies When, determine the starting point of the speech; when continuous Frame satisfies At that time, the termination point of the speech is determined, thereby extracting a complete and valid speech segment.
[0120] Subsequently, the system performs pre-emphasis, framing, and windowing processing on the effective speech segments, and extracts the filter bank energy features. Let the first... The frame signal is represented in the frequency domain as Then the first Energy of each filter bank channel for:
[0121]
[0122] in, Indicates the first Each filter at frequency index The response coefficient at that location, This represents the number of frequency points.
[0123] Taking the logarithm of the filter bank energy yields the logarithmic filter bank energy characteristics. :
[0124]
[0125] in, To prevent logarithmic operations from producing tiny constants with zero values.
[0126] After obtaining the acoustic feature sequence, the system uses a streaming end-to-end speech recognition network to incrementally decode the speech content, outputting partial recognition results before the speech has completely ended, and outputting the final recognized text after detecting the endpoint.
[0127] Preferably, the speech recognition link is deployed on the robot's local edge computing platform to reduce network dependence and improve recognition real-time performance and privacy security.
[0128] Preferably, when the recognition confidence level is lower than a preset threshold, the system outputs a confirmation prompt or requests the user to repeat the input, instead of directly proceeding to the subsequent response or control process.
[0129] In this embodiment of the invention, the system only processes the speech stream following the wake word.
[0130] The system calculates the short-time energy of each audio frame according to equation (6), calculates the zero-crossing rate according to equation (7), and determines whether the current frame is a speech frame according to equation (8). When several consecutive frames meet the speech entry condition, the speech start point is determined, and when several consecutive frames meet the silence exit condition, the speech end point is determined, and a valid speech segment is extracted.
[0131] Subsequently, the system extracts the energy features of the filter bank according to equation (9) and obtains the energy features of the logarithmic filter bank according to equation (10). Then, it inputs the data into the streaming speech recognition network for frame-by-frame decoding and outputs the text recognition results.
[0132] If the confidence level of the recognition result is lower than the set threshold, the robot will output a confirmation prompt such as "Please say it again" and will not proceed to the subsequent decision-making process.
[0133] Step 3: Verify the interaction intent and execution conditions through multimodal fusion confirmation;
[0134] The system combines the visual perception results, spatial relationship information, and target state information obtained in step one with the speech recognition text, speech recognition confidence, wake-up results, and interaction patterns obtained in step two to perform a fusion confirmation on whether the current input constitutes a valid interaction intent. When the fusion confirmation result indicates that the current input is inconsistent with the on-site visual target, the orientation of the interactive object, the distance relationship, the identity and permissions, or the robot's current working state, the system outputs a confirmation prompt, a clarification prompt, or a refusal to execute prompt, and pauses the subsequent response generation or action control process. When the fusion confirmation result meets the preset consistency conditions, the system enters the intelligent path decision-making step.
[0135] After obtaining the speech recognition text in step two, the system does not immediately enter the response generation or action execution process. Instead, it combines the speech recognition results, visual perception results, spatial positional relationships, target orientation information, wake-up mode information, permission information, and the robot's current working status to perform multimodal consistency confirmation of the input intent.
[0136] Multimodal conformance verification includes at least one or more of the following:
[0137] (1) Voice confidence confirmation
[0138] When the confidence level of speech recognition is lower than a preset threshold, the system determines that the current recognition result has a high degree of uncertainty and outputs a confirmation prompt instead of directly proceeding to the subsequent decision-making process.
[0139] (2) Visual presence confirmation
[0140] The system determines whether there are pedestrian targets in the preset interaction area in front of the robot based on the camera detection results. When no interactive object is detected, or the distance and orientation of the target do not meet the interaction conditions, the validity score of this voice input is reduced to avoid misjudging environmental noise, distant pedestrian voices, or voices not facing the robot as valid input.
[0141] (3) Confirmation of human-machine orientation
[0142] The system determines whether a pedestrian is facing the robot, approaching the robot, or has a clear intention to interact based on the direction the person is facing, the direction the head is facing, and the target's movement trend.
[0143] (4) Confirmation of wake-up mode consistency
[0144] When the voice wake-up result is in public interaction mode, the system verifies the recognized text according to public question-and-answer rules; when the voice wake-up result is in management interaction mode, the system further confirms whether the current input meets the management control conditions; if the mode and content are inconsistent, the system outputs a mode mismatch prompt or requests reconfirmation.
[0145] (5) Command-Scenario Consistency Confirmation
[0146] For management and control commands, the system will identify the intention of the action in the text and verify its consistency with the current robot operating status. When the identification result conflicts with the robot's current status, task stage, safety conditions, or execution conditions, the system will output a clarification prompt or a refusal to execute prompt to prevent miscontrol.
[0147] Furthermore, the system can construct a fusion confirmation score. This is used to evaluate whether the current input meets the execution conditions.
[0148]
[0149]
[0150] in, Indicates the visual presence consistency score. This indicates the reliability score of the speech recognition. Indicates the orientation and interaction intent score. This indicates a permission or pattern matching score. This indicates the consistency score between the command and the scenario state. to These are the corresponding weighting coefficients.
[0151]
[0152]
[0153] When equation (13) is satisfied, the system determines that the input has been confirmed through multimodal fusion and proceeds to the subsequent intelligent path decision-making process; when equation (14) is satisfied, the system outputs a confirmation prompt, a clarification prompt, or a refusal to execute prompt, and maintains the robot's current safety state unchanged. The threshold for fusion confirmation.
[0154] Preferably, when the identified text corresponds to a high-risk management action, motion control command, or task switching command, the system increases the fusion confirmation threshold to further reduce the risk of misidentification causing erroneous actions.
[0155] In this embodiment of the invention, after obtaining the text recognition result, the system further combines the pedestrian's position, orientation, relative distance, interaction area status, and current interaction mode detected by the current camera to perform multimodal fusion confirmation of the input.
[0156] Specifically, the system combines visual presence consistency score, speech recognition credibility score, orientation and interaction intent score, permission or pattern matching score, and command and scene state consistency score to calculate the fusion confirmation score according to equation (11). And constrain the weight of each scoring item according to formula (12).
[0157] If the fusion confirmation score satisfies equation (13), the current input is determined to be confirmed through multimodal fusion, and the system proceeds to the subsequent intelligent path decision-making step; if the fusion confirmation score satisfies equation (14), it indicates that the current input has a risk of misidentification, the interaction object is unclear, or the execution conditions are not met, and the system outputs a confirmation prompt, clarification prompt, or refusal to execute prompt, instead of proceeding to the subsequent response generation or action execution process.
[0158] For example, when the confidence level of speech recognition is insufficient, or the visual target does not exist, is not facing the robot, is beyond the interaction range, or the permission conditions in the management mode are not met, the system can determine that the input has not passed the fusion confirmation.
[0159] For high-risk management commands, such as returning to charging, switching tasks, pausing cleaning, or chassis motion control, the system can also increase the fusion confirmation threshold to further reduce the risk of erroneous actions caused by misidentification.
[0160] Step 4: Generate response content or robot action instructions through intelligent path decision-making;
[0161] If the current mode is public interaction mode, the system will select either the knowledge base response path or the edge language generation path to output a text response based on the matching result between the user's question and the local public knowledge base. If the current mode is management interaction mode, the system will generate the corresponding robot action sequence or output a clarification prompt based on the matching result between the management command and the local action library.
[0162] Intelligent decision-making paths include public interaction decision-making paths and management control decision-making paths.
[0163] 1. Public interaction decision-making path;
[0164] When the system is in public interaction mode, the user question text output in step three is input into the semantic decision module.
[0165] The semantic decision-making module first performs matching using a local public knowledge base. This local public knowledge base pre-stores multiple question-answer pairs, each pair including at least a pre-defined question text and a corresponding standard answer text.
[0166] To improve matching accuracy, the system adopts a dual-channel matching method that combines literal similarity and semantic similarity.
[0167] For user input text With the preset question text The system first calculates the edit distance. The literal similarity score was obtained by length normalization. :
[0168]
[0169] in, and Representing text respectively and text The length.
[0170] In some implementations, the system may further incorporate keyword overlap rates. This results in an enhanced literal matching score. :
[0171]
[0172] in, Let be the weighting coefficient, satisfying .
[0173] Meanwhile, the system uses a semantic coding model to generate user input text. With the preset question text semantic vector representation and Semantic similarity scores are calculated using cosine similarity. :
[0174]
[0175] Furthermore, the system weighted and fused the enhanced literal matching score with the semantic similarity score to obtain a comprehensive matching score. :
[0176]
[0177] in, To integrate the weighting coefficients, satisfy the following conditions: .
[0178]
[0179] When equation (19) is satisfied, the system determines that the user's question matches a candidate question in the local public knowledge base, where, Set the knowledge base hit threshold.
[0180] When the overall matching score is higher than the knowledge base hit threshold, the system retrieves the corresponding standard answer text from the local public knowledge base as the response content.
[0181] When the overall matching score falls below the knowledge base hit threshold, the system automatically switches to the client-side language generation path. This path inputs the user's question text and historical dialogue context into a lightweight language generation model, which then generates a natural language response that fits the current context.
[0182] Preferably, the lightweight language generation model employs one or more of the following: parameter compression, low-bit quantization, distillation transfer, efficient parameter adaptation, and computation graph optimization, in order to adapt to the limited computing resources on the robot's edge and reduce inference latency.
[0183] Preferably, the system introduces a content constraint mechanism during the language generation process to screen the generated content for security, friendliness, and task relevance, so as to avoid outputting inappropriate content.
[0184] Preferably, when the comprehensive matching score of the knowledge base is within a preset fuzzy range, the system prioritizes generating a clarifying response rather than directly providing a definitive answer.
[0185] 2. Management and control decision-making path;
[0186] When the system is in management interaction mode, the management command text output in step three is input into the management command parsing module.
[0187] The management instruction parsing module pre-stores a local preset action library, which includes multiple preset command texts and standard action sequences corresponding to each preset command text.
[0188] The system employs a dual-channel matching method that is the same as or equivalent to the public's interactive decision-making path to match management command text. With preset command text A comprehensive matching process is performed. The comprehensive matching score is expressed as follows:
[0189]
[0190] The determination formula is as follows:
[0191]
[0192] When equation (21) is satisfied, the system determines that the management command text matches a candidate command in the local action library, where, Identify thresholds for management commands.
[0193] When the overall matching score of the management command is higher than the management command threshold, the system outputs the corresponding standard action sequence as robot control instructions and sends it to the robot execution module for execution.
[0194] When the overall matching score of the management command is lower than the management command threshold, the system determines that the current input is not an executable control command, outputs a prompt message indicating that the command cannot be recognized, and keeps the robot's current running state unchanged to prevent miscontrol.
[0195] Preferably, the action sequence includes at least one or more of the following: start cleaning, pause task, return to charging, standby at a fixed point, status announcement, and task switching.
[0196] In this embodiment of the invention, if the current mode is public interaction mode, the system will identify the text and compare it one by one with the preset question texts in the local public knowledge base.
[0197] For each candidate question, the literal similarity is first calculated according to Equation (15), then the enhanced literal matching score is generated according to Equation (16), the semantic similarity is calculated according to Equation (17), and the comprehensive matching score is obtained according to Equation (18). When Equation (19) is satisfied, the system hits the local knowledge base and directly outputs the corresponding standard answer; otherwise, it will identify the text and historical context input side lightweight language generation model to generate natural language response content.
[0198] If the current system is in management interaction mode, it will perform a comprehensive match between the identified text and the preset command text in the local action library, and determine whether the management command is matched according to equations (20) and (21). If the match is successful, the corresponding action sequence will be output; otherwise, an unrecognizable command prompt will be output, and the robot's current working state will remain unchanged.
[0199] Step 5: Perform speech synthesis or robot actions;
[0200] The text response generated in step four can be converted into a playable voice response, or the action sequence generated in step four can be sent to the robot's underlying execution module for action control.
[0201] 1. Voice response output;
[0202] For the text response generated in step four, the system calls the on-device speech synthesis module to generate speech. The speech synthesis module converts the text response into an acoustic intermediate representation and further generates a playable speech waveform.
[0203] Preferably, the speech synthesis module adjusts the prosodic structure of the output speech by combining text pause positions, keyword stress, and speech rate control parameters during the generation process, so as to improve the naturalness and clarity of the broadcast.
[0204] To adapt to high-noise outdoor environments, the system can adjust according to the ambient noise intensity. Adaptive adjustment of broadcast volume and broadcasting speed For example, the broadcast volume can be represented as:
[0205]
[0206] in, Based on the basic volume, This is the noise compensation coefficient.
[0207] In some implementations, the broadcast speed can be further expressed as:
[0208]
[0209] in, Basic speaking speed, This is the speech rate adjustment coefficient.
[0210] When the length of the response text exceeds a preset threshold, the speech synthesis module supports a streaming output mode that generates and plays the response simultaneously to shorten the delay of the first packet broadcast.
[0211] 2. Execution of action sequence;
[0212] For the action sequence generated in step four, the system transmits the action sequence to the robot's underlying control module. The underlying control module parses the corresponding action according to the preset control protocol and drives the mobile chassis, cleaning mechanism, obstacle avoidance mechanism or status display module to perform the corresponding task.
[0213] Preferably, before the action is executed, the system also verifies the current safety status of the robot. If low battery, obstacle avoidance conflict, execution channel occupation, or permission mismatch is detected, the corresponding action is delayed or refused to be executed.
[0214] In this embodiment of the invention, for text response content, the system calls the local speech synthesis module to generate corresponding broadcast audio, and adaptively adjusts the broadcast volume and speech rate according to Equations (22) and (23) in combination with the current ambient noise level.
[0215] For the action sequence, the system sends it to the underlying control module, which then executes the corresponding chassis movement, job switching, or status feedback operation.
[0216] Step 6: Perform unified intelligent scheduling and playback of multi-source audio;
[0217] The system unifies the audio formats of interactive response audio, proactive safety prompt audio, and robot status broadcast audio, and performs priority arbitration, preemption control, and playback recovery to ensure that safety-related audio is broadcast first.
[0218] The robot has at least three types of audio broadcast sources during operation:
[0219] (1) Audio of public interactive responses;
[0220] (2) Active safety alert audio;
[0221] (3) Audio broadcast of robot status.
[0222] To avoid overlapping broadcasts, obscuring of key information, or confusion for users caused by concurrent audio from multiple sources, the system has a unified audio playback management module that centrally receives, standardizes, queues, and controls the output of broadcast requests initiated by various modules.
[0223] The unified audio playback management module first performs unified formatting processing on audio data from different sources. The formatting processing includes at least one or more of the following: unified sampling rate, quantization bit width conversion, channel mapping, buffer management, and target playback format encapsulation, in order to ensure the compatibility of multi-source audio in the same playback link.
[0224] Subsequently, the system categorizes and schedules broadcast requests according to preset priority rules, including:
[0225] (1) Active safety alert audio and emergency alarm audio are given high priority;
[0226] (2) Public interaction response audio is of medium priority;
[0227] (3) The robot status broadcast audio is of low priority.
[0228] For any audio request to be played The system calculates a scheduling priority score based on the event type, urgency, waiting time, and repetition rate. :
[0229]
[0230] in, Indicates the default type priority of audio requests. Indicates the urgency of the event. Indicates the waiting time. This indicates that the penalty item will be broadcast repeatedly. , , are the corresponding weighting coefficients.
[0231] The system prioritizes playing the audio request with the highest priority score. If the current playback channel is idle, newly arrived audio requests are played immediately; if the current playback channel is busy, newly arrived audio requests are placed in the playback queue and played according to their priority score. Determine the queuing order.
[0232]
[0233] When a new request arrives When equation (25) is satisfied, preemptive broadcasting is performed, where, Assign a priority rating to the currently playing audio. To seize the threshold.
[0234] After the high-priority audio finishes playing, the system resumes playback based on the remaining duration and original priority of the interrupted audio.
[0235] Preferably, the system also includes a repetitive broadcast suppression mechanism, which performs merging and deduplication, delayed broadcasting, or suppressed broadcasting processing on prompt audios that are repeatedly triggered by the same event source within a preset time window, in order to reduce the interference of invalid repetitive reminders to users.
[0236] Preferably, the system also includes an audio segment integrity protection mechanism, which supports resuming playback from the interrupted position or replaying from the current semantic integrity boundary for low-priority audio that has been interrupted due to preemption.
[0237] In this embodiment of the invention, the system sends public interaction response audio, proactive safety prompt audio, and robot status broadcast audio to the audio playback management module in a unified manner.
[0238] For any broadcast request, the system calculates its scheduling priority score according to equation (24) and sorts it into the waiting queue according to priority. When a new audio request arrives and satisfies equation (25), the system performs preemptive broadcast.
[0239] For example, if a new safety alert audio arrives while the robot is broadcasting status information and the preemption condition is met, the system will immediately interrupt the current status broadcast and prioritize the output of the safety alert; after the safety alert ends, the previously interrupted status broadcast content will be resumed or the broadcast will be re-broadcast from the semantically complete position.
[0240] like Figure 2 As shown, in another embodiment, based on step 3 of the above embodiments, the intelligent path decision-making process is as follows:
[0241] The system first determines whether to enter the public interaction decision path or the management control decision path based on the current interaction status.
[0242] In the public interaction decision-making process, the system uses a dual-channel matching method to match user questions with the local public knowledge base:
[0243] The first channel is the approximate text matching channel, which is used to measure the similarity of character-level and word-level expressions according to Equations (15) and (16);
[0244] The second channel is the semantic vector matching channel, which is used to measure the consistency of semantic expression according to equation (17).
[0245] The scores from the two channels are fused according to Equation (18). If the fused score satisfies Equation (19), the answer in the knowledge base is directly matched; otherwise, the end-side language generation module is called to generate the response.
[0246] In the management control decision path, the system compares the similarity between the management command text and the command template in the action library. If the comprehensive matching score satisfies equation (21), a standard action sequence is generated; if not, a clarification prompt is output, and the robot is prevented from performing ambiguous actions.
[0247] like Figure 3 As shown, in another embodiment, based on step 5 of the above embodiment, the multi-source audio scheduling process is as follows:
[0248] The audio playback management module includes at least the following components: an audio request receiving unit, an audio format unification unit, a priority determination unit, a playback queue management unit, and a preemption and recovery control unit.
[0249] When any module initiates a broadcast request, the audio request receiving unit receives the corresponding audio data or broadcast instruction.
[0250] The audio format unification unit converts audio from different sources into a unified playback format;
[0251] The priority determination unit calculates the current request priority score according to formula (24);
[0252] The playback queue management unit maintains the waiting queue based on priority scoring and arrival time.
[0253] The preemption recovery control unit interrupts low-priority broadcasting when it detects a new arrival request satisfying equation (25), and resumes the interrupted broadcasting after the high-priority broadcasting is completed.
[0254] In addition, for the same security event that is repeatedly triggered within a short time window, the system adopts a repetition suppression strategy to avoid information redundancy caused by continuous repeated broadcasts.
[0255] The method of this invention can effectively reduce the risks of false wake-up, voice misrecognition and miscontrol, and realize natural robot interaction, proactive safety prompts, on-site voice operation and maintenance control and orderly broadcasting of multi-source audio in complex outdoor environments. It has the advantages of high real-time performance, high reliability, low cloud dependence and good human-machine collaboration effect.
[0256] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multimodal embodied interaction method for outdoor cleaning robots, applied to outdoor cleaning robots, characterized in that: The interaction mode is activated through multimodal perception; voice detection is used to judge human voices, and public interaction mode and management interaction mode are triggered according to different wake words; pedestrians are perceived by the collected environmental image sequence, and active safety prompt mode is triggered according to the spatial relationship between pedestrians and robots. By using a speech recognition link with speech activity detection and speech endpoint detection, speech content is recognized after the wake word to obtain text recognition results; The interaction intent and execution conditions are verified through multimodal fusion confirmation; a fusion confirmation score is performed based on the speech confidence of the text recognition result, the spatial relationship, and the consistency of the text recognition result with the interaction mode and the robot's operating state, respectively. Voice confidence scores are used to identify highly confident speech as reliable speech and generate a speech recognition confidence score. The spatial relationships are used to confirm the visual presence of pedestrians within the interaction area, so as to generate a visual presence consistency score. Pedestrian interaction intent is determined by identifying pedestrian orientation information and movement trends from the environmental images, thereby generating an orientation and interaction intent score. When the public interaction mode is triggered, the text recognition result is verified according to the public question and answer rules. When the management interaction mode is triggered, it is further confirmed whether the current voice input meets the management control conditions, and an permission or pattern matching score is generated. For the management interaction mode, the recognized management control commands will be verified for consistency between the action intent involved in the text recognition results and the current robot operating state, and a consistency score between the command and the scene state will be generated. The fusion confirmation score is a weighted sum of the speech recognition credibility score, visual presence consistency score, orientation and interaction intent score, permission or pattern matching score, and command and scene state consistency score. When the fusion confirmation score meets the requirements, intelligent path decision-making generates response content or robot action instructions; if it is the public interaction mode, the robot responds according to the question; if it is the management interaction mode, the robot generates robot actions according to the management command.
2. The multimodal embodied interaction method for an outdoor cleaning robot according to claim 1, characterized in that: The robot uses a voice wake-up channel to perform frame segmentation, windowing, and frequency domain analysis on the acquired raw audio stream. Based on the background noise estimation results, it suppresses steady-state wind noise, mechanical noise, and broadband interference in the environment to obtain an enhanced voice signal. Subsequently, speech activity detection and wake word detection are performed based on the enhanced speech signal.
3. The multimodal embodied interaction method for an outdoor cleaning robot according to claim 1, characterized in that: The speech activity detection method uses a combination of frame-level statistical features and temporal continuity to determine whether each audio frame contains human voice. The frame-level statistical features include one or more of the following: short-time energy, spectral entropy, zero-crossing rate, and dominant frequency band energy distribution.
4. The multimodal embodied interaction method for an outdoor cleaning robot according to claim 1, characterized in that: The robot detects, associates, and continuously tracks pedestrian targets. Based on the positional changes of the target detection box in consecutive image frames, combined with camera calibration parameters and the robot's own pose information, it estimates the relative distance, relative orientation, and movement trend between the pedestrian and the robot. The motion trend is obtained by acquiring the position of the pedestrian target in the robot coordinate system in two adjacent frames to obtain the relative motion vector of the pedestrian. The robot combines the relative motion vector of the pedestrian with the robot's planned motion vector to obtain the angle between the pedestrian's motion direction and the robot's planned motion direction, so as to generate an active safety warning risk score. When the score is greater than the threshold, the active safety warning mode is triggered. When a pedestrian is detected entering a preset safety area and the proactive safety warning risk score is greater than a threshold, the robot triggers the proactive safety warning mode. Based on the pedestrian's relative position, direction of movement, the robot's current planned path, and preset safe distance threshold, it generates prompts that include detour directions, operational intentions, or avoidance reminders.
5. The multimodal embodied interaction method for an outdoor cleaning robot according to claim 1, characterized in that: The voice activity detection process divides the audio signal collected after the wake word into frames with a fixed frame length, and obtains the short-time energy and zero-crossing rate of each frame. A voice activity determination function is constructed based on the short-time energy and zero-crossing rate. When the short-time energy is greater than the energy threshold and the zero-crossing rate is less than the zero-rate threshold, the current frame is determined to be a voice frame; otherwise, it is a non-voice frame.
6. The multimodal embodied interaction method for an outdoor cleaning robot according to claim 5, characterized in that: The speech endpoint detection process determines the speech start point when multiple consecutive frames are identified as speech frames, and determines the speech end point when multiple consecutive frames are not speech frames, in order to extract complete and valid speech segments. The speech segments are then processed by pre-emphasis, framing, and windowing, and the energy features of the filter bank are extracted. Based on the frequency domain of a single frame audio signal and the response coefficient of a single filter at the frequency index, the filter bank energy is obtained by accumulating the frequency points and taking the logarithm. The resulting logarithmic filter bank energy is input into a streaming speech recognition network for frame-by-frame decoding, and the text recognition result is output.
7. The multimodal embodied interaction method for an outdoor cleaning robot according to claim 1, characterized in that: In the public interaction mode, the user question is obtained from the text recognition result and input into the semantic decision module. Local public knowledge base matching is performed, and the encoding distance between the user question text and the preset question text is calculated to obtain a literal similarity score. Combined with the keyword overlap rate, an enhanced literal matching score is obtained. Semantic vectors of the user question text and the preset question text are generated using a semantic encoding model, and a semantic similarity score is obtained through similarity. The enhanced literal matching score and the semantic similarity score are weighted and fused to obtain a comprehensive matching score. If the score reaches a threshold, it is determined that the user question matches a candidate question in the local public knowledge base, and a standard answer is retrieved based on the candidate question for response. Otherwise, the process switches to the terminal-side language generation path, and the user question text and historical dialogue context are input into the language generation model to generate natural language that conforms to the current context for response. The response text is generated by the edge speech synthesis module. The prosodic structure of the output speech is adjusted by combining the pause position of the response text, keyword stress and speech rate control parameters. The broadcast volume and broadcast speed are adaptively adjusted according to the ambient noise intensity. When the length of the response text exceeds the preset threshold, the speech synthesis module adopts a streaming output mode of generating and playing at the same time.
8. The multimodal embodied interaction method for an outdoor cleaning robot according to claim 1, characterized in that: In the management interaction mode, the management command text is obtained from the text recognition result and input into the management command parsing module. Local preset action library matching is performed, and the encoding distance between the management command text and the preset command text is calculated to obtain a literal similarity score. Combined with keyword overlap rate, an enhanced literal matching score is obtained. A semantic encoding model is used to generate semantic vectors for the management command text and the preset command text respectively, and a semantic similarity score is obtained through similarity. The enhanced literal matching score and the semantic similarity score are weighted and fused to obtain a comprehensive matching score. If the score reaches a threshold, it is determined that the management command matches a candidate command in the local preset action library, and a standard action sequence is retrieved and executed according to the candidate command. Otherwise, a prompt message indicating that the command cannot be recognized is output. The action sequence is parsed by the robot's underlying control module according to the preset control protocol, and drives the mobile chassis, cleaning mechanism, obstacle avoidance mechanism or status display module to perform the corresponding tasks. Before the action is executed, the current safety status of the robot is checked. If low battery, obstacle avoidance conflict, execution channel occupation or permission mismatch is detected, the corresponding action is delayed or refused to be executed. For voice prompts, the local speech synthesis module is invoked to generate corresponding audio, and the volume and speed of the audio are adaptively adjusted based on the current ambient noise level.
9. A multimodal embodied interaction method for an outdoor cleaning robot according to claim 1, characterized in that: Unified intelligent scheduling and broadcasting of multi-source audio; A unified audio playback management module is adopted to unify the multi-source audio formats of interactive response audio, proactive safety prompt audio, and robot status broadcast audio. The broadcast requests are classified and scheduled according to the preset priority rules, with the priority order being: proactive safety prompt audio and emergency alarm audio, public interaction response audio, and robot status broadcast audio. The scheduling priority score is calculated based on the audio event type, urgency, waiting time, and repetition rate, and the audio request with the highest score is played first. If the current playback channel is idle, the newly arrived audio request is played immediately. If the current playback channel is busy, the newly arrived audio request enters the waiting queue, and the queuing order is determined according to the priority score. A preemption threshold is set to adjust the sorting. After the high-priority audio finishes playing, playback is resumed based on the remaining duration and original priority of the interrupted audio.
Citation Information
Patent Citations
Multi-modal information fusion body-equipped intelligent robot control method
CN121374589A
Working system, method and equipment of intelligent distribution network hot-line work robot with body and storage medium
CN121870733A