A method and system for robot audio-visual fusion question and answer interaction

By linking streaming speech recognition with intent inference and visual perception for cross-modal verification and proactive clarification, the perception accuracy and task execution accuracy of the robot system are enhanced, solving the problems of instruction ambiguity and insufficient perception in robot systems under complex environments in existing technologies.

CN121315979BActive Publication Date: 2026-08-04ANHUI GUANGDING INTELLIGENT TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ANHUI GUANGDING INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2025-11-28
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing intelligent robot systems lack proactive clarification mechanisms when faced with complex, ambiguous, or environmentally sensitive instructions, resulting in insufficient precision in intent understanding and perception, making it difficult to accurately execute tasks.

Method used

It employs streaming speech recognition and intent inference in conjunction with visual perception, eliminates ambiguity through cross-modal verification, and enhances perception and spatial anchoring in the task to ensure execution accuracy.

Benefits of technology

It improves the fluency, robustness, and success rate of robot interaction and task execution, and ensures the accuracy of task planning and precise positioning of target objects through proactive clarification and multi-perspective information acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121315979B_ABST
    Figure CN121315979B_ABST
Patent Text Reader

Abstract

The application discloses a kind of robot audio-visual fusion's question and answer type interaction method and system, belong to artificial intelligence and man-machine interaction technical field.Its method includes: obtaining voice instruction and generating task hypothesis;Dynamic adjustment AI visual perception resource;Combining speech and visual information carries out cross-modal verification;Judge ambiguity or poor view, and initiate active clarification question and answer or instruction movement view transformation;Execute task and obtain multi-view image by enhanced recognition movement mode to carry out three-dimensional fusion;Combining robot pose carries out spatial anchoring processing;Use stable coordinate information to generate contextual feedback information and output.The application adopts the combination strategy of active clarification, motion enhanced perception and high-precision spatial anchoring, can accurately solve the problem of instruction ambiguity and insufficient perception in high dynamic complex environment, improve the accuracy of robot task execution, the robustness of spatial positioning and the efficiency and naturalness of human-computer interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and human-computer interaction technology, and in particular to a question-and-answer interaction method and system for robot audiovisual fusion. Background Technology

[0002] Intelligent robot technology is gradually being integrated into various fields such as industry, services, and daily life. In terms of human-computer interaction, the robot's ability to perform complex tasks is a key indicator for measuring its level of intelligence. Existing intelligent robot systems often rely on preset instructions or single-modal input, and their ability to understand intent, accurately perceive, and execute actions is significantly limited when faced with complex, ambiguous, or environmentally-related instructions from users.

[0003] In related technologies, Chinese Patent No. CN119927906A discloses an interaction method for collaborative control of the upper limbs of a humanoid robot. This method includes: performing task analysis on a received voice request; if there is a target object to be executed, then performing target object recognition processing based on the task analysis results; based on the target recognition information and target pose information output by the recognition processing, and the task analysis results, using a large language model to perform task planning processing according to the principle of dual-arm division of labor and cooperation, generating a task sequence and text instructions corresponding to each task; and sequentially performing the following operations on each target task in the task sequence: based on the target pose information and target recognition information, controlling the upper limbs of the humanoid robot to perform the operation action corresponding to the target task on the target object.

[0004] Regarding the aforementioned technologies, the inventors believe that: Existing technologies primarily focus on the collaborative control of the upper limbs of humanoid robots and task sequence generation based on large language models, but they do not comprehensively address the issues of instruction ambiguity, perception enhancement under limited perspective, and high-precision 3D spatial anchoring in multimodal question answering. In complex real-world environments, when voice commands and visual perception information do not perfectly match, this technology lacks an active clarification mechanism, easily leading to robot task planning errors. Simultaneously, when the robot's perspective is poor, it lacks the ability to execute enhanced recognition motion patterns to actively acquire high-precision perception data, limiting the accuracy of task execution. Therefore, there is an urgent need for a robot audiovisual fusion question-answering interaction method that can actively resolve instruction ambiguity and enhance environmental perception accuracy. Summary of the Invention

[0005] To address the aforementioned issues, this invention provides a question-and-answer interaction method and system for robot audiovisual fusion. It employs a technical solution that combines streaming speech recognition and intent inference with visual perception, eliminates ambiguity through cross-modal verification and proactive clarification, and ensures execution accuracy during tasks through enhanced perception and spatial anchoring. This approach can improve the fluency, robustness, and success rate of robot interaction and task execution.

[0006] To achieve the above objectives, this application adopts the following technical solution:

[0007] Firstly, a question-and-answer interaction method for robot audiovisual fusion is provided, including:

[0008] The system acquires user voice commands and performs real-time streaming recognition and conversion to generate preliminary recognized text.

[0009] Based on the initially identified text, it is processed by a preset model for inferring user intent to generate a preliminary task hypothesis model;

[0010] Receive the preliminary task hypothesis model and dynamically adjust the perceptual resources of AI visual recognition to generate visual candidate information;

[0011] By combining the fully recognized voice commands with the visual candidate information, cross-modal verification is performed to generate verification results;

[0012] If the verification result indicates ambiguity, an active clarification question-and-answer process is initiated, and a refined task plan is generated based on user feedback. If not, but the indicated perspective is poor, the robot is instructed to move to change the perspective, and a refined task plan is generated based on the changed perspective information.

[0013] The robot receives the refined task plan and controls the robot's movement to execute the task. At the same time, it acquires multi-view images by executing motion patterns for enhanced recognition and fuses the multi-view images to generate enhanced perception data.

[0014] Based on the enhanced perception data and combined with the robot's own pose change data, spatial anchoring processing is performed to generate stable coordinate information;

[0015] Using the stable coordinate information and task execution status, contextualized feedback information is generated and output to the user.

[0016] Based on the above technical solutions, in the question-and-answer interaction method of robot audiovisual fusion provided in this application, the technical solutions of using streaming speech recognition and intent inference linked with visual perception, eliminating ambiguity through cross-modal verification and active clarification, and ensuring execution accuracy through enhanced perception and spatial anchoring in the task can improve the fluency, robustness and success rate of robot interaction and task execution.

[0017] In conjunction with the first aspect above, in one possible implementation, the dynamic adjustment of the perceptual resources for AI visual recognition includes:

[0018] Based on the preliminary task hypothesis model, the hypothetical target area is determined, and the robot gimbal or body is controlled to face the hypothetical target area to achieve focused view.

[0019] The image recognition algorithm associated with the preliminary task hypothesis model is enabled to achieve algorithm focus;

[0020] Based on the results of the perspective focusing and algorithm focusing, visual candidate information is generated.

[0021] In conjunction with the first aspect above, in one possible implementation, the cross-modal verification by combining the fully recognized voice command with the visual candidate information includes:

[0022] Key entity information is extracted from the fully recognized voice command, and the key entity information is compared with the target attributes in the visual candidate information to generate a matching score.

[0023] When the matching score is lower than the threshold used to determine ambiguity, an active clarification question-and-answer process is triggered.

[0024] The visual candidate information is updated based on user feedback obtained from the proactive clarification question-and-answer process, and an updated verification result is generated.

[0025] In conjunction with the first aspect above, in one possible implementation, the process of initiating a proactive clarification-based question-and-answer process and generating a refined task plan based on user feedback includes:

[0026] Based on the visual candidate information, several candidate target options are generated;

[0027] The system uses speech synthesis to ask the user a clarifying question containing the candidate options, receives and parses the user's voice feedback, and obtains the user's selection intent.

[0028] Based on the user's selected intent, the initial task hypothesis model is optimized to generate a refined task plan.

[0029] In conjunction with the first aspect above, in one possible implementation, receiving the refined task plan and controlling the robot's movement to perform the task, while simultaneously acquiring multi-view images by executing motion patterns for enhanced recognition, and fusing the multi-view images includes:

[0030] Monitor the confidence parameter of AI visual recognition on the target object, and activate the motion mode for enhanced recognition when the confidence parameter is lower than the threshold used to trigger enhanced perception.

[0031] Perform arc motion around the target or vertical motion The system elevates and lowers the characters, acquires multi-view images, and performs 3D point cloud fusion processing on the multi-view images to generate enhanced perception data.

[0032] In conjunction with the first aspect above, in one possible implementation, the 3D point cloud fusion processing of multi-view images includes:

[0033] Feature points are extracted and matched from the multi-view images to generate a feature point set;

[0034] The feature point set is processed using the structure-reconstruction-motion algorithm to generate a sparse point cloud;

[0035] A multi-view stereo algorithm is applied to densify the sparse point cloud to generate a dense 3D model.

[0036] Based on the dense 3D model, the position and orientation information of the target object are updated to generate updated enhanced perception data.

[0037] In conjunction with the first aspect above, in one possible implementation, spatial anchoring processing based on the enhanced perception data and combined with the robot's own pose change data includes:

[0038] The robot acquires real-time pose change data from inertial measurement output, and calculates its own pose transformation matrix by combining it with a kinematic model describing the relationship between the robot's joints and end effectors.

[0039] The pose transformation matrix is ​​synchronized to the perception of AI visual recognition to correct target coordinate drift and output stable coordinate information.

[0040] In conjunction with the first aspect above, in one possible implementation, acquiring the real-time pose change data output by inertial measurement, combined with a kinematic model describing the relationship between the robot joints and the end effector, includes:

[0041] By performing the hand-eye calibration process, multiple sets of camera pose transformation data and robot end-effector pose transformation data are collected.

[0042] Construct hand-eye calibration equations, where the camera pose transformation is a matrix and the robot end effector pose transformation is a matrix.

[0043] The fixed transformation matrix is ​​solved using an algorithm for solving overdetermined systems of equations, and then the fixed transformation matrix is ​​used. Establish a precise spatial mapping relationship to achieve coordinate transformation from the camera coordinate system to the robot coordinate system.

[0044] In conjunction with the first aspect above, in one possible implementation, generating contextualized feedback information using the stable coordinate information and task execution status includes:

[0045] Visual context information is integrated from the enhanced perception data, and motion history is fused from the task execution state;

[0046] Using a large language model for semantic generation, the visual context information and the motion history are processed into a natural language description;

[0047] The natural language description and generative image content are combined to form contextualized feedback information.

[0048] Secondly, a question-and-answer interactive system for robot audiovisual fusion is provided, comprising: a voice command processing module, an intent inference module, an image perception and adjustment module, a cross-modal verification module, a task planning and clarification module, a motion control and enhanced perception module, a spatial anchoring module, and a feedback generation and information output module; wherein:

[0049] The voice command processing module is used to acquire user voice commands, perform real-time streaming recognition and conversion, and generate preliminary recognized text.

[0050] The intent inference module is used to process the pre-defined user intent model based on the pre-identified text to generate a preliminary task hypothesis model.

[0051] The image perception and adjustment module is used to receive the preliminary task hypothesis model and dynamically adjust the perception resources of AI visual recognition to generate visual candidate information.

[0052] The cross-modal verification module is used to combine the fully recognized voice command with the visual candidate information to perform cross-modal verification and generate verification results;

[0053] The task planning and clarification module is used to determine whether there is any ambiguity in the verification result indication. If yes, it initiates an active clarification question-and-answer process and generates a refined task plan based on user feedback. If no, but the display view is not good, it instructs the robot to move to change the view and generates a refined task plan based on the changed view information.

[0054] The motion control and enhanced perception module is used to receive the refined task plan and control the robot to perform the task. At the same time, it acquires multi-view images by executing motion patterns for enhanced recognition, and fuses the multi-view images to generate enhanced perception data.

[0055] The spatial anchoring module is used to perform spatial anchoring processing based on the enhanced perception data and combined with the robot's own pose change data to generate stable coordinate information;

[0056] The feedback generation and information output module is used to generate contextualized feedback information using the stable coordinate information and task execution status, and output the contextualized feedback information to the user.

[0057] Compared with the prior art, the present invention has the following advantages:

[0058] This invention achieves parallel processing of perception and understanding by performing real-time streaming recognition and intent inference during the user's voice command input process, and by adjusting the perception resources of AI visual recognition in conjunction with the input. This effectively shortens the delay from command input to robot response and improves the fluency and immediacy of human-computer interaction.

[0059] This invention introduces a cross-modal verification and proactive clarification question-and-answer process, enabling the robot to proactively initiate communication to eliminate uncertainty when faced with ambiguous instructions or uncertain visual information, or to obtain better observation conditions through autonomous perspective changes. This enhances the system's robustness to complex environments and ambiguous instructions, and ensures the accuracy of task planning.

[0060] This invention combines motion patterns and spatial anchoring processing for enhanced recognition during task execution, enabling the robot to proactively acquire multi-view information to construct a more accurate 3D model when perception confidence is insufficient. At the same time, it compensates for coordinate drift caused by its own motion, ensuring continuous and accurate positioning of the target object and improving the success rate of physical interaction tasks.

[0061] This invention generates contextualized feedback information, presenting the robot's task execution process, including its perception, decision-making, and motion history, to the user in easily understandable natural language and visual content. This constructs a complete interactive loop, enhances the transparency of the interaction and the user's sense of trust, and optimizes the overall human-machine collaboration experience.

[0062] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description

[0063] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0064] Figure 1 A structural architecture diagram of a robot audiovisual fusion question-and-answer interactive system provided in this application embodiment;

[0065] Figure 2 A flowchart illustrating a question-and-answer interaction method for robot audiovisual fusion provided in an embodiment of this application;

[0066] Figure 3 This is a comparison chart of AI visual confidence parameters and enhanced perception thresholds provided in the embodiments of this application.

[0067] Figure 4 This is an image showing the optimization effect of hand-eye calibration on spatial anchoring residuals provided in the embodiments of this application. Detailed Implementation

[0068] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.

[0069] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0070] The question-and-answer interaction method for robot audiovisual fusion provided in this application embodiment can be applied to, for example... Figure 1 In the robot audiovisual fusion question-and-answer interactive system 100 shown, such as Figure 1 As shown, the system includes: a voice command processing module, an intent inference module, an image perception and adjustment module, a cross-modal verification module, a task planning and clarification module, a motion control and enhanced perception module, a spatial anchoring module, and a feedback generation and information output module; wherein:

[0071] The voice command processing module is used to acquire user voice commands, perform real-time streaming recognition and conversion, and generate preliminary recognized text.

[0072] The intent inference module is used to process the pre-defined user intent model based on the pre-identified text to generate a preliminary task hypothesis model.

[0073] The image perception and adjustment module is used to receive the preliminary task hypothesis model and dynamically adjust the perception resources of AI visual recognition to generate visual candidate information.

[0074] The cross-modal verification module is used to combine the fully recognized voice command with the visual candidate information to perform cross-modal verification and generate verification results;

[0075] The task planning and clarification module is used to determine whether there is any ambiguity in the verification result indication. If yes, it initiates an active clarification question-and-answer process and generates a refined task plan based on user feedback. If no, but the display view is not good, it instructs the robot to move to change the view and generates a refined task plan based on the changed view information.

[0076] The motion control and enhanced perception module is used to receive the refined task plan and control the robot to perform the task. At the same time, it acquires multi-view images by executing motion patterns for enhanced recognition, and fuses the multi-view images to generate enhanced perception data.

[0077] The spatial anchoring module is used to perform spatial anchoring processing based on the enhanced perception data and combined with the robot's own pose change data to generate stable coordinate information;

[0078] The feedback generation and information output module is used to generate contextualized feedback information using the stable coordinate information and task execution status, and output the contextualized feedback information to the user.

[0079] like Figure 2 As shown in the figure, this application provides a question-and-answer interaction method for robot audiovisual fusion, including:

[0080] The system acquires user voice commands and performs real-time streaming recognition and conversion to generate preliminary recognized text.

[0081] Based on the initially identified text, it is processed by a preset model for inferring user intent to generate a preliminary task hypothesis model;

[0082] Receive the preliminary task hypothesis model and dynamically adjust the perceptual resources of AI visual recognition to generate visual candidate information;

[0083] By combining the fully recognized voice commands with the visual candidate information, cross-modal verification is performed to generate verification results;

[0084] If the verification result indicates ambiguity, an active clarification question-and-answer process is initiated, and a refined task plan is generated based on user feedback. If not, but the indicated perspective is poor, the robot is instructed to move to change the perspective, and a refined task plan is generated based on the changed perspective information.

[0085] The robot receives the refined task plan and controls the robot's movement to execute the task. At the same time, it acquires multi-view images by executing motion patterns for enhanced recognition and fuses the multi-view images to generate enhanced perception data.

[0086] Based on the enhanced perception data and combined with the robot's own pose change data, spatial anchoring processing is performed to generate stable coordinate information;

[0087] Using the stable coordinate information and task execution status, contextualized feedback information is generated and output to the user.

[0088] It should be noted that by streaming and processing user voice commands in real time, preliminary task hypotheses are generated in advance. Based on these hypotheses, the robot's visual perception resources are proactively and dynamically allocated to achieve pre-focusing on potential target areas. The complete auditory commands are then cross-modal validated with the preliminary visual perception results to identify potential ambiguities or adverse observation conditions. Faced with uncertainty, instead of passively waiting or failing, the robot actively initiates clarifying dialogues or autonomously changes perspectives to obtain the key information needed for decision-making, thereby generating a refined and executable task plan. During the task execution phase, this method closely integrates physical motion with enhanced perception. The robot collects multi-view information through specific motion patterns and fuses it into more accurate enhanced perception data. Simultaneously, it continuously uses its own pose change data to spatially anchor the target, eliminating coordinate drift caused by self-motion and ensuring continuous and stable positioning. By integrating stable coordinates, execution status, and visual context throughout the entire task process, contextualized multimodal feedback information is generated and output to the user, completing the entire interaction loop.

[0089] In one possible implementation of the embodiments of this application, combined with Figure 2 The above-mentioned dynamic adjustment of AI visual recognition perception resources can be explained in detail as follows:

[0090] Based on the preliminary task hypothesis model, the hypothetical target area is determined, and the robot gimbal or body is controlled to face the hypothetical target area to achieve focused view.

[0091] In some implementations, a preliminary task hypothesis model is received. This model is structured data obtained after parsing the initial text from streaming speech recognition, containing inferences about the user's possible intentions, such as "take," "point," and attribute descriptions including "red" and "on the table." Upon receiving the preliminary task hypothesis model, the perspective focusing process is immediately initiated. Descriptions related to spatial location, such as "on the table" and "to your left," are extracted from the preliminary task hypothesis model, and combined with the robot's prior knowledge map of the environment, a hypothetical target area in three-dimensional space is parsed. Subsequently, the robot's motion control system calculates the angle and direction that the robot's gimbal or body needs to rotate based on the coordinates of the hypothetical target area. Control commands are sent to the corresponding drive unit, driving the camera to physically face the area. This physical action allows the robot's vision sensors to prioritize acquiring scene images most likely containing the target object, greatly narrowing the scope of subsequent visual processing and avoiding ineffective scanning of the entire environment.

[0092] For example, when the robot receives the initial recognized text of the user's voice command as "take a red cup", the inference model generates an initial task hypothesis model {intent: take, attribute: color is red, area: desktop}. Based on the description of the "desktop" area, the robot calculates the instruction that its gimbal needs to tilt down 15 degrees and rotate to the left 30 degrees, driving the camera's field of view to accurately focus on the desktop area.

[0093] The image recognition algorithm associated with the preliminary task hypothesis model is enabled to achieve algorithm focus;

[0094] In some implementations, the algorithm focusing process is initiated in parallel while performing perspective focusing. The initial task hypothesis model is analyzed in depth, focusing on keywords related to object attributes, such as color, shape, and category. Based on these keywords, the algorithm most relevant to these attribute detections is dynamically selected and activated from the built-in AI vision algorithm library. For example, when the model contains the hypothesis of a "red cup," the object detection model and the color segmentation or classification model are loaded and run first. Meanwhile, algorithms unrelated to the current task, such as face recognition and text recognition, are placed with low priority or suspended. This approach concentrates limited computing resources on solving the most critical perception problem, achieving efficient utilization of computing resources.

[0095] For example, since the initial task assumption model contains the intentions of "color-red" and "take", the color recognition algorithm based on HSV color space conversion and the 3D object detection and recognition algorithm based on Faster R-CNN are immediately enabled, while the MPII human pose estimation algorithm and license plate recognition algorithm are paused.

[0096] Based on the results of the perspective focusing and algorithm focusing, visual candidate information is generated.

[0097] In some implementations, image data captured by a camera with focused field of view is fed into a specific image recognition algorithm activated after being focused by an algorithm. The algorithm's processing results, such as information on all red objects identified within a hypothetical target area, their locations, and confidence levels, are integrated and structured to form a list of visual candidate information. This visual candidate information, serving as a preliminary visual response to the user's ambiguous commands, is passed to subsequent processes for comparison with complete voice commands, providing crucial visual evidence for the final task decision.

[0098] For example, after focusing and algorithmic processing, three objects are identified on the table. The visual candidate information is structured as: "{{Target 1: Category: Cup, Color: Red, 3D Coordinates:}} Confidence level: 0.95}; {Target 2: Category: Box, Color: Red, 3D: Confidence level: 0.88}; {Target 3: Category: Pen holder, Color: Blue, 3D coordinates:} , confidence level: 0.92:}}” and pass this list to the subsequent cross-modal verification step.

[0099] In one possible implementation, combining Figure 2 The cross-modal verification, which combines the fully recognized voice command with the visual candidate information, can be explained as follows:

[0100] Key entity information is extracted from the fully recognized voice command, and the key entity information is compared with the target attributes in the visual candidate information to generate a matching score.

[0101] In some implementations, deep natural language understanding processing is applied to the fully recognized speech commands to extract key entity information. This information consists of words describing the specific attributes and relationships of the target object, typically including object category, object attributes, and spatial relationships. The extracted key entity information is structured to form a target description vector in the auditory modality. This auditory target description vector is then compared one by one with visual candidate information, generating a quantified matching score. The visual candidate information itself is a list, where each candidate target is accompanied by its visual attributes, such as category labels given by image recognition algorithms, dominant hues analyzed from color histograms, and spatial coordinates determined by 3D localization. The comparison process calculates a matching score for each visual candidate target. This score can be calculated using a weighted summation model, with the formula:

[0102] ;

[0103] in, This represents the final matching score. This represents key entity information extracted from voice commands. It represents a single candidate target attribute obtained from visual candidate information. , , These are functions used to calculate the degree of matching for categories, attributes, and spatial relationships, and their outputs are normalized scores. , , These are preset weighting coefficients used to adjust the importance of different types of information in the overall score. These weights are obtained through experimental tuning based on specific application scenarios and task types. For example, Scoring is achieved by comparing the nouns in the voice commands with the object labels recognized visually; Scoring is achieved by comparing the adjectives in the instructions with attributes such as color and shape detected visually.

[0104] For example, assuming the complete instruction is "Take away the red cup on the table," the robot extracts the key entity information. The existing visual candidate information is defined as follows: "Category: Cup", "Attribute: Red", "Space: Table". For a red cup, the category match is... Attribute matching degree , For a red box, the category matching degree Attribute matching degree If the weight is set to ,but The matching score is ; The matching score is .

[0105] When the matching score is lower than the threshold used to determine ambiguity, an active clarification question-and-answer process is triggered.

[0106] In some implementations, the calculated highest matching score is compared with a preset threshold used to determine ambiguity. The comparison is then performed. If the highest matching score is below the threshold, or if multiple candidate targets have very similar matching scores that are all above the threshold, then ambiguity is considered, and the user's intent cannot be uniquely determined. In this case, a proactive clarification question-and-answer process is triggered, sending a clarification request to the user.

[0107] For example, suppose the complete instruction is "Take that red thing". Since the instruction does not specify a category, the category weight... The weight is reduced to 0, focusing on color and spatial attributes. Calculations show that the red cup... and the red box Both scores had a match score higher than the absolute threshold of 0.9, but the difference between the two was significant. The value is less than the preset ambiguity threshold of 0.01. If the target cannot be uniquely identified, a proactive clarifying question-and-answer process is immediately triggered, asking the user: "There are two red things on the table: one is a cup, and the other is a box. Which one are you referring to?"

[0108] The visual candidate information is updated based on user feedback obtained from the proactive clarification question-and-answer process, and an updated verification result is generated.

[0109] In some implementations, additional information is obtained from the user's voice feedback during the proactive clarification question-and-answer process. For example, a user might answer, "It's the red one on the left." This user feedback, after being parsed, serves as a new constraint to update the visual candidate information. Candidate targets that do not meet the new conditions are eliminated, or the weight of candidate targets that do meet the new conditions is significantly increased. The matching score is then recalculated using the updated constraints until a unique, highly matching target is selected. This final result, confirmed by user feedback, constitutes the updated verification result, ensuring the accuracy of task planning.

[0110] For example, if a user answers, "Take the red box," the feedback is immediately parsed, and the new constraint "Category: Box" is added. The vector is processed, and the visual candidate information is re-verified. At this point, the matching score of the red box significantly surpasses that of the red cup, identifying the red box as the sole target object. An updated verification result is generated, and the precise 3D pose information of this target object is passed to subsequent task planning steps.

[0111] In one possible implementation, combining Figure 2 The aforementioned proactive clarification-based question-and-answer process, which generates a refined task plan based on user feedback, can be explained as follows:

[0112] Based on the visual candidate information, several candidate target options are generated;

[0113] In some implementations, when multiple potential target objects are identified as having a high degree of matching with the user's command, or when the matching degree of any target does not exceed a confidence threshold, concise and distinctive descriptive options—candidate target options—are generated for each potential target based on the current visual candidate information. When generating these options, the most differentiated features among the candidate targets are intelligently extracted. For example, if there are two cups in the scene, one red and one blue, the options will be generated based on color; if the two cups are the same color, their spatial relative positions will be used to generate options such as "the cup on your left" or "the cup near the window."

[0114] For example, after performing cross-modal validation, an ambiguity is identified where the visual candidate information contains two highly similar red targets: a red cup A, located on the left side of the table, and a red box B, located on the right side of the table. The system intelligently selects category and relative position as discriminative features, generating two candidate target options: "the cup closer to the left" and "the box closer to the right".

[0115] The system uses speech synthesis to ask the user a clarifying question containing the candidate options, receives and parses the user's voice feedback, and obtains the user's selection intent.

[0116] In some implementations, the built-in speech synthesis function is invoked to organize these candidate options into clarifying questions that conform to natural language logic. For example, the system might synthesize and announce, "I see two cups. Are you referring to the red one or the blue one?" This question is clearly conveyed to the user through the robot's speaker, guiding the user to make a choice. The system then enters listening mode, receiving the user's voice feedback through a microphone array. The user's answer, such as "the red one," is captured in real time and transmitted to speech recognition and natural language understanding for analysis, accurately extracting the user's selection intent from their response. This involves more than just simple keyword matching; it also includes understanding pronouns and requires combining the question context mentioned above to precisely map the user's ambiguous answer to a previously generated candidate option.

[0117] For example, the speech synthesis function is invoked to clarify the question to the user: "There are two red targets on the table. Do you mean the cup closer to the left or the box closer to the right?" The user's voice response is: "I want the box." The voice response is analyzed, the keyword "box" is extracted, and combined with the question context, it is determined that the user's intention is to select the candidate target "the box closer to the right".

[0118] Based on the user's selected intent, the initial task hypothesis model is optimized to generate a refined task plan.

[0119] In some implementations, once the user's selection intent is successfully parsed, crucial information to disambiguate the problem is obtained. This information is used to optimize the initial preliminary task hypothesis model. The unique goal selected by the user and all its associated precise information are solidified into the task model, replacing the original vague, multi-possible goal description. After this optimization step, the previously uncertain initial task hypothesis model evolves into a refined task plan with clearly defined instructions and a unique goal. This plan contains all the precise parameters required to execute the task.

[0120] For example, the user's selected intent "near the box on the right" is bound to its precise coordinates B to optimize the initial preliminary task hypothesis model {intent: take, attribute: color - red, area: desktop}, generating a refined task plan {intent: take, target ID: red box, target category: box, precise pose: coordinate B}. This plan contains all the precise parameters required to perform the task and can directly control the motion actuator to start the action.

[0121] In one possible implementation, combining Figure 2 The above-mentioned receiving of the refined task plan and control of robot movement to execute the task, while acquiring multi-view images by executing motion patterns for enhanced recognition, and fusing the multi-view images can be specifically explained as follows:

[0122] Monitor the confidence parameter of AI visual recognition on the target object, and activate the motion mode for enhanced recognition when the confidence parameter is lower than the threshold used to trigger enhanced perception.

[0123] In some implementations, after receiving the refined task plan, the process of executing the task and enhancing perception are executed simultaneously. This aims to address the perceptual uncertainty caused by changes in perspective, occlusion, or unclear object features when the robot approaches or manipulates a target. This process is initiated synchronously when the robot begins physical movement according to the refined task plan. As the robot moves towards the target object to perform the task, the AI ​​visual recognition continuously tracks the target and outputs key indicators in real time, namely the confidence parameter. This confidence parameter is a quantified value representing the degree of confidence the visual algorithm has in the detected object in the current frame image that it is indeed the target object. This confidence parameter is continuously compared with a preset threshold used to trigger enhanced perception. When the target is partially occluded, or its visual features become blurred due to poor lighting or perspective, the confidence parameter decreases. Once the confidence parameter is detected to be below the threshold, a motion mode for enhanced recognition is activated. This does not terminate the main task, but rather superimposes a fine-tuned sub-motion for perception purposes onto the main task path.

[0124] For example, the robot is performing the task of "grabbing the blue toolbox". When the robotic arm approaches the target, the shadow of the robotic arm itself obscures the side of the toolbox, causing the confidence score of the target category output by the AI ​​visual recognition algorithm to drop rapidly from the normal 0.92 to 0.78. Since 0.78 is lower than the preset enhanced perception threshold of 0.80, the motion control immediately receives the instruction and starts the motion mode for enhanced recognition without stopping the main task of grasping.

[0125] Perform arc motion around the target or vertical motion The system elevates and lowers the characters, acquires multi-view images, and performs 3D point cloud fusion processing on the multi-view images to generate enhanced perception data.

[0126] In some implementations, after initiating the motion mode, the motion control executes a pre-set specific trajectory designed to observe the target from multiple angles. For example, it performs an arc-shaped motion around the target, controlling the robot's end effector or moving chassis to move along an arc path while maintaining a relatively constant distance from the target, thereby acquiring a side image of the target object. Another mode is a vertical zigzag ascent and descent, controlling the robot arm to drive the camera in a zigzag trajectory of first rising, then translating, and finally descending, to acquire images of the object at different heights and horizontal offsets. While executing these specific motion trajectories, the robot's vision sensors continuously acquire a series of multi-view images. The acquired multi-view images are then subjected to 3D point cloud fusion processing. This process uses the depth information generated from each frame to construct a local, fragmented 3D point cloud, which is then combined with the motion trajectory data recorded by the robot itself to precisely align and stitch these local point clouds from different perspectives into a more complete and dense whole. This fusion process ultimately generates richer and more accurate augmented perception data, containing a more comprehensive description of the target object's 3D morphology.

[0127] For example, upon detecting a decrease in confidence, the motion control system immediately drives the robot's camera to execute a vertical zigzag lifting trajectory, with the trajectory height changing. Horizontal offset Within 3 seconds, 15 frames of RGB-D images were continuously acquired. Subsequently, a 3D point cloud fusion processing algorithm combined the depth information of these 15 frames with the robot's pose change data to accurately register and stitch the scattered point cloud data, ultimately generating a more complete dense 3D model of the toolbox than the initial perception. This model served as augmented perception data, improving the accuracy of edge and texture recognition for target objects. Figure 3As shown, the diagram illustrates the real-time changes in the confidence parameters of the AI ​​visual recognition of the target object during task execution, and marks the threshold line used to trigger the enhanced recognition motion mode. The results clearly demonstrate how the enhanced recognition motion mode is immediately activated and multi-view information is actively collected when the confidence parameters fall below the threshold due to environmental or perspective factors.

[0128] In one possible implementation, combining Figure 2 The above-mentioned 3D point cloud fusion processing of multi-view images can be explained in detail as follows:

[0129] Feature points are extracted and matched from the multi-view images to generate a feature point set;

[0130] In some implementations, feature points are extracted and matched from the acquired multi-view image sequence. Efficient feature point detection and description algorithms, such as ORB (Oriented Fast and Rotated BRIEF), are used to identify stable, recognizable local regions like corners and blobs as feature points in each image. By comparing the descriptors of these feature points, correspondences are found between different images, i.e., the projections of the same 3D point in different viewpoints are identified. All these successfully matched feature point pairs are aggregated to form a feature point set containing cross-image correspondences.

[0131] For example, the robot acquired 10 frames of images around the blue toolbox. The ORB algorithm detected 500 feature points on the edge of the toolbox in the first frame and 480 feature points in the second frame. After descriptor matching and RANSAC filtering, 350 reliable pairs of cross-frame corresponding feature points were successfully found. These data constituted the feature point set for the next step of processing.

[0132] The feature point set is processed using the structure-reconstruction-motion algorithm to generate a sparse point cloud;

[0133] In some implementations, the Structure for Motion (SfM) algorithm is used to process this set of feature points. SfM is a technique that can simultaneously estimate 3D structure and camera motion. Using the set of feature points as input, iterative optimization is performed through multi-view geometric constraints to calculate two key outputs: first, the 3D spatial coordinates of each feature point, which together constitute a sparse point cloud; and second, the precise camera pose for each image captured. This sparse point cloud provides the basic skeletal structure of the target object, while the precise camera pose is the foundation for subsequent denser reconstruction.

[0134] For example, 350 pairs of matching feature points are input into the SfM algorithm. After optimization processes such as bundle adjustment, the algorithm calculates the three-dimensional coordinates of the 350 feature points in the robot coordinate system, generates a sparse point cloud of the toolbox edge contour, and accurately determines the 10 precise poses and motion trajectories of the camera in space when these 10 frames of images are acquired.

[0135] A multi-view stereo algorithm is applied to densify the sparse point cloud to generate a dense 3D model.

[0136] In some implementations, after obtaining the sparse point cloud and camera pose, a multi-view stereo (MVS) algorithm is applied to densen the sparse point cloud. The MVS algorithm uses the known camera pose to estimate the depth of non-feature point regions in the image. For each pixel in the reference image, the best match is searched on the corresponding epipolar line in other view images, and the depth information of that pixel is calculated using triangulation principles. By repeating this process on a large number of pixels, the dense 3D points of the object's surface can be recovered, thus filling the sparse skeleton into a dense 3D model containing rich surface details, also known as a dense point cloud.

[0137] For example, the MVS algorithm is used, combined with the solved poses of 10 cameras, to perform depth estimation and geometric consistency verification on all non-feature point pixels in the image. Finally, the sparse point cloud is filled into a dense 3D model of a toolbox containing more than 20,000 3D points. This model not only has the outline of the toolbox, but also includes the texture and bump details of its surface.

[0138] Based on the dense 3D model, the position and orientation information of the target object are updated to generate updated enhanced perception data.

[0139] In some implementations, the final target information is updated based on the generated dense 3D model. By performing geometric analysis on this dense point cloud data, such as calculating its centroid, bounding box, or principal axis direction, the 3D position and orientation information of the target object in the robot coordinate system can be calculated with extremely high precision. This high-precision position and orientation information, along with the dense 3D model itself, is encapsulated into updated enhanced perception data, replacing the previous perception results with lower confidence, providing final and highly reliable data support for the robot's subsequent precise operations.

[0140] For example, a geometric analysis is performed on a model consisting of 20,000 3D points to calculate the toolbox's 3D centroid as its precise position, and the orientation of its minimum bounding box is calculated as its pose information. Updated augmented perception data indicates that the toolbox's precise 3D pose is... This data replaced the initial low-confidence data and can be directly used for the robotic arm's grasping planning.

[0141] In one possible implementation, combining Figure 2 The spatial anchoring process based on the enhanced perception data and combined with the robot's own pose change data can be specifically explained as follows:

[0142] The robot acquires real-time pose change data from inertial measurement output, and calculates its own pose transformation matrix by combining it with a kinematic model describing the relationship between the robot's joints and end effectors.

[0143] In some implementations, the robot's own pose change data is acquired in real time. This data is fused from two sources. The first is the robot's internal inertial measurement unit (IMU), which can output high-frequency linear acceleration and angular velocity information of the robot body. Through integration, the pose change of the robot base can be estimated in real time. The second is a kinematic model describing the relationship between the robot's joints and end effector, also known as forward kinematics. By reading the real-time angle values ​​of each joint encoder and substituting them into this model, the pose of the robot end effector with the camera mounted relative to the robot base can be accurately calculated. Combining these two, a precise transformation matrix representing the camera's current pose in the robot base coordinate system can be calculated.

[0144] For example, as the robot moves toward the target toolbox, its chassis IMU sensors report a slight vibration and linear acceleration within 0.1 seconds. Meanwhile, the robot's multi-axis attitude adjustment component encoder reported a 0.2-degree fine-tuning of the joint angle of the end effector where the camera is located. Immediately, the IMU data and the 0.2-degree joint angle were substituted into the kinematic model describing the relationship between the robot's joints and the end effector, and the camera coordinate system relative to the robot's base coordinate system was calculated. The total pose transformation matrix is:

[0145] ;

[0146] This matrix accurately reflects the camera's position at the current moment. The real-time position and attitude provide basic data for subsequent spatial anchoring.

[0147] The pose transformation matrix is ​​synchronized to the perception of AI visual recognition to correct target coordinate drift and output stable coordinate information.

[0148] In some implementations, the pose transformation matrix is ​​used as the core calculation for spatial anchoring. The stable coordinate information of the target object in the robot's base coordinate system is represented as... At any time The stable coordinate is calculated and maintained using the following formula:

[0149] ;

[0150] in, It is the perception of AI visual recognition. The target's coordinates in the current camera coordinate system are output at any given time. These are instantaneous observations obtained directly from augmented sensing data or through subsequent tracking. This is the pose transformation matrix calculated by fusing inertial measurement and kinematic models, representing the transformation relationship from the camera coordinate system to the robot's base coordinate system. Through computation, the unstable coordinates of the target object relative to the camera are determined. It was converted in real time to a stable coordinate system relative to the robot base. In the middle. Because the robot base is usually fixed during the mission, therefore This constitutes a constant positional anchor point in the robot's worldview. This calculated pose transformation matrix... It will be synchronized in real time to the AI ​​visual recognition perception system to correct and predict the position where the target should appear in the next frame of the image, thereby effectively suppressing the target coordinate drift and finally outputting continuously updated stable coordinate information.

[0151] For example, suppose the robot is in At any given moment, the toolbox for augmented perception data reporting uses the center coordinates in the camera coordinate system. for Using the pose transformation matrix Perform the following matrix multiplication:

[0152] ;

[0153] Calculation results This refers to the stable and precise spatial anchoring coordinates of the toolbox within the robot's base coordinate system. This stable coordinate information is used for the final task execution planning, ensuring that the target grasping coordinates remain accurate even if the robot itself undergoes slight pose changes.

[0154] In one possible implementation, combining Figure 2 The real-time pose change data obtained from the inertial measurement output, combined with the kinematic model describing the relationship between the robot's joints and end effectors, can be explained as follows:

[0155] By performing the hand-eye calibration process, multiple sets of camera pose transformation data and robot end-effector pose transformation data are collected.

[0156] In some implementations, this is achieved by directing the robot to execute a pre-set sequence of calibration actions. During this process, the robot moves its end effector to a series of different positions and poses. Simultaneously, a camera mounted on the robot's end effector or arm continuously observes a calibration object with known geometric dimensions, such as a checkerboard or dot array calibration board, fixed in the scene. After each new calibration pose, two sets of key data are simultaneously acquired. The first set is the robot's end effector pose transformation data, directly provided by the robot controller. Based on the joint encoder readings and the forward kinematics model, the controller accurately calculates the pose transformation of the end effector from the previous pose to the current pose; this data constitutes the robot's end effector pose transformation matrix. The second set is the camera pose transformation data, calculated by the AI ​​vision system. The vision system analyzes images of the calibration board taken by the camera at two different positions to calculate the pose transformation of the calibration board relative to the camera coordinate system; this constitutes the camera pose transformation matrix.

[0157] For example, the robot is programmed to execute 20 different sequences of "move-pause" actions. In the... During this movement, the robot's motion controller records the position of its end effector. to posture pose transformation matrix Simultaneously, the images of the calibration board captured by the camera are processed and their attitude is calculated to determine the distance the calibration board travels from the camera. Coordinate system to camera Relative pose transformation matrix of coordinate system Ultimately, 20 sets were obtained. and The corresponding data pairs.

[0158] Construct hand-eye calibration equations, where camera pose transformation is used as a matrix. The robot's end-effector pose transformation is a matrix ;

[0159] In some implementations, multiple sets, typically dozens, of collected data are used to construct and solve the hand-eye calibration equation. The classic form of this equation is:

[0160] ;

[0161] in, This represents the camera pose transformation matrix, which is the relative motion of the calibration board in the camera coordinate system. This matrix is ​​calculated by a vision algorithm from two consecutive frames of images. This represents the robot end-effector pose transformation matrix, which is the relative motion of the robot end-effector in the base coordinate system. This matrix is ​​directly provided by the robot motion controller. The unknowns to be solved are the fixed transformation matrices. This represents the constant transformation relationship from the robot's end-effector coordinate system to the camera coordinate system, and is the ultimate goal of this calibration process.

[0162] For example, suppose that during a certain movement, the robot controller records the end-effector transformation matrix. The camera transformation matrix solved by the vision system as follows:

[0163] ;

[0164] ;

[0165] Construct an overdetermined system of equations consisting of 20 matrix equations. The goal is to find a unique solution. matrix.

[0166] Solving the fixed transformation matrix using an algorithm for solving overdetermined systems of equations And using a fixed transformation matrix Establish a precise spatial mapping relationship to achieve coordinate transformation from the camera coordinate system to the robot coordinate system.

[0167] In some implementations, due to the collection of multiple sets of data, the equations constitute an overdetermined system of equations. Mature algorithms for solving overdetermined systems of equations, such as the Tsai-Lenz method or the split-method approach, are used to calculate the optimal fixed transformation matrix. This solution process effectively smooths out noise and errors from a single measurement, yielding high-precision results. Once the transformation matrix is ​​fixed... Once solved, the data is stored as a robust geometric binding between the camera and the robot body. From this point onward, a precise spatial mapping is established, enabling the camera to read the real-time pose of the robot's end effector and combine it with matrix data. It can accurately perform coordinate transformations from the camera coordinate system to the robot coordinate system at any time. For example... Figure 4 As shown, the original coordinate drift and the residual of the stable coordinates after optimization and solution of the overdetermined equations are compared in the hand-eye calibration process. The results intuitively demonstrate that solving the fixed transformation matrix using algorithms such as SVD is effective. This can effectively reduce the average coordinate residual, thereby achieving high-precision spatial anchoring.

[0168] For example, the Singular Value Decomposition (SVD) method is used to optimize and solve 20 sets of equations. After iterative convergence, a fixed camera-end transformation matrix is ​​obtained. :

[0169] ;

[0170] The matrix This includes the fixed translation and rotation relationship of the camera relative to the robot's end effector. In subsequent tasks, the robot's end effector pose will be read in real time. It can be used The matrix calculates the real-time pose of the camera in the base coordinate system. It achieves precise spatial mapping from the camera coordinate system to the robot base coordinate system, with an accuracy of sub-millimeter level.

[0171] In one possible implementation, combining Figure 2 The generation of contextualized feedback information using the aforementioned stable coordinate information and task execution status can be specifically explained as follows:

[0172] Visual context information is integrated from the enhanced perception data, and motion history is fused from the task execution state;

[0173] In some implementations, visual contextual information is integrated from augmented perception data. This includes high-resolution images of the final confirmed target object or its 3D reconstructed model, snapshots of the target object's surrounding environment, and visual records of key nodes during task execution. For example, the target state before grasping and the state of the arm carrying the target after successful grasping. Motion history data is also fused from the task execution state. This includes complete motion trajectory data of the robot from receiving instructions to the end of the task, motion logs of each joint, and key events encountered during task execution, such as triggering motion patterns for augmented recognition or performing obstacle avoidance maneuvers.

[0174] For example, after completing the "grab toolbox" task, the integrated data includes: 1) Visual context information: a dense 3D model of the toolbox, and high-resolution snapshots of the toolbox and robotic arm when the grasp is successful; 2) Motion history: motion logs show that during the approach to the target, the confidence level decreased and triggered a vertical movement. The enhanced motion sensing mode of the character's rise and fall was achieved, and the grasping end effector was successfully positioned in a stable coordinate system. Place.

[0175] Using a large language model for semantic generation, the visual context information and the motion history are processed into a natural language description;

[0176] In some implementations, a large language model for semantic generation is invoked, taking integrated visual context information and motion history as input and processing it into easily understandable natural language descriptions. This large language model is specifically fine-tuned to understand and describe the robot's behavior and perception. It transforms dry coordinate data and joint angles into vivid action descriptions; for example, describing "performing an arc motion around a target" as "turning half a circle around it to see it more clearly." It also transforms visual data into scene descriptions, such as "picking up the red cup on the table; there's a book next to it."

[0177] For example, the large language model receives "The toolbox has been fetched and executed". After processing structured data such as "graph movement," the system leverages its context awareness and semantic generation capabilities to output a natural feedback description: "I have successfully picked up that blue toolbox. When approaching it, because the light was a bit dim, I deliberately made a slight up-and-down movement to ensure a clearer view, ultimately reaching the precise location..." Fetch successful.

[0178] The natural language description and generative image content are combined to form contextualized feedback information.

[0179] In some implementations, generative image content technology is used in parallel to generate natural language descriptions and create visual materials that match the feedback content. This is not simply a reproduction of the original image, but rather allows for artistic processing or highlighting as needed. For example, the target object being manipulated can be highlighted or marked with arrows on a scene image, or a simplified animation can be generated to demonstrate the key motion paths performed by the robot. Combining the generated natural language descriptions with the generative image content creates a multimodal, information-rich, and contextualized feedback message. This information can be output to the user in various forms, such as displaying a graphic report on the robot's screen while simultaneously reading the natural language description through speech synthesis. Users will see not only the result, but also a vivid recap of the entire task process, understanding how the robot understood instructions, observed the environment, overcame difficulties, and ultimately completed the task.

[0180] For example, the natural language description "The blue toolbox has been successfully picked up..." is sent to the voice broadcast. Simultaneously, the generative image content function generates a highlighted image based on the captured snapshot: the picked toolbox is circled in green on the original image, and dashed arrows indicate the robot's actions. The trajectory of the character's movement is ultimately displayed synchronously on the screen as a detailed and illustrated review, forming a complete contextualized feedback for the user to understand and verify.

[0181] It should be noted that the electrical connections between the various units described above do not necessarily represent direct or indirect connections. Any indirect connection method can be applied to the embodiments of the present invention as long as it achieves the purpose of the present invention. The above descriptions are merely exemplary embodiments of the present invention and should not be construed as limiting the scope of the present invention.

[0182] All equivalent changes and modifications made in accordance with the teachings of this invention are still within the scope of this invention. Those skilled in the art will readily conceive of other embodiments of this invention upon considering the specification and the disclosure of practical truth. This application is intended to cover any variations, uses, or adaptations of this invention that follow the general principles of this invention and include common knowledge or conventional techniques in the art not described herein.

Claims

1. A question-and-answer interaction method for robot audiovisual fusion, characterized in that, The method includes: The system acquires user voice commands and performs real-time streaming recognition and conversion to generate preliminary recognized text. Based on the initially identified text, it is processed by a preset model for inferring user intent to generate a preliminary task hypothesis model; Receive the preliminary task hypothesis model and dynamically adjust the perceptual resources of AI visual recognition to generate visual candidate information; By combining the fully recognized voice command with the visual candidate information, cross-modal verification is performed to generate a verification result. This cross-modal verification includes: extracting key entity information from the fully recognized voice command; comparing the key entity information with target attributes in the visual candidate information to generate a matching score; triggering an active clarification question-answering process when the matching score is lower than a threshold for ambiguity judgment; and updating the visual candidate information based on user feedback obtained from the active clarification question-answering process to generate an updated verification result. If the verification result indicates ambiguity, a proactive clarification question-and-answer process is initiated. A refined task plan is generated based on user feedback. If not, but the indicated perspective is unfavorable, the robot is instructed to move to change the perspective, and a refined task plan is generated based on the changed perspective information. Initiating the proactive clarification question-and-answer process and generating a refined task plan based on user feedback includes: generating several candidate target options based on the visual candidate information; posing a clarification question containing the candidate target options to the user through speech synthesis; receiving and parsing the user's voice feedback to obtain the user's selection intention; and optimizing the preliminary task hypothesis model based on the user's selection intention to generate a refined task plan. The system receives the refined task plan, controls the robot to move and execute the task, and simultaneously acquires multi-view images by executing a motion mode for enhanced recognition. These multi-view images are then fused to generate enhanced perception data. The process of receiving the refined task plan, controlling the robot to move and execute the task, and simultaneously acquiring multi-view images by executing a motion mode for enhanced recognition and fusing the multi-view images includes: monitoring the confidence parameter of the AI ​​visual recognition on the target object; when the confidence parameter is lower than the threshold used to trigger enhanced perception, activating the motion mode for enhanced recognition; executing arc-shaped motion around the target or vertical Z-shaped lifting and lowering, acquiring multi-view images, and performing 3D point cloud fusion processing on the multi-view images to generate enhanced perception data. Based on the enhanced perception data and combined with the robot's own pose change data, spatial anchoring processing is performed to generate stable coordinate information; Using the stable coordinate information and task execution status, contextualized feedback information is generated and output to the user.

2. The question-and-answer interaction method for robot audiovisual fusion according to claim 1, characterized in that, The dynamically adjusted perceptual resources for AI visual recognition include: Based on the preliminary task hypothesis model, the hypothetical target area is determined, and the robot gimbal or body is controlled to face the hypothetical target area to achieve focused view. The image recognition algorithm associated with the preliminary task hypothesis model is enabled to achieve algorithm focus; Based on the results of the perspective focusing and algorithm focusing, visual candidate information is generated.

3. The question-and-answer interaction method for robot audiovisual fusion according to claim 1, characterized in that, The three-dimensional point cloud fusion processing of multi-view images includes: Feature points are extracted and matched from the multi-view images to generate a feature point set; The feature point set is processed using the structure-reconstruction-motion algorithm to generate a sparse point cloud; A multi-view stereo algorithm is applied to densify the sparse point cloud to generate a dense 3D model. Based on the dense 3D model, the position and orientation information of the target object are updated to generate updated enhanced perception data.

4. The question-and-answer interaction method for robot audiovisual fusion according to claim 1, characterized in that, Based on the enhanced perception data, and combined with the robot's own pose change data, spatial anchoring processing includes: The robot acquires real-time pose change data from inertial measurement output, and calculates its own pose transformation matrix by combining it with a kinematic model describing the relationship between the robot's joints and end effectors. The pose transformation matrix is ​​synchronized to the perception of AI visual recognition to correct target coordinate drift and output stable coordinate information.

5. The question-and-answer interaction method for robot audiovisual fusion according to claim 4, characterized in that, The acquisition of real-time pose change data from inertial measurement output, combined with a kinematic model describing the relationship between the robot joints and the end effector, includes: By performing the hand-eye calibration process, multiple sets of camera pose transformation data and robot end-effector pose transformation data are collected. Construct hand-eye calibration equations, where the camera pose transformation is a matrix and the robot end effector pose transformation is a matrix. The fixed transformation matrix is ​​solved by an algorithm for solving overdetermined equations, and then a precise spatial mapping relationship is established using the fixed transformation matrix to realize the coordinate transformation from the camera coordinate system to the robot coordinate system.

6. The question-and-answer interaction method for robot audiovisual fusion according to claim 1, characterized in that, Using the stable coordinate information and task execution status, contextualized feedback information is generated, including: Visual context information is integrated from the enhanced perception data, and motion history is fused from the task execution state; Using a large language model for semantic generation, the visual context information and the motion history are processed into a natural language description; The natural language description and generative image content are combined to form contextualized feedback information.

7. A question-and-answer interactive system for robot audiovisual fusion, characterized in that, The system is used in a question-and-answer interaction method for robot audiovisual fusion as described in any one of claims 1-6, the system comprising: The voice command processing module is used to acquire user voice commands, perform real-time streaming recognition and conversion, and generate preliminary recognized text. The intent inference module is used to process the pre-defined user intent model based on the pre-identified text to generate a preliminary task hypothesis model. The image perception and adjustment module is used to receive the preliminary task hypothesis model and dynamically adjust the perception resources of AI visual recognition to generate visual candidate information. A cross-modal verification module is used to perform cross-modal verification by combining the fully recognized voice command with the visual candidate information and generate a verification result. The cross-modal verification by combining the fully recognized voice command with the visual candidate information includes: extracting key entity information from the fully recognized voice command; comparing the key entity information with target attributes in the visual candidate information to generate a matching score; triggering an active clarification question-and-answer process when the matching score is lower than a threshold used to determine ambiguity; updating the visual candidate information based on user feedback obtained from the active clarification question-and-answer process, and generating an updated verification result. The task planning and clarification module is used to determine whether the verification result indication is ambiguous. If yes, it initiates an active clarification question-and-answer process to generate a refined task plan based on user feedback. If no, but the display perspective is poor, it instructs the robot to move to change the perspective and generates a refined task plan based on the changed perspective information. The initiation of the active clarification question-and-answer process and the generation of a refined task plan based on user feedback includes: generating several candidate target options based on the visual candidate information; asking the user a clarification question containing the candidate target options through speech synthesis; receiving and parsing the user's voice feedback to obtain the user's selection intention; and optimizing the preliminary task hypothesis model according to the user's selection intention to generate a refined task plan. The motion control and enhanced perception module is used to receive the refined task plan, control the robot's motion to execute the task, and simultaneously acquire multi-view images by executing motion modes for enhanced recognition. It then fuses these multi-view images to generate enhanced perception data. The steps of receiving the refined task plan, controlling the robot's motion to execute the task, acquiring multi-view images by executing motion modes for enhanced recognition, and fusing these images include: monitoring the confidence parameter of the AI ​​visual recognition on the target object; activating the motion mode for enhanced recognition when the confidence parameter is lower than a threshold used to trigger enhanced perception; executing arc-shaped motion around the target or vertical Z-shaped lifting and lowering; acquiring multi-view images; and performing 3D point cloud fusion processing on the multi-view images to generate enhanced perception data. The spatial anchoring module is used to perform spatial anchoring processing based on the enhanced perception data and combined with the robot's own pose change data to generate stable coordinate information; The feedback generation and information output module is used to generate contextualized feedback information using the stable coordinate information and task execution status, and output the contextualized feedback information to the user.