Question and answer mode interaction method and system for robot audio-visual fusion
By linking streaming speech recognition with visual perception to perform cross-modal verification and proactive clarification question-and-answer processes, the problem of insufficient intent understanding and perception of intelligent robots in complex environments has been solved, achieving high-precision task execution and a smooth interactive experience.
Patent Information
- Application Number
- CN202511769361.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2045-11-28
AI Technical Summary
Existing intelligent robot systems lack proactive clarification mechanisms when faced with complex, ambiguous, or environmentally sensitive instructions, resulting in insufficient precision in intent understanding and perception, making it difficult to accurately execute tasks.
It employs streaming speech recognition and intent inference in conjunction with visual perception, eliminates ambiguity through cross-modal verification, and ensures execution accuracy in the task through enhanced perception and spatial anchoring, including dynamically adjusting visual resources, proactively clarifying the question-and-answer process, and multi-view image fusion processing.
It improves the fluency and robustness of robot interaction, ensures the accuracy and success rate of task execution, and provides contextualized feedback to enhance user trust.
Smart Images

Figure CN121315979A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence and human-computer interaction, in particular to a robot audio-visual fusion question and answer interaction method and system. BACKGROUND
[0002] Intelligent robot technology is gradually integrated into many fields such as industry, service and life. In terms of human-computer interaction, the ability of robots to perform complex tasks is a key indicator of their intelligent level. Existing intelligent robot systems often rely on pre-set instructions or single-modal input, and when faced with complex, ambiguous or environment-judgment-related instructions from users, their intent understanding, accurate perception and action execution capabilities are obviously limited.
[0003] In related technologies, a patent with publication number CN119927906A discloses an interactive method for upper limb cooperative control of humanoid robots. The method includes: performing task analysis on the received voice request; if there is a target object that needs to be executed, performing target object recognition processing according to the task analysis result; based on the target recognition information and target pose information output by the recognition processing, and the task analysis result, using a large language model to perform task planning processing according to the principle of double-arm division of labor, generating a task sequence and the corresponding text instructions for each task; for each target task in the task sequence, the following operations are performed in turn: based on the target pose information and the target recognition information, controlling the upper limbs of the humanoid robot to perform the operation action corresponding to the target task on the target object.
[0004] For the related technologies in the above, the inventors believe that: the existing technology mainly focuses on the double-arm cooperative control of the upper limbs of humanoid robots and the generation of task sequences based on large language models, but does not fully solve the problems of instruction ambiguity in multi-modal question and answer, perception enhancement under limited visual angle, and high-precision three-dimensional space anchoring. In a complex real environment, when the voice instruction and visual perception information do not completely match, this technology lacks an active clarification mechanism, which can easily cause errors in robot task planning; at the same time, when the robot has a poor visual angle, it lacks the ability to execute enhanced recognition motion patterns to actively obtain high-precision perception data, limiting the accuracy of task execution. Therefore, there is an urgent need for a robot audio-visual fusion question and answer interaction method that can actively solve instruction ambiguity and enhance environmental perception accuracy. SUMMARY
[0005] To solve the above problems, the present application provides a robot audio-visual fusion question and answer interaction method and system, which adopts a technical solution of streaming voice recognition and intent speculation linked visual perception, cross-modal verification and active clarification to eliminate ambiguity, and enhanced perception and spatial anchoring in tasks to ensure execution accuracy, which can improve the fluency, robustness and success rate of robot interaction.
[0006] To achieve the above object, the application adopts the following technical solutions: In a first aspect, a robot audio-visual fusion question and answer interaction method is provided, comprising: Obtaining a user voice instruction and performing real-time streaming recognition conversion to generate a preliminary recognition text; Based on the preliminary recognition text, processing through a pre-set model for inferring user intent to generate a preliminary task hypothesis model; Receiving the preliminary task hypothesis model and dynamically adjusting the perception resources of AI visual recognition to generate visual candidate information; Combining the complete recognition voice instruction with the visual candidate information, performing cross-modal verification to generate a verification result; Judging whether the verification result indicates the existence of ambiguity, if yes, initiating an active clarification question and answer process, generating a refined task plan based on user feedback, if no, but indicating poor perspective, instructing the robot to move to change the perspective, and generating a refined task plan based on the changed perspective information; Receiving the refined task plan and controlling the robot to move to execute the task, while obtaining multi-perspective images through the execution of a motion mode for enhancing recognition, and fusing the multi-perspective images to generate enhanced perception data; Based on the enhanced perception data and combining the robot's own pose change data, performing spatial anchoring processing to generate stable coordinate information; Using the stable coordinate information and the task execution state to generate situational feedback information, and outputting the situational feedback information to the user.
[0007] Based on the above technical solutions, in the robot audio-visual fusion question and answer interaction method provided by the application, the technical solutions of streaming voice recognition and intent inference combined with visual perception, cross-modal verification and active clarification to eliminate ambiguity, and enhanced perception and spatial anchoring in the task to ensure execution accuracy can improve the fluency, robustness and success rate of task execution of robot interaction.
[0008] In combination with the first aspect described above, in a possible implementation manner, the dynamic adjustment of the perception resources of AI visual recognition comprises: According to the preliminary task hypothesis model, determining a hypothesis target area and controlling the robot cloud cover or body to face the hypothesis target area to realize perspective focusing; Enabling an image recognition algorithm related to the preliminary task hypothesis model to realize algorithm focusing; Based on the results of perspective focusing and algorithm focusing, generating visual candidate information.
[0009] In a possible implementation manner of the first aspect, the combining the complete recognized voice instruction with the visual candidate information to perform cross-modal verification comprises: extracting key entity information from the complete recognized voice instruction, comparing the key entity information with a target attribute in the visual candidate information to generate a matching score; when the matching score is lower than a threshold value for judging ambiguity, triggering an active clarification question-answering process; updating the visual candidate information based on user feedback obtained from the active clarification question-answering process to generate an updated verification result.
[0010] In a possible implementation manner of the first aspect, the initiating the active clarification question-answering process and generating a refined task plan based on user feedback comprises: generating a plurality of candidate target options based on the visual candidate information; asking the user a clarification question containing the candidate target options through a voice synthesis function, receiving and analyzing user voice feedback to obtain a user selection intention; optimizing the preliminary task assumption model according to the user selection intention to generate a refined task plan.
[0011] In a possible implementation manner of the first aspect, receiving the refined task plan and controlling the robot to move to perform a task, while obtaining multi-view images through execution of a motion mode for enhancing recognition, and performing fusion processing on the multi-view images comprises: monitoring a confidence parameter of AI visual recognition on a target object, and when the confidence parameter is lower than a threshold value for triggering enhanced perception, starting a motion mode for enhancing recognition; performing arc motion around the target or vertical characteristic shape lifting, collecting multi-view images, and performing three-dimensional point cloud fusion processing on the multi-view images to generate enhanced perception data.
[0012] In a possible implementation manner of the first aspect, the three-dimensional point cloud fusion processing on the multi-view images comprises: extracting and matching feature points from the multi-view images to generate a feature point set; processing the feature point set by using a motion recovery structure algorithm to generate a sparse point cloud; applying a multi-view stereoscopic algorithm to densify the sparse point cloud to generate a dense three-dimensional model; updating position and posture information of the target object based on the dense three-dimensional model to generate updated enhanced perception data.
[0013] In a possible implementation manner of the first aspect, based on the enhanced perception data, and in combination with the pose change data of the robot itself, the spatial anchoring processing comprises: obtaining real-time pose change data of inertial measurement output, combining a kinematics model describing the relationship between joints and the end of the robot, and calculating a pose transformation matrix of the robot itself; synchronizing the pose transformation matrix to the perception of AI visual recognition, to correct target coordinate drift, and output stable coordinate information.
[0014] In a possible implementation manner of the first aspect, the obtaining of the real-time pose change data of the inertial measurement output, in combination with the kinematics model describing the relationship between the joints and the end of the robot comprises: acquiring a plurality of sets of camera pose transformation data and robot end pose transformation data by performing a hand-eye calibration process; constructing a hand-eye calibration equation, wherein the camera pose transformation is a matrix, and the robot end pose transformation is a matrix; solving the fixed transformation matrix by an algorithm for solving overdetermined equations, and using the fixed transformation matrix to establish an accurate spatial mapping relationship, to realize coordinate conversion from the camera coordinate system to the robot coordinate system.
[0015] In a possible implementation manner of the first aspect, the generation of situational feedback information using the stable coordinate information and the task execution state comprises: integrating visual context information from the enhanced perception data, and fusing motion history records from the task execution state; using a large language model for semantic generation to process the visual context information and the motion history records into a natural language description; combining the natural language description and the generated image content to form situational feedback information.
[0016] The second aspect provides a robot audio-visual fusion question-and-answer interaction system, comprising: a voice instruction processing module, an intention inference module, an image perception and adjustment module, a cross-modal verification module, a task planning and clarification module, a motion control and enhanced perception module, a spatial anchoring module, and a feedback generation and information output module; wherein: The voice instruction processing module is configured to obtain a user voice instruction, and perform real-time streaming recognition conversion to generate a preliminary recognition text; The intention inference module is configured to, based on the preliminary recognition text, process by using a preset model for inferring user intention to generate a preliminary task hypothesis model; The image perception and adjustment module is configured to receive the preliminary task hypothesis model, and dynamically adjust perception resources of AI visual recognition to generate visual candidate information; a cross-modal verification module configured to combine the complete recognized voice instruction with the visual candidate information to perform cross-modal verification and generate a verification result; a task planning and clarification module configured to determine whether the verification result indicates an ambiguity, and if so, initiate an active clarification question and answer process, generate a refined task plan based on user feedback, or if not, display a poor view angle and instruct the robot to perform view angle transformation and generate a refined task plan based on the transformed view angle information; a motion control and enhanced perception module configured to receive the refined task plan, control the robot to perform a task, acquire multi-view images by performing a motion pattern for enhancing recognition, and fuse the multi-view images to generate enhanced perception data; a spatial anchoring module configured to perform spatial anchoring processing based on the enhanced perception data and combined with the robot's own pose change data to generate stable coordinate information; a feedback generation and information output module configured to generate situational feedback information using the stable coordinate information and task execution status, and output the situational feedback information to the user.
[0017] Compared with the prior art, the present application has the following advantages: The present application realizes parallel processing of perception and understanding by performing real-time streaming recognition and intent speculation during user voice instruction input, and adjusting the perception resources of AI visual recognition in parallel, effectively shortening the delay from instruction input to robot response, and improving the fluency and immediacy of human-computer interaction.
[0018] The present application introduces cross-modal verification and active clarification question and answer process, so that the robot can initiate communication to eliminate uncertainty or obtain better observation conditions through autonomous view angle transformation when facing instruction ambiguity or visual information uncertainty, thereby enhancing the robustness of the system to complex environments and ambiguous instructions and ensuring the accuracy of task planning.
[0019] The present application combines motion patterns for enhancing recognition and spatial anchoring processing during task execution, so that the robot can actively acquire multi-view information to construct a more accurate three-dimensional model when the perception confidence is insufficient, and at the same time compensate for the coordinate drift caused by its own motion, ensuring continuous and accurate positioning of the target object and improving the success rate of physical interaction tasks.
[0020] The present application presents the task execution process of the robot, including its perception, decision and motion history, to the user in an easy-to-understand natural language and visual content, constructs a complete interaction closed loop, improves the transparency of interaction and the user's trust, and optimizes the overall human-robot collaboration experience.
[0021] It should be understood that the description of technical features, technical solutions, beneficial effects or similar language in this application does not imply that all features and advantages can be realized in any single embodiment. On the contrary, it can be understood that the description of a feature or a beneficial effect means that the specific technical feature, technical solution or beneficial effect is included in at least one embodiment. Therefore, the description of technical features, technical solutions or beneficial effects in this specification does not necessarily refer to the same embodiment. Further, the technical features, technical solutions and beneficial effects described in this embodiment can also be combined in any appropriate manner. Those skilled in the art will understand that the embodiments can be implemented without one or more specific technical features, technical solutions or beneficial effects of a specific embodiment. In other embodiments, additional technical features and beneficial effects can be identified in specific embodiments that do not embody all embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings required to be used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application, and those skilled in the art can obtain other drawings from these drawings without creating labor.
[0023] Figure 1 A structural architecture diagram of a robot audio-visual fusion question and answer interactive system provided by an embodiment of the present application; Figure 2 A flowchart of a robot audio-visual fusion question and answer interactive method provided by an embodiment of the present application; Figure 3 An AI visual confidence parameter and enhanced perception threshold comparison diagram provided by an embodiment of the present application.
[0024] Figure 4 An optimization effect diagram of a hand-eye calibration on a space anchor residual provided by an embodiment of the present application. DETAILED DESCRIPTION
[0025] In the description of the present application, unless otherwise specified, " / " means "or", for example, A / B can mean A or B. "And / or" in this document is only a description of the relationship between the associated objects, which means that there can be three relationships, for example, A and / or B, which means that there are three cases of A alone, A and B together, and B alone. In addition, "at least one" means one or more, and "multiple" means two or more. "First", "second", etc. do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different.
[0026] It should be noted that in this application, "exemplary" or "for example" and the like are used to indicate examples, instances, or illustrations. Any embodiment or design presented as "exemplary" or "for example" in this application should not be interpreted as more preferred or advantageous than other embodiments or design solutions. Rather, the use of "exemplary" or "for example" is intended to present relevant concepts in a concrete manner.
[0027] The robot audio-visual fusion question and answer interaction method provided by the embodiments of the application can be applied to a robot audio-visual fusion question and answer interaction system 100 as shown in Figure 1 Figure 1 The system includes a voice instruction processing module, an intention inference module, an image perception and adjustment module, a cross-modal verification module, a task planning and clarification module, a motion control and enhanced perception module, a spatial anchoring module, and a feedback generation and information output module. Wherein: The voice instruction processing module is configured to obtain a user voice instruction and perform real-time streaming recognition conversion to generate a preliminary recognition text. The intention inference module is configured to process the preliminary recognition text based on a pre-set model for inferring user intention to generate a preliminary task hypothesis model. The image perception and adjustment module is configured to receive the preliminary task hypothesis model and dynamically adjust the perception resources of AI visual recognition to generate visual candidate information. The cross-modal verification module is configured to combine the complete recognition voice instruction and the visual candidate information to perform cross-modal verification to generate a verification result. The task planning and clarification module is configured to determine whether the verification result indicates ambiguity. If yes, initiate an active clarification question and answer process, generate a refined task planning based on user feedback. If no, but the display angle is poor, instruct the robot to move to change the angle, and generate a refined task planning based on the changed angle information. The motion control and enhanced perception module is configured to receive the refined task planning and control the robot motion to execute the task, acquire multi-angle images through the execution of the motion mode for enhancing recognition, fuse the multi-angle images, and generate enhanced perception data. The spatial anchoring module is configured to perform spatial anchoring processing based on the enhanced perception data and the pose change data of the robot itself to generate stable coordinate information. The feedback generation and information output module is configured to generate situational feedback information using the stable coordinate information and the task execution state, and output the situational feedback information to the user.
[0028] As shown in Figure 2 As shown in the figure, this application provides a question-and-answer interaction method for robot audiovisual fusion, including: The system acquires user voice commands and performs real-time streaming recognition and conversion to generate preliminary recognized text. Based on the initially identified text, it is processed by a preset model for inferring user intent to generate a preliminary task hypothesis model; Receive the preliminary task hypothesis model and dynamically adjust the perceptual resources of AI visual recognition to generate visual candidate information; By combining the fully recognized voice commands with the visual candidate information, cross-modal verification is performed to generate verification results; If the verification result indicates ambiguity, an active clarification question-and-answer process is initiated, and a refined task plan is generated based on user feedback. If not, but the indicated perspective is poor, the robot is instructed to move to change the perspective, and a refined task plan is generated based on the changed perspective information. The robot receives the refined task plan and controls the robot's movement to execute the task. At the same time, it acquires multi-view images by executing motion patterns for enhanced recognition and fuses the multi-view images to generate enhanced perception data. Based on the enhanced perception data and combined with the robot's own pose change data, spatial anchoring processing is performed to generate stable coordinate information; Using the stable coordinate information and task execution status, contextualized feedback information is generated and output to the user.
[0029] It should be noted that by streaming and processing user voice commands in real time, preliminary task hypotheses are generated in advance. Based on these hypotheses, the robot's visual perception resources are proactively and dynamically allocated to achieve pre-focusing on potential target areas. The complete auditory commands are then cross-modal validated with the preliminary visual perception results to identify potential ambiguities or adverse observation conditions. Faced with uncertainty, instead of passively waiting or failing, the robot actively initiates clarifying dialogues or autonomously changes perspectives to obtain the key information needed for decision-making, thereby generating a refined and executable task plan. During the task execution phase, this method closely integrates physical motion with enhanced perception. The robot collects multi-view information through specific motion patterns and fuses it into more accurate enhanced perception data. Simultaneously, it continuously uses its own pose change data to spatially anchor the target, eliminating coordinate drift caused by self-motion and ensuring continuous and stable positioning. By integrating stable coordinates, execution status, and visual context throughout the entire task process, contextualized multimodal feedback information is generated and output to the user, completing the entire interaction loop.
[0030] In one possible implementation of the embodiments of this application, combined with Figure 2The dynamic adjustment of the perception resources of the AI visual recognition can be specifically described as follows: According to the preliminary task hypothesis model, a hypothesis target area is determined, and the robot cloud cover or body is controlled to be directed to the hypothesis target area, so as to realize the focus of the visual angle; In some implementations, the preliminary task hypothesis model is received. This model is based on the structured data obtained by analyzing the preliminary text of the streaming speech recognition, which contains the speculation of the user's possible intention, such as "take", "point", and contains attribute descriptions such as "red" and "on the table". After receiving the preliminary task hypothesis model, the visual angle focusing process is immediately started. The descriptions related to the spatial position such as "on the table", "on your left" are extracted from the preliminary task hypothesis model, and the hypothesis target area in the three-dimensional space is analyzed in combination with the prior knowledge map of the robot itself to the environment. Subsequently, the robot motion control system calculates the angle and direction that the robot cloud cover or body needs to rotate according to the coordinates of the hypothesis target area. The control command is sent to the corresponding drive unit to drive the camera to physically point to the area. This physical action enables the robot's visual sensor to preferentially collect scene images that are most likely to contain target objects, greatly reducing the scope of subsequent visual processing and avoiding inefficient scanning of the entire environment.
[0031] For example, when the robot receives the preliminary recognition text of the user's voice instruction as "take a red cup", the speculation model generates the preliminary task hypothesis model {intention: take, attribute: color is red, area: table top}, and the robot calculates the instruction that the cloud cover needs to be pitched down by 15 degrees and rotated to the left by 30 degrees according to the area description "table top", and drives the camera view to accurately focus on the table top area.
[0032] An image recognition algorithm related to the preliminary task hypothesis model is enabled to realize algorithm focusing; In some implementations, while the visual angle focusing is performed, the algorithm focusing process is started in parallel. The content in the preliminary task hypothesis model is deeply analyzed, and the keywords related to the attributes of the object such as color, shape, category, etc. are focused on. Based on these keywords, the most relevant algorithms for detecting these attributes are dynamically selected and enabled from the built-in AI visual algorithm library. For example, when the model contains the hypothesis of "red cup", the object detection model and the color segmentation or classification model are preferentially loaded and run. At the same time, algorithms unrelated to the current task, such as face recognition and text recognition, are placed in a low priority or suspended. This approach concentrates limited computing resources on solving the most core perception problem at the moment, achieving efficient use of computing resources.
[0033] For example, since the preliminary task hypothesis model contains the "color-red" and "take" intent, the color recognition algorithm based on HSV color space conversion and the three-dimensional object detection and recognition algorithm based on Faster R-CNN are immediately enabled, while the MPII human pose estimation algorithm and the license plate recognition algorithm are temporarily suspended.
[0034] Based on the results of the view focusing and algorithm focusing, visual candidate information is generated.
[0035] In some implementations, the image data collected by the camera after view focusing is sent to the specific image recognition algorithm activated after algorithm focusing for processing. The processing results of the algorithm, such as all red objects and their positions, confidence, and other information identified in the assumed target area, are integrated and structured to form a list of visual candidate information. This visual candidate information, as a preliminary visual solution to the user's ambiguous instruction, will be passed to the subsequent process to be compared with the complete voice instruction to provide key visual basis for the final task decision.
[0036] For example, after focusing and algorithm processing, it is identified that there are 3 objects on the desktop, and the visual candidate information is structured as: "{Target 1: Category: Cup, Color: Red, 3D coordinates:, Confidence: 0.95}"; {Target 2: Category: Box, Color: Red, 3D coordinates:, Confidence: 0.88}"; {Target 3: Category: Pen container, Color: Blue, 3D coordinates:, Confidence: 0.92:}" and this list is passed to the subsequent cross-modal verification step.
[0037] In a possible implementation, in combination with Figure 2 The above cross-modal verification of the complete recognized voice instruction and the visual candidate information can be specifically explained as follows: Extracting key entity information from the complete recognized voice instruction, comparing the key entity information with the target attributes in the visual candidate information, and generating a matching degree score; In some implementations, the complete recognized speech instruction is processed by a deep natural language understanding to extract key entity information. These information are words that describe the specific attributes and relationships of the target object, usually including object category, object attribute, and spatial relationship. The extracted key entity information is structured to form a target description vector in the auditory modality. This auditory target description vector is compared with visual candidate information one by one, and a quantitative matching degree score is generated. The visual candidate information itself is a list, in which each candidate target is accompanied by its visual attributes, such as the category label given by the image recognition algorithm, the dominant color tone analyzed by the color histogram, the spatial coordinates determined by the three-dimensional positioning, etc. The comparison process calculates a matching degree score for each visual candidate target. The calculation of the score can use a weighted summation model, the formula is: ; wherein, represents the final generated matching degree score. represents the key entity information extracted from the speech instruction. represents the single candidate target attribute obtained from the visual candidate information. , , are functions for calculating the matching degree of category, attribute, and spatial relationship, respectively, and the output is a normalized score. , , are preset weight coefficients for adjusting the importance of different types of information in the total score, and these weights are obtained by experimental tuning according to specific application scenarios and task types. For example, score by comparing the nouns in the speech instruction with the object labels recognized by vision; score by comparing the adjectives in the instruction with the attributes such as color, shape, etc. detected by vision.
[0038] For example, assuming that the complete instruction is "take away the red cup on the table", the key entity information extracted by the robot is "category: cup", "attribute: red", and "space: on the table". The existing visual candidate information is a red cup, the category matching degree , the attribute matching degree , is a red box, the category matching degree , the attribute matching degree . If the weight setting is , then the matching degree score of is ; the matching degree score of .
[0039] when the matching score is lower than a threshold value for judging ambiguity, triggering an active clarification question-answering process; In some implementations, the highest matching score calculated is compared with a preset threshold value for judging ambiguity. If the highest matching score is lower than the threshold value, or there are multiple candidate targets with matching scores very close to each other and all higher than the threshold value, it is determined that there is ambiguity and the user's intention cannot be uniquely determined. At this time, an active clarification question-answering process is triggered to initiate a clarification request to the user.
[0040] For example, assuming the complete instruction is "take that red thing", since the instruction does not specify the category, the category weight is 0, and the weights are concentrated on the color and spatial attributes. After calculation, the matching scores of the red cup and the red box are both higher than the absolute threshold value 0.9, but the difference between them is less than the preset ambiguity judgment threshold value 0.01. It is determined that the target cannot be uniquely determined, and an active clarification question-answering process is immediately triggered to ask the user: "There are two red things on the table: a cup and a box. Which one do you mean?"
[0041] updating the visual candidate information based on the user feedback obtained from the active clarification question-answering process, to generate an updated verification result.
[0042] In some implementations, in the active clarification question-answering process, additional information is obtained according to the user's voice feedback. For example, the user may answer "the red one on the left". After this part of the user feedback is analyzed, it will be used as a new constraint condition to update the visual candidate information. The candidate targets that do not meet the new condition are eliminated, or the weights of the candidate targets that meet the new condition are greatly increased. The matching score calculation is performed again using the updated constraints until a unique target with a high matching score is selected. This final result confirmed by the user feedback constitutes the updated verification result, ensuring the accuracy of the task planning.
[0043] For example, the user answers: "take that red box". The feedback is immediately analyzed, and the new constraint condition "category: box" is added to the vector, and the visual candidate information is re-verified. At this time, the matching score of the red box greatly exceeds that of the red cup, and the red box is uniquely locked as the target object, generating an updated verification result, and the precise three-dimensional pose information of the target object is passed to the subsequent task planning step.
[0044] In one possible implementation, the active clarification question-answering process is combined Figure 2The above initiation of the active clarification question flow can be based on user feedback to generate a refined task planning, which can be described as follows: Based on the visual candidate information, a plurality of candidate target options are generated. In some implementations, when it is identified that the matching degrees of a plurality of potential target objects with the user instruction are all high, or none of the matching degrees of the targets exceeds the confidence threshold, a concise and recognizable descriptive option, i.e., a candidate target option, is generated for each potential target based on the current visual candidate information. When generating these options, the most differentiated features between the candidate targets are intelligently extracted. For example, if there are two cups in the scene, one red and one blue, the options are generated based on the color; if the two cups are the same color, the relative spatial positions of the two cups are used to generate options such as "the cup on your left" or "the cup close to the window".
[0045] For example, after performing the cross-modal verification, it is determined that there is ambiguity, and the visual candidate information includes two highly similar red targets: a red cup A on the left side of the table and a red box B on the right side of the table. The class and relative position are intelligently selected as the differentiated features to generate two candidate target options: "the cup close to the left" and "the box close to the right".
[0046] A clarification question containing the candidate target options is proposed to the user through a voice synthesis function, user voice feedback is received and analyzed, and a user selection intention is obtained; In some implementations, a built-in voice synthesis function is called to organize the candidate target options into a clarification question that conforms to the logic of natural language. For example, a voice "two cups are seen, do you mean the red one or the blue one?" can be synthesized and played. This question is clearly conveyed to the user through the speaker of the robot, guiding the user to make a selection. Then, the robot enters a listening state and receives the user's voice feedback through the microphone array. The user's answer, such as "the red one", is captured in real time and transmitted to the voice recognition and natural language understanding for analysis. The selection intention of the user is accurately extracted from the user's answer. This is not just a simple keyword matching, but also includes an understanding of pronouns, which requires combining the context of the question to accurately map the user's ambiguous answer to one of the previously generated candidate target options.
[0047] For example, a voice synthesis function is called to play a clarification question to the user: "There are two red targets on the table, do you mean the cup close to the left or the box close to the right?" The user's voice feedback is "I want the box." The keyword "box" is extracted from the voice feedback, and the selection intention of the user is determined to be "the box close to the right" based on the context of the question.
[0048] According to the user selection intention, the preliminary task assumption model is optimized to generate a refined task plan.
[0049] In some implementations, once the user's selection intention is successfully resolved, the key information that disambiguates the intention is obtained. This information is used to optimize the initial preliminary task assumption model. The unique target selected by the user and all its associated precise information is solidified into the task model, replacing the original ambiguous, multi-possibility target description. After this step of optimization, the originally uncertain preliminary task assumption model evolves into a refined task plan with clear instructions and a unique target. This plan contains all the precise parameters required to execute the task.
[0050] For example, the user selection intention "box near the right side" is bound to its precise coordinates B to optimize the initial preliminary task assumption model {intention: pick up, attribute: color-red, area: tabletop} to generate a refined task plan {intention: pick up, target ID: red box, target category: box, precise pose: B coordinates}, which contains all the precise parameters required to execute the task and can directly control the motion execution mechanism to start action.
[0051] In one possible implementation, the refined task plan is obtained by combining Figure 2 The above receiving the refined task plan and controlling the robot to execute the task, while acquiring multi-view images through the motion pattern for enhancing recognition and fusing the multi-view images can be specifically explained as follows: Monitoring the confidence parameter of the perception of the target object by AI visual recognition, and starting the motion pattern for enhancing recognition when the confidence parameter is lower than the threshold for triggering enhanced perception; In some implementations, after receiving the refined task plan, the process of executing the task and enhancing perception is performed synchronously, aiming to solve the problem of perception uncertainty caused by changes in viewing angle, occlusion, or unclear features of the object itself when the robot approaches or operates the target. This process is started simultaneously when the robot starts physical motion according to the refined task plan. During the process of the robot moving towards the target object to execute the task, the perception of AI visual recognition continuously tracks the target and outputs key indicators in real time, i.e., the confidence parameter. This confidence parameter is a quantitative value representing the degree of confidence that the visual algorithm is detecting the object in the current frame of image as the target object. When the target is partially occluded or its visual features become blurred due to poor lighting or viewing angle, the confidence parameter will decrease. Once it is detected that the confidence parameter is lower than the threshold for triggering enhanced perception, the motion pattern for enhancing recognition is started. This is not a suspension of the main task, but a fine-tuned, perception-oriented sub-motion superimposed on the main task path.
[0052] For example, the robot is performing a task of "grabbing the blue toolbox". When the robot arm approaches the target, the shadow of the robot arm itself blocks the side of the toolbox, causing the target class confidence output by the AI vision recognition algorithm to rapidly drop from the normal 0.92 to 0.78; since 0.78 is lower than the preset enhanced perception threshold 0.80, the motion control immediately receives the instruction and starts the motion mode for enhanced recognition while not stopping the main task of grabbing.
[0053] Performing an arc motion around the target or a vertical Z-shaped lifting, collecting multi-view images, and performing three-dimensional point cloud fusion processing on the multi-view images to generate enhanced perception data.
[0054] In some implementations, after starting the mode, the motion control performs a preset specific trajectory designed to observe the target from multiple angles. For example, performing an arc motion around the target, that is, controlling the robot end or mobile chassis to move along a circular path while keeping the distance to the target substantially unchanged, so as to collect side images of the target object. Another mode is vertical Z-shaped lifting, that is, controlling the robot arm to drive the camera to perform Z-shaped trajectory motion of first lifting, then translating, and finally descending, so as to obtain images of the object at different heights and horizontal offsets. While performing these specific motion trajectories, the vision sensor of the robot continuously collects a series of multi-view images. Subsequently, three-dimensional point cloud fusion processing is performed on the collected multi-view images. This processing process constructs local and scattered three-dimensional point clouds using the depth information generated by each frame of image, and then combines the motion trajectory data recorded by the robot to accurately align and splice these local point clouds from different angles into a more complete and dense whole. This fusion process finally generates enhanced perception data that is more rich and accurate, which contains a more comprehensive description of the three-dimensional shape of the target object.
[0055] For example, after monitoring the confidence drop, the motion control system immediately drives the camera of the robot to perform a vertical Z-shaped lifting trajectory, the trajectory height changes , the horizontal offset , and 15 frames of RGB-D images are continuously collected within 3 seconds; subsequently, a three-dimensional point cloud fusion processing algorithm combines the depth information of the 15 frames of images and the pose change data of the robot to accurately register and splice the scattered point cloud data, and finally generates a more complete toolbox dense three-dimensional model than the initial perception, as enhanced perception data, which improves the recognition accuracy of the edges and textures of the target object. Figure 3As shown, the confidence parameter of the AI visual recognition on the target object during the task execution of the robot is displayed in real time, and a threshold line for triggering the enhanced recognition motion mode is marked. The figure clearly shows how to immediately start the enhanced recognition motion mode when the confidence parameter drops below the threshold due to environmental or perspective factors, and actively collect multi-perspective information.
[0056] In a possible implementation manner, in combination with Figure 2 The three-dimensional point cloud fusion processing on the multi-perspective images can be specifically described as follows. Feature points are extracted and matched from the multi-perspective images to generate a feature point set. In some implementation manners, feature points are extracted and matched from the acquired multi-perspective image sequence. Efficient feature point detection and description algorithms such as ORB (Oriented FAST and Rotated BRIEF) are used to identify corner points, spot points, and other local regions with stable recognition degrees in each image as feature points. By comparing the descriptors of these feature points, corresponding relationships between different images are found, that is, the projections of the same three-dimensional space point in different perspective images are identified. All these successfully matched feature point pairs are collected to form a feature point set containing cross-image corresponding relationships.
[0057] For example, the robot collects 10 frames of images around the blue toolbox 10. The ORB algorithm detects 500 feature points of the toolbox edge on the first frame of image and 480 feature points on the second frame of image. After descriptor matching and RANSAC filtering, 350 pairs of reliable cross-frame corresponding feature points are successfully found, which constitute the feature point set for the next step of processing.
[0058] The feature point set is processed by using a motion recovery structure algorithm to generate a sparse point cloud. In some implementation manners, the feature point set is processed by using a motion recovery structure (SfM) algorithm. The motion recovery structure algorithm is a technology that can estimate three-dimensional structure and camera motion simultaneously. With the feature point set as input, iterative optimization is performed through multi-view geometric constraints to solve two key outputs: one is the three-dimensional space coordinates of each feature point, and these three-dimensional points collectively constitute a sparse point cloud; the other is the accurate pose of the camera when each image is taken. The sparse point cloud provides the basic skeletal structure of the target object, and the accurate camera pose is the basis for subsequent densification reconstruction.
[0059] Exemplarily, 350 pairs of matched feature points are input into the SfM algorithm, and after optimization processes such as bundle adjustment, the algorithm calculates the three-dimensional coordinates of the 350 feature points in the robot coordinate system, generates a sparse point cloud of the edge contour of the tool box, and accurately determines the 10 precise poses and motion trajectories of the camera in space when collecting the 10 images.
[0060] The sparse point cloud is densified by applying a multi-view stereo algorithm to generate a dense three-dimensional model. In some implementations, after obtaining the sparse point cloud and the camera poses, a multi-view stereo (MVS) algorithm is applied to densify the sparse point cloud. The MVS algorithm uses the known camera poses to estimate the depth of non-feature point regions in the images. For each pixel in the reference image, the best match is searched on the corresponding epipolar line of other view images, and the depth information of the pixel is calculated by triangulation principle. By repeating this process for a large number of pixels, the dense three-dimensional points of the object surface can be recovered, and the sparse skeleton can be filled into a dense three-dimensional model containing rich surface details, also known as a dense point cloud.
[0061] Exemplarily, using the MVS algorithm, combined with the calculated 10 camera poses, the depth of all non-feature point pixels in the image is estimated and geometric consistency is checked. Finally, the sparse point cloud is filled into a tool box dense three-dimensional model containing more than 20,000 three-dimensional points, which not only has the contour of the tool box, but also contains the texture and concave-convex details of its surface.
[0062] Based on the dense three-dimensional model, the position and attitude information of the target object is updated to generate updated augmented perception data.
[0063] In some implementations, based on the generated dense three-dimensional model, the final target information update is performed. By geometric analysis of this dense point cloud data, such as calculating its centroid, bounding box or principal axis direction, the three-dimensional position and attitude information of the target object in the robot coordinate system can be extremely accurately calculated. This high-precision position and attitude information, together with the dense three-dimensional model itself, is encapsulated into updated augmented perception data and replaces the previous perception result with low confidence, providing final and highly reliable data support for subsequent precise operation of the robot.
[0064] Exemplarily, the model composed of 20,000 three-dimensional points is geometrically analyzed to calculate the three-dimensional centroid of the tool box as the precise position and to calculate the orientation of the minimum circumscribed rectangle bounding box as the attitude information. The updated augmented perception data indicates that the precise 3D pose of the tool box is , which replaces the initial low-confidence data and can be directly used for the planning of the robot arm grasping.
[0065] In one possible implementation, combined withFigure 2 The above space anchoring based on the enhanced perception data and combined with the pose change data of the robot itself can be specifically described as follows: Real-time pose change data of inertial measurement output is obtained, and a kinematics model describing the relationship between the joints and the end of the robot is combined to calculate the pose transformation matrix of the robot itself; In some implementations, real-time pose change data of the robot itself is obtained. This part of data is fused from two sources. One is from the inertial measurement unit (IMU) inside the robot, which can output linear acceleration and angular velocity information of the robot body at a high frequency, and the pose change of the robot base can be estimated in real time through integration operation. The second is the kinematics model describing the relationship between the joints and the end of the robot, also known as forward kinematics. By reading the real-time angle values of each joint encoder and substituting them into the model, the pose of the robot end with the camera mounted relative to the robot base can be accurately calculated. Combined with the two, the accurate transformation matrix representing the current pose of the camera in the robot base coordinate system can be calculated.
[0066] For example, during the movement of the robot to the target tool box, the IMU sensor of the chassis reports a slight vibration within 0.1 seconds, and the linear acceleration At the same time, the multi-axis attitude adjustment component encoder of the robot reports that the joint angle of the end effector where the camera is located is fine-tuned by 0.2 degrees. Immediately substitute the IMU data and the 0.2-degree joint angle into the kinematics model describing the relationship between the joints and the end of the robot to calculate the total pose transformation matrix of the camera coordinate system relative to the robot base coordinate system The total pose transformation matrix is: ; This matrix accurately reflects the real-time position and attitude of the camera at the current time , providing basic data for subsequent space anchoring.
[0067] The pose transformation matrix is synchronized to the perception of AI visual recognition to correct the target coordinate drift and output stable coordinate information.
[0068] In some implementations, the pose transformation matrix is used to implement the core calculation of space anchoring. The stable coordinate information of the target object in the robot base coordinate system is represented as At any time , the stable coordinate is calculated and maintained by the following formula: ; Where, is the perception of AI visual recognition at The time-varying output is the coordinate of the target in the current camera coordinate system, which is a momentary observation value directly obtained from the augmented perception data or through subsequent tracking. is a pose transformation matrix calculated by fusing the inertial measurement and the kinematic model, representing the conversion relationship from the camera coordinate system to the robot base coordinate system. Through the execution operation, the unstable coordinate of the target object relative to the camera is converted in real time to the stable coordinate relative to the robot base . Since the robot base is usually fixed during the task , it constitutes a constant position anchor in the robot world view. This calculated pose transformation matrix is synchronized in real time to the perception system of the AI visual recognition to correct and predict the position where the target should appear in the next frame of image, thereby effectively suppressing the target coordinate drift, and finally outputting continuously updated stable coordinate information.
[0069] For example, assuming that at time , the augmented perception data reports the center coordinate of the toolbox in the camera coordinate system is . Using the pose transformation matrix , the following matrix multiplication is performed: ; The calculation result is the stable and accurate spatial anchor coordinate of the toolbox in the robot base coordinate system. This stable coordinate information is used for the final task execution planning, ensuring that even if the robot itself has slight pose changes, the grasping coordinate of the target is still accurate.
[0070] In one possible implementation, in combination with Figure 2 , the above-mentioned acquisition of real-time pose change data of inertial measurement output, in combination with the kinematic model describing the relationship between the robot joints and the end, can be specifically described as follows: Through the execution of the hand-eye calibration process, a plurality of camera pose transformation data and robot end pose transformation data are collected; In some implementations, this is achieved by directing the robot to execute a pre-set sequence of calibration actions. During this process, the robot moves its end effector to a series of different positions and poses. Simultaneously, a camera mounted on the robot's end effector or arm continuously observes a calibration object with known geometric dimensions, such as a checkerboard or dot array calibration board, fixed in the scene. After the robot moves to a new calibration pose, two sets of key data are collected synchronously. The first set is the robot's end effector pose transformation data, directly provided by the robot controller. Based on the joint encoder readings and the forward kinematics model, the controller accurately calculates the pose transformation of the end effector from the previous pose to the current pose; this data constitutes the robot's end effector pose transformation matrix. The second set is the camera pose transformation data, calculated by the AI vision system. The vision system analyzes images of the calibration board taken by the camera at two different positions to calculate the pose transformation of the calibration board relative to the camera coordinate system; this constitutes the camera pose transformation matrix.
[0071] For example, the robot is programmed to execute 20 different sequences of "move-pause" actions. In the... During this movement, the robot's motion controller records the position of its end effector. to posture pose transformation matrix Simultaneously, the images of the calibration board captured by the camera are processed and their attitude is calculated to determine the distance the calibration board travels from the camera. Coordinate system to camera Relative pose transformation matrix of coordinate system Ultimately, 20 sets were obtained. and The corresponding data pairs.
[0072] Construct hand-eye calibration equations, where camera pose transformation is used as a matrix. The robot's end-effector pose transformation is a matrix ; In some implementations, multiple sets, typically dozens, of collected data are used to construct and solve the hand-eye calibration equation. The classic form of this equation is: ; in, This represents the camera pose transformation matrix, which is the relative motion of the calibration board in the camera coordinate system. This matrix is calculated by a vision algorithm from two consecutive frames of images. This represents the robot end-effector pose transformation matrix, which is the relative motion of the robot end-effector in the base coordinate system. This matrix is directly provided by the robot motion controller. The unknowns to be solved are the fixed transformation matrices. , which represents the constant transformation relationship from the robot end coordinate system to the camera coordinate system, is the final goal of the calibration process.
[0073] For example, assume that in a certain movement, the robot controller records the end transformation matrix and the vision system calculates the camera transformation matrix as follows: ; ; A set of 20 matrix equations is constructed , and the goal is to solve a unique matrix.
[0074] The fixed transformation matrix is solved by an algorithm for solving overdetermined equations, and the fixed transformation matrix is used to establish an accurate spatial mapping relationship, realizing coordinate conversion from the camera coordinate system to the robot coordinate system.
[0075] In some implementations, since multiple sets of data are collected, the equations form an overdetermined equation set. A mature algorithm for solving overdetermined equations, such as the Tsai-Lenz method or the separation method, is used to calculate the optimal fixed transformation matrix . This solving process can effectively smooth the noise and errors of single measurement, obtaining high-precision results. Once the fixed transformation matrix is solved, it is stored as a firm geometric binding relationship between the camera and the robot body. From now on, an accurate spatial mapping relationship is established, and coordinate conversion from the camera coordinate system to the robot coordinate system can be accurately completed at any time by reading the real-time pose of the robot end and combining the matrix . As shown in Figure 4 , the residual errors of the original coordinate drift and the stable coordinate after optimization and solving of the overdetermined equation set in the hand-eye calibration process are compared. The results of the figure directly show that solving the fixed transformation matrix by SVD and other algorithms can effectively reduce the average coordinate residual error, thereby realizing high-precision spatial anchoring.
[0076] For example, the singular value decomposition (SVD) method is used to optimize and solve the 20 equations, and after iterative convergence, the fixed camera-end transformation matrix is obtained: ; This matrix contains the fixed translation and rotation relationship of the camera relative to the robot end effector. In subsequent tasks, by reading the real-time pose of the robot end i.e. can be used matrix to calculate the real-time pose of the camera in the base coordinate system , which realizes accurate spatial mapping from the camera coordinate system to the robot base coordinate system, and the accuracy can reach sub-millimeter level.
[0077] In one possible implementation manner, in combination with Figure 2 The above generation of situational feedback information using the stable coordinate information and the task execution state can be specifically described as follows: Integrating visual context information from the augmented perception data, and fusing motion history records from the task execution state; In some implementation manners, the visual context information integrated from the augmented perception data. This includes the high-definition image of the finally confirmed target object or its three-dimensional reconstruction model, the snapshot of the environment around the target object, and the visual record of the key nodes in the process of executing the task. For example, the state before grasping, and the state of the arm carrying the target after successful grasping. The motion history records fused from the task execution state. This includes the complete motion trajectory data of the robot from receiving the instruction to the end of the task, the motion log of each joint, and the key events encountered in the process of task execution, such as triggering the motion pattern for enhancing recognition, or performing obstacle avoidance action, etc.
[0078] For example, after completing the “grasp the toolbox” task, the integrated data includes: 1) visual context information: three-dimensional dense model of the toolbox, high-definition snapshot of the toolbox and the robotic arm when grasping is successful; 2) motion history records: the motion log shows that during the approach to the target, the confidence level once decreased and triggered the enhanced perception motion pattern of vertical characteristic lifting, and the end effector successfully positioned at the stable coordinate .
[0079] Using a large language model for semantic generation, the visual context information and the motion history records are processed into natural language descriptions; In some implementation manners, a large language model for semantic generation is called, and the integrated visual context information and motion history records are taken as input and processed into natural language descriptions that are easy to understand. The large language model is specially fine-tuned to understand and describe the behavior and perception of the robot. It can convert dry coordinate data and joint angles into vivid action descriptions, for example, describing “performing an arc motion around the target” as “turning around it for half a circle to get a better view”. It can also convert visual data into scene descriptions, such as “picked up the red cup on the table, and there is a book next to it”.
[0080] For example, the large language model receives “the toolbox has been grasped, and the enhanced perception motion pattern of vertical After structured data such as "character motion" and the like, through its contextual perception ability and semantic generation ability, a natural feedback description is output: "The blue toolbox has been successfully picked up. When approaching it, a slight up-down movement was performed to ensure a clearer view due to the dim light, and finally the precise position was reached."
[0081] The natural language description is combined with the generated image content to form situational feedback information.
[0082] In some implementations, while generating the natural language description, the generated image content technology is also used in parallel to create visual materials that match the feedback content. This is not simply a reproduction of the original image, but can be artistically processed or highlighted as needed. For example, the target object being operated can be highlighted or marked with an arrow on a scene picture, or a simplified animation can be generated to demonstrate the key action path of the robot. The generated natural language description and generated image content are combined to form a multi-modal, information-rich situational feedback information. This information can be output to the user in various forms, such as through the robot's screen displaying the report with text and images, and through the speech synthesis function playing the natural language description part. What the user sees is not just the result, but a lively review of the entire task process, understanding how the robot understands the instructions, how it observes the environment, how it overcomes difficulties and finally completes the task.
[0083] For example, the natural language description "The blue toolbox has been successfully picked up..." is sent to the voice broadcast. At the same time, the generated image content function generates a highlighted annotation image based on the snapshot of the grab: a green box is drawn around the grabbed toolbox on the original image, and a dotted arrow is used to mark the trajectory of the character motion, and finally the annotated image is displayed on the screen to form complete situational feedback information for the user to understand and verify.
[0084] It should be noted that the electrical connection between the above-mentioned units does not necessarily mean direct connection of the line, indirect connection mode, as long as the purpose of the application is achieved, which can be applied to the embodiments of the application. The above-described is only an exemplary embodiment of the application, which cannot limit the scope of the application.
[0085] intended to encompass any and all embodiments of the application with equivalents as would be ascertained by those skilled in the art to which the application pertains. Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit of the application being indicated by the following claims.
Claims
1. A question-and-answer interaction method for robot audiovisual fusion, characterized in that, The method includes: The system acquires user voice commands and performs real-time streaming recognition and conversion to generate preliminary recognized text. Based on the initially identified text, it is processed by a preset model for inferring user intent to generate a preliminary task hypothesis model; Receive the preliminary task hypothesis model and dynamically adjust the perceptual resources of AI visual recognition to generate visual candidate information; By combining the fully recognized voice commands with the visual candidate information, cross-modal verification is performed to generate verification results; If the verification result indicates ambiguity, an active clarification question-and-answer process is initiated, and a refined task plan is generated based on user feedback. If not, but the indicated perspective is poor, the robot is instructed to move to change the perspective, and a refined task plan is generated based on the changed perspective information. The robot receives the refined task plan and controls the robot's movement to execute the task. At the same time, it acquires multi-view images by executing motion patterns for enhanced recognition and fuses the multi-view images to generate enhanced perception data. Based on the enhanced perception data and combined with the robot's own pose change data, spatial anchoring processing is performed to generate stable coordinate information; Using the stable coordinate information and task execution status, contextualized feedback information is generated and output to the user.
2. The question-and-answer interaction method for robot audiovisual fusion according to claim 1, characterized in that, The dynamically adjusted perceptual resources for AI visual recognition include: Based on the preliminary task hypothesis model, the hypothetical target area is determined, and the robot gimbal or body is controlled to face the hypothetical target area to achieve focused view. The image recognition algorithm associated with the preliminary task hypothesis model is enabled to achieve algorithm focus; Based on the results of the perspective focusing and algorithm focusing, visual candidate information is generated.
3. The question-and-answer interaction method for robot audiovisual fusion according to claim 1, characterized in that, The cross-modal verification, which combines the fully recognized voice command with the visual candidate information, includes: Key entity information is extracted from the fully recognized voice command, and the key entity information is compared with the target attributes in the visual candidate information to generate a matching score. When the matching score is lower than the threshold used to determine ambiguity, an active clarification question-and-answer process is triggered. The visual candidate information is updated based on user feedback obtained from the proactive clarification question-and-answer process, and an updated verification result is generated.
4. The question-and-answer interaction method for robot audiovisual fusion according to claim 1, characterized in that, The process of initiating a proactive clarification-based question and answer session, which generates a detailed task plan based on user feedback, includes: Based on the visual candidate information, several candidate target options are generated; The system uses speech synthesis to ask the user a clarifying question containing the candidate options, receives and parses the user's voice feedback, and obtains the user's selection intent. Based on the user's selected intent, the initial task hypothesis model is optimized to generate a refined task plan.
5. The question-and-answer interaction method for robot audiovisual fusion according to claim 1, characterized in that, The process includes receiving the refined task plan, controlling the robot's movement to execute the task, acquiring multi-view images by executing motion patterns for enhanced recognition, and fusing the multi-view images, including: Monitor the confidence parameter of AI visual recognition on the target object, and activate the motion mode for enhanced recognition when the confidence parameter is lower than the threshold used to trigger enhanced perception. Perform arc-shaped motion around the target or vertical Z-shaped ascent and descent to acquire multi-view images, and perform 3D point cloud fusion processing on the multi-view images to generate enhanced perception data.
6. The question-and-answer interaction method for robot audiovisual fusion according to claim 5, characterized in that, The three-dimensional point cloud fusion processing of multi-view images includes: Feature points are extracted and matched from the multi-view images to generate a feature point set; The feature point set is processed using the structure-reconstruction-motion algorithm to generate a sparse point cloud; A multi-view stereo algorithm is applied to densify the sparse point cloud to generate a dense 3D model. Based on the dense 3D model, the position and orientation information of the target object are updated to generate updated enhanced perception data.
7. The question-and-answer interaction method for robot audiovisual fusion according to claim 1, characterized in that, Based on the enhanced perception data, and combined with the robot's own pose change data, spatial anchoring processing includes: The robot acquires real-time pose change data from inertial measurement output, and calculates its own pose transformation matrix by combining it with a kinematic model describing the relationship between the robot's joints and end effectors. The pose transformation matrix is synchronized to the perception of AI visual recognition to correct target coordinate drift and output stable coordinate information.
8. The question-and-answer interaction method for robot audiovisual fusion according to claim 7, characterized in that, The acquisition of real-time pose change data from inertial measurement output, combined with a kinematic model describing the relationship between the robot joints and the end effector, includes: By performing the hand-eye calibration process, multiple sets of camera pose transformation data and robot end-effector pose transformation data are collected. Construct hand-eye calibration equations, where the camera pose transformation is a matrix and the robot end effector pose transformation is a matrix. The fixed transformation matrix is solved by an algorithm for solving overdetermined equations, and then a precise spatial mapping relationship is established using the fixed transformation matrix to realize the coordinate transformation from the camera coordinate system to the robot coordinate system.
9. The question-and-answer interaction method for robot audiovisual fusion according to claim 1, characterized in that, Using the stable coordinate information and task execution status, contextualized feedback information is generated, including: Visual context information is integrated from the enhanced perception data, and motion history is fused from the task execution state; Using a large language model for semantic generation, the visual context information and the motion history are processed into a natural language description; The natural language description and generative image content are combined to form contextualized feedback information.
10. A question-and-answer interactive system for robot audiovisual fusion, characterized in that, The system is used in a question-and-answer interaction method for robot audiovisual fusion as described in any one of claims 1-9, the system comprising: The voice command processing module is used to acquire user voice commands, perform real-time streaming recognition and conversion, and generate preliminary recognized text. The intent inference module is used to process the pre-defined user intent model based on the pre-identified text to generate a preliminary task hypothesis model. The image perception and adjustment module is used to receive the preliminary task hypothesis model and dynamically adjust the perception resources of AI visual recognition to generate visual candidate information. The cross-modal verification module is used to combine the fully recognized voice command with the visual candidate information to perform cross-modal verification and generate verification results; The task planning and clarification module is used to determine whether there is any ambiguity in the verification result indication. If yes, it initiates an active clarification question-and-answer process and generates a refined task plan based on user feedback. If no, but the display view is not good, it instructs the robot to move to change the view and generates a refined task plan based on the changed view information. The motion control and enhanced perception module is used to receive the refined task plan and control the robot to perform the task. At the same time, it acquires multi-view images by executing motion patterns for enhanced recognition, and fuses the multi-view images to generate enhanced perception data. The spatial anchoring module is used to perform spatial anchoring processing based on the enhanced perception data and combined with the robot's own pose change data to generate stable coordinate information; The feedback generation and information output module is used to generate contextualized feedback information using the stable coordinate information and task execution status, and output the contextualized feedback information to the user.
Citation Information
Patent Citations
Interaction method and device for upper limb cooperative control of humanoid robot
CN119927906A
Indoor mobile service robot interaction task execution method and device, storage medium and indoor mobile service robot system
CN119772883A
Power grid monitoring bionic robot with multi-mode sensing and voice interaction functions
CN120412583A
Humanoid robot multi-mode instruction analysis system
CN120516701A
Posture recognition algorithm for any object under monocular camera and application system
CN120707630A