Intelligent companion robot and interaction method thereof
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-18
- Publication Date
- 2026-08-11
AI Technical Summary
对于激光电视,也仅仅是围绕视觉体验,音响体验不佳,没有AI交互
[0014]经由上述的技术方案可知,与现有技术相比,本发明公开提供了一种智能陪伴机器人及其交互方法,基于市场客户需求及当前行业产品现状,推出了一体化的智能陪伴机器人产品。该机器人产品,具有极为出色的视觉和声音呈现,满足家庭影院的影音能力,同时具备AI大模型核心驱动的智慧家庭空间Agent,融合多模态交互(语音+视觉+触控+投影)与AI大模型,又集成移动底盘与传感器,实现自主导航、避障及家庭场景全适配,提升人机交互自然度支撑家庭陪伴、智慧安防、智能家居等未来智慧化家庭场景业务。此机器人带有AI处理单元,搭载多模态大模型,实现语音语义解析、情绪识别及具身智能决策,联动智能家居控制模块。基于AI大模型的核心基础能力,面向智慧家庭的空间智能Agent是多维空间智能体。通过获取感知家庭空间的多维空间因素,在家庭空间大模型的计算和推理下,通过空间具身智能体实现意图识别解析和决策与执行。
Smart Images

Figure CN122546733A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robotics, and more specifically to an intelligent companion robot and its interaction method. Background Technology
[0002] With the development of various artificial intelligence, Internet of Things and robotics technologies, and the accompanying new demands for smart home living scenarios, there is a greater need for smart products and services in the home environment.
[0003] In response to this situation, the industry has seen the emergence of many new products designed for specific scenarios, such as portable projectors, chatbots, large-screen laser TVs, and smart speakers. However, these products are essentially focused on a single, specific scenario and are largely disconnected from each other. Portable projectors only offer projection-related functions, failing to provide a comparable audio experience or AI interaction. Chatbots merely utilize language and voice interaction capabilities based on large AI models, with sluggish and unresponsive visual interaction. Laser TVs also prioritize visual experience, offering poor audio quality and lacking AI interaction. Smart speakers or traditional speakers either focus solely on large AI model interaction or solely on sound quality, lacking visual interaction or sufficient intelligent capabilities.
[0004] Therefore, how to provide an intelligent companion robot that integrates multimodal interaction and AI large-scale model and its interaction method is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, the present invention provides an intelligent companion robot and its interaction method, which solves the problems existing in the background art.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: On one hand, this invention discloses an intelligent companion robot, comprising: a perception module, an execution module, a control module, and a smart home space intelligent agent; the perception module is used to acquire user instructions or multi-dimensional multimodal environmental data; the control module receives the processing results of the smart home space intelligent agent and sends control instructions to the execution module; the execution module receives the control instructions to drive the robot's actuators or control smart home devices; the smart home space intelligent agent, based on an AI big model, acquires multi-dimensional spatial factors of the perceived home space, and under the calculation and reasoning of the home space big model, realizes intent recognition, parsing, decision-making, and execution in different working modes.
[0007] Preferably, in the above-mentioned intelligent companion robot, the sensing module includes: Microphone module, microphone array, long-distance sound pickup, sound source localization and directional sound pickup, to obtain user voice commands; The camera module is used for image and video capture. Touchscreen, allowing users to issue commands via touch; Multimodal sensors collect multimodal environmental data.
[0008] Preferably, in the above-mentioned intelligent companion robot, the execution module includes: The speaker module, combined with AI voice synthesis capabilities, enables personalized and ultimate sound output; The projection module automatically focuses and corrects, projecting onto a giant screen. The smart home module responds to control commands and executes operations.
[0009] Preferably, in the above-mentioned intelligent companion robot, different working modes include, but are not limited to: The proactive home theater mode recognizes user commands, understands user intentions, and has the ability to remember the needs of members in a smart home. It searches online resources, formulates the best viewing plan, dims the lights, draws the curtains, and plays 4K content; it also coordinates various modules to create an immersive environment. Multimodal teaching support mode: The perception module captures images and videos; the visual model recognizes image content; the large language model generates a story about the image and suggests projecting it; and the educational interaction mode is activated. The environmental adaptive butler mode uses online weather alerts combined with temperature and humidity sensors to determine the indoor temperature of the home and take tiered safety actions. In seamless video call mode, the user says "Make a video call to XX" and walks to the sofa; the system understands the command, recognizes the user's location, adjusts the camera and projection angles for the best composition and viewing experience, and initiates the video call. Personalized wake-up experience mode: The large language model generates wake-up plans based on the schedule and user preferences, while checking weather and traffic information to provide a one-stop morning service.
[0010] Preferably, in the above-mentioned intelligent companion robot, the control module receives the reasoning and decision-making results of the intelligent agent in the smart home space, queries the status of each execution module, and performs task scheduling, behavior planning, and escapes and sends execution instructions to each execution module to complete the issuance of control instructions.
[0011] Preferably, in the aforementioned intelligent companion robot, the smart home space intelligent agent receives human-computer interaction or multi-dimensional multimodal environmental data input from the perception module. Through natural language processing (NLP) or large language model (LLM), it performs speech, semantic, or emotion analysis and intent recognition on the input information. Combining the short-term context and long-term memory retrieval capabilities of the space intelligent agent, it performs computational reasoning and task decomposition planning based on the AI Agent architecture through a local large model or an AI cloud large model, and decides whether to call external tools or sub-intelligent agents. Then, based on the returned results, it makes intelligent decisions and executes actions, and finally outputs feedback to the user and updates the status according to the environmental response.
[0012] Preferably, the aforementioned intelligent companion robot further includes a chassis; the chassis is equipped with corresponding functional modules that can autonomously plan paths, dynamically avoid obstacles, and accurately dock at target points based on self-built maps and real-time positioning.
[0013] On the other hand, this invention discloses an interaction method for intelligent companion robots, applied to intelligent companion robots, with the following specific steps: Step 1: Command Input and Sensing. The system receives user commands, performs initial updates and optimizations of the commands through the sensing module, and ensures proper handling in case of no response or abnormality. Step Two: Core Thinking and Decision-Making. This involves task planning for the identified structured instructions, initiating a query, and generating an executable task sequence. Step 3: Instruction classification and scheduling. Based on the task content, instructions are classified into different control domains and tasks are scheduled according to priority. Step 4: Execution of multi-dimensional data, system call response, handling data storage and retrieval, media playback and user feedback respectively, to achieve collaborative execution of multi-dimensional data; Step 5: Complete data processing, organize the execution results, combine them with basic physical environment information, conduct multi-objective design and user experience evaluation to ensure that the system response meets the actual scenario and user expectations.
[0014] As can be seen from the above technical solution, compared with the prior art, this invention discloses an intelligent companion robot and its interaction method. Based on market customer needs and the current status of industry products, an integrated intelligent companion robot product has been launched. This robot product has excellent visual and sound presentation, meeting the audio-visual capabilities of a home theater. It also possesses a smart home space agent driven by an AI large-scale model, integrating multimodal interaction (voice + vision + touch + projection) with the AI large-scale model. Furthermore, it integrates a mobile chassis and sensors to achieve autonomous navigation, obstacle avoidance, and full adaptation to home scenarios, improving the naturalness of human-computer interaction and supporting future smart home scenarios such as family companionship, smart security, and smart homes. This robot has an AI processing unit, equipped with a multimodal large-scale model, to achieve voice semantic analysis, emotion recognition, and embodied intelligent decision-making, linking with the smart home control module. Based on the core capabilities of the AI large-scale model, the spatial intelligent agent for smart homes is a multi-dimensional spatial intelligent agent. By acquiring and perceiving multi-dimensional spatial factors of the home space, and under the calculation and reasoning of the home space large-scale model, the spatial embodied intelligent agent achieves intent recognition, analysis, decision-making, and execution. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0016] Figure 1 This is a structural schematic diagram provided for the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] This invention discloses an intelligent companion robot, comprising: a perception module, an execution module, a control module, and a smart home space intelligent agent; the perception module is used to acquire user commands or multi-dimensional, multi-modal environmental data; the control module receives the processing results of the smart home space intelligent agent and sends control commands to the execution module; the execution module receives the control commands to drive the robot's actuators or control smart home devices; the smart home space intelligent agent, based on an AI big data model, acquires multi-dimensional spatial factors of the perceived home space, and under the calculation and reasoning of the big data model, realizes intent recognition, parsing, decision-making, and execution in different working modes.
[0019] To further optimize the above technical solution, the sensing module includes: Microphone module, microphone array, long-distance sound pickup, sound source localization and directional sound pickup, to obtain user voice commands; The camera module is used for image and video capture. Touchscreen allows users to issue commands by touching the screen.
[0020] The execution module includes: The speaker module, combined with AI voice synthesis capabilities, enables personalized and ultimate sound output; The projection module automatically focuses and corrects, projecting onto a giant screen. The smart home module responds to control commands and executes operations.
[0021] Furthermore, the audio module features high-fidelity HiFi audio output capabilities, completely revolutionizing the current industry's audio experience and perception, achieving cinema-level sound effects, and combining AI voice synthesis capabilities to achieve personalized and ultimate sound output.
[0022] Projection Module: A customized high-brightness, high-definition laser projection module that surpasses ordinary TV screens in terms of visual brightness, clarity, contrast, and color saturation, while also offering greater flexibility and portability.
[0023] Touchscreen: A high-definition, highly sensitive multi-touch screen that not only displays content clearly but also enables convenient and easy-to-use touch management.
[0024] Camera Module: High-definition camera module, enabling high-definition image and video capture, better serving image and video-related business applications.
[0025] Microphone module: High-performance noise-canceling microphone array with long-distance sound pickup capability, as well as sound source localization and directional sound pickup capabilities, enabling better voice interaction.
[0026] Smart Home Module: The smart home module integrates the companion robot with the smart home system, allowing the companion robot to function as a mobile smart home control system.
[0027] To further optimize the above technical solution, different working modes include: The proactive home theater mode recognizes user commands, understands user intentions, and has the ability to remember the needs of members in a smart home. It searches online resources, formulates the best viewing plan, dims the lights, draws the curtains, and plays 4K content; it also coordinates various modules to create an immersive environment. Specifically, in active home theater mode, the modules work together as follows: 1. The microphone module receives commands.
[0028] 2. The intelligent agent in the smart home space understands and plans tasks.
[0029] 3. The smart home module turns off the living room lights and curtains.
[0030] 4. The projection module automatically focuses and corrects, projecting a giant screen.
[0031] 5. The audio module activates surround sound, and the AI generates a voice message: "It's ready for you. Enjoy the movie." Multimodal teaching support mode: The perception module captures images and videos; the visual model recognizes image content; the large language model generates a story about the image and suggests projecting it; and the educational interaction mode is activated. Specifically, in the multimodal teaching support mode, the modules collaborate as follows: 1. The camera module captures an image of the painting, identifies the painting in the hand, and asks what it is.
[0032] 2. The intelligent agent in the smart home space identifies content and generates stories.
[0033] 3. The audio module tells stories using vivid AI-synthesized voices.
[0034] 4. The projection module simultaneously projects the 3D model of the Stegosaurus and related information onto the wall for interactive teaching.
[0035] 5. The touchscreen displays a "Learn More" button.
[0036] The environmental adaptive butler mode uses online weather alerts combined with temperature and humidity sensors to determine the indoor temperature of the home and take tiered safety actions. Specifically, in the environment adaptive steward mode, module collaboration is as follows: 1. The smart home module detects abnormal sounds and triggers sensor alarms.
[0037] 2. The smart home space intelligent agent conducts risk assessment.
[0038] 3. The speaker module plays a reminder directed towards the kitchen: "Master, a red alert for high temperature has been issued. Do you need to draw the curtains and turn on the air conditioner?" 4. If the user does not respond, the Space Smart Agent will remotely draw the curtains via the smart home module.
[0039] 5. The screen displays a "Processed" security message.
[0040] In seamless video call mode, the user says "Make a video call to XX" and walks to the sofa; the system understands the command, recognizes the user's location, adjusts the camera and projection angles for the best composition and viewing experience, and initiates the video call. Specifically, in seamless video call mode, module collaboration is as follows: 1. The microphone module receives commands.
[0041] 2. The smart home space's intelligent agent makes phone calls and uses cameras to track the user's body, keeping the user centered in the frame.
[0042] 3. The projection module projects the other party's video onto the wall opposite the user, creating a "face-to-face" feeling.
[0043] 4. The audio module provides clear call quality, and the microphone array suppresses ambient noise to ensure the other party can hear clearly.
[0044] Personalized wake-up experience mode: The large language model generates wake-up plans based on the schedule and user preferences, while checking weather and traffic information to provide a one-stop morning service.
[0045] Specifically, in the personalized wake-up experience mode, the modules work together as follows: 1. Smart home space intelligent agent trigger wake-up process.
[0046] 2. The audio module gradually increases the volume, playing the user's favorite morning music and news summaries.
[0047] 3. The projection module projects the time, weather, today's schedule, and traffic route map onto the ceiling or wall.
[0048] 4. Simultaneously, the smart home module slowly opens the curtains and starts the coffee machine.
[0049] To further optimize the above technical solution, the control module receives the reasoning and decision-making results of the smart home space intelligent agent, queries the status of each execution module, and performs task scheduling, behavior planning, and escapes and sends execution instructions to each execution module to complete the issuance of control instructions.
[0050] To further optimize the above technical solution, the smart home space intelligent agent receives human-computer interaction or multi-dimensional multimodal environmental data input from the sensing module. Through natural language processing (NLP) or large language model (LLM), it performs speech, semantic, or emotion parsing and intent recognition on the input information. Combining the short-term context and long-term memory retrieval capabilities of the space intelligent agent, it performs computational reasoning and task decomposition planning based on the AI Agent architecture through a local large model or an AI cloud large model, and decides whether to call external tools or sub-intelligent agents. Then, based on the returned results, it makes intelligent decisions and executes actions, and finally outputs feedback to the user and updates the status according to the environmental response.
[0051] To further optimize the above technical solution, it also includes: a chassis; the chassis has a built-in self-built map and point unit, spatial positioning unit, path planning unit, adaptive navigation unit and automatic obstacle avoidance unit; according to the target point set by the user; the path planning unit's global planner generates a path from the current position to the target point; the controller drives the motor, and the robot moves along the path; the automatic obstacle avoidance unit, the lidar detects obstacles, and the local planner generates a detour trajectory; the robot stops precisely at the target point and completes the task.
[0052] Another embodiment of the present invention discloses an interaction method for an intelligent companion robot, applied to an intelligent companion robot, the specific steps of which are as follows: Step 1: Command Input and Sensing. Receive user commands, and perform initial updates and optimizations of the commands through the sensing module. At the same time, ensure appropriate handling in case of no response or abnormality. Step Two: Core Thinking and Decision-Making. This involves task planning for the identified structured instructions, initiating a query, and generating an executable task sequence. Step 3: Instruction classification and scheduling. Based on the task content, instructions are classified into different control domains and tasks are scheduled according to priority. Step 4: Execution of multi-dimensional data, system call response, handling data storage and retrieval, media playback and user feedback respectively, to achieve collaborative execution of multi-dimensional data; Step 5: Complete data processing, organize the execution results, combine them with basic physical environment information, conduct multi-objective design and user experience evaluation to ensure that the system response meets the actual scenario and user expectations.
[0053] Specifically, take voice commands as an example; Step 1: Command Input and Sensing (Data Acquisition and Preliminary Processing); This stage is the data entry point, and the goal is to convert the user's voice commands into structured information that the machine can understand.
[0054] Data acquisition (microphone module): Raw data: The microphone array continuously collects mid-frequency ambient sound. When a wake word (such as "Xiao X Xiao X") or a specific command pattern is detected, the user's statement "I want to watch Oppenheimer" is recorded, generating an audio stream (PCM encoded, etc.).
[0055] Data processing: Preprocessing such as noise reduction, echo cancellation, and voice separation is performed to improve the signal-to-noise ratio.
[0056] Speech recognition (ASR - Automatic Speech Recognition Module): Input: Processed audio stream data.
[0057] Processing: Use deep learning models (such as end-to-end ASR models) to map audio features to text.
[0058] Output: Generate raw text data – “I want to read Oppenheimer”.
[0059] Semantic understanding (NLU - Natural Language Understanding Module): Input: Raw text data.
[0060] Processing: Intent Recognition: Determines whether the user's intent is to start watching a movie or play a movie. Slot Filling: Extracts key entity information from the sentence, such as the movie title: Oppenheimer.
[0061] Output: Generates structured instruction data Step Two: Core Thinking and Decision Making (The Brain of the Spatial Intelligence Agent); This stage is crucial, where the spatial intelligence agent performs planning, decision-making, and resource retrieval based on the information it understands.
[0062] Task planning (Planner module): Input: Structured instruction data.
[0063] Processing: Based on the preset "best viewing plan" logic, the spatial intelligence agent decomposes the macro command "play movie" into a series of ordered atomic tasks.
[0064] Output: Generates an executable sequence of tasks, for example: Task 1: Adjust the ambient lighting to "Cinema Mode"; Task 2: Close the curtains; Task 3: Find the 4K version of the film "Oppenheimer"; Task 4: Turn on the projector and project the video; Task 5: Start the audio system and configure surround sound; Task 6: Play TTS voice feedback; Resource Query (Knowledge Base / Network module): Input: Tasks in the task sequence that require specific data (e.g., Task 3).
[0065] Processing: The agent accesses the local media library or the API of an authorized online streaming service, and queries using movie_name as the keyword.
[0066] Output: Obtain the unique identifier (such as URL, file path) and metadata (resolution, audio track format, etc.) of the source and add them to the task sequence (for example, inform task 4 of the specific resource address of task 3).
[0067] Step 3: Instruction Distribution and Dispatch (Command Center); This stage is responsible for distributing the task plan generated by the brain to the various execution units.
[0068] Input: The refined sequence of executable tasks.
[0069] Processing: The scheduler of the spatial intelligence agent sends control commands to the corresponding intelligent device modules through different communication protocols (such as HTTPREST, MQTT, Zigbee, Bluetooth, etc.) according to the task type.
[0070] Output: Generates a series of standard, device-recognizable control command data.
[0071] Step 4: Multi-module collaborative execution (physical world changes); Each module receives and executes instructions, producing actual effects. This stage is where the data flow triggers physical actions.
[0072] Smart home module: Input: The specific control command received (e.g., {"device":"main_light","action":"set_state","value":"10%"}).
[0073] Execution: The smart lighting controller adjusts the brightness, and the smart curtain motor closes the curtains.
[0074] Data flow: After execution is complete, a status confirmation message may be sent to the Agent.
[0075] Projection / playback module: Input: The received playback command (including the source address).
[0076] Execution: The projector starts up, performs automatic focus and keystone correction; the media player loads the video source.
[0077] Data stream: The module may provide status data such as "Player ready" or "Projector ready".
[0078] Audio module: Input: Received audio control commands (such as {"action":"set_mode","mode":"surround"} and TTS text data {"text":"Ready for you, enjoy the movie."}).
[0079] Execution: The audio processor configures the sound field mode; the TTS engine synthesizes the text into a speech audio stream and plays it.
[0080] Data stream: Audio data stream output.
[0081] Step 5: Experience Presentation and Closure (Final Output); Once all modules have been executed, the target state has been achieved.
[0082] Final output: An immersive viewing environment tailored specifically for Oppenheimer.
[0083] Data closed loop (optional but important): Environmental monitoring: Light sensors and cameras may continuously monitor ambient brightness, whether users have left the area, etc., to achieve more intelligent control (such as automatic pause when no one is present).
[0084] To further optimize the above technical solution, the following control steps for the chassis are also included: Mapping steps: The robot starts SLAM, autonomously navigates, and builds an environmental map; Location steps: Load map, match in real time with sensors, and determine your own location; Task assignment: User sets target point; Path planning: The global planner generates a path from the current location to the target point; Motion execution: The controller drives the motor, and the robot moves along the path; Real-time obstacle avoidance: When the lidar detects an obstacle, the local planner generates a detour trajectory; Reaching the target: The robot stops precisely at the target point, completing the task.
[0085] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0086] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An intelligent companion robot, characterized by, include: The system comprises a sensing module, an execution module, a control module, and a smart home space intelligent agent; the sensing module is used to acquire user commands or multi-dimensional, multi-modal environmental data. The control module receives the processing results of the smart home space intelligent agent and sends control commands to the execution module; the execution module receives the control commands to drive the robot body actuator or control the smart home devices; the smart home space intelligent agent, based on an AI big model, acquires multi-dimensional spatial factors of the perceived home space, and under the calculation and reasoning of the home space big model, realizes intent recognition, parsing, decision-making and execution in different working modes. 2.The intelligent companion robot of claim 1, wherein, The sensing module includes: Microphone module, microphone array, long-distance sound pickup, sound source localization and directional sound pickup, to obtain user voice commands; The camera module is used for image and video capture. Touchscreen, allowing users to issue commands via touch; Multimodal sensors collect multimodal environmental data. 3.The intelligent companion robot of claim 1, wherein, The execution module includes: The speaker module, combined with AI voice synthesis capabilities, enables personalized and ultimate sound output; The projection module automatically focuses and corrects, projecting onto a giant screen. The smart home module responds to control commands and executes operations; The robot body actuators are driven to perform actions according to control commands. 4.The intelligent companion robot of claim 1, wherein, Different working modes include, but are not limited to: The proactive home theater mode recognizes user commands, understands user intentions, and has the ability to remember the needs of members in a smart home. It searches for online resources, formulates the best viewing plan, dims the lights, draws the curtains, and plays 4K content; it also coordinates various modules to create an immersive environment. The multimodal teaching companion mode involves a perception module that captures images and videos; a visual model that identifies image content; and a large language model that generates a story about the image and suggests projecting it. Launch an interactive educational model; The environmental adaptive butler mode uses online weather alerts combined with temperature and humidity sensors to determine the indoor temperature of the home and take tiered safety actions. In seamless video call mode, the user says "Make a video call to XX" and walks to the sofa; the system understands the command, recognizes the user's location, adjusts the camera and projection angles for the best composition and viewing experience, and initiates the video call. Personalized wake-up experience mode: The large language model generates wake-up plans based on the schedule and user preferences, while checking weather and traffic information to provide a one-stop morning service. 5.The intelligent companion robot of claim 1, wherein, The control module receives the reasoning and decision-making results of the smart home space intelligent agent, queries the status of each execution module, and performs task scheduling, behavior planning, and escapes and sends execution instructions to each execution module to complete the issuance of control instructions. 6.The intelligent companion robot of claim 1, wherein, The intelligent home space agent receives human-computer interaction or multi-dimensional multimodal environmental data input from the perception module. Through natural language processing (NLP) or large language model (LLM), it performs speech, semantic, or emotion parsing and intent recognition on the input information. Combining the short-term context and long-term memory retrieval capabilities of the space agent, it performs computational reasoning and task decomposition planning based on the AI Agent architecture through a local large model or an AI cloud large model, and decides whether to call external tools or sub-agents. Then, based on the returned results, it makes intelligent decisions and executes actions, and finally outputs feedback to the user and updates the status according to the environmental response.
7. An intelligent companion robot interaction method applied to the intelligent companion robot of any one of claims 1-6, characterized in that, The specific steps are as follows: Step 1: Command Input and Sensing. The system receives user commands, performs initial updates and optimizations of the commands through the sensing module, and ensures proper handling in case of no response or abnormality. Step Two: Core Thinking and Decision-Making. This involves task planning for the identified structured instructions, initiating a query, and generating an executable task sequence. Step 3: Instruction classification and scheduling. Based on the content of the task sequence, instructions are classified into different control domains and tasks are scheduled according to priority. Step 4: Execution of multi-dimensional data, system call response, handling data storage and retrieval, media playback and user feedback respectively, to achieve collaborative execution of multi-dimensional data; Step 5: Complete data processing, organize the execution results, and combine them with basic information about the physical environment to conduct multi-objective design and user experience evaluation to ensure that the system response meets the actual scenario and user expectations.