Control method of body-aware device, electronic device, and storage medium

CN122837643APending Publication Date: 2026-09-29ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611357268.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-09-03
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

但是大模型的每次动作决策仅依赖当前感知输入,导致在需要连续推理和自适应调整的复杂任务中,大模型输出的动作序列对任务的适配性较差

Benefits of technology

[0009]上述方案,通过将与当前任务请求相关的当前图像特征、文本特征与历史缓存信息共同输入包含第一大模型和动作专家网络的动作预测模型,使得语义理解与动作决策解耦且协同工作,通过引入历史缓存信息中的历史图像特征,使得动作预测模型能够获取任务执行过程中的时间上下文与状态演变信息,从而在面临物体遮挡或位置移动时,能够基于历史观测推断与当前任务请求相关的任务对象的当前物体状态,提高了视觉感知在动态环境下的鲁棒性;通过引入历史关节角度,使得动作预测模型能够感知具身智能设备自身的运动轨迹与物理状态,有利于后续生成的预测动作具备连续性与平滑性,避免了因具身智能设备的状态跳跃导致的机械冲击或动作失控,从而提升了具身智能设备在复杂任务下的动作预测精度与当前任务请求执行成功率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122837643A_ABST
    Figure CN122837643A_ABST
Patent Text Reader

Abstract

This application discloses a control method, electronic device, and storage medium for an embodied intelligent device. The control method includes: in response to receiving a current task request, extracting features from the acquired current image and the target task description of the current task request to obtain current image features and current text features; inputting the current image features, current text features, and historical cache information associated with the current image into an action prediction model; processing the current image features, current text features, and historical image features through a first-level model to obtain target fusion information output by the first-level model; processing the target fusion information and historical joint angles through an action expert network to obtain the predicted action of the current task request; and controlling the embodied intelligent device to execute the current task request with the predicted action. This approach can improve the accuracy of the predicted actions executed by the embodied intelligent device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of embodied intelligence technology, and in particular to a control method, electronic device, and storage medium for an embodied intelligence device. Background Technology

[0002] In recent years, the field of end-to-end imitation learning has developed rapidly. Large models generally possess excellent versatility and environmental adaptability, enabling direct deployment of large models on embodied intelligent devices. This allows for the direct output of motion control signals from embodied intelligent devices in an end-to-end manner based on perceptual data during the execution of actual task instructions, significantly simplifying the deployment process of embodied intelligent systems. However, each action decision of a large model relies solely on the current perceptual input, resulting in poor adaptability of the action sequences output by the large model to complex tasks requiring continuous reasoning and adaptive adjustment.

[0003] Therefore, there is an urgent need for an effective control method for embodied intelligent devices. Summary of the Invention

[0004] This application provides at least one control method, electronic device, and storage medium for a holographic smart device, which can improve the accuracy of the predicted actions performed by the holographic smart device.

[0005] This application provides a control method for an embodied intelligent device. The control method includes: in response to receiving a current task request, extracting features from the acquired current image and the target task description of the current task request to obtain current image features and current text features; inputting the current image features, current text features, and historical cache information associated with the current image into an action prediction model; processing the current image features, current text features, and historical image features through a first model to obtain target fusion information output by the first model; processing the target fusion information and historical joint angles through an action expert network to obtain the predicted action of the current task request; and controlling the embodied intelligent device to execute the current task request with the predicted action.

[0006] This application provides a control device for an embodied intelligent device, comprising: an extraction module, an input module, a first processing module, a second processing module, and a control module; the extraction module is used to extract features from the acquired current image and the target task description of the current task request in response to receiving a current task request, respectively, to obtain current image features and current text features; the input module is used to input the current image features, current text features, and historical cache information associated with the current image into an action prediction model; the first processing module is used to process the current image features, current text features, and historical image features through a first model to obtain target fusion information output by the first model; the second processing module is used to process the target fusion information and historical joint angles through an action expert network to obtain the predicted action of the current task request; the control module is used to control the embodied intelligent device to execute the current task request with the predicted action.

[0007] This application provides an electronic device, including a memory and a processor, wherein the processor is used to execute program instructions stored in the memory to implement the control method of the aforementioned embodied intelligent device.

[0008] This application provides a computer-readable storage medium storing program instructions thereon, which, when executed by a processor, implement the control method of the aforementioned embodied intelligent device.

[0009] The above scheme decouples semantic understanding and action decision-making while allowing them to work collaboratively by inputting current image features, text features, and historical cache information related to the current task request into an action prediction model that includes a first-level model and an action expert network. By introducing historical image features from the historical cache information, the action prediction model can obtain temporal context and state evolution information during task execution. Thus, when faced with object occlusion or positional movement, it can infer the current object state of the task object related to the current task request based on historical observations, improving the robustness of visual perception in dynamic environments. By introducing historical joint angles, the action prediction model can perceive the motion trajectory and physical state of the embodied intelligent device itself, which is beneficial for the continuity and smoothness of the subsequently generated predicted actions. This avoids mechanical shocks or loss of control caused by state jumps of the embodied intelligent device, thereby improving the action prediction accuracy and the success rate of current task request execution of the embodied intelligent device under complex tasks.

[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description

[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.

[0012] Figure 1 This is a flowchart illustrating an exemplary embodiment of the control method for an embodied intelligent device according to this application; Figure 2 This is a first framework schematic diagram of an exemplary embodiment of the control method for an embodied intelligent device of this application; Figure 3 This is a schematic diagram of the spatial perception network structure in an exemplary embodiment of the control method for an embodied intelligent device of this application; Figure 4 This is a second framework schematic diagram of an exemplary embodiment of the control method for an embodied intelligent device of this application; Figure 5a This is a schematic diagram of the third frame of an exemplary embodiment of the control method for an embodied intelligent device of this application; Figure 5b This is a schematic diagram of the first current image in an exemplary embodiment of the control method for an embodied smart device of this application; Figure 6 This is a schematic diagram of the structure of an embodiment of the control device for the intelligent device of this application; Figure 7 This is a schematic diagram of the structure of an embodiment of the electronic device of this application; Figure 8 This is a schematic diagram of the structure of an embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It is understood that the specific embodiments described herein are only for explaining this application and not for limiting it. Furthermore, it should be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all structures. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0014] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.

[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus. The term "and / or" is merely a description of the association of related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, "many" in this document means two or more. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of elements, for example, including at least one of A, B, and C, which can mean including any one or more elements selected from the set consisting of A, B, and C. Additionally, the term "several" in this document means one or more.

[0016] This application provides control methods and control devices for embodied intelligent devices. The application scenarios of these control methods include, but are not limited to, control scenarios for embodied intelligent devices, such as all scenarios of embodied intelligent applications. Specific scenarios can be application scenarios for various robots, such as home assistants, logistics, and manufacturing. The task capabilities of the embodied intelligent devices cover simple tasks, complex tasks, and long-range tasks. The executing entity of the control methods can be the control device itself, for example, the control device can be located within the embodied intelligent device. For example, the control device can be located within a terminal device, server, or other processing device. The terminal device can be user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, etc. In some possible implementations, the control methods can be implemented by a processor calling computer-readable instructions stored in memory. For example, embodied intelligent devices can refer to: humanoid robots, single-arm / dual-arm robots, bipedal / quadrupedal robots, robotic arms, etc.

[0017] Please see Figure 1 , Figure 1 This is a flowchart illustrating an exemplary embodiment of the control method for an embodied smart device according to this application. Specifically, the control method for an embodied smart device may include the following steps: S11: In response to receiving the current task request, perform feature extraction on the acquired current image and the target task description of the current task request to obtain the current image features and the current text features.

[0018] The current task request represents the control information for the current task received by the embodied intelligent device at the current moment. For example, the current task request carries location information of the task environment and indication information of the task object. The location information of the task environment represents the indication information of the scene in which the embodied intelligent device is located when performing the current task. The indication information of the task object represents the target object that the embodied intelligent device needs to contact when performing the current task. The current image represents the image acquired by the embodied intelligent device for the current task at the current moment. For example, the current image includes the image acquired by the end effector of the embodied intelligent device at the current moment and the image acquired by the task environment at the current moment. The target task description represents the descriptive text for the current task. The current image features represent the image features obtained by feature extraction from the current image. The current text features represent the text features obtained by feature extraction from the target task description.

[0019] In some application scenarios, the current task request can be a task request generated based on the user-input task description text / default description text, and then sent to the processing end. For example, the processing end can be an embodied smart device or a cloud device that communicates and connects with the embodied smart device.

[0020] In some application scenarios, feature extraction is performed on the current image to obtain the current image features. This includes inputting the current image into an image embedding model and obtaining the current image features output by the image embedding model. The current image can be acquired by an image acquisition device of the embodied intelligent device, or by an image acquisition device that is communicatively connected to the embodied intelligent device.

[0021] In some application scenarios, feature extraction is performed on the target task description to obtain the current text features, including: inputting the target task description into a text embedding model to obtain the current text features output by the text embedding model.

[0022] S12: Input the current image features, current text features, and historical cache information associated with the current image into the action prediction model.

[0023] The action prediction model consists of a primary model and a motion expert network. The action prediction model can be a Vision-Language-Action (VLA) model or other pre-trained models for predicting actions. The primary model can be a Vision-Language Model (VLM) or other pre-trained models capable of reasoning and fusing features from input images and text. The primary model is used to generate target fusion information based on current image features, current text features, and historical image features. The motion expert network is used to generate the predicted action for the current task request based on the target fusion information and historical joint angles.

[0024] The historical cache information associated with the current image includes historical image features of historical images and historical joint angles of the embodied intelligent device. The historical cache information associated with the current image represents sub-information from several cached moments prior to the acquisition time of the current image (the current moment). Cached moments are earlier than the current moment. The sub-information of each cached moment represents the image features of past observation data at the same cached moment and the motion state of the embodied intelligent device. Historical images are observation images from cached moments prior to the current moment. Historical image features of historical images represent image features obtained by feature extraction from historical images. Historical joint angles represent the joint angles of the embodied intelligent device at the acquisition time of the historical image. Historical joint angles are used to describe the motion state of the embodied intelligent device at the acquisition time of the historical image.

[0025] S13: The first major model processes the current image features, current text features, and historical image features to obtain the target fusion information output by the first major model.

[0026] The target fusion information includes the target fusion features obtained by fusing current image features, current text features, and historical image features.

[0027] In some application scenarios, the internal processing logic of the first main model can be: directly fusing the current image features, current text features, and historical image features to obtain the target fused feature; or, preprocessing the current image features, current text features, and historical image features separately to obtain preprocessed current image features, preprocessed current text features, and preprocessed historical image features. Preprocessing can include data cleaning, feature scaling, and feature selection of the data to be processed. The target fused feature is obtained by fusing the preprocessed current image features, preprocessed current text features, and preprocessed historical image features.

[0028] S14: The target fusion information and historical joint angles are processed by a motion expert network to obtain the predicted motion requested by the current task.

[0029] The predicted action for the current task request represents the action that the embodied intelligent device, as output by the action expert network, predicts it should perform in the future after the current moment. The predicted action for the current task request consists of a sequence of actions composed of joint angles associated with several future moments.

[0030] In some application scenarios, the internal processing logic of a motion expert network (MAN) can be as follows: Perform motion prediction processing on the target fusion information and historical joint angles to obtain the predicted motion requested by the current task. Alternatively, preprocess the target fusion information and historical joint angles separately to obtain preprocessed target fusion information and preprocessed historical joint angles; the preprocessing can refer to the above content. Perform motion prediction processing on the preprocessed target fusion information and preprocessed historical joint angles to obtain the predicted motion requested by the current task.

[0031] S15: Control the embodied intelligent device to predict actions and perform the current task request.

[0032] After the embodied intelligent device obtains the predicted action for the current task request, the embodied intelligent device uses the predicted action to drive the completion of the current task request.

[0033] The above scheme decouples semantic understanding and action decision-making while allowing them to work collaboratively by inputting current image features, text features, and historical cache information related to the current task request into an action prediction model that includes a first-level model and an action expert network. By introducing historical image features from the historical cache information, the action prediction model can obtain temporal context and state evolution information during task execution. Thus, when faced with object occlusion or positional movement, it can infer the current object state of the task object related to the current task request based on historical observations, improving the robustness of visual perception in dynamic environments. By introducing historical joint angles, the action prediction model can perceive the motion trajectory and physical state of the embodied intelligent device itself, which is beneficial for the continuity and smoothness of the subsequently generated predicted actions. This avoids mechanical shocks or loss of control caused by state jumps of the embodied intelligent device, thereby improving the action prediction accuracy and the success rate of current task request execution of the embodied intelligent device under complex tasks.

[0034] In some embodiments, the current task request includes an initial task description. The initial task description represents the task description text input by the user. The target task description represents the task description text optimized from the initial task description. Before the steps of extracting features from the acquired current image and the target task description of the current task request to obtain current image features and current text features, respectively, the control method of the embodied intelligent device further includes: obtaining historical task descriptions associated with the current image; generating task plans based on the historical task descriptions; and processing the initial task description in the current task request, the task plans of the historical task descriptions, and the current image through a second model to obtain the target task description output by the second model.

[0035] The historical task description associated with the current image was generated earlier than the acquisition time of the current image. The historical task description associated with the current image represents the target task description of the previous task request preceding the current task request. The historical task description associated with the current image can be the task description text entered by the user in the previous task request, or an optimized task description text obtained by optimizing the task description text entered by the user in the previous task request. The historical task description associated with the current image can be obtained by either retrieving the target task description of a historical moment adjacent to the current moment from the task description database, or by directly using the task description text entered by the user in the previous task request.

[0036] The historical task description is a textual description of the action flow that instructs the embodied intelligent device to execute the previous task request.

[0037] In some application scenarios, the task planning of historical task descriptions can be generated in the following ways: directly using the historical task description as the task planning of historical task descriptions; or, optimizing the historical task description with preset text to obtain the task planning of historical task descriptions. The preset text optimization can be inputting the historical task description and preset prompt words into a large language model to obtain the task planning of historical task descriptions output by the large language model; or, calling common sense and experience bases to reason about the historical task description to obtain the task planning of historical task descriptions.

[0038] The second major model could be the Vision-Language Model (VLM), or other pre-trained models that can reason about input images and text and output optimized task descriptions.

[0039] In some application scenarios, the internal processing logic of the second major model can be: directly optimizing the initial task description, the task planning of the historical task description, and the current image to obtain the target task description; or, preprocessing the initial task description, the task planning of the historical task description, and the current image respectively to obtain the preprocessed initial task description, the preprocessed task planning of the historical task description, and the preprocessed current image. The preprocessing can refer to the above content, and image preprocessing can include image noise reduction, region of interest cropping, and sharpness adjustment. The target task description is obtained by optimizing the preprocessed initial task description, the preprocessed task planning of the historical task description, and the preprocessed current image.

[0040] For example, such as Figure 2 As shown, after obtaining the historical task description, common sense and experience base are called to obtain the task plan of the historical task description. The initial task description in the current task request, the task plan of the historical task description, and the current image are input into the second large model for processing to obtain the target task description output by the second large model. The target task description is input into the text embedding model to obtain the current text features; the current image is input into the image feature extraction module to obtain the current image features output by the image feature extraction module. In some application scenarios, the image feature extraction module can be equipped with a preset feature extraction network. Specifically, the preset feature extraction network can be a convolutional neural network, a recurrent neural network, a long short-term memory network, a gated recurrent unit, a BiGRU neural network, etc. The current text features, the current image features, and the historical cache information are sequentially passed through the first large model and the action expert network in the action prediction model to obtain the predicted action output by the action prediction model. For example, the initial task description in the current task request can be "tidy up the desktop". The task plan of the historical task description can be "I close the drawer. Put the pen in the pen holder". The target task description corresponding to the current task request can be "close the notebook".

[0041] For example, if an embodied intelligent device has already completed the historical task of "organizing the desktop" ("I close the drawer and put the pen in the pen holder"), then at the current moment, the initial task description, the task planning of the completed historical task, and the current image are fed into the second main model to generate the current target task description: "Close the notebook." During this process, the planning is generated by connecting to a common-sense and experiential knowledge base. The core function of this knowledge base is to provide cognition for task planning, making the task more aligned with cognition, and to achieve personalized customization, providing personalized and experiential context for task planning, thus making the task more in line with user habits, executed at the right time, and with a higher success rate.

[0042] This application improves the success rate of long-term complex tasks by optimizing the initial task description to obtain the target task description. It leverages the skills possessed by the embodied intelligent device itself, combining high-success-rate skills with the matching degree of the current task to further enhance the task success rate. Furthermore, upon completion of a complex task, a complete task planning chain forms action experience that guides future tasks. This allows the embodied intelligent device to learn during continuous task execution, improving its adaptability to unknown scenarios. In some embodiments, the embodied intelligent device includes an end effector for contacting a task object. The task object represents a target object that needs to be moved / contacted in the task scenario of the current task request. The end effector is a component of the embodied intelligent device. The current image includes a first current image acquired towards the end effector at the time of acquisition and second current images acquired from at least two perspectives towards the task environment to which the current task request belongs. The current image features include first image features corresponding to the first current image and second image features obtained based on the second current images from each perspective. Specifically, the steps described above, which extract features from the acquired current image and the target task description of the current task request to obtain current image features and current text features, may include the following steps: Image encoding is performed on the first current image to obtain first image features. Second current images from each viewpoint are input into a depth estimation network to obtain spatial features of the target object in the task environment output by the depth estimation network. Second current images from at least two viewpoints are input into a semantic alignment network and a detection network respectively to obtain semantic features of the target object in the task environment output by the semantic alignment network and geometric features of the target object in the task environment output by the detection network. Second image features are determined based on the spatial features, semantic features, and geometric features.

[0043] Each second current image is an image taken from a different perspective of the same task environment. The first image features characterize the image features extracted from the first current image. For example, the task environment is the specific location or position indicated by the current task request when performing the task. The target object in the task scene is at least one physical object contained in the actual scene where the task is performed. At least one physical object includes the task object indicated by the current task request (i.e., the target object that the end effector needs to move / contact during task execution) and other objects. For example, the task environment could be a bedroom or an area within a bedroom; the target object in the task scene could include a desk drawer, a pen on the desk, a pen holder, and a notebook; in the case where the target task is described as "closing the notebook," the task object is the notebook.

[0044] Image encoding processing can involve inputting a first current image into an image embedding model to obtain the first image features output by the image embedding model. The image embedding model may include the aforementioned preset feature extraction network.

[0045] The second image feature represents the image features obtained by sequentially extracting and fusing features from the second current image at each viewpoint.

[0046] Spatial features represent the image features obtained by the depth estimation network from the second current image input for each viewpoint. Semantic features represent the image features obtained by the semantic alignment network from the second current image input for one viewpoint. Geometric features represent the image features obtained by the detection network from the second current image input for one viewpoint.

[0047] Semantic alignment networks are used for image semantic alignment; for example, a semantic alignment network is an image semantic alignment network. Detection networks are used for image perception; for example, a detection network is an image perception network.

[0048] In some application scenarios, the second image feature can be determined by: performing feature fusion processing on spatial features, semantic features, and geometric features to obtain candidate fused features, and using the candidate fused features as the second image feature; or, performing post-processing on the candidate fused features to obtain the second image feature. Post-processing includes, but is not limited to, mapping processing, or further fusing the mapping processing result with the candidate fused features.

[0049] For example, the spatial perception network is used to extract features from the second current image from various viewpoints, and the module / network that obtains the features of the second image is a spatial perception network (i.e., a 3D spatial perception network). A schematic diagram of its spatial perception network structure is shown below. Figure 3 As shown. Taking each viewpoint as two views as an example, for each viewpoint, the image acquired from that viewpoint towards the task environment to which the current task request belongs is used as the second current image for that viewpoint. For example, each viewpoint includes a first viewpoint and a second viewpoint. Specifically, the second current image for each viewpoint includes, as shown below. Figure 3 The second current image from the first viewpoint and the second current image from the second viewpoint are shown. For example, the inputs to the depth estimation network can be as follows: Figure 3 The second current image from the first perspective and the second current image from the second perspective are shown; the input to the semantic alignment network can be as follows: Figure 3 The second current image shown is from a second perspective; the input to the detection network can be as follows: Figure 3The second current image shown is from a second perspective. Further details regarding the specific inputs to each network in the spatial perception network will not be elaborated upon hereafter. It is understood that both the first and second perspectives are acquisition perspectives used when acquiring data about the task environment. They can be acquisition perspectives from the same image acquisition device (such as a binocular camera) or from different image acquisition devices. They can be flexibly set in practical application scenarios, and this application does not limit their use.

[0050] The 3D spatial perception network is the core perception component for achieving precise operation in this application. Its basic function is to generate pseudo-3D semantic features (i.e., the second image features mentioned above) based on the multi-view images at the current moment, so as to explicitly solve the object positioning deviation caused by the viewing distance error in single-view images and the resulting task failure problem.

[0051] Existing depth estimation networks typically aim to output dense depth maps or relative depth values, focusing primarily on the accuracy of geometric measurements while neglecting the embedding and alignment of semantic information. This makes it difficult for downstream policy networks to effectively associate depth signals with task semantics. To address this, this application proposes a 3D-like perception architecture that fuses semantic, geometric, and depth information. Instead of directly outputting depth values, it generates 3D-like semantic features that combine accurate spatial perception with strong semantic representation capabilities. These features reflect the target's position and geometric structure in 3D space and are well-aligned with textual instructions in the semantic space, fundamentally differentiating it from traditional depth estimation paradigms that only output geometric information. Figure 3 As shown, the spatial perception network receives binocular visual input for a business scenario, processes it in parallel through three complementary feature extraction links, and finally fuses it into a unified pseudo-3D semantic feature (i.e., the second image feature mentioned above).

[0052] The first link can include the following: The first link is a depth perception branch based on binocular vision (i.e., the depth estimation network mentioned above). This branch takes the left and right views of the task scene (i.e., the second current images of each viewpoint) as input, and perceives the depth distribution of different objects in the scene relative to the camera through stereo matching and disparity calculation mechanisms. The output of this branch is a high-dimensional feature containing the depth spatial structure (i.e., the spatial feature mentioned above), and this high-dimensional feature is represented as F. depth It preserves the relative positional relationships and depth levels between objects, providing spatial geometric priors for subsequent fusion.

[0053] The second link may include the following: The second link is a semantic alignment branch based on a single-view image (i.e., the second current image from one viewpoint) (i.e., the semantic alignment network mentioned above). This branch uses a visual coding network (such as CLIP visual encoder) with cross-modal image-text alignment capabilities to extract the semantic features of the image, and represents the semantic features as F. semThe advantage of this semantic feature lies in its embedding space being naturally aligned with the natural language semantic space, allowing the fused pseudo-3D semantic features to be directly semantically matched with the text task description without the need for an additional modality alignment adaptation layer. The unique value of this branch is that it explicitly introduces the semantic channels missing in traditional depth estimation, bridging the gap between geometric perception and task semantics.

[0054] The third link may include the following: The third link is a geometric detail-aware branch based on a single-view image (i.e., the second current image from one of the viewpoints) (i.e., the detection network mentioned above). This branch uses a segmentation or detection backbone network that is good at pixel-level detail perception to extract the location, edges, and local geometric features of the target, and represents the geometric features as F. seg Geometric features can supplement candidate fusion features with fine-grained spatial cues, enabling pseudo-3D semantic features to have a more sensitive ability to depict the precise outline, local deformation and spatial occupancy of objects.

[0055] like Figure 3 As shown, the characteristic F output by the above three links depth F sem F seg The fusion module performs channel-level concatenation to obtain fused features, and then feeds the candidate fused features into a fusion network composed of a multilayer perceptron (MLP). After nonlinear transformation, the final pseudo-3D semantic features are generated, and these pseudo-3D semantic features are represented as F. 3D-sem This pseudo-3D semantic feature possesses three key attributes: spatial 3D perception capability, strong semantic representation capability, and cross-modal alignment capability. The spatial 3D perception capability is ensured by both depth and geometry branches, enabling the differentiation of object distance, front-back relationships, and occlusion. The strong semantic representation capability is guaranteed by the semantic alignment branch, exhibiting high sensitivity to semantic information such as object category, attributes, and functions. The cross-modal alignment capability benefits from the embedding spatial characteristics of the semantic alignment branch, allowing this pseudo-3D semantic feature to directly perform similarity calculations or attention interactions with text modal features, achieving fine-grained visual-linguistic alignment.

[0056] Through the aforementioned three-stream fusion design, the spatial perception network of this application, in terms of output characteristics, outputs a composite feature representation that integrates "spatial geometry, semantic identity, and fine-grained contours," unlike networks that only output geometric values ​​reflecting distance. This allows downstream policy modules to directly perform joint reasoning on "where the object is, what it is, and how to operate it" within a unified pseudo-3D semantic space, eliminating the need for separate queries of depth and semantic maps and heuristic alignment when generating action sequences. This fundamentally reduces the risk of task failure due to viewing distance estimation bias and semantic fragmentation. To further enhance the spatial-semantic consistency of pseudo-3D semantic features, this application introduces cross-viewpoint semantic consistency constraints during the training phase. Specifically, for observations of the same task scene from different viewpoints, their representations in the pseudo-3D semantic feature space must remain similar. This constraint can be achieved through contrastive learning or feature distillation, enabling the spatial perception network to learn to map images of the same object with significant appearance differences from different viewpoints to neighboring regions in the feature space. This allows the pseudo-3D semantic features of the current image to be naturally robust to viewpoint changes in practical applications, maintaining stable semantic recognition and spatial localization capabilities even when the embodied intelligent device moves or the relative position of the object changes. Furthermore, the depth perception branch does not rely on expensive dense depth ground truth for supervision and can be jointly trained with the semantic alignment branch using multimodal contrastive loss and downstream task reward signals for weak or self-supervised training. This lowers the barrier to obtaining training data and improves the engineering feasibility of the solution.

[0057] In some embodiments, prior to S12, the control method for the embodied smart device further includes: acquiring a first historical image and a second historical image of the current image. In response to a situation where the total number of cached information in the historical cache information associated with the second historical image exceeds a threshold, the historical cache information associated with the second historical image is updated based on the state information of the embodied smart device at a first moment and the image features of the second historical image, thereby obtaining the historical cache information associated with the current image.

[0058] The first historical image was acquired at time 1, the second historical image was acquired at time 2, and time 1 is earlier than time 2, which is earlier than the acquisition time of the current image. For example, the first historical image, the second historical image, and the current image can be images acquired sequentially over time for the same task scenario. The first historical image, the second historical image, and the current image are images acquired consecutively at different times.

[0059] In some application scenarios, the first and second historical images of the current image can be obtained by retrieving two historical images from the historical image database that are before the current time and adjacent to each other, and using them as the first and second historical images respectively.

[0060] The historical cache information associated with the second historical image represents sub-information of several candidate times prior to the acquisition time (second time) of the second historical image. When the sub-information of the second time is updated to the historical cache information associated with the second historical image, at least some of the cached times are the same as at least some of the candidate times. When the sub-information of the second time does not need to be updated to the historical cache information associated with the second historical image, each cached time is the same as each candidate time, differing only in the naming of the same time in the historical cache information associated with different images.

[0061] The system determines whether the total number of cached information entries in the historical cache information associated with the second historical image exceeds a threshold. The historical cache information associated with the second historical image includes sub-information of several candidate time points. This sub-information includes candidate image features and candidate joint angles belonging to the same candidate time point. The candidate images were acquired earlier than the second time point. Candidate images are images acquired for the task scenario at the corresponding candidate time point. Candidate image features represent image features obtained by feature extraction from the candidate images. The method for feature extraction from candidate images is the same as that for the current image. Candidate joint angles represent the joint angles of the intelligent device itself at the time the candidate image was acquired (i.e., the candidate time point). The total number of cached information entries in the historical cache information is equal to the total number of candidate time points, the total number of candidate image features, or the total number of candidate joint angles in the historical cache information associated with the second historical image.

[0062] In some application scenarios, when the total number of cached information in the historical cache information associated with the second historical image is less than or equal to a threshold, the image features of the second historical image are associated with the joint angle of the acquired embodied intelligent device at the second moment to obtain sub-information at the second moment. This sub-information at the second moment is then stored in the historical cache information associated with the second historical image to obtain the historical cache information associated with the current image.

[0063] In other application scenarios, when the total number of cached information in the historical cache information associated with the second historical image exceeds a threshold, based on the state information of the embodied intelligent device at the first moment and the image features of the second historical image, it is determined whether to update the sub-information of the second moment to the historical cache information associated with the second historical image, thus obtaining the historical cache information associated with the current image. For example, the state information, the image features of the second historical image, and preset prompt words are input into a large language model to obtain the specificity score of the sub-information of the second moment output by the large language model; in response to the specificity score of the sub-information of the second moment being greater than or equal to the scoring threshold, the sub-information of the second moment is updated to the historical cache information associated with the second historical image, thus obtaining the historical cache information associated with the current image. In response to the specificity score of the sub-information of the second moment being less than the scoring threshold, the historical cache information associated with the second historical image is directly retained as the historical cache information associated with the current image.

[0064] The state information represents the observation information and execution state of the embodied intelligent device at the first moment. The image feature representation of the second historical image is the image feature obtained by feature extraction from the second historical image. The method for feature extraction from the second historical image is the same as the method for feature extraction from the current image. The historical cache information associated with the current image represents the historical cache information obtained by updating the historical cache information associated with the second historical image.

[0065] In some embodiments, the state information includes image compression features of a first historical image and the historical actions of the embodied intelligent device at a first moment. The observation information of the embodied intelligent device at the first moment is the image compression features obtained by compressing the first historical image. The execution state characterizes the historical actions of the embodied intelligent device at the first moment. Specifically, the step of updating the historical cache information associated with the second historical image based on the state information of the embodied intelligent device at the first moment and the image features of the second historical image to obtain the historical cache information associated with the current image includes: performing prediction processing on the image compression features and historical actions of the first historical image to obtain the potential state information of the embodied intelligent device at the second moment; performing decoding processing on the potential state information to obtain the predicted image features of the embodied intelligent device at the second moment; determining the prediction error associated with the second historical image based on the feature difference between the predicted image features and the image features of the second historical image; and updating the historical cache information associated with the second historical image based on the prediction error associated with the second historical image to obtain the historical cache information associated with the current image.

[0066] Historical actions represent the actions performed by the embodied intelligent device at the first moment. Image compression features refer to the low-dimensional, high-semantic-density feature representation extracted by the encoding network, used to reduce computational overhead and preserve core visual semantics. Historical actions refer to the joint angle sequence or end-effector pose increment executed by the embodied intelligent device at the first moment, used to drive the world model to predict state transitions. Latent state information refers to the compact vector representation within the world model used to summarize the environmental dynamics at the second moment, situated between the original observation and the decoded output. Predicted image features refer to the multimodal observation features reconstructed from the latent state information at the second moment, consistent with the dynamic cognition of the world model. Prediction error refers to the difference between the true historical image features and the predicted image features at the second moment, reflecting the surprise of the observation or the information increment. The process of updating the historical cache information associated with the second historical image can be a process of eliminating redundant frames or retaining keyframes to optimize the historical storage information based on the magnitude of the prediction error. It can be understood that each frame represents a sub-information associated with a timestamp / moment, and each frame includes the image features and joint angles of the historical image at the corresponding moment, which will not be elaborated further below.

[0067] For example, the image compression features and historical actions of the first historical image are input into the world model to obtain the predicted image features of the second moment output by the world model.

[0068] If the prediction error at the second time step is small, it indicates that the environmental change at that time step conforms to the dynamic laws, the visual content can be accurately predicted, and it is redundant information, so it does not need to be updated in the historical cache information associated with the second historical image. If the prediction error is large, it indicates that a visual abrupt change has occurred that is inconsistent with expectations (such as object contact, sliding, or the appearance of a new object), which is a keyframe, and the sub-information associated with the second time step needs to be updated in the historical cache information associated with the second historical image. Based on this prediction error, the cache is updated, prioritizing the retention of frames with high errors. For example, the prediction error threshold can be set to a fixed value or a statistical value within a sliding window. When the prediction error at the second time step is lower than the error threshold, it is marked as a redundant candidate frame; when the prediction error at the second time step is higher than the error threshold, it is marked as a keyframe. When constructing a short-term memory cache containing n time steps, image features marked as keyframes are preferentially selected to fill the cache slots. If there are not enough keyframes, they are supplemented from the redundant candidate frames according to the prediction error associated with the corresponding time step from large to small. When the cache is full and a new frame arrives, if the new frame is a keyframe, it replaces the frame in the cache with the smallest (most redundant) prediction error associated with the corresponding time step, thereby achieving dynamic optimization of the cache content. For example, the world model can be set as a Transformer-based sequence prediction network. The input is the compressed features and action sequences of historical images from the past k frames or the corresponding time of the previous frame, and the output is the predicted image features of the next frame. The world model is trained by minimizing the mean squared error between the predicted features and the true features, so that it can accurately predict visual changes under steady motion. Thus, during the inference stage, abnormal events (such as failed grabbing or falling objects) can be captured by abnormally high prediction errors, and the historical frames corresponding to these abnormal events can be stored in short-term memory for reference in the first major model of the subsequent action prediction model.

[0069] In some application scenarios, the error threshold can be a pre-set fixed value.

[0070] In some embodiments, the historical cache information associated with the second historical image includes sub-information of several candidate times. The sub-information includes candidate image features and candidate joint angles belonging to candidate images at the same candidate time, where the candidate images were acquired earlier than the second time. Specifically, the step of updating the historical cache information associated with the second historical image based on the prediction error of the second historical image association to obtain the historical cache information associated with the current image may include the following steps: determining an error threshold based on the prediction errors of the candidate image associations at several acquired candidate times and the prediction error of the second historical image association; in response to the prediction error of the second historical image association being greater than or equal to the error threshold, removing one of the sub-information of the several candidate times to obtain the removed historical cache information; associating the image features of the second historical image with the acquired joint angles of the embodied intelligent device at the second time to obtain the sub-information of the second time; and storing the sub-information of the second time in the removed historical cache information to obtain the historical cache information associated with the current image.

[0071] In some application scenarios, the error threshold can be determined by: calculating the mean and / or standard deviation between the prediction error associated with candidate images at several candidate time points and the prediction error associated with the second historical image; and determining the error threshold based on the mean and / or standard deviation.

[0072] The sub-information at the second moment includes the image features of the second historical image and the joint angle of the embodied intelligent device at the second moment. The historical cache information after elimination processing represents the historical cache information obtained by eliminating one of the sub-information from several candidate moments in the historical cache information associated with the second historical image. One of the sub-information from several candidate moments can be any sub-information associated with any candidate moment in the historical cache information associated with the second historical image, or the sub-information associated with the earliest candidate moment, or the sub-information associated with the candidate moment with the smallest prediction error, etc.

[0073] If the prediction error associated with the second historical image is greater than or equal to the error threshold, first remove one of the sub-information in the historical cache information associated with the second historical image, and then add the sub-information at the second moment to the historical cache information to obtain the historical cache information associated with the current image.

[0074] For example, the historical cache information associated with the current image in this application is constructed based on a short-term observation redundancy removal mechanism of the world model. To further enhance the information density of short-term memory (i.e., the historical cache information associated with the current image) within a limited window, this application introduces a short-term observation redundancy removal mechanism based on the world model. This short-term observation redundancy removal mechanism utilizes a pre-trained or online-adapted world model to determine the information importance of historical observations preceding the current image, actively removing sub-information of redundant historical moments with highly predictable dynamic evolution, thereby allocating the limited historical slots in the historical cache information to sub-information of key historical moments that have higher information gain for action decisions. Specifically, the world model receives the pseudo-3D semantic features of historical moments, the F1 scene features of the end effector, and the corresponding joint actions in sequence, learning the dynamic transfer of the environment.

[0075] An adaptive or fixed error threshold τ is set. If the prediction error associated with a historical image at a certain historical moment is less than the error threshold, it is considered that the observation at that historical moment can be completely inferred from the neighboring states by the world model, containing little new information, and is marked as sub-information of redundant candidate moments. Conversely, if the prediction error at a certain historical moment is large, it indicates that a visual or state change (such as contact, slippage, appearance or disappearance of objects, etc.) has occurred at that historical moment that is inconsistent with the dynamic prediction, and has significant reference value for task execution, and is marked as sub-information of key moments. Among them, the sub-information of key moments includes the image features and joint angles corresponding to the key moment.

[0076] During the construction of historical cache information, when short-term memories of n historical moments need to be saved, priority is given to selecting sub-information from the marked key moments to fill the cache slots. These cache slots are used to store historical cache information associated with the current image. If the number of sub-information from key moments is less than n, it is supplemented from the sub-information of redundant candidate moments in descending order of prediction error, until all n slots are filled. If the number of sub-information from key moments exceeds n, it is retained based on a combination of recent and distant time and prediction error magnitude, ensuring that the most recent and informative frames are not discarded.

[0077] For example, at each moment, the current historical cache information is updated to achieve cache renewal. When a new moment arrives, the world model evaluates the new observation in real time, generating prediction errors. If the short-term memory cache is full, a "first-in, first-out" strategy combined with redundancy weighting is used for updating, prioritizing the elimination of historical moments with the smallest prediction errors (i.e., the most redundant historical moment's sub-information), or prioritizing those that are older and have low information content, thereby ensuring that high-information short-term memory continuously resides in the historical cache information. This application can also degenerate into using the sub-information of historical moments within a fixed time window before the current moment as the historical cache information associated with the current image. However, compared to the strategy of simply storing all historical moments of n consecutive frames, this short-term memory caching mechanism based on world model prediction errors effectively eliminates redundant segments caused by stationary, stable movement, or no significant changes in the scene of the embodied intelligent device, allowing the same length n of short-term memory to cover task-critical nodes over a longer time span, improving the effectiveness of short-term memory, and partially alleviating the problem of insufficient effective information caused by time window limitations.

[0078] To balance implementation complexity, this application retains a simplification strategy. When the world model is not enabled or computational resources are limited, short-term memory can directly store sub-information (including image features and joint angles at corresponding moments) from n consecutive historical moments and update the historical cache information through a first-in-first-out queue. However, by enabling redundancy removal using the world model, the n moments with the most information can be selected from a longer historical interval to construct the memory, resulting in a significant enhancement in both the temporal coverage and decision relevance of "short-term memory".

[0079] Specifically, the state transition and observation prediction process of the world model is as follows: The world model in short-term memory summarizes the environmental dynamics with compact latent states. For time t-1, given the state information z of the embodied intelligent device at the previous time (t-2). t 2. Robotic Actions a t 2. For the process of the world model predicting the prior of the current potential state at time t-1, please refer to formula (1): Z t-1 =f dyn (z t 2,a t 2; dyn ) formula (1); Where t represents the current time, t-2 represents the first time, and t-1 represents the second time. The potential state information of the embodied intelligent device at the second time is predicted using the state information from the first historical time. t-1This represents the potential state information of the embodied intelligent device at a second moment. t 2 represents the image compression features of the first historical image. a t 2 represents the historical actions (joint angle increments or end poses) of the embodied intelligent device at the first moment, which are recorded by the embodied intelligent device itself and provided to the world model. dyn This represents preset parameters or random environmental variables that change with the first historical image at the first moment. dyn This represents the preset function, which is the internal computation function of the world model. The process of decoding the predicted values ​​of the current multimodal observation features from this latent state, i.e., obtaining the predicted image features at the second time step, can be referred to formula (2): o t-1 =g dec (Z t-1 ; dec ) formula (2); Among them, o t-1 Z represents the predicted image features at the second time step. t-1 This represents the potential state information of an embodied intelligent device at a second moment. dec This represents the preset parameters. dec This represents the preset function, which is the internal computational function of the world model. The image features of the second historical image are represented as o. t-1 , and can be represented as o t-1 =[First Image Features, Second Image Features], that is, the image features of the second historical image are obtained based on the first image features and the second image features of the second historical image. The image features of the second historical image are a concatenated vector of the pseudo-3D semantic features (second image features) at the second time step and the end effector scene features (first image features).

[0080] Current real observation features o t-1 With predictive features o t-1 The difference between the two times measures the surprise or information content of the observation at that moment and is used as the prediction error at the second moment. The process of determining the prediction error at the second moment can refer to the following formula (3): e t-1 =||o t-1 o t-1 || 2 Formula (3); Among them, e t-1 This represents the prediction error at the second time step. t-1 The predicted image features represent the second time step. t-1Image features representing the second historical image. Formula (3) uses measures such as cosine distance or L2 norm, surprise or prediction error e t-1 The larger the value, the less the visual-state change of the frame conforms to the existing dynamic understanding of the world model, and the higher the amount of decision-making information it contains.

[0081] The binary discrimination process at critical moments can refer to the following formula (4): k t-1 ={1,e t-1 ≥τ0,e t-1 <τ1} Formula (4); Here, a threshold τ1 is set to obtain the keyframe label k. t-1 The value of τ1 represents the critical moment at time t-1. τ1 can be a fixed hyperparameter or an adaptive threshold, such as τ1=μ+σ, where μ and σ are the mean and standard deviation of the prediction errors for N_e historical moments (i.e., several cached moments) within the sliding window, respectively. τ0 represents the preset prediction error, which is a preset value dynamically set in advance, i.e., the lower limit of the preset error. It can be understood that the critical moment can be determined through formula (4), and then the sub-information of the critical moment can be used as the historical cache information associated with the current image.

[0082] Taking the task of "retrieving an item from a drawer" as an example, without short-term memory, the embodied intelligent device's subject relies solely on the current moment's sub-information to predict the action sequence, thus failing to perceive other objects previously observed in the drawer. While long-term memory provides common-sense and experiential priors, it is not suitable for this scenario. The items in the drawer may change in real time with each operation. Requiring long-term memory to follow these frequent and subtle changes necessitates complex real-time perception update logic. Furthermore, excessively frequent rewriting of long-term memory can lead to confusion in common-sense memory and higher forward reasoning latency. By introducing short-term memory, when predicting future action sequences in the current moment, the embodied intelligent device's subject can directly reference the visual content actually observed by the subject within the past n key moments, thereby obtaining immediate contextual information. For example, during the action of "opening the drawer to retrieve a pen," the wrist camera captures a pair of headphones in the drawer in a historical frame. Although the current task is unrelated to the headphones, this observation is recognized as a high-information keyframe by the world model and retained in short-term memory. If a command related to headphones appears in a subsequent task, the entity can quickly reference the prior visual evidence in short-term memory, such as "the headphones are in the drawer," to directly generate an efficient action decision, without needing to re-explore or rely on the slow retrieval of long-term memory. The long-term memory module of this application (which optimizes the initial task description to obtain the target task description) is mainly used for summarizing common-sense knowledge, accumulating experience, and planning tasks, thereby providing a task description that is more conducive to improving the success rate of tasks. In some embodiments, the historical cache information associated with the current image includes sub-information of several cache times. Historical image features include cached image features of cached images at several cache times. Historical joint angles include cached joint angles at several cache times. Sub-information includes cached image features and cached joint angles belonging to the same cache time. Cached image features and cached joint angles belonging to the same cache time are associated and stored as sub-information of one cache time. Target fusion information includes a target fusion feature representing the feature fusion result between the current image features, current text features, and historical image features. The feature fusion result between the current image features, current text features, and historical image features is used as the target fusion feature. Specifically, S13 above may include the following steps: performing fusion processing on the current image features and current text features to obtain initial fusion features. Determining the fusion weight of each cached image feature based on the current text features and cached image features associated with several cache times. Determining the target fusion feature based on the initial fusion feature, cached image features associated with each cache time, and the fusion weight of the corresponding cached image features.

[0083] The initial fusion feature represents the fusion feature between the current image features and the current text features. In some application scenarios, the initial fusion feature can be determined by inputting the current image features and the current text features into a feature fusion module, which is equipped with a preset feature fusion network. The preset feature fusion network may include, but is not limited to, a feature concatenation network, an attention-based fusion network, or a bilinear pooling-based fusion network.

[0084] The fusion weight of cached image features represents the image correlation between the current text features and the cached image features associated at each cache time. The greater the image correlation, the greater the fusion weight of the cached image features. In some application scenarios, for each cached image feature associated at each cache time, the similarity between the current text features and the cached image features is normalized to obtain the fusion weight of the cached image features.

[0085] In some application scenarios, the initial fusion features, the cached image features associated with each cache time step, and the fusion weights of the corresponding cached image features are input into a preset fusion module to obtain the target fusion features output by the preset fusion module. The preset feature fusion network set on the preset fusion module is described above.

[0086] In some embodiments, the above S14 may include the following steps: obtaining the predicted action based on the target fusion features and the cache joint angle at several cache times.

[0087] The target fusion features and cached joint angles at several cached times are input into the motion expert network to obtain the predicted action output by the motion expert network.

[0088] In some embodiments, the step of determining the fusion weight of each cached image feature based on the current text features and cached image features associated with several cached times may include the following steps: using the similarity between the cached image features at each cached time and the current text features as the similarity associated with the cached image features at the corresponding cached time; using the sum of the similarities associated with each cached image feature as the target similarity; and obtaining the fusion weight of the corresponding cached image feature based on the similarity associated with each cached image feature and the target similarity.

[0089] The similarity of cached image features represents the similarity between cached image features at the current cache time and current text features. This similarity can be calculated using methods such as cosine similarity or Euclidean distance. Target similarity represents the sum of the similarities of all cached image feature associations.

[0090] In some application scenarios, the fusion weight of cached image features can be determined as follows: for each cached image feature, the ratio between the similarity associated with the cached image feature and the target similarity is used as the fusion weight of the cached image feature.

[0091] In other application scenarios, for each cached image feature, the similarity associated with that cached image feature is substituted into a preset exponential function to obtain the exponential function value associated with the cached image feature; the sum of the exponential function values ​​associated with each cached image feature is used as the target exponential function value. For each cached image feature, the ratio between the exponential function value associated with that cached image feature and the target exponential function value is used as the fusion weight of the cached image feature. The contribution of sub-information at different times in the short-term memory sequence (i.e., the historical cached information associated with the current image) to action generation varies significantly. The introduction of some redundant or irrelevant sub-information not only increases computational overhead but may also introduce noise that interferes with decision-making. Therefore, this application introduces a sequence keyness perception mechanism within the action prediction model, enabling the action prediction model to autonomously learn to evaluate the decision relevance of each historical frame. Specifically, in the feature interaction stage, the first major model learns a keyness weight (fusion weight of cached image features) for each cached image feature in the sub-information at each cached time. This keyness weight can be represented as αt, which is jointly determined by the current text feature requested by the current task and the cached image feature at the corresponding cached time. The current text feature can be represented as F2. Specifically, the process of determining the fusion weights of cached image features can refer to the following formula (5): Formula (5); in, F2 represents the cached image features at cache time t. F2 represents the current text features. The cached image features represent any cached time in the historical cache information associated with the current image, where This represents any cached moment. This represents any given cached time point belonging to a number of cached times points within the historical cached information associated with the current image. This represents a set of cached times from the historical cache information associated with the current image, i.e., a set of time indices in the short-term memory cache. s represents the similarity function. exp represents the preset exponential function. This represents the fusion weight or critical weight of the cached image features at cache time t. Through this attention-based weight allocation, the first major model automatically focuses on critical cache times that are highly relevant to the task semantics, suppressing the interference of cached image features from irrelevant cache times, and achieving adaptive selection of sequence-level features.

[0092] like Figure 4 As shown, the first image features, the second image features, and the current text features are input into the first master model of the action prediction model to obtain the initial fusion features output by the first master model. The initial fusion features and historical image features in the short-term memory cache are then input into the first master model to obtain the target fusion features output by the first master model. The target fusion features and historical joint angles are then input into the action expert network to obtain the predicted action output by the action expert network. It can be understood that the initial fusion features are intermediate feature representations output by the first master model, and they need to be processed with the historical image features input into the first master model to obtain the target fusion features.

[0093] In some embodiments, the target fusion information further includes entity matching results between cached image features at each cache time and entity objects in the target task description. After the above steps of fusing current image features and current text features to obtain initial fused features, the control method of the embodied intelligent device further includes: extracting key information from the current text features to obtain entity semantic features of the target task description; and determining entity matching results based on the entity semantic features and cached image features at each cache time.

[0094] Entity matching results refer to the association mapping relationship between the key objects or operation objects (i.e., task objects) specified in the target task description and the specific visual targets appearing in the cached images at each cache time. This can be a Boolean value or confidence score obtained by calculating the cosine similarity threshold between the text entity embedding vector and the image region feature vector, or it can be a specific bounding box coordinate mapping. Entity semantic features refer to the semantic vectors extracted from the current text features that are related to the core operation objects of the current task. This can be achieved by extracting the semantic representations of noun entities in the target task description through named entity recognition, dependency parsing, or prompt word engineering of large language models. Predicted actions refer to the joint angle sequence or end effector pose changes required for the embodied intelligent device to perform the current task.

[0095] In some embodiments, the step of obtaining the predicted action based on the target fusion features and the cache joint angles at several cache times includes: obtaining the predicted action based on the target fusion features, the cache joint angles at several cache times, and the entity matching results.

[0096] Specifically, the target fusion features, cached joint angles at several cached times, and entity matching results are input into the action expert network to obtain the predicted action output by the action expert network.

[0097] In some embodiments, the entity matching result includes entity association features between entity objects in the target task description and cached image features at several cached time points, and / or fusion weights of cached joint angles at each cached time point. The fusion weights of cached joint angles characterize the degree of association between entity objects in the target task description and cached joint angles. The step of determining the entity matching result based on entity semantic features and cached image features at each cached time point includes: matching the entity semantic features with each sub-feature in the cached image features at each cached time point to obtain the feature matching result for the corresponding cached time point; and determining the entity association features and / or fusion weights of each cached joint angle based on the feature matching results at each cached time point.

[0098] Entity association features refer to the similarity or correspondence mapping between key semantic entities (such as "water cup" or "door handle") in the target task description and cached image features at each cache time step. This can be determined by calculating the cosine similarity between entity semantic features and image region feature vectors, cross-attention weights, or confidence scores based on bounding box IoU. Cache joint angle fusion weights refer to the association strength coefficient between robot joint angle data at each cache time step and the task object or entity operation object, used to quantify the importance or relevance of historical action states to the task object. Sub-features refer to the local feature vectors obtained after feature decomposition or slicing of cached image features, such as depth components, semantic components, edge components, or target components in image features. The feature matching results at each cache time step represent the matching results between entity semantic features and at least some sub-features in the cached image features at that time step. For example, the feature matching results at each cache time step indicate whether each sub-feature in the cached image features at that time step contains the task object in the entity semantic features.

[0099] For example, the entity semantic features and the cached image features at each cache time are decomposed into multiple sub-features. Then, the entity semantic features are matched one by one with each sub-feature in the cached image features at each cache time. For example, the semantic features of "water cup" are matched with the "geometric depth sub-feature" in the cached image features to locate the three-dimensional position of the water cup, matched with the "texture semantic sub-feature" to confirm the material or category of the water cup, and matched with the "edge contour sub-feature" to obtain the precise contour of the water cup. Then, the system calculates an overall "entity association feature" based on the matching results of these sub-features (such as similarity score and confidence). This feature can be a weighted similarity vector. At the same time, the system analyzes the temporal relationship between the joint angle sequence at each cache time and the entity appearance. If the joint angle change at a certain time is highly synchronized with the significant change in the entity position, then the joint angle at that time is given a higher fusion weight.

[0100] In some embodiments, the feature matching result at each cached time includes the matching degree between the cached image features and the entity semantic features at that cached time. The step of determining the fusion weights of entity association features and / or each cached joint angle based on the feature matching results at each cached time may include the following steps: determining the fusion weights of the cached joint angles at the corresponding cached time based on the matching degree between the cached image features and the entity semantic features at each cached time. And / or, in response to the feature matching result at least one cached time indicating that the entity semantic features and at least some sub-features in the cached image features at the corresponding cached time are successfully matched, the spatial locations of the successfully matched sub-features in the cached image features at at least one cached time are combined to obtain a target spatial location, and the at least one cached time is combined to obtain a target time. The target spatial location and the target time are aggregated to obtain entity association features.

[0101] The matching degree between cached image features and entity semantic features at each cached time refers to a numerical measure of the similarity or correlation between the cached image features and entity semantic features. It can be calculated using mathematical methods such as cosine similarity, inner product, or normalized cross-entropy. Its numerical range is usually normalized to a specific interval for subsequent weighted processing. Sub-features refer to the local feature vectors obtained after feature decomposition or slicing of cached image features, such as depth components, texture semantic components, and edge contour components in cached image features, or word vectors and syntactic vectors in text entity features. The target spatial location represents the set of spatial locations of sub-features that successfully match entity semantic features in the cached image features at each cached time. The target time represents the set of cached times to which the sub-features that successfully match entity semantic features belong.

[0102] In some application scenarios, the fusion weight for the cache joint angle can be determined by: directly using the matching degree between the cached image features and entity semantic features at each cache time as the fusion weight for the corresponding cache joint angle; or, using the sum of the matching degrees between the cached image features and entity semantic features at each cache time as the target matching degree; and for each cache time, using the ratio between the matching degree between the cached image features and entity semantic features at that cache time and the target matching degree as the fusion weight for the corresponding cache joint angle.

[0103] In some embodiments, the step of obtaining the predicted action based on the target fusion features, the cache joint angles at several cache times, and the entity matching results includes: obtaining the predicted action based on the target fusion features, the cache joint angles at each cache time, the entity association features, and / or the fusion weights of each cache joint angle.

[0104] Specifically, the target fusion features, cached joint angles at several cached time points, and entity association features are input into the action expert network to obtain the predicted action output by the action expert network. Alternatively, the target fusion features, cached joint angles at several cached time points, and the fusion weights of each cached joint angle are input into the action expert network to obtain the predicted action output by the action expert network. Or, the target fusion features, cached joint angles at several cached time points, the fusion weights of each cached joint angle, and entity association features are input into the action expert network to obtain the predicted action output by the action expert network.

[0105] In some embodiments, the control method for an embodied intelligent device further includes a training step for an action prediction model. The training step includes: acquiring sample image features, sample text features, historical cache information associated with the sample image, and the actual action of the sample task request; masking the historical cache information associated with the sample image to obtain masked historical cache information; masking the sample text features to obtain masked sample text features; processing the sample image features, masked sample text features, and masked historical cache information using the action prediction model to obtain the sample action of the sample task request output by the action prediction model; and training the action prediction model based on the difference between the sample action and the actual action until a fully trained action prediction model is obtained.

[0106] In the training process of the action prediction model, masking refers to randomly masking or discarding certain dimensions of the data, sub-information from certain moments in the historical cache information, or semantic units before the training data is input into the action prediction model. This simulates observation noise, missing information, or incomplete representations that may occur in the real environment. The sample task request is similar to the current task request; it is the task request for the training process of the action prediction model. The sample image is similar to the current image, representing an image captured by the embodied intelligent device for the current task at the sample moment. The sample image features are obtained in the same way as the current image features; feature extraction is performed on the sample image to obtain the sample image features. The sample text features are obtained by extracting features from the target task description of the sample task request; the method for obtaining the sample text features is the same as the method for obtaining the current text features. The real action representation of the sample task request is the pre-annotated action of the sample task request. The sample action of the sample task request is similar to the predicted action of the current task request; it represents the predicted action of the sample task request and is generated by the action prediction model.

[0107] In the action prediction model, the internal processing logic for sample image features, masked sample text features, and masked historical cache information is the same as the internal processing logic for current image features, current text features, and obtaining historical cache information associated with the current image. The sample image features are similar to the current image features, the masked sample text features are similar to the current text features, and the masked historical cache information is similar to the historical cache information associated with the current image. For the specific implementation logic, please refer to the above content, which will not be repeated here.

[0108] Specifically, based on the difference between the sample action and the real action, the target loss is determined, and the action prediction model is trained with the goal of reducing the target loss, until a fully trained action prediction model is obtained.

[0109] The method involves masking to force the action prediction model to learn to recover correct semantic alignment and action mapping from incomplete cross-modal inputs. Specifically, when constructing training samples, data pairs containing sample task requests, sample image features, sample text features, and historical cached data associated with the sample images are first acquired. Next, a temporal mask (randomly discarding some historical frames) or a spatial mask (randomly obscuring some feature dimensions) is applied to the historical cache information to simulate memory cache updates or observation occlusion. Simultaneously, a semantic mask (randomly discarding or replacing some words) is applied to the sample text features to simulate instruction expression variations. The masked multimodal data is then input into the action prediction model, which predicts action sequences based on this incomplete context to obtain sample actions. Finally, the predicted sample actions are compared with pre-labeled real actions, the loss function is calculated, and backpropagation is used to update the action prediction model parameters. Taking a drawer retrieval task as an example, a historical cache containing keyframes from the past 5 seconds can be constructed. During training, 2 seconds of frame data are randomly discarded, and the semantic information of the word "pen" in the target task description of the sample task request is masked. The action prediction model needs to combine the remaining visual cues and task intent (such as "take something") to infer the correct grasping action of the embodied intelligent device for the sample task, thus learning to maintain decision-making stability even with incomplete information. For another example, in long-term tasks, a scenario where keyframes are missing due to redundancy removal in the historical cache can be simulated. The action prediction model is trained to use common-sense experience stored in long-term memory to fill in the gaps in short-term memory, ensuring the coherence of action generation.

[0110] To further enhance the robustness of alignment between text features and image sequence features, this application introduces a dynamic semantic masking strategy. During the training phase, randomized, asymmetric masking perturbations are applied to the historical cache information associated with sample images in the short-term memory sequence (e.g., historical image features associated with sample images). The masking of the historical cache information associated with sample images can be spatial masking and / or temporal masking. Specifically, spatial masking can be: randomly masking a portion of the image feature vectors in the historical cache information associated with sample images, simulating observation noise and local occlusion. Temporal masking can be: randomly discarding sub-information at certain moments in the historical cache information associated with sample images, simulating information gaps caused by cache replacement or redundancy removal in short-term memory. Masking of sample text features can be semantic masking, which can be: randomly discarding or replacing sample text features requested by the sample task at the word level, simulating scenarios with incomplete task descriptions or variations in expression. Under the above masking conditions, the action prediction model is forced to learn to recover the correct semantic alignment and action mapping from incomplete cross-modal inputs, thereby achieving strong robustness to input perturbations. This strategy works in synergy with the aforementioned short-term memory redundancy removal mechanism. Even if short-term memory discards some historical information from before the sample time due to redundancy removal, the dynamic mask training of the action prediction model has enabled it to make stable inferences from sparse, incomplete sequences.

[0111] This application designs a cross-granularity semantic association enhancement mechanism to explicitly enhance the association between the target task description and the historical cache information associated with the current image, where similar semantic objects appear. This avoids the action prediction model relying solely on the current image and task description, thus ignoring sub-information from the short-term memory sequence (historical cache information associated with the current image) where the same or related objects may have appeared at different cache times. Specifically, an object-level semantic matching pathway is introduced: firstly, task-related object semantic vectors are explicitly or implicitly extracted from the current text features of the current task and used as entity semantic features. Entity semantic features can be represented as S... obj Subsequently, the entity semantic features are subjected to sliding matching on the local feature blocks of each frame in the short-time memory sequence (i.e., the local feature blocks in the cached image features of each cached time in the historical cache information associated with the current image). This identifies the spatial locations and corresponding times of all objects or semantically similar objects that have appeared in each cached image feature, which are used as the feature matching results for the corresponding cached time. The feature matching results of each cached time are aggregated into task-object association features (i.e., the entity association features mentioned above), which can be represented as F. asso Furthermore, entity association features are injected as additional conditions into the action expert network to enhance the expressive power of the action sequence or predicted action output by the action expert network.

[0112] It can be argued that by introducing entity association features, when the target task description mentions a certain object (such as "put the pen on the table into the pen holder"), even if the object is not significant in the current observation or has been occluded, the action prediction model can still recall the cached moment when the object was last clearly observed in short-term memory and incorporate its spatial location information as a priori into action planning, thereby avoiding localization failure caused by visual field occlusion or object removal.

[0113] This application can be considered to improve the accuracy of predicted actions output by the enhanced action expert network by introducing a sequence key perception mechanism to generate fusion weights for cached image features, dynamic semantic masks to train the action prediction model, and fusion weights for cached joint angles and / or entity association features through cross-granular semantic association. These three mechanisms work synergistically, enabling the action prediction model to achieve three key leaps in capabilities under short-term memory sequence conditions: it can automatically identify and utilize sub-information from a few key cached moments that are most valuable to the current task, rather than indiscriminately processing sub-information from all historical moments; furthermore, even in cases of partial observation loss or text description variations, the output target fusion features maintain a stable visual-linguistic semantic correspondence, achieving robust alignment; and it can proactively retrieve semantically relevant historical information of objects in the target task description requested by the current task from short-term memory, compensating for the perceptual blind spots of the current observation. Ultimately, this allows the action prediction model to establish a faster mapping relationship between task semantics and action space when facing complex task operation scenarios, outputting effective action sequences (i.e., predicted actions) that are highly adapted to the current context, directly improving the success rate and generalization ability of task execution from the action prediction model level.

[0114] For example, such as Figure 5a The implementation framework shown addresses the issues of poor planning ability and repeated failures due to lack of memory. This application employs a long-term memory module and a short-term memory cache to solve these problems. Based on the long-term memory module, the embodied intelligent device can manage the entire lifecycle of task planning, execution, and completion, accumulating common-sense and experiential conclusions to assist task completion. Based on the short-term memory module, the embodied intelligent device can summarize the experience of current task execution details (e.g., if turning a knob to the right fails to open a door, try turning it to the left), improving the success rate of current task execution. The ability to perceive memory should be learned by an action prediction model, enabling the rational use of memory features and the output of action sequences with high success rates. Regarding the problem of weak spatial awareness, this application... Figure 3 The spatial perception network shown is based on dual-view construction of pseudo-3D semantic features that mimic binocular vision, thereby enhancing perception capabilities and improving the robustness of visual features.

[0115] The short-term memory module is responsible for remembering and providing operational context information within a short time window, used to guide the generation of historical cache information associated with the current image when the embodied intelligent device performs tasks. At the current time t0, the system receives multi-source input: two second current images from different perspectives captured by the left and right cameras for the same task scene. After processing each second current image through a spatial perception network, pseudo-3D semantic features are obtained, i.e., the aforementioned second image features. The first current image captured by the wrist camera (usually one for a single arm, two for dual arms) is processed by an image encoder (i.e., an image embedding model) to generate end effector scene features F1, i.e., the aforementioned first image features. The first current image is as follows: Figure 5b As shown, the first current image is a real-time image acquired by the end effector, wherein... Figure 5b In the image content, the bottom region B1 represents the end effector of the embodied intelligent device, used for contact with the task object. For example, the first current image was captured by a wrist camera. The target task description is then processed by a text encoder (i.e., a text embedding model) to generate the current text features F2.

[0116] Inside the short-term memory module, the current text features are fed together with the short-term memory cache features into the first large model (VLM) to extract enhanced multimodal image-text features, which are then fed into the action expert network to generate the action sequence of the next s timestamps, i.e., the predicted action, where s is the preset prediction step size.

[0117] The information cached by the short-term memory module consists of image feature sequences and corresponding joint sequences traced back from the current moment to n previous moments. The image feature sequences include end effector scene features and pseudo-3D semantic features, while the joint sequences are the joint angles of the embodied intelligent device at each cached moment. n is a time window hyperparameter that controls the length of the history that the short-term memory focuses on, and is preset by algorithm developers based on task characteristics.

[0118] The above scheme decouples semantic understanding and action decision-making while allowing them to work collaboratively by inputting current image features, text features, and historical cache information related to the current task request into an action prediction model that includes a first-level model and an action expert network. By introducing historical image features from the historical cache information, the action prediction model can obtain temporal context and state evolution information during task execution. Thus, when faced with object occlusion or positional movement, it can infer the current object state of the task object related to the current task request based on historical observations, improving the robustness of visual perception in dynamic environments. By introducing historical joint angles, the action prediction model can perceive the motion trajectory and physical state of the embodied intelligent device itself, which is beneficial for the continuity and smoothness of the subsequently generated predicted actions. This avoids mechanical shocks or loss of control caused by state jumps of the embodied intelligent device, thereby improving the action prediction accuracy and the success rate of current task request execution of the embodied intelligent device under complex tasks.

[0119] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of a control device for an embodied intelligent device according to an embodiment of this application. The control device 60 of the embodied intelligent device includes an extraction module 61, an input module 62, a first processing module 63, a second processing module 64, and a control module 65. The extraction module 61 is used to extract features from the acquired current image and the target task description of the current task request in response to receiving a current task request, to obtain current image features and current text features. The input module 62 is used to input the current image features, current text features, and historical cache information associated with the current image into an action prediction model. The first processing module 63 is used to process the current image features, current text features, and historical image features through a first large model to obtain target fusion information output by the first large model. The second processing module 64 is used to process the target fusion information and historical joint angles through an action expert network to obtain the predicted action of the current task request. The control module 65 is used to control the embodied intelligent device to execute the current task request with the predicted action.

[0120] For details on the functions performed by each module, please refer to the control methods for embodied intelligent devices; these details will not be repeated here.

[0121] The above scheme decouples semantic understanding and action decision-making while allowing them to work collaboratively by inputting current image features, text features, and historical cache information related to the current task request into an action prediction model that includes a first-level model and an action expert network. By introducing historical image features from the historical cache information, the action prediction model can obtain temporal context and state evolution information during task execution. Thus, when faced with object occlusion or positional movement, it can infer the current object state of the task object related to the current task request based on historical observations, improving the robustness of visual perception in dynamic environments. By introducing historical joint angles, the action prediction model can perceive the motion trajectory and physical state of the embodied intelligent device itself, which is beneficial for the continuity and smoothness of the subsequently generated predicted actions. This avoids mechanical shocks or loss of control caused by state jumps of the embodied intelligent device, thereby improving the action prediction accuracy and the success rate of current task request execution of the embodied intelligent device under complex tasks.

[0122] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of an embodiment of the electronic device of this application. The electronic device 70 includes a memory 71 and a processor 72. The processor 72 is used to execute program instructions stored in the memory 71 to implement the steps in the control method embodiment of the above-described intelligent device. In a specific implementation scenario, the electronic device 70 may include, but is not limited to, a microcomputer or a server. In addition, the electronic device 70 may also include mobile devices such as laptops and tablets, which are not limited here.

[0123] Specifically, processor 72 controls itself and memory 71 to implement the steps in the control method embodiment of the aforementioned embodied intelligent device. Processor 72 can also be referred to as a CPU (Central Processing Unit). Processor 72 may be an integrated circuit chip with signal processing capabilities. Processor 72 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 72 can be implemented using integrated circuit chips.

[0124] Please see Figure 8 , Figure 8This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present application. The computer-readable storage medium 80 stores program instructions 801 thereon, which, when executed by a processor, implement the steps in any of the above-described embodiments of the control method for a personal intelligent device.

[0125] In some embodiments, the functions or modules of the apparatus provided in the embodiments disclosed in this application can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0126] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0127] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0128] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0129] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0130] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

Claims

1. A control method for an embodied intelligent device, characterized in that, The control method for the embodied intelligent device includes: In response to receiving a current task request, feature extraction is performed on the acquired current image and the target task description of the current task request to obtain current image features and current text features; The current image features, the current text features, and historical cache information associated with the current image are input into the action prediction model. The action prediction model includes a first large model and an action expert network. The historical cache information associated with the current image includes historical image features of historical images and historical joint angles of the embody intelligent device. The first large model processes the current image features, the current text features, and the historical image features to obtain the target fusion information output by the first large model; The motion expert network processes the target fusion information and the historical joint angles to obtain the predicted motion for the current task request. Control the embodied intelligent device to execute the current task request with the predicted action.

2. The control method for a embodied intelligent device according to claim 1, characterized in that, The historical cache information associated with the current image includes sub-information of several cache times, the historical image features include cache image features of cached images at the several cache times, the historical joint angles include cache joint angles at the several cache times, the sub-information includes cached image features and cache joint angles belonging to the same cache time, and the target fusion information includes target fusion features that characterize the feature fusion result between the current image features, the current text features, and the historical image features. The step of processing the current image features, the current text features, and the historical image features using the first large model to obtain the target fusion information output by the first large model includes: The current image features and the current text features are fused to obtain initial fused features; Based on the current text features and the cached image features associated with the several cached times, the fusion weight of each cached image feature is determined respectively; The target fusion feature is determined based on the initial fusion feature, the cached image features associated with each cache time step, and the fusion weights of the corresponding cached image features; The step of processing the target fusion information and the historical joint angles through the action expert network to obtain the predicted action of the current task request includes: obtaining the predicted action based on the target fusion features and the cached joint angles at several cached times.

3. The control method for a embodied intelligent device according to claim 2, characterized in that, The target fusion information also includes entity matching results between cached image features at each cache time and entity objects in the target task description; After the step of fusing the current image features and the current text features to obtain initial fused features, the control method of the embodied intelligent device further includes: Key information is extracted from the current text features to obtain the entity semantic features describing the target task; Based on the entity semantic features and the cached image features at each cache time, the entity matching result is determined; The step of obtaining the predicted action based on the target fusion features and the cache joint angles at the plurality of cache times includes: obtaining the predicted action based on the target fusion features, the cache joint angles at the plurality of cache times, and the entity matching result.

4. The control method for the embodied intelligent device according to claim 3, characterized in that, The entity matching result includes entity association features between entity objects in the target task description and cached image features at several cached times, and / or fusion weights of cached joint angles at each cached time. The fusion weights of cached joint angles characterize the degree of association between entity objects in the target task description and cached joint angles. The step of determining the entity matching result based on the entity semantic features and the cached image features at each cache time includes: The entity semantic features are matched with each sub-feature in the cached image features at each cache time to obtain the feature matching results for the corresponding cache time. Based on the feature matching results at each cache time, determine the fusion weights of the entity association features and / or each cache joint angle; The step of obtaining the predicted action based on the target fusion features, the cache joint angles at several cache times, and the entity matching results includes: obtaining the predicted action based on the target fusion features, the cache joint angles at each cache time, the entity association features, and / or the fusion weights of each cache joint angle.

5. The control method for a embodied intelligent device according to claim 4, characterized in that, The feature matching result at the cache time includes the matching degree between the cached image features at the cache time and the entity semantic features; The step of determining the fusion weights of the entity association features and / or the angles of each cache joint based on the feature matching results at each cache time includes: Based on the matching degree between the cached image features and the entity semantic features at each cache time, the fusion weight of the cached joint angle at the corresponding cache time is determined; and / or, In response to the feature matching result at at least one cached time, the entity semantic features are successfully matched with at least some sub-features in the cached image features at the corresponding cached time. The spatial locations of the successfully matched sub-features in the cached image features at the at least one cached time are combined to obtain the target spatial location, and the at least one cached time is combined to obtain the target time. The target spatial location and the target time are aggregated to obtain the entity association features.

6. The control method for a embodied intelligent device according to claim 2, characterized in that, The step of determining the fusion weight of each cached image feature based on the current text features and the cached image features associated with the several cached times includes: The similarity between the cached image features at each cache time and the current text features is used as the similarity associated with the cached image features at the corresponding cache time. The sum of the similarities associated with the features of each cached image is used as the target similarity. The fusion weights of the corresponding cached image features are obtained based on the similarity between the features associated with each cached image and the target similarity.

7. The control method for a embodied intelligent device according to claim 1, characterized in that, Before the step of inputting the current image features, the current text features, and historical cache information associated with the current image into the action prediction model, the control method of the embodied intelligent device further includes: Acquire a first historical image and a second historical image of the current image. The acquisition time of the first historical image is a first moment, and the acquisition time of the second historical image is a second moment. The first moment is earlier than the second moment, and the second moment is earlier than the acquisition time of the current image. In response to the total number of cached information in the historical cache information associated with the second historical image being greater than a threshold, the historical cache information associated with the second historical image is updated according to the state information of the embodied smart device at the first moment and the image features of the second historical image, so as to obtain the historical cache information associated with the current image.

8. The control method for a embodied intelligent device according to claim 7, characterized in that, The status information includes the image compression features of the first historical image and the historical actions of the embodied intelligent device at the first moment; The step of updating the historical cache information associated with the second historical image based on the state information of the embodied smart device at the first moment and the image features of the second historical image to obtain the historical cache information associated with the current image includes: The image compression features of the first historical image and the historical actions are used for prediction processing to obtain the potential state information of the embodied intelligent device at the second moment. The potential state information is decoded to obtain the predicted image features of the embodied intelligent device at the second moment; Based on the feature differences between the predicted image features and the image features of the second historical image, the prediction error associated with the second historical image is determined; Based on the prediction error associated with the second historical image, the historical cache information associated with the second historical image is updated to obtain the historical cache information associated with the current image.

9. The control method for a embodied intelligent device according to claim 8, characterized in that, The historical cache information associated with the second historical image includes sub-information of several candidate moments. The sub-information includes candidate image features and candidate joint angles of candidate images belonging to the same candidate moment. The acquisition time of the candidate image is earlier than the second moment. The step of updating the historical cache information associated with the second historical image based on the prediction error associated with the second historical image to obtain the historical cache information associated with the current image includes: An error threshold is determined based on the prediction error associated with the candidate images at the aforementioned candidate times and the prediction error associated with the second historical image. In response to the prediction error associated with the second historical image being greater than or equal to the error threshold, one of the sub-information of the plurality of candidate moments is removed to obtain the historical cache information after removal. The image features of the second historical image are correlated with the joint angle of the embodied intelligent device at the second moment to obtain the sub-information of the second moment; The sub-information at the second moment is stored in the historical cache information after the removal process to obtain the historical cache information associated with the current image.

10. The control method for a embodied intelligent device according to claim 1, characterized in that, The current task request includes an initial task description; Before the step of extracting features from the acquired current image and the target task description of the current task request to obtain current image features and current text features, the control method of the embodied intelligent device further includes: Obtain the historical task description associated with the current image, wherein the historical task description was generated earlier than the time the current image was acquired; Based on the historical task description, a task plan for the historical task description is generated; The second major model processes the initial task description in the current task request, the task planning of the historical task description, and the current image to obtain the target task description output by the second major model.

11. The control method for a embodied intelligent device according to claim 1, characterized in that, The embodied intelligent device includes an end effector for contacting the task object, the current image includes a first current image acquired at the time of acquisition of the current image towards the end effector and a second current image acquired from at least two perspectives towards the task environment to which the current task request belongs, and the current image features include a first image feature corresponding to the first current image and a second image feature obtained based on the second current image from each perspective. The step of extracting features from the acquired current image and the target task description of the current task request to obtain current image features and current text features includes: The first current image is subjected to image encoding processing to obtain the first image features; The second current image from each viewpoint is input into the depth estimation network to obtain the spatial features of the target object in the task environment output by the depth estimation network. The second current image of one of the at least two perspectives is input into the semantic alignment network and the detection network respectively to obtain the semantic features of the target object in the task environment output by the semantic alignment network and the geometric features of the target object in the task environment output by the detection network. The second image feature is determined based on the spatial features, the semantic features, and the geometric features.

12. The control method for a embodied intelligent device according to claim 1, characterized in that, The control method for the embodied intelligent device further includes a training step for the action prediction model, the training step including: Obtain information about the sample task request, the sample image features associated with the sample image, the sample text features, the historical cache information associated with the sample time, and the actual action of the sample task request; The historical cache information associated with the sample image is masked to obtain the masked historical cache information; The sample text features are masked to obtain the masked sample text features. The action prediction model processes the sample image features, the masked sample text features, and the masked historical cache information to obtain the sample action of the sample task request output by the action prediction model. The action prediction model is trained based on the difference between the sample action and the real action until a fully trained action prediction model is obtained.

13. An electronic device, characterized in that, include: A memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to execute the control method of the embodied intelligent device as described in any one of claims 1-12.

14. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they are used to implement the control method of the embodied intelligent device as described in any one of claims 1-12.