Three-dimensional embodied intelligent building control method fusing visual and language information

CN122528689BActive Publication Date: 2026-09-15SHENZHEN SENSING DATA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611011824.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-15
Estimated Expiration
2046-07-08

AI Technical Summary

Technical Problem

然而,现有技术普遍存在仅依赖二维视觉信息、缺乏三维空间理解能力以及仿真与现实脱节的问题,会导致三维空间约束缺失、施工目标定位与姿态估计误差累积、路径规划在复杂结构环境中不稳定、仿真结果难以迁移到真实施工现场,从而降低自动化施工的安全性与可靠性,并增加误操作与碰撞风险

Benefits of technology

(1)本发明引入三维的空间提示词元序列注入机制,通过三维适配器根据任务指令和三维语义世界模型生成空间提示词元序列T3D。再将T3D与ViT视觉编码器得到的图像词元序列Timg、Transformer语言编码器得到的语言词元序列Tlang,通过跨模态Transformer融合模块记性融合生成到融合特征Atoken,可以有效提升机器人对空间结构与任务语义的联合理解能力。这种方法不仅改善了复杂场景下的感知质量,还确保了三维空间推理的一致性与精度,为建筑机器人自主决策提供了更强的数据支持。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122528689B_ABST
    Figure CN122528689B_ABST
Patent Text Reader

Abstract

The application discloses a three-dimensional embodied intelligent building control method fusing visual and language information, and belongs to the technical field of computer models, and comprises the following steps: constructing samples corresponding to construction data of an embodied intelligent agent and task instructions; generating a three-dimensional semantic world model corresponding to the samples based on an RGB-D graph and a BIM model; constructing an action generation model; obtaining a specific intelligent action control model through online verification of digital twinning and closed-loop update of a three-dimensional error field, and the specific intelligent action control model is used for generating a skill parameter sequence according to the task instruction. Through the three-dimensional spatial prompt word sequence, the BIM prior and the RGB-D perception fusion method, the digital twinning simulation verification mechanism and the like, the application can significantly improve the understanding ability and execution efficiency of a robot on a complex construction scene, and also ensures the spatial positioning precision and operation consistency, and further promotes the landing application of intelligent construction technology in the building industry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer modeling technology, and in particular to a three-dimensional embodied intelligent building control method that integrates visual and linguistic information. Background Technology

[0002] With the rapid development of the construction industry and 3D intelligent construction technology, the vision-language-action method based on RGB-D perception and 3D semantic modeling has shown great potential in fields such as intelligent construction robot control, digital twin simulation, and building operation and maintenance automation.

[0003] In the field of construction, ensuring the accuracy of three-dimensional spatial data and semantic information is crucial for high-precision assembly, path planning, and safety control. Practical applications typically require high-precision positioning, robust perception, and adaptability to complex environments to meet the needs of construction automation and safe production. However, existing technologies generally rely solely on two-dimensional visual information, lack three-dimensional spatial understanding capabilities, and suffer from a disconnect between simulation and reality. This leads to a lack of three-dimensional spatial constraints, accumulated errors in construction target positioning and attitude estimation, instability in path planning in complex structural environments, and difficulty in transferring simulation results to real construction sites. Consequently, the safety and reliability of automated construction are reduced, and the risks of misoperation and collisions are increased. Summary of the Invention

[0004] The purpose of this invention is to provide a three-dimensional embodied intelligent building control method that integrates visual and linguistic information, which not only takes into account real-time perception and design prior knowledge, but also accurately models key elements such as construction structure and constrained areas, and significantly improves the effect of automated building operations.

[0005] To achieve the above objectives, the technical solution adopted by this invention is as follows: a three-dimensional embodied intelligent building control method integrating visual and linguistic information, comprising the following steps: S1, construct dataset D, including S11~S12; S11, construct samples corresponding to task instructions based on the construction data of the embodied intelligent agent, and combine all samples into a dataset D, where the m-th sample D in D is... m ={L m RGB-D m PC m BIM m A m ,K m}, L m For task instructions, RGB-D m PC m BIM m Execute L respectively mPrevious construction site RGB-D drawings, 3D point clouds and BIM models, A m K m Execute L respectively m The set of actions and skill parameters at that time; in, , , For A m The b-th three-dimensional semantic action, where B is... Total, a b o b p b R b C b These are the action type, the target instance corresponding to the action, the optimal point of action on the target instance, the action's area of ​​effect, and the action's constraints. , for Reference skill parameters; S12, Obtain the initial state S1 and the state S after executing the b-th action of the embodied agent. b+1 ; S2, based on RGB-D m and BIM m Generate a 3D semantic world model W of the construction site m ; S3 constructs an action generation model, which includes a ViT visual encoder, a Transformer language encoder, a cross-modal Transformer fusion module, an action decoder, a 3D adapter, a 3D action parser, and a skill mapping layer. The first four constitute the basic model, and the latter three constitute the training model. The ViT visual encoder is used to encode RGB images into image word sequences T. img and based on T img Generate semantic feature maps; The Transformer language encoder is used to encode task instructions into a sequence of language lexical units T. lang And attention converges into a semantic vector z L ; 3D adapter is used to adapt to L m and W m Generate the three-dimensional semantic features Z of the construction area 3D And mapped to a spatial cue lexical sequence T 3D ; The cross-modal Transformer fusion module will integrate T img T lang T 3D Cross-modal fusion is performed to obtain fused feature A token ; The action decoder is used to determine the fusion feature A token Output B predicted action types; The 3D action parser is used to generate the 3D semantic action for each predicted action type. The 3D semantic action for the b-th predicted action type is... , a, o, p*, R*, and C represent a b o b p b R b C b The predicted value; The skill mapping layer is used to... The robot's current state q and C b Mapped to the predicted skill parameters of the embodied intelligent agent Where xyz is 3D coordinates F and v represent the target posture, force, and velocity of the embodied intelligent agent performing the action, and T represents the duration of the action; S4, based on online verification of digital twins and closed-loop update training of three-dimensional error field, the action generation model is used as the specific intelligent action control model after the update is completed; S5 uses a specific intelligent motion control model to generate a sequence of skill parameters based on task instructions.

[0006] Preferably, S2 includes S21 to S24; S21, Obtain RGB-D m This includes RGB images and depth maps. The ViT visual encoder is used to extract semantic feature maps from the RGB images, and the semantic features of each pixel location are mapped to a PC by combining the depth map. m The corresponding spatial point, where the semantic features of spatial point i are: ; S22, based on BIM m Generate the BIM prior feature vector of spatial point i , It contains the BIM component category, BIM geometric parameters, material properties, construction status, and process constraints corresponding to spatial point i; S23, will and Adaptive fusion into fused semantic features f of spatial point i i ; S24, Generate a 3D semantic world model W of the construction site. m , In the formula, N is the total number of spatial points, and dc i s i o iThese represent the spatial coordinates, semantic category, and target instance of spatial point i, respectively.

[0007] As a preferred option, in S22 and S23, , , Among them, Enc BIM (∙) represents the BIM feature encoder, which consists of a category embedding layer and a multilayer perceptron. , , , , These represent the BIM component category, BIM geometric parameters, material properties, construction status, and process constraints corresponding to spatial point i, respectively, and α. i for The weight.

[0008] Preferably, the 3D adapter includes a PointNet++ encoder, a sparse 3D convolutional network, a cross-attention filtering module, and a mapping layer; The PointNet++ encoder is used as input to W and outputs multi-scale local features W1 of W. The sparse 3D convolutional network is used as input W1 and outputs sparse 3D features W2. The cross-attention filtering module is used for semantic vector z L W2 serves as the query matrix, and W2 serves as both the key and value matrix. Z is generated based on a cross-attention mechanism. 3D ; The mapping layer is used to convert Z... 3D Mapping to a preset dimension yields spatial cue words.

[0009] Preferably, the action constraints include BIM construction accuracy constraints, robot joint constraints, safety distance constraints, and no-operation zone constraints; the skills include grasping skills, handling skills, placement skills, alignment skills, insertion skills, clamping skills, obstacle avoidance movement skills, drilling and positioning skills, scanning and detection skills, compliant contact skills, tightening and fixing skills, and pushing and adjusting skills.

[0010] Preferably, S4 includes S41 to S43; S41, on the simulation platform, press Simulation is performed, and the embodied intelligent agent executes K b After that, the state changed from become ; S42, Construct the total training loss L all ; , In the formula, Lspace For the loss of three-dimensional spatial positioning error, L sim The simulation consistency loss is used to constrain the consistency between the simulated predicted state and the actual execution state; L skill For skill generation loss, γ1 and γ2 are respectively L sim L skill The weights;

[0011] S43, based on dataset D, to minimize L all Update the parameters of the training model, and after the update is complete, construct the specific intelligent action control model from the action generation model.

[0012] As a preferred option, in S42, L space L sim L skill Calculate according to the following formulas respectively: In the formula, The squared distance is the L2 norm.

[0013] As a preferred option, in S43, the training model is updated by guiding the model to adjust parameters, including Sa1~Sa3, based on the three-dimensional error field and error categories; Sa1, computation of embodied intelligent agent execution K b The state error e after b , A three-dimensional error field is generated based on M sampling points, and the three-dimensional error of any spatial point p in the three-dimensional error field is E(p). , , In the formula, Assign weights to the error of the j-th sampling point. Let d be the spatial weight of the j-th sampling point to the spatial point p. j σ is the distance from the j-th sampling point to p*, σ is the distance attenuation coefficient, and Softmax(⋅) is the Softmax function; Sa2 predefines four types of errors and their corresponding thresholds, and represents the optimal deviation of the point of application, D. loc BIM and site inconsistency deviation D bim State deviation D sim Trajectory tracking or force control instability deviation D ctrl ; , In the formula, It is a 2-norm. This represents the local structure of the BIM model within a three-dimensional error field. Let t be the distance from point i in space to B, and t be the distance from point i to point B. bTotal duration, q t , These represent the robot's simulated posture at time t and its reference posture in the sample, respectively. Sa3 calculates four types of errors within E(p); If D loc If the threshold is exceeded, adjust the network parameters of the 3D adapter and the 3D motion parser. If D bim If the threshold is exceeded, the BIM model will be adjusted or an alarm will be triggered. If D sim Or D ctrl If the threshold is exceeded, the skill mapping layer will be adjusted.

[0014] Compared with the prior art, the advantages of the present invention are as follows: (1) This invention introduces a three-dimensional spatial cue word sequence injection mechanism, which generates a spatial cue word sequence T based on task instructions and a three-dimensional semantic world model through a three-dimensional adapter. 3D Then put T 3D The image word sequence T obtained by the ViT visual encoder img The language lexical sequence T obtained by the Transformer language encoder lang The fused feature A is generated through cross-modal Transformer fusion module. token This approach can effectively enhance a robot's ability to jointly understand spatial structure and task semantics. It not only improves perception quality in complex scenarios but also ensures the consistency and accuracy of 3D spatial reasoning, providing stronger data support for autonomous decision-making in construction robots.

[0015] (2) Introducing the BIM prior and RGB-D perception fusion method as a key link in the 3D modeling process improves the stability and reliability of scene understanding. This method can not only take into account real-time perception and design prior, but also accurately model key elements such as construction structure and constrained areas, significantly improving the effect of building automation operations, especially significantly improving the engineering application value of the system in complex construction environments.

[0016] (3) Introduce a digital twin online verification and a three-dimensional error field closed-loop update mechanism. Due to the high cost of trial and error in construction, the system does not directly execute the generated skills, but first performs online verification in the digital twin environment, reducing the collision risk and execution failure rate to a controllable range before execution, which significantly improves the safety and reliability of construction robots in complex construction environments.

[0017] (4) The present invention divides the action generation model into a basic model and a training model. During training, the basic model is frozen and only the network parameters of the training model are updated. This method can achieve three-dimensional capability expansion with only a few new modules, reducing the training cost by more than 80%, while maintaining the general semantic understanding capability of the original model.

[0018] In summary, this invention, through mechanisms such as three-dimensional spatial cue word sequences, BIM prior and RGB-D perception fusion method, and digital twin simulation verification, can significantly improve the robot's understanding of complex construction scenarios and execution efficiency, while also ensuring spatial positioning accuracy and operational consistency, thereby promoting the application of intelligent construction technology in the construction industry. Attached Figure Description

[0019] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0020] The present invention will be further described below with reference to the embodiments and accompanying drawings.

[0021] Example 1: See Figure 1 A three-dimensional embodied intelligent building control method integrating visual and linguistic information includes the following steps: S1, construct dataset D, including S11~S12; S11, construct samples corresponding to task instructions based on the construction data of the embodied intelligent agent, and combine all samples into a dataset D, where the m-th sample D in D is... m ={L m RGB-D m PC m BIM m A m ,K m}, L m For task instructions, RGB-D m PC m BIM m Execute L respectively m Previous construction site RGB-D drawings, 3D point clouds and BIM models, A m K m Execute L respectively m The set of actions and skill parameters at that time; in, , , For A m The b-th three-dimensional semantic action, where B is... Total, a b o b p b R bC b These are the action type, the target instance corresponding to the action, the optimal point of action on the target instance, the action's area of ​​effect, and the action's constraints. , for Reference skill parameters; S12, Obtain the initial state S1 and the state S after executing the b-th action of the embodied agent. b+1 ; S2, based on RGB-D m and BIM m Generate a 3D semantic world model W of the construction site m ; S3 constructs an action generation model, which includes a ViT visual encoder, a Transformer language encoder, a cross-modal Transformer fusion module, an action decoder, a 3D adapter, a 3D action parser, and a skill mapping layer. The first four constitute the basic model, and the latter three constitute the training model. The ViT visual encoder is used to encode RGB images into image word sequences T. img and based on T img Generate semantic feature maps; The Transformer language encoder is used to encode task instructions into a sequence of language lexical units T. lang And attention converges into a semantic vector z L ; 3D adapter is used to adapt to L m and W m Generate the three-dimensional semantic features Z of the construction area 3D And mapped to a spatial cue lexical sequence T 3D ; The cross-modal Transformer fusion module will integrate T img T lang T 3D Cross-modal fusion is performed to obtain fused feature A token ; The action decoder is used to determine the fusion feature A token Output B predicted action types; The 3D action parser is used to generate the 3D semantic action for each predicted action type. The 3D semantic action for the b-th predicted action type is... , a, o, p*, R*, and C represent a b o b p b R b C b The predicted value; The skill mapping layer is used to... The robot's current state q and C b Mapped to the predicted skill parameters of the embodied intelligent agent Where xyz is 3D coordinates F and v represent the target posture, force, and velocity of the embodied intelligent agent performing the action, and T represents the duration of the action; S4, based on online verification of digital twins and closed-loop update training of three-dimensional error field, the action generation model is used as the specific intelligent action control model after the update is completed; S5 uses a specific intelligent motion control model to generate a sequence of skill parameters based on task instructions.

[0022] In this embodiment, S2 includes S21 to S24; S21, Obtain RGB-D m This includes RGB images and depth maps. The ViT visual encoder is used to extract semantic feature maps from the RGB images, and the semantic features of each pixel location are mapped to a PC by combining the depth map. m The corresponding spatial point, where the semantic features of spatial point i are: ; S22, based on BIM m Generate the BIM prior feature vector of spatial point i , It contains the BIM component category, BIM geometric parameters, material properties, construction status, and process constraints corresponding to spatial point i. Enc BIM (∙) represents the BIM feature encoder, which consists of a category embedding layer and a multilayer perceptron. , , , , These are the BIM component category, BIM geometric parameters, material properties, construction status, and process constraints corresponding to spatial point i, respectively. S23, will and Adaptive fusion into fused semantic features f of spatial point i i , α i for The weights; S24, Generate a 3D semantic world model W of the construction site. m , In the formula, N is the total number of spatial points, and dc i s i o i These represent the spatial coordinates, semantic category, and target instance of spatial point i, respectively.

[0023] The 3D adapter includes a PointNet++ encoder, a sparse 3D convolutional network, a cross-attention filtering module, and a mapping layer. The PointNet++ encoder takes W as input and outputs multi-scale local features W1 of W. The sparse 3D convolutional network takes W1 as input and outputs sparse 3D features W2. The cross-attention filtering module is used for semantic vector z. L W2 serves as the query matrix, and W2 serves as both the key and value matrix. Z is generated based on a cross-attention mechanism. 3D The mapping layer is used to map Z... 3D Mapping to a preset dimension yields spatial cue words.

[0024] The action constraints include BIM construction accuracy constraints, robot joint constraints, safety distance constraints, and no-operation zone constraints; the skills include grasping skills, handling skills, placement skills, alignment skills, insertion skills, clamping skills, obstacle avoidance movement skills, drilling and positioning skills, scanning and detection skills, compliant contact skills, tightening and fixing skills, and pushing and adjusting skills.

[0025] S4 includes S41~S43; S41, on the simulation platform, press Simulation is performed, and the embodied intelligent agent executes K b After that, the state changed from become ; S42, Construct the total training loss L all ; , , In the formula, L space For the loss of three-dimensional spatial positioning error, L sim The simulation consistency loss is used to constrain the consistency between the simulated predicted state and the actual execution state; L skill For skill generation loss, γ1 and γ2 are respectively L sim L skill The weights; The squared distance is the norm 2. S43, based on dataset D, to minimize L all Update the parameters of the training model. After the update is completed, the action generation model is used to form a specific intelligent action control model. Specifically, updating the training model involves adjusting the parameters of the model based on the three-dimensional error field and error categories, including Sa1~Sa3. Sa1, computation of embodied intelligent agent execution K b The state error e after b , A three-dimensional error field is generated based on M sampling points, and the three-dimensional error of any spatial point p in the three-dimensional error field is E(p). , In the formula, Assign weights to the error of the j-th sampling point. Let d be the spatial weight of the j-th sampling point to the spatial point p. j σ is the distance from the j-th sampling point to p*, σ is the distance attenuation coefficient, and Softmax(⋅) is the Softmax function; Sa2 predefines four types of errors and their corresponding thresholds, and represents the optimal deviation of the point of application, D. loc BIM and site inconsistency deviation D bim State deviation D sim Trajectory tracking or force control instability deviation D ctrl ; , In the formula, It is a 2-norm. This represents the local structure of the BIM model within a three-dimensional error field. Let t be the distance from point i in space to B, and t be the distance from point i to point B. b Total duration, q t , These represent the robot's simulated posture at time t and its reference posture in the sample, respectively. Sa3 calculates four types of errors within E(p); If D loc If the threshold is exceeded, adjust the network parameters of the 3D adapter and 3D motion parser; if D bim If the threshold is exceeded, adjust the BIM model or trigger an alarm; if D sim Or D ctrl If the threshold is exceeded, the skill mapping layer will be adjusted.

[0026] The word sequences described in this invention are all token sequences.

[0027] Example 2: A three-dimensional embodied intelligent building control method integrating visual and linguistic information, comprising the following steps: S1. Construct dataset D, following the same steps as step S1 in Example 1. Taking the task instruction "The robot installs the window frame into the specified reserved opening" as an example, each task instruction can correspond to one or more samples. The first type is the window frame installation simulation data automatically generated in Isaac Sim based on the BIM model, including the RGB-D map of the construction site, 3D point cloud, BIM model, and the action set and skill parameter set when executing the "window frame installation" task instruction. The second type is a small amount of manual teaching data, such as the trajectory of a human remotely operating embodied intelligent agent to complete "aligning the opening, low-speed insertion, and pressing and positioning". The third type is real construction site trial operation data, used to correct the differences between simulation and real environment. Assuming that for this task instruction, we use the first type of data to obtain a sample D. m ={L m RGB-D m PC m BIM m A m ,K m}, then execute L m The action set can be broken down into three actions: "alignment, insertion, and compression." Each action corresponds to a reference skill parameter. The reference skill parameter has the same structure as the predicted skill parameter, except that the values ​​of each element in the predicted skill parameter are the predicted values ​​of the corresponding elements in the reference skill parameter. After each action is executed, the state of the embodied agent can be obtained. Since it contains three actions, the sample can be divided into four states: the initial state S1, and states S2 to S4 after executing "alignment, insertion, and compression" respectively.

[0028] S2, generates D m The corresponding three-dimensional semantic world model includes S21~S24; S21, the semantic features of generating spatial point i are: Specifically, RGB-D images of the construction site are acquired, including RGB images and depth maps. The RGB images are used to identify window frames, walls, reserved openings, and scaffolding, while the depth maps are used to obtain the 3D positions of these objects. The ViT visual encoder extracts semantic feature maps from the RGB images and combines them with the depth maps to map the semantic features of each pixel position to a PC. m The corresponding spatial point, where the semantic features of spatial point i are: This can generate semantically meaningful 3D point clouds, such as one point cloud being identified as the "edge of a reserved opening" and another point cloud being identified as the "window frame boundary". S22, Generate the BIM prior feature vector of spatial point i. The data, including BIM component categories, BIM geometric parameters, material properties, construction status, and technological constraints, is extracted from the BIM model. In this embodiment, the design coordinates, opening dimensions, wall thickness, installation depth, and allowable errors of the opening are encoded into vectors. The BIM model has been aligned with the site coordinate system using a total station. S23, will and Adaptive fusion into fused semantic features f i ; S24, Generate a 3D semantic world model W of the construction site. m .

[0029] S3 constructs an action generation model, including a ViT visual encoder, a Transformer language encoder, a cross-modal Transformer fusion module, an action decoder, a 3D adapter, a 3D action parser, and a skill mapping layer; The ViT visual encoder, Transformer language encoder, and 3D adapter respectively output T img T lang T 3D The cross-modal Transformer fusion module is based on T img T lang T 3D Generate fusion feature A token The motion decoder will A token Decoding is performed into three predicted action types: alignment, insertion, and compaction; a 3D action parser generates a 3D semantic action for each predicted action. ~ The skill mapping layer then... The robot's current state q and C b Mapped to the predicted skill parameters of the embodied intelligent agent Taking the "insert" action as an example, during the execution of the "insert" action: The system first determines the three-dimensional coordinates (x, y, z) of the predicted value p* of the optimal action point on the target instance based on the center coordinates of the opening; The target posture Θ for performing the action is generated by combining the normal constraint of the opening. The normal constraint of the opening is such that the normal of the window frame is consistent with the normal of the opening. The target posture Θ is specifically the posture of the end effector such as the robotic arm after performing the action. Subsequently, based on the spatial and attitude information, a phased motion trajectory was planned, including the processes of approaching, attitude correction, low-speed insertion, and clamping positioning; the resulting motion trajectory was: approaching the opening → attitude correction → low-speed insertion → clamping positioning; Decompose the motion trajectory and dynamically adjust the speed v of the action according to different stages, such as medium speed operation in the "approaching the hole" stage and low speed operation in the "low speed insertion" stage. Analysis of force control parameters: Force feedback limits the force F that performs the action to a set threshold to ensure structural safety; at the same time, the duration T of the action is determined according to the overall task plan to control the execution rhythm; Set termination conditions: such as the position error, attitude error and insertion depth all satisfying the BIM constraints.

[0030] Based on the above analysis, the predicted skill parameters corresponding to the "insertion" action are finally generated. The analysis of the other actions is similar.

[0031] S4 and S5 are the same as steps S4 and S5 in Example 1.

[0032] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A three-dimensional embodied intelligent building control method integrating visual and linguistic information, characterized in that, Includes the following steps: S1, construct dataset D, including S11~S12; S11, construct samples corresponding to task instructions based on the construction data of the embodied intelligent agent, and combine all samples into a dataset D, where the m-th sample D in D is... m ={L m RGB-D m PC m BIM m A m ,K m }, L m For task instructions, RGB-D m PC m Execute L respectively m Previous construction site RGB-D drawings, 3D point clouds and BIM models, A m K m Execute L respectively m The set of actions and skill parameters at that time; in, , , For A m The b-th three-dimensional semantic action, where B is... Total, a b o b p b R b C b These are the action type, the target instance corresponding to the action, the optimal point of action on the target instance, the action's area of ​​effect, and the action's constraints. , for Reference skill parameters; S12, Obtain the initial state S1 and the state S after executing the b-th action of the embodied agent. b+1 ; S2, based on RGB-D m and BIM m Generate a 3D semantic world model W of the construction site m ; S3 constructs an action generation model, which includes a ViT visual encoder, a Transformer language encoder, a cross-modal Transformer fusion module, an action decoder, a 3D adapter, a 3D action parser, and a skill mapping layer. The first four constitute the basic model, and the latter three constitute the training model. The ViT visual encoder is used to encode RGB images into image word sequences T. img and based on T img Generate semantic feature maps; The Transformer language encoder is used to encode task instructions into a sequence of language lexical units T. lang And attention converges into a semantic vector z L ; 3D adapter is used to adapt to L m and W m Generate the three-dimensional semantic features Z of the construction area 3D And mapped to a spatial cue lexical sequence T 3D ; The cross-modal Transformer fusion module will integrate T img T lang T 3D Cross-modal fusion is performed to obtain fused feature A token ; The action decoder is used to determine the fusion feature A token Output B predicted action types; The 3D action parser is used to generate the 3D semantic action for each predicted action type. The 3D semantic action for the b-th predicted action type is... , a, o, p*, R*, and C represent a b o b p b R b C b The predicted value; The skill mapping layer is used to... The robot's current state q and C b Mapped to the predicted skill parameters of the embodied intelligent agent Where xyz is 3D coordinates F and v represent the target posture, force, and velocity of the embodied intelligent agent performing the action, and T represents the duration of the action; S4, based on online verification of digital twins and closed-loop update training of three-dimensional error field, the action generation model is used as the embodied intelligent action control model after the update is completed; S5 uses an embodied intelligent motion control model to generate skill parameter sequences based on task instructions; S2 includes S21 to S24; S21, Obtain RGB-D m This includes RGB images and depth maps. The ViT visual encoder is used to extract semantic feature maps from the RGB images, and the semantic features of each pixel location are mapped to a PC by combining the depth map. m The corresponding spatial point, where the semantic features of spatial point i are: ; S22, based on BIM m Generate the BIM prior feature vector of spatial point i , It contains the BIM component category, BIM geometric parameters, material properties, construction status, and process constraints corresponding to spatial point i; , Among them, Enc BIM (∙) represents the BIM feature encoder, which consists of a category embedding layer and a multilayer perceptron. , , , , These are the BIM component category, BIM geometric parameters, material properties, construction status, and process constraints corresponding to spatial point i, respectively. S23, will and Adaptive fusion into fused semantic features f of spatial point i i , α i for The weights; S24, Generate a 3D semantic world model W of the construction site. m , In the formula, N is the total number of spatial points, and dc i s i o i These are the spatial coordinates, semantic category, and target instance of spatial point i, respectively. S4 includes S41~S43; S41, on the simulation platform, press Simulation is performed, and the embodied intelligent agent executes K b After that, the state changed from become ; S42, Construct the total training loss L all ; , In the formula, L space For the loss of three-dimensional spatial positioning error, L sim The simulation consistency loss is used to constrain the consistency between the simulated predicted state and the actual execution state; L skill For skill generation loss, γ1 and γ2 are respectively L sim L skill The weight, L space L sim L skill Calculate according to the following formulas respectively: , In the formula The squared distance is the norm 2. S43, based on dataset D, to minimize L all The parameters of the training model are updated. After the update is completed, the action generation model is used to form an embodied intelligent action control model. The updated training model is to guide the model to adjust parameters based on the three-dimensional error field and error category, including Sa1~Sa3. Sa1, computation of embodied intelligent agent execution K b The state error e after b , A three-dimensional error field is generated based on M sampling points, and the three-dimensional error of any spatial point p in the three-dimensional error field is E(p). , , In the formula, Assign weights to the error of the j-th sampling point. Let d be the spatial weight of the j-th sampling point to the spatial point p. j σ is the distance from the j-th sampling point to p*, σ is the distance attenuation coefficient, and Softmax(⋅) is the Softmax function; Sa2 predefines four types of errors and their corresponding thresholds, and represents the optimal deviation of the point of application, D. loc BIM and site inconsistency deviation D bim State deviation D sim Trajectory tracking or force control instability deviation D ctrl ; , In the formula, It is a 2-norm. This represents the local structure of the BIM model within a three-dimensional error field. Let t be the distance from point i in space to B, and t be the distance from point i to point B. b Total duration, q t , These represent the robot's simulated posture at time t and its reference posture in the sample, respectively. Sa3 calculates four types of errors within E(p); If D loc If the threshold is exceeded, adjust the network parameters of the 3D adapter and the 3D motion parser. If D bim If the threshold is exceeded, the BIM model will be adjusted or an alarm will be triggered. If D sim Or D ctrl If the threshold is exceeded, the skill mapping layer will be adjusted.

2. The three-dimensional embodied intelligent building control method integrating visual and linguistic information according to claim 1, characterized in that, The 3D adapter includes a PointNet++ encoder, a sparse 3D convolutional network, a cross-attention filtering module, and a mapping layer; The PointNet++ encoder is used as input to W and outputs multi-scale local features W1 of W. The sparse 3D convolutional network is used as input W1 and outputs sparse 3D features W2. The cross-attention filtering module is used for semantic vector z L W2 serves as the query matrix, and W2 serves as both the key and value matrix. Z is generated based on a cross-attention mechanism. 3D ; The mapping layer is used to convert Z... 3D Mapping to a preset dimension yields spatial cue words.

3. The three-dimensional embodied intelligent building control method integrating visual and linguistic information according to claim 1, characterized in that, The motion constraints include BIM construction accuracy constraints, robot joint constraints, safety distance constraints, and no-operation zone constraints. The skills include grasping skills, handling skills, placement skills, alignment skills, insertion skills, clamping skills, obstacle avoidance movement skills, drilling and positioning skills, scanning and detection skills, compliant contact skills, tightening and fixing skills, and pushing and adjusting skills.

Citation Information

Patent Citations

  • Multi-modal fusion body-equipped intelligent system in complex scene and use method of multi-modal fusion body-equipped intelligent system

    CN121615071A

  • Multi-modal task understanding and collaborative planning method and system oriented to agent with body

    CN122044801A