Building scene robot operation method based on large model technology
By applying big model technology in construction robots, building the underlying instruction library and operation thinking chain, combining BIM information and industry knowledge, the existing construction robots lack autonomy and difficulty in adapting to environmental changes is solved, and more efficient and accurate task operations are achieved.
Patent Information
- Application Number
- CN202510223278.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Existing construction robots lack autonomy during task operations, need to manually specify the order of action and behavioral logic, and it is difficult to adapt to environmental changes and cannot effectively utilize building scene information, resulting in inaccurate and inefficient task operations.
Using a building scene robot operation method based on large-model technology, we can screen the basic movements and tasks of the robot, build an underlying instruction library, combine BIM information and industry knowledge to identify the workable areas, use a large language model to build an operation thinking chain, generate action instructions, and find the most similar skills in the skill library through the SimCSE model, execute skills, and finally correct the robot's end trajectory to ensure smoothness and safety.
It realizes the independent task reasoning and execution of robots in building scenarios, improves the robustness and generalization ability of task operations, reduces human intervention, and improves the accuracy and efficiency of operations.
Smart Images

Figure CN120056106A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent construction, and specifically to a method for a construction scene robot operation based on large model technology. Background Art
[0002] When existing construction robots perform task operations, after generating task operations for the operation area based on BIM information, it is necessary to artificially specify the action sequence and behavior logic of the robots, lacking autonomy.
[0003] At the same time, when the construction scene environment and the given tasks change dynamically, the robots need to be artificially intervened and reconfigured to adapt to the new environment.
[0004] And the task planning method based on large language model technology often adopts an end-to-end method to give an action sequence that conforms to the constraints of the robot and the environment.
[0005] However, the existing robot task planning technology based on large language models is often limited to specific robot tasks, difficult to adapt to the environment, and lacks effective utilization of construction scene information, and cannot perform task operations accurately and efficiently.
[0006] Therefore, a new solution needs to be proposed for the above problems. Summary of the Invention
[0007] The purpose of the present invention is to provide a method for a construction scene robot operation based on large model technology to solve the technical problems raised in the background art.
[0008] To achieve the above purpose, the present invention provides the following technical solution: A method for a construction scene robot operation based on large model technology, at least including the following steps:
[0009] S1: By screening the basic actions and tasks of the robot, collecting trajectory data and training with the Diffusion model to obtain a robot instruction model, thereby constructing a bottom instruction library to optimize the robot action execution;
[0010] S2: Combining BIM information and industry knowledge, identifying the operable areas in the construction scene, and generating a specific set of operation tasks through discretizing the operation areas and path planning algorithms;
[0011] S3: Using the large language model to construct an operation thinking chain;
[0012] S4: Given a task prompt, using the large language model to form action instructions for each operation task, using the SimCSE model to find the most similar skills in the skill library, and executing the skills;
[0013] S5: Correct the trajectory of the robot's end to ensure the smoothness and safety of the robot's movement path while adhering to the motion constraints of the actuator.
[0014] Further, the S1 at least includes the following steps:
[0015] According to the actual robot structure, screen the basic actions and tasks of the robot and construct the robot's underlying instruction library.
[0016] Collect and clean the trajectory data during the actions or tasks of the robotic arm in real or simulated scenarios. For each action or task obtained by screening, use the Diffusion model to train the robot instruction model. where I i is the i-th underlying instruction, see the following formula:
[0017]
[0018] where k is the iteration step. is the robot's action at the (k - 1)-th iteration at time t, O t is the robot's observation at time t, is the language description of the instruction I i N(0, δ 2 I) represents a Gaussian distribution with an expected value of 0 and a variance of δ 2 I, is the pose of the robot's starting and ending points, and α, γ, and δ are parameters;
[0019] The loss function of the robot instruction model is:
[0020]
[0021] where is the logarithmic gradient of A t , is the predicted action logarithmic gradient.
[0022] Further, the robot's underlying instruction library at least includes actuator action or task action instructions.
[0023] Further, the trajectory data at least includes the robot's pose and the robot's image data.
[0024] Further, the S2 at least includes the following steps:
[0025] Combined with BIM information and prior industry knowledge, find the workable areas of the robot within the building scene;
[0026] For each work area S k, for \(k\in[1,K]\), use the discrete point path planning and task generation algorithm to generate a set of operation tasks on the operation area. Each operation task includes an operation start point, an end point, and operation content;
[0027] The discrete point path planning and task generation algorithm at least includes the following steps:
[0028] For each piece of operation area \(S\) k , uniformly take discrete points on the surface, denoted as \(P = \{P\) 1 ,…, \(P\) i ,…, \(P\) M \}, \(M\) is the number of discrete points, and \(P\) i is the discrete point pose;
[0029] Connect the surface points in the way of zig-zag or Hilbert curve, etc., then the task \(T\) i = \{P\) i , \(P\) i+1 , \(C\}, where \(C\) is the task.
[0030] Furthermore, the operation thought chain in S3 at least includes trajectory information, action information, and historical information. The trajectory information represents the actions that have been executed currently. The action information at least includes the underlying instructions of the robot that can be used and the meanings represented by the instructions. The historical information at least includes the action information of the robot executed previously.
[0031] Furthermore, S4 at least includes the following steps:
[0032] First, the robot uses its own sensors to sense the environment and determine the positional relationship between the robot and the operation area;
[0033] For each action, construct geometric and dynamic functions according to the operation points and the operation area, and correct the smoothness and dynamic safety of the trajectory;
[0034] For each task \(T\) i , obtain the coordinates \(p\) i of the operation point \(P\) i in the operation area through a three-dimensional sensor. According to the relative relationship between \(P\) i+1 and \(P\) i , obtain the coordinates \(p\) i+1 of \(P\) i+1 in the operation area, and use the large model to decompose the operation task into the corresponding action instruction set \(I'\);
[0035] For each action instruction in the operation instruction set, the action logarithmic gradient estimated by the robot instruction model satisfies:
[0036]
[0037] Among them
[0038] The movement of the robot can be obtained by the following formula:
[0039]
[0040] where β is a parameter.
[0041] Furthermore, the S5 at least includes the following steps:
[0042] Correct the obtained end trajectory of the robot e = (e 1 , …, e j , …, e N );
[0043] where: e 1 = p i , e N = p i+1 , e j is the j-th pose point, and N is the number of poses;
[0044] Therefore, the actuator motion constraint function is as follows:
[0045]
[0046] where λ snap (e j ) is the trajectory smooth term from e j-1 to e j , represents the distance that violates the safety threshold from the obstacle in the movement path from e j-1 to e j .
[0047] Compared with the prior art, the beneficial effects of the present invention are:
[0048] The present invention proposes a task operation method based on a large model. After a given task, using the prior information of the building scene, the long-range task is decoupled into multiple simple subtasks, improving the robustness of task operation;
[0049] At the same time, the present invention does not require manual specification of the behavior logic of the robot. The robot can autonomously infer the behavior logic and complete the task according to its underlying skills, improving the generalization ability of the robot task. Description of the Drawings
[0050] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0051] Figure 1 It is a schematic diagram of the execution process of a single time-step task of the present invention;
[0052] Figure 2 It is a schematic diagram of the robot's underlying skill library of the present invention;
[0053] Figure 3 It is a schematic diagram of the test area of the present invention;
[0054] Figure 4 It is a schematic diagram of the measuring point sequence of the present invention;
[0055] Figure 5 It is a schematic diagram of the bolt tightening skill library of the present invention;
[0056] Figure 6 It is a schematic diagram of the bolt tightening skill library of the present invention;
[0057] Figure 7 It is a schematic diagram of the robot of the present invention tightening bolts. Specific embodiments
[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0059] Embodiment 1:
[0060] Please refer to Figure 1 and Figure 2 , a building scene robot operation method based on large model technology, at least including the following steps:
[0061] S1: By screening the basic actions and tasks of the robot, collecting trajectory data and training with the Diffusion model to obtain a robot instruction model, thereby constructing an underlying instruction library and optimizing the robot's action execution;
[0062] S2: Combining BIM information and industry knowledge, identifying the operable areas in the building scene, and generating a specific set of operation tasks through discretizing the operation areas and path planning algorithms;
[0063] S3: Using a large language model to construct an operation thinking chain;
[0064] S4: Given the task prompt, use the large language model to form action instructions for each job task, use the SimCSE model to find the most similar skills in the skill library, and execute the skills;
[0065] S5: Correct the trajectory of the robot end to ensure the smoothness and safety of the robot motion path, while following the action constraints of the actuator.
[0066] S1 includes at least the following steps:
[0067] According to the actual robot structure, screen the basic actions and tasks of the robot, and construct the robot's underlying instruction library.
[0068] Collect and clean the trajectory data during the actions or tasks of the robotic arm in real or simulated scenarios. For each action or task obtained by screening, use the Diffusion model to train the robot instruction model. Where I i is the i-th underlying instruction, see the following formula:
[0069]
[0070] where k is the iteration step, is the robot action at the (k - 1)-th iteration at time t, O t is the observation of the robot at the t-th moment, L Ii is the language description of the instruction I i N(0, δ 2 I) represents a Gaussian distribution with an expectation of 0 and a variance of δ 2 I, is the pose of the robot's starting and ending points, and α, γ, and δ are parameters;
[0071] The loss function of the robot instruction model is:
[0072]
[0073] where is the logarithmic gradient of A t and is the predicted action logarithmic gradient.
[0074] The robot's underlying instruction library includes at least the action instructions of the actuator or task actions.
[0075] The trajectory data includes at least the robot pose and the robot image data.
[0076] S2 includes at least the following steps:
[0077] Combined with BIM information and prior industry knowledge, find the operable areas of the robot in the building scene;
[0078] For each operation area S k , k ∈ [1, K], use the discrete point path planning and task generation algorithm to generate a set of operation tasks on the operation area. Each operation task includes an operation start point, an end point, and operation content;
[0079] The discrete point path planning and task generation algorithm at least includes the following steps:
[0080] For each operation area S k , uniformly take discrete points on the surface, denoted as P = {P 1 , …, P i , …, P M}, M is the number of discrete points, and P i is the pose of the discrete point;
[0081] Connect the surface points in the way of zig-zag or Hilbert curve, then the task T i = {P i , P i+1 , C}, where C is the task.
[0082] The operation thought chain in S3 at least includes trajectory information, action information, and historical information. The trajectory information represents the actions that have been executed currently. The action information at least includes the underlying instructions of the robot that can be used and the meanings represented by the instructions. The historical information at least includes the action information of the robot executed previously.
[0083] S4 at least includes the following steps:
[0084] First, the robot uses its own sensors to perceive the environment and determine the positional relationship between the robot and the operation area;
[0085] For each action, construct geometric and dynamic functions according to the operation points and the operation area to correct the smoothness and dynamic safety of the trajectory;
[0086] For each task T i , obtain the coordinate p i of the operation point P i in the operation area through a three-dimensional sensor. According to the relative relationship between P i+1 and P i , obtain the coordinate p i+1 of P i+1 in the operation area, and use the large model to decompose the operation task into a corresponding action instruction set I′;
[0087] For each action instruction in the operation instruction set, the action pair gradient estimated by the robot instruction model satisfies:
[0088]
[0089] Among them
[0090] The movement of the robot can be obtained by the following formula:
[0091]
[0092] where β is a parameter.
[0093] S5 includes at least the following steps:
[0094] Correct the obtained end trajectory of the robot e = (e 1 , …, e j , …, e N );
[0095] where: e 1 = p i , e N = p i+1 , e j is the j-th pose point, and N is the number of poses;
[0096] Therefore, the actuator motion constraint function is as follows:
[0097]
[0098] where λ snap (e j ) is the trajectory smooth term from e j-1 to e j , represents the distance that violates the safety threshold from the obstacle in the moving path from e j-1 to e j .
[0099] Embodiment 2:
[0100] Based on the above Embodiment 1, a specific application to a concrete strength testing robot is proposed;
[0101] Refer to Figure 2 to set the underlying skill library of the robot;
[0102] Refer to Figure 3 – Figure 4 , according to the BIM information, the following is obtained for a certain test area. According to the existing specifications, 12 measuring points can be set for each measuring point area, and the measuring point sequence is as follows;
[0103] After the strength testing robot moves to the measuring area, calculate the position of each measuring point in the robot coordinate system according to the distance from the measuring point to the wall edge. According to the measuring point sequence, the large model gives the robot motion sequence and conducts dynamic checks.
[0104] Embodiment 3:
[0105] Based on the above Embodiment 1, a specific application to a bolt tightening robot is proposed;
[0106] Refer to Figure 5 - Figure 6 , set the underlying skill library of the robot;
[0107] Refer to Figure 7 , and combine to achieve bolt tightening of the robot.
[0108] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
Claims
1. A construction scene robot operation method based on large model technology, characterized by: At least the following steps are included: S1: By screening the basic actions and tasks of the robot, collecting trajectory data and using the Diffusion model to train the robot instruction model, the underlying instruction library is constructed to optimize the robot action execution; S2: Combine BIM information and industry knowledge to identify the workable areas in the building scene, and generate a specific set of work tasks through discretized work areas and path planning algorithms; S3: Use the big language model to build a task thinking chain; S4: Given a task prompt, use the large language model to form action instructions for each task, use the SimCSE model to find the most similar skills in the skill library, and execute the skills; S5: Correct the robot's end trajectory to ensure the smoothness and safety of the robot's motion path while complying with the actuator's motion constraints.
2. The construction scene robot operation method based on large model technology according to claim 1 is characterized by: The S1 at least comprises the following steps: According to the actual robot structure, filter the basic actions and tasks of the robot and build the robot's underlying instruction library. Collect and clean the trajectory data of the robot arm's actions or tasks in real or simulated scenes, and use the Diffusion model to train the robot instruction model for each action or task obtained. Among them I i is the ith underlying instruction, see the following formula: Where k is the iteration step, is the robot action at time t at k-1 iterations, O t is the robot’s observation at time t, For instruction I i The language description of N(0,δ 2 I) means the expectation is 0 and the variance is δ 2 Gaussian distribution of I, are the starting and ending positions of the robot, α, γ and δ are parameters; The loss function of the robot instruction model is: in A t The logarithmic gradient of is the predicted action log gradient.
3. The construction scene robot operation method based on large model technology according to claim 2 is characterized in that: The robot bottom-level instruction library at least includes actuator action or task action instructions.
4. The construction scene robot operation method based on large model technology according to claim 2 is characterized by: The trajectory data at least includes robot posture and robot image data.
5. The construction scene robot operation method based on large model technology according to claim 2 is characterized by: The S2 at least comprises the following steps: Combine BIM information with prior industry knowledge to find out the robot’s operable area within the building scene; For each work area S k ,k∈[1,K], using discrete point path planning and task generation algorithm to generate a set of work tasks in the work area, each work task contains the work start point, end point and work content; The discrete point path planning and task generation algorithm comprises at least the following steps: For each work area S k , uniformly select discrete points on the surface, denoted as P = {P1,…,P i ,…,P M }, M is the number of discrete points, P i is the discrete point pose; If the surface points are connected by zig-zag or Hilbert curve, then task T i = {P i ,P i+1 ,C}, where C is the task.
6. The construction scene robot operation method based on large model technology according to claim 1 is characterized by: The operation thinking chain in S3 includes at least trajectory information, action information and history information. The trajectory information represents the action currently executed, the action information includes at least the robot underlying instructions that can be used and the meaning of the instructions, and the history information includes at least the robot action information executed previously.
7. The construction scene robot operation method based on large model technology according to claim 2 is characterized by: The S4 at least comprises the following steps: First, the robot uses its own sensors to perceive the environment and determine the positional relationship between the robot and the work area; For each action, geometric and dynamic functions are constructed according to the working point and working area to correct the smoothness and dynamic safety of the trajectory; For each task T i , the coordinates of the working point P are obtained through the three-dimensional sensor i Coordinate p in the working area i , according to P i+1 With P i The relative relationship between i+1 Coordinate p in the working area i+1 , using the large model to decompose the task into the corresponding action instruction set I′; For each action instruction in the job instruction set, the action logarithmic gradient estimated by the robot instruction model satisfies: in The robot's motion can be obtained by the following formula: Where β is a parameter.
8. The construction scene robot operation method based on large model technology according to claim 7 is characterized by: The S5 at least comprises the following steps: The robot terminal trajectory e=(e1,…,e j ,…,e N ) for amendment; Where: e1 = p i , e N =p i+1 , e j is the jth pose point, N is the number of poses; Therefore, the actuator action constraint function is as follows: Among them, λ snap (e j ) is e j-1 To e j The trajectory smoothing term, It means that in e j-1 To e j The distance along the motion path that violates the obstacle safety threshold.
Citation Information
Patent Citations
Mechanical arm assembly task planing method and building assembly method based on BIM
CN109760059A
Intelligent climbing frame and exterior wall operation robot control method based on BIM
CN111764664A
Multi-robot distributed collaborative operation method and system
CN117452932A
Robot control method based on multi-modal large model
CN117944052A
Robot motion planning optimization method
CN118809583A