Position estimation method and device for curtain wall panel installation and anti-collision warning
Through binocular vision and deep learning technology, the problem of estimating the position position of the curtain wall installation area in complex environments is solved, and the precise installation and anti-collision warning of curtain wall robots outdoors is realized.
Patent Information
- Application Number
- CN202310392217.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-13
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-04-13
AI Technical Summary
The prior art is difficult to accurately estimate the position of the curtain wall installation area under complex environments, especially when the light changes are large, the curtain wall is made of glass and the installation plane is not perpendicular.
Binocular vision and deep learning technology are used to obtain the precise grasping position of curtain wall panels and train the pose estimation network model based on computer vision and deep learning to achieve accurate estimation of the pose position of curtain wall installation area.
It improves the accuracy of attitude prediction in curtain wall installation areas in complex environments, overcomes the problem of insufficient robustness of traditional computer vision methods in lighting changes and complex environments, and realizes anti-collision warning and precise installation of curtain wall robots outdoor installation.
Smart Images

Figure CN116433765B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of building curtain wall installation, and relates to a posture estimation method and device for curtain wall panel installation and anti-collision warning. Background Art
[0002] With the development of construction technology and the improvement of architectural aesthetics, in recent years, large curtain walls are used more and more in building design to increase the design sense of the building, and glass curtain walls are used to improve the lighting of the building. However, the curtain wall is large in size and weight and is easily damaged, making it difficult to move. Some curtain walls are not installed perpendicular to the ground, and the location is not convenient for manual installation. Manual installation consumes a lot of manpower and is also dangerous. Therefore, a method for installing curtain walls using robots is needed to achieve the installation of panels and anti-collision warning. The key step is to achieve accurate grasping of curtain wall panels and obtain the relative position of the curtain wall area and the center position of the robot gripper.
[0003] At present, the research on the installation and positioning of curtain walls is relatively limited. In similar fields, the binocular ranging principle is used to determine the posture relationship between automobile glass and frame. Structured light sources are used to create features, and feature matching is performed through artificial feature description to obtain depth. The installation plane is fitted by computer image processing to determine the relative posture of the installation plane and the calibration plane. However, for curtain walls, it is necessary to consider that the installation position of each curtain wall is different, and installation in outdoor environments may be affected by factors such as lighting and occlusion. In addition, since some curtain walls are made of glass, changes in outdoor weather and light reflection and refraction of other glass curtain walls will change the lighting conditions, thereby affecting image acquisition. Traditional computer vision methods are difficult to be robust in complex environments. In terms of depth acquisition of the installation area, the most critical step in traditional binocular ranging is feature matching, but factors such as narrow glass curtain wall frames, large lighting changes, and high feature repetition in curtain wall installation scenes make it undesirable to use structured light to create feature points for feature matching.
[0004] In the research on pose estimation based on deep learning, most of the current work focuses on monocular pose estimation of small objects, which is difficult to directly apply to curtain wall installation scenarios. Also, due to the complex environment, it is difficult to directly obtain accurate depth using a camera, and it is also difficult to use RGB-D images as data for deep learning to train the network.
[0005] Therefore, there is an urgent need for a method that can estimate the posture of the curtain wall installation area in complex environments. Summary of the invention
[0006] In view of this, the purpose of the present invention is to provide a posture estimation method for (unconventional / multiple / different) posture curtain wall panel installation and anti-collision warning, to solve the problem of the relative posture relationship between the panel and the installation area when using a curtain wall robot to install the curtain wall, and to improve the accuracy of the posture prediction of the curtain wall installation area in a complex environment.
[0007] In order to achieve the above object, the present invention provides the following technical solutions:
[0008] A posture estimation method for curtain wall panel installation and anti-collision warning specifically comprises the following steps:
[0009] S1: Get the precise grabbing position of the curtain wall panels: Use a curtain wall panel storage device with special marking points to store the curtain wall panels, and use binocular vision to accurately locate them;
[0010] S2: Get the dataset and divide it into training set, validation set and test set;
[0011] Training set: mainly used for network parameter training;
[0012] Validation set: Mainly used to evaluate the model, confirm the termination time of network training, and prevent overfitting. When the training loss decreases and the validation loss increases, the network training is terminated.
[0013] Test set: Use the network model with the best parameters retained during training to evaluate the generalization ability of the model on the test set. If the test index performance meets the requirements, it can be transplanted to the actual scenario application; if it does not meet the requirements, optimize and adjust the network structure design and restart training.
[0014] S3: Input the training set into the pose estimation network model based on computer vision and deep learning for training. The training termination time is confirmed by the loss curve of the validation set and the training set. Finally, the trained network model is used to estimate the relative pose relationship (T, q) between the camera and the installation area.
[0015] S4: Real-time pose estimation: First, the relative pose relationship between the camera and the robot gripper, i.e., the curtain wall panel, is calibrated to obtain the pose (T) from the camera to the robot gripper (curtain wall panel). 0 ,q 0 ); then use the calibrated data (T 0 ,q 0 ) and the data estimated by the pose estimation network model (T, q) are converted into the pose information (T e ,q e ).
[0016] Further, in step S1, the curtain wall panel storage device includes a frame 1, an adjustment device and a fixing device 2; a special marking point 201 is provided on the fixing device 2; the curtain wall panel is placed upright on the frame 1 through the fixing device 2, and the special marking point 201 is positioned using binocular vision to achieve accurate positioning of the curtain wall panel.
[0017] Further, in step S1, the robot gripper accurately grasps the position of the curtain wall panel and obtains it through binocular vision positioning; the special marking point 201 on the fixing device is used for feature matching, and the coordinates of the marking point on the left and right images are respectively (u l ,v l ) and (u r ,v r );
[0018] Take the left camera as the main camera, according to the focal length f of the camera itself, the horizontal baseline b, and the pixel density m of the image in the X and Y directions of the pixel coordinate system x 、m y and the origin of the image physical coordinate system (o x ,o y ), get the three-dimensional coordinates (x, y, z) of the special marker point in the camera coordinate system;
[0019]
[0020]
[0021] Among them, f x 、f y Represents the focal length in the X and Y directions respectively.
[0022] The accurate grabbing position estimation of the panel is achieved through the relative position relationship between the coordinates of the special marking points in the camera coordinate system and the center of the curtain wall panel.
[0023] Further, in step S2, a data collection system 3 is used to acquire a data set, wherein the data collection system 3 includes a binocular camera 301 and a tilt sensor 302, and the tilt sensor is attached to the side of the binocular camera;
[0024] Obtaining the dataset specifically includes the following steps:
[0025] S21: before shooting, place the camera shooting plane parallel to the plane of the installation area to calibrate the tilt sensor 302;
[0026] S22: Based on the installation area and camera parameters, a shooting distance range that can capture the entire installation area is selected, and shooting is performed under different weather, lighting conditions, distances, angles, dust occlusion and other influencing factors to obtain a large number of binocular RGB images of the installation area;
[0027] S23: It is necessary to obtain the value of the tilt sensor 302 at the same time of each shooting. The value is usually the Euler angle (α, β, γ) rotated in the order of ZYX, where α, β, and γ respectively represent the rotation angles of the tilt sensor 302 relative to the X, Y, and Z axes of the original coordinate system; the value recorded here represents the rotation amount of the camera shooting plane relative to the plane of the curtain wall installation area;
[0028] S24: Perform pixel-level segmentation on the left and right binocular RGB images respectively, and use straight lines to divide the installation area as a label for the semantic segmentation task;
[0029] S25: Enhance the binocular RGB image as the image of the data set and increase the training samples;
[0030] S26: In extreme cases, the rotation angle of the shooting plane relative to the plane of the installation area may reach 90°. Using Euler angles to represent the rotation may cause the problem of universal joint deadlock. Therefore, the Euler angles (α, β, γ) are converted into quaternion q m =w+xi+yj+zk=[w,(x,y,z)]=[w,v] represents the rotation amount, where the real part w determines the size of the rotation angle, and the imaginary part v determines the direction of the rotation transformation, where i, j, k are imaginary parts of different units, satisfying i 2 =j 2 =k 2 =ijk=-1. m Taking the inverse we get Used to indicate the rotation amount of the curtain wall panel installation area plane relative to the shooting plane;
[0031]
[0032]
[0033] S27: Using Quaternions The posture information represented is used as the label of posture estimation to complete the establishment of the data set; each data includes two RGB images obtained by the binocular camera, the installation area segmentation label and the posture estimation label
[0034] S28: After labeling, take 70% of the dataset as the training set, 20% as the validation set, and 10% as the test set.
[0035] Further, in step S3, training a posture estimation network model based on computer vision and deep learning specifically includes the following steps:
[0036] S31: Train the network through supervised learning so that the network can obtain preliminary results of the installation area through semantic segmentation (the segmentation result at this time contains multiple areas, and the area boundaries are blurred);
[0037] S32: Screening the segmentation prediction results (area comparison, whether the center point is included) to obtain the target segmentation area;
[0038] S33: fitting the target segmentation region to obtain a region composed of straight lines, thereby obtaining a final segmentation region;
[0039] S34: According to the divided area, the two-dimensional coordinates of a corner point of the area and its diagonal point in the pixel coordinate system can be obtained, and the three-dimensional coordinates of the corner point in the camera coordinate system can be obtained through the binocular principle; the midpoint of the two is calculated, that is, the midpoint of the installation area, and the coordinates of the point are the translation matrix T from the camera to the center of the installation area;
[0040] S35: Splice the two segmented and processed left and right images, train the network through supervised learning, and design the loss function RLoss so that the network can estimate the rotation amount q.
[0041] Further, in step S33, the RANSAC method is used to fit the target segmentation area.
[0042] Further, in step S4, the position information (T e ,q e ) is calculated as: e =TT 0 , in, Indicates q 0 The conjugate quaternion of .
[0043] Furthermore, T e and q e Feedback is given to the curtain wall grabbing robot so that the robot can adjust the movement of the robotic arm according to the feedback amount; first, according to q e Adjust the posture of the curtain wall panel so that the plane of the curtain wall panel is parallel to the plane of the installation area; then adjust the position of the curtain wall panel according to T e The provided position information moves the curtain wall panel toward the center O of the installation area until it is 20 cm away, issues an anti-collision warning message and limits the movement speed of the robotic arm to prevent collisions that may damage the panel or frame structure, and enables possible fine-tuning during the installation process, thereby achieving precise installation.
[0044] The beneficial effects of the present invention are:
[0045] 1) Compared with traditional machine vision, the present invention has better robustness, can cope with complex interference factors under different environmental conditions, is more suitable for outdoor work scenes, and can solve complex situations such as the plane of the curtain wall installation area is not perpendicular to the ground. It overcomes the limitations of using structured light sources to manufacture features and the inconvenience caused by the continuous use of precise measurement sensors.
[0046] 2) The present invention uses a binocular camera to obtain image data and uses the splicing of the viewing angle dimension for the deep learning network. Compared with the single-eye pose estimation, the deep learning can better reduce the error and have better performance in accuracy by connecting the binocular RGB images in the viewing angle dimension to form a four-dimensional input, performing 3D convolution to extract features, and regressing the pose information.
[0047] 3) The present invention can achieve anti-collision warning during the installation of panels by controlling the mechanical arm of the curtain wall robot by estimating the position, thereby improving the safety of the construction process.
[0048] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:
[0050] Figure 1 This is a schematic diagram of a curtain wall panel storage device;
[0051] Figure 2 is a schematic diagram of a fixing device for a clamping structure;
[0052] Figure 3 It is a schematic diagram of the data acquisition system;
[0053] Figure 4 This is a schematic diagram of binocular system coordinate calculation;
[0054] Figure 5 Schematic diagram of the steps for training the pose estimation network;
[0055] Figure 6 This is a schematic diagram of the network structure for installing region segmentation and pose estimation;
[0056] Figure 7 This is a schematic diagram of the effect of RANSAC fitting on one edge of the installation area;
[0057] Figure 8Schematic diagram of further processing of segmentation results;
[0058] Fig. 9 Schematic diagram of the relative posture relationship of each part in the prediction process;
[0059] Figure numerals: 1-frame, 2-fixing device, 201-special marking point, 202-rotation axis, 203-flexible clamp, 3-data acquisition system, 301-binocular camera, 302-tilt sensor. DETAILED DESCRIPTION
[0060] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0061] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0062] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0063] See also Figures 1 to 9 The present invention provides a posture estimation method for curtain wall panel installation and anti-collision warning, which mainly includes three parts: obtaining the precise grasping position of the curtain wall panel, training the posture estimation network and real-time posture estimation.
[0064] 1. Get the precise grabbing position of the curtain wall panels.
[0065] Purpose: To make the gripping center and plane coincide with the center and plane of the plate, eliminating the need for subsequent measurement of the relative position relationship between the gripper and the plate.
[0066] Method: Use a curtain wall panel storage device with special marking points to store the curtain wall panels, and use binocular vision to accurately locate them.
[0067] In order to facilitate the curtain wall robot to grab the curtain wall, a curtain wall panel storage device is designed (such as Figure 1 As shown in FIG. 1 , the main body is composed of a frame 1, an adjusting device and a fixing device 2. Since some curtain walls are made of transparent glass, it is very difficult to identify the glass curtain wall body by vision, so a special marking point 201 is added on the above basis. The curtain wall panel is placed upright on the frame through the fixing device, and the special marking point is positioned by binocular vision, so that the position of the curtain wall panel can be accurately positioned.
[0068] The frame 1 adopts a nested structure and can be extended and retracted to change the width to adapt to panels of different sizes.
[0069] The fixing device 2 adopts a clamping structure, such as Figure 2 As shown, the curtain wall panel is kept vertical to facilitate the curtain wall robot to grab the panel. The flexible clamp 203 is used to clamp the panel to prevent damage to the curtain wall panel. The movement of the clamp adopts a rotation method. After the robot grabs the panel, the clamp is released through the rotating shaft 202 to avoid the movement route of the curtain wall panel. The special marking point 201 is located on the fixing device 2.
[0070] The special marking point 201 can also be added to the existing curtain wall storage device of applicable structure as the plate storage device required by this system, and is not limited to the structure described in this embodiment.
[0071] The grabbing position of the curtain wall plate is determined by the precise positioning of the special marking point 201 on the plate storage device by binocular vision. The data acquisition system 3 where the binocular camera 301 is located, such as Figure 3 As shown, it includes a binocular camera 301 and a tilt sensor 302.
[0072] Through this device and precise visual positioning, the accurate grasping position of the plate can be guaranteed, so that in the subsequent posture estimation part, only the posture of the installation area needs to be estimated, and there is no need to estimate the posture of the curtain wall plate.
[0073] The precise grasping position of the robot gripper is obtained through binocular vision positioning.
[0074] The two cameras are calibrated separately to solve the distortion problem of the camera itself.
[0075] The binocular camera 301 obtains a binocular RGB image of the location of the plate storage device and performs feature matching.
[0076] Get the coordinates of the special mark point 201 in the pixel coordinate system on the left and right images (u l ,v l ) and (u r ,v r ). Based on the focal length f of the camera itself, the horizontal baseline b, and the pixel density m in the x and y directions x 、m y and the origin of the image physical coordinate system (o x ,o y ) calculates the three-dimensional coordinates (x, y, z) of the special marker point 201 in the camera coordinate system. The geometric conversion relationship of the coordinates is as follows: Figure 4 shown.
[0077] Take the left camera as the main camera, according to the focal length f of the camera itself, the horizontal baseline b, and the pixel density m of the image in the X and Y directions of the pixel coordinate system x 、m y and the origin of the image physical coordinate system (o x ,o y ) to obtain the three-dimensional coordinates (x, y, z) of the special marker point in the camera coordinate system.
[0078]
[0079]
[0080] The precise gripping of the plate is achieved through the relative position relationship between the gripper center position P and the special marking point 201 in the camera coordinate system.
[0081] 2. Train the pose estimation network and use computer vision and deep learning models to estimate the relative pose relationship (T,q) from the camera to the installation area.
[0082] Purpose: To train a network model, the function of which is to input two RGB images taken in real time and output the pose information (T, q).
[0083] The specific steps of the method are:
[0084] 1) Establish a data set, mainly including: (1) Calibrate the camera and tilt sensor. (2) Take a large number of photos at various angles in different scenes and record the tilt information at the same time. (Because the frame does not move when taking data, but the camera does, the tilt information recorded at this time is the rotation of the shooting plane relative to the installation area plane.) (3) Make labels. Use annotation software to divide the installation area as semantic segmentation labels. Convert the recorded tilt information into quaternion representation and invert it (convert it into the rotation of the installation area plane relative to the shooting plane) as the posture label. (4) Data enhancement. (5) Divide the data set.
[0085] The specific steps are:
[0086] (1) Preparing a data set: First, data needs to be collected. The data collection system 3 is composed of a binocular camera 301 and a tilt sensor 302 . The tilt sensor 302 is attached to one side of the binocular camera 301 .
[0087] (2) Align the image plane of the camera parallel to the plane of the installation area, and set the value of the tilt sensor 302 at this time to zero to achieve zero-point calibration of the rotation angle.
[0088] (3) The data acquisition system 3 is moved to a certain distance in front of the curtain wall installation area to be photographed, and the binocular camera 301 is used to obtain binocular RGB images of the installation area at various angles under different weather conditions, lighting conditions, distances, dust cover and other shooting environments.
[0089] (4) While shooting, it is necessary to obtain the value of the tilt sensor 302, which represents the rotation amount of the camera shooting plane relative to the curtain wall installation area plane.
[0090] (5) Use the labeling tool to perform pixel-level segmentation on the installation area to achieve accurate division of the installation area, which serves as the semantic segmentation label for subsequent training of the network.
[0091] (6) Perform data enhancement on the image, add data with different brightness and blur levels, and increase training samples.
[0092] (7) In actual situations, the shooting plane may be perpendicular to the plane of the installation area. If Euler angles are used to represent the rotation, a universal joint deadlock may occur, causing the rotation to lose one degree of freedom. Therefore, quaternions are used to represent the rotation. In one embodiment, the rotation angle obtained by the tilt sensor 302 is represented by Euler angles (α, β, γ) in the order of ZYX, which is converted to quaternion q m Indicates that for the quaternion q m Taking the inverse we get Represents the inverse transformation of the rotation, that is, the rotation of the curtain wall installation area relative to the camera shooting plane. As labels for pose estimation.
[0093]
[0094]
[0095] (8) Each set of data in the dataset contains two binocular RGB images and the corresponding installation area semantic segmentation label and pose estimation label.
[0096] (9) Take 70% of the dataset as the training set, 20% as the validation set, and 10% as the test set.
[0097] 2) Training the network, mainly including: (1) Training the network through supervised learning so that the network can obtain the preliminary results of the installation area through semantic segmentation (at this time, the segmentation results contain multiple areas and the area boundaries are blurred). (2) Screening the segmentation prediction results (area comparison, whether it contains the center point) to obtain the target area. (3) Fitting the target segmented area to obtain an area composed of straight lines, and obtaining the final divided area. (4) According to the division results, the coordinates of a corner point of the area and its diagonal pixel coordinate system (two-dimensional) can be obtained, and the coordinates of the corner point in the camera coordinate system (three-dimensional) can be obtained through the binocular principle. The midpoint of the two is calculated as the midpoint of the installation area. The coordinates of this point are the translation matrix T from the camera to the center of the installation area. (5) The left and right pictures obtained by segmentation and processing are spliced, and the network is trained through supervised learning. The loss function RLoss is designed so that the network can estimate the rotation amount q.
[0098] The specific steps include:
[0099] 1) First, semantic segmentation of the input RGB image is required. The left and right images of the binocular RGB image are sent to the network for training respectively. The image features of the binocular RGB image are obtained through the feature extraction network, and then deconvolution or upsampling is performed to restore the original image size to achieve semantic segmentation of the curtain wall installation area.
[0100] 2) Segmentation area screening. In the curtain wall installation scenario, multiple curtain wall borders will appear in one shooting range at the same time, so multiple areas may be divided in one RGB image through semantic segmentation. Compare the areas of the segmented areas, calculate the number of pixels contained in each area, and filter out the largest segmented area. Then judge each area to determine whether it contains the pixel center point, and select the segmented area containing the center point. Combine the two methods to get the target installation area.
[0101] 3) Refitting the installation area. Affected by factors such as lighting, the installation area obtained by segmentation may have errors with the actual area. The segmentation results are linearly fitted using pixel points. The RANSAC method can be used to randomly sample k points from the N pixels on each edge, fit the model to the k points, calculate the distance from other pixels to the fitting model and set a threshold, and count the points that are less than the threshold, i.e., the number of inliers. Based on the inliers, repeat the above steps to obtain a new fitting model, iterate M times, and select the model with the most inliers.
[0102]
[0103] Among them, p is the proportion of internal points, that is, the probability of internal points, ninliers is the number of internal points, n outliers The number of external points.
[0104] z=1-(1-p k ) M (6)
[0105] Where z is the probability that at least one of the k selected points is an inlier in M iterations. Usually, the model is considered accurate when z is greater than 0.95.
[0106] Thus, the error of the predicted installation area caused by environmental factors such as light is eliminated, and the final installation area consisting of straight line segments is obtained by fitting.
[0107] 4) Calculate the translation matrix T. Based on the segmentation results, select a corner point and the corner point opposite to it, and obtain the three-dimensional coordinates of the two corner points in the camera coordinate system through geometric transformation based on the two-dimensional coordinates of the corner points in the pixel coordinate system and the parameters of the binocular camera system. The midpoint of the two is the three-dimensional coordinate of the installation area center O. Then the translation matrix T from the camera to the installation area center O is obtained.
[0108] 5) Estimate the posture. Connect the left and right RGB images after segmentation in the viewing dimension V to form a four-dimensional input, with the dimension represented by H×W×C×V, where H, W, C, and V represent the height, width, channel, and viewing angle of the image, respectively. Use three-dimensional convolution for feature extraction, and use the loss function RLoss regression through the fully connected layer to obtain the rotation amount represented by the quaternion q.
[0109]
[0110] Among them, R(q) is the rotation matrix corresponding to the quaternion q estimated by the network, is the quaternion in the label The corresponding rotation matrix, n is the size of the training batch, and m is the number of samples in each batch.
[0111] 3. During the real-time pose estimation process, the pose relationship from the camera to the installation area is converted into the pose relationship from the plate to the installation area.
[0112] Purpose: This part is to convert the relative pose relationship (T, q) from the camera to the installation area into the relative pose relationship (T e ,q e ), converted into information that can be used by the curtain wall robot.
[0113] Method: Calibrate the position and posture of the camera to the gripper (plate) (T 0 ,q 0). The binocular image is obtained by real-time shooting, (T,q) is obtained through the network model, and (T e ,q e ). The specific steps are as follows:
[0114] 1) Deploy the network model to the curtain wall robot system.
[0115] 2) Calibrate the relative position relationship between the binocular camera 301 and the robot gripper, and obtain the translation matrix T of the gripper center P relative to the camera 0 , the quaternion representation of the rotation of the grabbing plane relative to the camera shooting plane q 0 .
[0116] 3) After the plate is grasped by binocular vision, a suitable shooting position is selected and the installation area is photographed using the binocular camera 301 to obtain a binocular RGB image. The installation area in the image is segmented using the trained network model, and the position of the installation area relative to the camera is estimated, including the translation matrix T and the rotation amount q represented by the quaternion.
[0117] 4) Use the calibration data and the network estimated data to convert the position information of the robot gripper to the installation area, including the translation matrix T e and the rotation q e , that is, the amount of posture that the robot needs to adjust. The relative posture relationship of each part is as follows Fig. 9 shown.
[0118] T e =TT 0 (8)
[0119]
[0120] 5) T e and q e Feedback is given to the curtain wall robot so that the robot can adjust the movement of the robotic arm according to the feedback amount. e Adjust the plate's posture so that the plate plane is parallel to the plane of the installation area; then adjust the plate's posture according to T e The provided position information moves the plate toward the center O of the installation area until it is 20 cm away, issues an anti-collision warning message and limits the movement speed of the robot arm to prevent collisions that may damage the plate or frame structure, and enables possible fine-tuning during the installation process, thereby achieving precise installation.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.
Claims
1. A posture estimation method for curtain wall panel installation and anti-collision warning, It is characterized in that The method specifically comprises the following steps: S1: Obtaining the precise grabbing position of the curtain wall panel: using a curtain wall panel storage device with special marking points to store the curtain wall panel, and using binocular vision to accurately locate the panel; the curtain wall panel storage device comprises a frame (1), an adjustment device and a fixing device (2); the fixing device (2) is provided with a special marking point (201); the curtain wall panel is placed upright on the frame (1) by the fixing device (2), and the special marking point (201) is located using binocular vision to achieve precise positioning of the curtain wall panel; S2: Use the data acquisition system to obtain the data set and divide it into training set, validation set and test set; The data acquisition system (3) comprises a binocular camera (301) and an inclination sensor (302), wherein the inclination sensor is attached to the side of the binocular camera; The specific steps to obtain the data set are: S21: before shooting, the camera shooting plane is placed parallel to the plane of the installation area, and the tilt sensor (302) is calibrated; S22: combining the installation area and the camera parameters, selecting a shooting distance range that can capture the entire installation area, shooting under different influencing factors, and acquiring a large number of binocular RGB images of the installation area; S23: It is necessary to obtain the value of the tilt sensor (302) at the same time of each shooting. The value is usually the Euler angle (α, β, γ) rotated in the order of ZYX, where α, β, γ respectively represent the angles of rotation of the tilt sensor (302) relative to the X, Y, and Z axes of the original coordinate system; the value recorded here represents the rotation amount of the camera shooting plane relative to the plane of the curtain wall installation area; S24: Perform pixel-level segmentation on the left and right binocular RGB images respectively, and use straight lines to divide the installation area as a label for the semantic segmentation task; S25: Enhance the binocular RGB image as the image of the data set and increase the training samples; S26: Convert Euler angles (α, β, γ) to quaternion q m To express the amount of rotation; for q m Taking the inverse we get Used to indicate the rotation amount of the curtain wall panel installation area plane relative to the shooting plane; Where w is q m The real part of the rotation angle determines the size of the rotation angle; x, y, z are q m The imaginary part determines the direction of the rotation transformation; (x, y, z) is the three-dimensional coordinate of the special marker point in the camera coordinate system; S27: Using Quaternions The posture information represented is used as the label of posture estimation to complete the establishment of the data set; each data includes two RGB images obtained by the binocular camera, the installation area segmentation label and the posture estimation label S28: After labeling, take 70% of the dataset as the training set, 20% as the validation set, and 10% as the test set; S3: Input the training set into the pose estimation network model based on computer vision and deep learning for training. The training termination time is confirmed by the loss curve of the validation set and the training set. Finally, the trained network model is used to estimate the relative pose relationship (T, q) between the camera and the installation area. S4: Real-time pose estimation: First, the relative pose relationship between the camera and the curtain wall robot gripper, i.e., the curtain wall panel, is calibrated to obtain the pose from the camera to the robot gripper (T 0 ,q 0 ); then use the calibrated data (T 0 ,q 0 ) and the data estimated by the pose estimation network model (T,q) are converted into the pose information (T e ,q e ).
2. The method for estimating a posture according to claim 1, It is characterized in that In step S1, the robot gripper accurately grasps the curtain wall panel position through binocular vision positioning; the special marking point (201) on the fixing device is used for feature matching, and the coordinates of the marking point on the left and right images are obtained in the pixel coordinate system (u l ,v l ) and (u r ,v r ); Take the left camera as the main camera, according to the focal length f of the camera itself, the horizontal baseline b, and the pixel density m of the image in the X and Y directions of the pixel coordinate system x 、m y and the origin of the image physical coordinate system (o x ,o y ), get the three-dimensional coordinates (x, y, z) of the special marker point in the camera coordinate system; Among them, f x 、f y Respectively represent the focal length in the X and Y directions; The accurate grabbing position estimation of the panel is achieved through the relative position relationship between the coordinates of the special marking points in the camera coordinate system and the center of the curtain wall panel.
3. The method for estimating a posture according to claim 1, It is characterized in that In step S3, training a posture estimation network model based on computer vision and deep learning specifically includes the following steps: S31: Train the network through supervised learning so that the network can obtain preliminary results of the installation area through semantic segmentation; S32: Screening the segmentation prediction results to obtain a target segmentation area; S33: fitting the target segmentation region to obtain a region composed of straight lines, thereby obtaining a final segmentation region; S34: according to the divided area, obtain the two-dimensional coordinates of a corner point of the area and its diagonal point in the pixel coordinate system, and obtain the three-dimensional coordinates of the corner point in the camera coordinate system through the binocular principle; calculate the midpoint of the two, that is, the midpoint of the installation area, and the coordinates of the point are the translation matrix T from the camera to the center of the installation area; S35: Splice the two segmented and processed left and right images, train the network through supervised learning, and design the loss function RLoss so that the network can estimate the rotation amount q.
4. The method for estimating a posture according to claim 3, It is characterized in that In step S33, the RANSAC method is used to fit the target segmentation area.
5. The method for posture estimation according to claim 1, It is characterized in that In step S4, the position information (T e ,q e ) is calculated as: in, Indicates q 0 The conjugate quaternion of .
6. The method for estimating a posture according to any one of claims 1 to 5, It is characterized in that T e and q e Feedback is given to the curtain wall grabbing robot, so that the robot adjusts the movement of the robotic arm according to the feedback amount; first according to q e Adjust the posture of the curtain wall panel so that the plane of the curtain wall panel is parallel to the plane of the installation area; then adjust the position of the curtain wall panel according to T e The provided position information moves the curtain wall panel toward the center O of the installation area until it is 20 cm away, issues an anti-collision warning message and limits the movement speed of the robotic arm to prevent collisions that may damage the panel or frame structure, and enables fine-tuning during the installation process to achieve precise installation.