Adaptive modular integrated building feature pose estimation method, device, electronic equipment and storage medium

CN121482148BActive Publication Date: 2026-09-01CHINA STATE CONSTR HAILONG TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511525318.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-09-01
Estimated Expiration
2045-10-23

AI Technical Summary

Technical Problem

一些相关技术融合红外热成像与可见光图像增强,但热特征受天气、日照、材料影响,稳定性差,额外传感器增加系统复杂度与能耗,实用性有限

Benefits of technology

[0012] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the adaptive modular integrated building feature pose estimation method described in any one of the first aspects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121482148B_ABST
    Figure CN121482148B_ABST
Patent Text Reader

Abstract

This application relates to the field of pose estimation technology, and in particular to an adaptive modular integrated building (MiC) feature pose estimation method and apparatus, comprising: acquiring multimodal data of the current modular integrated building (MiC) scene, the multimodal data including RGB images and depth data; determining target features and corresponding reference models adapted to the current MiC scene based on the multimodal data; locating the target features in the multimodal data through a feature detection network to obtain localization bounding boxes; generating segmentation masks of the target features through a segmentation network based on the localization bounding boxes; and inputting the reference model and segmentation mask into a pose estimation network to output the pose parameters of the target features. This allows for the introduction of cross-modal attention fusion, automatically matching the most suitable feature from candidate CAD models according to the environment, solving automation problems under multiple scenes and multiple features, improving the robustness of MiC across multiple scenes, achieving high-precision pose estimation for MiC features, and realizing end-to-end pose estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of pose estimation technology, and in particular to an adaptive modular integrated building feature pose estimation method, device, electronic device, and computer-readable storage medium. Background Technology

[0002] Currently, the assembly of modular integrated buildings (MiC) mainly relies on manual operation. Workers visually assess the position and orientation of the upper and lower modules and manually adjust the upper module to complete the assembly. The process is as follows: after the prefabricated modules are transported to the site, workers use cranes to lift the modules, visually judge the alignment of the connection points (such as steel beams or slots), and manually pull or fine-tune them to achieve precise alignment before completing the fixed connection. However, this position estimation is not accurate enough and the efficiency is relatively low.

[0003] In related technologies, some have attempted to apply computer vision methods to MiC assembly, but these methods have significant drawbacks. For example, segmentation techniques are used to distinguish different modules, components, or construction areas in a MiC factory, but this method requires a large amount of precise pixel-level annotation, such as the annotation of hundreds of images to train a Mask R-CNN model. Some related technologies propose a 3D pose estimation system that relies on traditional machine vision's 2D feature matching and Iterative Closest Point (ICP) algorithm for 3D model registration. However, compared to deep learning-based methods, these traditional vision methods perform poorly and cannot select an appropriate 3D model based on the scene. Some related machine vision methods utilize multi-camera images and invariant physical relationships to solve nonlinear equations for 3D pose estimation, which is sensitive to environment and hardware, requires multiple views and calibration, and has high calibration requirements. Some related technologies propose an improved ICP algorithm using edges and normals for 3D pose estimation, but it is prone to getting trapped in local optima, and fails when the continuous iteration error is significant and exceeds a threshold. Some related technologies employ binocular vision and marker-based localization for three-dimensional positioning, but they rely on image processing and recognition of marker points, and their effective ranging range is limited by the binocular baseline distance. Other technologies combine LiDAR with computer-aided design model matching for building structure monitoring and deviation detection. However, the system's component identification and mapping processes may be affected by the quality of point cloud data, and setting deviation thresholds requires high-precision data processing. Still other technologies integrate infrared thermal imaging with visible light image enhancement, but thermal characteristics are affected by weather, sunlight, and materials, resulting in poor stability. Additional sensors increase system complexity and energy consumption, limiting their practicality. Summary of the Invention

[0004] This application provides an adaptive modular integrated building feature pose estimation method, apparatus, electronic device, and computer-readable storage medium, which can improve the efficiency and accuracy of adaptive modular integrated building feature pose estimation.

[0005] To achieve the above objectives, this application adopts the following technical solution: Firstly, an adaptive modular integrated building feature pose estimation method is provided, including: Acquire multimodal data of the current modular integrated building (MiC) scene, including RGB images and depth data; Based on the multimodal data, target features and corresponding reference models adapted to the current MiC scene are determined; The target features in the multimodal data are located using a preset feature detection network to obtain the location bounding box; Based on the localization bounding box, a segmentation mask for the target feature is generated by a segmentation network;

[0006] The reference model and the segmentation mask are input into the pose estimation network, which outputs the pose parameters of the target feature.

[0007] Secondly, an adaptive modular integrated building feature pose estimation device is provided, comprising: The acquisition module is used to acquire multimodal data of the current modular integrated building (MiC) scene, including RGB images and depth data; The determination module is used to determine the target features and corresponding reference models that are adapted to the current MiC scene based on the multimodal data.

[0008] The localization module is used to locate target features in the multimodal data through a preset feature detection network to obtain a localization bounding box; The generation module is used to generate a segmentation mask for the target feature based on the localized bounding box through a segmentation network;

[0009] The output module is used to input the reference model and the segmentation mask into the pose estimation network and output the pose parameters of the target feature.

[0010] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the adaptive modular integrated building feature pose estimation method as described in any one of the first aspects above.

[0011] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the adaptive modular integrated building feature pose estimation method as described in any one of the first aspects above.

[0012] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the adaptive modular integrated building feature pose estimation method described in any one of the first aspects.

[0013] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.

[0014] In this embodiment, multimodal data of the current Modular Integrated Building (MiC) scene is first acquired, including RGB images and depth data. Then, based on the multimodal data, target features and corresponding reference models adapted to the current MiC scene are determined. Next, a pre-defined feature detection network is used to locate the target features in the multimodal data, obtaining localization bounding boxes. Based on these bounding boxes, a segmentation network generates a segmentation mask for the target features. Finally, the reference model and segmentation mask are input into a pose estimation network, which outputs the pose parameters of the target features. Therefore, by acquiring RGB images and depth data of the MiC scene, determining target features and reference models, locating them through a detection network, generating masks through a segmentation network, and then combining the reference model with the pose estimation network to output pose parameters, the accuracy of feature recognition can be improved, achieving precise pose calculation.

[0015] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0016] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiments below. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This is a flowchart illustrating the adaptive modular integrated building feature pose estimation method provided in the embodiments of this application; Figure 2 This is a schematic diagram of the adaptive scene feature selection network provided in the embodiments of this application; Figure 3 This is a schematic diagram of the system flow for controlling the adjustment of the hanger provided in an embodiment of this application; Figure 4 This is a structural block diagram of the adaptive modular integrated building feature pose estimation device provided in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0017] The embodiments of the technical solutions of this application will now be described in detail with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of this application, and are therefore merely examples and should not be used to limit the scope of protection of this application. When the following description relates to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Various changes, modifications, and equivalents of the methods, apparatus, and / or systems described herein will become apparent upon understanding this disclosure. For example, the order of operations described herein is merely illustrative and is not limited to those orders set forth herein, but can be changed as will become apparent upon understanding this disclosure, except for operations that must be performed in a specific order. Furthermore, for clarity and conciseness, descriptions of features known in the art may be omitted.

[0018] The embodiments described in the following examples of this disclosure are not representative of all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0019] It should be noted that the subject executing the adaptive modular integrated building feature pose estimation method in this embodiment can be an adaptive modular integrated building feature pose estimation device, hereinafter referred to as "device". This device can be configured in any type of electronic device, and this application embodiment does not limit it.

[0020] See Figure 1 This is a flowchart illustrating the adaptive modular integrated building feature pose estimation method provided in the first embodiment of this application. Figure 1 As shown, the adaptive modular integrated building feature pose estimation method may include the following steps: Step 101: Obtain multimodal data of the current modular integrated building (MiC) scene. The multimodal data includes RGB images and depth data.

[0021] The current modular integrated building (MiC) scenario refers to a building operation scenario in which the building is disassembled into standardized prefabricated modules (such as wall modules, floor modules, kitchen and bathroom modules, etc.) in the factory, and assembled on the construction site through hoisting and splicing. It includes factory prefabrication testing, on-site assembly construction, quality acceptance and other links.

[0022] Optionally, multimodal data of the MiC scene can be acquired using a depth camera (such as RealSense or ZED2), specifically RGB images and depth data.

[0023] (Capture the appearance, texture, and color of features) (Obtain the spatial structure and distance of features) Step 102: Based on multimodal data, determine the target features and corresponding reference models that are suitable for the current MiC scenario.

[0024] In this context, target features refer to physical structures or components that are adapted to the current MiC scenario and play a crucial role in the accuracy of module assembly. For example, in a factory MiC scenario, target features are "reinforcing bar connectors" and "module interface embedded parts." In a construction site MiC scenario, target features are "concrete module splicing joints" and "lifting point connection structures." In a laboratory simulation scenario, the target feature is "standard convex assembly connectors."

[0025] The reference model refers to a pre-constructed physical model that matches the target feature. For example, it can be a CAD model (computer-aided design model) of the target feature.

[0026] Optionally, multimodal data can be input into a pre-trained adaptive scene optimal feature selection network to obtain target features and corresponding reference models that are adapted to the current scene.

[0027] Specifically, the collected multimodal scene data including MiC is input into a pre-trained adaptive scene optimal feature selection network, which automatically performs scene analysis and feature matching. The adaptive scene optimal feature selection network first parses the input RGB image and depth data separately. It extracts the core features of both unimodal data through a built-in self-attention mechanism, such as the scene and MiC color and texture features in the RGB data, and the scene and MiC 3D structural features in the depth data. Then, a cross-modal attention mechanism fuses the two types of features to form multimodal features. Finally, a fully connected layer classifies and matches the multimodal features, outputting the most suitable target feature for the current scene and the corresponding reference model based on different scenes and MiCs.

[0028] Optionally, a training dataset can be obtained first, which includes MiC images from multiple scenes. These scenes cover different MiC module types, different lighting conditions, and different degrees of feature occlusion. Then, the MiC images from the training dataset are input into an initial adaptive feature selection network to extract unimodal features through a self-attention mechanism. Subsequently, the unimodal features are fused through a cross-modal attention mechanism to obtain multimodal features. The network outputs the most suitable feature and its corresponding CAD model through a fully connected layer until the accuracy of the output of the initial adaptive feature selection network meets the preset conditions. The initial adaptive feature selection network is then used as the optimal feature selection network for the adapted scene after training.

[0029] Specifically, it can collect MiC images and corresponding depth data from multiple scenes, covering different MiC module types, different lighting conditions (normal lighting, darkness, strong light), and different degrees of feature occlusion (no occlusion, partial occlusion, severe occlusion). The adaptive scene optimal feature selection network first collects the CAD (Design Alignment) of feature objects in different MiC scenes, and then selects the most suitable feature based on the scene and uses its CAD. During training, the labeling of the most suitable feature comes from: 1. If there is only one known feature in the scene image, then that feature is labeled as a positive sample, and the others as negative samples. 2. If there are more than one feature in the scene, then the segmentation mask of the feature is manually drawn, and FoundationPose is used to predict the pose estimation; the one with the higher accuracy is selected as the positive sample. The judgment method can be to visually observe the pose result (when the result is significantly better). If it cannot be judged visually (the pose results are close or difficult to judge), 2D points can be manually labeled and converted using PnP (Proof-of-Placement), and the closest one is selected as the positive sample.

[0030] Specifically, multimodal data from the training dataset can be input into the initial adaptive feature selection network. The network then iteratively trains by extracting unimodal features using a self-attention mechanism, fusing unimodal features using a cross-modal attention mechanism to obtain multimodal features, and finally outputting the most suitable feature and its corresponding CAD model from the fully connected layer. During training, data augmentation (such as generating nighttime scene data using the Cut method) is used to supplement the training data. After training, this network is used as the trained adaptive scene optimal feature selection network for subsequent determination of target features and reference models.

[0031] Figure 2This demonstrates the process of selecting a suitable CAD model based on multimodal data in a modular integrated building scenario. First, sensors acquire RGB images and depth maps from different MiC scenes. Next, this multimodal data enters the "modal fusion" module, integrating the features of the two datasets (the appearance and texture of the RGB images and the spatial structure of the depth maps). Then, the fused features are processed by a neural network, ultimately selecting the CAD model that best matches the current scene from multiple candidate CAD models on the right. This provides a standard reference model for subsequent target feature detection and pose estimation.

[0032] Step 103: Locate the target features in the multimodal data using a preset feature detection network to obtain the localization bounding box.

[0033] The feature detection network can be the YOLOv12 detection network.

[0034] Specifically, multimodal data containing target features can be input into a preset detection network. The feature detection network uses its pre-trained MiC feature recognition capability to locate the target features in the image and output a bounding box to define the location range of the feature in the image.

[0035] Optionally, the bounding box confidence score of the localized bounding box can be extracted. If the bounding box confidence score is less than a preset confidence threshold, it is determined to be an anomaly, the pose estimation process is stopped, and a pause action command is sent to the hybrid pose adjustment robot via the RS485 protocol.

[0036] Specifically, the bounding box confidence score corresponding to the localization bounding box can be extracted from the detection results. If the bounding box confidence score is less than a preset confidence threshold (e.g., 0.75), it is considered an anomaly, possibly due to lighting interference or feature occlusion causing inaccurate network recognition, and the subsequent pose estimation process is stopped. A pause command is sent to the hybrid pose adjustment robot via the RS485 protocol to control the robot to stop the current lifting or alignment operation, preventing assembly deviations or equipment malfunctions caused by positioning anomalies, and ensuring the safety and accuracy of operations in MiC scenarios.

[0037] Step 104: Based on the localized bounding box, generate a segmentation mask for the target features using a segmentation network.

[0038] The segmentation network can be the SAM2 segmentation network.

[0039] SAM2 (SegmentAnythingModel2) is an image segmentation model developed by Meta. After training on large-scale data, it can segment various types of image content, such as natural scenes, man-made objects, and medical images. It can achieve segmentation tasks in different image domains without requiring extensive customized training for specific datasets or tasks.

[0040] Specifically, the bounding box is used as input to the segmentation network as a cue. Taking SAM2 as an example, SAM2 relies on a pre-trained Transformer model, requiring no additional training for MiC scenes. Using bounding boxes as constraints, and combining multimodal data (texture and color of RGB images, spatial structure of depth data), it distinguishes target features from the background (such as module surfaces and environmental clutter). The segmentation network outputs a segmentation mask. The segmentation mask is a binary image with the same dimensions as the original image. "1" (or a specific identifier) ​​marks the pixel region of the target feature, and "0" (or a background identifier) ​​marks the non-feature region.

[0041] Step 105: Input the reference model and segmentation mask into the pose estimation network and output the pose parameters of the target features.

[0042] Specifically, the reference model (CAD model) of the target feature and the segmentation mask generated by the segmentation network are input together into the pose estimation network (such as FoundationPose). Through feature matching and error optimization, the pose parameters of the feature (including spatial coordinates and rotation angles) are output, providing accurate data support for the automated assembly of MiC modules.

[0043] In the MiC scene pose estimation stage, the reference model of the target feature and the segmentation mask (pixel-level markers of feature regions, which, combined with depth data, can be mapped to the actual 3D point cloud) are first input into the pose estimation network. The network first extracts and matches the key features of the reference model and the actual 3D point cloud, and obtains the initial pose through algorithms such as PnP; then, using the segmentation mask as a constraint, the initial pose is iteratively optimized to reduce the error, and finally outputs the pose parameters of the target feature, including spatial position and orientation. These parameters can be transmitted to the robot controller via the RS485 protocol to guide the precise assembly of the MiC module.

[0044] Optionally, an RS485 protocol is used to establish a communication connection with the hybrid pose adjustment robot, converting the pose parameters into data that the hybrid pose adjustment robot can recognize. The recognizable data includes: hanger rotation control parameters and translation control parameters. The hanger rotation control parameters at least cover the rotation direction and rotation angle; the translation control parameters at least include the translation distance in six directions: forward, backward, left, right, up, and down. Based on the recognizable data, the hybrid pose adjustment robot is driven to perform pose adjustment actions on the target feature.

[0045] Optionally, this solution establishes a stable communication connection with the hybrid pose adjustment robot via the RS485 protocol to realize the conversion and execution of pose parameters into robot control commands: the pose parameters (position coordinates, rotation angle) of the target feature output by the pose estimation network are converted into control parameters that the robot can recognize. The hanger rotation control parameters include at least the rotation direction (e.g., clockwise / counterclockwise) and rotation angle, used to adjust the spatial orientation of the hanger. The translation control parameters cover the translation distance (accurate to 0.1mm) in six directions: front, back, left, right, up, and down, used to correct the spatial position of the hanger. The control parameters are transmitted to the robot controller in real time via the RS485 protocol. After parsing the parameters, the controller drives the actuator of the hybrid pose adjustment robot, driving the hanger and MiC module to complete the pose adjustment, achieving precise alignment between the module and the target feature.

[0046] Optionally, after continuously outputting N frames of pose parameters, the deviation value of the pose parameters between adjacent frames or between consecutive N frames is calculated, where N is a preset value. If the deviation value is greater than the preset deviation threshold, it is determined that the pose is unstable, the pose parameter output is stopped, and a stop check command is sent to the hybrid pose adjustment robot through the RS485 protocol until the deviation value is restored to within the preset deviation threshold.

[0047] Optionally, to ensure assembly safety and accuracy, the solution adds a pose deviation monitoring mechanism, the specific process of which is as follows: After continuously outputting N frames (N is a preset value, such as 3-5 frames) of pose parameters, the deviation value of the pose parameters between adjacent frames or within N consecutive frames (including position coordinate deviation and rotation angle deviation) is calculated. If the deviation value is greater than the preset deviation threshold (such as position deviation > 1cm, angle deviation > 1°), it is determined that the pose is unstable, and the pose parameter output is stopped immediately; a stop check command is sent to the hybrid pose adjustment robot through the RS485 protocol, and the robot pauses the current adjustment action; after the interference factors are investigated and the deviation value is restored to within the preset threshold, the pose estimation and adjustment are restarted.

[0048] Figure 3 This diagram illustrates the system flow for target feature detection, segmentation, pose estimation, and scaffold adjustment control based on multimodal data in a modular integrated building (MiC) scenario. Sensors collect data from the MiC scene, which is first filtered by an adaptive scene feature selection network to select suitable target features. Next, a feature detection and segmentation network detects and segments the target features, obtaining the detection and segmentation results. Simultaneously, the feature CAD (reference model of the target feature) is input into a pose estimation network to calculate the pose results (module position and orientation). Finally, the pose results are transmitted to the controller, which issues commands to adjust the scaffold via motors to achieve precise assembly of the MiC modules.

[0049] In this embodiment, multimodal data of the current Modular Integrated Building (MiC) scene is first acquired, including RGB images and depth data. Then, based on the multimodal data, target features and corresponding reference models adapted to the current MiC scene are determined. Next, a preset feature detection network is used to locate the target features in the multimodal data, obtaining localization bounding boxes. Based on these bounding boxes, a segmentation network generates a segmentation mask for the target features. Finally, the reference model and segmentation mask are input into a pose estimation network, which outputs the pose parameters of the target features. Therefore, by acquiring RGB images and depth data of the MiC scene, determining target features and reference models, locating them through a detection network, generating masks through a segmentation network, and then combining the reference model with the pose estimation network to output pose parameters, the accuracy of feature perception can be improved, enabling precise pose calculation and adapting to the industrial automation needs of MiC.

[0050] In some embodiments, the automatic estimation of the pose of feature objects such as the MiC module is achieved by sequentially using an adaptive scene feature selection network, a YOLOv12 detection network, a SAM2 segmentation network, and a FoundationPose pose estimation network, which has new innovative points: 1. Adaptive scene feature selection mechanism: Introducing cross-modal attention fusion (RGB image + depth image), automatically matching the most suitable feature from candidate CAD models according to the environment, solving the automation problem under multiple scenes and multiple features, improving the robustness of MiC in multiple scenes, and applying it to building assembly pose estimation.

[0051] 2. Highly efficient cascaded deep learning framework: Combining YOLOv12 real-time detection, SAM2 zero-shot segmentation model and FoundationPose model to drive pose estimation model, achieving high-precision pose estimation for MiC features, realizing end-to-end 6D pose estimation, with strong practicality.

[0052] 3. Low-cost multimodal perception integration: It only relies on RGB-D depth cameras (such as RealSense), without the need for high-cost LiDAR, multi-camera systems and additional manual markers; it does not require pixel-level annotation training of traditional segmentation networks, significantly reducing data preparation costs.

[0053] 4. Safety and fault tolerance and robot closed-loop control: The robot drives the hybrid pose adjustment through real-time transmission of rotation and translation parameters via RS485 protocol; it sets confidence thresholds and multi-frame deviation monitoring, and stops working when unreasonable predictions occur, thus preventing abnormal risks during the precise assembly of MiC modules and improving safety.

[0054] Corresponding to the adaptive modular integrated building feature pose estimation method described in the above embodiments, Figure 4This is a structural block diagram of the adaptive modular integrated building feature pose estimation device provided in the embodiments of this application.

[0055] Reference Figure 4 The adaptive modular integrated building feature pose estimation device 400 includes: The acquisition module 410 is used to acquire multimodal data of the current modular integrated building (MiC) scene, wherein the multimodal data includes RGB images and depth data; The determination module 420 is used to determine the target feature objects and corresponding reference models that are adapted to the current MiC scene based on the multimodal data;

[0056] The positioning module 430 is used to locate the target features in the multimodal data through a preset feature detection network to obtain the positioning bounding box; The generation module 440 is used to generate a segmentation mask for the target feature based on the positioning bounding box and through a segmentation network.

[0057] The output module 450 is used to input the reference model and the segmentation mask into the pose estimation network and output the pose parameters of the target feature.

[0058] Optionally, the determining module is specifically used for: The multimodal data is input into a pre-trained adaptive scene optimal feature selection network to obtain the target feature and the corresponding reference model that are adapted to the current scene.

[0059] Optionally, the reference model is a CAD model, and the determining module is further configured to: Obtain a training dataset, which includes MiC images from multiple scenes, covering scenes with different MiC module types, different lighting conditions, and different degrees of feature occlusion. The MiC images in the training dataset are input into an initial adaptive feature selection network to extract unimodal features through a self-attention mechanism; Multimodal features are obtained by fusing the single-modal features through a cross-modal attention mechanism. The most suitable feature object and its corresponding CAD model are output through a fully connected layer until the accuracy of the output result of the initial adaptive feature object selection network meets the preset conditions. The initial adaptive feature object selection network is then used as the trained optimal feature object selection network for the adaptive scene.

[0060] Optionally, the output module is also used for:

[0061] A communication connection was established with the hybrid pose adjustment robot using the RS485 protocol. The pose parameters are converted into data that the hybrid pose adjustment robot can recognize. The recognizable data includes: hanger rotation control parameters and translation control parameters, wherein the hanger rotation control parameters at least cover the rotation direction and rotation angle; the translation control parameters at least include translation distances in six directions: forward, backward, left, right, up, and down. Based on the identifiable data, the hybrid pose adjustment robot is driven to perform pose adjustment actions on the target feature.

[0062] Optionally, the positioning module 430 is also used for: Extract the bounding box confidence score of the localized bounding box; If the confidence level of the bounding box is less than the preset confidence threshold, it is determined to be an anomaly, and the pose estimation process is stopped. A pause command is sent to the hybrid pose adjustment robot via the RS485 protocol.

[0063] Optionally, the output module is also used for:

[0064] After continuously outputting N frames of pose parameters, the deviation of pose parameters between adjacent frames or between consecutive N frames is calculated, where N is a preset value. If the deviation value is greater than the preset deviation threshold, it is determined that the pose is unstable, and the pose parameter output is stopped. A stop check command is sent to the hybrid pose adjustment robot via RS485 protocol until the deviation value returns to within the preset deviation threshold.

[0065] In this embodiment, multimodal data of the current Modular Integrated Building (MiC) scene is first acquired, including RGB images and depth data. Then, based on the multimodal data, target features and corresponding reference models adapted to the current MiC scene are determined. Next, a preset feature detection network is used to locate the target features in the multimodal data, obtaining localization bounding boxes. Based on these bounding boxes, a segmentation network generates a segmentation mask for the target features. Finally, the reference model and segmentation mask are input into a pose estimation network, which outputs the pose parameters of the target features. Therefore, by acquiring RGB images and depth data of the MiC scene, determining target features and reference models, locating them through a detection network, generating masks through a segmentation network, and then combining the reference model with the pose estimation network to output pose parameters, the accuracy of feature perception can be improved, enabling precise pose calculation and adapting to the industrial automation needs of MiC.

[0066] in addition, Figure 4The adaptive modular integrated building feature pose estimation device shown can be a software unit, hardware unit, or a combination of software and hardware built into existing electronic devices, or it can be integrated into the electronic devices as an independent component, or it can exist as an independent electronic device.

[0067] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0068] Figure 5 This is a schematic diagram of the structure of the electronic device provided in an embodiment of this application. For example... Figure 5 As shown, the electronic device 5 of this embodiment includes: at least one processor 50 ( Figure 5 (Only one is shown in the diagram) a processor, a memory 51, and a computer program 52 stored in the memory 51 and executable on the at least one processor 50, wherein the processor 50 executes the computer program 52 to implement the steps in any of the above embodiments of the adaptive modular integrated building feature pose estimation method.

[0069] The electronic device may be a desktop computer, laptop, handheld computer, or cloud server, etc. This electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 5 This is merely an example of electronic device 5 and does not constitute a limitation on electronic device 5. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0070] The processor 50 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0071] In some embodiments, the memory 51 may be an internal storage unit of the electronic device 5, such as a hard disk or memory of the electronic device 5. In other embodiments, the memory 51 may be an external storage device of the electronic device 5, such as a plug-in hard disk, SmartMediaCard (SMC), SecureDigital (SD) card, or FlashCard equipped on the electronic device 5. Furthermore, the memory 51 may include both internal and external storage units of the electronic device 5. The memory 51 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory 51 can also be used to temporarily store data that has been output or will be output.

[0072] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps in the above-described method embodiments.

[0073] This application provides a computer program product that, when run on an electronic device, enables the electronic device to implement the steps described in the various method embodiments above.

[0074] If the integrated unit is implemented as a software functional unit and used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.

[0075] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0076] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0077] In the embodiments provided in this application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings or direct couplings or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0078] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0079] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.

[0080] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.

[0081] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0082] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.

[0083] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), similarly, "multiple sets" refers to two or more (including two sets), and "multiple pieces" refers to two or more (including two pieces).

[0084] In the description of the embodiments of this application, the technical terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "radial," and "circumferential" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing the embodiments of this application and simplifying the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the embodiments of this application.

[0085] In the description of the embodiments of this application, unless otherwise expressly specified and limited, technical terms such as "installation," "connection," "joining," and "fixing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. For those skilled in the art, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.

[0086] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. An adaptive modular integrated building feature pose estimation method, characterized in that, include: Acquire multimodal data of the current modular integrated building (MiC) scene, including RGB images and depth data; Based on the multimodal data, target features and corresponding reference models adapted to the current MiC scene are determined; The target features in the multimodal data are located using a preset feature detection network to obtain the location bounding box; Based on the localization bounding box, a segmentation mask for the target feature is generated by a segmentation network; The reference model and the segmentation mask are input into the pose estimation network to output the pose parameters of the target feature. The step of determining the target feature and its corresponding reference model adapted to the current MiC scene based on the multimodal data includes: The multimodal data is input into a pre-trained adaptive scene optimal feature selection network to obtain the target feature and the corresponding reference model that are adapted to the current scene. The reference model is a CAD model. Before inputting the multimodal data into a pre-trained adaptive scene optimal feature selection network to obtain the target feature and corresponding reference model that are adapted to the current target feature, the method further includes: Obtain a training dataset, which includes MiC images from multiple scenes, covering scenes with different MiC module types, different lighting conditions, and different degrees of feature occlusion. The MiC images in the training dataset are input into an initial adaptive feature selection network to extract unimodal features through a self-attention mechanism; Multimodal features are obtained by fusing the single-modal features through a cross-modal attention mechanism. The most suitable feature object and its corresponding CAD model are output through a fully connected layer until the accuracy of the output result of the initial adaptive feature object selection network meets the preset conditions. The initial adaptive feature object selection network is then used as the trained optimal feature object selection network for the adaptive scene.

2. The method according to claim 1, characterized in that, After inputting the reference model and the segmentation mask into the pose estimation network and outputting the pose parameters of the target feature, the method further includes: A communication connection was established with the hybrid pose adjustment robot using the RS485 protocol. The pose parameters are converted into data that the hybrid pose adjustment robot can recognize. The recognizable data includes: hanger rotation control parameters and translation control parameters, wherein the hanger rotation control parameters at least cover the rotation direction and rotation angle; the translation control parameters at least include translation distances in six directions: forward, backward, left, right, up, and down. Based on the identifiable data, the hybrid pose adjustment robot is driven to perform pose adjustment actions on the target feature.

3. The method according to claim 1, characterized in that, After locating the target features in the multimodal data using a preset feature detection network to obtain the localization bounding box, the method further includes: Extract the bounding box confidence score of the localized bounding box; If the confidence level of the bounding box is less than the preset confidence threshold, it is determined to be an anomaly, and the pose estimation process is stopped. A pause command is sent to the hybrid pose adjustment robot via the RS485 protocol.

4. The method according to claim 1, characterized in that, After inputting the reference model and the segmentation mask into the pose estimation network and outputting the pose parameters of the target feature, the method further includes: After continuously outputting N frames of pose parameters, the deviation of pose parameters between adjacent frames or between consecutive N frames is calculated, where N is a preset value. If the deviation value is greater than the preset deviation threshold, it is determined that the pose is unstable, and the pose parameter output is stopped. A stop check command is sent to the hybrid pose adjustment robot via RS485 protocol until the deviation value returns to within the preset deviation threshold.

5. An adaptive modular integrated building feature pose estimation device, characterized in that, include: The acquisition module is used to acquire multimodal data of the current modular integrated building (MiC) scene, including RGB images and depth data; The determination module is used to determine the target features and corresponding reference models that are adapted to the current MiC scene based on the multimodal data. The localization module is used to locate target features in the multimodal data through a preset feature detection network to obtain a localization bounding box; The generation module is used to generate a segmentation mask for the target feature based on the localized bounding box through a segmentation network; The output module is used to input the reference model and the segmentation mask into the pose estimation network and output the pose parameters of the target feature. The determination module is specifically used for: The multimodal data is input into a pre-trained adaptive scene optimal feature selection network to obtain the target feature and the corresponding reference model that are adapted to the current scene. The reference model is a CAD model. Before inputting the multimodal data into a pre-trained adaptive scene optimal feature selection network to obtain the target feature and corresponding reference model that are adapted to the current target feature, the method further includes: Obtain a training dataset, which includes MiC images from multiple scenes, covering scenes with different MiC module types, different lighting conditions, and different degrees of feature occlusion. The MiC images in the training dataset are input into an initial adaptive feature selection network to extract unimodal features through a self-attention mechanism; Multimodal features are obtained by fusing the single-modal features through a cross-modal attention mechanism. The most suitable feature object and its corresponding CAD model are output through a fully connected layer until the accuracy of the output result of the initial adaptive feature object selection network meets the preset conditions. The initial adaptive feature object selection network is then used as the trained optimal feature object selection network for the adaptive scene.

6. An electronic device, characterized in that, It includes a memory and a processor; the memory stores a computer program, and the processor is configured to run the computer program in the memory to perform the method according to any one of claims 1 to 4.

7. A storage medium, characterized in that, The storage medium is used to store a computer program, which is loaded by a processor to perform the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Industrial part assembling method and equipment and medium

    CN120244532A