Scene segmentation method based on man-machine cooperation, terminal equipment and storage medium
Through a scene segmentation method based on human-machine collaboration, data is collected using multiple robots and drones, combined with semantic segmentation models and manual annotation, efficient three-dimensional mapping and scene segmentation are achieved, solving the problems of low efficiency, high cost and data processing limitations in the existing technology.
Patent Information
- Application Number
- CN202510120819.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-25
- Publication Date
- 2025-06-20
AI Technical Summary
The existing three-dimensional mapping and scene segmentation technologies have problems such as time-consuming and labor-intensive, difficult to adapt to dynamic and complex indoor environments, high costs and point cloud data processing limitations, and it is difficult to effectively carry out three-dimensional mapping and improve scene segmentation performance.
A scene segmentation method based on human-machine collaboration is adopted, and observation data is collected through multiple robots, and an improved multi-robot active graph building algorithm is used to build an initial three-dimensional point cloud. Combined with the data of the hole area of the drone, the semantic segmentation model is used for segmentation, and the model is optimized by manual labeling to realize three-dimensional graph building and scene segmentation of human-machine collaboration.
It significantly improves the efficiency of three-dimensional map construction, ensures the integrity of map construction, and combines human-computer collaboration to improve scene segmentation performance, solving the problems of data obsolete, missing labels and incomplete scene maps.
Smart Images

Figure CN120182589A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of scene segmentation, and in particular, to a scene segmentation method, a terminal device, and a storage medium based on human-machine collaboration. Background Art
[0002] Indoor 3D mapping and scene segmentation are regarded as fundamental elements for many applications including digital twins. They are playing an increasingly important role in digital twins, military security, and building information modeling (BIM). In addition, semantically rich indoor 3D maps are the basis for various upstream applications, including indoor navigation, human-computer interaction, and intelligent space applications. Given the growing demand for these applications, it has become increasingly important to develop an efficient and effective mapping and scene segmentation method.
[0003] However, in terms of 3D mapping, traditional manual measurement and computer-aided design (CAD) modeling methods are not only time-consuming and laborious, but also difficult to adapt to dynamic and complex indoor environments. Widely used 3D laser point cloud scanning devices are costly, require manual operation by engineers in indoor environments, and are subject to limitations such as floor installation and height, often resulting in missing point clouds. Current scene segmentation techniques mainly rely on deep neural networks to extract semantic features of point cloud data to achieve more accurate segmentation results. However, due to the irregular and unstructured nature of point clouds, there are significant limitations in using deep networks to process 3D point cloud data. In addition, existing scene segmentation algorithms usually rely on limited public datasets and are difficult to handle complex indoor scenes in the real world. How to effectively perform indoor 3D mapping and improve the performance of scene segmentation has become a major challenge faced by the current academic and industrial communities. Against this background, there is an urgent need for an efficient, low-cost, and comprehensive 3D mapping and scene segmentation method. Summary of the Invention
[0004] To solve the above problems, the present invention proposes a scene segmentation method, a terminal device, and a storage medium based on human-machine collaboration.
[0005] The specific solutions are as follows:
[0006] A scene segmentation method based on human-machine collaboration includes the following steps:
[0007] S1: Collect observation data of the research scene through multiple robots, and construct an initial 3D point cloud of the research scene through an improved multi-robot active mapping algorithm based on neural bipartite graph matching; the specific improvements are as follows:
[0008] Add a dimension representing height to the collected observation data on the basis of the 2D map;
[0009] Replace the 2D convolutional layer used in the original mGNN with a 3D convolutional layer;
[0010] Replace the observation space of the original reinforcement learning from a two-dimensional plane with a three-dimensional space;
[0011] Replace the calculation of the reward function of the original reinforcement learning based on the grid increment of the two-dimensional map with the calculation based on the voxel increment of the three-dimensional map;
[0012] Add a three-dimensional cuboid bounding box based on the size of the robot, and dynamically update the position of the bounding box. Filter out dynamic obstacles by deleting the data of each voxel inside the bounding box in the point cloud;
[0013] S2: Based on the initial three-dimensional point cloud, calculate the point cloud holes therein, and scan the holes by a drone. After complementing the point cloud holes based on the scan data, obtain a complete three-dimensional point cloud;
[0014] S3: Perform semantic segmentation on the complete three-dimensional point cloud through a semantic segmentation model;
[0015] S4: Manually annotate the point cloud data with low confidence in the semantic segmentation results;
[0016] S5: Based on the manual annotation results, optimize the semantic segmentation model through active learning, and then return to S3 to perform semantic segmentation on the complete three-dimensional point cloud again through the optimized semantic segmentation model until there is no point cloud data with low confidence, and then output the semantic segmentation results.
[0017] Furthermore, when scanning the hole area by the drone, the scanning order is sorted based on the priority of all hole areas; the priority setting method of the hole area is as follows:
[0018] First, set the priority according to the size of the hole area. The larger the hole area, the higher the priority;
[0019] Second, set the priority according to the confidence level of the point cloud around the hole. The lower the confidence level of the hole area, the higher the priority;
[0020] Finally, set the priority according to the path length of the drone to reach the hole area. The shorter the path length of the hole area, the higher the priority.
[0021] Furthermore, when scanning the hole area by the drone, the scanning order is sorted based on the priority of all hole areas; the priority setting method of the hole area is as follows:
[0022] First, set the priority according to the size of the hole area. The larger the hole area, the higher the priority;
[0023] Second, set the priority according to the confidence level of the point cloud around the hole. The lower the confidence level of the hole area, the higher the priority;
[0024] Finally, the priority is set according to the path length of the drone reaching the hole area, and the hole area with a shorter path length has a higher priority.
[0025] Furthermore, the multiple robots adopted are multi-source heterogeneous robots.
[0026] A scene segmentation terminal device based on human-machine collaboration includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method described above in the embodiments of the present invention are implemented.
[0027] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described above in the embodiments of the present invention are implemented.
[0028] The present invention adopts the above technical solutions, which can significantly improve the mapping efficiency on the premise of ensuring the integrity of mapping in 3D mapping. In terms of scene segmentation, the active supplementary scanning of point cloud data is combined with manual annotation, realizing human-machine collaboration in indoor 3D mapping and scene segmentation, and solving problems such as data obsolescence, label missing, incomplete scene maps, and limited open-source data sets. Description of the Drawings
[0029] Figure 1 The flowchart of the method in Embodiment 1 of the present invention is shown. Detailed Embodiments
[0030] To further illustrate each embodiment, the present invention provides drawings. These drawings are part of the disclosure of the present invention, mainly used to illustrate the embodiments, and can be used to explain the operating principle of the embodiments in combination with the relevant descriptions in the specification. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention.
[0031] The present invention will be further described below in combination with the drawings and specific embodiments.
[0032] Embodiment 1:
[0033] The embodiment of the present invention provides a scene segmentation method based on human-machine collaboration, as Figure 1 shown, the method includes the following steps:
[0034] S1: Collect the observation data of the research scene through multiple robots, and construct the initial 3D point cloud of the research scene through an improved multi-robot active mapping algorithm based on neural bipartite graph matching.
[0035] In this embodiment, the multiple robots used are multi-source heterogeneous robots, including types such as guiding robots, floor-sweeping robots, and robotic dogs. The observation data collected by the robots includes the RGB images and depth maps captured by the cameras at each time step t.
[0036] The traditional multi-robot active mapping algorithm based on neural bipartite graph matching is the algorithm adopted in the paper "Multi-Robot Active Mapping via Neural Bipartite Graph Matching" jointly proposed by the research teams of Peking University, Shandong University, Tencent AI Lab, Tsinghua University, and Stanford University. Compared with the traditional algorithm, the following improvements are made in this embodiment:
[0037] (1) Modification of the data unit format
[0038] The traditional algorithm only retains the two-dimensional map data of the plane information (x-axis and y-axis), ignoring the height information. Now, through improvement, three-dimensional map data is introduced, and a new z-axis is added to represent the height dimension, upgrading to a three-dimensional map structure including the x-axis, y-axis, and z-axis. This improvement retains more detailed environmental features and information, providing a more comprehensive data basis for subsequent mapping and decision-making.
[0039] (2) Improvement of the network structure
[0040] In the mGNN network structure adopted by the original algorithm, two-dimensional convolutional layers (Conv2d) are used to process two-dimensional map data, which is suitable for simple scenarios that only contain plane information. This network gradually extracts two-dimensional features through multiple layers of two-dimensional convolution (the typical convolution kernel size is 6×6, stride = 2), and reduces the dimension to the final output through activation functions and fully connected layers. Considering the limited expressive ability of two-dimensional convolution, it is difficult to capture the depth information and complex structures in three-dimensional space, so it performs inadequately when dealing with three-dimensional scenarios.
[0041] In this embodiment, to address the above problems, three-dimensional convolutional layers (Conv3d) are used to replace the original two-dimensional convolutional layers, which can synchronously extract features in length, width, and height (i.e., the three-dimensional space structure). For example, in Conv3d(6, 64, (6, 6, 3), stride = (2, 2, 1), padding = (2, 2, 2)), the convolution kernel size is designed as (6, 6, 3), and stride and padding are (2, 2, 1) and (2, 2, 2) respectively, so that the three-dimensional space features can maintain continuity while avoiding excessive compression of the depth dimension (z-axis). This improvement significantly enhances the network's adaptability to three-dimensional data, can capture richer spatial features, and effectively process complex three-dimensional scenarios containing depth information.
[0042] (3) Improvement of the observation space for reinforcement learning
[0043] The observation space of the traditional algorithm is a two-dimensional plane (120×120). In this embodiment, it is improved to a three-dimensional observation space (120×120×30) to fully reflect the geometric structure and object distribution of the three-dimensional environment, thereby better supporting the learning and exploration of the intelligent agent.
[0044] (4) Improvement of the calculation method of reinforcement learning reward
[0045] The reward function of the traditional algorithm is calculated based on the grid increment of the two-dimensional map, reflecting the progress of exploring the plane area. In this embodiment, it is improved to calculate the reward based on the voxel increment of the three-dimensional map, which not only takes into account the exploration situation on the plane, but also introduces the change of the depth dimension, which can more comprehensively measure the exploration efficiency in the three-dimensional space.
[0046] (5) Improvement of dynamic obstacle filtering
[0047] Multi-robot scenarios create a highly dynamic environment, and moving robots may be observed by other robots. This brings new challenges to obstacle-free path planning and point cloud registration. In terms of point cloud registration, too many dynamic obstacles may affect the accuracy and robustness of registration, or even cause registration failure. During the mapping stage, dynamic obstacles may leave false "ghosts" (shadows left when the robot moves) in the map, which will affect subsequent map processing and path planning. Dynamic obstacle filtering based on RGB image object detection is usually limited by distance factors and camera resolution. For example, when the robot appears small in the RGB image, the object detection algorithm may not be able to correctly identify the object. However, in the depth image, the robot is visible, so "ghosts" may appear in the generated point cloud.
[0048] In the process of mapping, the traditional algorithm robot will mark the plane grid where it is located as occupied, and restore the grid to an unoccupied state after the robot moves away. However, in a three-dimensional environment, filtering out dynamic obstacles is more complicated. Therefore, this embodiment sets a three-dimensional rectangular bounding box based on the size of the robot, dynamically updates the position of the bounding box, and deletes the voxel data in the corresponding bounding box in the world coordinate system point cloud (that is, filters out the point cloud information of all robots in the point cloud merged at the current time step), thereby achieving accurate filtering of three-dimensional dynamic obstacles. This method effectively reduces the occurrence of mislabeling and improves the accuracy of map modeling and the ability to adapt to dynamic environments.
[0049] S2: Based on the initial 3D point cloud, the point cloud holes are calculated and scanned by a drone. The point cloud holes are completed based on the scanned data to obtain a complete 3D point cloud.
[0050] In the above-mentioned active mapping process, due to the limitation of the robot's own field of view, some visual blind spots will appear, forming blank areas on the map, mainly concentrated on the surfaces of furniture such as desktops, beds and stoves. The appearance of these hole areas will have an adverse effect on the subsequent semantic segmentation. First, the insufficient point cloud data in the hole area destroys the integrity of the scene, interrupts the continuity of some object surfaces, and brings greater difficulties to semantic understanding. These areas cannot be effectively segmented and errors will occur. Secondly, the sudden change of the point cloud distribution near the hole will also interfere with the judgment of the segmentation algorithm and lead to misclassification. To solve this problem, it is necessary to enhance the visual coverage of the robot mapping, scan from multiple angles, and obtain a continuous, complete and accurate map.
[0051] In order to fill in the blank areas left in the above-mentioned mapping process, a drone is introduced in this embodiment to perform targeted vulnerability scanning. Three factors are considered when planning the drone's re-scanning path. First, consider the size of the hole area; larger areas should be given priority. Second, consider the confidence level of the point cloud around the hole; low confidence means that the semantic information is not clear enough, and these areas should be scanned first to enhance the accuracy of the point cloud. Third, consider the length of the path that the drone takes to reach the hole area. When the first two conditions are met, select the order of shorter path lengths for scanning to shorten the time for a complete re-scan. In this way, the drone can efficiently fill in the holes left by the above-mentioned mapping process, obtain a complete point cloud model of the scene, provide reliable and accurate point cloud data support for subsequent semantic segmentation, make up for the robot's own visual limitations, and realize model construction and comprehensive optimization.
[0052] S3: Perform semantic segmentation on the complete 3D point cloud through the semantic segmentation model.
[0053] S4: Manually annotate point cloud data with low confidence in semantic segmentation results.
[0054] S5: Based on the manual annotation results, after optimizing the semantic segmentation model through active learning, return to S3 and re-segment the complete 3D point cloud through the optimized semantic segmentation model until there is no point cloud data with low confidence, and then output the semantic segmentation result.
[0055] The method of judging low confidence is to regard the confidence lower than the set confidence threshold as low confidence, and the threshold can be determined according to the actual segmentation effect.
[0056] Active completion through manual annotation results and continuous injection of fresh data and labels greatly improves the accuracy of the semantic segmentation model.
[0057] Embodiment 2:
[0058] The present invention also provides a scene segmentation terminal device based on human-machine collaboration, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the above method embodiment of Embodiment 1 of the present invention are implemented.
[0059] Further, as an executable solution, the scene segmentation terminal device based on human-machine collaboration may be a computing device such as a desktop computer, a notebook, a palm computer, or a cloud server. The scene segmentation terminal device based on human-machine collaboration may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the composition structure of the above scene segmentation terminal device based on human-machine collaboration is only an example of the scene segmentation terminal device based on human-machine collaboration, and does not constitute a limitation on the scene segmentation terminal device based on human-machine collaboration. It may include more or fewer components than the above, or combine some components, or different components. For example, the scene segmentation terminal device based on human-machine collaboration may further include input / output devices, network access devices, a bus, etc. The embodiments of the present invention do not make limitations in this regard.
[0060] Further, as an executable solution, the so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the scene segmentation terminal device based on human-machine collaboration, and connects various parts of the entire scene segmentation terminal device based on human-machine collaboration through various interfaces and lines.
[0061] The memory can be used to store the computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory, the processor can implement various functions of the human-machine collaboration-based scene segmentation terminal device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include high-speed random access memory and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.
[0062] The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above method in the embodiments of the present invention are implemented.
[0063] If the modules / units integrated in the human-machine collaboration-based scene segmentation terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution media, etc.
[0064] Although the present invention is specifically shown and described in combination with the preferred embodiments, those skilled in the art should understand that various changes can be made to the present invention in form and detail without departing from the spirit and scope of the present invention defined by the appended claims, and all are within the protection scope of the present invention.
Claims
1. A scene segmentation method based on human-computer collaboration, characterized in that: The following steps are involved: S1: Multiple robots are used to collect observation data of the research scene, and the initial 3D point cloud of the research scene is constructed through an improved multi-robot active mapping algorithm based on neural bipartite graph matching. The following improvements are specifically adopted: The collected observation data is added with a dimension representing height on the basis of the two-dimensional map; Replace the two-dimensional convolutional layer used in the original mGNN with a three-dimensional convolutional layer; Replace the original reinforcement learning observation space from a two-dimensional plane to a three-dimensional space; The original reinforcement learning reward function is replaced by the calculation based on the grid increment of the two-dimensional map and the calculation based on the voxel increment of the three-dimensional map. Add a 3D cuboid bounding box based on the size of the robot and dynamically update the position of the bounding box. Filter out dynamic obstacles by deleting the data of each voxel in the bounding box in the point cloud. S2: Based on the initial 3D point cloud, calculate the point cloud holes, scan the holes through the drone, and complete the point cloud holes based on the scanned data to obtain a complete 3D point cloud; S3: semantic segmentation of the complete 3D point cloud is performed through the semantic segmentation model; S4: Manually annotate the point cloud data with low confidence in the semantic segmentation results; S5: Based on the manual annotation results, after optimizing the semantic segmentation model through active learning, return to S3 and re-segment the complete 3D point cloud through the optimized semantic segmentation model until there is no point cloud data with low confidence, and then output the semantic segmentation result.
2. The scene segmentation method based on human-machine collaboration according to claim 1, characterized in that: When scanning the hole area with a drone, the scanning order is based on the priority of all hole areas; the priority of the hole area is set as follows: First, the priority is set according to the size of the hole area. The larger the hole area, the higher the priority. Secondly, the priority is set according to the confidence level of the point cloud around the hole. The lower the confidence level, the higher the priority of the hole area. Finally, the priority is set according to the path length of the drone to the hole area. The shorter the path length, the higher the priority of the hole area.
3. The scene segmentation method based on human-machine collaboration according to claim 1, characterized in that: The multiple robots used are multi-source heterogeneous robots.
4. A scene segmentation terminal device based on human-machine collaboration, characterized in that: The method comprises a processor, a memory and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of the method according to any one of claims 1 to 3 when executing the computer program.
5. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 3 are implemented.