Three-dimensional object semantic automatic annotation method based on Gaussian splashing
By using a Gaussian splatter-based automatic semantic labeling method for three-dimensional objects, combined with three-dimensional point cloud data and multimodal feature evaluation, and dynamically adjusting labeling parameters, the imbalance between efficiency and resource allocation of traditional labeling technology in complex scenarios is solved, and accurate labeling and resource optimization are achieved to adapt to complex and changeable smart park and enterprise data processing scenarios.
Patent Information
- Application Number
- CN202510727438.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional labeling technology has difficulty in efficiently integrating geometric structures, material textures, and spatial relationships in complex scenarios where multimodal features are intertwined, resulting in insufficient accuracy and completeness of semantic labeling, and a lack of real-time perception and adaptive adjustment of large-scale dynamic data, leading to an imbalance between labeling efficiency and resource allocation.
A Gaussian splatter-based automatic semantic labeling method for three-dimensional objects is adopted. By collecting three-dimensional point cloud data, the global distribution and neighborhood search range are dynamically adjusted. Combined with a multi-scale feature fusion network and a reinforcement learning framework, the feature quality is monitored in real time, and the labeling results are adjusted through an interactive interface to form an automated closed-loop labeling.
It achieves precise semantic annotation of three-dimensional objects, optimizes resource allocation, improves the accuracy and reliability of annotation, adapts to complex and changing scenarios, and improves the management efficiency and security of smart park and enterprise data processing.
Smart Images

Figure CN120689858A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method for automatic semantic annotation of three-dimensional objects based on Gaussian splattering. Background Art
[0002] Traditional annotation technologies struggle to efficiently integrate multi-dimensional information, including geometric structure, material texture, and spatial relationships, in complex scenes with interwoven multimodal features, resulting in inaccurate and incomplete semantic annotation. In the 3D landscape of a smart campus, traditional methods for annotating complex buildings within the campus (multi-story office buildings with glass curtain walls intersecting with vegetation landscapes) fail to coordinate analysis of geometric curvature and material reflectance. This leads to mislabeling of building facades as transparent materials rather than building structures, and confuses the semantic boundaries between vegetation and surrounding facilities, leading to discrepancies in the spatial analysis of the subsequent campus management system.
[0003] Furthermore, when processing large-scale dynamic data, the lack of real-time perception and adaptive adjustment mechanisms for the dynamic characteristics of data streams leads to an imbalance between labeling efficiency and resource allocation in complex scenarios. In urban traffic, faced with high-density traffic and pedestrians, traditional methods struggle to dynamically optimize labeling resources based on real-time data volume fluctuations. This can lead to labeling delays or simplified labeling (labeling different types of vehicles as a single vehicle) in densely populated areas (e.g., intersections) due to insufficient computing resources, while open areas consume excessive resources for redundant labeling, making it difficult to balance overall labeling efficiency and accuracy. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method for automatic semantic labeling of three-dimensional objects based on Gaussian splashing, which combines three-dimensional probability and multimodal features to evaluate scene complexity, dynamically adjusts Gaussian splashing parameters to propagate semantic labels, and achieves accurate labeling.
[0005] In order to solve the above technical problems, the technical solutions of the present invention are as follows:
[0006] In a first aspect, a method for automatic semantic annotation of three-dimensional objects based on Gaussian splatting is provided, the method comprising:
[0007] Step 1: Collect 3D point cloud data of the park's building surfaces, road facilities, green vegetation, and dynamic objects. Based on the density of the 3D point cloud data, scene complexity, and dynamic object characteristics, the global distribution, number of loops, and neighborhood search range are dynamically adjusted to generate a 3D spatial probabilistic description.
[0008] Step 2: Input the 3D spatial probabilistic description into a multi-scale feature fusion network to extract a multimodal scene feature group, including geometric curvature, material reflectance, and semantic association features. The feature quality is monitored in real time through a traceability diagnosis unit to generate a feature quality assessment report.
[0009] Step 3: Generate scene complexity spatial distribution information based on the feature quality assessment report, and perform Gaussian splattering on the geometrically dense areas identified in the scene complexity spatial distribution information to obtain initial annotation results.
[0010] Step 4: Input the initial annotation results into the reinforcement learning framework and obtain correction instructions through the interactive interface to adjust the initial annotation results to obtain a corrected annotation solution;
[0011] Step 5: Platform-adaptive processing is performed on the improved annotation scheme, including optimizing resource allocation and dynamically adjusting the level of detail based on the viewpoint to obtain the optimized annotation configuration;
[0012] In step 6, the optimized annotation configuration is verified in a 3D semantic consistency space, and the verification results are injected into the knowledge base to update the feature extraction rules and 3D spatial probability description parameters to form an automated closed-loop annotation.
[0013] Furthermore, 3D point cloud data of the park's building surfaces, road facilities, green vegetation, and dynamic objects are collected. Based on the density of the 3D point cloud data, scene complexity, and dynamic object characteristics, the global distribution, number of loops, and neighborhood search range are dynamically adjusted to generate a 3D spatial probabilistic description, including:
[0014] Divide the space where the 3D point cloud data is located into multiple identical areas, count the number of point clouds in each area, and calculate the point cloud density of each area by the ratio of the number of point clouds to the area volume;
[0015] Based on the point cloud density of each area, the complexity of building edges and corners, the density of vegetation, and the distribution density of road facilities are evaluated. A complexity threshold is set to distinguish complex areas with overlapping objects from open areas with simple structures.
[0016] Based on complex areas with multiple objects overlapping and open areas with simple structures, the trajectory of moving vehicles and pedestrians is tracked to determine the speed of movement and the frequency of direction changes, thereby obtaining the characteristics of dynamic objects.
[0017] According to the data density, scene complexity and dynamic object characteristics, the distribution range of geometric structure and dynamic features is adjusted, and the number of geometric structure and dynamic feature extraction cycles is determined. According to the point cloud distribution and object characteristics of complex areas with multiple overlapping objects and open areas with a single structure, the neighborhood search range is adjusted to generate a three-dimensional spatial probabilistic description.
[0018] Furthermore, the space where the 3D point cloud data is located is divided into multiple identical regions, the number of point clouds in each region is counted, and the point cloud density of each region is calculated by the ratio of the number of point clouds to the region volume, including:
[0019] Obtain the number of 3D point clouds and the volume of each region, and determine the initial density of each region;
[0020] According to the initial density, the number of point clouds in the current area and the global average number of point clouds, it is determined whether the initial density value should be amplified and adjusted, and a correction value reflecting the ratio of the number of point clouds to the global average level is obtained;
[0021] Obtain the geometric center coordinates and sensor position coordinates of each area, and obtain the distance attenuation coefficient that reflects the influence of the distance between the area and the sensor based on the preset distance attenuation parameter;
[0022] The initial density value, ratio correction value and distance attenuation coefficient are fused to obtain the final point cloud density of each area.
[0023] Furthermore, the three-dimensional spatial probabilistic description is input into a multi-scale feature fusion network to extract a multimodal scene feature group, including geometric curvature, material reflectance, and semantic association features. The feature quality is monitored in real time through a traceability diagnosis unit to generate a feature quality assessment report, including:
[0024] Based on the position distribution of each point in the three-dimensional spatial probability description, the shape of the building surface and road facility objects is analyzed to obtain the geometric curvature feature information of the object shape structure; based on the reflection of the scanning signal by the point cloud in the three-dimensional spatial probability description, the material of different objects is distinguished and the reflection feature information of the material properties is extracted; based on the position distribution of the point cloud in the three-dimensional spatial probability description and the semantic information of the objects in the park scene, the category of the object to which each point belongs is determined, and the corresponding relationship between the spatial position and the object semantics is established to obtain the semantic association feature information;
[0025] The geometric curvature feature information, material reflection feature information and semantic association feature information are monitored in real time and compared with the features in the historical data of the same scene to determine the feature items with abnormal fluctuations and generate a feature quality assessment report.
[0026] Furthermore, based on the feature quality assessment report, the scene complexity spatial distribution information is generated, and the areas with dense geometric structures identified in the scene complexity spatial distribution information are annotated using the Gaussian splattering algorithm to obtain the initial annotation results, including:
[0027] Combining the probabilistic description of three-dimensional space and the multimodal scene feature group, the scene complexity of the three-dimensional space is evaluated. Based on the characteristic performance of complex areas with multiple overlapping objects and empty areas with simple structures, the location of the scene complexity in the three-dimensional space is distributed to generate the scene complexity spatial distribution information;
[0028] In the spatial distribution information of scene complexity, a complexity threshold is set, and the area with scene complexity greater than the complexity threshold is regarded as a geometric structure dense area. The clustering center of the point cloud in the geometric structure dense area is used as the splash center. The Gaussian kernel standard deviation is adjusted according to the regional point cloud density and geometric characteristics, so that the Gaussian distribution propagates semantic labels to the neighboring points to obtain the initial labeling result.
[0029] Furthermore, the initial annotation results are input into the reinforcement learning framework, and correction instructions are obtained through the interactive interface to adjust the initial annotation results to obtain a corrected annotation solution, including:
[0030] The initial annotation results are fed into a pre-set reinforcement learning processing framework. Based on the correlation between the position information and scene features in the 3D spatial probabilistic description, the initial annotation results are evaluated to identify label mismatches and range deviations.
[0031] Collect correction information for label mismatches and range deviations through an interactive interface, convert the correction information for label mismatches and range deviations into execution adjustment signals, and generate adjustment strategies;
[0032] According to the adjustment strategy, the problem parameters in the initial annotation results are corrected to obtain the corrected annotation scheme.
[0033] Furthermore, the revised annotation scheme is subjected to platform adaptive processing, including optimizing resource allocation and dynamically adjusting the level of detail according to the viewpoint to obtain an optimized annotation configuration, including:
[0034] The revised annotation scheme was evaluated, and the annotation complexity and data volume distribution of each area were analyzed to find the difference in resource requirements during the annotation process between complex areas with multiple overlapping objects and open areas with simple structures.
[0035] Based on the differences in resource requirements during the annotation process between complex areas with multiple overlapping objects and open areas with simple structures, adjust resource allocation to match the annotation requirements of each area and determine the adjusted resource allocation plan;
[0036] Integrate resource allocation schemes with dynamic level-of-detail strategies to obtain optimized annotation configurations.
[0037] Furthermore, the optimized annotation configuration is verified for 3D semantic consistency in space, and the verification results are injected into the knowledge base to update the feature extraction rules and 3D spatial probability description parameters, forming an automated closed-loop annotation process, including:
[0038] Compare and verify the annotation content in the optimized annotation configuration in the three-dimensional space scene to obtain the conflicting annotation areas;
[0039] During the verification process, semantic label errors, annotation range deviations, and spatial relationship contradictions are integrated to obtain a verification result report;
[0040] The verification result report is transmitted to the knowledge base to update the feature extraction rules and three-dimensional space probability description parameters;
[0041] After the feature extraction rules and three-dimensional space probability description parameters are updated, the new rules and parameters are automatically applied to the three-dimensional object semantic labeling process to form an automated closed-loop labeling.
[0042] In a second aspect, a computing device includes:
[0043] one or more processors;
[0044] The storage device is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method.
[0045] According to a third aspect, a computer-readable storage medium stores a program, which implements the method described above when executed by a processor.
[0046] The above solution of the present invention includes at least the following beneficial effects:
[0047] By combining probabilistic descriptions of three-dimensional space with multimodal scene features, the scene complexity of the three-dimensional space is evaluated, and the Gaussian splash algorithm is used to annotate areas with dense geometric structures, enabling the precise identification and annotation of the semantics of three-dimensional objects. This facilitates a deeper understanding of the scene, enabling a more accurate grasp of the situation, whether it is the identification of buildings, road facilities, and green vegetation in smart parks, or the classification of various data objects involved in enterprise data center processing, providing a high-quality data foundation for the data center. Based on the complexity of the scene and annotation requirements, resource allocation is optimized, and computing and storage resources are rationally adjusted to avoid resource waste and improve resource utilization efficiency. In the management of smart parks, resources can be precisely allocated based on the importance and data volume of complex areas with overlapping multiple objects and open areas with simple structures. In the process of enterprise data processing, resources can be rationally allocated for different types of data tasks, thereby improving overall management efficiency and reducing operating costs.
[0048] The annotation results are corrected through a reinforcement learning framework and an interactive interface, and the feature extraction rules and three-dimensional spatial probability description parameters are updated based on the verification results, forming an automated closed-loop annotation. This method can continuously adapt to complex and changing scenarios, continuously improving the accuracy and reliability of annotations. It can quickly adapt to new facility construction or layout adjustments, as well as changes in corporate business data, in smart parks. The accurate annotation results generated by this annotation method can be directly applied to dynamic monitoring of operational data, infrastructure visualization, emergency command visualization, and data monitoring and alarming. It helps to grasp the operational status of the park in real time, promptly identify and address various problems, calmly deal with emergencies, and improve emergency response efficiency, thereby realizing the comprehensive intelligent management of smart parks and improving park safety and service quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a flowchart of a method for automatic semantic annotation of three-dimensional objects based on Gaussian splatting provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0050] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0051] like Figure 1 As shown, an embodiment of the present invention provides a method for automatic semantic annotation of three-dimensional objects based on Gaussian splatting, the method comprising the following steps:
[0052] Step 1: Collect 3D point cloud data of the park's building surfaces, road facilities, green vegetation, and dynamic objects. Based on the density of the 3D point cloud data, scene complexity, and dynamic object characteristics, the global distribution, number of loops, and neighborhood search range are dynamically adjusted to generate a 3D spatial probabilistic description.
[0053] Step 2: Input the 3D spatial probabilistic description into a multi-scale feature fusion network to extract a multimodal scene feature group, including geometric curvature, material reflectance, and semantic association features. The feature quality is monitored in real time through a traceability diagnosis unit to generate a feature quality assessment report.
[0054] Step 3: Generate scene complexity spatial distribution information based on the feature quality assessment report, and perform Gaussian splattering on the geometrically dense areas identified in the scene complexity spatial distribution information to obtain initial annotation results.
[0055] Step 4: Input the initial annotation results into the reinforcement learning framework and obtain correction instructions through the interactive interface to adjust the initial annotation results to obtain a corrected annotation solution;
[0056] Step 5: Platform-adaptive processing is performed on the improved annotation scheme, including optimizing resource allocation and dynamically adjusting the level of detail based on the viewpoint to obtain the optimized annotation configuration;
[0057] In step 6, the optimized annotation configuration is verified in a 3D semantic consistency space, and the verification results are injected into the knowledge base to update the feature extraction rules and 3D spatial probability description parameters to form an automated closed-loop annotation.
[0058] In an embodiment of the present invention, a variety of three-dimensional point cloud data are collected, and processing parameters are dynamically adjusted according to data characteristics to generate a three-dimensional spatial probability description, which can more accurately reflect the distribution probability of different objects in the park in space and improve adaptability to complex and changeable park environments. Comprehensively obtain the geometric, material and semantic feature information of the park scene to provide rich and high-quality data for scene complexity assessment and annotation, while also helping to promptly discover and solve problems that may arise in the feature extraction process and ensure data reliability. Based on the feature quality assessment report, the scene complexity spatial distribution information is generated, and the Gaussian splash algorithm is used to annotate the areas with dense geometric structures to obtain initial annotation results. This method can focus on annotating complex areas in a targeted manner, improve annotation efficiency, and avoid wasting resources in simple areas.
[0059] The reinforcement learning framework can continuously optimize annotation results based on existing data, while the interactive interface can quickly correct annotation deviations, making the annotation scheme more in line with actual needs and improving the practicality and accuracy of annotation. The annotation scheme is platform-adaptive, optimizing resource allocation and dynamically adjusting the level of detail based on the viewpoint. This enables annotation to better adapt to different platform environments and, given limited resources, rationally allocate resources to required areas. At the same time, the level of detail is dynamically adjusted based on changes in viewpoint, providing more refined annotations in areas of interest, improving annotation visualization and user experience. Verification ensures the semantic consistency and accuracy of annotation results, while updates to the knowledge base enable continuous learning and improvement, enhancing adaptability and annotation capabilities for different scenarios, and enabling the self-evolution and continuous optimization of annotation methods.
[0060] In a preferred embodiment of the present invention, step 1, collecting 3D point cloud data of the surfaces of campus buildings, road facilities, green vegetation, and dynamic objects, and dynamically adjusting the global distribution, number of loops, and neighborhood search range based on the density of the 3D point cloud data, scene complexity, and characteristics of the dynamic objects to generate a 3D spatial probabilistic description, may include:
[0061] Step 110 , dividing the space where the three-dimensional point cloud data resides into a plurality of identical regions, counting the number of point clouds in each region, and calculating the point cloud density of each region by the ratio of the number of point clouds to the region volume;
[0062] Step 111 , based on the point cloud density of each area, evaluate the building angular complexity, vegetation density, and road facility distribution density, and set a complexity threshold to separate complex areas with multiple overlapping objects from open areas with a single structure.
[0063] Step 112: Track the trajectories of moving vehicles and pedestrians based on the complex areas with multiple objects overlapping and the open areas with simple structures, determine the speed of movement and the frequency of direction changes, and obtain the characteristics of the dynamic objects;
[0064] Step 113, based on the data density, scene complexity and dynamic object characteristics, adjust the distribution range of the geometric structure and dynamic features, and determine the number of geometric structure and dynamic feature extraction cycles. Based on the point cloud distribution and object characteristics of the complex area with multiple overlapping objects and the open area with a single structure, adjust the neighborhood search range to generate a three-dimensional spatial probabilistic description.
[0065] In an embodiment of the present invention, after obtaining the three-dimensional point cloud data, the spatial range covered by the data is determined. Imagine that this space is a large box, and the spatial range is cut into many small boxes of the same size according to certain rules. These small boxes are the divided areas. Then, check each small box in turn and count the number of point clouds contained in it. After counting the number of point clouds, calculate the volume of the small box and divide the number of point clouds by the volume of the small box. In this way, the point cloud density of each area can be calculated. For example, if there are 100 point clouds in a small box and the volume of the small box is 1 cubic meter, then the point cloud density of this area is 100 point clouds per cubic meter.
[0066] In step 111, for building areas, observe areas with high point cloud density. If you notice significant variation in the point cloud distribution—some areas are densely packed with points, others are sparse, and these variations are concentrated within a relatively small spatial area—this indicates that the buildings in that area have many angular features and a high degree of angular complexity. For vegetation areas, if the point clouds are numerous and disorganized, lacking any regular pattern, you can infer that the vegetation is densely populated. For road infrastructure areas, if the point clouds exhibit a regular linear distribution with minimal density variation, the road infrastructure distribution is relatively uniform. Based on multiple experiments, a complexity threshold is determined. The calculated complexity of each area is compared to this threshold. Areas with a greater complexity than the threshold are complex areas with multiple overlapping objects, such as areas where buildings and vegetation are adjacent. Areas with a lesser complexity than the threshold are open areas with a simple structure, such as large, empty parking lots.
[0067] Step 112: Use point cloud data collected at different times to track moving vehicles and pedestrians within the identified complex and open areas. By comparing their positions at different times, the movement trajectory is determined. The speed of movement is derived from the trajectory. For example, if the object moves 20 meters in 10 seconds, the speed is 2 meters per second. Furthermore, by observing how many times the direction changes during movement and calculating the frequency of these changes, the characteristics of the dynamic object are determined, indicating whether the vehicle or pedestrian is moving fast or slow, and whether the direction changes are frequent.
[0068] Step 113 makes adjustments based on the obtained data density, scene complexity, and dynamic object features. If the data density is high, the scene is complex, and there are many fast-moving, directional-changing dynamic objects, the focus on geometric structures and dynamic features is expanded to provide a more comprehensive understanding of the object's features. If the opposite is true, this range is narrowed. The number of loops for extracting geometric structures and dynamic features is determined based on the scene complexity. The more complex the scene, the more loops are used, allowing for more detailed feature extraction. Simpler scenes require fewer loops, saving time. In areas with dense point clouds and complex object features, the neighborhood search range is narrowed to allow for more accurate observation of local features. In areas with sparse point clouds or simple object features, the neighborhood search range is expanded to prevent missing important information. Finally, the adjusted parameters are combined to generate a three-dimensional spatial probability description, which is used to represent the probability of different objects and features appearing in each area.
[0069] Imagine a large smart campus with several tall buildings, extensive green areas, and a crisscrossing network of roads. The three-dimensional space of the campus is divided into cubic regions with sides of 2 meters. In one of these regions, 300 point clouds are counted. The volume of this cubic region is 8 cubic meters, so the point cloud density in this region is 37.5 points per cubic meter. The area near the tall buildings has a high point cloud density and noticeable variations in distribution, indicating a high degree of building complexity. In the green areas, the point clouds are numerous and chaotic, with dense vegetation. In the area along a straight main road, the point cloud distribution is regular, and the density of road infrastructure is relatively stable. Assuming a complexity threshold of 30 is set, the complexity of the area surrounding the tall buildings and green areas exceeds the threshold, indicating a complex area with multiple overlapping objects. The complexity of some sections of the main road is less than the threshold, indicating an open area with a simple structure.
[0070] At the entrance to the park, there were vehicles and pedestrians moving around. By comparing point cloud data at different times, it was tracked that a vehicle moved 50 meters in 20 seconds, at a speed of 2.5 meters per second, and changed direction once during this process; a pedestrian moved 15 meters in 15 seconds, at a speed of 1 meter per second, and changed direction three times. This shows that vehicles move faster and change direction less frequently, while pedestrians move slower and change direction more frequently. Because the overall data density of the park is high, the scene is complex, and there are many dynamic objects, the distribution range of geometric structures and dynamic features was expanded from the original 3-meter range centered on a certain point to 5 meters; the number of geometric structure and dynamic feature extraction cycles was determined to be 4; in areas with dense point clouds near high-rise buildings, the neighborhood search range was reduced from 1 meter to 0.5 meters, and in open parking lots, the neighborhood search range was expanded from 1 meter to 2 meters. The final three-dimensional spatial probability description shows that the probability of building features appearing in high-rise building areas is high, the probability of vegetation features appearing in green areas is high, and the probability of vehicles, pedestrians, and road facility features appearing in road areas varies depending on the road section.
[0071] By dividing three-dimensional space into regions and calculating point cloud density, complex point cloud data is converted into specific numerical values. This allows for understanding the density of each region's point cloud, enabling a preliminary assessment of its characteristics, including whether it is a concentrated area of objects or an open area. This unified region division approach streamlines data processing, simplifying both analyzing a single region and comparing data from complex areas with multiple overlapping objects with open areas with simple structures, improving data processing efficiency and accuracy. By assessing the complexity associated with different objects and dividing the regions, complex and simple areas within the scene can be clearly identified. Different strategies can be adopted for complex areas with multiple overlapping objects and open areas with simple structures, focusing analysis on complex areas and streamlining processing on simple areas, thereby improving overall processing efficiency. Assessing the complexity of buildings, vegetation, and road infrastructure provides a better understanding of the campus scene. Tracking the trajectories of moving vehicles and pedestrians and analyzing their speed and direction changes accurately captures the behavioral characteristics of dynamic objects. This is highly beneficial for traffic management and security monitoring within the campus, enabling timely detection of abnormal behavior and ensuring safety and order within the campus. The resulting dynamic object characteristics, when used to generate a probabilistic description of 3D space, can more accurately describe the likelihood and behavior of dynamic objects in complex areas with overlapping objects and in open areas with simple structures, enhancing the simulation capabilities of dynamic scenes. The distribution range, number of loops, and neighborhood search range are adjusted based on the data and scene characteristics, enabling adaptive feature extraction for different scenes and objects. By fully exploiting various features in complex scenes and avoiding unnecessary calculations in simple scenes, the accuracy and efficiency of feature extraction are improved, making the generated probabilistic description of 3D space more consistent with real-world scenarios.
[0072] In another preferred embodiment of the present invention, the space where the three-dimensional point cloud data resides is divided into a plurality of identical regions, the number of point clouds in each region is counted, and the point cloud density of each region is calculated by the ratio of the number of point clouds to the volume of the region, which may include:
[0073] Obtain the number of 3D point clouds and the volume of each region, and determine the initial density of each region;
[0074] According to the initial density, the number of point clouds in the current area and the global average number of point clouds, it is determined whether the initial density value should be amplified and adjusted, and a correction value reflecting the ratio of the number of point clouds to the global average level is obtained;
[0075] Obtain the geometric center coordinates and sensor position coordinates of each area, and obtain the distance attenuation coefficient that reflects the influence of the distance between the area and the sensor based on the preset distance attenuation parameter;
[0076] The initial density value, ratio correction value and distance attenuation coefficient are fused to obtain the final point cloud density of each area.
[0077] In the embodiment of the present invention, the three-dimensional space is divided into several regions (uniform grids in step 110), and the number of point clouds n is counted for each region i. i , that is, the total number of point clouds falling into the area. Calculate the volume based on the area geometry (The volume of a cubic region is the cube of its side length, and irregular regions can be calculated using approximate methods.) Give each region a weight w i , reflecting the importance of the region (the weight of the core area of the building is higher than that of the open grassland). Assume that the weight w i The value range is between 0.5-2.0, and the core area of the building is w i =1.8, open grassland takes w i = 0.6. Add up the number of point clouds in all regions and divide it by the total number of regions M to get the average number of point clouds in all regions. Then, calculate the standard deviation The number of regional point clouds n i Relative to the global average number of point clouds The difference is expressed by the standard deviation σ n After normalization, the standard deviation can reflect the degree of dispersion of the point cloud number distribution, that is, whether the number of point clouds is concentrated or dispersed. Calculate the average value of all point cloud coordinates in each area to obtain the geometric center coordinates c of the area i =(x i ,y i , z i ). At the same time, determine the coordinates c of the sensor s =(x s ,ys , z s ), and then set a reference distance R, which can be understood as the effective collection radius of the sensor. For example, in a campus scene, the sensor is installed at a certain location, the reference distance R is set to 50 meters, and the adjustment factors α and β are set. α is used to control the influence intensity of the global distribution difference E, and β controls the speed of distance attenuation. The value range of α is between 0-1. In complex scenes such as urban campuses, in order to highlight the differences in high-density areas, α can be set to 0.7; in simple scenes such as indoors where data distribution is relatively uniform, α is set to 0.3. The value range of β is generally between 0.5-1.5. For scenes where the sensor accuracy drops significantly with distance, β is set to 1.2 to speed up the attenuation rate at long distances; in close-range high-precision scanning scenes, such as scanning in a small room with a handheld scanner, β is set to 0.6 to retain more long-distance details.
[0078] calculate Indicates the number of weighted point clouds within a unit volume, which can reflect the density of point clouds based on the region. i =600, w i =1.5, Then the basic density is Points / m 3 . Calculate first Get the standardized deviation of the number of regional point clouds relative to the global one. If this result is a positive number, it means that the point cloud density in the area is higher than the average level. This positive number is processed by the global distribution difference E (E(x) = x when x>0, otherwise 0), multiplied by α and then added 1 to get the adjustment coefficient. For example, if α=0.6, then the adjustment coefficient is 1.78, which means that the density of the area is enhanced by 78%. Calculate the Euclidean distance between the center of the area and the sensor Then substitute the exponential formula For example, if the distance is 30m, R=60m, and β=1, the attenuation factor is 0.6065, which means that the reliability of the density in the distant area decreases by 39.35%. The final point cloud density is obtained by multiplying the results of the basic density term, the distribution difference adjustment term, and the distance attenuation term.
[0079] By weighting the number of point clouds (n i ×w i ) and regional volume Calculate the basic density, combine the global distribution difference (E term) and the spatial distance attenuation (exponential term), and realize the multi-dimensional characterization of the regional point cloud density. The density weight is enhanced by the E term, and the noise effect is reduced by exponential decay in the distant area, so that ρ i Closer to the actual distribution of data richness and sensor collection reliability. Only for areas with > average point cloud density Positive adjustments are made to effectively highlight key areas with complex geometry and dense objects (building corners, dense vegetation areas), while avoiding interference from low-density open areas. For example, if the point cloud density at a road intersection is greater than the surrounding area, the density weight will be increased after E adjustment. The sensor's distance attenuation mechanism reduces the impact of noise caused by reduced acquisition accuracy in long-distance point clouds (fuzzy data at the edge of the sensor's field of view), making density calculation more dependent on reliable data at close range. For example, the density value of a fuzzy point cloud area far from the sensor will decrease due to exponential decay, avoiding misjudgment of high-value areas. α controls the sensitivity to density differences: in complex scenes with large differences in point cloud distribution (urban campuses), increasing α can enhance the distinction between high-density and low-density areas; in simple scenes with uniform data distribution (a single indoor space), decreasing α can avoid over-amplifying noise differences. In scenes where sensor accuracy decreases significantly with distance, increasing β can accelerate the long-range attenuation rate and focus the effective acquisition range; in close-range, high-precision scanning scenes, decreasing β can preserve more distant details.
[0080] In a preferred embodiment of the present invention, in step 2, the three-dimensional spatial probabilistic description is input into a multi-scale feature fusion network to extract a multimodal scene feature group, including geometric curvature, material reflectance, and semantic association features. The feature quality is monitored in real time by a traceability diagnosis unit to generate a feature quality assessment report, which may include:
[0081] Step 220: Based on the position distribution of each point in the three-dimensional spatial probabilistic description, the shapes of the building surfaces and road facilities are analyzed to obtain geometric curvature feature information of the object's shape structure. Based on the reflection of the scanning signal by the point cloud in the three-dimensional spatial probabilistic description, the materials of different objects are distinguished and reflection feature information of the material properties is extracted. Based on the position distribution of the point cloud in the three-dimensional spatial probabilistic description and the semantic information of the objects in the campus scene, the category of the object to which each point belongs is determined, and a correspondence between the spatial position and the object's semantics is established to obtain semantic association feature information.
[0082] Step 221 , monitor the geometric curvature feature information, material reflection feature information and semantic association feature information in real time, and compare them with the features in the historical data of the same scene to determine the feature items of abnormal fluctuations to generate a feature quality assessment report.
[0083] In an embodiment of the present invention, building surfaces and road infrastructure objects are identified and segmented based on the positional distribution of each point in a three-dimensional spatial probabilistic description. For buildings, the point cloud set belonging to the building is determined, and similar processing is performed for road infrastructure objects. For a curved surface object, a point is selected in the point cloud, and a local surface model is formed by fitting the points in its neighborhood. Typically, a quadratic surface is fitted using the least squares method. Based on this fitted surface, the Gaussian curvature and mean curvature geometric curvature indices are calculated for that point according to the definition of curvature. These indices can reflect the degree of curvature and shape variation of the object's surface. For road infrastructure with more regular shapes, simple curvature information can be obtained by calculating the changes in the tangent and normal directions of the point cloud. Reflection intensity data is collected for each point cloud based on the reflection of the scanning signal from the point cloud in the three-dimensional spatial probabilistic description. Different materials have different reflection characteristics for the scanning signal (the laser beam of the lidar). Metal materials generally have higher reflection intensities, while vegetation materials have relatively lower reflection intensities. Cluster analysis is performed on the reflection intensity data, and the point cloud is divided into different categories based on the differences in reflection intensities, with each category corresponding to a possible material. During the clustering process, it is necessary to pre-set the number of clusters (K value) and determine the appropriate K value. For example, in a campus scene, after multiple attempts, we found that when K = 5, the masonry and metal materials of buildings, the asphalt material of roads, the green vegetation material, and the water material can be better distinguished.
[0084] Based on the point cloud position distribution in the three-dimensional spatial probabilistic description and the semantic information of objects in the campus scene, the point cloud is matched with the object category. For example, in the campus, point clouds close to the road and with a regular linear distribution, combined with semantic information, are likely to belong to road facilities (street lights, railings); while point clouds concentrated in a large area and at a certain height may belong to buildings. Using the position, geometric shape features, and reflectance features of the point cloud as input, the category of the object to which each point belongs is classified. A corresponding relationship between spatial position and object semantics is established to form semantically associated feature information. The coordinates of the point cloud and the object category can be stored in a data structure.
[0085] Step 221: Establish a real-time monitoring system to continuously acquire geometric curvature feature information, material reflection feature information, and semantic association feature information. This feature information is collected and recorded at regular intervals (once per second). Simultaneously, historical stored feature data for the same scene is read from the data store.
[0086] Compare the feature information collected in real time with the features in the stored historical data. For geometric curvature features, compare the current curvature value with the historical average curvature value and the fluctuation range. If the current curvature value is greater than the historical fluctuation range, it is marked as an anomaly. For material reflection features, compare the distribution of reflection intensity and clustering results. If the clustering result changes significantly or the reflection intensity is greater than the historical range, it is considered an anomaly. For semantic association features, check whether the category prediction results of the point cloud are consistent with the historical situation. If there are a large number of non-conformities, such as point clouds that originally belonged to buildings being mistakenly identified as road facilities, record them as an anomaly. Based on the comparison results, determine the feature items of abnormal fluctuations. Summarize the abnormal geometric curvature point cloud areas, the categories of material reflection anomalies, and the point cloud information with semantic association errors to generate a feature quality assessment report. The report should include a detailed description of the abnormal features (the location of the abnormal point cloud, the category of objects involved, the specific manifestation of the anomaly, etc.), the time when the anomaly occurred, and the possible impact of the anomaly on subsequent tasks (scene complexity assessment, object labeling).
[0087] Suppose 3D data collection and feature extraction are being performed in a smart park. Within the park, a circular building is located. A point cloud of the building is determined using a probabilistic 3D spatial description. A point on the building's surface is identified, and the points in its neighborhood are fitted to an approximate sphere using the least squares method. Based on the geometric properties of the sphere, the Gaussian curvature of this point is assumed to be 0.005 (hypothetically) and the mean curvature is assumed to be 0.01 (hypothetically). For a straight road guardrail within the park, the curvature is approximately zero (the curvature of an ideal straight line) based on the rate of change of the point cloud's tangent direction. LiDAR scan data shows that the point cloud in a certain area has high reflection intensity, concentrated within a specific range. Processing the reflection intensity data reveals that the point cloud in this area is classified as a metal material. Further analysis of the reflection spectral characteristics indicates that the metal material is likely aluminum alloy. In a corner of the park, several point clouds are located near the road and distributed linearly. Combining this with the park's semantic information knowledge base, these point clouds are identified as streetlights. The coordinates of these point clouds are associated with the semantic category of streetlights and stored in the database.
[0088] The real-time monitoring system collects feature information every second. At one point, it discovered that the geometric curvature of a certain area of the building suddenly increased to 0.015 (exceeding the historical average fluctuation range). This was due to partial wall modifications in this area, resulting in a shape change. Simultaneously, in another area, the point cloud, previously associated with greenery, experienced a change in material reflectance characteristics, with a significant increase in reflection intensity. Comparison with historical data revealed an anomaly. Inspection revealed that this was due to the installation of new reflective floor tiles in that area. Regarding semantic association, a portion of the point cloud, previously identified as a building, was suddenly misclassified as road infrastructure. Further analysis revealed that nearby construction had interfered with the point cloud data. Based on these anomalies, a feature quality assessment report was generated. The report details the location and timing of the abnormal increase in geometric curvature in a certain area of the building, as well as the potential impact on building structural analysis; the location and manifestation of the abnormal material reflectance area, as well as the potential impact on the material distribution assessment within the campus; the location and type of point clouds with semantic association errors, and the potential problems they may cause in object identification and management within the campus.
[0089] Through in-depth analysis of the probabilistic description of three-dimensional space, feature information is extracted from three aspects: geometric curvature, material reflectance, and semantic association. This multimodal feature extraction approach comprehensively describes objects in the scene. For example, geometric curvature features help accurately identify the shape and structural changes of buildings, material reflectance features help distinguish different object materials, and semantic association features establish a connection between point clouds and object semantics, facilitating semantic understanding and analysis. The acquired feature information provides a deeper understanding of the campus scene. Geometric curvature and material reflectance features can be used to understand the physical characteristics of buildings and road facilities. Semantic association features combine these physical characteristics with the semantic meaning of objects, enabling semantic understanding of the scene and determining whether an area is a parking lot or an ancillary facility of a building, providing a powerful basis for campus planning, management, and maintenance. Real-time monitoring of feature information and comparison with stored historical data can promptly detect abnormal feature fluctuations. This helps ensure the quality of feature data and avoid analysis and annotation errors caused by data anomalies. The generated feature quality assessment report provides a valuable reference for campus decision-making and management, and the abnormal information in the report can help managers take timely measures.
[0090] In a preferred embodiment of the present invention, the above step 3 generates scene complexity spatial distribution information based on the feature quality assessment report, and performs Gaussian splattering on the geometrically dense regions identified in the scene complexity spatial distribution information to obtain initial labeling results, which may include:
[0091] Step 330 : Combining the three-dimensional spatial probabilistic description and the multimodal scene feature set, the scene complexity of the three-dimensional space is evaluated. Based on the characteristic representations of complex areas with multiple objects overlapping and open areas with simple structures, the scene complexity is distributed in the three-dimensional space to generate scene complexity spatial distribution information.
[0092] Step 331: Set a complexity threshold in the scene complexity spatial distribution information, regard the area where the scene complexity is greater than the complexity threshold as a geometric structure dense area, use the cluster center of the point cloud in the geometric structure dense area as the splash center, adjust the Gaussian kernel standard deviation according to the regional point cloud density and geometric characteristics, and make the Gaussian distribution propagate semantic labels to the neighboring points to obtain the initial labeling result.
[0093] In an embodiment of the present invention, the point cloud distribution and density information in the three-dimensional space probability description are fused and analyzed with the geometric curvature, material reflection and semantic association features in the multimodal scene feature group. For each area, the scene complexity is determined. For example, if an area has a high point cloud density, large changes in geometric curvature (indicating that the shape of the object is complex), a variety of material types (judged from the material reflection characteristics), and rich semantic associations (there are interactions between multiple different types of objects), then the scene complexity score of the area will be higher. According to the scene complexity of each area, the scene complexity is mapped to the corresponding position in the three-dimensional space. By generating the scene complexity spatial distribution information, it is possible to intuitively see the complexity distribution of complex areas where multiple objects overlap and empty areas with a single structure in the three-dimensional space.
[0094] Step 331, in the generated scene complexity spatial distribution information, set a complexity threshold. The areas where the scene complexity is greater than this threshold are determined as geometrically dense areas. These areas have complex geometric shapes and a large number of objects. For the determined geometrically dense areas, first find the clustering center of the point cloud in the area. Adjust the Gaussian kernel standard deviation according to the regional point cloud density and geometric features. If the point cloud density is large and the geometric features are complex (there are a large number of irregular shapes), then appropriately reduce the Gaussian kernel standard deviation to make the Gaussian distribution more concentrated, so that the local area can be marked more accurately; if the point cloud density is small and the geometric features are relatively simple, then appropriately increase the Gaussian kernel standard deviation to make the Gaussian distribution wider to cover more point clouds.
[0095] The cluster center of the point cloud within the dense geometric structure area is used as the splash center, and the shape of the Gaussian distribution is determined based on the adjusted standard deviation of the Gaussian kernel. Starting from the splash center, semantic labels are propagated to neighboring points according to the probability of the Gaussian distribution. For example, if the object category to which the cluster center belongs is a building, then according to the Gaussian distribution, there is a certain probability that the neighboring points will also be labeled as buildings. The closer the point is to the cluster center, the higher the probability of being labeled as a building. In this way, the point cloud within the dense geometric structure area is semantically annotated, and the initial annotation results are finally obtained.
[0096] Consider a scene annotation operation within a large industrial park. Within the park, one area includes several large, complex-shaped factory buildings, various pipelines, and equipment made of different materials. The 3D spatial probabilistic description of this area shows a high point cloud density. Within the multimodal scene feature set, the geometric curvature features indicate numerous bends and corners in the factory building walls and roofs, with significant geometric curvature variation. Material reflectance features indicate the presence of various materials, including metal and concrete. Semantic association features reveal diverse semantic information, such as production equipment and transportation corridors, resulting in a high scene complexity. Meanwhile, the open parking lot within the park has a low point cloud density, minimal geometric curvature variation, a single material structure, and simple semantic associations, resulting in low scene complexity. The scene complexity of each area is mapped to 3D space to generate a spatial distribution of scene complexity. A complexity threshold of 60 is set. In the spatial distribution of scene complexity, the factory building area has a scene complexity greater than the threshold, identifying it as a region with dense geometric structure.
[0097] The cluster center of the point cloud within the factory area is used as the splash center, assuming the cluster center is identified as the main structure of the factory. The semantic label of the main structure of the factory is propagated to neighboring points according to the Gaussian distribution determined by the adjusted Gaussian kernel standard deviation. Points closer to the cluster center have a higher probability of being labeled as the main structure of the factory. As the distance increases, the probability gradually decreases, but points within a certain range still have a certain probability of being labeled as the main structure of the factory. After propagation, the initial labeling results for the area are obtained.
[0098] Combining a probabilistic description of 3D space with a multimodal scene feature set for scene complexity assessment, this approach comprehensively considers factors such as point cloud distribution, geometry, material, and semantics, providing a more comprehensive and accurate reflection of scene complexity. By displaying scene complexity in 3D space, the resulting spatial distribution of scene complexity intuitively illustrates the location and extent of complex and simple regions. This helps quickly understand the overall scene context, clarify annotation priorities, and improve annotation efficiency. By setting a complexity threshold to identify areas with dense geometric structure, it is possible to precisely identify areas requiring key annotation, avoiding wasting excessive annotation resources on simpler areas and improving both targeted and efficient annotation. Adjusting the Gaussian kernel standard deviation based on regional point cloud density and geometric features enables the Gaussian splattering algorithm to adapt to both complex areas with overlapping objects and open areas with simpler structures. Using a smaller standard deviation for precise annotation in complex areas and a larger standard deviation for improved efficiency in relatively simple areas achieves a balance between accuracy and efficiency. Taking the cluster center as the splash center and propagating semantic labels according to Gaussian distribution, we can utilize the spatial relationship and density information between point clouds to reasonably assign semantic labels to neighborhood points, making the labeling results more consistent with the object distribution patterns in the actual scene and improving the accuracy and rationality of the initial labeling results.
[0099] In a preferred embodiment of the present invention, the above step 4, inputting the initial annotation results into the reinforcement learning framework and obtaining correction instructions through the interactive interface to adjust the initial annotation results to obtain a corrected annotation solution, may include:
[0100] Step 440: Input the initial annotation results into a preset reinforcement learning processing framework. Based on the relationship between the position information and scene features in the three-dimensional spatial probabilistic description, the initial annotation results are evaluated to identify label mismatches and range deviations.
[0101] Step 441 , collecting correction information for label mismatch and range deviation through an interactive interface, and converting the correction information for label mismatch and range deviation into an execution adjustment signal to generate an adjustment strategy;
[0102] Step 442: According to the adjustment strategy, the problem parameters in the initial annotation result are corrected to obtain a corrected annotation solution.
[0103] In an embodiment of the present invention, a preset reinforcement learning processing framework is constructed to integrate the position information and scene feature associations in the three-dimensional spatial probabilistic description to evaluate and make decisions on the initial annotation results. The initial annotation results are input into the reinforcement learning processing framework, and based on the position information and scene feature associations, the initial annotation results are checked point by point and region by region. For the position information, it is determined whether the position of the annotated object is consistent with the actual position in the three-dimensional spatial probabilistic description; for the scene feature associations, it is analyzed whether the semantic and geometric relationships between the annotated object and the surrounding objects are reasonable. The initial annotation results are evaluated to identify any label mismatches and range deviations.
[0104] Step 441, through an interactive interface, displays the annotated areas with label mismatch and range deviation problems. You can intuitively see the places where the annotation is wrong, and input correction information based on your understanding of the scene. The input correction information is converted into an execution adjustment signal, including the modified target label and the corresponding annotated area information. Based on the converted execution adjustment signal, combined with the rules of the reinforcement learning framework, an adjustment strategy is generated. If the annotation label of an area is modified, the adjustment strategy will also involve the linkage adjustment of the annotations of related objects around the area to ensure the rationality of the association relationship between scene features. The adjustment strategy will ensure that the adjusted annotation results can obtain higher rewards based on the reinforcement learning framework, that is, they are more in line with the actual scene situation.
[0105] Step 442, based on the generated adjustment strategy, correct the problem parameters in the initial annotation results. If the adjustment strategy is to modify the annotation label, the annotation label of the corresponding area will be directly updated; if it is to adjust the annotation range, the boundaries of the annotation area will be redefined. For example, if the adjustment strategy requires the range of a certain annotation area to be expanded, the boundary points of the area in three-dimensional space will be re-determined according to the parameters in the adjustment strategy, thereby achieving the adjustment of the annotation range. After completing the correction of all problem parameters, these corrected annotation information are integrated to form a corrected annotation scheme. This annotation scheme is based on the initial annotation results, and is evaluated and corrected by the reinforcement learning framework, which more accurately reflects the actual situation of the objects in the scene.
[0106] Suppose that in the three-dimensional scene annotation of a smart park, a reinforcement learning framework is built to integrate the three-dimensional spatial probabilistic description information of the park, including the location distribution of buildings, roads, and green objects, as well as the spatial relationships and semantic associations between them. The annotation results are evaluated by continuously learning information. The initial annotation results are input into the reinforcement learning processing framework, and it is found that an area is labeled as a building. However, from the location information in the three-dimensional spatial probabilistic description, it is located in the lake area of the park, and the correlation relationship of the surrounding scene features also shows that the area is more consistent with the characteristics of the lake, such as the reflection characteristics of the water body. Therefore, it is identified as a label mismatch problem. At the same time, it is also found that the range of another road annotation area only covers part of the road and does not include the entire road. This is a range deviation problem.
[0107] Through the interactive interface, we can see the lake area that was incorrectly labeled as a building and the road label area with a range deviation. On the interface, we change the label of the lake area to lake, and manually expand the scope of the road label area to include the complete road. The correction is completed and converted into an execution adjustment signal. For the modification of the lake area, the converted signal contains the target label lake and the corresponding label area information; for the adjustment of the road area, the signal contains the coordinate information of the expanded area boundary. The reinforcement learning framework generates an adjustment strategy based on these execution adjustment signals. For the lake area, the adjustment strategy is to directly update the label label to lake and check the consistency of the surrounding area labeling, such as whether the labeling of the area adjacent to the lake is reasonable; for the road area, the adjustment strategy is to redefine the labeling range according to the boundary coordinates set by the user, and at the same time check the reasonableness of the spatial relationship between the road and surrounding buildings and green objects to ensure that the adjusted labeling conforms to the overall logic of the scene. To correct these two problems, the corrected labeling information is integrated to form a corrected labeling scheme.
[0108] By leveraging a reinforcement learning framework, combined with the rich information from the probabilistic description of three-dimensional space, we can more deeply identify label mismatches and range deviations in the initial annotation results. Compared to simple manual inspection, the reinforcement learning framework can evaluate the annotation results from multiple dimensions and in a more comprehensive manner, improving the accuracy and completeness of problem identification. Accurately identifying annotation issues provides a reliable basis for correction, ensuring the correctness of the correction direction. Precise evaluation of the annotation results avoids inaccurate annotations caused by undetected errors, laying the foundation for ultimately obtaining high-quality annotation solutions. Correction information collected through the interactive interface is converted into execution adjustment signals to generate adjustment strategies, ensuring that adjustments to the annotation results are not isolated but rather comprehensively consider the consistency and rationality of the entire scene. The adjustment strategy not only corrects the current annotation problem but also considers the relationship with other object annotations, avoiding the impact of local adjustments on the overall annotation accuracy. By modifying the problem parameters based on the adjustment strategy, problems in the initial annotation results can be targeted and improved, thereby improving the quality of the annotation solution. The revised labeling scheme more accurately reflects the actual situation of objects in the scene and better meets the needs of practical applications. For example, in the management of smart parks, accurate labeling schemes can provide more reliable data support for park planning and facility management, thereby improving the practicality and application value of the labeling results.
[0109] In a preferred embodiment of the present invention, step 5, performing platform adaptive processing on the revised annotation scheme, including optimizing resource allocation and dynamically adjusting the level of detail according to the viewpoint to obtain an optimized annotation configuration, may include:
[0110] Step 550 , evaluating the revised annotation scheme, analyzing the annotation complexity and data volume distribution of each region, and determining the difference in resource requirements between the annotation process for complex regions with multiple overlapping objects and for open regions with a single structure;
[0111] Step 551 , adjusting resource allocation based on the difference in resource requirements during the labeling process between complex areas with multiple overlapping objects and open areas with a single structure to match the labeling requirements of each area, and determining an adjusted resource allocation plan;
[0112] Step 552 : Integrate the resource allocation scheme with the dynamic level of detail strategy to obtain an optimized annotation configuration.
[0113] In an embodiment of the present invention, the revised annotation scheme is used to divide the entire annotation scene into multiple sub-regions. For each sub-region, the annotation complexity is evaluated based on the object types being annotated, the spatial relationships between objects, and the required annotation accuracy. For example, if a region contains multiple different types of objects with complex overlapping and occlusion relationships between these objects, and high annotation accuracy requirements (such as annotation of fine structures inside buildings), then the annotation complexity of this region is high. Conversely, if the region contains a single type of objects with simple spatial relationships and low annotation accuracy requirements (such as annotation of open spaces), then the annotation complexity is low. The amount of annotated data in each sub-region is counted, including the amount of point cloud data and the number of annotation labels. The data volume is analyzed for complex areas with multiple overlapping objects and open areas with simple structures to determine which areas have larger and smaller data volumes. For example, in a scene containing a large building and a large open space, the building area, due to its complex structure, requires more point cloud data for accurate annotation, resulting in a larger data volume; while the open space area has a simpler structure and a relatively smaller data volume. By comprehensively analyzing the complexity of annotation and the distribution of data volume, we determined the differences in resource requirements between complex areas with multiple overlapping objects and open areas with simple structures. Areas with high annotation complexity and large data volumes require more computing resources (CPU and GPU computing power), storage resources (for storing large amounts of point cloud data and annotation information), and time resources (complex annotation tasks require longer processing time); whereas areas with simpler annotation and smaller data volumes require fewer resources.
[0114] Step 551: Develop a resource allocation strategy based on the differences in resource requirements between complex areas with overlapping objects and open areas with simple structures. If multiple computing nodes are available, higher-performance computing nodes can be allocated to areas with high annotation complexity and large data volumes, while lower-performance computing nodes can be allocated to simpler areas. Regarding storage resources, more storage space should be allocated to areas with large data volumes to ensure complete data storage. Regarding time resource allocation, priority should be given to ensuring sufficient computing time for annotation tasks in complex areas. Existing resources are adjusted according to the established resource allocation strategy. For example, in a distributed computing environment, computing tasks can be reallocated to assign annotation tasks for complex areas to high-performance computing nodes. Storage resource allocation can also be adjusted to ensure rapid access to relevant data by computing nodes. During this adjustment process, resource limits must be considered to ensure rational resource allocation and avoid over-allocation of resources in some areas while under-allocating resources in others. After resource adjustments are completed, detailed records of resource allocation are kept, including information on the computing, storage, and time resources allocated to each area, to form an adjusted resource allocation plan. This plan determines the resources available to each area during the annotation process.
[0115] In step 552, when the viewpoint approaches a certain area, the annotation level of detail for that area is increased, for example, by increasing annotation accuracy and displaying more object details. When the viewpoint moves away from the area, the annotation level of detail is decreased to reduce unnecessary detail display. Different levels of detail can be set, such as low, medium, and high, each corresponding to different annotation accuracy and display content. The adjusted resource allocation scheme is integrated with the dynamic level of detail strategy. Resources are allocated appropriately at different levels of detail. For example, at a high level of detail, to ensure high annotation accuracy and rich detail display, more computing and storage resources are allocated to the corresponding area; at a low level of detail, fewer resources are allocated. Furthermore, resource allocation and level of detail are dynamically adjusted based on real-time changes in the viewpoint. For example, when the viewpoint moves rapidly from one area to another, resource allocation and level of detail are adjusted promptly to ensure a fast response and provide appropriate annotation display. By integrating the resource allocation scheme with the dynamic level of detail strategy, a final optimized annotation configuration is formed.
[0116] Consider a 3D annotation scenario for a large smart campus. The campus comprises multiple functional areas, such as offices, commercial areas, green areas, and parking lots. The office area has numerous buildings with complex internal structures, including various facilities and rooms. The diverse object types and complex spatial relationships require high annotation accuracy, resulting in a high annotation complexity. The commercial area, while also home to numerous buildings, has relatively simple structures and moderate annotation complexity. The green area, primarily composed of vegetation and a single object type, has a lower annotation complexity. The parking lot, primarily composed of parking spaces and vehicles, has simple spatial relationships and also a lower annotation complexity. Statistics show that the office area, due to its complex structure, has a very large point cloud data volume and a high number of annotation labels. The commercial area has the second highest data volume, while the green area and parking lot have relatively small data volumes. Considering the annotation complexity and data volume distribution, the office area has the highest resource requirements, requiring significant computing resources to handle the complex annotation tasks, significant storage resources to store the rich point cloud data and annotation information, and a considerable amount of time to complete the annotation. The commercial area has moderate resource requirements, while the green area and parking lot require fewer resources. In terms of computing resources, there are three computing nodes, of which high-performance computing nodes are allocated to the office area, medium-performance computing nodes are allocated to the commercial area, and low-performance computing nodes are allocated to the green area and parking lot. In terms of storage resources, a larger storage space is allocated to the office area, an appropriate amount of storage space is allocated to the commercial area, and a smaller amount of storage space is allocated to the green area and parking lot. In terms of time resources, priority is given to ensuring sufficient computing time for labeling tasks in the office area, followed by the commercial area, and the green area and parking lot are scheduled based on the remaining time. Labeling tasks for the office area are sent to the high-performance computing nodes for execution, tasks for the commercial area are sent to the medium-performance computing nodes, and tasks for the green area and parking lot are sent to the low-performance computing nodes. At the same time, storage resource allocation is adjusted to ensure that data in the office area can be quickly accessed by the high-performance computing nodes, and data in the commercial area and other areas can also be smoothly accessed by the corresponding computing nodes.
[0117] The resources allocated to each area are recorded. For example, the office area is allocated 80% of the computing resources and 50% of the storage resources of the high-performance computing nodes; the commercial area is allocated all the computing resources and 30% of the storage resources of the medium-performance computing nodes; the green area and parking lot are jointly allocated all the computing resources and 20% of the storage resources of the low-performance computing nodes, forming an adjusted resource allocation plan. Three levels of detail are set. The high level of detail is used for the viewpoint within 50 meters of the area. At this time, the internal structure of the building and detailed information on the type of vegetation are displayed; the medium level of detail is used for the viewpoint within 50-100 meters. It displays the outline of the building and the main facilities and the general distribution of the vegetation; the low level of detail is used for the viewpoint beyond 100 meters. It only displays the location and general shape of the building, the scope of the green area and parking lot. When the viewpoint is near an office area and the distance is less than 50 meters, more computing resources are allocated to the office area to handle high-LOD annotation tasks, while ensuring that storage resources can accommodate the reading and storage of large amounts of detailed data. When the viewpoint moves away from the office area and enters a commercial area with a distance of 50-100 meters, resource allocation is adjusted, reducing resources in the office area and increasing them in the commercial area to meet the high-LOD annotation needs in the commercial area. When the viewpoint moves to a green area or parking lot, resource allocation and LOD are adjusted based on the viewpoint distance. By integrating the resource allocation scheme with the dynamic LOD strategy, an optimized annotation configuration is formed. In complex areas with overlapping objects, open areas with simple structures, and when the viewpoint changes, resources can be allocated reasonably to provide appropriate annotation details.
[0118] By assessing the complexity of annotation and analyzing the data volume distribution, we can understand the differences in resource requirements between complex areas with multiple overlapping objects and open areas with simple structures during the annotation process. This avoids blind resource allocation, improves resource utilization efficiency, and ensures that limited resources are prioritized for areas with the greatest need, thereby improving the quality and efficiency of the overall annotation process. Identifying resource requirement differences provides a scientific basis for resource allocation. Resource allocation strategies based on these differences better align with actual annotation needs, making resource allocation more rational and helping to improve annotation accuracy and completeness. Adjusting resource allocation based on resource requirement differences ensures more efficient resource utilization. Allocating high-performance resources to complex areas and low-performance resources to simple areas avoids resource waste, improves overall operational efficiency, and reduces the time and cost of annotation. Reasonable resource allocation ensures that each area's annotation tasks receive sufficient resources, preventing labeling tasks from being stalled or incomplete due to insufficient resources, and ensuring the smooth progress of the annotation process. Dynamically adjust the level of detail based on the viewpoint. When the viewpoint is close to a certain area, high-level detail annotation is provided to meet the demand for detailed information. When the viewpoint is far away, the level of detail is reduced, improving responsiveness, avoiding lags caused by excessive detail display, and enhancing the user experience. Integrating resource allocation solutions with dynamic level of detail strategies can automatically adjust resource allocation and level of detail based on different annotation needs and viewpoint changes, enhancing adaptability and flexibility, and improving the practicality of the annotation system.
[0119] In a preferred embodiment of the present invention, step 6, performing spatial verification of the 3D semantic consistency of the optimized annotation configuration and injecting the verification results into the knowledge base to update the feature extraction rules and 3D spatial probability description parameters to form an automated closed-loop annotation, may include:
[0120] Step 660 , comparing and verifying the annotation content in the optimized annotation configuration in the three-dimensional space scene to obtain conflicting annotation areas;
[0121] Step 661 , during the verification process, semantic label errors, annotation range deviations, and spatial relationship contradictions are integrated to obtain a verification result report;
[0122] Step 662: Transmit the verification result report to the knowledge base to update the feature extraction rules and three-dimensional space probability description parameters;
[0123] In step 663 , after the feature extraction rules and the three-dimensional space probability description parameters are updated, the new rules and parameters are automatically applied to the three-dimensional object semantic annotation process to form an automated closed-loop annotation.
[0124] In an embodiment of the present invention, a complete three-dimensional space scene is constructed using the annotation information in the optimized annotation configuration and the original three-dimensional space data (point cloud data, terrain data). Each annotation content in the optimized annotation configuration is compared and verified in the constructed three-dimensional space scene. Specifically, check whether the position, shape, and semantic label information of each annotated object are consistent with the actual situation in the three-dimensional space scene. For example, for an object marked as a street lamp, check whether its position in the three-dimensional space is consistent with the regular position of the street lamp beside the road, and whether its shape is consistent with the actual shape of the street lamp. By comparing one by one, find out those annotated areas that are inconsistent with the actual scene.
[0125] Step 661: Identify and classify the problems found during the comparison and verification process. Determine whether it is a semantic label error, that is, the labeled semantics do not match the actual semantics of the object; a label range deviation, such as the labeled object range is too large or too small; or a spatial relationship contradiction, such as the spatial relationship between two labeled objects does not conform to the actual situation (the relative position of the building and the road is incorrect). Summarize and integrate the identified semantic label errors, label range deviations, and spatial relationship contradictions. Record the labeled area, problem type, and specific performance information of each problem to form a verification result report.
[0126] In step 662, the generated verification result report is transferred to a knowledge base for storing and managing various knowledge and information related to the semantic annotation of three-dimensional objects. The data in the verification result report is accurately imported into the knowledge base through a data transmission protocol. Based on the verification result report, the feature extraction rules and three-dimensional spatial probability description parameters in the knowledge base are updated. If semantic label errors are frequently found in a certain area, the feature extraction rules related to that area may need to be adjusted. For issues with annotation range deviation, the three-dimensional spatial probability description parameters may need to be readjusted.
[0127] Step 663, when the feature extraction rules and three-dimensional space probability description parameters are updated, these new rules and parameters will be automatically applied to the three-dimensional object semantic annotation process. In the annotation task, the scene features are extracted according to the new feature extraction rules, and the annotation prediction and decision are made according to the updated three-dimensional space probability description parameters. For example, in the new annotation process, according to the updated feature extraction rules, the texture features of the object are analyzed more carefully, and combined with the adjusted three-dimensional space probability description parameters, the position and range of the object are determined more accurately. By feeding the verification results back to the knowledge base and updating the rules and parameters in the annotation process, an automated closed-loop annotation system is formed. Each verification of the annotation results can provide a more accurate basis for the next annotation. As time goes by and the number of verifications increases, the accuracy and reliability of the annotation system will continue to improve, and it will gradually adapt to the annotation needs of various complex three-dimensional scenes.
[0128] Consider a virtual city block 3D scene annotation project. Based on an optimized annotation configuration, a 3D city block model is constructed. The annotated buildings, streets, streetlights, and trees are modeled according to their annotated locations and shapes. The original terrain data and point cloud data are combined to ensure the scene model is as realistic as possible. During verification of the areas annotated as buildings, it was discovered that the shape of one annotated area differed from the actual building shape. The boundary of this annotated area was greater than the actual building boundary, representing an annotation range deviation. Continuing to examine other annotations, it was discovered that an object labeled as a streetlight was located inside the building, clearly inconsistent with its actual location. This represents a spatial relationship discrepancy. Through this review, multiple discrepant annotation areas were identified. The identified issues were analyzed and determined to be either annotation range deviation or spatial relationship discrepancies. These issues were compiled into a verification report. The report included the following information: "The building annotation at coordinates (100, 200, 5) exceeds its actual boundary by 5 meters; the streetlight at coordinates (150, 250, 3) is located inside the building, resulting in a spatial relationship discrepancy." Through the data interface, the verification result report is transmitted to the project's knowledge base, and the feature extraction rules in the knowledge base are updated.
[0129] For building annotation, we added rules for extracting detailed features of building outline edges to more accurately determine building boundaries. For the three-dimensional spatial probability description parameters, we adjusted the probability distribution parameters of streetlight locations so that the probability of streetlights appearing inside buildings approaches 0. When annotating other areas of the city block, they are automatically annotated according to the updated feature extraction rules and three-dimensional spatial probability description parameters. During the annotation process, we place greater emphasis on analyzing the details of building outlines, and when predicting streetlight locations, we avoid annotating streetlights inside buildings based on the updated parameters. With continuous annotation and verification, the rules and parameters in the knowledge base are continuously optimized, and the annotation system's annotation of urban block scenes becomes increasingly accurate, forming a continuously evolving closed-loop annotation process.
[0130] By comparing and verifying annotation content within a 3D scene, inconsistent areas can be promptly identified, effectively avoiding mislabeling errors, improving the accuracy of annotation results, and ensuring that the annotation content matches the actual scene. Constructing a 3D scene and performing comparison and verification helps to more accurately recreate the actual scene. Semantic label errors, annotation range deviations, and spatial relationship inconsistencies are integrated to comprehensively record various issues encountered during the annotation process. Verification results are presented in a report format, allowing for quick identification of annotation issues, targeted analysis and resolution, and improving the quality of the annotation system. Updating feature extraction rules and 3D spatial probabilistic description parameters based on verification results continuously optimizes the annotation system's parameter settings, making it more adaptable to annotation needs. Verification results are stored in a knowledge base, enabling knowledge accumulation and transfer. Automatically applying updated rules and parameters creates an automated closed-loop annotation system, enabling the system to continuously optimize and improve itself. Over time, the annotation results become increasingly accurate and adaptable.
[0131] An embodiment of the present invention further provides a computing device comprising: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the above-described method. All implementations in the above-described method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0132] The embodiment of the present invention further provides a computer-readable storage medium storing instructions, which, when executed on a computer, causes the computer to execute the above-described method. All implementations in the above-described method embodiment are applicable to this embodiment and can achieve the same technical effects.
[0133] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A method for automatic semantic annotation of three-dimensional objects based on Gaussian splatting, characterized in that: The method comprises: Step 1: Collect 3D point cloud data of the park's building surfaces, road facilities, green vegetation, and dynamic objects. Based on the density of the 3D point cloud data, scene complexity, and dynamic object characteristics, the global distribution, number of loops, and neighborhood search range are dynamically adjusted to generate a 3D spatial probabilistic description. Step 2: Input the 3D spatial probabilistic description into a multi-scale feature fusion network to extract a multimodal scene feature group, including geometric curvature, material reflectance, and semantic association features. The feature quality is monitored in real time through a traceability diagnosis unit to generate a feature quality assessment report. Step 3: Generate scene complexity spatial distribution information based on the feature quality assessment report, and perform Gaussian splattering on the geometrically dense areas identified in the scene complexity spatial distribution information to obtain initial annotation results. Step 4: Input the initial annotation results into the reinforcement learning framework and obtain correction instructions through the interactive interface to adjust the initial annotation results to obtain a corrected annotation solution; Step 5: Platform-adaptive processing is performed on the improved annotation scheme, including optimizing resource allocation and dynamically adjusting the level of detail based on the viewpoint to obtain the optimized annotation configuration; In step 6, the optimized annotation configuration is verified in a 3D semantic consistency space, and the verification results are injected into the knowledge base to update the feature extraction rules and 3D spatial probability description parameters to form an automated closed-loop annotation.
2. The method for automatic semantic annotation of three-dimensional objects based on Gaussian splatting according to claim 1, characterized in that: Collect 3D point cloud data of the park's building surfaces, road facilities, green vegetation, and dynamic objects. Based on the density of the 3D point cloud data, scene complexity, and dynamic object characteristics, dynamically adjust the global distribution, number of loops, and neighborhood search range to generate a 3D spatial probabilistic description, including: Divide the space where the 3D point cloud data is located into multiple identical areas, count the number of point clouds in each area, and calculate the point cloud density of each area by the ratio of the number of point clouds to the area volume; Based on the point cloud density of each area, the complexity of building edges and corners, the density of vegetation, and the distribution density of road facilities are evaluated. A complexity threshold is set to distinguish complex areas with overlapping objects from open areas with simple structures. Based on complex areas with multiple objects overlapping and open areas with simple structures, the trajectory of moving vehicles and pedestrians is tracked to determine the speed of movement and the frequency of direction changes, thereby obtaining the characteristics of dynamic objects. According to the data density, scene complexity and dynamic object characteristics, the distribution range of geometric structure and dynamic features is adjusted, and the number of geometric structure and dynamic feature extraction cycles is determined; according to the point cloud distribution and object characteristics of complex areas with overlapping multiple objects and open areas with a single structure, the neighborhood search range is adjusted to generate a three-dimensional spatial probabilistic description.
3. The method for automatic semantic annotation of three-dimensional objects based on Gaussian splatting according to claim 2, characterized in that: Divide the space where the 3D point cloud data is located into multiple identical areas, count the number of point clouds in each area, and calculate the point cloud density of each area by the ratio of the number of point clouds to the area volume, including: Obtain the number of 3D point clouds and the volume of each region, and determine the initial density of each region; According to the initial density, the number of point clouds in the current area and the global average number of point clouds, it is determined whether the initial density value should be amplified and adjusted, and a correction value reflecting the ratio of the number of point clouds to the global average level is obtained; Obtain the geometric center coordinates and sensor position coordinates of each area, and obtain the distance attenuation coefficient that reflects the influence of the distance between the area and the sensor based on the preset distance attenuation parameter; The initial density value, ratio correction value and distance attenuation coefficient are fused to obtain the final point cloud density of each area.
4. The method for automatic semantic annotation of three-dimensional objects based on Gaussian splatting according to claim 3, characterized in that: The 3D spatial probabilistic description is input into a multi-scale feature fusion network to extract a multimodal scene feature group, including geometric curvature, material reflectance, and semantic association features. The feature quality is monitored in real time through a traceability diagnosis unit to generate a feature quality assessment report, including: Based on the position distribution of each point in the three-dimensional spatial probability description, the shape of the building surface and road facility objects is analyzed to obtain the geometric curvature feature information of the object shape structure; based on the reflection of the scanning signal by the point cloud in the three-dimensional spatial probability description, the material of different objects is distinguished and the reflection feature information of the material properties is extracted; based on the position distribution of the point cloud in the three-dimensional spatial probability description and the semantic information of the objects in the park scene, the category of the object to which each point belongs is determined, and the corresponding relationship between the spatial position and the object semantics is established to obtain the semantic association feature information; The geometric curvature feature information, material reflection feature information and semantic association feature information are monitored in real time and compared with the features in the historical data of the same scene to determine the feature items with abnormal fluctuations and generate a feature quality assessment report.
5. The method for automatic semantic annotation of three-dimensional objects based on Gaussian splatting according to claim 4, characterized in that: Based on the feature quality assessment report, the scene complexity spatial distribution information is generated. The areas with dense geometric structures identified in the scene complexity spatial distribution information are annotated using the Gaussian splashing algorithm to obtain the initial annotation results, including: Combining the probabilistic description of three-dimensional space and the multimodal scene feature group, the scene complexity of the three-dimensional space is evaluated. Based on the characteristic performance of complex areas with multiple overlapping objects and empty areas with simple structures, the location of the scene complexity in the three-dimensional space is distributed to generate the scene complexity spatial distribution information; In the spatial distribution information of scene complexity, a complexity threshold is set, and the area with scene complexity greater than the complexity threshold is regarded as a geometric structure dense area. The clustering center of the point cloud in the geometric structure dense area is used as the splash center. The Gaussian kernel standard deviation is adjusted according to the regional point cloud density and geometric characteristics, so that the Gaussian distribution propagates semantic labels to the neighboring points to obtain the initial labeling result.
6. The method for automatic semantic annotation of three-dimensional objects based on Gaussian splatting according to claim 5, characterized in that: Input the initial annotation results into the reinforcement learning framework and obtain correction instructions through the interactive interface to adjust the initial annotation results to obtain a revised annotation solution, including: The initial annotation results are fed into a pre-set reinforcement learning processing framework. Based on the correlation between the position information and scene features in the 3D spatial probabilistic description, the initial annotation results are evaluated to identify label mismatches and range deviations. Collect correction information for label mismatches and range deviations through an interactive interface, convert the correction information for label mismatches and range deviations into execution adjustment signals, and generate adjustment strategies; According to the adjustment strategy, the problem parameters in the initial annotation results are corrected to obtain the corrected annotation scheme.
7. The method for automatic semantic annotation of three-dimensional objects based on Gaussian splatting according to claim 6, characterized in that: The revised annotation scheme is subjected to platform adaptation, including optimizing resource allocation and dynamically adjusting the level of detail based on the viewpoint to obtain an optimized annotation configuration, including: The revised annotation scheme was evaluated, and the annotation complexity and data volume distribution of each area were analyzed to find the difference in resource requirements during the annotation process between complex areas with multiple overlapping objects and open areas with simple structures. Based on the differences in resource requirements during the annotation process between complex areas with multiple overlapping objects and open areas with simple structures, adjust resource allocation to match the annotation requirements of each area and determine the adjusted resource allocation plan; Integrate resource allocation schemes with dynamic level-of-detail strategies to obtain optimized annotation configurations.
8. The method for automatic semantic annotation of three-dimensional objects based on Gaussian splatting according to claim 7, characterized in that: The optimized annotation configuration is verified for 3D semantic consistency in space, and the verification results are injected into the knowledge base to update the feature extraction rules and 3D spatial probability description parameters, forming an automated closed-loop annotation process, including: Compare and verify the annotation content in the optimized annotation configuration in the three-dimensional space scene to obtain the conflicting annotation areas; During the verification process, semantic label errors, annotation range deviations, and spatial relationship contradictions are integrated to obtain a verification result report; The verification result report is transmitted to the knowledge base to update the feature extraction rules and three-dimensional space probability description parameters; After the feature extraction rules and three-dimensional space probability description parameters are updated, the new rules and parameters are automatically applied to the three-dimensional object semantic labeling process to form an automated closed-loop labeling.
9. A computing device, characterized in that include: one or more processors; A storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program, which implements the method according to any one of claims 1 to 8 when executed by a processor.
Citation Information
Cited By
Low-computing-power rapid three-dimensional modeling method and system based on 3DGS
CN122023680A