Starting point annotation correction method based on panoramic image and scene graph matching

By using deep learning and triangle feature matching algorithms, the problems of scene adaptability and computational efficiency in 3D layout and depth prediction of indoor panoramic images are solved, and accurate starting point annotation and efficient annotation correction are achieved in complex scenes.

CN120747967BActive Publication Date: 2025-11-14LANJIAN (SUZHOU) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511261590.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-11-14
Estimated Expiration
2045-09-05

AI Technical Summary

Technical Problem

Existing technologies for 3D layout and depth prediction in indoor panoramic images suffer from problems such as strong assumption dependence, insufficient scene adaptability, fragmented research on layout and depth, limited ability to handle complex scenes, high computational cost, and limitations in data and annotation. These issues result in limited prediction accuracy and efficiency in non-Manhattan structures and complex layout scenes.

Method used

By analyzing panoramic images using a deep learning model and combining connected component analysis and triangle feature matching algorithms, corner point sets and feature point sets of obstacles are generated. Local obstacle alignment and affine transformation are then performed between the scene map and the actual scene to achieve accurate correction of the starting coordinates.

Benefits of technology

It improves the accuracy and efficiency of starting point annotation, is suitable for complex scenarios, reduces manual intervention, adapts to the scene map annotation needs of various indoor scenarios, and is compatible with scene map applications of different precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747967B_ABST
    Figure CN120747967B_ABST
Patent Text Reader

Abstract

This invention provides a starting point annotation correction method based on panoramic image and scene graph matching, relating to the fields of computer vision and building information processing technology. The method includes: generating sets of triangles with the same orientation combination—a scene graph triangle set and an actual scene triangle set—based on the actual scene corner point set and the scene graph feature point set, respectively; matching the triangle pair with the highest similarity between the two sets by calculating the similarity of the triangle's interior angles and side length ratios, establishing a correspondence between local obstacles in the scene graph and the actual scene; mapping the physical starting point position of the actual scene to the scene graph coordinate system based on the affine transformation parameters determined by the triangle pair, and outputting the corrected starting point coordinates. This invention can achieve accurate and efficient restoration of the current layer layout of a construction site.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and building information processing technology, and in particular to a starting point annotation correction method based on matching panoramic images with scene graphs. Background Technology

[0002] Early research relied heavily on hand-designed geometric features and prior assumptions to interpret panoramic images. For example, PanoContext (ECCV 2014), a pioneering work in panoramic scene understanding, extracted geometric cues from panoramic images using techniques such as line segment detection (LSD), vanishing point detection, and alignment, laying the foundation for subsequent layout analysis. These methods typically use the "Manhattan world hypothesis" as a core constraint, assuming that room layouts and objects are aligned along three principal axes, simplifying the layout estimation problem through vanishing point detection and line segment direction analysis. Furthermore, RoomNet was the first to realize the function of recovering room structure from perspective images, while PanoContext further cropped panoramic images into multiple perspective views, achieving layout estimation through multi-view prediction fusion, reflecting the reliance of early methods on perspective transformation and multi-view fusion.

[0003] With the development of deep learning, researchers have begun to leverage the powerful feature learning capabilities of neural networks to solve layout and depth prediction problems. LayoutNet (CVPR 2018) combines deep learning with geometric knowledge, preprocessing to calculate vanishing points and geometric constraints, and then using post-processing to optimize the network output, directly regressing 3D layout parameters and improving the prediction accuracy of Manhattan layouts. DuLa-Net (CVPR 2019) proposes a dual-branch end-to-end learning framework, which directly outputs a probability map of a 2D plan view by fusing surface semantic masks of regular rectangular views with features of projected ground / ceiling views, reducing post-processing dependencies and leveraging the low noise of ceiling views to improve learning efficiency. HorizonNet (CVPR 2019) innovatively uses 1D representation to model room layouts, combining PanoStretch data augmentation technology to quickly generate training samples, and achieves low-computational-cost layout prediction through an efficient post-processing process, making it particularly suitable for complex layouts such as non-cubic or L-shaped layouts. IndoorNet (Advanced Computing Technologies and Applications 2020) uses panoramic images and Manhattan lines as input to build an end-to-end model to predict corners and top and bottom wall lines, marking the gradual replacement of traditional geometric methods by fully convolutional neural networks (FCNN) as the mainstream approach. MyDLNote-360camera (ECCV 2020) further explores the joint modeling of layout and depth, arguing that the two are complementary as key information for 3D scene understanding, and attempts to improve overall prediction performance by fusing layout and depth maps.

[0004] Despite significant progress in indoor panoramic 3D layout and depth prediction, existing methods still have the following limitations:

[0005] (1) Strong assumption dependence and insufficient scene adaptability: Most methods rely on geometric constraints such as the "Manhattan world hypothesis" (e.g., DuLa-Net, LayoutNet), which can easily lead to deviations in non-Manhattan structures (e.g., sloping walls, curved rooms) or complex layout scenes; some methods (e.g., DuLa-Net) are sensitive to mirrors and large object occlusion, which can easily lead to misjudgments.

[0006] (2) Separation of layout and depth research: Previous work has often regarded layout prediction and depth estimation as independent tasks, ignoring their complementarity (such as depth information can reduce the interference of object clutter and occlusion on layout, and layout information can reduce the ambiguity of depth estimation), resulting in limited consistency and accuracy of overall scene understanding.

[0007] (3) Limited ability to handle complex scenes: For complex room layouts such as non-cubic, concave corners, and multi-region splicing, existing models (especially those that have not been specially trained) tend to simplify them into regular shapes (such as the prediction bias of HorizonNet for non-cubic rooms before fine-tuning), resulting in insufficient generalization ability.

[0008] (4) The contradiction between computational cost and model efficiency: Some methods based on dense prediction (such as early models that relied on full-image pixel classification) have an output dimension of O(HW), which consumes a lot of computational resources; although some methods (such as HorizonNet) improve efficiency by reducing the output dimension (O(W)), there is still room for improvement in accuracy in complex scenarios.

[0009] (5) Limitations of data and annotation: In existing datasets (such as PanoContext and Stanford2D-3D), non-cubic rooms are often labeled as cuboids, which leads to a lack of learning samples for complex layouts, thus affecting the prediction performance in real complex scenes. Summary of the Invention

[0010] The technical problem to be solved by the present invention is to provide a starting point annotation correction method based on the matching of panoramic images and scene maps, which can achieve accurate and efficient restoration of the current layer layout of a construction site.

[0011] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0012] Firstly, a starting point annotation correction method based on matching panoramic images with scene graphs, the method comprising:

[0013] Step 1: Use a deep learning model to process the input panoramic image of the actual scene and output the set of corner coordinates of key obstacles in the actual physical space, denoted as the actual scene corner set.

[0014] Step 2: For the initial marked points on the scene map, use the connected component analysis algorithm to extract the obstacle contour feature points in multiple key directions around them to generate a scene map feature point set.

[0015] Step 3: Based on the orientation of the initial annotation point, align the local coordinate system of the scene graph (i.e., with the initial annotation point as the origin) with the local coordinate system of the actual scene (i.e., with the physical space starting point as the origin).

[0016] Step 4: Based on the actual scene corner point set output in Step 1 and the scene map feature point set output in Step 2, generate triangle sets of scenes with the same orientation combination. Figure 3 The set of triangles is compared with the set of triangles in the actual scene. By calculating the similarity of the ratio of the interior angles and side lengths of the triangles, the pair of triangles with the highest similarity in the two sets is matched to establish the correspondence between the scene map and the local obstacles in the actual scene.

[0017] Step 5: Based on the affine transformation parameters determined by the triangle pair matched in Step 3, map the physical starting point position of the actual scene to the scene graph coordinate system and output the corrected starting point coordinates.

[0018] Further, in step 1, the input panoramic image of the actual scene is processed using a deep learning model to output the set of corner coordinates of key obstacles in the actual physical space, denoted as the actual scene corner set, which includes:

[0019] Step 11: Analyze the panoramic image of the actual scene using a pre-trained indoor 3D layout prediction model, and generate 3D layout parameters through the pre-trained indoor 3D layout prediction model to extract the coordinates of wall corners, column vertices, and wall boundary intersections; generate the actual scene corner point set based on the extracted coordinates of wall corners, column vertices, and wall boundary intersections according to the 3D layout parameters; wherein, the 3D layout parameters include the position information of the camera in physical space.

[0020] Further, in step 2, for the initial labeled points on the scene graph, the contour feature points of obstacles in multiple key directions around them are extracted using a connected component analysis algorithm to generate a scene graph feature point set, including:

[0021] Step 21: Using the initial annotation point as the center, establish a local coordinate system for the scene graph according to its annotation orientation, and divide it into four directional regions: front region, back region, left region, and right region.

[0022] Step 22: Based on the connectivity of obstacle pixels in the scene graph, the 4-neighborhood judgment rule is used to identify the set of wall and pillar obstacle pixels that are spatially connected to the initial annotation point.

[0023] Step 23: For the set of obstacle pixels in each directional region, extract its minimum bounding polygon contour, and take the concave points, convex points and extreme points based on curvature changes in the geometric vertices of the contour as feature points.

[0024] Step 24: Aggregate the coordinates of all feature points extracted from the four directional regions into a scene graph feature point set.

[0025] Further, in step 3, based on the orientation of the initial annotation points, the local coordinate system of the scene graph, i.e., with the initial annotation points as the origin, is aligned with the local coordinate system of the actual scene, with the physical space starting point as the origin. This includes:

[0026] Step 31: Using the initial annotation point as the origin, determine the positive direction axis based on its annotation orientation, and establish a local coordinate system for the scene graph;

[0027] Step 32: Using the physical space starting point as the origin, determine the positive direction axis based on the actual spatial orientation corresponding to the panoramic image, and establish a local coordinate system for the actual scene;

[0028] Step 33: Rotate the positive direction axis of the local coordinate system of the scene graph and the positive direction axis of the local coordinate system of the actual scene to the same reference direction to complete the orientation alignment; wherein the reference direction is determined by the approximate orientation of the initial annotation point.

[0029] Furthermore, in step 4, based on the actual scene corner point set output in step 1 and the scene map feature point set output in step 2, triangle sets with the same orientation are generated respectively. Figure 3 The triangle set and the actual triangle set in the scene include:

[0030] Step 41: Divide the feature points in the scene graph feature point set into front point group, rear point group, left point group and right point group according to their respective directional regions; at the same time, group the actual scene corner point set according to the equivalent directional regions in physical space.

[0031] Step 42: For the feature point set of the scene graph, select one feature point from each of the different orientation point groups to form a set of three points, generating all possible triangle combinations to constitute the scene. Figure 3 Triangle set; Perform the same operation on the actual scene corner set to generate the actual scene triangle set; where the three vertices of each triangle must belong to three different orientation point groups.

[0032] Furthermore, by calculating the similarity of the interior angles and side length ratios of triangles, the pair of triangles with the highest similarity in the two sets is matched to establish the correspondence between the scene graph and the local obstacles in the actual scene, including:

[0033] Step 43, for the scene Figure 3 Calculate the pairs of triangles belonging to the same orientation point group in both the triangle set and the actual triangle set:

[0034] Interior angle similarity is the reciprocal of the sum of the absolute values ​​of the differences between the three interior angles.

[0035] Side length ratio similarity, that is, the cosine similarity of the vectors representing the ratio of the lengths of the three sides;

[0036] Step 44: The similarity of the interior angles and the similarity of the side length ratio are weighted and summed according to the preset weights to obtain the comprehensive similarity of the triangle pair.

[0037] Step 45: Traverse all triangle pairs with the same orientation and select the triangle pair with the highest overall similarity as the matching result.

[0038] Further, in step 5, based on the affine transformation parameters determined by the triangle pair matched in step 3, the physical starting point position of the actual scene is mapped to the scene graph coordinate system, and the corrected starting point coordinates are output, including:

[0039] Step 51, centering the scene based on the matched triangles. Figure 3 Find the correspondence between the coordinates of the three vertices of the triangle and the coordinates of the three vertices of the triangle in the actual scene, and calculate the affine transformation matrix from the actual scene coordinate system to the scene graph coordinate system.

[0040] Step 52: Based on the camera position information contained in the 3D layout parameters, determine the coordinates of the physical starting point corresponding to the panoramic image in the actual scene in the actual scene coordinate system.

[0041] Step 53: Map the physical starting point coordinates to the scene graph coordinate system through the affine transformation matrix to obtain the corrected starting point coordinates.

[0042] Secondly, a starting point annotation correction system based on panoramic image and scene graph matching includes:

[0043] The processing module is used to process the input panoramic image of the actual scene using a deep learning model and output the set of corner coordinates of key obstacles in the actual physical space, denoted as the actual scene corner set.

[0044] The analysis module is used to extract the outline feature points of obstacles in multiple key directions around the initial marked points on the scene map through the connected component analysis algorithm, and generate a scene map feature point set.

[0045] The alignment module is used to align the local coordinate system of the scene graph (i.e., with the initial annotation point as the origin) with the local coordinate system of the actual scene (i.e., with the physical space starting point as the origin) based on the orientation of the initial annotation point.

[0046] The module is used to generate a set of triangles with the same orientation based on the actual scene corner point set and the scene graph feature point set. Figure 3 The set of triangles is compared with the set of triangles in the actual scene. By calculating the similarity of the ratio of the interior angles and side lengths of the triangles, the pair of triangles with the highest similarity in the two sets is matched to establish the correspondence between the scene map and the local obstacles in the actual scene.

[0047] The correction module is used to map the physical starting position of the actual scene to the scene graph coordinate system based on the affine transformation parameters determined by the matched triangle pair, and output the corrected starting coordinates.

[0048] The above-described solution of the present invention has at least the following beneficial effects:

[0049] By analyzing the 3D layout of actual panoramic images using deep learning models, this technology binds abstract scene icon annotations to real geometric features of physical space, compared to traditional annotation methods that rely on human experience or simple rules. This improves the accuracy of starting point annotations and avoids annotation errors caused by human estimation bias or scene image simplification.

[0050] A triangle feature matching algorithm is used to match obstacles between the scene graph and the actual scene. Compared with simple point-to-point matching, triangle features have stronger geometric stability and can effectively resist feature deviations caused by missing local obstacles, noise interference, or scene graph simplification. At the same time, by aligning the coordinate system and matching obstacles from multiple directions, the reliability of matching in complex scenes is further ensured.

[0051] From 3D layout analysis of panoramic images and obstacle extraction from scene graphs, to coordinate system alignment, feature matching, and final annotation correction, everything is implemented programmatically, significantly reducing reliance on manual adjustments. It is particularly suitable for large-scale scene annotation or dynamically updated scenes (such as construction sites and renovation sites), significantly improving annotation efficiency and reducing labor costs.

[0052] The technical solution does not rely on specific scene map formats or panoramic image acquisition equipment. As long as the 3D layout of the scene (including obstacle corner points) can be obtained through a deep learning model, it can be applied to scene map annotation correction for various indoor scenes (such as office buildings, residences, shopping malls, construction sites, etc.). At the same time, the triangle matching mechanism has a certain tolerance for the simplification of the scene map (such as abstracting obstacles into lines or polygons), making it adaptable to the application requirements of scene maps with different levels of precision. Accurate starting point annotation is the foundation for downstream applications such as path planning, navigation and positioning, and spatial analysis based on the scene map. Attached Figure Description

[0053] Figure 1 This is a flowchart illustrating the starting point annotation correction method based on matching panoramic images and scene graphs provided in an embodiment of the present invention.

[0054] Figure 2 This is a schematic diagram of the location of the starting point marked on the scene diagram M of the present invention and the approximate orientation of M, wherein the horizontal and vertical axes represent physical distances in meters.

[0055] Figure 3 This is a schematic diagram showing the approximate coordinates of multiple obstacles in the front, back, left, and right directions of the marked point on the scene diagram M of this invention, where the horizontal and vertical coordinates represent pixels.

[0056] Figure 4 This is a schematic diagram of the matched points (yellow) in the scene diagram M(a) of the present invention, where the horizontal and vertical coordinates represent pixels.

[0057] Figure 5 This is a schematic diagram of the points (yellow) matched in the actual scene coordinate system (b) of the present invention, where the horizontal and vertical coordinates represent physical distances in meters.

[0058] Figure 6 This is a schematic diagram of the actual starting point of the present invention on the scene graph, where the horizontal and vertical axes represent pixels. Detailed Implementation

[0059] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0060] like Figure 1 As shown, embodiments of the present invention propose a starting point annotation correction method based on matching panoramic images with scene graphs. The method includes the following steps:

[0061] Step 1: Use a deep learning model to process the input panoramic image of the actual scene and output the set of corner coordinates of key obstacles in the actual physical space, denoted as the actual scene corner set.

[0062] Step 2: For the initial marked points on the scene map, use the connected component analysis algorithm to extract the obstacle contour feature points in multiple key directions around them to generate a scene map feature point set.

[0063] Step 3: Based on the orientation of the initial annotation point, align the local coordinate system of the scene graph (i.e., with the initial annotation point as the origin) with the local coordinate system of the actual scene (i.e., with the physical space starting point as the origin).

[0064] Step 4: Based on the actual scene corner point set output in Step 1 and the scene map feature point set output in Step 2, generate triangle sets of scenes with the same orientation combination. Figure 3 The set of triangles is compared with the set of triangles in the actual scene. By calculating the similarity of the ratio of the interior angles and side lengths of the triangles, the pair of triangles with the highest similarity in the two sets is matched to establish the correspondence between the scene map and the local obstacles in the actual scene.

[0065] Step 5: Based on the affine transformation parameters determined by the triangle pair matched in Step 3, map the physical starting point position of the actual scene to the scene graph coordinate system and output the corrected starting point coordinates.

[0066] In this embodiment of the invention, precise mapping between physical space and scene map coordinates is achieved through fine feature matching (such as interior angles of triangles and side length ratios) of panoramic images and scene maps, combined with affine transformation, effectively correcting initial annotation errors and improving the accuracy of starting point positioning. Matching based on the geometric features (triangle set) of local obstacles reduces the impact of global environmental differences or noise on the matching results, and has better adaptability to scene scale changes, local occlusion, and other situations. Through deep learning feature extraction, automatic algorithm matching, and coordinate transformation, manual intervention is reduced, improving the efficiency and consistency of starting point annotation correction, making it suitable for scenarios requiring rapid and accurate positioning (such as robot navigation, indoor positioning, etc.). Through coordinate system alignment and local feature matching, a reliable association between the scene map and the actual physical space is established.

[0067] In a preferred embodiment of the present invention, step 1 involves processing the input panoramic image of the actual scene using a deep learning model to output a set of corner coordinates of key obstacles in the actual physical space, denoted as the actual scene corner set, including:

[0068] Step 11: Analyze the panoramic image of the actual scene using a pre-trained indoor 3D layout prediction model, and generate 3D layout parameters through the pre-trained indoor 3D layout prediction model to extract the coordinates of wall corners, column vertices, and wall boundary intersections; generate the actual scene corner point set based on the extracted coordinates of wall corners, column vertices, and wall boundary intersections according to the 3D layout parameters; wherein, the 3D layout parameters include the position information of the camera in physical space.

[0069] In this embodiment of the invention, a panoramic image of the actual scene is input and fed into a pre-trained indoor 3D layout prediction model. The model analyzes the panoramic image and, by identifying the visual features of key obstacles such as walls and pillars in the image, and combining the 3D structural rules of the indoor scene that the model has learned, infers the key corner positions of these obstacles in the actual physical space. Finally, the model outputs a set of coordinates of these key corner points, i.e., the actual scene corner point set.

[0070] The implementation process of step 11 above is as follows:

[0071] The pre-trained indoor 3D layout prediction model performs deep analysis on the input panoramic image. By extracting features related to the indoor space structure (such as wall texture, object edges, spatial perspective, etc.) from the image, and combining them with the model's built-in 3D scene reconstruction logic, it generates 3D layout parameters that include the specific location information of the camera in the physical space.

[0072] Based on the generated 3D layout parameters, and using the camera position as a spatial reference, the corners between walls, the vertices of columns, and the intersections of different wall boundaries in the physical space are located and extracted. The coordinates of the extracted wall corners, column vertices, and wall boundary intersections in the physical space are determined, and these coordinates are integrated into a set to generate the actual scene corner point set.

[0073] This invention utilizes a pre-trained 3D layout prediction model, combined with the camera's position information in physical space, to accurately identify and extract the corner points of key obstacles from panoramic images, ensuring that the corner point coordinates match the actual physical space height. The extracted corner points (wall corners, column vertices, etc.) are stable structural features in the scene, providing reliable basic data for subsequent matching with the scene map and reducing interference from non-structural features (such as temporary objects). The generated set of actual scene corner points directly corresponds to key obstacles in physical space, which is consistent with the matching logic of obstacle outline feature points in the scene map, laying a good foundation for establishing a correspondence through geometric features (such as triangles).

[0074] The construction and training process of the pre-trained indoor 3D layout prediction model is as follows:

[0075] Data input layer design:

[0076] The input is defined as a panoramic image of an indoor scene (e.g., a 360° panoramic image). The image resolution is set (e.g., 512×1024 pixels), and the image pixel values ​​are standardized (e.g., normalized to the range of 0-1) as the input data format for the model. A deep convolutional neural network (CNN) is used as the basic feature extraction module, such as a variant structure based on ResNet, VGG, or Transformer. Through multi-layer convolution and pooling operations, local texture features (e.g., wall color, object edges), global structural features (e.g., spatial perspective relationships, wall connection methods), and contextual features (e.g., the relative positions of furniture and walls) in the image are extracted step by step.

[0077] 3D layout parameter prediction layer design:

[0078] After feature extraction, the extracted features are integrated globally and spatially by connecting fully connected layers or recurrent neural networks (RNNs) and attention mechanisms.

[0079] The output layer is designed to predict core parameters of the 3D layout, including:

[0080] Boundary coordinates of walls, ceilings, and floors in physical space (e.g., the position of a corner in a 3D coordinate system); the 6-DOF pose of the camera in physical space (position coordinates and orientation angle); and the dimensional information of the interior space (e.g., room length, width, and height).

[0081] Model structure integration:

[0082] The input layer, feature extraction network, and 3D layout parameter prediction layer are connected in series to form an end-to-end indoor 3D layout prediction model, ensuring a complete mapping from panoramic image input to 3D layout parameter output.

[0083] The model training process is as follows:

[0084] Collect a large-scale labeled dataset containing panoramic images of indoor scenes and their corresponding real 3D layout parameters. The dataset should cover diverse indoor scenes (such as offices, residences, and shopping malls) and be labeled with the following information:

[0085] Pixel information of panoramic images;

[0086] The three-dimensional coordinates of key corner points such as wall corners and column apexes in the physical space; the actual position and orientation of the camera when taking panoramic images; and the actual dimensions and structure of the room.

[0087] Loss function design:

[0088] Design a multi-objective loss function to measure the difference between the model's predicted 3D layout parameters and the ground truth annotations:

[0089] Geometric loss: Calculate the Euclidean distance between the predicted coordinates of wall corners and column vertices and the actual coordinates to ensure spatial positioning accuracy;

[0090] Pose loss: measures the deviation between the predicted camera pose and the true pose by angle difference and position difference;

[0091] Structural consistency loss: constrains the predicted connection relationships between walls, floors, and ceilings (e.g., walls perpendicular to the floor) to ensure that the layout structure conforms to physical laws.

[0092] Training parameter settings:

[0093] Choose an optimizer (such as Adam or SGD), set the initial learning rate to 0.001 (which can be dynamically adjusted, such as with a learning rate decay strategy); set the batch size (e.g., 16 or 32), and adjust it according to the GPU memory; divide the dataset into training set (80%), validation set (10%), and test set (10%) to avoid overfitting.

[0094] Model training and iteration:

[0095] The model is fed with panoramic images from the training set, and it outputs predicted 3D layout parameters. The loss between the predicted parameters and the ground truth annotations is calculated, and the model weights (convolutional kernels, fully connected layer parameters, etc.) are updated through backpropagation. After each training round, the model performance is evaluated using a validation set (e.g., the trend of loss value decrease and parameter prediction accuracy). If the loss on the validation set no longer decreases, training is stopped (early stopping strategy) to avoid overfitting. For scenarios with concentrated errors found during training (e.g., complex multi-room layouts), the training weights of the corresponding samples are increased to improve the model's adaptability to special scenarios.

[0096] Model fine-tuning and optimization:

[0097] After training, the model's generalization ability is evaluated on the test set. If there are problems with large prediction errors in specific scenarios, data for such scenarios are added for fine-tuning. The network structure (such as adding an attention module to strengthen key features) or the weight of the loss function is adjusted to further optimize the prediction accuracy, and finally a pre-trained indoor 3D layout prediction model is obtained.

[0098] This invention enables the model to accurately infer 3D layout parameters from panoramic images through large-scale labeled data training and multi-objective loss function constraints. The training data and optimization strategies covering diverse indoor scenes enable the model to adapt to different types of indoor environments, reducing prediction errors caused by scene differences. The model directly maps images to 3D parameters, avoiding complex manual feature engineering and improving the automation and efficiency of 3D layout parsing.

[0099] In a preferred embodiment of the present invention, step 2, for the initial marked point on the scene map, extracts obstacle contour feature points in multiple key directions around it using a connected component analysis algorithm to generate a scene map feature point set, including:

[0100] Step 21: Using the initial annotation point as the center, establish a local coordinate system for the scene graph according to its annotation orientation, and divide it into four directional regions: front region, back region, left region, and right region.

[0101] Step 22: Based on the connectivity of obstacle pixels in the scene graph, the 4-neighborhood judgment rule is used to identify the set of wall and pillar obstacle pixels that are spatially connected to the initial annotation point.

[0102] Step 23: For the set of obstacle pixels in each directional region, extract its minimum bounding polygon contour, and take the concave points, convex points and extreme points based on curvature changes in the geometric vertices of the contour as feature points.

[0103] Step 24: Aggregate the coordinates of all feature points extracted from the four directional regions into a scene graph feature point set.

[0104] This invention employs a connected component analysis algorithm to perform pixel-level connectivity analysis on obstacles (such as walls and pillars) around the initial annotation point, identifying obstacle regions spatially associated with the initial annotation point. For these obstacle regions, feature points (such as vertices, curvature extrema, etc.) on their contours are extracted at multiple key locations (such as front, back, left, and right) around the initial annotation point. The coordinates of all extracted feature points are integrated into a set, thus generating a scene graph feature point set.

[0105] The implementation process of step 21 above is as follows:

[0106] Establish a local coordinate system for the scene graph (such as a two-dimensional rectangular coordinate system, where the x-axis and y-axis represent the horizontal and vertical directions, respectively) with the initial annotation point on the scene graph as the origin. Determine the orientation reference of the coordinate system based on the annotation orientation of the initial annotation point (such as the direction of movement of the annotation). Based on this reference, divide the space around the initial annotation point into four orientation regions: the front region (directly in front of the annotation orientation), the rear region (in the opposite direction to the front), the left region (to the left of the annotation orientation), and the right region (to the right of the annotation orientation). Each region covers a certain angular range (such as 90° each).

[0107] The implementation process of step 22 above is as follows:

[0108] Traverse the local area centered on the initial marker point in the scene image, identify pixels representing obstacles such as walls and pillars (usually distinguished by pixel values ​​or preset labels, such as obstacle pixels having specific grayscale values); use the 4-neighborhood judgment rule (i.e., determine whether the four adjacent pixels above, below, left, and right of a pixel are pixels of the same obstacle) to perform connectivity analysis on obstacle pixels: if two obstacle pixels are adjacent and have the same attributes (belonging to the same obstacle), they are determined to be connected; aggregate all obstacle pixels connected to the spatial location of the initial marker point to form multiple obstacle pixel sets (each set corresponds to a complete obstacle, such as a section of wall or a pillar).

[0109] The implementation process of step 23 above is as follows:

[0110] For each directional region (front, back, left, right) obtained in step 22, the set of obstacle pixels in that region is selected. Contour extraction is performed on the set of obstacle pixels in each region. The boundary pixels of the obstacle are determined by edge detection. Then, the smallest bounding polygon that can surround these boundary pixels (i.e., the simplest polygon that covers all boundary pixels) is fitted. The geometric features of the polygon contour are analyzed to identify concave points (vertices where the contour is concave inward) and convex points (vertices where the contour is convex outward). By calculating the curvature change of the contour curve, the curvature extrema (points with the maximum or minimum curvature) are found. The concave points, convex points, and curvature extrema are determined as feature points of the obstacle in that directional region.

[0111] The implementation process of step 24 above is as follows:

[0112] Collect all feature points extracted from the four directional regions (front, back, left, and right) in step 23, record the coordinates of each feature point in the local coordinate system of the scene graph, and organize these feature point coordinates (such as deduplication and classification by direction) to form a complete set, that is, generate the scene graph feature point set.

[0113] This invention precisely locates feature regions by dividing the area into four key directional regions, focusing on core obstacles around the initial annotation point, avoiding interference from irrelevant areas, and ensuring that the extracted feature points are more targeted in matching with the corner points of the actual scene. The extracted concave points, convex points, and curvature extrema points are geometrically critical nodes of the obstacle contour, possessing uniqueness and stability. Through connected component analysis and minimum bounding polygon extraction, the contour information of the obstacle is fully covered. Combined with feature aggregation from multiple directional regions, this ensures that the feature point set of the scene map can comprehensively reflect the obstacle structure around the initial annotation point, improving the robustness of matching with the actual scene. The local coordinate system established based on the orientation of the initial annotation point is consistent with the alignment logic of the local coordinate system of the actual scene, reducing coordinate transformation errors.

[0114] In a preferred embodiment of the present invention, step 3, based on the orientation of the initial annotation point, aligns the local coordinate system of the scene graph (i.e., with the initial annotation point as the origin) with the local coordinate system of the actual scene (with the physical space starting point as the origin), including:

[0115] Step 31: Using the initial annotation point as the origin, determine the positive direction axis based on its annotation orientation, and establish a local coordinate system for the scene graph;

[0116] Step 32: Using the physical space starting point as the origin, determine the positive direction axis based on the actual spatial orientation corresponding to the panoramic image, and establish a local coordinate system for the actual scene;

[0117] Step 33: Rotate the positive direction axis of the local coordinate system of the scene graph and the positive direction axis of the local coordinate system of the actual scene to the same reference direction to complete the orientation alignment; wherein the reference direction is determined by the approximate orientation of the initial annotation point.

[0118] This invention clearly defines the orientation references of the local coordinate system of the scene graph (with the initial annotation point as the origin) and the local coordinate system of the actual scene (with the physical space starting point as the origin); with the orientation of the initial annotation point as a reference, the positive direction axes of the two local coordinate systems are adjusted by rotation to make their orientations consistent; after the orientation alignment is completed, it is ensured that the correspondence between the two coordinate systems in spatial orientation (such as front-back, left-right directions) matches, laying the foundation for subsequent feature point matching and coordinate mapping.

[0119] The implementation process of step 31 above is as follows:

[0120] In the scene graph, the initial annotation point is set as the origin (0,0) of the coordinate system. Based on the orientation of the initial annotation point (such as the direction indicated by the arrow or the preset forward direction), the orientation is defined as the positive direction axis of the local coordinate system of the scene graph (such as the positive x-axis). Using the positive direction axis as the reference, another coordinate axis (such as the y-axis) is set perpendicular to the positive direction axis to form a complete local two-dimensional rectangular coordinate system of the scene graph, clarifying the correspondence between each direction in the coordinate system and the spatial orientation in the scene graph (such as the positive direction axis being forward, and the vertical axis being left on one side and right on the other side).

[0121] The implementation process of step 32 above is as follows:

[0122] In the actual physical space, the physical starting point position to be corrected is set as the origin (0,0) of the local coordinate system of the actual scene; based on the actual spatial orientation corresponding to the panoramic image (such as the direction of physical space movement inferred from the perspective relationship of objects in the panoramic image, the shooting direction of the camera, etc.), this orientation is defined as the positive direction axis of the local coordinate system of the actual scene (such as the positive x-axis); with the positive direction axis as the reference, another coordinate axis (such as the y-axis) is set perpendicular to the positive direction axis to form a local two-dimensional rectangular coordinate system of the actual scene, and the correspondence between each direction in the coordinate system and the front-back and left-right positions in the physical space is clarified.

[0123] The implementation process of step 33 above is as follows:

[0124] Determine the approximate orientation of the initial annotation point relative to the reference direction defined by the annotation (e.g., the unified reference direction corresponding to "forward" in the annotation, such as geographic north, right side of the image, etc.); calculate the angle between the positive direction axis of the scene map's local coordinate system and the reference direction, and the angle between the positive direction axis of the actual scene's local coordinate system and the reference direction; rotate and adjust the scene map's local coordinate system or the actual scene's local coordinate system so that the positive direction axes of both coordinate systems coincide with the reference direction, that is, the positive direction axes of both point in the same direction; verify the orientation correspondence between the two coordinate systems after adjustment (e.g., "left" in the scene map corresponds to "left" in the actual scene), and confirm that the orientation alignment is complete.

[0125] This invention solves the problem of orientation misalignment between the scene graph and the actual scene caused by initial orientation labeling errors or spatial mapping differences by aligning the positive direction axes of two local coordinate systems to the same reference direction, ensuring consistent correspondence in directions such as front-back, left-right, etc. After orientation alignment, the feature point set of the scene graph and the corner point set of the actual scene have direct spatial orientation correspondence (feature points in the same orientation region are in the same quadrant of the coordinate system), reducing directional interference during subsequent feature matching and improving matching speed and accuracy. A unified orientation reference provides a spatially consistent basis for subsequent mapping of the actual physical starting point to the scene graph coordinate system through affine transformation, avoiding coordinate transformation errors caused by orientation misalignment and improving the final accuracy of starting point correction. Based on the approximate orientation of the initial labeling points to determine the reference direction, it adapts to orientation labeling habits in different scenarios (such as arrows, text descriptions, etc.), making the coordinate system alignment logic more universal.

[0126] In a preferred embodiment of the present invention, step 4 involves generating a set of triangles with the same orientation based on the actual scene corner point set output in step 1 and the scene map feature point set output in step 2. Figure 3 The triangle set and the actual triangle set in the scene include:

[0127] Step 41: Divide the feature points in the scene graph feature point set into front point group, rear point group, left point group and right point group according to their respective directional regions; at the same time, group the actual scene corner point set according to the equivalent directional regions in physical space.

[0128] Step 42: For the feature point set of the scene graph, select one feature point from each of the different orientation point groups to form a set of three points, generating all possible triangle combinations to constitute the scene. Figure 3 Triangle set; Perform the same operation on the actual scene corner set to generate the actual scene triangle set; where the three vertices of each triangle must belong to three different orientation point groups.

[0129] This invention obtains the actual scene corner point set obtained in step 1 and the scene map feature point set obtained in step 2, respectively; processes the two point sets according to the same orientation combination rule to generate corresponding triangle sets; ensuring the scene Figure 3 The generation logic of the triangle set and the actual scene triangle set is the same. Both are composed of triangles formed by points in different directions, which prepares for subsequent similarity matching.

[0130] The implementation process of step 41 above is as follows:

[0131] For the feature point set of the scene graph, based on the front area, back area, left area and right area divided in step 21, each feature point is assigned to the corresponding directional area, thus dividing it into front point group, back point group, left point group and right point group. Each point group contains all the feature points in that directional area.

[0132] For the actual scene corner point set, based on the actual scene local coordinate system after orientation alignment in step 3, determine the physical space equivalent orientation region corresponding to the front, back, left, and right regions in the scene diagram. Then, group the corner points in the actual scene corner point set according to their respective equivalent orientation regions to obtain the actual scene front point group, back point group, left point group, and right point group corresponding to the point group in the scene diagram.

[0133] The implementation process of step 42 above is as follows:

[0134] Processing the feature point set of the scene graph: From the front, back, left, and right point groups of the scene graph, select one feature point from each of the three different directional point groups to form a triangular set. Iterate through all possible combinations of different directional point groups to generate all triangles that meet the conditions. Combine these triangles to form the scene graph. Figure 3 Corner set; processing real-world scene corner set: using and generating scene Figure 3The same operation method is used for triangle sets. From the four directional point groups in the actual scene, one corner point is selected from each of the three different directional point groups to form a triangle. After traversing all possible combinations, the actual scene triangle set is generated. Throughout the process, the rule that the three vertices of each triangle must belong to three different directional point groups is strictly followed to ensure that the structure of the two triangle sets is consistent.

[0135] This invention generates a set of triangles by grouping areas with the same orientation, thus improving the scene. Figure 3 The triangle set and the actual scene triangle set are structurally and source-corresponding, providing a unified basis for subsequent similarity calculation and ensuring the consistency of matching logic. Triangles are composed of points in different orientations, and their shape and geometric features can reflect the spatial relative relationships of obstacles in different orientations. Compared with a single point or line segment, the features of triangles are more unique, reducing ambiguity in the matching process and improving the accuracy of matching. Generating all possible triangle combinations covers multiple spatial relationships between points in different orientations, making the feature differences and similarity performance between two point sets more comprehensively reflected, which helps to find the most similar triangle pairs. Grouping based on the orientation-aligned regions makes full use of the orientation alignment results in step 3, allowing the generated triangles to accurately reflect the spatial layout relationship of obstacles in two coordinate systems, providing a reliable basis for establishing correspondence through geometric features in the future.

[0136] In a preferred embodiment of the present invention, by calculating the similarity of the interior angles and side length ratios of triangles, matching the triangle pairs with the highest similarity in two sets, a correspondence between the scene graph and the local obstacles in the actual scene is established, including:

[0137] Step 43, for the scene Figure 3 Calculate the pairs of triangles belonging to the same orientation point group in both the triangle set and the actual triangle set:

[0138] Interior angle similarity is the reciprocal of the sum of the absolute values ​​of the differences between the three interior angles.

[0139] Side length ratio similarity, that is, the cosine similarity of the vectors representing the ratio of the lengths of the three sides;

[0140] Step 44: The similarity of the interior angles and the similarity of the side length ratio are weighted and summed according to the preset weights to obtain the comprehensive similarity of the triangle pair.

[0141] Step 45: Traverse all triangle pairs with the same orientation and select the triangle pair with the highest overall similarity as the matching result.

[0142] The implementation process of step 43 of the present invention is as follows:

[0143] Filtering triangle pairs with the same orientation points: From the scene Figure 3From the triangle set and the actual scene triangle set, select triangle pairs whose three vertices belong to the same combination of orientation points (e.g., triangles with the "front + left + right" combination in the scene diagram form candidate pairs with triangles of the same combination in the actual scene); extract the scene from each triangle pair. Figure 3 The values ​​of the three interior angles of the triangle and the actual triangle in the scene; calculate the difference between corresponding interior angles of the two triangles (e.g., in the scene). Figure 3 The interior angle similarity is calculated by taking the absolute value of each difference and summing them, then taking the reciprocal of the sum. (The larger the value, the closer the interior angles of the two triangles are.)

[0144] Calculate the similarity of side length ratios:

[0145] Calculate the lengths of the three sides of the two triangles in the triangle pair (based on the coordinates in their respective coordinate systems); for each triangle, calculate the ratio of the lengths of the three sides (e.g., side 1: side 2, side 2: side 3, side 1: side 3) to form a ratio vector; calculate the cosine similarity of the ratio vectors of the two triangles (the closer the value is to 1, the more similar the side length ratios are), and use this as the side length ratio similarity.

[0146] The implementation process of step 44 above is as follows:

[0147] Preset the weights for interior angle similarity and side length ratio similarity (e.g., interior angle similarity weight 0.6, side length ratio similarity weight 0.4, which can be adjusted according to scene characteristics); multiply the interior angle similarity calculated in step 42 by its corresponding weight, and multiply the side length ratio similarity by its corresponding weight; add the two weighted results to obtain the comprehensive similarity of the triangle pair, which comprehensively reflects the overall similarity of the two triangles in terms of geometric features.

[0148] The implementation process of step 45 above is as follows:

[0149] Traverse all triangle pairs that have been filtered in step 43 and have the same orientation point combination; calculate the comprehensive similarity of each triangle pair one by one (using the method in step 44); compare the comprehensive similarity values ​​of all triangle pairs, select the triangle pair with the largest value, and determine it as the triangle pair with the highest matching degree between the scene map and the actual scene; based on the matching triangle pair, establish the association between the feature points of the scene map corresponding to its three vertices and the corner points of the actual scene, that is, form the correspondence of local obstacles.

[0150] This invention calculates similarity based on dual geometric features—interior angles and side length ratios—combined with a weighted comprehensive evaluation. Compared to single-feature matching, this method more accurately reflects the geometric consistency of triangles, reducing mismatches caused by scale differences or local deformations. Matching is performed only on triangles with the same orientation point group, narrowing the candidate range and avoiding invalid comparisons across orientations, thus improving matching efficiency and accuracy. Since both interior angles and side length ratios are geometrically invariant (unaffected by translation, rotation, or scaling), stable matching is ensured even when there are scale differences or slight deformations between the scene graph and the actual scene, providing a reliable basis for subsequent affine transformation parameter calculations. Through the matching results of the highest similarity triangle pairs, the correspondence between the scene graph and key obstacle feature points in the actual scene is directly established, providing precise spatial mapping anchors for starting coordinate correction.

[0151] In a preferred embodiment of the present invention, step 5, based on the affine transformation parameters determined by the triangle pair matched in step 3, maps the physical starting position of the actual scene to the scene graph coordinate system, and outputs the corrected starting coordinates, including:

[0152] Step 51, centering the scene based on the matched triangles. Figure 3 Find the correspondence between the coordinates of the three vertices of the triangle and the coordinates of the three vertices of the triangle in the actual scene, and calculate the affine transformation matrix from the actual scene coordinate system to the scene graph coordinate system.

[0153] Step 52: Based on the camera position information contained in the 3D layout parameters, determine the coordinates of the physical starting point corresponding to the panoramic image in the actual scene in the actual scene coordinate system.

[0154] Step 53: Map the physical starting point coordinates to the scene graph coordinate system through the affine transformation matrix to obtain the corrected starting point coordinates.

[0155] This invention utilizes the matched triangle pairs in step 3 to determine the affine transformation parameters between the actual scene and the scene graph; combining the coordinates of the physical starting point in the actual scene and the affine transformation parameters, the physical starting point is mapped to the scene graph coordinate system; the mapped coordinates, i.e. the corrected starting point coordinates, are output to complete the coordinate calibration from the actual physical space to the scene graph.

[0156] The implementation process of step 51 above is as follows:

[0157] Extracting the matched triangle pairs, scene Figure 3The coordinates of the three vertices of the triangle in the scene graph coordinate system and the coordinates of the three vertices of the actual scene triangle in the actual scene coordinate system are used to establish a one-to-one correspondence between the two sets of coordinates (e.g., vertex A in the scene graph corresponds to vertex A' in the actual scene, and so on). Based on the coordinate relationship of the three corresponding vertices, an affine transformation matrix that can transform points in the actual scene coordinate system to the scene graph coordinate system is calculated through geometric transformation reasoning. This matrix includes transformation parameters such as scaling, rotation, and translation to ensure that the spatial relationship of corresponding points in the two sets of coordinate systems remains consistent.

[0158] The implementation process of step 52 above is as follows:

[0159] Call the 3D layout parameters generated in step 11 to extract the specific position information (such as three-dimensional coordinates) of the camera in the actual physical space; using the camera position as a reference, and combining the mapping relationship between the panoramic image and the physical space, determine the coordinate value of the physical starting point (i.e. the initial physical position to be corrected) of the panoramic image in the actual scene in the local coordinate system.

[0160] The implementation process of step 53 above is as follows:

[0161] Obtain the affine transformation matrix calculated in step 51, and the coordinates of the physical starting point in the actual scene coordinate system determined in step 52; substitute the coordinates of the physical starting point into the affine transformation matrix, and transform it from the actual scene coordinate system to the scene graph coordinate system through matrix operations to obtain the corresponding coordinates of the physical starting point in the scene graph; these coordinates are the corrected starting point coordinates, and output these coordinates as the final result.

[0162] This invention establishes a coordinate transformation relationship between the actual scene and the scene graph through an affine transformation matrix. Combined with the geometric constraints of matching triangle pairs, it ensures that the mapping error from the physical starting point to the scene graph coordinates is minimized, significantly improving the accuracy of starting point labeling. The physical starting point coordinates are determined using camera position information and then mapped to the scene graph through affine transformation, establishing a coordinate connection between the actual physical space and the scene graph. This ensures that the corrected starting point reflects the true physical location while conforming to the drawing specifications of the scene graph. The affine transformation matrix is ​​calculated based on the correspondence of matching triangle pairs and can simultaneously handle spatial transformations such as scaling, rotation, and translation, ensuring that the relative positional relationship between obstacles in the physical space and the scene graph remains consistent after mapping, providing a reliable coordinate reference for subsequent applications such as path planning. The entire process from transformation matrix calculation to coordinate mapping requires no manual intervention, reducing subjective errors and improving the efficiency and consistency of starting point correction.

[0163] This invention achieves accurate correction of starting point annotations in the scene map by fusing panoramic image information from the actual scene with the geometric features of the scene map. Its core principle is based on matching the geometric features of obstacles in the scene map and the actual scene, correcting deviations in the initial annotations through coordinate system alignment and feature similarity calculation. The specific algorithm consists of four steps: 3D layout analysis of the actual scene and obstacle feature extraction; obstacle feature extraction and orientation analysis in the scene map; obstacle feature matching between the scene map and the actual scene; and starting point annotation correction based on the matching relationship.

[0164] The first step is to process the input panoramic image of the actual scene using publicly available deep learning models for indoor layout analysis (such as 3D layout prediction models based on panoramic images).

[0165] The model's construction and training process is as follows: In model construction, a data input layer is designed, with the input being a panoramic image of an indoor scene. The image resolution is set and pixel values ​​are standardized. A feature extraction network based on a deep convolutional neural network is built to extract local texture, global structure, and contextual features from the image. A 3D layout parameter prediction layer is designed, integrating features with fully connected layers and other modules to output core 3D layout parameters. Finally, the layers are concatenated to form an end-to-end model. During training, a large-scale labeled dataset containing panoramic images and corresponding real 3D layout parameters is prepared. A multi-objective loss function is designed, and training parameters such as optimizer and batch size are set. Through model training and iteration, fine-tuning and optimization, a pre-trained model is obtained.

[0166] When processing panoramic images of real-world scenes using this model, after inputting a panoramic image, the model analyzes the visual features of key obstacles such as walls and pillars in the image. Combining this with learned 3D structural patterns of indoor scenes, it infers the key corner locations of these obstacles in the actual physical space. Specifically, the model performs depth analysis on the panoramic image, extracting features related to the indoor spatial structure and generating 3D layout parameters containing camera position information in the physical space. Based on these parameters, using the camera position as a spatial reference, it locates and extracts wall corners, pillar vertices, and intersections of wall boundaries in the physical space, determines their coordinates, and integrates them into a set of corner points in the actual scene (denoted as the actual scene obstacle feature set B). These corner points directly correspond to key obstacles affecting the spatial structure in the actual scene and are crucial for subsequent matching.

[0167] The second step involves using a connected component analysis algorithm to detect the distribution of obstacles around the initial marked points on the scene map.

[0168] First, using the initial annotation point as the center, establish a local coordinate system for the scene graph based on its annotation orientation, and divide it into four directional regions: the front region, the back region, the left region, and the right region. That is, set the initial annotation point as the origin, define the annotation orientation as the positive direction axis, set another coordinate axis perpendicular to the positive direction axis to form a coordinate system, and then divide it into four directional regions, each approximately 90° apart, based on the positive direction axis.

[0169] Then, based on the connectivity of obstacle pixels in the scene graph, a 4-neighborhood determination rule is used to identify the set of wall and pillar obstacle pixels that are spatially connected to the initial annotation point. The local region centered on the initial annotation point is traversed to identify obstacle pixels. By determining whether the upper, lower, left, and right adjacent pixels of a pixel are the same obstacle pixel, connected obstacle pixels are aggregated to form a set.

[0170] Next, for the obstacle pixel set in each directional region, its minimum bounding polygon contour is extracted, and the concave points, convex points, and extreme points based on curvature changes in the geometric vertices of this contour are used as feature points. Edge detection is performed on the obstacle pixel set of each region to determine the boundary pixels, the minimum bounding polygon is fitted, and then the contour is analyzed to identify the aforementioned feature points.

[0171] Finally, the coordinates of all feature points extracted from the four directional regions are aggregated into a scene graph feature point set (denoted as scene graph obstacle feature set A). This set is formed by collecting feature point coordinates from each direction, removing duplicates, and classifying them. This is to obtain local geometric environment features related to the initial annotation points from the scene graph, providing a reference for subsequent matching with the actual scene.

[0172] The third step is the core of annotation correction. By aligning the coordinate system and calculating the similarity of geometric features, a correspondence is established between the scene graph and the obstacle features in the actual scene. This step specifically includes coordinate system alignment and triangle feature matching algorithms.

[0173] Coordinate system alignment involves initially aligning the local coordinate system A of the scene map (with the initial annotation point as the origin and the initial orientation as the positive direction) with the local coordinate system B of the actual scene (with the physical space origin corresponding to the panoramic image as the origin and the actual orientation as the positive direction) based on the approximate orientation markings of the starting point in the scene map. This ensures that the two coordinate systems maintain consistency in their orientation references, reducing spatial deviations in subsequent matching. Specifically, first, the positive direction axis is determined based on the initial annotation point as the origin and its orientation, establishing the local coordinate system of the scene map; then, the positive direction axis is determined based on the actual spatial orientation corresponding to the panoramic image, establishing the local coordinate system of the actual scene, with the physical space origin as the origin; finally, the positive direction axes of both coordinate systems are rotated to the same reference direction determined by the approximate orientation markings of the initial annotation point, completing the orientation alignment.

[0174] The triangle feature matching algorithm utilizes the geometric stability of triangles (angles and side length ratios are invariant under affine transformations), thus employing triangle features as the matching unit in this invention. For the scene graph obstacle feature set A and the actual scene corner point set B, they are first grouped according to their directional regions. The scene graph feature point set is divided into front, rear, left, and right point groups, while the actual scene corner point set is simultaneously grouped according to their equivalent physical spatial directional regions. Then, one feature point is selected from each of the different directional point groups to form a three-point set, generating all possible triangle combinations, which respectively constitute the scene. Figure 3 Angular set (tri a (set) and actual scene triangle set (triangle) b (a set), where the three vertices of each triangle belong to three different orientation point groups. Then, for tri... a with tri b For triangle pairs that belong to the same orientation combination (such as all consisting of three points: "front + left + right"), calculate the similarity of their interior angles (the reciprocal of the sum of the absolute values ​​of the differences of the three interior angles) and the similarity of their side length ratios (the cosine similarity of the vectors of the ratios of the lengths of the three sides). Then, sum the two values ​​according to a preset weight to obtain the comprehensive similarity. Iterate through all triangle pairs with the same orientation combination and select the triangle pair with the highest comprehensive similarity as the matching result. This result represents the best correspondence between the scene graph and the local obstacle geometry in the actual scene.

[0175] The final step, after determining the triangular feature matching relationship between the scene graph and the actual scene, uses the spatial transformation (affine transformation) parameters implicit in the matching to map the physical starting point position corresponding to the panoramic image in the actual scene (which can be derived from the camera position or corner point relative relationship output by the 3D layout model) to the scene graph coordinate system.

[0176] First, based on the matching triangle centering scenario Figure 3 Determine the correspondence between the coordinates of the three vertices of the triangle and the coordinates of the three vertices of the triangle in the actual scene, and calculate the affine transformation matrix from the actual scene coordinate system to the scene graph coordinate system. Establish the one-to-one correspondence between the two sets of coordinates, and based on this, deduce the matrix containing parameters such as scaling, rotation, and translation through geometric transformations.

[0177] Then, based on the camera position information contained in the 3D layout parameters, the coordinates of the physical starting point corresponding to the panoramic image in the actual scene are determined in the actual scene coordinate system. The position information of the camera in the actual physical space is extracted and used as a reference, combined with the mapping relationship between the panoramic image and the physical space, to determine the coordinates.

[0178] Finally, the physical starting point coordinates are mapped to the scene graph coordinate system through an affine transformation matrix to obtain the corrected starting point coordinates, which are then output. Ultimately, this mapping relationship yields the corrected coordinates of the starting point in the scene graph, achieving accurate correction of the initial annotation.

[0179] In summary, this algorithm effectively integrates the geometric constraints of the virtual scene graph and the real physical space by analyzing the geometric features of the actual scene through deep model analysis, analyzing local obstacles in the scene graph, and cross-domain matching based on triangle features. This enables precise correction of the starting point annotation and provides a reliable coordinate benchmark for the accurate application of the scene graph.

[0180] When applied in specific situations:

[0181] The first step is to input a panoramic image L of the actual scene at the starting point. This panoramic image will serve as the basis for subsequent analysis of the 3D layout of the actual scene and extraction of obstacle features, providing visual information of the actual physical space for the entire starting point annotation and correction process.

[0182] The second step involves processing the panoramic image L using a pre-trained indoor 3D layout prediction model to calculate the camera's position in the actual physical space as (0,0), and simultaneously acquiring the position information of surrounding obstacles, such as... Figure 2 The locations of these obstacles are represented by the coordinates of key corner points, forming a set of actual scene corner points. These corner points include wall corners, pillar vertices, and intersections of wall boundaries, accurately reflecting the distribution of obstacles in the actual physical space. Simultaneously, the scene graph M is marked with the starting point's position within M and M's approximate orientation. This initial marked position is the object of subsequent correction, while the approximate orientation provides an initial reference for coordinate system alignment.

[0183] The third step involves using the initial annotations on scene graph M to detect obstacles and their locations near the annotation points using a connected component analysis algorithm. First, a local coordinate system is established in scene graph M, centered on the initial annotation points and based on their approximate orientation, dividing the graph into four directional regions: front, back, left, and right. Then, based on the connectivity of obstacle pixels in the scene graph, a 4-neighborhood judgment rule is used to identify the pixel sets of obstacles such as walls and pillars spatially connected to the initial annotation points. For each directional region's obstacle pixel set, its minimum bounding polygon contour is extracted, and the concave points, convex points, and extreme points based on curvature changes in the geometric vertices of this contour are used as feature points. Finally, the coordinates of all feature points extracted from the four directional regions are aggregated into a scene graph feature point set, obtaining the approximate coordinates of obstacles in multiple directions (front, back, left, and right) on scene graph M, such as... Figure 3The blue circle represents the approximate starting point marked by the user (with the camera facing downwards), and the red circle represents obstacles around the analyzed starting point. These feature points provide geometric feature references from the scene graph for subsequent matching with the actual scene.

[0184] The fourth step is to match the coordinates of nearby obstacles on scene map M (i.e., the feature point set of the scene map) with the coordinates of nearby obstacles in the actual scene (i.e., the corner point set of the actual scene). First, based on the approximate orientation of the starting point in scene map M, the local coordinate system A of the scene map (with the initial annotation point as the origin and the initial orientation as the positive direction) is aligned with the local coordinate system B of the actual scene (with the physical starting point corresponding to the panoramic image as the origin and the actual orientation as the positive direction). This ensures that the two coordinate systems are consistent in orientation, reducing spatial deviations in subsequent matching. Next, the feature point set of the scene map is divided into front, back, left, and right point groups according to their orientation regions. One feature point is selected from each orientation group to form a three-point set, generating all possible triangle combinations to constitute the scene. Figure 3 Triangle set tri_a; The same operation is performed on the actual scene corner set, i.e., after grouping according to the equivalent orientation region in physical space, one corner point is selected from each orientation group to form a set of three points, generating the actual scene triangle set tri_b, where the three vertices of each triangle belong to three different orientation groups. Then, all triangles belonging to the same orientation combination in tri_a and tri_b are matched. During matching, the similarity of the interior angles (the reciprocal of the sum of the absolute values ​​of the differences of the three interior angles) and the similarity of the side length ratio (the cosine similarity of the vectors of the ratios of the lengths of the three sides) of each triangle pair are calculated. Then, these two similarities are weighted and summed according to preset weights to obtain the comprehensive similarity. Finally, the triangle with the highest comprehensive similarity from the two sets is selected as the matching result. These two triangles represent three points in the scene diagram and the actual scene, respectively. Figure 4 The points (yellow) that are matched in scene graph M(a) Figure 5 Points (yellow) are matched in the actual scene coordinate system (b). Based on this matching result, a correspondence between the scene graph and local obstacles in the actual scene is established.

[0185] Fifth, based on the affine transformation parameters determined by the matching triangle pair obtained in step four, the camera position in the actual scene is directly mapped to the scene graph M through affine transformation. Since the relationship between the camera position and surrounding obstacles in the actual scene was clarified in step two, this relationship and the affine transformation matrix can be used to accurately transform the camera position from the actual scene coordinate system to the scene graph coordinate system, thereby achieving the annotation of the true starting point on the scene graph, such as... Figure 6 The annotation result is the corrected starting point coordinates, accurately reflecting the actual starting point's position in the scene graph.

[0186] A starting point annotation correction system based on panoramic image and scene graph matching includes:

[0187] The processing module is used to process the input panoramic image of the actual scene using a deep learning model and output the set of corner coordinates of key obstacles in the actual physical space, denoted as the actual scene corner set.

[0188] The analysis module is used to extract the outline feature points of obstacles in multiple key directions around the initial marked points on the scene map through the connected component analysis algorithm, and generate a scene map feature point set.

[0189] The alignment module is used to align the local coordinate system of the scene graph (i.e., with the initial annotation point as the origin) with the local coordinate system of the actual scene (i.e., with the physical space starting point as the origin) based on the orientation of the initial annotation point.

[0190] The module is used to generate a set of triangles with the same orientation based on the actual scene corner point set and the scene graph feature point set. Figure 3 The set of triangles is compared with the set of triangles in the actual scene. By calculating the similarity of the ratio of the interior angles and side lengths of the triangles, the pair of triangles with the highest similarity in the two sets is matched to establish the correspondence between the scene map and the local obstacles in the actual scene.

[0191] The correction module is used to map the physical starting position of the actual scene to the scene graph coordinate system based on the affine transformation parameters determined by the matched triangle pair, and output the corrected starting coordinates.

[0192] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A starting point annotation correction method based on panoramic image and scene graph matching, characterized in that, The method includes: Step 1: Process the input panoramic image of the actual scene using a deep learning model, and output the set of corner coordinates of key obstacles in the actual physical space, denoted as the actual scene corner set. This includes: Step 11: Analyze the panoramic image of the actual scene using a pre-trained indoor 3D layout prediction model, and generate 3D layout parameters through the pre-trained indoor 3D layout prediction model to extract the coordinates of wall corners, column vertices, and wall boundary intersections; generate the actual scene corner set based on the extracted coordinates of wall corners, column vertices, and wall boundary intersections according to the 3D layout parameters; wherein, the 3D layout parameters include the position information of the camera in the physical space; Step 2: For the initial annotation point on the scene graph, extract the obstacle contour feature points in multiple key directions around it using the connected component analysis algorithm to generate a scene graph feature point set. This includes: Step 21: Establish a local coordinate system of the scene graph centered on the initial annotation point and based on its annotation orientation, and divide it into four directional regions: front region, back region, left region, and right region; Step 22: Based on the connectivity of obstacle pixels in the scene graph, use the 4-neighborhood judgment rule to identify the set of wall and pillar obstacle pixels that are spatially connected to the initial annotation point; Step 23: For the set of obstacle pixels in each directional region, extract its minimum bounding polygon contour, and use the concave points, convex points, and extreme points based on curvature changes in the geometric vertices of the contour as feature points; Step 24: Aggregate the coordinates of all feature points extracted from the four directional regions into a scene graph feature point set. Step 3: Based on the orientation of the initial annotation point, align the local coordinate system of the scene graph (i.e., with the initial annotation point as the origin) with the local coordinate system of the actual scene (with the physical space starting point as the origin). This includes: Step 31: Using the initial annotation point as the origin, determine the positive direction axis according to its annotation orientation to establish the local coordinate system of the scene graph; Step 32: Using the physical space starting point as the origin, determine the positive direction axis based on the actual space orientation corresponding to the panoramic image to establish the local coordinate system of the actual scene; Step 33: Rotate the positive direction axis of the local coordinate system of the scene graph and the positive direction axis of the local coordinate system of the actual scene to the same reference direction to complete the orientation alignment; wherein the reference direction is determined by the approximate orientation of the initial annotation point. Step 4: Based on the actual scene corner point set output in Step 1 and the scene map feature point set output in Step 2, generate scene map triangle sets and actual scene triangle sets with the same orientation combination, respectively; by calculating the similarity of the interior angles and side length ratios of the triangles, match the triangle pairs with the highest similarity in the two sets to establish the correspondence between the local obstacles in the scene map and the actual scene, including: Step 41: Divide the feature points in the scene map feature point set into front point group, rear point group, left point group, and right point group according to their orientation region; simultaneously group the actual scene corner point set according to the equivalent orientation region in physical space; Step 42: For the scene map feature point set, select one feature point from each of the different orientation point groups to form a three-point set, generating all possible triangles. The shapes are combined to form a triangle set in the scene graph; the same operation is performed on the corner point set of the actual scene to generate a triangle set in the actual scene; the three vertices of each triangle must belong to three different orientation point groups; step 43, for triangle pairs belonging to the same orientation point group combination in the scene graph triangle set and the actual scene triangle set, calculate the following respectively: interior angle similarity, which is the reciprocal of the sum of the absolute values ​​of the differences of the three interior angles; side length ratio similarity, which is the cosine similarity of the vector of the ratio of the lengths of the three sides; step 44, the interior angle similarity and side length ratio similarity are weighted and summed according to preset weights to obtain the comprehensive similarity of the triangle pair; step 45, traverse all triangle pairs with the same orientation combination and select the triangle pair with the highest comprehensive similarity as the matching result; Step 5: Based on the affine transformation parameters determined by the triangle pair matched in Step 3, map the physical starting point position of the actual scene to the scene graph coordinate system and output the corrected starting point coordinates.

2. The starting point annotation correction method based on panoramic image and scene graph matching according to claim 1, characterized in that, Step 5: Based on the affine transformation parameters determined by the triangle pair matched in Step 3, map the physical starting position of the actual scene to the scene graph coordinate system, and output the corrected starting coordinates, including: Step 51: Based on the correspondence between the coordinates of the three vertices of the scene graph triangle and the coordinates of the three vertices of the actual scene triangle, calculate the affine transformation matrix from the actual scene coordinate system to the scene graph coordinate system. Step 52: Based on the camera position information contained in the 3D layout parameters, determine the coordinates of the physical starting point corresponding to the panoramic image in the actual scene in the actual scene coordinate system. Step 53: Map the physical starting point coordinates to the scene graph coordinate system through the affine transformation matrix to obtain the corrected starting point coordinates.

3. A starting point annotation correction system based on panoramic image and scene graph matching, characterized in that, The system is used to perform the method as described in any one of claims 1 to 2, comprising: The processing module is used to process the input panoramic image of the actual scene using a deep learning model and output the set of corner coordinates of key obstacles in the actual physical space, denoted as the actual scene corner set. The analysis module is used to extract the outline feature points of obstacles in multiple key directions around the initial marked points on the scene map through the connected component analysis algorithm, and generate a scene map feature point set. The alignment module is used to align the local coordinate system of the scene graph (i.e., with the initial annotation point as the origin) with the local coordinate system of the actual scene (i.e., with the physical space starting point as the origin) based on the orientation of the initial annotation point. A module is established to generate a set of triangles in the scene map and a set of triangles in the actual scene based on the set of corner points in the actual scene and the set of feature points in the scene map, respectively. By calculating the similarity of the interior angles and side length ratios of the triangles, the highest similarity pair of triangles in the two sets is matched to establish the correspondence between the local obstacles in the scene map and the actual scene. The correction module is used to map the physical starting position of the actual scene to the scene graph coordinate system based on the affine transformation parameters determined by the matched triangle pair, and output the corrected starting coordinates.

Citation Information

Patent Citations

  • Recognition processing method and image processing device using the same

    CN101542520A