Furniture real-time space positioning and scene generation method and system fusing ar

CN122574313BActive Publication Date: 2026-09-22SHANGHAI JIANGFENG FURNITURE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611000552.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-22
Estimated Expiration
2046-07-07

AI Technical Summary

Technical Problem

[0002]随着移动终端摄像头、深度感知和增强现实技术的发展,利用AR方式在室内场景中预览家具摆放效果已得到广泛应用;现有方案通常根据相机位姿、平面检测结果或深度信息确定地面、桌面等放置面,再将虚拟家具模型叠加显示于真实图像中;该类方案在空旷场景或放置面可被充分观测时能够取得一定效果,但在实际家居更换场景中,旧沙发、床、柜体、餐桌等家具往往仍处于原位置,其底部地面、墙脚区域及部分墙面会被遮挡,导致系统难以直接获得完整的可支撑区域

Benefits of technology

[0025]本申请通过在房间坐标系中建立支撑证据图,将地面或桌面的直接观测结果、由旧家具接触边推断得到的间接支撑结果、墙体及固定障碍物对应的排除结果统一记录,并结合旧家具实例分割、深度信息和平面信息恢复被遮挡的底面候选区域,可在旧家具未搬离的情况下判断新家具的可放置位置;同时,通过新家具底面接触区域与支撑证据图的匹配,按覆盖面积和证据类型计算支撑评分,并以排除证据筛除冲突位置,能够降低遮挡、深度缺失和墙体干扰造成的误判,减少虚拟家具偏移、悬空或穿插现象,使连续底座家具和支脚类家具均能获得较稳定的空间位置和渲染姿态,从而提高AR家具替换预览的准确性、实时性和使用便利性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122574313B_ABST
    Figure CN122574313B_ABST
Patent Text Reader

Abstract

The application discloses a kind of furniture real-time space positioning and scene generation method and system fused AR, which comprises: obtaining the RGB image of indoor scene, camera pose, plane information and depth information, old furniture instance segmentation is carried out to RGB image, and old furniture class and instance mask are obtained;Support evidence graph is established in room coordinate system, and direct support evidence, indirect support evidence, exclusion evidence and unknown state are recorded;According to old furniture instance mask, depth information and plane information, old furniture contact edge is extracted, and old furniture bottom surface candidate area is generated in combination with old furniture class and furniture size data, and conflict area is screened out by exclusion evidence;Bottom surface contact area and candidate position are generated based on new furniture model data, support score is determined according to area weight and each type of evidence, and then new furniture space position and rendering posture are determined.The application can improve the accuracy of AR furniture positioning and scene generation under the shielding scene of old furniture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of augmented reality and image processing technology, and more specifically, to a method and system for real-time spatial positioning and scene generation of furniture that integrates AR. Background Technology

[0002] With the development of mobile terminal cameras, depth perception, and augmented reality technologies, the use of AR to preview furniture placement effects in indoor scenes has been widely applied. Existing solutions typically determine the placement surfaces such as the ground and tabletops based on camera pose, plane detection results, or depth information, and then overlay virtual furniture models onto real images. This type of solution can achieve certain results in open scenes or when the placement surfaces can be fully observed. However, in actual home replacement scenarios, old sofas, beds, cabinets, dining tables, and other furniture are often still in their original positions, and their bottom surfaces, base areas, and parts of the walls may be obscured, making it difficult for the system to directly obtain the complete support area.

[0003] In the above situations, relying solely on visible ground or single-frame plane detection results can easily lead to the identification of real placeable areas obscured by old furniture as unknown areas. It may also misidentify areas near walls, door frames, or fixed obstacles as placeable areas, resulting in problems such as virtual furniture position shift, conflict with walls or obstacles, unsupported bottom surfaces, and unstable posture. Some existing methods require users to move old furniture, manually calibrate the placement range, or repeatedly adjust the model position, which is costly and difficult to meet the needs of real-time preview and natural interaction. Other methods can use image segmentation or depth estimation to identify furniture outlines, but they usually do not uniformly associate old furniture categories, instance masks, contact edges, depth information, and planar information, and lack a stable judgment mechanism for occluded bottom areas and excluded areas. Summary of the Invention

[0004] This application provides a method and system for real-time spatial positioning and scene generation of furniture that integrates AR, so as to at least solve some of the technical problems existing in the related technologies described above.

[0005] According to a first aspect of the embodiments of this application, a method for real-time spatial positioning and scene generation of furniture integrating AR is provided, including:

[0006] Acquire RGB images, camera pose, planar information, and depth information of the indoor scene; segment the RGB images into old furniture instances to obtain old furniture categories and old furniture instance masks.

[0007] Based on the camera pose, the planar information, and the depth information, a supporting evidence map is established in the room coordinate system. The grid cells of the supporting evidence map record direct supporting evidence, indirect supporting evidence, excluded evidence, and unknown states.

[0008] Based on the old furniture instance mask, the depth information, and the planar information, the contact edges of the old furniture are extracted. The old furniture category and furniture size data are combined to generate candidate areas for the bottom surface of the old furniture. The candidate areas for the bottom surface of the old furniture are written into the indirect supporting evidence of the supporting evidence diagram, and conflict areas are screened out with the exclusion evidence.

[0009] Based on the new furniture model data, a new furniture bottom surface contact area is generated, and a new furniture candidate position is generated within the old furniture bottom surface candidate area after screening.

[0010] The contact area of ​​the bottom surface of the new furniture is projected onto the support evidence map to obtain a covering grid cell. The support score is determined based on the area weight of the covering grid cell, direct support evidence, indirect support evidence, and exclusion evidence. The spatial position and rendering posture of the new furniture are determined based on the support score and the exclusion evidence.

[0011] As an optional approach, old furniture instance segmentation of the RGB image includes: inputting the RGB image into an instance segmentation model, the instance segmentation model comprising a backbone network, a feature fusion network, a detection branch, a mask prototype branch, and a mask coefficient branch connected in sequence; the detection branch outputs the old furniture category and the instance bounding region, the mask prototype branch outputs the full-image mask prototype, and the mask coefficient branch outputs instance mask coefficients; combining the full-image mask prototype according to the instance mask coefficients, and cropping the combination result according to the instance bounding region to obtain the old furniture instance mask.

[0012] As an optional approach, establishing the supporting evidence map includes: transforming the observation points representing the ground plane or tabletop plane in the planar information to the room coordinate system according to the camera pose and assigning them to the corresponding grid cells to form the direct supporting evidence; assigning the projection areas representing the wall plane or fixed obstacles in the planar information to the corresponding grid cells to form the exclusion evidence; and recording the grid cells that do not form the direct supporting evidence and the exclusion evidence as the unknown state.

[0013] As an optional approach, extracting the old furniture contact edges includes: extracting a mask contour from the old furniture instance mask, and obtaining contour sampling points along the mask contour; calculating the three-dimensional position of the contour sampling points in the room coordinate system based on the depth information and the camera pose; grouping contour sampling points whose height satisfies the height difference condition with the ground plane and are continuously distributed into ground contact edges; and grouping contour sampling points whose distance satisfies the distance condition with the wall plane and are directionally continuous into adjacent wall edges; the old furniture contact edges include the ground contact edges and the adjacent wall edges.

[0014] As an optional approach, generating the candidate area for the bottom surface of the old furniture includes: selecting a bottom surface representation method according to the category of the old furniture; when the bottom surface representation method is a continuous bottom surface, taking the ground contact edge as the first boundary and the adjacent edge of the wall or the depth direction range defined by the furniture size data as the second boundary to generate the candidate area for the bottom surface of the old furniture; when the bottom surface representation method is a leg bottom surface, generating the candidate area for the bottom surface of the old furniture according to the leg contact position corresponding to the ground contact edge.

[0015] As an optional approach, writing the candidate area of ​​the old furniture bottom surface into the indirect supporting evidence of the supporting evidence diagram, and using the exclusion evidence to screen out conflict areas, includes: determining the grid cells covered by the candidate area of ​​the old furniture bottom surface as grid cells to be written; reading the exclusion evidence of each grid cell to be written; writing the indirect supporting evidence into the grid cells to be written that do not meet the exclusion criteria; and removing the grid cells to be written that meet the exclusion criteria from the candidate area of ​​the old furniture bottom surface.

[0016] As an optional approach, generating the bottom contact area of ​​the new furniture includes: reading the bottom contour and contact structure from the new furniture model data; when the contact structure is a continuous base, determining the bottom contour as the bottom contact area of ​​the new furniture; when the contact structure is a leg structure, merging the leg contact areas corresponding to each leg into the bottom contact area of ​​the new furniture.

[0017] As an optional approach, determining the support score includes: identifying the covering grid cells in the support evidence diagram that the contact area of ​​the new furniture's bottom surface covers; reading the direct supporting evidence, indirect supporting evidence, and exclusionary evidence from the covering grid cells; and performing a weighted calculation on the direct supporting evidence, the indirect supporting evidence, and the exclusionary evidence according to the area weight of the covering grid cells in the contact area of ​​the new furniture's bottom surface and the evidence weight coefficients predetermined based on the sample scenario to obtain the support score; the area weight is determined by the overlap area between the covering grid cells and the contact area of ​​the new furniture's bottom surface.

[0018] As an optional approach, determining the spatial location of the new furniture and the rendering pose includes: determining a key contact area from the contact area of ​​the bottom surface of the new furniture; when the key contact area covers a coverage mesh cell that meets the exclusion criteria, excluding the corresponding candidate location of the new furniture; when the key contact area does not cover a coverage mesh cell that meets the exclusion criteria and the support score meets the placement criteria, determining the corresponding candidate location of the new furniture as the spatial location of the new furniture; and generating the rendering pose based on the spatial location of the new furniture and the new furniture model data.

[0019] According to a second aspect of the embodiments of this application, a real-time spatial positioning and scene generation system for furniture integrating AR is also provided, comprising:

[0020] The instance segmentation module is used to acquire RGB images, camera pose, planar information and depth information of the indoor scene, and to perform old furniture instance segmentation on the RGB images to obtain old furniture categories and old furniture instance masks;

[0021] The supporting evidence map building module is used to build a supporting evidence map in the room coordinate system based on the camera pose, the planar information and the depth information. The grid cells of the supporting evidence map record direct supporting evidence, indirect supporting evidence, excluded evidence and unknown states.

[0022] The old furniture bottom surface area generation module is used to extract the contact edge of the old furniture based on the old furniture instance mask, the depth information and the plane information, generate candidate areas of the old furniture bottom surface by combining the old furniture category and furniture size data, write the candidate areas of the old furniture bottom surface into the indirect support evidence of the support evidence diagram, and filter out conflict areas with the exclusion evidence;

[0023] The new furniture candidate location generation module is used to generate the contact area of ​​the bottom surface of the new furniture based on the new furniture model data, and to generate the candidate location of the new furniture in the candidate area of ​​the bottom surface of the old furniture after screening.

[0024] The spatial positioning and rendering posture determination module is used to project the contact area of ​​the bottom surface of the new furniture onto the support evidence map to obtain a covering grid cell, determine the support score based on the area weight of the covering grid cell, direct support evidence, indirect support evidence and exclusion evidence, and determine the spatial position and rendering posture of the new furniture based on the support score and the exclusion evidence.

[0025] This application establishes a support evidence map in the room coordinate system, uniformly recording direct observations of the ground or tabletop, indirect support results inferred from the contact edges of old furniture, and exclusion results corresponding to walls and fixed obstacles. It also combines old furniture instance segmentation, depth information, and planar information to reconstruct occluded candidate bottom areas, allowing for the determination of suitable placement locations for new furniture even without moving the old furniture. Simultaneously, by matching the contact area of ​​the new furniture's bottom surface with the support evidence map, a support score is calculated based on coverage area and evidence type. Conflicting locations are then eliminated by excluding evidence. This reduces misjudgments caused by occlusion, missing depth, and wall interference, minimizing virtual furniture offset, suspension, or interweaving. It enables continuous base furniture and legged furniture to achieve more stable spatial positions and rendering postures, thereby improving the accuracy, real-time performance, and ease of use of AR furniture replacement previews.

[0026] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Furthermore, no embodiment in this disclosure is required to achieve all the effects described above. Attached Figure Description

[0027] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0028] Figure 1 This is a schematic diagram of a method for real-time spatial positioning and scene generation of furniture that integrates AR, provided in an embodiment of this disclosure.

[0029] Figure 2 This is a schematic diagram illustrating the process of establishing supporting evidence diagrams for embodiments of this disclosure.

[0030] Figure 3 This is a schematic diagram of the process for generating candidate areas for the bottom surface of old furniture, provided in an embodiment of this disclosure.

[0031] Figure 4 This is a schematic diagram illustrating the process of determining the spatial location and rendering posture of new furniture as provided in an embodiment of this disclosure.

[0032] Figure 5 This is a schematic diagram of a real-time spatial positioning and scene generation system for furniture that integrates AR, provided as an embodiment of this disclosure.

[0033] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0034] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0035] This disclosure applies to augmented reality (AR) furniture localization and scene generation in indoor scenes using mobile terminals. The mobile terminal can be a mobile phone, tablet, or other electronic device with a camera, pose calculation capabilities, and graphics display capabilities. This method is suitable for scenarios where a user previews new furniture while old furniture is still in the room, such as when an old sofa, bed, TV cabinet, dining table, or cabinet obscures the floor, baseboard, or part of a wall. The mobile terminal acquires color images (RGB images), camera pose, planar information, and depth information, and processes the visible outline of the old furniture, its support area, and the model data of the new furniture in the room coordinate system. The depth information can be a dense depth map or sparse depth points formed by feature points and planar detection results; the planar information includes at least the ground plane, and when a wall or tabletop is detected, it also includes the wall plane or tabletop plane.

[0036] The implementation process of the method described in this application will be described in detail below with reference to specific embodiments. It should be noted that this embodiment is only used to explain this application and is not intended to limit the scope of protection of this application. Conventional adjustments or substitutions of each step by those skilled in the art without departing from the concept of this application should be included in the scope of protection of this application.

[0037] Please see Figure 1 , Figure 1 A flowchart illustrating the real-time spatial positioning and scene generation method for furniture integrating AR according to an embodiment of the present invention is shown, as follows: Figure 1 As shown, the method includes S1-S5:

[0038] In S1, the RGB image of the indoor scene, camera pose, planar information and depth information are acquired, and the old furniture instance segmentation is performed on the RGB image to obtain the old furniture category and old furniture instance mask.

[0039] In this embodiment, the room coordinate system is established by the AR spatial computing unit based on the camera pose. Pixel positions in the RGB image, 3D points in the depth information, planes in the planar information, and geometric data in the furniture model are all converted to the room coordinate system before subsequent processing. Using the same room coordinate system avoids inconsistencies in position representation caused by changes in viewing angle between different image frames. Especially when old furniture obscures the floor, the system can compare the support evidence obtained from multiple frames, the contact edges of the old furniture, and the candidate positions of the new furniture in the same spatial reference.

[0040] The supporting evidence map used in this embodiment refers to two-dimensional grid data established in a room coordinate system. Each grid cell corresponds to a local area on the ground or tabletop, and records evidence related to the support of the furniture's bottom surface in that local area. Each grid cell in the supporting evidence map has at least direct supporting evidence, indirect supporting evidence, exclusionary evidence, and an unknown state. Direct supporting evidence indicates that the grid cell has been directly observed by the camera on the ground or tabletop plane; indirect supporting evidence indicates that the grid cell has not been directly seen, but it is inferred from the contact edge of the old furniture, the type of old furniture, and the furniture size data that it belongs to the area covered by the bottom surface of the old furniture; exclusionary evidence indicates that the grid cell corresponds to a wall, fixed obstacle, door frame, or other space that cannot serve as the bottom surface support of the furniture; an unknown state indicates that the grid cell has not yet obtained sufficient direct supporting evidence, indirect supporting evidence, or exclusionary evidence. The above four types of records are stored in the same grid cell, where direct supporting evidence, indirect supporting evidence, and exclusionary evidence can be represented by confidence values, and the unknown state is represented by the state where each confidence value does not meet the corresponding judgment condition.

[0041] In practical applications, the processor in the mobile terminal performs image acquisition, AR spatial calculation, instance segmentation, supporting evidence map updating, contact edge extraction, generation of candidate regions for the bottom of old furniture, verification of candidate positions for new furniture, and rendering and display. These processes can be completed by a single application or by the mobile terminal in conjunction with a server. For ease of explanation, the following description uses the example of the mobile terminal performing the main calculations locally.

[0042] The image acquisition unit acquires RGB images sequentially according to video frames and sends each RGB image frame to the AR spatial computing unit and the instance segmentation unit. The AR spatial computing unit obtains the camera pose based on feature points in adjacent image frames, inertial measurement data, and camera intrinsic parameters, specifically using existing visual inertial odometry or simultaneous localization and mapping (SMR) methods. This unit also detects ground planes, wall planes, and table planes based on image features, depth information, and multi-frame observation results. For mobile terminals that support depth output, the depth information comes from the depth map provided by the camera system or operating system; for mobile terminals that do not have the function of outputting dense depth maps, the system uses sparse feature points and detected planes to calculate local 3D points as sparse depth information.

[0043] The instance segmentation unit segments the RGB image into old furniture instances, outputting old furniture categories and old furniture instance masks. The old furniture category represents the type of the identified furniture, such as sofa, bed, cabinet, TV stand, dining table, or chair; the old furniture instance mask represents the set of pixels in the RGB image belonging to the same old furniture instance. Optionally, the instance segmentation unit can also output instance bounding regions, which are used to define the processing range of the old furniture instance mask in the image.

[0044] In one implementation, the instance segmentation unit employs an instance segmentation model. This model can adopt a YOLOv8-seg structure, taking an RGB image scaled and normalized to a predetermined size as input and outputting the old furniture category, the instance's bounding region, and an old furniture instance mask. The model includes a backbone network, a feature fusion network, a detection branch, a mask prototype branch, and a mask coefficient branch. The backbone network first extracts low-level texture and edge features through convolutional layers, then obtains multi-scale semantic features through a C2f feature extraction module. The C2f feature extraction module consists of branch convolutions, a bottleneck layer, and feature concatenation, used to extract furniture outlines and local structural features while keeping computationally manageable. A fast spatial pyramid pooling module is placed at the end of the backbone network. This module extracts contextual information at different ranges through multiple pooling scales, and its output serves as the input to the feature fusion network.

[0045] The feature fusion network employs a combination of a feature pyramid network and a path aggregation network. The feature pyramid network upsamples high-level semantic features and concatenates them with low-level spatial features, while the path aggregation network then propagates the fused low-level features to higher levels. After passing through this network, the model generates fused feature maps at multiple scales. The detection branch performs category prediction and bounding box regression on the fused feature maps at each scale, outputting the old furniture category and instance bounding region. The mask prototype branch outputs a set of full-image mask prototypes from the high-resolution fused feature maps. The mask coefficient branch outputs a set of instance mask coefficients for each old furniture instance obtained by the detection branch. During inference, the system linearly combines the instance mask coefficients with the full-image mask prototypes and clips the combination result according to the instance bounding region to obtain the old furniture instance mask. The outputs of the detection branch, mask prototype branch, and mask coefficient branch are all processed in subsequent steps. The old furniture category is used to generate the old furniture bottom candidate region, the old furniture instance mask is used to extract the old furniture contact edges, and the instance bounding region is used to limit the mask processing range.

[0046] In some embodiments, when training the instance segmentation model, the training samples include indoor RGB images, old furniture category annotations, instance bounding region annotations, and instance mask annotations. The training losses include classification loss, bounding box regression loss, distribution focus loss, and mask loss; the classification loss constrains the old furniture category prediction results, the bounding box regression loss constrains the instance bounding region, the distribution focus loss refines the boundary positions, and the mask loss constrains the old furniture instance masks. The weight coefficients corresponding to each loss term are non-negative parameters, and their values ​​can be configured between 0 and 1, or they can be configured as a normalized set of positive numbers; these weight coefficients are determined statistically through a validation set, for example, by comparing category accuracy, bounding box overlap, and mask intersection-union ratio in the validation set and selecting a set of weight coefficients with a smaller overall error. After the model training is completed, the above weight coefficients are fixed during the inference phase and are not updated in individual room sessions.

[0047] In S2, based on the camera pose, the planar information, and the depth information, a supporting evidence map is established in the room coordinate system. The grid cells of the supporting evidence map record direct supporting evidence, indirect supporting evidence, excluded evidence, and unknown states.

[0048] Specifically, please refer to Figure 2 , Figure 2 A schematic diagram illustrating the supporting evidence graph creation process provided in this disclosure is shown. For example... Figure 2 As shown, at 201, the planar information and depth information corresponding to the current image frame are transformed to the room coordinate system.

[0049] For observation points on the ground plane or tabletop plane, the system calculates their 3D position in the room coordinate system based on the camera pose, and then projects this 3D position onto the grid cells of the supporting evidence map; the projected grid cells form direct supporting evidence. If the same grid cell is observed as the ground plane or tabletop plane in multiple frames, and its height is within the height tolerance relative to the corresponding plane, the system increases the confidence value corresponding to the direct supporting evidence of that grid cell. The height tolerance is used to absorb depth measurement errors from ordinary mobile terminals, and its value can be obtained by device calibration or statistically determined based on depth fluctuations in stable ground areas within the current room session; in one example, the height tolerance is determined based on the distance distribution from stable ground points to the fitted ground plane and is updated as stable ground samples increase within the same room session.

[0050] In frame 202, for spatial areas that cannot serve as furniture supports, such as wall surfaces, fixed obstacles, or door frames, the system assigns their projected areas on the ground to corresponding grid cells in the support evidence map, forming exclusion evidence. If a grid cell corresponds to a stable vertical structure across multiple frames, or if its depth and height consistently conflict with the ground plane, the system increases the confidence value corresponding to the exclusion evidence for that grid cell. For human legs, pets, or temporary objects appearing only in a few frames, the system determines whether they are fixed obstacles based on the stability of their position in the room coordinate system; if the area moves significantly over time, the relevant grid cell remains in its original state or becomes unknown; if the area maintains the same spatial position in multiple views and conflicts with the ground support relationship, the system further increases the confidence value corresponding to the exclusion evidence. This process ensures that short-term occlusion does not overwrite previously obtained direct support evidence.

[0051] In section 203, for grid cells that lack both direct supporting evidence and exclusionary evidence, the system records them as an unknown state. Grid cells in an unknown state do not participate in the direct placement decision, nor are they directly determined as unplaceable. Floors obscured by old furniture, low-texture floors, reflective tile areas, carpet edges, and areas lacking depth may all enter an unknown state at this stage.

[0052] In one implementation, each grid cell in the supporting evidence map stores a direct support confidence value, an indirect support confidence value, and an exclusion confidence value. All three confidence values ​​range from 0 to 1. The direct support confidence value is updated based on observations of the ground plane or tabletop plane; the indirect support confidence value is written from candidate areas on the underside of old furniture; and the exclusion confidence value is updated from walls, fixed obstacles, door frames, or stable non-supported areas. The system can also store the number of direct observations, the number of indirect inferences, and the number of conflicting observations as auxiliary data for confidence value updates. The number of direct observations indicates how many times the grid cell was directly observed as a supporting area by the camera; the number of indirect inferences indicates how many times the grid cell was covered by candidate areas on the underside of old furniture; and the number of conflicting observations indicates how many times the grid cell was observed as a fixed obstacle or a non-supported area. The supporting evidence map generation unit can update the confidence values ​​using a recursive method. When the current frame provides direct support observations, the system calculates the observation confidence level based on the plane fitting residual, depth noise, and observation angle, and adds this confidence level to the direct support confidence value. When the current frame provides excluded observations, the system adds the excluded observation confidence level to the excluded confidence value. The observation confidence level is a value between 0 and 1, which can be obtained from validation set statistics or device calibration. During recursive updates, the system prunes the upper limit of the confidence value to ensure it does not exceed 1. If a grid cell has not been observed for a long period, the system retains its existing confidence value to avoid changing the support judgment simply because the viewpoint is not covered.

[0053] Specifically, the determination of direct support confidence values ​​involves a default value of 0 (no supporting evidence). When a grid cell is observed as ground or a tabletop for the first time in the current frame (i.e., forming direct supporting evidence), the system calculates the observation confidence (a value between 0 and 1) by weighted summing the normalized plane fitting residual, depth noise, and observation angle of that frame. Since there is no accumulated base value beforehand, the observation confidence of this observation becomes the direct support confidence value of that grid cell (equivalent to accumulating this value from 0). In subsequent frames, if the grid cell is observed as a supporting region again, the system continues to calculate the observation confidence of the current frame and adds it to the original value (i.e., the original value plus the new observation confidence), while performing upper limit pruning to ensure it does not exceed 1. If the grid cell has not been observed for a long period of time, the original value remains unchanged.

[0054] The determination of the confidence value of indirect supporting evidence includes the following steps: First, when generating candidate areas for the bottom surface of old furniture, the system calculates the candidate confidence level of the area (detailed below). This confidence level is determined by the weighted average of contact edge stability, planar consistency, and consistency of supporting evidence (all three are between 0 and 1). During writing, the grid cells to be written that are covered by the candidate areas for the bottom surface of old furniture are identified as objects. Only grid cells that do not meet the exclusion criteria are given an increased indirect supporting confidence value. At the same time, the number of indirect inferences is used as auxiliary data for updating. If the area is excluded by the evidence, the corresponding cell is not written or is removed.

[0055] Regarding the determination of the exclusion confidence value, the default value is 0, indicating that no exclusion observations have been received. When a grid cell is observed for the first time in the current frame as an unsupportable area such as a wall, fixed obstacle, or door frame (i.e., exclusion evidence is formed), the system calculates the observation confidence (a value between 0 and 1) by weighted summing the normalized plane fitting residual, depth noise, and observation angle of that frame. Since there is no accumulated base value beforehand, the observation confidence value of this observation becomes the exclusion confidence value of that grid cell (equivalent to accumulating this value from 0). In subsequent frames, if the grid cell is observed as an excluded area again, the system first determines the stability of the area's position in the room coordinate system: If the area maintains the same spatial position from multiple perspectives and continuously conflicts with the ground support relationship (such as a stable vertical wall), the system calculates the observation confidence of the current frame (the calculation method is the same as above, and will not be repeated), and adds it to the original exclusion confidence value (i.e., the original value plus the new observation confidence), while performing upper limit pruning to ensure that it does not exceed 1; if the area moves significantly over time (such as a human leg or a temporary object), the system does not accumulate the exclusion confidence value, and the relevant grid cell maintains its original state or becomes an unknown state (equivalent to removing this observation); if a grid cell has not been covered by new excluded observations for a long time, the system maintains its existing exclusion confidence value, which will not automatically decay or return to zero, in order to preserve historical stability evidence. In S3, the contact edges of the old furniture are extracted based on the old furniture instance mask, the depth information and the planar information. The old furniture category and furniture size data are combined to generate candidate areas for the bottom surface of the old furniture. The candidate areas for the bottom surface of the old furniture are written into the indirect supporting evidence of the supporting evidence diagram, and conflict areas are screened out with the exclusion evidence.

[0056] Specifically, please refer to Figure 3 , Figure 3 A schematic diagram illustrating the process for generating candidate areas for the bottom surface of old furniture, provided in an embodiment of this disclosure, is shown. Figure 3 As shown, at 301, the mask outline is extracted from the old furniture instance mask.

[0057] In some embodiments, the mask contour can be obtained through boundary tracking, and the system acquires contour sampling points along the mask contour. The sampling interval is determined based on the image resolution and the size of the instance's bounding region; optionally, a larger sampling interval is used for a larger instance bounding region, and a smaller sampling interval is used for a smaller instance bounding region, so that the number of contour sampling points is within a range suitable for mobile computing. After acquiring the contour sampling points, the system calculates the 3D position of each contour sampling point in the room coordinate system based on depth information and camera pose. If a contour sampling point lacks a depth value, the system searches for a reliable depth point in the neighborhood within the old furniture instance mask, and estimates the 3D position of the contour sampling point using the neighborhood depth point and the camera pose; if the neighborhood also lacks a reliable depth, the system marks the contour sampling point as an invalid sampling point.

[0058] In step 302, the height difference between the effective contour sampling points and the ground plane is calculated. If the height difference of a set of consecutive effective contour sampling points satisfies the height difference condition, and these sampling points are continuously distributed along the same local direction in the room coordinate system, the system merges this set of consecutive sampling points into a ground contact edge. The height difference condition adopts the judgment range corresponding to the aforementioned height tolerance, thus ensuring that the tolerance source used to determine ground support remains consistent within the same room session. For floor furniture such as sofas, cabinets, and beds, the ground contact edge usually corresponds to its front edge, side edge, or local bottom contour. For table and chair furniture, the ground contact edge can be represented as a short contour segment near the legs. The system only retains ground contact edges that are spatially close in multiple frames, reducing the reliability of short boundary segments formed by shadows, fabric sagging, or segmentation burrs in a single frame.

[0059] In step 303, the distance from the valid contour sampling points to the wall plane is calculated, and it is determined whether the direction of the corresponding contour segment matches the projection direction of the wall plane onto the ground. If a set of consecutive valid contour sampling points satisfies the distance condition and their contour directions are continuous, the system merges this set of sampling points into adjacent wall edges. The range of the distance condition can be determined based on the wall plane fitting residual and depth noise. This range uses the same calculation method within the same room session to avoid instability of adjacent wall edges caused by using different standards in different frames. Adjacent wall edges are used to constrain the boundary of the bottom surface of old furniture that is close to the wall. When both the ground contact edge and the adjacent wall edge exist, the system uses both to generate the bottom surface range; when the adjacent wall edge does not exist, the system generates the bottom surface range based on the ground contact edge, the type of old furniture, and the furniture size data.

[0060] In section 304, the system selects a bottom surface representation method based on the type of old furniture and generates candidate areas for the bottom surface of the old furniture. For sofas, beds, TV cabinets, and floor cabinets, the system uses a continuous bottom surface representation method, which represents the coverage area of ​​the bottom surface of the old furniture as a rectangular or approximately rectangular area. For dining tables, chairs, and other legged furniture, the system uses a leg bottom surface representation method, which represents the coverage area of ​​the bottom surface of the old furniture as several leg contact areas and their outer boundaries.

[0061] In some embodiments, under the continuous bottom surface representation method, the system takes the ground contact edge as the first boundary. If the adjacent wall edge exists, the system takes the adjacent wall edge or its projection on the ground plane as the second boundary, and forms the old furniture bottom surface candidate area with the area between the first boundary and the second boundary. If the adjacent wall edge does not exist, the system determines the depth direction range according to the furniture size data corresponding to the old furniture category, and then forms the old furniture bottom surface candidate area along the normal direction of the ground contact edge.

[0062] Furniture size data can come from user-selected new furniture model data, old furniture product data, or furniture category size statistics. When the user has already selected a new furniture model before generating candidate areas for the bottom surface of old furniture, and the dimensions of that new furniture model are used as a replacement reference, the furniture size data can also come from the new furniture model data. Category size statistics are obtained from the furniture product database before application release; for furniture with large size differences, such as beds or sofas, the system can retain multiple candidate areas for the bottom surface of old furniture corresponding to multiple depth directions, which are then filtered out using supporting evidence and exclusionary evidence.

[0063] In some embodiments, under the support leg bottom surface representation method, the system generates candidate areas for the bottom surface of old furniture based on the support leg contact position corresponding to the ground contact edge. Each support leg contact position corresponds to a local support area, and the outer range of multiple local support areas forms the candidate areas for the bottom surface of old furniture; if the system only sees part of the support legs, the bottom surface range inference unit generates one or more candidate outer ranges based on the seen support leg contact positions, old furniture category, and furniture size data.

[0064] In step 305, after each candidate area of ​​the bottom surface of old furniture is generated, the candidate confidence level of the area is calculated, and the candidate areas of the bottom surface of old furniture that meet the candidate confidence level requirements are written into the indirect supporting evidence of the supporting evidence diagram.

[0065] In some embodiments, candidate credibility is jointly determined by contact edge stability, planar consistency, and supporting evidence consistency. Contact edge stability indicates whether the ground contact edge or adjacent wall edge corresponds to the same room coordinate position in multiple frames; planar consistency indicates whether the geometric relationship between the contact edge and the ground plane or wall plane satisfies the corresponding conditions; supporting evidence consistency indicates whether the directly visible part of the candidate area of ​​the old furniture bottom surface is consistent with the direct supporting evidence in the supporting evidence diagram, while avoiding existing excluded evidence.

[0066] The three quantities mentioned above all take values ​​between 0 and 1. The system obtains the candidate credibility through a weighted average. The candidate credibility can be expressed as the weighted average of contact edge stability, planar consistency, and supporting evidence consistency. The weighting coefficients used are predetermined non-negative numbers with values ​​ranging from 0 to 1, and are normalized during the calculation so that the sum of the three weighting coefficients is 1. These weighting coefficients can be obtained statistically from a validation set with actual furniture footprint markings. If the contact edge stability in the validation set has a significant impact on footprint recovery, the corresponding weight will be higher. If the wall detection error of a certain device is large, the planar consistency weight can be reduced through device calibration. The weighting coefficients are fixed during the inference phase and do not change temporarily with the results of a single frame.

[0067] In some embodiments, when writing indirect supporting evidence to the supporting evidence map, the grid cells to be written are first determined to be covered by the candidate area of ​​the old furniture bottom surface, and then the exclusion evidence for each grid cell to be written is read. For grid cells to be written that do not meet the exclusion criteria, the system increases the confidence value corresponding to their indirect supporting evidence. For grid cells to be written that meet the exclusion criteria, the system removes the grid cell from the candidate area of ​​the old furniture bottom surface and does not write indirect supporting evidence to it. The exclusion criteria are determined by the exclusion threshold, which is a value between 0 and 1 and can be obtained statistically from the samples of misplaced fixed obstacles in the validation set. The exclusion threshold remains consistent in the same room session to avoid using different exclusion criteria for the same grid cell in consecutive frames.

[0068] When the candidate area of ​​the old furniture bottom surface after rejection still has a sufficiently continuous coverage, the system outputs it as the candidate area of ​​the old furniture bottom surface to the next processing step; if the rejected area is broken, or the grid cells corresponding to the boundary used to form the candidate area of ​​the old furniture bottom surface are all covered by excluded evidence, the system discards the candidate area of ​​the old furniture bottom surface; for multiple candidate areas of the old furniture bottom surface, the system performs writing and rejection processing respectively, and outputs one or more candidate areas of the old furniture bottom surface after rejection; candidate areas that fail to meet the credibility requirements or are rejected will not be included in the generation of new furniture candidate positions.

[0069] Therefore, by utilizing the visible outline of the old furniture and the geometric relationship between the ground and the wall, a portion of the obscured area is transformed into indirect supporting evidence, and the parts that conflict with the wall and fixed obstacles are removed by excluding evidence before writing; this process transforms the old furniture instance mask into spatial evidence that can participate in the three-dimensional placement judgment.

[0070] In S4, a new furniture bottom contact area is generated based on the new furniture model data, and a new furniture candidate position is generated within the old furniture bottom candidate area after filtering.

[0071] In some embodiments, the new furniture model data includes model dimensions, bottom contour, contact structure, and model coordinate system; the bottom contour represents the outer contour of the new furniture on the placement surface, and the contact structure represents the structural type of the new furniture that supports the ground or tabletop. If the contact structure is a continuous base, the system determines the bottom contour as the bottom contact area of ​​the new furniture; if the contact structure is a leg structure, the system merges the leg contact areas corresponding to each leg into the bottom contact area of ​​the new furniture; each leg contact area can be obtained from model preprocessing or from the bottom geometry of the leg components in the model; the bottom contact area of ​​the new furniture is then used for support score calculation and exclusion determination, and decorative components or upper structures that do not enter this area do not participate in the bottom support verification.

[0072] New furniture candidate locations are generated within the candidate areas on the bottom surfaces of the old furniture that have been filtered out. For each candidate area on the bottom surface of the old furniture that has been filtered out, the system uses the location and orientation of that area as an initial reference to place the contact area of ​​the bottom surface of the new furniture within that area. If the size of the new furniture is the same as or close to the size of the candidate area on the bottom surface of the old furniture, the system generates an initial candidate location for the new furniture. If the new furniture is wider or deeper, the system performs a local search near the candidate area on the bottom surface of the old furniture. The local search is only performed around the candidate area on the bottom surface of the old furniture that has been filtered out, and the search variables include the position and orientation of the new furniture on the ground plane. The search range is jointly defined by the boundary of the candidate area on the bottom surface of the old furniture, the wall plane, and the supporting evidence diagram to avoid generating irrelevant locations throughout the entire room.

[0073] When there are multiple candidate areas for the bottom surface of old furniture after screening, the system generates corresponding candidate positions for new furniture. Each candidate position for new furniture records its source region, position, orientation, and the corresponding new furniture bottom surface contact area covering mesh cell. The source region is used to trace which candidate area for the bottom surface of old furniture it comes from when selecting a position later. The position and orientation are used to generate the rendering pose. The covering mesh cell is used to calculate the support score and make the exclusion judgment. If a candidate area for the bottom surface of old furniture after screening cannot accommodate the contact area of ​​the bottom surface of new furniture, and there is no placement method that meets the basic geometric boundary within the local search range, the system discards the candidate position for new furniture corresponding to that area.

[0074] In S5, the contact area of ​​the bottom surface of the new furniture is projected onto the support evidence diagram. The support score is determined based on the direct support evidence, the indirect support evidence, and the exclusion evidence. The spatial position and rendering posture of the new furniture are determined based on the support score and the exclusion evidence.

[0075] In some embodiments, please refer to Figure 4 , Figure 4 A schematic diagram illustrating the process of determining the spatial location and rendering posture of new furniture according to an embodiment of this disclosure is shown. Figure 4 As shown, at 401, the contact area of ​​the bottom surface of the new furniture at each candidate position is projected onto the support evidence map to obtain the covered grid cell.

[0076] For each covered grid cell, the system reads direct supporting evidence, indirect supporting evidence, and exclusionary evidence; among them, exclusionary evidence enters the exclusion judgment first, and direct supporting evidence and indirect supporting evidence enter the support score calculation together with exclusionary evidence; the key contact area is determined from the contact area of ​​the bottom surface of the new furniture; for furniture with continuous base, the key contact area includes the bottom edge area and the area near the center of gravity projection; for furniture with leg structure, the key contact area includes the contact area of ​​each leg.

[0077] In section 402, the support score is calculated based on the area weight of the covering grid cell, direct supporting evidence, indirect supporting evidence, and exclusionary evidence. The area weight is determined by the overlap area between the covering grid cell and the bottom surface of the new furniture; the larger the overlap area, the greater the influence of the covering grid cell on the support score. As an example, the following calculation method can be used:

[0078]

[0079] in, Indicates a new furniture candidate location. This represents the set of cover grid cells representing the contact area of ​​the bottom surface of the new furniture at the candidate location. This indicates one of the covering grid cells. Indicates the covering grid cell The overlapping area of ​​the contact area with the bottom surface of the new furniture. Indicates the covering grid cell The confidence value corresponding to the direct supporting evidence. Indicates the covering grid cell The confidence value corresponding to the indirect supporting evidence, Indicates the covering grid cell The confidence value corresponding to the exclusion of evidence, Indicates the candidate positions for new furniture Support rating.

[0080] Weighting coefficient , , These represent the contribution coefficients of direct supporting evidence, indirect supporting evidence, and exclusionary evidence in the support score, respectively. All three are non-negative numbers, and their values ​​can be configured between 0 and 1. When participating in the support score calculation, they satisfy α+β+γ=1. The above weight coefficients can be obtained by statistically analyzing sample scenarios labeled with placeable and non-placeable positions. In one example, the system iterates through multiple sets of weight coefficients in the validation set and selects a set of parameters with fewer misplaced items and stable identification of placeable positions.

[0081] At step 403, the spatial location of the new furniture is determined based on the support score and the exclusion evidence. If the critical contact area covers a coverage grid cell that meets the exclusion criteria, the system discards the corresponding candidate location for the new furniture and no longer compares its support score. If the critical contact area does not cover a coverage grid cell that meets the exclusion criteria, the system compares the support score of the candidate location for the new furniture with the placement criteria. The placement criteria are represented by a placement threshold, which is a value between 0 and 1. This value is obtained through statistical analysis of sample scenarios and can also be set before application release based on misplacement rate requirements.

[0082] New furniture candidate locations whose supporting scores meet the placement criteria are added to the retention set, while those that do not are discarded. A supporting score meeting the placement criteria means that the supporting score used for placement determination is not lower than the placement threshold. If multiple new furniture candidate locations exist in the retention set, the system selects the location with the highest supporting score as the new furniture's spatial location. If the supporting scores are the same, the system can select a location with a smaller difference in the direction of its contact edge with the old furniture. This directional difference is only used when the supporting scores are the same to avoid overriding the supporting evidence judgment.

[0083] At position 404, a rendering pose is generated based on the spatial location and model data of the new furniture. The spatial location of the new furniture provides the translation component of the model in the room coordinate system, and the orientation corresponding to the candidate position provides the rotation component. The model data of the new furniture provides the relationship between the model coordinate system and the model geometry. The rendering and display unit places the new furniture model in the room coordinate system according to the translation and rotation components, and then projects it onto the RGB image for display according to the current camera pose. The old furniture instance mask can participate in the occlusion processing at the display level. For example, when the old furniture still blocks part of the new furniture model, the rendering and display unit clips the occluded virtual pixels according to the old furniture instance mask and depth information.

[0084] In some embodiments, after the spatial location of the new furniture is determined, the system continues to acquire subsequent RGB images, subsequent camera poses, subsequent planar information, and subsequent depth information. The AR spatial calculation unit converts the subsequent planar information and subsequent depth information to the room coordinate system, and the supporting evidence generation unit updates the direct supporting evidence, indirect supporting evidence, excluded evidence, and unknown states accordingly. If a grid cell covered by the contact area of ​​the bottom surface of the new furniture changes from an unknown state to direct supporting evidence, the system recalculates the support score for the spatial location of the new furniture. If the grid cell changes to excluded evidence, the system re-executes the exclusion judgment. The system only recalculates the corresponding support score when the updated grid cell belongs to the covering grid cell covered by the contact area of ​​the bottom surface of the new furniture. If the updated grid cell is unrelated to the current spatial location of the new furniture, the system maintains the current rendering posture. This process reduces the impact of irrelevant spatial changes on real-time rendering.

[0085] In a specific scenario, such as a user previewing a replacement sofa in their living room; the mobile terminal acquires an RGB image of the area where the old sofa is located, the instance segmentation unit outputs the old furniture category as sofa, and outputs an old furniture instance mask; the AR spatial calculation unit detects the ground plane and the wall plane behind the old sofa, but the ground below the old sofa is obscured; the supporting evidence map generation unit generates direct supporting evidence based on the visible ground and excludes evidence based on the wall projection; the grid units below the old sofa that are not directly seen remain in an unknown state; the contact edge extraction unit extracts the ground contact edge from the lower edge of the old furniture instance mask and extracts the adjacent wall edges based on the wall plane; the bottom surface range inference unit... The ground contact edge serves as the first boundary, and the adjacent wall edge serves as the second boundary, generating candidate areas for the bottom surface of old furniture. Grid cells that meet the exclusion criteria are then removed, and the remaining areas are written into the indirect support evidence of the support evidence diagram. After the user selects a new sofa model, the candidate position generation unit reads the bottom contour from the new furniture model data and determines the contact area of ​​the new furniture's bottom surface. New furniture candidate positions are generated within the filtered old furniture bottom surface candidate areas, and the spatial position of the new furniture is determined based on the support score and exclusion evidence. In this scenario, the furniture size data used to generate the old furniture bottom surface candidate areas can come from old furniture category size statistics or old furniture product data.

[0086] In some embodiments, the grid size of the supporting evidence map can be configured according to the device's computing power and the furniture's positioning accuracy; the smaller the grid size, the finer the match between the bottom contact area and the supporting evidence map; the larger the grid size, the lower the computational load; in one example, the system jointly determines the grid size based on the width of common furniture legs, the effective resolution of the depth map, and rendering error requirements; once the grid size is determined, it remains unchanged in the same room session to avoid changes in the correspondence between grid cells; if the user restarts the room scan, the system reinitializes the room coordinate system and the supporting evidence map.

[0087] In actual implementation, the system can store in memory for each old furniture instance the old furniture category, old furniture instance mask, ground contact edge, wall adjacent edge, old furniture bottom candidate area, and filtered old furniture bottom candidate area; the old furniture instance mask generates contact edges and is still used for rendering occlusion; ground contact edge and wall adjacent edge are used to generate old furniture bottom candidate area; after filtering out evidence, qualified areas of old furniture bottom candidate area are used to generate new furniture candidate positions, and unqualified areas are discarded; each new furniture candidate position stores the source area, position, orientation, new furniture bottom contact area, cover mesh cell, and support score; after the support score is calculated, candidate positions that do not meet the judgment conditions are discarded, and candidate positions that meet the judgment conditions are added to the retention set; finally, the selected new furniture spatial position and rendering posture are entered into the rendering display unit.

[0088] Through the above implementation method, when old furniture obscures the supporting surface, the mobile terminal saves directly observed ground or tabletop information, indirect support information inferred from the contact edges of the old furniture, and exclusion information formed by walls and fixed obstacles into a unified support evidence map. When new furniture is placed, the system uses the grid cells covered by the contact area of ​​the new furniture's bottom surface as the judgment object, and determines the spatial position and rendering posture of the new furniture based on the support score and exclusion evidence. This method reduces misjudgments of placeable areas caused by old furniture obstruction and reduces the situation where virtual furniture deviates from the area occupied by the original furniture, making it suitable for AR furniture replacement previews without requiring users to move the old furniture.

[0089] Please see Figure 5 , Figure 5 This is a schematic diagram of a real-time spatial positioning and scene generation system for furniture integrating AR, provided in an embodiment of this application. As shown in the figure, the system includes:

[0090] The instance segmentation module 501 is used to acquire RGB images, camera pose, planar information and depth information of an indoor scene, and to perform old furniture instance segmentation on the RGB images to obtain old furniture categories and old furniture instance masks.

[0091] The supporting evidence map building module 502 is used to build a supporting evidence map in the room coordinate system based on the camera pose, the planar information and the depth information. The grid cells of the supporting evidence map record direct supporting evidence, indirect supporting evidence, excluded evidence and unknown states.

[0092] The old furniture bottom surface area generation module 503 is used to extract the old furniture contact edge based on the old furniture instance mask, the depth information and the plane information, generate the old furniture bottom surface candidate area by combining the old furniture category and furniture size data, write the old furniture bottom surface candidate area into the indirect support evidence of the support evidence diagram, and filter out conflict areas with the exclusion evidence;

[0093] The new furniture candidate location generation module 504 is used to generate a new furniture bottom surface contact area based on the new furniture model data, and to generate a new furniture candidate location within the old furniture bottom surface candidate area after filtering.

[0094] The spatial positioning and rendering posture determination module 505 is used to project the contact area of ​​the bottom surface of the new furniture onto the support evidence map to obtain a covering grid cell, determine the support score based on the area weight of the covering grid cell, direct support evidence, indirect support evidence and exclusion evidence, and determine the spatial position and rendering posture of the new furniture based on the support score and the exclusion evidence.

[0095] Each processing unit and / or module in the embodiments of this application can be implemented by an analog circuit that implements the functions described in the embodiments of this application, or by software that executes the functions described in the embodiments of this application.

[0096] Please see Figure 6 It shows a schematic diagram of the structure of an electronic device according to an embodiment of this application, which can be used to implement... Figure 1 The method in the illustrated embodiment. (As shown) Figure 6 As shown, the electronic device may include:

[0097] The system includes at least one processor 601, at least one network interface 604, a user interface 603, a memory 605, and at least one communication bus 602. The communication bus 602 is used to enable connection and communication between the components. The user interface 603 may include buttons, and optionally include a standard wired or wireless interface. The network interface 604 may include, but is not limited to, a Bluetooth module, an NFC module, a Wi-Fi module, etc.

[0098] The processor 601 may include one or more processing cores and connect to various parts within the electronic device through various interfaces and lines. It implements various functions and data processing of the electronic device by running or executing instructions, programs, code sets, or instruction sets stored in the memory 605, and by accessing data in the memory 605. Optionally, the processor 601 may be implemented using at least one hardware form of DSP, FPGA, or PLA. The processor 601 may also integrate one or more combinations of CPU, GPU, and modem.

[0099] Memory 605 may include random access memory (RAM) or read-only memory (ROM). Optionally, memory 605 includes a non-transitory computer-readable medium for storing instructions, programs, code, code sets, or instruction sets. Memory 605 may be divided into a program storage area and a data storage area, wherein the program storage area can be used to store instructions for implementing an operating system and instructions for implementing the foregoing method embodiments; the data storage area can be used to store data related to the relevant method embodiments. Memory 605 may also be at least one storage device located remotely from processor 601. Figure 6 As shown, the memory 605, which serves as a computer storage medium, may contain an operating system, a network communication module, a user interface module, and program instructions.

[0100] In particular, the methods and / or embodiments in this application can be implemented as computer software programs. For example, the embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. When the computer program is executed by processor 601, it performs the functions defined in the methods of this application.

[0101] Another embodiment of this application provides a storage medium storing computer program instructions thereon, which can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of this application described above.

[0102] In the above embodiments, the descriptions of each embodiment have different focuses. Parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. The above descriptions are merely preferred embodiments of this application and explanations of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to the technical solutions formed by specific combinations of the above technical features, but should also cover other technical solutions formed by arbitrary combinations of the above technical features or their equivalent features without departing from the inventive concept.

Claims

1. A method for real-time spatial positioning and scene generation of furniture integrating AR, characterized in that, include: Acquire RGB images, camera pose, planar information, and depth information of the indoor scene; segment the RGB images into old furniture instances to obtain old furniture categories and old furniture instance masks. Based on the camera pose, the planar information, and the depth information, a supporting evidence map is established in the room coordinate system. The grid cells of the supporting evidence map record direct supporting evidence, indirect supporting evidence, excluded evidence, and unknown states. Establishing the supporting evidence map includes: transforming the observation points representing the ground plane or tabletop plane in the planar information to the room coordinate system according to the camera pose and assigning them to the corresponding grid cells to form the direct supporting evidence; assigning the projection areas representing the wall plane or fixed obstacles in the planar information to the corresponding grid cells to form the excluded evidence; and recording the grid cells that do not form direct supporting evidence or excluded evidence as the unknown states. Based on the old furniture instance mask, the depth information, and the planar information, the contact edges of the old furniture are extracted. Combined with the old furniture category and size data, candidate areas for the bottom surface of the old furniture are generated. These candidate areas are then written into the indirect supporting evidence of the supporting evidence diagram. Conflict areas are filtered out using the exclusion evidence. Specifically, this includes: determining the grid cells covered by the candidate areas for the bottom surface of the old furniture as the grid cells to be written; reading the exclusion evidence for each grid cell to be written; writing the indirect supporting evidence into grid cells that do not meet the exclusion criteria; and removing grid cells that meet the exclusion criteria from the candidate areas for the bottom surface of the old furniture. Based on the new furniture model data, a new furniture bottom surface contact area is generated, and a new furniture candidate position is generated within the old furniture bottom surface candidate area after screening. The contact area of ​​the bottom surface of the new furniture is projected onto the support evidence map to obtain a covering grid cell. The support score is determined based on the area weight of the covering grid cell, direct support evidence, indirect support evidence, and exclusion evidence. The spatial position and rendering posture of the new furniture are determined based on the support score and the exclusion evidence.

2. The method according to claim 1, characterized in that, The old furniture instance segmentation of the RGB image includes: inputting the RGB image into an instance segmentation model, which includes a backbone network, a feature fusion network, a detection branch, a mask prototype branch, and a mask coefficient branch connected in sequence; the detection branch outputs the old furniture category and the instance bounding region, the mask prototype branch outputs the full-image mask prototype, and the mask coefficient branch outputs the instance mask coefficients; combining the full-image mask prototype according to the instance mask coefficients, and cropping the combination result according to the instance bounding region to obtain the old furniture instance mask.

3. The method according to claim 1, characterized in that, Extracting the old furniture contact edges includes: extracting a mask contour from the old furniture instance mask, and obtaining contour sampling points along the mask contour; calculating the three-dimensional position of the contour sampling points in the room coordinate system based on the depth information and the camera pose; grouping contour sampling points whose height satisfies the height difference condition with the ground plane and are continuously distributed into ground contact edges; grouping contour sampling points whose distance satisfies the distance condition with the wall plane and are continuously oriented into adjacent wall edges; the old furniture contact edges include the ground contact edges and the adjacent wall edges.

4. The method according to claim 3, characterized in that, Generating the candidate area for the bottom surface of the old furniture includes: selecting a bottom surface representation method according to the category of the old furniture; when the bottom surface representation method is a continuous bottom surface, taking the ground contact edge as the first boundary and the adjacent edge of the wall or the depth direction range defined by the furniture size data as the second boundary to generate the candidate area for the bottom surface of the old furniture; when the bottom surface representation method is a leg bottom surface, generating the candidate area for the bottom surface of the old furniture according to the leg contact position corresponding to the ground contact edge.

5. The method according to claim 4, characterized in that, Generating the bottom contact area of ​​the new furniture includes: reading the bottom contour and contact structure from the new furniture model data; when the contact structure is a continuous base, determining the bottom contour as the bottom contact area of ​​the new furniture; when the contact structure is a leg structure, merging the leg contact areas corresponding to each leg into the bottom contact area of ​​the new furniture.

6. The method according to claim 5, characterized in that, Determining the support score includes: identifying the covering grid cells in the supporting evidence diagram that the contact area of ​​the bottom surface of the new furniture covers; reading the direct supporting evidence, indirect supporting evidence, and exclusionary evidence of the covering grid cells; and performing a weighted calculation on the direct supporting evidence, the indirect supporting evidence, and the exclusionary evidence according to the area weight of the covering grid cells in the contact area of ​​the bottom surface of the new furniture, and the evidence weight coefficients predetermined according to the sample scenario, to obtain the support score; the area weight is determined by the overlap area between the covering grid cells and the contact area of ​​the bottom surface of the new furniture.

7. The method according to claim 6, characterized in that, Determining the spatial location of the new furniture and the rendering pose includes: identifying a key contact area from the contact area on the bottom surface of the new furniture; excluding the corresponding candidate location of the new furniture when the key contact area covers a coverage mesh cell that meets the exclusion criteria; determining the corresponding candidate location of the new furniture as the spatial location of the new furniture when the key contact area does not cover a coverage mesh cell that meets the exclusion criteria and the support score meets the placement criteria; and generating the rendering pose based on the spatial location of the new furniture and the new furniture model data.

8. A real-time spatial positioning and scene generation system for furniture integrating AR, characterized in that, include: The instance segmentation module is used to acquire RGB images, camera pose, planar information and depth information of the indoor scene, and to perform old furniture instance segmentation on the RGB images to obtain old furniture categories and old furniture instance masks; A supporting evidence map building module is used to build a supporting evidence map in a room coordinate system based on the camera pose, the planar information, and the depth information. The grid cells of the supporting evidence map record direct supporting evidence, indirect supporting evidence, excluded evidence, and unknown states. Building the supporting evidence map includes: transforming the observation points representing the ground plane or tabletop plane in the planar information to the room coordinate system according to the camera pose and assigning them to the corresponding grid cells to form the direct supporting evidence; assigning the projection areas representing the wall plane or fixed obstacles in the planar information to the corresponding grid cells to form the excluded evidence; and recording the grid cells that do not form direct supporting evidence or excluded evidence as the unknown states. The old furniture bottom surface region generation module is used to extract the contact edges of the old furniture based on the old furniture instance mask, the depth information, and the planar information; generate candidate regions for the bottom surface of the old furniture by combining the old furniture category and furniture size data; write the candidate regions for the bottom surface of the old furniture into the indirect support evidence of the support evidence diagram; and filter out conflict areas using the exclusion evidence. Specifically, this includes: determining the grid cells covered by the candidate regions for the bottom surface of the old furniture as the grid cells to be written; reading the exclusion evidence for each grid cell to be written; writing the indirect support evidence into the grid cells to be written that do not meet the exclusion criteria; and removing the grid cells to be written that meet the exclusion criteria from the candidate regions for the bottom surface of the old furniture. The new furniture candidate location generation module is used to generate the contact area of ​​the bottom surface of the new furniture based on the new furniture model data, and to generate the candidate location of the new furniture in the candidate area of ​​the bottom surface of the old furniture after screening. The spatial positioning and rendering posture determination module is used to project the contact area of ​​the bottom surface of the new furniture onto the support evidence map to obtain a covering grid cell, determine the support score based on the area weight of the covering grid cell, direct support evidence, indirect support evidence and exclusion evidence, and determine the spatial position and rendering posture of the new furniture based on the support score and the exclusion evidence.

Citation Information

Patent Citations

  • AR scene furniture identification and dynamic removal system based on artificial intelligence

    CN121010916A

  • Road scene dynamic updating method and system based on 4DGS, terminal and storage medium

    CN121482738A