Three-dimensional scene reconstruction method and device, equipment and storage medium

By using the depth diffusion technology of color images and depth images in three-dimensional scene reconstruction, the smallest enclosure box of the object is generated, which solves the problems of large computing power consumption and inaccurate reconstruction in the prior art, and achieves efficient and accurate three-dimensional scene reconstruction.

CN120219602APending Publication Date: 2025-06-27BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311808802.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing three-dimensional scene reconstruction method consumes a lot of computing power on the scene image sequence, and cannot guarantee the accuracy of three-dimensional scene reconstruction.

Method used

By obtaining the color image and depth image of the three-dimensional scene, the prior depth of the object pixel points after depth clustering in the object detection box is determined, and the depth diffuses are carried out based on the color information and prior depth to generate the smallest enclosure box of the object to achieve efficient and accurate three-dimensional scene reconstruction.

Benefits of technology

The calculation overhead during reconstruction of various objects in a three-dimensional scene is reduced, the reconstruction reliability of various objects in a three-dimensional scene is improved, and efficient and accurate three-dimensional scene reconstruction is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219602A_ABST
    Figure CN120219602A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a three-dimensional scene reconstruction method and device, equipment and a storage medium. The method comprises the following steps: acquiring a color image and a depth image of a three-dimensional scene at the same moment; according to the depth image, determining the prior depth of object pixel points subjected to depth clustering in a detection frame of each object in the color image; and performing depth diffusion on the pixel points in the detection frame according to the color information of each pixel point in the detection frame and the priori depth of the pixel points of the object to obtain the actual depth of each pixel point in the detection frame so as to generate a minimum bounding box of the object. According to the embodiment of the invention, the minimum bounding box with a simple geometric structure can be used for approximately replacing each complex object, the efficient and accurate reconstruction of each object in the three-dimensional scene is realized, the calculation overhead during the reconstruction of each object in the three-dimensional scene is greatly reduced, and the reconstruction reliability of each object in the three-dimensional scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of data processing, and in particular, to a three-dimensional scene reconstruction method, apparatus, device, and storage medium. Background Art

[0002] With the rapid development of Extended Reality (XR) technology, three-dimensional scene reconstruction technology can provide users with more and more virtual interaction scenes to enhance the immersive interaction experience of users in three-dimensional scenes.

[0003] Generally, multiple scene images at different perspectives can be taken of a three-dimensional scene to form a corresponding scene image sequence. Then, the Structure From Motion (SFM) method is used to analyze the scene image sequence to determine the sparse three-dimensional structure of each object in the three-dimensional scene. Furthermore, through the Multi-View Stereo (MVS) reconstruction method, the sparse three-dimensional structure of each object is processed to obtain the three-dimensional structure mesh after densification of each object, thereby realizing the three-dimensional reconstruction of each object in the three-dimensional scene.

[0004] However, the above three-dimensional reconstruction method consumes a large amount of computing power for the scene image sequence and cannot guarantee the accuracy of three-dimensional scene reconstruction. Summary of the Invention

[0005] Embodiments of the present application provide a three-dimensional scene reconstruction method, apparatus, device, and storage medium, which can realize the efficient and accurate reconstruction of each object in the three-dimensional scene, reduce the computational overhead during the reconstruction of each object in the three-dimensional scene, and improve the reconstruction reliability of each object in the three-dimensional scene.

[0006] In a first aspect, embodiments of the present application provide a three-dimensional scene reconstruction method, which includes:

[0007] Obtain a color image and a depth image of a three-dimensional scene at the same moment;

[0008] According to the depth image, determine the prior depth of the object pixel points after depth clustering within the detection frame of each object in the color image;

[0009] According to the color information of each pixel point within the detection frame and the prior depth of the object pixel points, perform depth diffusion on the pixel points within the detection frame to obtain the actual depth of each pixel point within the detection frame, so as to generate the minimum bounding box of the object.

[0010] In a second aspect, embodiments of the present application provide a three-dimensional scene reconstruction apparatus, which includes:

[0011] An image acquisition module, configured to acquire a color image and a depth image of a three-dimensional scene at the same moment;

[0012] A prior depth determination module, configured to determine, according to the depth image, the prior depth of the object pixel points after depth clustering within the detection frame of each object in the color image;

[0013] A depth diffusion module, configured to perform depth diffusion on the pixel points within the detection frame according to the color information of each pixel point within the detection frame and the prior depth of the object pixel points, so as to obtain the actual depth of each pixel point within the detection frame, and generate a minimum bounding box of the object.

[0014] In a third aspect, an embodiment of the present application provides an electronic device, which includes:

[0015] A processor and a memory, where the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the three-dimensional scene reconstruction method provided in the first aspect of the present application.

[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which is used to store a computer program, and the computer program enables a computer to execute the three-dimensional scene reconstruction method provided in the first aspect of the present application.

[0017] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program / instructions, and the computer program / instructions enable a computer to execute the three-dimensional scene reconstruction method provided in the first aspect of the present application.

[0018] Through the technical solution of the present application, first, a color image and a depth image that are calibrated and aligned at the same moment of a three-dimensional scene are acquired, and according to the pixel point matching between the two, the prior depth of the object pixel points after depth clustering within the detection frame of each object in the color image is determined. Then, according to the color information of each pixel point within the detection frame of each object and the prior depth of the object pixel points, depth diffusion is performed on each pixel point within the detection frame of the object to generate a minimum bounding box of the object, that is, a minimum bounding box with a simple geometric structure can be used to approximately replace each complex object, so as to realize the efficient and accurate reconstruction of each object within the three-dimensional scene, and depth diffusion is performed on each pixel point within the detection frame of each object, without analyzing the depth information of each pixel point within the entire color image, greatly reducing the calculation overhead during the reconstruction of each object within the three-dimensional scene and improving the reconstruction reliability of each object within the three-dimensional scene. Description of the Drawings

[0019] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0020] Figure 1 It is a flowchart of a three-dimensional scene reconstruction method provided by an embodiment of the present application;

[0021] Figure 2 It is a flowchart of another three-dimensional scene reconstruction method provided by an embodiment of the present application;

[0022] Figure 3 It is an exemplary schematic diagram of the pixel point matching process between a depth image and a color image provided by an embodiment of the present application;

[0023] Figure 4 It is a method flowchart of the determination process of object pixel points and the prior depth of object pixel points within the detection frame of each object provided by an embodiment of the present application;

[0024] Figure 5 It is a principle block diagram of a three-dimensional scene reconstruction device provided by an embodiment of the present application;

[0025] Figure 6 It is a schematic block diagram of an electronic device provided by an embodiment of the present application. Specific Embodiments

[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0028] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or advantageous than other embodiments or solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.

[0029] To enhance the diverse interactions between users and various objects in a 3D scene, a corresponding virtual scene can usually be reconstructed for the 3D scene, and a corresponding virtual model can be reconstructed for each object in the 3D scene within the virtual scene, so as to support the user to perform various interaction operations with the virtual models of the objects within the virtual scene, thereby improving the immersive interaction experience between the user and the various objects in the 3D scene.

[0030] Before introducing the specific technical solutions of the present application, the specific application scenarios of the virtual scene reconstructed for the 3D scene in the present application and the virtual models reconstructed for the various objects in the 3D scene can be described accordingly:

[0031] The virtual scene reconstructed for any 3D scene in the present application and the virtual models reconstructed for the various objects within the virtual scene can be supported to be presented on mobile phones, tablet computers, personal computers, servers, and smart wearable devices, so that the user can view the virtual models reconstructed for the various objects within the virtual model reconstructed for the 3D scene. When the virtual scene reconstructed for the 3D scene and the virtual models reconstructed for the various objects within the virtual scene are presented on an XR device, the user can enter the virtual scene reconstructed for the 3D scene by wearing the XR device, so as to perform various interaction operations with the virtual models reconstructed for the various objects within the virtual scene, thereby realizing the diverse interactions between the user and the various objects in the 3D scene.

[0032] XR refers to combining the real and the virtual through a computer to create a virtual environment for human-computer interaction. XR is also a general term for various technologies such as virtual reality (VR for short), augmented reality (AR for short), and mixed reality (MR for short). By integrating the visual interaction technologies of the three, it brings an "immersive feeling" of seamless conversion between the virtual world and the real world to the experiencer. XR devices are usually worn on the user's head, so XR devices are also called head-mounted devices.

[0033] VR: A technology for creating and experiencing virtual worlds. It computationally generates a virtual environment, which is a multi-source information (the virtual reality mentioned in this article includes at least visual perception, and may also include auditory perception, tactile perception, motion perception, and even taste perception, olfactory perception, etc.). It realizes an integrated and interactive three-dimensional dynamic visual scene and simulation of entity behavior in the virtual environment, enabling users to immerse themselves in the simulated virtual reality environment and realizing applications in various virtual environments such as maps, games, videos, education, medical care, simulation, collaborative training, sales, assisting manufacturing, maintenance, and repair.

[0034] A VR device refers to a terminal that realizes virtual reality effects. It usually can be provided in the form of glasses, head-mounted displays (HMDs), contact lenses, etc., for realizing visual perception and other forms of perception. Of course, the form of the virtual reality device is not limited to this, and it can be further miniaturized or enlarged according to needs.

[0035] AR: An AR scene refers to a simulated scene in which at least one virtual object is superimposed on a physical scene or its representation. For example, an electronic system may have an opaque display and at least one imaging sensor for capturing images or videos of the physical scene, which are representations of the physical scene. The system combines the images or videos with virtual objects and displays the combination on the opaque display. An individual uses the system to indirectly view the physical scene via the images or videos of the physical scene and observes the virtual objects superimposed on the physical scene. When the system uses one or more image sensors to capture images of the physical scene and uses those images to present the AR scene on the opaque display, the displayed images are called video pass-through. Alternatively, the electronic system for displaying the AR scene may have a transparent or semi-transparent display, and an individual can directly view the physical scene through the display. The system can display virtual objects on the transparent or semi-transparent display, enabling the individual to use the system to observe the virtual objects superimposed on the physical scene. Another example is that the system may include a projection system for projecting virtual objects onto the physical scene. The virtual objects can be projected, for example, on a physical surface or as a hologram, enabling the individual to use the system to observe the virtual objects superimposed on the physical scene. Specifically, a technology that calculates the camera pose information parameters of a camera in the real world (or three-dimensional world, real world) in real time during the process of the camera capturing images, and adds virtual elements to the images captured by the camera according to the camera pose information parameters. The virtual elements include but are not limited to: images, videos, and 3D models. The goal of AR technology is to interactively overlay the virtual world on the real world on the screen.

[0036] MR: By presenting virtual scene information in a real - world scenario, an information loop for interactive feedback is established among the real world, the virtual world, and the user to enhance the realism of the user experience. For example, integrating sensory inputs created by a computer (such as virtual objects) with sensory inputs from a physical set or their representations in a simulated set. In some MR sets, the sensory inputs created by the computer can adapt to changes in the sensory inputs from the physical set. Additionally, some electronic systems for presenting MR sets can monitor orientation and / or position information relative to the physical set so that virtual objects can interact with real objects (i.e., physical elements from the physical set or their representations). For example, the system can monitor movement so that a virtual plant appears stationary relative to a physical building.

[0037] Optionally, the XR device described in the embodiments of the present application, also known as a virtual reality device, may include but is not limited to the following types:

[0038] 1) Mobile virtual reality device, which supports setting a mobile terminal (such as a smartphone) in various ways (such as a head - mounted display with a dedicated card slot). Through a wired or wireless connection with the mobile terminal, the mobile terminal performs relevant calculations for virtual reality functions and outputs data to the mobile virtual reality device. For example, watching virtual reality videos through an APP on the mobile terminal.

[0039] 2) All - in - one virtual reality device, which has a processor for performing relevant calculations for virtual functions, and thus has independent virtual reality input and output functions. It does not need to be connected to a PC or a mobile terminal and has a high degree of freedom of use.

[0040] 3) PC - based virtual reality (PCVR) device, which uses the PC to perform relevant calculations for virtual reality functions and data output. The externally connected PCVR device uses the data output by the PC to achieve the virtual reality effect.

[0041] After introducing the specific application scenarios that the virtual scene after 3D scene reconstruction in the present application can support for presentation, a 3D scene reconstruction method provided by the embodiments of the present application will be specifically described below with reference to the accompanying drawings.

[0042] Currently, traditional 3D scene reconstruction methods consume excessive computing power and cannot guarantee the accuracy of 3D scene reconstruction.

[0043] To solve the above problems, the inventive concept of this application is as follows: First, obtain the color image and depth image of the three-dimensional scene after calibration and alignment at the same moment, and determine the prior depth of the object pixels after depth clustering within the detection box of each object in the color image based on the pixel point matching between the two. Then, based on the color information of each pixel point within the detection box of each object and the prior depth of the object pixels, perform depth diffusion on each pixel point within the detection box of the object to generate the minimum bounding box of the object. That is, the minimum bounding box with a simple geometric structure can be used to approximately replace each complex object, realizing the efficient and accurate reconstruction of each object within the three-dimensional scene. Through the depth diffusion of each pixel point within the detection box of each object, the computational overhead during the reconstruction of each object within the three-dimensional scene is greatly reduced, and the reconstruction reliability of each object within the three-dimensional scene is improved.

[0044] Figure 1 The flowchart of a three-dimensional scene reconstruction method provided by an embodiment of this application. This method can be executed by the three-dimensional scene reconstruction device provided by this application, where the three-dimensional scene reconstruction device can be implemented in any software and / or hardware manner. Exemplarily, the three-dimensional scene reconstruction device can be configured in any electronic device such as an XR device, a server, a mobile phone, a tablet computer, a personal computer, a smart wearable device, etc. This application does not impose any restrictions on the specific type of the electronic device.

[0045] Specifically, as Figure 1 shown, this method may include the following steps:

[0046] S110, obtain the color image and depth image of the three-dimensional scene at the same moment.

[0047] The three-dimensional scene in this application can be any real scene where the user is located, such as a certain game scene, an indoor scene, etc. Usually, various real objects can be set within the three-dimensional scene, such as tables, chairs, sofas, coffee tables, cabinets, refrigerators, televisions, etc., to support the user's diverse interactions with each object within the three-dimensional scene. Then, when reconstructing the three-dimensional scene, it is necessary to reconstruct each object within the three-dimensional scene. And the reconstruction of each object within the three-dimensional scene requires the help of the point cloud information of each object in the three-dimensional scene.

[0048] Therefore, to ensure the accurate reconstruction of each object within the three-dimensional scene, for the electronic device (such as an XR device) used for three-dimensional scene reconstruction, this application can respectively configure a normal camera and a depth camera on the electronic device, and pre-calibrate and align the normal camera and the depth camera to obtain the coordinate transformation relationship between the normal camera and the depth camera after calibration and alignment, so as to accurately analyze the position matching of the same object in the two images captured by the normal camera and the depth camera subsequently.

[0049] Then, in order to accurately obtain the point cloud data of each object in the three-dimensional scene, the present application can use an ordinary camera and a depth camera to scan the real environment information in the three-dimensional scene at a certain shooting frequency (for example, shooting once every 0.5 ms), so as to obtain a color image of the three-dimensional scene captured by the ordinary camera at a certain moment, and a depth image of the three-dimensional scene captured by the depth camera at a certain moment. Thus, the present application can obtain the color image and the depth image of the three-dimensional scene at the same moment, so as to subsequently analyze the spatial position information of the same object in the three-dimensional scene and obtain the point cloud data corresponding to each object.

[0050] Among them, each pixel point in the color image can carry corresponding color information, while each pixel point in the depth image can carry corresponding depth information.

[0051] It can be understood that after the ordinary camera and the depth camera are calibrated and aligned, the captured color image and depth image are also calibrated and aligned, and the coordinate transformation relationship between the color image and the depth image can be determined, so as to accurately analyze the matching of each pixel point in the color image and the depth image.

[0052] S120. According to the depth image, determine the prior depth of the object pixel points after depth clustering in the detection frame of each object in the color image.

[0053] In order to reduce the computational cost during the reconstruction of each object in the three-dimensional scene, after obtaining the color image of the three-dimensional scene, the present application will first use a pre-configured object detection algorithm to perform corresponding object detection processing on the color image, so as to frame out each object from the color image to obtain the detection frame of each object in the color image.

[0054] Exemplarily, the detection frame of each object can be a rectangular frame that can completely enclose the object, so as to describe the position information of the object in the color image.

[0055] Since each object in the color image usually appears in the detection frame of the object and does not appear in other areas of the color image. Therefore, when reconstructing each object, the present application can perform depth analysis on the pixel points within the detection frame of each object in the color image, without analyzing the depth information of each pixel point in the entire color image, greatly reducing the computational cost during the reconstruction of each object in the three-dimensional scene.

[0056] For the detection box of each object in the color image, the present application can first use the coordinate transformation relationship after calibration and alignment of an ordinary camera and a depth camera to match each pixel point in the depth image with the pixel points within the detection box of the object, so as to determine the partial pixel points within the detection box of the object that match the pixel points in the depth image. Then, according to the depth information of each pixel point in the depth image, the depth information of the partial pixel points within the detection box of the object that match the pixel points in the depth image can be determined.

[0057] It can be understood that the pixel points within the detection box of each object in the color image can be divided into two types: object pixel points inside the object and non-object pixel points outside the object. Among them, the non-object pixel points can be the outlier points representing the background information of the three-dimensional scene within the detection box of the object.

[0058] Since the depth differences among the object pixel points inside the same object in the color image are relatively small, while the depths of the pixel points corresponding to a certain object and other objects or the background in the color image facing the user are relatively large. Therefore, in order to ensure the efficient reconstruction of each object in the three-dimensional scene, for the detection box of each object, the present application can analyze the depth information of the partial pixel points within the detection box that match the pixel points in the depth image to perform clustering processing on the partial pixel points within the detection box, so as to divide the partial pixel points within the detection box of the object into two categories: object pixel points and non-object pixel points. Moreover, for each object pixel point within the detection box of the object, the present application can use the depth information of the pixel point in the depth image that matches the object pixel point as the prior depth of the object pixel point.

[0059] In the same way as above, the prior depths of the object pixel points after depth clustering within the detection box of each object in the color image can be determined.

[0060] S130. According to the color information of each pixel point within the detection box and the prior depth of the object pixel points, perform depth diffusion on the pixel points within the detection box to obtain the actual depth of each pixel point within the detection box, so as to generate the minimum bounding box of the object.

[0061] Since the pixel resolution of the color image is usually higher than that of the depth image, the number of pixel points in the color image is more than that in the depth image. It can be seen that the object pixel points with determined prior depths within the detection box of each object in the color image belong to the partial pixel points within the detection box of the object, and the object pixel points within the detection box of the object are evenly distributed.

[0062] Then, in order to achieve the accurate reconstruction of each object in the three-dimensional scene, based on the prior depth of the object pixel points in the detection frame of each object, this application also needs to comprehensively analyze the depth information of each pixel point in the detection frame of this object.

[0063] Considering that the depth difference between two pixel points belonging to the same object in a color image is not significant, and if the colors of two pixel points in the detection frame of this object are not significantly different, it indicates that these two pixel points are pixel points on the same object, then the depth difference between these two pixel points is also not significant. Moreover, any two adjacent pixel points in the detection frame of a certain object in a color image usually may approximately represent the same object, making the depth difference between any two adjacent pixel points as small as possible.

[0064] Therefore, for the detection frame of each object in a color image, this application can first determine the color information of each pixel point in this detection frame and analyze the color change situation of each pixel point in this detection frame. Furthermore, with the prior depth of each object pixel point already determined in the detection frame of this object as a reference, starting from each object pixel point in the detection frame of this object, according to the color change situation of each pixel point in the detection frame of this object, depth diffusion is performed on each pixel point in the detection frame of this object to determine the actual depth of each pixel point in the detection frame of this object.

[0065] Then, for the detection frame of each object, this application can determine the spatial position information corresponding to each pixel point in the detection frame of this object in the three-dimensional scene by analyzing the pixel coordinates and actual depth of each pixel point in this detection frame. Then, by analyzing the spatial position information corresponding to each pixel point in the detection frame of this object in the three-dimensional scene, the spatial boundary information of this object in the three-dimensional scene can be judged, and the minimum bounding box of this object is generated accordingly.

[0066] In the same way as above, the minimum bounding box of each object in the three-dimensional scene can be generated, and when reconstructing the three-dimensional scene, the minimum bounding box with a simple geometric structure can be used to approximately replace each complex object to achieve the efficient and accurate reconstruction of each object in the three-dimensional scene.

[0067] The technical solution provided by the embodiments of the present application first obtains a color image and a depth image after calibration and alignment at the same moment of a three-dimensional scene, and determines the prior depth of the object pixels after depth clustering in the detection frame of each object in the color image according to the pixel point matching between the two. Then, according to the color information of each pixel point in the detection frame of each object and the prior depth of the object pixels, depth diffusion is performed on each pixel point in the detection frame of the object to generate the minimum bounding box of the object, that is, the minimum bounding box with a simple geometric structure can be used to approximately replace each complex object, realizing the efficient and accurate reconstruction of each object in the three-dimensional scene, and performing depth diffusion on each pixel point in the detection frame of each object, without analyzing the depth information of each pixel point in the entire color image, greatly reducing the computational cost when reconstructing each object in the three-dimensional scene and improving the reconstruction reliability of each object in the three-dimensional scene.

[0068] As an alternative implementation solution in the present application, in order to ensure the efficient and accurate reconstruction of each object in the three-dimensional scene, the present application can give a detailed explanation of the clustering process and depth diffusion process of each pixel point in the detection frame of each object in the color image.

[0069] Figure 2 It is a flowchart of another three-dimensional scene reconstruction method provided by the embodiments of the present application. As Figure 2 shown, the method may specifically include the following steps:

[0070] S210, obtain a color image and a depth image of the three-dimensional scene at the same moment.

[0071] S220, determine the matching pixel point of each pixel point in the depth image in the color image and the prior depth of the matching pixel point.

[0072] After obtaining the color image and the depth image of the three-dimensional scene at the same moment, in order to accurately obtain the point cloud data of any object, the present application can first obtain the coordinate transformation relationship between a common camera and a depth camera after calibration and alignment. Since the pixel resolution of the color image is usually higher than that of the depth image, the number of pixel points in the color image is more than that in the depth image. Then, each pixel point in the depth image will have a matching pixel point in the color image.

[0073] Therefore, the present application can use the coordinate transformation relationship after calibration alignment of an ordinary camera and a depth camera to perform coordinate transformation on the pixel coordinates of each pixel point in the depth image, so as to obtain the transformed pixel coordinates of each pixel point. Then, in the color image, a certain pixel point with the same pixel coordinates as the transformed pixel coordinates of each pixel point in the depth image can be found as the matching pixel point of this pixel point in the depth image. Thus, in the color image, the matching pixel point of each pixel point in the depth image can be determined. Then, the depth information of each pixel point in the depth image is used as the prior depth of the matching pixel point of this pixel point in the color image.

[0074] Exemplarily, the pixel resolution of the depth image is less than that of the color image. Then, as Figure 3 shown, for each pixel point in the depth image, the matching pixel point of this pixel point can be determined in the color image, and the depth information of each pixel point in the depth image is used as the prior depth of the matching pixel point of this pixel point. Among them, Figure 3 the pixel points represented by in the color image are the matching pixel points of each pixel point in the depth image.

[0075] S230. For the detection frame of each object in the color image, perform depth clustering on the matching pixel points within the detection frame to obtain the object pixel points and non-object pixel points within the detection frame.

[0076] To reduce the computational cost during the reconstruction of each object in the three-dimensional scene, for the color image, the present application can adopt a corresponding object detection algorithm to perform object detection processing on each object in the color image to obtain the detection frame of each object in the color image.

[0077] Since the pixel points in the color image can be divided into two types: the matching pixel points of each pixel point in the depth image and the unmatched ordinary pixel points. Then, as Figure 3 shown, according to the boundary position where the detection frame of each object in the color image is located, the respective pixel points within the detection frame of each object can be determined, and the respective pixel points within the detection frame also include two types: the matching pixel points of the corresponding pixel points in the depth image and the unmatched ordinary pixel points.

[0078] Then, considering that in addition to the object itself being framed within the detection frame of each object, a part of the background outside the object is also framed, so that the pixel points within the detection frame of each object can be divided into two types: the object pixel points inside the object and the non-object pixel points outside the object, indicating that the respective matching pixel points within the detection frame of each object can also be divided into two types: the object pixel points inside the object and the non-object pixel points outside the object.

[0079] Then, for the detection box of each object in the color image, the present application can perform depth clustering processing on each matching pixel point within the detection box of the object according to the depth information of each matching pixel point within the detection box of the object, so as to divide each matching pixel point within the detection box of the object into two types: object pixel points and non-object pixel points. Thus, in order to ensure the efficiency of object reconstruction, the present application can directly determine the object pixel points and non-object pixel points after the division of each matching pixel point within the detection box of each object.

[0080] In some implementable manners, for the object pixel points within the detection box of each object, as Figure 4 shown, the present application can be determined through the following steps:

[0081] S410, perform object detection on the color image to obtain the detection box of each object in the color image.

[0082] For the object detection of the color image, the present application can pre-train an object detection model, and the object detection model is trained using a corresponding object detection algorithm to accurately detect each object in the color image.

[0083] For each obtained color image, the present application will input the color image into the pre-trained object detection model to detect and identify each object in the color image through the object detection model, so as to output the detection box of each object in the color image.

[0084] It can be understood that when performing object detection on each object in the color image, the specific information of each object will also be recognized to output the semantic information of the detection box of each object to represent the specific identification meaning of the framed object.

[0085] S420, according to the prior depth difference between every two adjacent matching pixel points within the detection box, perform depth clustering on the matching pixel points within the detection box to obtain the object pixel points within the detection box.

[0086] Considering that the depths of the pixel points within the same object in the color image are relatively close, while the depths of the pixel points corresponding to a certain object and other objects or the background in the color image facing the user are relatively large. Moreover, for each matching pixel point within the detection box of each object in the color image, the matching pixel point includes two types: object pixel points within the object and non-object pixel points outside the object.

[0087] Therefore, in order to ensure the efficient reconstruction of each object in the three-dimensional scene, for the detection box of each object in the color image, the present application can first find a matching pixel point that must be inside the object from each matching pixel point within the detection box of the object. Then, starting from this matching pixel point, by analyzing whether the prior depth difference between this matching pixel point and its adjacent matching pixel points is greater than a preset depth difference threshold. If the prior depth difference between this matching pixel point and its adjacent matching pixel points is less than or equal to the preset depth difference threshold, it indicates that this matching pixel point and its adjacent matching pixel points are pixel points within the same object, and then this matching pixel point and its adjacent matching pixel points are aggregated into one category as the object pixel points inside the object. If the prior depth difference between this matching pixel point and its adjacent matching pixel points is greater than the preset depth difference threshold, it indicates that the adjacent matching pixel points of this matching pixel point are different from this matching pixel point and do not belong to the pixel points inside the object, and then the adjacent matching pixel points of this matching pixel point can be aggregated into one category as the non-object pixel points outside the object.

[0088] Then, taking each of the matching pixel points adjacent to this matching pixel point as a new starting pixel point, and again analyzing whether the prior depth difference between this new starting pixel point and its adjacent matching pixel points is greater than the preset depth difference threshold, and looping in turn to continuously analyze the prior depth difference between each two adjacent matching pixel points, so as to divide each adjacent matching pixel point into one of the object pixel points and the non-object pixel points until all the matching pixel points within the detection box of the object are traversed, and then the object pixel points within the detection box of the object can be obtained.

[0089] Exemplarily, the present application can adopt a breadth-first search algorithm to perform depth clustering on each matching pixel point within the detection box of each object. The specific process is as follows: For the detection box of each object, the central pixel point in this detection box can be determined as the pixel point inside the object. Then, the present application can start from the central matching pixel point within the detection box of the object and traverse each neighborhood of this central matching pixel point to determine each adjacent matching pixel point of this central matching pixel point. If the prior depth difference between this central matching pixel point and an adjacent matching pixel point is less than or equal to the preset depth difference threshold (for example, 0.1 m), it is determined that these two adjacent matching pixel points belong to the same category, that is, they are pixel points inside the object, and they are aggregated into the corresponding object pixel points. If the prior depth difference between this central matching pixel point and an adjacent matching pixel point is greater than the preset depth difference threshold (for example, 0.1 m), it is determined that these two adjacent matching pixel points belong to different categories, and this adjacent matching pixel point is classified as a non-object pixel point outside the object.

[0090] Then, take each pair of adjacent matching pixel points as new starting pixel points. Starting from these new starting pixel points, traverse the neighborhoods of these new starting pixel points again to determine their respective adjacent matching pixel points, and perform the same pixel clustering steps as above. Repeat this process in sequence until all the matching pixel points within the detection bounding box of the object are traversed. Finally, according to the clustering results of each matching pixel point within the detection bounding box of each object, the object pixel points within the detection bounding box of the object can be obtained.

[0091] It should be noted that for the object pixel points and non-object pixel points within the detection bounding box of each object, in addition to using the depth clustering algorithm to divide them, this application can also use other clustering algorithms to divide them according to other pixel differences between the object pixel points and non-object pixel points. The specific clustering method adopted by this application for dividing the object pixel points and non-object pixel points within the detection bounding box of each object is not limited.

[0092] S240, determine the prior depth of the object pixel points and set the prior depth of the non-object pixel points to a fixed value.

[0093] For the detection bounding box of each object, the prior depth of each object pixel point within the detection bounding box of the object can be directly obtained. Considering that when reconstructing the object, it is not necessary to refer to the depth of the non-object pixel points to reduce the computational overhead during object reconstruction. Therefore, this application can set the prior depth of each non-object pixel point within the detection bounding box of each object to a certain fixed value. Among them, this fixed value can be zero.

[0094] S250, determine the corresponding depth diffusion target according to the pixel color difference and the actual depth difference between every two adjacent pixel points within the detection bounding box, and the difference between the actual depth and the prior depth of the object pixel points within the detection bounding box.

[0095] Considering that the depths of two pixel points belonging to the same object in a color image do not differ much, and if the colors of two pixel points within the detection bounding box of the object differ little, it indicates that these two pixel points are pixel points on the same object, then the depths of these two pixel points do not differ much either. Moreover, any two adjacent pixel points within the detection bounding box of a certain object in a color image usually may approximately represent the same object, making the depth difference between any two adjacent pixel points as small as possible.

[0096] As can be seen from the above, when performing depth diffusion on each pixel point within the detection bounding box of each object, every two adjacent pixel points within the detection bounding box of the object should satisfy that the actual depth difference between these two adjacent pixel points should be as small as possible and be consistent with the change situation of the pixel color difference between these two adjacent pixel points.

[0097] Moreover, since there is already a corresponding prior depth for some of the object pixel points within the detection frame of each object. Then, to ensure the accuracy of object reconstruction, for each object pixel point within the detection frame of each object, the difference between the actual depth and the prior depth of this object pixel point should be minimized as much as possible.

[0098] Therefore, based on the above two conditions that each pixel point within the detection frame of each object should satisfy, the present application can limit the actual depth difference between every two adjacent pixel points according to the pixel color difference between these two adjacent pixel points within the detection frame of the object, and limit the actual depth of each object pixel point according to the prior depth of each object pixel point within the detection frame of the object, so as to construct a depth diffusion target corresponding to each pixel point within the detection frame of each object.

[0099] In some implementable ways, for the depth diffusion target corresponding to each pixel point within the detection frame of each object, the present application can determine it in the following way: determine the corresponding depth diffusion smooth term according to the pixel color difference and the actual depth difference between each pixel point within the detection frame and its adjacent pixel points; determine the corresponding depth diffusion regularization term according to the difference between the actual depth and the prior depth of the object pixel points within the detection frame; determine the corresponding depth diffusion target according to the depth diffusion smooth term and the depth diffusion regularization term.

[0100] That is to say, for each pixel point within the detection frame of each object, since this pixel point and its adjacent pixel points can usually be approximately represented as the same object, the actual depths of these two adjacent pixel points differ as little as possible. Moreover, the pixel color difference between these two adjacent pixel points will also have a corresponding positive impact on the actual depth difference between these two adjacent pixel points. From this, it can be known that if the pixel color difference between this pixel point and its adjacent pixel points is smaller, it means that the actual depth difference between these two adjacent pixel points is also smaller, and the proportion of the actual depth difference between these two adjacent pixel points in depth diffusion is higher.

[0101] Thus, for each pixel point within the detection frame of each object, the present application can set the proportion of the actual depth difference between this pixel point and its adjacent pixel points in depth diffusion according to the pixel color difference between this pixel point and its adjacent pixel points, so as to obtain the depth diffusion smooth term corresponding to this pixel point.

[0102] Exemplarily, the depth diffusion smooth term can be:

[0103] E1 = {Q ij1 (x i,j -x i+1,j ) 2 +Q ij2 (xi,j -x i-1,j ) 2 +Q ij3 (x i,j -x i,j+1 ) 2 +Q ij4 (x i,j -x i,j-1 ) 2}

[0104] Among them, the actual depth of each pixel point within the detection frame of each object can be x i,j , then x i+1,j , x i-1,j , x i,j+1 and x i,j-1 are respectively the actual depths of the respective adjacent pixel points of this pixel point, and Q ij1 , Q ij2 , Q ij3 and Q ij4 are respectively the reciprocals of the pixel color differences between this pixel point and the respective adjacent pixel points, indicating that when the pixel color difference between two adjacent pixel points is smaller, the proportion of the actual depth difference between these two adjacent pixel points in depth diffusion is larger, and the actual depth difference between these two adjacent pixel points is also made as small as possible.

[0105] Moreover, since prior depths already exist for some object pixel points within the detection frame of each object, and the object pixel points should satisfy that the difference between their actual depths and the prior depths should be as small as possible. Therefore, for each object pixel point within the detection frame of each object, the present application can set the difference between the actual depth and the prior depth of this object pixel point to obtain the depth diffusion regularization term corresponding to this pixel point.

[0106] Exemplarily, the depth diffusion regularization term can be: E2 = W ij (x i,j -y i,j ) 2 .

[0107] Among them, the actual depth of each pixel point within the detection frame of each object can be x i,j , then y i,j can be the prior depth of this pixel point.

[0108] Moreover, indicates that when the prior depth y i,j of this pixel point is valid (there is a depth value), the corresponding depth diffusion regularization term is used. And when the prior depth y i,j of this pixel point is invalid (there is no depth value), the depth diffusion regularization term can be directly set to 0 and does not need to participate in depth diffusion.

[0109] Then, after determining the corresponding depth diffusion smoothing term and depth diffusion regularization term, the present application can directly set the sum value of the depth diffusion smoothing term and the depth diffusion regularization term to be the minimum to determine the corresponding depth diffusion target.

[0110] Exemplarily, the depth diffusion target can be:

[0111] where N and M respectively represent the pixel length and width information of the detection frame of each object.

[0112] S260. According to the depth diffusion target, perform depth diffusion on the pixel points within the detection frame to obtain the actual depth of each pixel point within the detection frame.

[0113] The depth diffusion target of the present application is to solve the optimization problem of Laplace, and the optimization variable is the actual depth x of each pixel point within the detection frame of each object. i,j 。

[0114] Therefore, by using the solution method of the optimization problem of Laplace to perform corresponding solution on the above depth diffusion target, the actual depth of each pixel point within the detection frame of the object can be calculated.

[0115] S270. According to the pixel coordinates and actual depth of each pixel point within the detection frame, determine the point cloud data of the object.

[0116] For each pixel point within the detection frame of each object in the color image, the present application can uniformly process the pixel coordinates and actual depth of the pixel point to transform the pixel point to a certain spatial point in the three-dimensional scene. Accordingly, according to the various spatial points obtained by transforming each pixel point within the detection frame of each object to the three-dimensional scene, the point cloud data of the object can be determined.

[0117] S280. Perform principal component analysis on the point cloud data of the object to determine the eigenvalues of the object in at least three principal axis directions to generate the minimum bounding box of the object.

[0118] For each object in the three-dimensional scene, different-shaped spatial bounding boxes can be used to approximately replace the object. For example, the spatial bounding box can include a cube, a polyhedron composed of more than four polygons, etc. And different spatial bounding boxes will have different principal axis directions for the suitable spatial coordinate systems. For example, the spatial coordinate system corresponding to a cube can include three principal axis directions: the x-axis, the y-axis, and the z-axis. Moreover, in the principal component analysis algorithm, each principal component can be used to represent each principal axis direction corresponding to the object when it is enclosed.

[0119] Therefore, the present application can perform corresponding principal component analysis on the point cloud data of each object according to the characteristics of the object bounding box used, transform the point cloud data of the object into at least three feature vectors, and determine the eigenvalues of the object on each feature vector. Among them, the at least three feature vectors can represent at least three principal axis directions corresponding to the bounding box, and the eigenvalues of the object on each feature vector can represent the lengths of the bounding boxes corresponding to the object on each principal axis direction.

[0120] Then, according to the lengths of the bounding boxes corresponding to each object in each principal axis direction, the minimum bounding box of the object can be generated, and the geometrically simple minimum bounding box can be used to approximately replace each complex object to achieve efficient and accurate reconstruction of each object in the three-dimensional scene.

[0121] S290, generate a semantic map of the three-dimensional scene according to the minimum bounding box and the detection frame semantic information of each object in the three-dimensional scene.

[0122] In the three-dimensional scene, the minimum bounding box of each object can be generated, and the geometrically simple minimum bounding box can be used to approximately replace each complex object to achieve efficient reconstruction of each object in the three-dimensional scene. Then, according to the detection frame semantic information of each object, the specific object information represented by each minimum bounding box in the three-dimensional scene is determined, so that the corresponding object information is marked on each minimum bounding box generated in the three-dimensional scene to generate a semantic map of the three-dimensional scene.

[0123] The technical solution provided by the embodiment of the present application first obtains the color image and the depth image after calibration and alignment at the same moment in the three-dimensional scene, and determines the prior depth of the object pixel points after depth clustering in the detection frame of each object in the color image according to the pixel point matching between the two. Then, according to the color information of each pixel point in the detection frame of each object and the prior depth of the object pixel points, depth diffusion is performed on each pixel point in the detection frame of the object to generate the minimum bounding box of the object, and the geometrically simple minimum bounding box can be used to approximately replace each complex object to achieve efficient and accurate reconstruction of each object in the three-dimensional scene. Moreover, depth diffusion is performed on each pixel point in the detection frame of each object, without analyzing the depth information of each pixel point in the entire color image, greatly reducing the computational cost during the reconstruction of each object in the three-dimensional scene and improving the reconstruction reliability of each object in the three-dimensional scene.

[0124] Figure 5 It is a principle block diagram of a three-dimensional scene reconstruction device provided by an embodiment of the present application. As Figure 5 shown, the three-dimensional scene reconstruction device 500 may include:

[0125] An image acquisition module 510, configured to acquire a color image and a depth image of a three-dimensional scene at the same moment;

[0126] A prior depth determination module 520, configured to determine the prior depth of the object pixel points after depth clustering within the detection frame of each object in the color image according to the depth image;

[0127] A depth diffusion module 530, configured to perform depth diffusion on the pixel points within the detection frame according to the color information of each pixel point within the detection frame and the prior depth of the object pixel points, so as to obtain the actual depth of each pixel point within the detection frame, and generate a minimum bounding box of the object.

[0128] In some implementable manners, the prior depth determination module 520 may include:

[0129] A pixel point matching unit, configured to determine the matching pixel point of each pixel point in the depth image in the color image and the prior depth of the matching pixel point;

[0130] A pixel point clustering unit, configured to perform depth clustering on the matching pixel points within the detection frame for each object in the color image, so as to obtain the object pixel points and non-object pixel points within the detection frame;

[0131] A prior depth determination unit, configured to determine the prior depth of the object pixel points and set the prior depth of the non-object pixel points to a fixed value.

[0132] In some implementable manners, the pixel point clustering determination unit may specifically be configured to:

[0133] Perform object detection on the color image to obtain the detection frame of each object in the color image;

[0134] Perform depth clustering on the matching pixel points within the detection frame according to the difference in prior depth between every two adjacent matching pixel points within the detection frame, so as to obtain the object pixel points within the detection frame.

[0135] In some implementable manners, the depth diffusion module 530 may include:

[0136] A depth diffusion target determination unit, configured to determine a corresponding depth diffusion target according to the difference in pixel color and the difference in actual depth between every two adjacent pixel points within the detection frame, and the difference between the actual depth and the prior depth of the object pixel points within the detection frame;

[0137] A depth diffusion unit, configured to perform depth diffusion on the pixel points within the detection frame according to the depth diffusion target, so as to obtain the actual depth of each pixel point within the detection frame.

[0138] In some implementable manners, the depth diffusion target determination unit may specifically be used for:

[0139] Determine a corresponding depth diffusion smoothing term according to the pixel color difference and the actual depth difference between each pixel point in the detection box and its adjacent pixel points;

[0140] Determine a corresponding depth diffusion regularization term according to the difference between the actual depth and the prior depth of the object pixel points in the detection box;

[0141] Determine a corresponding depth diffusion target according to the depth diffusion smoothing term and the depth diffusion regularization term.

[0142] In some implementable manners, the depth diffusion module 530 may further include a bounding box generation unit. The bounding box generation unit may be used for:

[0143] Determine the point cloud data of the object according to the pixel coordinates and the actual depth of each pixel point in the detection box;

[0144] Perform principal component analysis on the point cloud data of the object to determine the eigenvalues of the object in at least three principal axis directions, so as to generate the minimum bounding box of the object.

[0145] In some implementable manners, the three-dimensional scene reconstruction device 500 may further include:

[0146] A semantic map generation module, configured to generate a semantic map of the three-dimensional scene according to the minimum bounding box of each object in the three-dimensional scene and the detection box semantic information.

[0147] In the embodiments of the present application, first, a color image and a depth image after calibration alignment at the same moment of the three-dimensional scene are obtained, and according to the pixel point matching between the two, the prior depth of the object pixel points after depth clustering in the detection box of each object in the color image is determined. Then, according to the color information of each pixel point in the detection box of each object and the prior depth of the object pixel points, depth diffusion is performed on each pixel point in the detection box of the object to generate the minimum bounding box of the object, that is, the minimum bounding box with a simple geometric structure can be used to approximately replace each complex object, so as to realize the efficient and accurate reconstruction of each object in the three-dimensional scene, and depth diffusion is performed on each pixel point in the detection box of each object, without analyzing the depth information of each pixel point in the entire color image, greatly reducing the calculation overhead during the reconstruction of each object in the three-dimensional scene and improving the reconstruction reliability of each object in the three-dimensional scene.

[0148] It should be understood that the device embodiments and the method embodiments in the present application can correspond to each other, and similar descriptions can refer to the method embodiments in the present application. To avoid repetition, they will not be elaborated here.

[0149] Specifically, Figure 5 the illustrated device 500 can execute any method embodiment provided by the present application, and Figure 5 the foregoing and other operations and / or functions of the respective modules in the illustrated device 500 respectively implement the corresponding processes of the above method embodiments. For the sake of brevity, they will not be elaborated here.

[0150] The above method embodiments of the present application have been described from the perspective of functional modules in combination with the accompanying drawings. It should be understood that the functional modules can be implemented in the form of hardware, can also be implemented by instructions in the form of software, or can be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in the hardware of the processor and / or instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or can be executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0151] Figure 6 It is a schematic block diagram of an electronic device provided by an embodiment of the present application.

[0152] As Figure 6 shown, the electronic device 600 may include:

[0153] a memory 610 and a processor 620. The memory 610 is used to store a computer program and transmit the program code to the processor 620. In other words, the processor 620 can call and run the computer program from the memory 610 to implement the method in the embodiments of the present application.

[0154] For example, the processor 620 can be used to execute the above method embodiments according to the instructions in the computer program.

[0155] In some embodiments of the present application, the processor 620 may include but is not limited to:

[0156] General-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the like.

[0157] In some embodiments of the present application, the memory 610 includes, but is not limited to:

[0158] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0159] In some embodiments of the present application, the computer program may be divided into one or more modules, and the one or more modules are stored in the memory 610 and executed by the processor 620 to complete the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device 600.

[0160] As Figure 6 shown, the electronic device may further include:

[0161] A transceiver 630, which can be connected to the processor 620 or the memory 610.

[0162] Among them, the processor 620 can control the transceiver 630 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 630 can include a transmitter and a receiver. The transceiver 630 can further include an antenna, and the number of antennas can be one or more.

[0163] It should be understood that the various components in the electronic device 600 are connected through a bus system. Among them, the bus system includes not only a data bus, but also a power bus, a control bus, and a status signal bus.

[0164] This application also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer can execute the methods in the above method embodiments.

[0165] The embodiments of this application also provide a computer program product containing computer programs / instructions. When the computer programs / instructions are executed by a computer, the computer executes the methods in the above method embodiments.

[0166] When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)).

[0167] The above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A three-dimensional scene reconstruction method, characterized in that, Including: Obtaining a color image and a depth image of a three-dimensional scene at the same moment; Determining the prior depth of the object pixel points after depth clustering within the detection frame of each object in the color image according to the depth image; Performing depth diffusion on the pixel points within the detection frame according to the color information of each pixel point within the detection frame and the prior depth of the object pixel points, to obtain the actual depth of each pixel point within the detection frame, so as to generate the minimum bounding box of the object.

2. The method according to claim 1, wherein The determining the prior depth of the object pixel points after depth clustering within the detection frame of each object in the color image according to the depth image includes: Determining the matching pixel point of each pixel point in the depth image within the color image and the prior depth of the matching pixel point; For the detection frame of each object in the color image, performing depth clustering on the matching pixel points within the detection frame to obtain the object pixel points and non-object pixel points within the detection frame; Determining the prior depth of the object pixel points and setting the prior depth of the non-object pixel points to a fixed value.

3. The method according to claim 2, wherein The performing depth clustering on the matching pixel points within the detection frame for the detection frame of each object in the color image to obtain the object pixel points within the detection frame includes: Performing object detection on the color image to obtain the detection frame of each object in the color image; Performing depth clustering on the matching pixel points within the detection frame according to the difference in prior depth between every two adjacent matching pixel points within the detection frame, to obtain the object pixel points within the detection frame.

4. The method according to claim 1, wherein The performing depth diffusion on the pixel points within the detection frame according to the color information of each pixel point within the detection frame and the prior depth of the object pixel points to obtain the actual depth of each pixel point within the detection frame includes: Determining the corresponding depth diffusion target according to the difference in pixel color and the difference in actual depth between every two adjacent pixel points within the detection frame, and the difference between the actual depth and the prior depth of the object pixel points within the detection frame; Performing depth diffusion on the pixel points within the detection frame according to the depth diffusion target to obtain the actual depth of each pixel point within the detection frame.

5. The method according to claim 4, wherein The determining the corresponding depth diffusion target according to the difference in pixel color and the difference in actual depth between every two adjacent pixel points within the detection frame, and the difference between the actual depth and the prior depth of the object pixel points within the detection frame includes: Determining the corresponding depth diffusion smooth term according to the difference in pixel color and the difference in actual depth between each pixel point within the detection frame and its adjacent pixel points; Determining the corresponding depth diffusion regularization term according to the difference between the actual depth and the prior depth of the object pixel points within the detection frame; Determining the corresponding depth diffusion target according to the depth diffusion smooth term and the depth diffusion regularization term.

6. The method according to claim 1, characterized in that, The generating the minimum bounding box of the object includes: Determining the point cloud data of the object according to the pixel coordinates and the actual depth of each pixel point within the detection frame; Perform principal component analysis on the point cloud data of the object to determine the eigenvalues of the object in at least three principal axis directions, so as to generate the minimum bounding box of the object.

7. The method according to claim 1, characterized in that The method further includes: Generating a semantic map of the three-dimensional scene according to the minimum bounding box and detection box semantic information of each object in the three-dimensional scene.

8. A three-dimensional scene reconstruction device, characterized in that, Including: An image acquisition module, configured to acquire a color image and a depth image of the three-dimensional scene at the same moment; A prior depth determination module, configured to determine the prior depth of the object pixel points after depth clustering in the detection box of each object in the color image according to the depth image; A depth diffusion module, configured to perform depth diffusion on the pixel points in the detection box according to the color information of each pixel point in the detection box and the prior depth of the object pixel points, to obtain the actual depth of each pixel point in the detection box, so as to generate the minimum bounding box of the object.

9. An electronic device, characterized in that, Including: A processor; And A memory, configured to store the executable instructions of the processor; Wherein, the processor is configured to execute the three-dimensional scene reconstruction method according to any one of claims 1-7 by executing the executable instructions.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the three-dimensional scene reconstruction method according to any one of claims 1-7.

11. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions included in the computer program product run on an electronic device, the electronic device is caused to execute the three-dimensional scene reconstruction method according to any one of claims 1-7.